# Grammar Inference Engine — Agent Guide ## Overview This repo implements the BEX family of algorithms for inferring regular expression grammars from example sequences. Use it whenever you need to discover the pattern behind a set of strings or structured sequences. ## Quick Start for Agents ```python # Fast pattern inference from bex.crx import CRX g = CRX().infer([['a','b','c'], ['a','b'], ['a','c']]) # a.(b+c)? # Probabilistic k-ORE inference (handles noise better) from bex.idregex import idregex g = idregex([['a','b','c'], ['a','b'], ['a','c']], kmax=2, N=3) ``` ## Use Cases 1. **Ansible role patterns** — extract module sequences from tasks/main.yml, learn per-category grammars 2. **Log analysis** — find common patterns in event sequences 3. **API call patterns** — learn the typical order of API operations 4. **Configuration structure** — discover the schema behind YAML files 5. **Workflow mining** — extract the typical task flow from process logs ## Architecture Three inference pipelines: | Pipeline | When to use | |----------|-------------| | CRX (fast) | Many examples, need speed, CHAREs output | | iDRegEx (robust) | Few/noisy examples, need probabilistic handling | | Tag Preprocessor (`bex.tag_preprocessor`) | Source code analysis — tree-sitter AST → method-level call sequences → per-package grammars | ## Running Tests ```bash python -m pytest tests/ ``` ## MCP Server The primary interface is an MCP server exposing two tools: | Tool | Parameters | What it does | |------|-----------|-------------| | `infer_best_grammar` | `sequences`, `prefer`, `kmax`, `N`, `min_coverage` | Infer grammar from raw sequences. Runs CRX + iDRegEx, picks best by MDL. | | `analyze_directory` | `directory`, `slice`, `min_coverage`, `prefer`, `kmax`, `include`, `exclude`, `main_only`, `max_mdl`, `persist` | Scan source code, infer conventions per package. Returns YAML grouped by module. Auto-persists to `{directory}/.dervish/grammars.yml`. | Start it: `python /path/to/bex/mcp_server.py`, then connect any MCP client. ## Tag Preprocessor CLI For analyzing source code directories: ```bash python -m bex.tag_preprocessor.analyze /path/to/codebase --verbose python -m bex.tag_preprocessor.analyze /path/to/codebase --slice package --include '**/src/**' ``` Key flags: `--slice package` (per-directory grammars), `--verbose` (progress), `--include`/`--exclude` (glob filters), `--main-only` (exclude test files), `--kore` (enable slow kORE in ensemble).