grammar-inference-engine/AGENTS.md
tobjend 830104b399
Some checks failed
ci/woodpecker/push/woodpecker Pipeline failed
ci/woodpecker/pr/woodpecker Pipeline failed
feat: add analyze_directory MCP tool — scan codebase, infer conventions, persist to .dervish/
2026-07-11 21:28:35 +02:00

2.4 KiB

Grammar Inference Engine — Agent Guide

Overview

This repo implements the BEX family of algorithms for inferring regular expression grammars from example sequences. Use it whenever you need to discover the pattern behind a set of strings or structured sequences.

Quick Start for Agents

# Fast pattern inference
from bex.crx import CRX
g = CRX().infer([['a','b','c'], ['a','b'], ['a','c']])  # a.(b+c)?

# Probabilistic k-ORE inference (handles noise better)
from bex.idregex import idregex
g = idregex([['a','b','c'], ['a','b'], ['a','c']], kmax=2, N=3)

Use Cases

  1. Ansible role patterns — extract module sequences from tasks/main.yml, learn per-category grammars
  2. Log analysis — find common patterns in event sequences
  3. API call patterns — learn the typical order of API operations
  4. Configuration structure — discover the schema behind YAML files
  5. Workflow mining — extract the typical task flow from process logs

Architecture

Three inference pipelines:

Pipeline When to use
CRX (fast) Many examples, need speed, CHAREs output
iDRegEx (robust) Few/noisy examples, need probabilistic handling
Tag Preprocessor (bex.tag_preprocessor) Source code analysis — tree-sitter AST → method-level call sequences → per-package grammars

Running Tests

python -m pytest tests/

MCP Server

The primary interface is an MCP server exposing two tools:

Tool Parameters What it does
infer_best_grammar sequences, prefer, kmax, N, min_coverage Infer grammar from raw sequences. Runs CRX + iDRegEx, picks best by MDL.
analyze_directory directory, slice, min_coverage, prefer, kmax, include, exclude, main_only, max_mdl, persist Scan source code, infer conventions per package. Returns YAML grouped by module. Auto-persists to {directory}/.dervish/grammars.yml.

Start it: python /path/to/bex/mcp_server.py, then connect any MCP client.

Tag Preprocessor CLI

For analyzing source code directories:

python -m bex.tag_preprocessor.analyze /path/to/codebase --verbose
python -m bex.tag_preprocessor.analyze /path/to/codebase --slice package --include '**/src/**'

Key flags: --slice package (per-directory grammars), --verbose (progress), --include/--exclude (glob filters), --main-only (exclude test files), --kore (enable slow kORE in ensemble).