6.9 KiB
Handover — Behavioral Grammar Inference Project
Current Status
Branch: feature/treesitter-tag-queries
Last commit: 8b3a454 — feat: enable idregex_refine by default + AST rewrite of _count_optionals
Tests: 313 passed, 8 warnings, 0 failures
What We Built
Source Code Analysis Pipeline
A language-agnostic pipeline that infers per-package calling conventions from any codebase:
- Tree-sitter AST → extract method-level behavioral sequences (call chains, control flow)
- Algorithm 7 (CRX) — generalized regular expression inference from examples
- AST grammar nodes — Symbol, Concat, Alt, Optional, Plus, Star, Empty
- Language Size scoring — ranks grammars by compressed description length
- YAML/GBNF output — structured grammars grouped by package, ready for constrained decoding
Key Changes Since Last Handover
- AST migration complete — all SORE string operations purged from pipeline
- Scoring: Language Size (
lang_size_score) chosen over MDL (ADR-13) - Decomposition ON by default —
decompose=True,max_seq_length=4 - idregex_refine ON by default — iDRefEx runs on small groups where it helps
_COUNT_CAPraised to 10^30 — no longer clamps scorer values- Memoization fixed —
_count_concathas@lru_cache, RAGSAK 54.9s→6.7s
Default Parameters (Golden Config)
{
'decompose': True,
'max_seq_length': 4,
'min_structure': 0.5,
'idregex_refine': True,
'min_methods': 3,
'min_coverage': 0.05,
}
What Works
High-Confidence Findings
- Decomposition increases grammar count 3× and reduces pure bags
- Language Size scoring discriminates between tight and bag grammars (20 vs 9975)
- CRX is fast and deterministic — always produces a grammar
- Package grouping — per-directory grammars are the right abstraction level
- Memoization is critical —
_count_concatwithout cache = exponential blowup
Grammar Quality Reality
~85% of grammars are orderless bags (A|B|C)+. ~15% are structured sequential flows:
- Web controller tests:
post→jsonPath→isEqualTo→exchange→expectStatus - API client patterns:
request→header→send→statusCode→jsonPath - Builder chains:
builder→field→value→build→validate
Bags survive because:
- CRX emits one grammar deterministically (no alternative to compare)
lang_size_scoreonly ranks between algorithms (CRX vs iDRegEx), not within CRX's own output- Methods in large packages don't share sequential patterns — they're genuinely unrelated
What Doesn't Work
iDRegEx on Small Bags
iDRegEx achieves only 3.8× tighter on 3-method groups (below the 10× gate threshold).
The gate correctly rejects it. The storage group (91× tighter) was an outlier.
MDL Scoring
Abandoned (ADR-13). MDL rewards short expressions, so generic info+ beats specific
a.b.c.d.e+ (21% vs 98% success in Bex paper).
Cross-Package Grouping
Grouping by first 3 symbols gives 12% coverage but most groups are too sparse. Package-specific patterns are the norm.
Files to Know
| File | Purpose |
|---|---|
bex/tag_preprocessor/analyze.py |
Pipeline: analyze_directory() → analyze_by_package() → _infer_group() |
bex/grammar.py |
AST nodes, _count_concat memoization, _COUNT_CAP = 10^30 |
bex/crx.py |
CRX algorithm (AST-based) |
bex/mdl.py |
lang_size_score(), model_cost(), data_cost() |
bex/idregex.py |
iDRegEx algorithm |
bex/decompose.py |
Decomposition forest |
bex/gbnf.py |
GBNF converter, grammar_structure_score() |
bex/ensemble.py |
infer_ensemble() — combine multiple algorithms |
bex/mcp_server.py |
MCP server with analyze_directory, get_grammar |
bex/tag_preprocessor/code.py |
preprocess_by_method() — AST to behavioral sequences |
How to Run
# Basic analysis (all defaults ON)
python -m bex.tag_preprocessor.analyze /path/to/codebase --verbose
# With custom settings
python -m bex.tag_preprocessor.analyze /path/to/codebase \
--decompose --idregex-refine --min-structure 0.5 --slice package
# Disable idregex refinement
python -m bex.tag_preprocessor.analyze /path/to/codebase --no-idregex-refine
# MCP server
python bex/mcp_server.py --port 8080
Test Coverage
tests/test_analyze.py: Pipeline integration teststests/test_grammar.py: AST node tests, count_words memoizationtests/test_mdl.py: Language Size scoring, MDL scoringtests/test_crx.py: CRX algorithmtests/test_idregex.py: iDRegEx algorithmtests/test_decompose.py: Decomposition foresttests/test_distributional.py: Distributional clusteringtests/test_gbnf.py: GBNF conversiontests/test_crx_refined.py: Refined CRXtests/test_grammar_index.py: Grammar indextests/test_reduce.py: Algorithm 4
Total: 313 tests passing
Decision Log
| Decision | Choice | Rationale |
|---|---|---|
| Scoring | Language Size (not MDL) | MDL rewards short over specific (ADR-13) |
| Decomposition | ON by default | 3× more grammars, fewer pure bags |
| idregex_refine | ON by default | Gate limits to groups where it helps |
| AST representation | Full AST nodes | Type safety, memoization, no string parsing |
| Algorithm | CRX (default) | Fast, deterministic, always produces output |
Open Questions
1. Grammar Usefulness for LLM Code Generation
MCP tools are ready but haven't validated if grammars help an LLM during generation. Need to test:
- Does constrained decoding with GBNF improve code quality?
- Do grammars reduce hallucination in call chains?
2. The 85% Bag Problem
Most grammars are orderless bags. Two possible directions:
- Accept it: Bags represent real diversity in method usage. Not a bug.
- Better grouping: If we group methods by semantic role (not just directory), we might find ordering within sub-groups. Requires understanding method semantics.
3. Cross-Codebase Grammar Reuse
Can grammars from one project inform another? (e.g., "Spring Boot service patterns")
Experiments Summary
| Round | What | Result |
|---|---|---|
| 1-5 | Context strategies | Package grouping wins (12% coverage) |
| 6-10 | Reduce, clustering | Reduce merges states, not packages |
| 11-15 | Distributional, ensemble | CRX is sufficient, no ensemble needed |
| 16-17 | Refined CRX | Better ~78% of the time when useful, but trivial ~36% |
| 18 | Decomposition | 3× more grammars, fewer bags |
| 19 | AST migration | 54.9s→6.7s after memoization fix |
| 20 | Scoring + defaults | Language Size works, decomposition ON, idregex ON |
Next Steps
- Validate grammar usefulness — test constrained decoding with llama.cpp
- Auto-detect decomposition — skip if codebase already structured
- Cross-project grammar reuse — share patterns across codebases
- IDE integration — grammar-aware code completion
- No further algorithmic changes on bags — they're a feature, not a bug