- ACHIEVEMENT_SUMMARY.md: what we built and learned - Round 14: Full pipeline run on 4 codebases (clean symbols) - Round 15: kORE/iDRegEx vs CRX comparison Next: Crucio-inspired clustering for better grouping
120 lines
5.1 KiB
Markdown
120 lines
5.1 KiB
Markdown
# Achievement Summary — Grammar Inference Pipeline
|
|
|
|
## What We Built
|
|
|
|
A source code analysis pipeline that discovers per-package calling conventions
|
|
from any codebase using tree-sitter AST → behavioral sequences → BEX grammar
|
|
inference. Zero per-language code, one pipeline for all 10 supported languages.
|
|
|
|
### Pipeline Components
|
|
|
|
1. **Preprocessing** (`code.py`)
|
|
- tree-sitter AST parsing with correct byte/char offset handling
|
|
- Behavioral prefix extraction (BEHAVIORAL_PREFIXES)
|
|
- Token coarsening (RETURN, IF, EXCEPTION, LOOP)
|
|
- Method-level call sequence extraction
|
|
|
|
2. **Grouping** (`analyze.py`)
|
|
- Package slicing by directory
|
|
- Split-by-first-symbol (`_recursive_split()`)
|
|
- Frequency filtering (remove rare symbols)
|
|
|
|
3. **Inference** (`crx.py`, `crx_refined.py`)
|
|
- CRX: deterministic, fast (2ms), always produces output
|
|
- Refined CRX: cluster-then-infer, tighter grammars on flat bags
|
|
- Optional: iDRegEx (`--idregex-refine`), kORE (`--kore`)
|
|
|
|
4. **Validation** (`gbnf.py`)
|
|
- SORE validation (`validate_sore()`)
|
|
- Structure scoring (`grammar_structure_score()`)
|
|
- GBNF conversion for constrained generation
|
|
|
|
5. **Scoring** (`mdl.py`)
|
|
- Language Size (Bex et al.) — primary metric
|
|
- MDL — fallback metric
|
|
- Model cost — grammar complexity
|
|
|
|
6. **Runtime** (`grammar_index.py`, `mcp_server.py`)
|
|
- GrammarIndex for lookup by (package, context_symbol)
|
|
- MCP tools: `get_grammar`, `get_package_grammars`, `analyze_directory`
|
|
|
|
### Key Fixes
|
|
|
|
- **Byte/char offset mismatch** — tree-sitter returns byte offsets, we indexed into strings with byte offsets. Fixed by encoding to bytes first. 653 truncated symbols → 0.
|
|
- **GBNF parser** — `+` inside `()` = alternation, outside = repetition. Fixed.
|
|
- **Symbol sanitization** — removed entirely (dead code after byte/char fix).
|
|
|
|
### Experiments Conducted
|
|
|
|
| Round | What | Finding |
|
|
|-------|------|---------|
|
|
| 1 | Context strategies | Package grouping is best |
|
|
| 2 | Reduce algorithm | Doesn't help (contexts too specific) |
|
|
| 3-4 | kORE, frequency filtering | kORE too slow, sweet spot 0.01-0.05 |
|
|
| 5-6 | SORE/GBNF conversion | Fixed parser, 28 tests |
|
|
| 7-8 | Token coarsening | Coarsened tokens → fewer flat bags |
|
|
| 9-10 | CRX over-approximation | 24% over-approximated, refined helps |
|
|
| 11 | Pipeline speed | 19s → 2.8s (6.8x) via parallelism |
|
|
| 12 | Structure scoring | min_structure=0.2 drops flat bags |
|
|
| 13 | Recursive split | +67% more grammars, avg score 0.45→0.71 |
|
|
| 14 | Byte/char fix | 653→0 truncated symbols, 3.4x more grammars |
|
|
| 15 | kORE/iDRegEx vs CRX | iDRegEx/kORE return None on flat bags |
|
|
| 16 | iDRegEx refinement | Heuristic: n<=10, opt>50%, 477x improvement |
|
|
| 17 | CRX vs refined CRX | Refined wins 78% when useful, trivial 36% |
|
|
|
|
### What We Achieved
|
|
|
|
- **Pipeline works**: 462 Kotlin files → 27 structured grammars (RAGSAK)
|
|
- **Multi-language**: Tested on Kotlin, Python, TypeScript
|
|
- **Fast**: 13s on RAGSAK (with --idregex-refine), 74s without
|
|
- **Clean symbols**: 0 truncated (was 653)
|
|
- **28 GBNF tests**: Parser handles real-world grammars
|
|
- **234 total tests**: Full test suite passes
|
|
|
|
### What We Learned
|
|
|
|
1. **CRX is the right default** — fast, reliable, always produces something
|
|
2. **Refined CRX helps on flat bags** — cluster-then-infer finds tighter groupings
|
|
3. **iDRegEx is too situational** — returns None on most real data
|
|
4. **kORE adds nothing** — same as iDRegEx but slower
|
|
5. **Flat bags are grouping problems** — no algorithm can find structure where there is none
|
|
6. **lang_size is the right metric** — counts words at each length, prefers tighter grammars
|
|
|
|
### Current State
|
|
|
|
**Pipeline**: Fully functional, tested on 4 codebases
|
|
**Default**: CRX only, refined opt-in via `--crx-method refined`
|
|
**Optional**: iDRegEx refinement (`--idregex-refine`), kORE (`--kore`)
|
|
**Output**: YAML by module, GBNF for constrained generation, MCP tools for runtime
|
|
|
|
### What's Next
|
|
|
|
1. **Integrate with LLM** — test if grammars actually help code completion
|
|
2. **Tune min_structure** — find the sweet spot for each use case
|
|
3. **Add more languages** — currently 10 supported via tree-sitter
|
|
4. **Package as skill** — reusable agent skill for wiki/MCP integration
|
|
|
|
### Git History
|
|
|
|
```
|
|
b2c1263 docs: ASCII diagrams for pipeline, parameters, and decision matrix
|
|
1d94c09 experiment: CRX vs refined CRX across 3 codebases (Round 17)
|
|
b92b765 feat: iDRegEx refinement for CRX flat bags (Round 16)
|
|
e9f672c experiment: kORE/iDRegEx vs CRX on flat bags (Round 15)
|
|
93d53f1 fix: byte/char offset mismatch in tree-sitter text extraction
|
|
...
|
|
```
|
|
|
|
### Key Files
|
|
|
|
- `bex/tag_preprocessor/analyze.py` — main pipeline orchestrator
|
|
- `bex/tag_preprocessor/code.py` — tree-sitter preprocessing
|
|
- `bex/crx.py` — CRX algorithm
|
|
- `bex/crx_refined.py` — refined CRX (cluster-then-infer)
|
|
- `bex/gbnf.py` — SORE→GBNF converter
|
|
- `bex/mdl.py` — scoring (lang_size, mdl)
|
|
- `bex/grammar_index.py` — runtime grammar lookup
|
|
- `bex/mcp_server.py` — MCP server with tools
|
|
- `experiments/DECISION_MATRIX.md` — when to use what
|
|
- `experiments/DIAGRAMS.md` — ASCII pipeline diagrams
|
|
- `experiments/EXPERIMENT_LOG.md` — full experiment history
|