grammar-inference-engine/docs/adr/0008-bex-ensemble-for-grammar-inference.md
tobjend e23922a1b7
Some checks failed
ci/woodpecker/push/woodpecker Pipeline failed
feat: adaptive multi-assignment clustering; add ADRs 1-10
- Multi-assignment clustering (no greedy 'used' set)
- Adaptive ngram fallback (shrink when (other) > 60%)
- Add docs/adr/ with 10 architecture decision records
- Fix ADR 1 (query modification description)
- Fix ADR 3 (multi-assignment + adaptive shrink)
- Fix ADR 5 (import sort order clarification)
- Fix ADR 6 (remove Kotlin call_suffix references)
- New ADR 9 (adaptive clustering rationale)
- New ADR 10 (universal package mapping via relpath)
2026-07-03 22:58:09 +02:00

60 lines
3 KiB
Markdown

# 8. BEX ensemble for grammar inference
**Date:** 2026-07-03
**Status:** Accepted
## Context
Given a set of symbol sequences (e.g. `["every", "assertEquals", "verify"]`), we need to infer a grammar that concisely describes the pattern. Three algorithms are available:
- **CRX**: Fast, produces unordered CHAREs (e.g. `(a+b+c)+`). Best for vocabulary discovery.
- **iDRegEx**: Slower, produces ordered regex with alternation and optionality (e.g. `a.b.(c|d)?`). Best for small, clean sequences.
- **kOREInference**: Probabilistic, handles noise well (e.g. `a.b.(b?(a|c))`). Best for diverse sequences with outliers.
No single algorithm works best for all codebases. We need to pick the right one for each cluster automatically.
## Decision
Run all three algorithms (ensemble), compute MDL (Minimum Description Length) for each, and select the one with the lowest MDL score.
MDL = grammar_length + sum of per-example encoding costs. Lower is better — the grammar explains the data most compactly.
Ensemble logic in `infer_ensemble()`:
```python
def infer_ensemble(sequences, kmax=2, N=3, prefer=None):
best = None
best_score = float('inf')
for name, fn in [('CRX', crx), ('iDRegEx', idregex), ('kOREInference', kore)]:
if prefer and name.lower() != prefer.lower():
continue
grammar = fn(sequences, ...)
mdl = compute_mdl(grammar, sequences)
if mdl < best_score:
best_score = mdl
best = {'algorithm': name, 'grammar': grammar, 'mdl_score': mdl}
return {'best': best, 'all': all_results, 'why': {...}}
```
Default `kmax=2`, `N=3` (max k for k-ORE, random trials).
## Consequences
**Positive:**
- CRX handles large clusters with diverse vocabulary — produces useful vocabulary bags.
- iDRegEx fires on small, focused clusters (3-12 methods) — produces ordered grammars with exact subsequences.
- kOREInference handles noisy clusters where methods share a theme but vary in exact call order.
- MDL provides a principled, automatic selection criterion.
**Negative:**
- k-ORE algorithms fail on real code when sequences are too diverse (per-file sequences differ more than per-log sequences they were designed for).
- Clustering helps by grouping similar methods before inference.
- iDRegEx can produce overfit grammars on very small clusters (3 methods) — e.g. `every.every.verify.(assertEquals)?` for 3 methods that happen to share an exact sequence.
- MDL comparison assumes grammars are comparable — CRX CHAREs and iDRegEx regex use different notation, so length comparison is approximate.
## Alternatives Considered
- **Single algorithm (CRX only)**: Fast but produces only unordered vocab — misses ordering conventions entirely.
- **Single algorithm (iDRegEx only)**: Produces ordered grammars but fails on diverse inputs (returns `ε`).
- **Single algorithm (kORE only)**: Most robust to noise but slowest, and still fails on highly diverse code sequences.
- **Algorithm per cluster size**: Manual heuristic (CRX for >20 methods, iDRegEx for <10). Harder to tune than MDL-driven selection.