- ACHIEVEMENT_SUMMARY.md: what we built and learned - Round 14: Full pipeline run on 4 codebases (clean symbols) - Round 15: kORE/iDRegEx vs CRX comparison Next: Crucio-inspired clustering for better grouping
13 KiB
13 KiB
Pipeline Architecture — Grammar Inference Engine
Full Pipeline: Source Code → GBNF Grammar
Source Code (directory of .py/.kt/.ts files)
│
▼
┌─────────────────────────────────────────────────────┐
│ 1. SCAN (scan_directory) │
│ Glob *.py/*.kt/*.ts, filter --include/--exclude │
│ Output: {ext: [file_paths]} │
└─────────────────────┬───────────────────────────────┘
│
▼
┌─────────────────────────────────────────────────────┐
│ 2. PREPROCESS (preprocess_by_method) │
│ Tree-sitter parse → extract method-level call │
│ sequences. Coarsen tokens (RETURN/IF/LOOP/EXC). │
│ Output: [(capture, text, line)] per method │
│ │
│ Key: code_bytes = code.encode() for correct │
│ byte/char offset handling (Round 14 fix) │
└─────────────────────┬───────────────────────────────┘
│
▼
┌─────────────────────────────────────────────────────┐
│ 3. GROUP (--slice mode) │
│ │
│ flat │ One group per language (no split) │
│ package │ Group by directory path, merge small │
│ reduce │ Group by dir + Reduce similarity │
│ ilocal │ Group by iLocal context extraction │
│ │
│ Output: {group_label: [sequences]} │
└─────────────────────┬───────────────────────────────┘
│
▼
┌─────────────────────────────────────────────────────┐
│ 4. FILTER (frequency_filter) │
│ Remove symbols in < min_coverage fraction of │
│ methods. Default min_coverage = 0.05 (5%). │
└─────────────────────┬───────────────────────────────┘
│
▼
┌─────────────────────────────────────────────────────┐
│ 5. SPLIT (--split-mixed) │
│ _recursive_split(): group by first symbol, │
│ split until all subgroups have one first-symbol. │
│ max_depth=3. Returns leaf subgroups. │
└─────────────────────┬───────────────────────────────┘
│
▼
┌─────────────────────────────────────────────────────┐
│ 6. INFER (infer_ensemble) │
│ │
│ CRX (always) → 2ms, deterministic │
│ iDRegEx (opt) → 100-700ms, probabilistic │
│ kORE (opt) → 400-700ms, Baum-Welch │
│ │
│ Pick best by lang_size score (lower = tighter). │
│ Output: SORE grammar string │
│ │
│ ┌───────────────────────────────────────────────┐ │
│ │ 6b. iDRegEx REFINEMENT (Round 16, opt-in) │ │
│ │ │ │
│ │ Only triggers when ALL: │ │
│ │ - n_methods ≤ 10 │ │
│ │ - CRX grammar has >50% top-level optionals │ │
│ │ │ │
│ │ If iDRegEx grammar is >10x tighter by │ │
│ │ lang_size: use iDRegEx. Else keep CRX. │ │
│ └───────────────────────────────────────────────┘ │
└─────────────────────┬───────────────────────────────┘
│
▼
┌─────────────────────────────────────────────────────┐
│ 7. VALIDATE (validate_sore + grammar_structure) │
│ - Parse SORE → check syntax │
│ - Compute structure score (0-1) │
│ - Filter by min_structure (default 0.0) │
│ - Reject malformed grammars │
└─────────────────────┬───────────────────────────────┘
│
▼
┌─────────────────────────────────────────────────────┐
│ 8. OUTPUT │
│ Text: Package → Grammar + Score + Args │
│ YAML: Persisted to .dervish/grammars.yml │
│ JSON: Machine-readable export │
│ │
│ GrammarIndex (runtime): for MCP tool queries │
└─────────────────────────────────────────────────────┘
CLI Parameters
python -m bex.tag_preprocessor.analyze <directory> [options]
REQUIRED:
directory Path to scan
GROUPING:
--slice flat|package|reduce|ilocal (default: flat)
--split-mixed Split mixed first-symbol groups
--reduce-threshold FLOAT Reduce similarity (default: 0.15)
--context-strategy STRATEGY iLocal strategy (default: dir)
FILTERING:
--include GLOB Include files matching pattern
--exclude GLOB Exclude files matching pattern
--main-only Exclude test files
--min-coverage FLOAT Symbol frequency threshold (default: 0.05)
--min-methods INT Min methods per group (default: 3)
--min-structure FLOAT Min structure score to keep (default: 0.0)
ALGORITHMS:
--prefer crx|idregex Force single algorithm
--kore Include kORE in ensemble
--idregex Include iDRegEx in ensemble
--idregex-refine Refine CRX flat bags with iDRegEx (Round 16)
--kmax INT Max k for k-ORE (default: 2)
--crx-method standard|refined (default: standard)
--scoring-method langsize|mdl (default: langsize)
OUTPUT:
--format text|json Output format
--json Shortcut for --format json
--verbose Print progress
Decision Matrix: Which Algorithm When?
┌─────────────────────────┬──────────────┬───────────────┬──────────────┐
│ Scenario │ CRX │ iDRegEx │ kORE │
│ │ (0-9ms) │ (100-700ms) │ (400-700ms) │
├─────────────────────────┼──────────────┼───────────────┼──────────────┤
│ Speed │ 1x │ 100-350x │ 200-350x │
│ │ │ slower │ slower │
├─────────────────────────┼──────────────┼───────────────┼──────────────┤
│ Small structured │ ✅ Best │ ✅ Better │ ✅ Same as │
│ (3-5 methods, │ a.b?.c?.d? │ a.(b|c).(d|e) │ iDRegEx │
│ clear branching) │ score=1.0 │ score=0.86 │ │
├─────────────────────────┼──────────────┼───────────────┼──────────────┤
│ Flat CRX bag │ ✅ Always │ ⚠️ Sometimes │ ❌ Often │
│ (many optional parts, │ produces │ None. When it │ None. │
│ small group ≤10) │ something │ works: >10x │ Adds nothing │
│ │ │ tighter! │ over iDRegEx │
├─────────────────────────┼──────────────┼───────────────┼──────────────┤
│ Medium/Large diverse │ ✅ Default │ ❌ None or │ ❌ None or │
│ (>10 methods, │ always works │ trivial │ trivial │
│ diverse patterns) │ │ │ │
├─────────────────────────┼──────────────┼───────────────┼──────────────┤
│ Very diverse │ ✅ Only │ ✅ But trivial │ ✅ But trivial│
│ (no shared structure) │ reliable │ (a.b.c|x.y.z) │ (same) │
│ │ option │ score=0.0 │ │
└─────────────────────────┴──────────────┴───────────────┴──────────────┘
DECISION LOGIC (--idregex-refine enabled):
1. Run CRX → always, fast
2. Count top-level optional parts
3. IF n_methods ≤ 10
AND optionals/parts > 50% → CRX produced flat optional chain
THEN run iDRegEx
IF iDRegEx returns grammar
AND lang_size improvement > 10x
→ use iDRegEx (477x tighter on RAGSAK example!)
ELSE keep CRX
4. ELSE → keep CRX
SCORING:
lang_size_score (Bex et al.): Lower = better
Counts how many words the grammar accepts at each input length.
CRX flat chains: 9432 words (a?.b?.c?.d?.e?.f? → many combos)
iDRegEx disjunctions: 60 words (a.(b|c).(d|e) → exact paths)
grammar_structure_score: Higher = more ordering info
Measures dots, optionals, repetition vs disjunction ratio.
Used for filtering (min_structure), NOT for algorithm selection.
mdl_score: Lower = better (model + data cost)
Fallback scoring method. lang_size is preferred.
Speed Profile (empirical, RAGSAK .kt 462 files)
Phase │ Time │ Notes
──────────────────────────┼─────────┼──────────────────────────────────
Scan + Preprocess │ 2.8s │ ProcessPoolExecutor parallel
Group + Filter │ <0.1s │ In-memory
Split (recursive) │ <0.1s │ Pure Python
CRX Inference (27 groups) │ ~54s │ 27 groups × ~2s each
iDRefinement (1 group) │ ~0.7s │ 1 eligible candidate
Validate + Output │ <0.1s │
──────────────────────────┼─────────┼──────────────────────────────────
Total (--split-mixed) │ ~74s │
Total (+ --idregex-refine)│ ~74.7s │ +0.7s negligible
Without --split-mixed: │ ~13s │ Fewer groups, faster