# Pipeline Diagrams ## 1. Full Pipeline ``` Source Code Directory │ ▼ ┌─────────────────┐ │ tree-sitter AST │ Parse each file, extract behavioral prefixes │ (code.py) │ coarsen_token(): RETURN, IF, EXCEPTION, LOOP └────────┬────────┘ │ ▼ ┌─────────────────┐ │ Preprocessing │ preprocess_by_method(): extract (capture, text, line) tuples │ (per file) │ frequency_filter(): remove symbols in < 5% of files └────────┬────────┘ │ ▼ ┌─────────────────┐ │ Grouping │ --slice package: group by directory │ (analyze.py) │ --split-mixed: _recursive_split() by first symbol └────────┬────────┘ │ ▼ ┌─────────────────┐ │ Inference │ CRX (always, 2ms) │ (per group) │ Refined CRX (--crx-method refined, ~50ms) │ │ iDRegEx (--idregex, ~700ms, opt-in) │ │ kORE (--kore, ~400ms, opt-in) └────────┬────────┘ │ ▼ ┌─────────────────┐ │ Validation │ validate_sore(): check SORE is parseable │ │ grammar_structure_score(): penalize flat bags │ │ model_cost >= 2: filter trivial grammars └────────┬────────┘ │ ▼ ┌─────────────────┐ │ Scoring │ lang_size_score(): count words at each length │ (mdl.py) │ mdl_score(): model_cost + data_cost └────────┬────────┘ │ ▼ ┌─────────────────┐ │ Output │ YAML grouped by module │ (gbnf.py) │ GBNF conversion for constrained generation │ │ GrammarIndex for runtime lookup └─────────────────┘ ``` ## 2. Inference Decision Tree ``` Group of sequences │ ▼ ┌───────────────────┐ │ CRX (always run) │──→ grammar └────────┬──────────┘ │ ▼ ┌───────────────────┐ ┌──────────────────┐ │ structure >= 0.7? │─Yes─│ Keep CRX │ └────────┬──────────┘ │ (already good) │ │ No └──────────────────┘ ▼ ┌───────────────────┐ ┌──────────────────┐ │ --crx-method │─Yes─│ Refined CRX │ │ = refined? │ │ (cluster-then- │ └────────┬──────────┘ │ infer) │ │ No └────────┬─────────┘ ▼ │ ┌───────────────────┐ ▼ │ --idregex-refine? │ ┌──────────────────┐ │ (n<=10, opt>50%)? │─Yes─│ model_cost >= 2? │ └────────┬──────────┘ └────────┬─────────┘ │ No Yes │ No ▼ ┌─────┘ │ ┌──────────────────┐ ▼ ▼ │ Keep CRX │ Use refined Keep CRX │ (default) │ (tighter) (trivial) └──────────────────┘ ``` ## 3. Parameters Reference ``` Required: directory Path to source code Grouping: --slice flat | package | reduce | ilocal --split-mixed Split groups by first symbol before inference --min-methods N Skip groups with < N methods (default: 3) Filtering: --min-coverage F Remove symbols in < F% of files (default: 0.05) --min-structure S Drop grammars with structure < S (default: 0.0) --include GLOB Only include matching files --exclude GLOB Skip matching files --main-only Exclude test files Inference: --crx-method standard | refined --kmax K Max k for k-ORE algorithms (default: 2) --prefer algo Skip ensemble, use only this algorithm --kore Include kORE in ensemble (slow, off by default) --idregex Include iDRegEx in ensemble (slow, off by default) --idregex-refine Run iDRegEx on small flat bags (off by default) Output: --format text | json --json Shortcut for --format json --verbose Print progress ``` ## 4. Decision Matrix ``` ┌─────────────────────────────────────────────────┐ │ Which algorithm? │ ├──────────┬──────────┬──────────┬───────────────┤ │ Speed │ Quality │ Reliable │ Best for │ ┌───────────────────┼──────────┼──────────┼──────────┼───────────────┤ │ CRX (default) │ 2ms │ Medium │ Always │ Most cases │ │ Refined CRX │ 50ms │ High │ ~64%* │ Flat bags │ │ iDRegEx │ 700ms │ High │ ~30%** │ Rarely helps │ │ kORE │ 400ms │ High │ ~20%** │ Don't use │ └───────────────────┴──────────┴──────────┴──────────┴───────────────┘ * Refined CRX produces useful grammar 64% of the time (trivial 36%) ** iDRegEx/kORE return None on most real-world data When to use what: Default: CRX (--crx-method standard) Want tighter grammars: Refined CRX (--crx-method refined) Drop flat bags: --min-structure 0.3 Split mixed groups: --split-mixed ``` ## 5. Grammar Quality Spectrum ``` Score: 0.0 0.2 0.5 0.7 1.0 │ │ │ │ │ ▼ ▼ ▼ ▼ ▼ ┌───────┐ ┌───────┐ ┌───────┐ ┌───────┐ ┌───────┐ │ FLAT │ │ SEMI │ │ MIXED │ │STRUCT.│ │ PURE │ │ BAG │ │ │ │ │ │ │ │SEQ. │ └───────┘ └───────┘ └───────┘ └───────┘ └───────┘ │ │ │ │ │ │ │ │ │ │ (a+b+c) a?.(b+c) a?.(b+c).d a.b?.c.d a.b.c.d │ │ │ │ │ ▼ ▼ ▼ ▼ ▼ NOISE PARTIAL USEFUL USEFUL EXACT (drop) (keep) (keep) (keep) (keep) min_structure thresholds: 0.0 = keep all (default) 0.2 = drop flat bags (Round 12) 0.3 = drop semi-structured (recommended) 0.5 = only keep clearly structured ``` ## 6. Package Slicing vs Split-by-Symbol ``` Package Slicing (--slice package): Group by directory path ┌──────────────┐ │ src/flask/ │──→ [app.py, views.py, ...] ──→ one grammar per dir │ src/auth/ │──→ [login.py, register.py] ──→ one grammar per dir │ tests/ │──→ [test_*.py, ...] ──→ one grammar per dir └──────────────┘ Split by First Symbol (--split-mixed): Within each package, group by first captured symbol ┌──────────────┐ │ tests/ │ │ ├─ [return] │──→ methods starting with "return" │ ├─ [if] │──→ methods starting with "if" │ ├─ [def] │──→ methods starting with "def" │ └─ [other] │──→ everything else └──────────────┘ Combined (--slice package --split-mixed): 1. Group by directory 2. Within each directory, split by first symbol 3. Infer grammar for each sub-group 4. Keep all that pass filtering ```