grammar-inference-engine/experiments/DIAGRAMS.md

9.1 KiB

Pipeline Diagrams

1. Full Pipeline

Source Code Directory
        │
        ▼
┌─────────────────┐
│  tree-sitter AST │  Parse each file, extract behavioral prefixes
│  (code.py)       │  coarsen_token(): RETURN, IF, EXCEPTION, LOOP
└────────┬────────┘
         │
         ▼
┌─────────────────┐
│  Preprocessing   │  preprocess_by_method(): extract (capture, text, line) tuples
│  (per file)      │  frequency_filter(): remove symbols in < 5% of files
└────────┬────────┘
         │
         ▼
┌─────────────────┐
│  Grouping        │  --slice package: group by directory
│  (analyze.py)    │  --split-mixed: _recursive_split() by first symbol
└────────┬────────┘
         │
         ▼
┌─────────────────┐
│  Inference       │  CRX (always, 2ms)
│  (per group)     │  Refined CRX (--crx-method refined, ~50ms)
│                  │  iDRegEx (--idregex, ~700ms, opt-in)
│                  │  kORE (--kore, ~400ms, opt-in)
└────────┬────────┘
         │
         ▼
┌─────────────────┐
│  Validation      │  validate_sore(): check SORE is parseable
│                  │  grammar_structure_score(): penalize flat bags
│                  │  model_cost >= 2: filter trivial grammars
└────────┬────────┘
         │
         ▼
┌─────────────────┐
│  Scoring         │  lang_size_score(): count words at each length
│  (mdl.py)        │  mdl_score(): model_cost + data_cost
└────────┬────────┘
         │
         ▼
┌─────────────────┐
│  Output          │  YAML grouped by module
│  (gbnf.py)       │  GBNF conversion for constrained generation
│                  │  GrammarIndex for runtime lookup
└─────────────────┘

2. Inference Decision Tree

Group of sequences
        │
        ▼
┌───────────────────┐
│ CRX (always run)  │──→ grammar
└────────┬──────────┘
         │
         ▼
┌───────────────────┐     ┌──────────────────┐
│ structure >= 0.7? │─Yes─│ Keep CRX         │
└────────┬──────────┘     │ (already good)   │
         │ No             └──────────────────┘
         ▼
┌───────────────────┐     ┌──────────────────┐
│ --crx-method      │─Yes─│ Refined CRX      │
│   = refined?      │     │ (cluster-then-   │
└────────┬──────────┘     │  infer)          │
         │ No             └────────┬─────────┘
         ▼                         │
┌───────────────────┐              ▼
│ --idregex-refine? │     ┌──────────────────┐
│ (n<=10, opt>50%)? │─Yes─│ model_cost >= 2? │
└────────┬──────────┘     └────────┬─────────┘
         │ No                   Yes │   No
         ▼                    ┌─────┘    │
┌──────────────────┐          ▼          ▼
│ Keep CRX         │   Use refined   Keep CRX
│ (default)        │   (tighter)     (trivial)
└──────────────────┘

3. Parameters Reference

Required:
  directory              Path to source code

Grouping:
  --slice                flat | package | reduce | ilocal
  --split-mixed          Split groups by first symbol before inference
  --min-methods N        Skip groups with < N methods (default: 3)

Filtering:
  --min-coverage F       Remove symbols in < F% of files (default: 0.05)
  --min-structure S      Drop grammars with structure < S (default: 0.0)
  --include GLOB         Only include matching files
  --exclude GLOB         Skip matching files
  --main-only            Exclude test files

Inference:
  --crx-method           standard | refined
  --kmax K               Max k for k-ORE algorithms (default: 2)
  --prefer algo          Skip ensemble, use only this algorithm
  --kore                 Include kORE in ensemble (slow, off by default)
  --idregex              Include iDRegEx in ensemble (slow, off by default)
  --idregex-refine       Run iDRegEx on small flat bags (off by default)

Output:
  --format               text | json
  --json                 Shortcut for --format json
  --verbose              Print progress

4. Decision Matrix

                    ┌─────────────────────────────────────────────────┐
                    │              Which algorithm?                    │
                    ├──────────┬──────────┬──────────┬───────────────┤
                    │  Speed   │ Quality  │ Reliable │ Best for      │
┌───────────────────┼──────────┼──────────┼──────────┼───────────────┤
│ CRX (default)     │  2ms     │ Medium   │ Always   │ Most cases    │
│ Refined CRX       │  50ms    │ High     │ ~64%*    │ Flat bags     │
│ iDRegEx           │  700ms   │ High     │ ~30%**   │ Rarely helps  │
│ kORE              │  400ms   │ High     │ ~20%**   │ Don't use     │
└───────────────────┴──────────┴──────────┴──────────┴───────────────┘

 * Refined CRX produces useful grammar 64% of the time (trivial 36%)
** iDRegEx/kORE return None on most real-world data

When to use what:
  Default:                    CRX (--crx-method standard)
  Want tighter grammars:      Refined CRX (--crx-method refined)
  Drop flat bags:             --min-structure 0.3
  Split mixed groups:         --split-mixed

5. Grammar Quality Spectrum

Score:  0.0         0.2         0.5         0.7         1.0
        │           │           │           │           │
        ▼           ▼           ▼           ▼           ▼
    ┌───────┐  ┌───────┐  ┌───────┐  ┌───────┐  ┌───────┐
    │ FLAT  │  │ SEMI  │  │ MIXED │  │STRUCT.│  │ PURE  │
    │  BAG  │  │       │  │       │  │       │  │SEQ.   │
    └───────┘  └───────┘  └───────┘  └───────┘  └───────┘
       │          │          │          │          │
       │          │          │          │          │
    (a+b+c)   a?.(b+c)   a?.(b+c).d  a.b?.c.d   a.b.c.d
       │          │          │          │          │
       ▼          ▼          ▼          ▼          ▼
    NOISE     PARTIAL    USEFUL     USEFUL     EXACT
    (drop)    (keep)     (keep)     (keep)     (keep)

min_structure thresholds:
  0.0  = keep all (default)
  0.2  = drop flat bags (Round 12)
  0.3  = drop semi-structured (recommended)
  0.5  = only keep clearly structured

6. Package Slicing vs Split-by-Symbol

Package Slicing (--slice package):
  Group by directory path
  ┌──────────────┐
  │ src/flask/   │──→ [app.py, views.py, ...] ──→ one grammar per dir
  │ src/auth/    │──→ [login.py, register.py] ──→ one grammar per dir
  │ tests/       │──→ [test_*.py, ...]         ──→ one grammar per dir
  └──────────────┘

Split by First Symbol (--split-mixed):
  Within each package, group by first captured symbol
  ┌──────────────┐
  │ tests/       │
  │  ├─ [return] │──→ methods starting with "return"
  │  ├─ [if]     │──→ methods starting with "if"
  │  ├─ [def]    │──→ methods starting with "def"
  │  └─ [other]  │──→ everything else
  └──────────────┘

Combined (--slice package --split-mixed):
  1. Group by directory
  2. Within each directory, split by first symbol
  3. Infer grammar for each sub-group
  4. Keep all that pass filtering