grammar-inference-engine/experiments/ROUND19_PLAN.md
tobjend 9d206f586f docs: Round 19 evaluation plan + golden config
- Plan: match rate audit, decomposition A/B, distributional A/B, fragment quality
- Golden config: best-known values for all heuristic parameters
- Single source of truth in bex/golden_config.py
2026-07-12 20:18:33 +02:00

4.6 KiB
Raw Blame History

Round 19: Crucio Evaluation & Golden Config

Goal: Quantify whether Crucio-inspired changes (distributional clustering, decomposition forest) actually improved grammar quality. Define the "golden config" with best heuristic values.

Open questions:

  1. Do the 21 RAGSAK / 22 FastAPI / 11 kotlinx patterns that survive min_structure=0.5 actually match their input sequences?
  2. Did decomposition improve pattern quality, or just create more patterns?
  3. Did distributional clustering help when it differs from first-symbol?
  4. What's the compression ratio (grammar size vs input size)?

Experiment 1: Match Rate Audit

What: For each YAML entry, check what fraction of input sequences the grammar actually accepts via _matches().

Method:

  1. Run pipeline with golden config on all 3 codebases
  2. For each surviving pattern (min_structure=0.5), reproduce the input sequences
  3. Check _matches(grammar, seq) for each input sequence
  4. Report: match rate, grammar size, compression ratio

Metrics:

  • match_rate = sequences accepted / total sequences (1.0 = perfect)
  • compression = len(grammar) / total_input_chars (< 1.0 = good)
  • structure_score = grammar_structure_score(grammar)

Experiment 2: Decomposition A/B

What: Compare pipeline output WITH decomposition ON vs OFF.

Method:

  1. Run with --decompose --max-seq-length 4 (current config)
  2. Run without --decompose
  3. Compare: pattern count, avg match rate, avg structure score, avg compression

Hypothesis: Decomposition helps diverse codebases (RAGSAK, FastAPI) but hurts already-structured ones (kotlinx.coroutines).


Experiment 3: Distributional vs First-Symbol A/B

What: Compare --cluster-method distributional vs --cluster-method first-symbol.

Method:

  1. Run with --cluster-method distributional
  2. Run with --cluster-method first-symbol (current default)
  3. Compare on the groups where they differ

Hypothesis: Distributional clustering doesn't help because the contexts are already too specific per-package.


Experiment 4: Decomposition Fragment Quality

What: Are decomposition fragments meaningful sub-patterns or noise?

Method:

  1. Take decomposed fragments from a diverse package (e.g., RAGSAK agents/rag/embabel)
  2. Check match rate of each fragment's grammar
  3. Check if fragments capture real sub-patterns (e.g., "return path" vs "error path")

Golden Config

The "golden config" captures our best-known heuristic values:

GOLDEN_CONFIG = {
    # Core pipeline
    "min_coverage": 0.05,      # BEX outlier threshold (was 0.8, too aggressive)
    "min_methods": 3,           # Min methods per group (was 5, lost too many)
    "method": "langsize",       # Scoring: Language Size (Bex et al.)
    
    # Grouping
    "slice": "package",         # Per-directory (not flat, not reduce)
    "split_mixed": True,        # Recursive split by first symbol
    "max_depth": 3,             # Max recursion depth for split
    "cluster_method": "first-symbol",  # Split method (distributional = no improvement)
    
    # Quality filter
    "min_structure": 0.5,       # Drop flat bags (noise)
    "max_mdl": 200.0,           # Drop high-score grammars
    
    # Decomposition (Crucio Phase 2)
    "decompose": True,          # Break long sequences into fragments
    "max_seq_length": 4,        # Max fragment length (5 = too aggressive, 4 = sweet spot)
    
    # Algorithms
    "crx_method": "standard",   # Standard CRX (refined = trivial on large groups)
    "include_kore": False,      # kORE = slow, no improvement
    "include_idregex": False,   # iDRegEx = slow, rare benefit
    "idregex_refine": False,    # iDRegEx refinement = rare benefit
}

Rationale for each value:

  • min_coverage=0.05: At 0.8, almost all symbols filtered out. 0.05 sees real vocabulary.
  • min_methods=3: At 5, lost 9 FastAPI grammars. 3 is safe minimum.
  • min_structure=0.5: Below this, patterns are flat bags (noise).
  • decompose=True: Helps RAGSAK 7×, FastAPI 1.8×. Hurts kotlinx.coroutines.
  • max_seq_length=4: At 5, fragments too short. 4 captures meaningful sub-patterns.
  • cluster_method="first-symbol": Distributional clustering showed no improvement.
  • crx_method="standard": Refined CRX produces trivial output 36% of the time.

Execution Order

  1. Write plan + golden config → commit
  2. Experiment 1: Match rate audit (all 3 codebases)
  3. Experiment 2: Decomposition A/B
  4. Experiment 3: Distributional vs first-symbol A/B
  5. Experiment 4: Decomposition fragment quality
  6. Update golden config if needed
  7. Summarize results and implications