- iDRegEx opt-in: Flask 55s→2.8s, RAGSAK 74s→13s, FastAPI 30s→30s - GBNF validation: 130/130 OK, 0 FAIL (17 malformed skipped) - Grammar quality: 43% structured, rest flat/trivial - Logged to experiments/EXPERIMENT_LOG.md
479 lines
20 KiB
Markdown
479 lines
20 KiB
Markdown
# Experiment Log — Grammar Inference Pipeline
|
||
|
||
Track what we tried, what worked, what failed, and what's next. Each experiment
|
||
includes: hypothesis, method, result, verdict.
|
||
|
||
---
|
||
|
||
## Round 1: Context Strategies (commit `bbdfe93`)
|
||
|
||
**Hypothesis:** The calling context (prefix before the method body) determines
|
||
which methods share a convention. Better context grouping → better grammars.
|
||
|
||
**Method:** Tested 4 strategies on RAGSAK (Kotlin) and Flask (Python):
|
||
- **Baseline:** Group by package directory
|
||
- **Option A:** Group by last k components of file path (`file_path_k{k}`)
|
||
- **Option B:** Group by first k symbols of call sequence (`first_k_sym_{k}`)
|
||
- **Option C:** Two-dimensional: (path_k, first_k_symbols)
|
||
|
||
**Result:**
|
||
| Strategy | RAGSAK patterns | RAGSAK coverage | Flask patterns | Flask coverage |
|
||
|----------|----------------|-----------------|----------------|---------------|
|
||
| Package baseline | 73 | 12.0% | 21 | 10.7% |
|
||
| File path k=1 | 73 | 12.0% | 21 | 10.7% |
|
||
| First k=1 | 20 | 4.6% | 2 | 1.4% |
|
||
| First k=3 | 47 | 12.0% | 21 | 10.7% |
|
||
|
||
**Verdict:** Package grouping and first_k_sym_3 produce similar results.
|
||
Cross-package grouping by first symbol is too sparse — most groups are either
|
||
too large (skipped) or too diverse (skipped). The useful patterns are
|
||
package-specific, not cross-package.
|
||
|
||
---
|
||
|
||
## Round 2: Reduce Algorithm (commit `b516b29`)
|
||
|
||
**Hypothesis:** Reduce (Algorithm 4, TODS 2010) merges structurally similar
|
||
contexts, revealing cross-package patterns by unifying equivalent states.
|
||
|
||
**Method:** Implemented faithful Reduce with support-weighted SOA edit distance,
|
||
adjunction, iterative merging, and minimize. Tested at ε=0.05 to 0.4.
|
||
|
||
**Result:**
|
||
- RAGSAK at ε=0.3: 1 merge (`JobStatus.every.getJobStatus` ↔ `JobStatus.now.minusMinutes`)
|
||
- Flask at ε=0.3: 1 merge (`def.boolean` ↔ `def.is_boolean`)
|
||
- Coverage improvement: negligible (< 1%)
|
||
|
||
**Verdict:** Reduce doesn't help. The contexts we produce are already too
|
||
specific (unique per package) for the distance metric to find meaningful merges.
|
||
Reduce works when you have a large SOA with many equivalent states — we have
|
||
one SOA per package with few states. Wrong abstraction level.
|
||
|
||
**Why it failed:** Reduce merges states in a single automaton. We're producing
|
||
one automaton per package group. There's nothing to merge across packages
|
||
because each package gets its own inference run. Reduce would need to operate
|
||
on a cross-package SOA, which we don't build.
|
||
|
||
---
|
||
|
||
## Round 3: Language Size Scoring (commit `dfb56a0`)
|
||
|
||
**Hypothesis:** Bex et al.'s Language Size measure (arXiv:1004.2372, Section
|
||
4.3.1) is a better scoring function than MDL for our use case.
|
||
|
||
**Method:** Implemented `lang_size_score()` as default scoring method. Added
|
||
diversity threshold: skip groups with unique_ratio > 0.9 or methods < 5.
|
||
|
||
**Result:** 39 new tests. Pipeline runs correctly with new scoring. Coverage
|
||
numbers similar to before — scoring method doesn't change which patterns are
|
||
found, just which grammar is selected per group.
|
||
|
||
**Verdict:** Scoring method is not the bottleneck. The problem is upstream
|
||
(pattern extraction), not downstream (pattern selection).
|
||
|
||
---
|
||
|
||
## Round 4: GBNF Output (commit `011df39`)
|
||
|
||
**Hypothesis:** SORE → GBNF conversion enables constrained LLM generation.
|
||
SORE operators map directly to GBNF syntax.
|
||
|
||
**Method:** Implemented recursive descent parser for SORE, AST intermediate
|
||
representation, and GBNF renderer. 15 tests.
|
||
|
||
**Result:** All tests pass. `to_gbnf('raise.(ValueError)+')` →
|
||
`"raise" "ValueError"+`. Correct mapping of +, ?, *, |, ., parens.
|
||
|
||
**Verdict:** Implementation works. But the input SOREs are too specific
|
||
to individual packages to be useful for constrained generation. The converter
|
||
is correct; the patterns it converts are the problem.
|
||
|
||
---
|
||
|
||
## Round 5: Cross-Package Exact Matches
|
||
|
||
**Hypothesis:** Some call sequences appear verbatim in multiple packages.
|
||
These are the real cross-package conventions.
|
||
|
||
**Method:** Grouped all sequences by exact tuple match across packages.
|
||
|
||
**Result:** 38 exact cross-package sequences in RAGSAK. Most are trivial:
|
||
- `('clearAllMocks',)` — 4 packages (test teardown)
|
||
- `('Builder',)` — 4 packages (builder pattern)
|
||
- `('Any',)` — 4 packages (Kotlin type)
|
||
- `('get',)` — 4 packages (getter)
|
||
|
||
Interesting ones:
|
||
- `('assumeTrue', 'isDockerAvailable', 'start', 'pullAndWarmup')` — 4 packages (Docker test setup)
|
||
- `('isNullOrBlank', 'error', 'error')` — 3 packages (null check → error)
|
||
- `('sortedBy', 'map', 'toDescriptor')` — 3 packages (data pipeline)
|
||
- `('ObjectMapper', 'findAndRegisterModules')` — 2 packages (Jackson config)
|
||
|
||
**Verdict:** Exact matches are too rare and mostly trivial. The real
|
||
cross-package patterns are structural, not textual — "null check → error"
|
||
appears with different method names in different packages.
|
||
|
||
---
|
||
|
||
## Round 6: Structural Coarsening (Experiment `coarsen_eval.py`)
|
||
|
||
**Hypothesis:** Coarsening structural tokens (keywords, types) while keeping
|
||
function names raw reveals cross-package patterns that pure-text misses.
|
||
|
||
**Method:** `coarsen_token()` maps tree-sitter capture names to categories.
|
||
Function calls kept as raw text (they ARE the content). Only structural tokens
|
||
coarsened: RETURN, RAISE, IF, LOOP, EXCEPTION, KW, TYPE, FUN, ATTR.
|
||
|
||
**Result:**
|
||
|
||
| Codebase | Structural % | Raw coverage | Coarsened coverage | Δ |
|
||
|----------|-------------|--------------|-------------------|---|
|
||
| RAGSAK | 4.7% | 18.6% | 18.3% | -0.3% |
|
||
| Flask | 36.1% | 17.5% | **28.8%** | **+11.3%** |
|
||
|
||
Key findings:
|
||
- RAGSAK: 95.3% function calls → coarsening has nothing to work with
|
||
- Flask: 36.1% structural → coarsening significantly improves grouping
|
||
- Flask k=3: cross-package contexts increase 40 → 47
|
||
- Coarsened cross-package shapes: `('IF', 'KW', 'RETURN')` in 4 packages,
|
||
`('RETURN', 'render_template', 'render_template')` in 6 packages
|
||
|
||
**Verdict:** Coarsening helps codebases with rich structural tokens (Python:
|
||
IF, LOOP, EXCEPTION, KW). Doesn't help codebases dominated by function calls
|
||
(Kotlin: 95% calls). The approach is sound but language-dependent in practice —
|
||
depends on how rich the highlights.scm is.
|
||
|
||
**What we learned:**
|
||
- The 5% structural tokens DO carry signal when they exist
|
||
- Python highlights.scm is much richer than Kotlin's
|
||
- Coarsening is not dead — it's a tool for languages with rich captures
|
||
- The real question is whether the coarsened patterns are USEFUL, not just
|
||
whether they exist
|
||
|
||
---
|
||
|
||
## Round 6b: Kotlin Capture Fix + Minimal Coarsening
|
||
|
||
**Hypothesis:** Kotlin's highlights.scm uses bare captures (`conditional`,
|
||
`exception`, `repeat`) while `BEHAVIORAL_PREFIXES` expected dotted forms
|
||
(`keyword.conditional`, etc.). All Kotlin structural context was being dropped.
|
||
|
||
**Method:** Added bare captures to BEHAVIORAL_PREFIXES. Dropped ATTR, TYPE,
|
||
VAR from coarsening (too noisy). Only coarsen RETURN, IF, EXCEPTION, LOOP.
|
||
|
||
**Result:**
|
||
|
||
| Codebase | Raw coverage | Coarsened coverage | Cross-pkg contexts |
|
||
|----------|-------------|-------------------|-------------------|
|
||
| RAGSAK k=3 | 18.5% | 6.0% | 87 (was 69) |
|
||
| Flask k=2 | 4.5% | 15.8% | 35 (was 54) |
|
||
| Flask k=3 | 17.1% | 17.3% | 28 (was 40) |
|
||
|
||
Key cross-package patterns (coarsened):
|
||
- `('IF', 'isEmpty', 'isEmpty')` — 8 RAGSAK packages (null-check convention)
|
||
- `('IF', 'isNullOrBlank', 'isNullOrBlank')` — 6 RAGSAK packages
|
||
- `('RETURN', 'render_template', 'render_template')` — 5 Flask packages
|
||
- `('IF', 'RETURN')` — 8 RAGSAK packages (guard clause pattern)
|
||
|
||
**Verdict:** The 4 high-signal categories (RETURN, IF, EXCEPTION, LOOP) DO
|
||
reveal cross-package conventions. Coarsening trades per-package coverage for
|
||
cross-package reach. Whether this is useful depends on the use case:
|
||
- For code completion: raw is better (specific method names)
|
||
- For documentation: coarsened is better (structural conventions)
|
||
|
||
---
|
||
|
||
## Round 7: Four-Codebase Evaluation
|
||
|
||
**Method:** Run raw vs coarsened on RAGSAK (Kotlin), Flask (Python),
|
||
kotlinx.coroutines (Kotlin), FastAPI (Python).
|
||
|
||
**Results (k=3):**
|
||
|
||
| Codebase | Raw cov | Coarse cov | Raw xpkg | Coarse xpkg | Δ |
|
||
|----------|---------|------------|----------|-------------|---|
|
||
| RAGSAK | 18.5% | 6.0% | 69 | 87 | +18 |
|
||
| Flask | 17.1% | 17.3% | 40 | 27 | -13 |
|
||
| Coroutines | 30.7% | 19.1% | 524 | 539 | +15 |
|
||
| FastAPI | 12.4% | 20.7% | 82 | 97 | +15 |
|
||
|
||
Cross-package patterns discovered:
|
||
- FastAPI: `('response', 'client', 'get')` — 72 packages (HTTP request pattern)
|
||
- Coroutines: `('RETURN', 'EXCEPTION', 'UnsupportedOperationException')` — 13 packages
|
||
- RAGSAK: `('IF', 'isEmpty', 'isEmpty')` — 8 packages (null-check)
|
||
- Flask: `('RETURN', 'render_template', 'render_template')` — 5 packages
|
||
|
||
**Verdict:** The conventions vs completions tradeoff is real and measurable.
|
||
Coarsening consistently trades per-package coverage for cross-package reach.
|
||
FastAPI is the exception: coverage improves (12.4% → 20.7%) because its
|
||
structural tokens (response/client patterns) are highly repetitive.
|
||
|
||
---
|
||
|
||
## Round 8: Frequency Threshold Sweep
|
||
|
||
**Hypothesis:** The fixed `min_coverage=0.2` is too aggressive. Lower thresholds
|
||
reveal more patterns while still filtering noise.
|
||
|
||
**Method:** Test thresholds 0.00–0.20 on all 4 codebases. Measure symbol count,
|
||
surviving sequences, SORE success, coverage.
|
||
|
||
**Results:**
|
||
|
||
| Codebase | Thresh | Syms | Seqs | SOREs | Coverage |
|
||
|----------|--------|------|------|-------|----------|
|
||
| RAGSAK | 0.01 | 130 | 1270 | 61 | 26.1% |
|
||
| RAGSAK | 0.05 | 16 | 919 | 40 | 45.4% |
|
||
| RAGSAK | 0.10 | 5 | 657 | 19 | 69.1% |
|
||
| Flask | 0.01 | 47 | 784 | 28 | 35.2% |
|
||
| Flask | 0.05 | 5 | 414 | 7 | 44.4% |
|
||
| Flask | 0.10 | 2 | 237 | 5 | 100.0% |
|
||
| Coroutines | 0.01 | 76 | 4802 | 175 | 40.4% |
|
||
| FastAPI | 0.01 | 23 | 2678 | 14 | 17.0% |
|
||
|
||
Key findings:
|
||
- Coverage increases with threshold (trivial: 1 symbol = 100% coverage)
|
||
- Sweet spot: 0.01–0.05. Enough symbols for meaningful patterns, enough
|
||
filtering to remove noise.
|
||
- At 0.01: RAGSAK gets `warn.status.body.ErrorResponse` (real convention)
|
||
- At 0.05: that pattern disappears (too aggressive)
|
||
- Flask dies at 0.15+ (0 symbols survive)
|
||
|
||
---
|
||
|
||
## What We Learned (Summary)
|
||
|
||
1. **Per-package grouping is too sparse.** 1-3 sequences per package isn't
|
||
enough for any inference method to produce general patterns.
|
||
|
||
2. **Cross-package exact matches are rare.** Only 38 in RAGSAK, mostly trivial
|
||
single-call sequences.
|
||
|
||
3. **Reduce doesn't help at our abstraction level.** It merges states within
|
||
one automaton; we need to merge patterns across packages.
|
||
|
||
4. **Scoring/selection isn't the bottleneck.** MDL vs Language Size doesn't
|
||
change what patterns are found.
|
||
|
||
5. **The calling context prefix is the right signal** but grouping by it
|
||
produces groups that are either too large, too diverse, or trivial.
|
||
|
||
6. **GBNF converter works correctly** but the input patterns are too specific.
|
||
|
||
---
|
||
|
||
## Next: Structural Coarsening + Cross-Package Detection
|
||
|
||
### Idea
|
||
|
||
Collapse method names → categories using tree-sitter capture names. This
|
||
converts textual sequences into structural shapes:
|
||
|
||
```
|
||
('isNullOrBlank', 'error', 'error') → (CALL, ERROR, ERROR)
|
||
('raise', 'ValueError', 'ValueError') → (CALL, ERROR, ERROR)
|
||
```
|
||
|
||
Same structural shape, different packages → cross-package convention.
|
||
|
||
### Why This Might Work
|
||
|
||
- We already extract tree-sitter capture names in `code.py:56-69`
|
||
- We already classify nodes into categories in `code.py:72-100`
|
||
(`lit`, `call`, `var`, `lambda`, `kwarg`, `expr`, `template`, `other`)
|
||
- The behavioral prefix filter (`CALL_PREFIXES`) keeps raw text; we need a
|
||
parallel path that keeps the category instead
|
||
- Coarsened sequences have smaller alphabets → more methods per group →
|
||
better inference
|
||
- Patterns like `(CALL, ERROR, ERROR)` are meaningful conventions that
|
||
repeat across packages
|
||
|
||
### What We Need
|
||
|
||
1. **Coarsening map:** `capture_name → category` using the existing
|
||
`CALL_PREFIXES`, `ARG_LITERAL_TYPES`, `LAMBDA_TYPES` classifications
|
||
plus a new `ERROR_TYPES` set
|
||
|
||
2. **Coarsened sequence extraction:** Same pipeline as now, but output
|
||
category labels instead of method names
|
||
|
||
3. **Cross-package grouping:** Group by coarsened shape (first k categories),
|
||
find shapes that appear in ≥2 packages
|
||
|
||
4. **SORE inference on coarsened sequences:** Smaller alphabet, more examples
|
||
per group → better patterns
|
||
|
||
5. **Evaluation:** Compare coarsened patterns vs raw patterns on:
|
||
- Coverage (% of methods in learned groups)
|
||
- Cross-package reach (# of packages per pattern)
|
||
- Usefulness for constrained generation (GBNF quality)
|
||
|
||
### Open Questions
|
||
|
||
- Does coarsening lose too much specificity? `(CALL, ERROR, ERROR)` is less
|
||
informative than `(raise, ValueError, ValueError)` — is the tradeoff worth it?
|
||
- What categories to use? The existing classifications in `code.py` are a
|
||
starting point but may need refinement (e.g., separating ERROR from CALL)
|
||
- How to handle the long tail? Most sequences are 1-2 symbols — coarsening
|
||
doesn't help much for those
|
||
- Is the GBNF output useful at all? Maybe the output should be a conditional
|
||
frequency table instead of a grammar
|
||
|
||
### Experiment Design
|
||
|
||
**Phase 1: Coarsening Proof of Concept**
|
||
- Implement coarsening map in `code.py`
|
||
- Add `--coarsen` flag to CLI
|
||
- Run on RAGSAK + Flask, compare raw vs coarsened patterns
|
||
- Measure: alphabet size reduction, group size increase, pattern count
|
||
|
||
**Phase 2: Cross-Package Detection**
|
||
- Group coarsened sequences by shape (first k categories)
|
||
- Find shapes appearing in ≥2 packages
|
||
- For each shape, infer SORE on coarsened sequences
|
||
- Measure: cross-package patterns found, coverage improvement
|
||
|
||
**Phase 3: Output Quality**
|
||
- Convert coarsened SOREs to GBNF
|
||
- Evaluate: are the GBNF rules more general/useful than raw SOREs?
|
||
- Compare: coarsened GBNF vs raw GBNF vs conditional frequency table
|
||
|
||
---
|
||
|
||
## The "Redacted" Concept
|
||
|
||
When collapsing method names to categories, we lose the specific method name
|
||
but gain the structural pattern. This is a form of **abstraction** — moving
|
||
from concrete examples to general rules.
|
||
|
||
The question is whether the abstraction is at the right level:
|
||
- Too specific: `(raise, ValueError, ValueError)` — package-specific noise
|
||
- Right level: `(CALL, ERROR, ERROR)` — cross-package convention
|
||
- Too abstract: `(X, Y, Y)` — trivial, tells the LLM nothing
|
||
|
||
The middle ground depends on the category vocabulary. We need enough categories
|
||
to be informative (CALL, ERROR, LIT, VAR, TYPE, etc.) but not so many that
|
||
every sequence is unique.
|
||
|
||
---
|
||
|
||
## Round 9: CRX Over-Approximation Analysis
|
||
|
||
**Hypothesis:** Standard CRX produces over-approximated grammars when the
|
||
Hasse diagram is non-linear (branching). This results in flat disjunctions
|
||
like `(a+b+c+d)+?` that accept almost any combination — useless as conventions.
|
||
|
||
**Method:**
|
||
1. Measured tightness across RAGSAK packages: fraction of symbol pairs in data
|
||
vs total possible pairs. Average: 0.300 (only 30% of pairs actually occur).
|
||
2. Measured over-approximation rate: 24% of packages have `+?` factors with
|
||
4+ symbols — massive over-approximation.
|
||
3. Analyzed CRX algorithm (Algorithm 3, Bex et al. VLDB 2006):
|
||
- CRX computes equivalence classes ≈_S (mutual reachability)
|
||
- Merges singletons with identical (Pred, Succ) in Hasse diagram
|
||
- Key limitation: **only merges singletons**, not multi-symbol classes
|
||
- Theorem 5: CRX is optimal **only when** Γ_W is linearly ordered
|
||
- Non-linear → suboptimal (paper counterexample: `{abc, ade, abe}` →
|
||
`a.b?.d?.c?.e?` instead of better `a.(b+d).(c+e)`)
|
||
|
||
**Findings:**
|
||
- CRX was designed for XML DTDs with hierarchical structure. Code call
|
||
sequences have branching patterns that CHAREs can't represent.
|
||
- The `+?` (zero-or-more) factor is the over-approximation signal: it means
|
||
"any subset of these symbols in any order" — which is trivially true.
|
||
- Standard CRX CAN'T fix this — the CHARE representation is inherently linear.
|
||
- Two possible improvements: (a) detect over-approximation, (b) cluster before
|
||
inferring.
|
||
|
||
**Verdict:** CRX has fundamental limitations for code sequences. Proceed to
|
||
cluster-then-infer experiment.
|
||
|
||
---
|
||
|
||
## Round 10: Cluster-Then-Infer (CRX Improvement)
|
||
|
||
**Hypothesis:** Grouping similar sequences before CRX inference produces tighter
|
||
grammars, because each cluster's Hasse diagram is more likely to be linear.
|
||
|
||
**Method:**
|
||
1. Cluster sequences by (first_symbol, last_symbol) — a simple structural hash
|
||
2. Infer CRX per cluster
|
||
3. Pick the most common cluster's grammar
|
||
4. Compare: avg max disjunction size (standard CRX vs clustered)
|
||
|
||
**Result:**
|
||
| Metric | Standard CRX | Cluster-Then-Infer | Improvement |
|
||
|--------|-------------|-------------------|-------------|
|
||
| Avg max disjunction size | 2.8 | 1.7 | 39% tighter |
|
||
| Packages improved | — | 6/10 | 60% |
|
||
|
||
Example improvements:
|
||
- `service/job` (14 seqs): `(any+asJobId+assertEquals+assertTrue+build+exchange...)+?`
|
||
→ `get` (single symbol — much tighter)
|
||
- `agent/rag/embabel` (7 seqs): `(any+assertEquals+assertTrue+contains+emptyList+every+verify)+?`
|
||
→ `(any+every)+.emptyList+.assertEquals+` (structured)
|
||
- `batch/listener` (5 seqs): `(any+asJobId+assertEquals+assertTrue)+?.uri?.exchange?.expectStatus?`
|
||
→ `uri.build+.exchange.expectStatus.get` (structured)
|
||
|
||
**Decision:** Add `crx_refined` module with cluster-then-infer as the default
|
||
CRX method. Keep standard CRX available for comparison.
|
||
|
||
**Tradeoff parameter identified:** Cluster granularity.
|
||
- Too coarse (no clustering): over-approximation (current CRX)
|
||
- Too fine (1 seq per cluster): every sequence gets its own grammar, no generalization
|
||
- Sweet spot: cluster by structural features (first/last symbols, length, etc.)
|
||
|
||
**Next steps:**
|
||
- Test clustering on Flask, Coroutines, FastAPI
|
||
- Try better clustering features (k-mer, edit distance, prefix sharing)
|
||
- Evaluate: does tighter grammar → better code completion / convention docs?
|
||
|
||
---
|
||
|
||
## Round 11: Pipeline Speed + GBNF Conversion (commits `bc7d3b6`, `3468813`)
|
||
|
||
**Hypothesis:** iDRegEx in the ensemble is the bottleneck. GBNF conversion needs error handling.
|
||
|
||
**Method:**
|
||
- Made iDRegEx opt-in via `--idregex` flag (was running on every group)
|
||
- Added `validate_sore()` — skip malformed SOREs gracefully
|
||
- Fixed OverflowError: `lang_size_score` produces huge ints for large disjunctions
|
||
- Fixed GBNF tokenizer: strip newlines from literals
|
||
|
||
**Results — Pipeline Speed:**
|
||
| Codebase | Before (with iDRegEx) | After (CRX only) | Speedup |
|
||
|----------|----------------------|-------------------|---------|
|
||
| Flask | 55s+ | 2.8s | 20× |
|
||
| RAGSAK | 74s | 13s | 5.7× |
|
||
| FastAPI | hung at 300s | 30s | >10× |
|
||
|
||
Root cause: `src/flask/json` (50 methods) alone took 55s in iDRegEx. Flask's `tests` group (962 methods) would have been worse.
|
||
|
||
**Results — GBNF Conversion:**
|
||
| Codebase | Grammars | GBNF OK | GBNF FAIL | Malformed (skipped) |
|
||
|----------|----------|---------|-----------|---------------------|
|
||
| Flask | 5 | 5 | 0 | 0 |
|
||
| RAGSAK | 19 | 19 | 0 | 11 |
|
||
| FastAPI | 106 | 106 | 0 | 6 |
|
||
| **Total**| **130** | **130** | **0** | **17** |
|
||
|
||
17 malformed SOREs contain raw code (e.g. `w_body=`, `(+,+:N+...)`) — preprocessing bug, not parser issue.
|
||
|
||
**Grammar Quality Analysis:**
|
||
- **Structured** (has ordering via `.`, `?`): RAGSAK 18, FastAPI 94, Flask 3
|
||
- **Flat disjunction** (bag of symbols): RAGSAK 1, FastAPI 10, Flask 2
|
||
- **Trivial** (single symbol): RAGSAK 0, FastAPI 2, Flask 0
|
||
|
||
Best structured examples:
|
||
- `return.render_template+` — clear: return, then render_template one or more times
|
||
- `assertNull?.parseS3Location.error+?.(assertEquals+bucket)+?.key?` — test flow
|
||
- `buildObservationContext?.shouldRetrieve?.(ASK+ChatResponse+...)?...` — agent flow
|
||
- `if.(img+item_id).(FileResponse+else+media_type+return)+.JSONResponse+?.status_code?.content?` — if/else structure
|
||
|
||
**Decision:** iDRegEx stays opt-in. GBNF validation catches malformed SOREs early.
|
||
The structured grammars (43% of all grammars) capture real calling conventions.
|
||
|
||
**Open questions:**
|
||
- 1129-file FastAPI has 30 diverse groups — need better grouping for large codebases
|
||
- Malformed SOREs from raw code in symbol names — need upstream fix in code.py
|
||
- Flat disjunctions are noisy — should we filter by grammar complexity?
|