Sweet spot: 0.01-0.05. At 0.01 RAGSAK gets 61 SOREs (26.1% cov) with real conventions like warn.status.body.ErrorResponse. At 0.05 coverage jumps to 45.4% but that pattern disappears. Flask dies at 0.15+. Current default min_coverage=0.2 is too aggressive for most codebases.
356 lines
14 KiB
Markdown
356 lines
14 KiB
Markdown
# Experiment Log — Grammar Inference Pipeline
|
||
|
||
Track what we tried, what worked, what failed, and what's next. Each experiment
|
||
includes: hypothesis, method, result, verdict.
|
||
|
||
---
|
||
|
||
## Round 1: Context Strategies (commit `bbdfe93`)
|
||
|
||
**Hypothesis:** The calling context (prefix before the method body) determines
|
||
which methods share a convention. Better context grouping → better grammars.
|
||
|
||
**Method:** Tested 4 strategies on RAGSAK (Kotlin) and Flask (Python):
|
||
- **Baseline:** Group by package directory
|
||
- **Option A:** Group by last k components of file path (`file_path_k{k}`)
|
||
- **Option B:** Group by first k symbols of call sequence (`first_k_sym_{k}`)
|
||
- **Option C:** Two-dimensional: (path_k, first_k_symbols)
|
||
|
||
**Result:**
|
||
| Strategy | RAGSAK patterns | RAGSAK coverage | Flask patterns | Flask coverage |
|
||
|----------|----------------|-----------------|----------------|---------------|
|
||
| Package baseline | 73 | 12.0% | 21 | 10.7% |
|
||
| File path k=1 | 73 | 12.0% | 21 | 10.7% |
|
||
| First k=1 | 20 | 4.6% | 2 | 1.4% |
|
||
| First k=3 | 47 | 12.0% | 21 | 10.7% |
|
||
|
||
**Verdict:** Package grouping and first_k_sym_3 produce similar results.
|
||
Cross-package grouping by first symbol is too sparse — most groups are either
|
||
too large (skipped) or too diverse (skipped). The useful patterns are
|
||
package-specific, not cross-package.
|
||
|
||
---
|
||
|
||
## Round 2: Reduce Algorithm (commit `b516b29`)
|
||
|
||
**Hypothesis:** Reduce (Algorithm 4, TODS 2010) merges structurally similar
|
||
contexts, revealing cross-package patterns by unifying equivalent states.
|
||
|
||
**Method:** Implemented faithful Reduce with support-weighted SOA edit distance,
|
||
adjunction, iterative merging, and minimize. Tested at ε=0.05 to 0.4.
|
||
|
||
**Result:**
|
||
- RAGSAK at ε=0.3: 1 merge (`JobStatus.every.getJobStatus` ↔ `JobStatus.now.minusMinutes`)
|
||
- Flask at ε=0.3: 1 merge (`def.boolean` ↔ `def.is_boolean`)
|
||
- Coverage improvement: negligible (< 1%)
|
||
|
||
**Verdict:** Reduce doesn't help. The contexts we produce are already too
|
||
specific (unique per package) for the distance metric to find meaningful merges.
|
||
Reduce works when you have a large SOA with many equivalent states — we have
|
||
one SOA per package with few states. Wrong abstraction level.
|
||
|
||
**Why it failed:** Reduce merges states in a single automaton. We're producing
|
||
one automaton per package group. There's nothing to merge across packages
|
||
because each package gets its own inference run. Reduce would need to operate
|
||
on a cross-package SOA, which we don't build.
|
||
|
||
---
|
||
|
||
## Round 3: Language Size Scoring (commit `dfb56a0`)
|
||
|
||
**Hypothesis:** Bex et al.'s Language Size measure (arXiv:1004.2372, Section
|
||
4.3.1) is a better scoring function than MDL for our use case.
|
||
|
||
**Method:** Implemented `lang_size_score()` as default scoring method. Added
|
||
diversity threshold: skip groups with unique_ratio > 0.9 or methods < 5.
|
||
|
||
**Result:** 39 new tests. Pipeline runs correctly with new scoring. Coverage
|
||
numbers similar to before — scoring method doesn't change which patterns are
|
||
found, just which grammar is selected per group.
|
||
|
||
**Verdict:** Scoring method is not the bottleneck. The problem is upstream
|
||
(pattern extraction), not downstream (pattern selection).
|
||
|
||
---
|
||
|
||
## Round 4: GBNF Output (commit `011df39`)
|
||
|
||
**Hypothesis:** SORE → GBNF conversion enables constrained LLM generation.
|
||
SORE operators map directly to GBNF syntax.
|
||
|
||
**Method:** Implemented recursive descent parser for SORE, AST intermediate
|
||
representation, and GBNF renderer. 15 tests.
|
||
|
||
**Result:** All tests pass. `to_gbnf('raise.(ValueError)+')` →
|
||
`"raise" "ValueError"+`. Correct mapping of +, ?, *, |, ., parens.
|
||
|
||
**Verdict:** Implementation works. But the input SOREs are too specific
|
||
to individual packages to be useful for constrained generation. The converter
|
||
is correct; the patterns it converts are the problem.
|
||
|
||
---
|
||
|
||
## Round 5: Cross-Package Exact Matches
|
||
|
||
**Hypothesis:** Some call sequences appear verbatim in multiple packages.
|
||
These are the real cross-package conventions.
|
||
|
||
**Method:** Grouped all sequences by exact tuple match across packages.
|
||
|
||
**Result:** 38 exact cross-package sequences in RAGSAK. Most are trivial:
|
||
- `('clearAllMocks',)` — 4 packages (test teardown)
|
||
- `('Builder',)` — 4 packages (builder pattern)
|
||
- `('Any',)` — 4 packages (Kotlin type)
|
||
- `('get',)` — 4 packages (getter)
|
||
|
||
Interesting ones:
|
||
- `('assumeTrue', 'isDockerAvailable', 'start', 'pullAndWarmup')` — 4 packages (Docker test setup)
|
||
- `('isNullOrBlank', 'error', 'error')` — 3 packages (null check → error)
|
||
- `('sortedBy', 'map', 'toDescriptor')` — 3 packages (data pipeline)
|
||
- `('ObjectMapper', 'findAndRegisterModules')` — 2 packages (Jackson config)
|
||
|
||
**Verdict:** Exact matches are too rare and mostly trivial. The real
|
||
cross-package patterns are structural, not textual — "null check → error"
|
||
appears with different method names in different packages.
|
||
|
||
---
|
||
|
||
## Round 6: Structural Coarsening (Experiment `coarsen_eval.py`)
|
||
|
||
**Hypothesis:** Coarsening structural tokens (keywords, types) while keeping
|
||
function names raw reveals cross-package patterns that pure-text misses.
|
||
|
||
**Method:** `coarsen_token()` maps tree-sitter capture names to categories.
|
||
Function calls kept as raw text (they ARE the content). Only structural tokens
|
||
coarsened: RETURN, RAISE, IF, LOOP, EXCEPTION, KW, TYPE, FUN, ATTR.
|
||
|
||
**Result:**
|
||
|
||
| Codebase | Structural % | Raw coverage | Coarsened coverage | Δ |
|
||
|----------|-------------|--------------|-------------------|---|
|
||
| RAGSAK | 4.7% | 18.6% | 18.3% | -0.3% |
|
||
| Flask | 36.1% | 17.5% | **28.8%** | **+11.3%** |
|
||
|
||
Key findings:
|
||
- RAGSAK: 95.3% function calls → coarsening has nothing to work with
|
||
- Flask: 36.1% structural → coarsening significantly improves grouping
|
||
- Flask k=3: cross-package contexts increase 40 → 47
|
||
- Coarsened cross-package shapes: `('IF', 'KW', 'RETURN')` in 4 packages,
|
||
`('RETURN', 'render_template', 'render_template')` in 6 packages
|
||
|
||
**Verdict:** Coarsening helps codebases with rich structural tokens (Python:
|
||
IF, LOOP, EXCEPTION, KW). Doesn't help codebases dominated by function calls
|
||
(Kotlin: 95% calls). The approach is sound but language-dependent in practice —
|
||
depends on how rich the highlights.scm is.
|
||
|
||
**What we learned:**
|
||
- The 5% structural tokens DO carry signal when they exist
|
||
- Python highlights.scm is much richer than Kotlin's
|
||
- Coarsening is not dead — it's a tool for languages with rich captures
|
||
- The real question is whether the coarsened patterns are USEFUL, not just
|
||
whether they exist
|
||
|
||
---
|
||
|
||
## Round 6b: Kotlin Capture Fix + Minimal Coarsening
|
||
|
||
**Hypothesis:** Kotlin's highlights.scm uses bare captures (`conditional`,
|
||
`exception`, `repeat`) while `BEHAVIORAL_PREFIXES` expected dotted forms
|
||
(`keyword.conditional`, etc.). All Kotlin structural context was being dropped.
|
||
|
||
**Method:** Added bare captures to BEHAVIORAL_PREFIXES. Dropped ATTR, TYPE,
|
||
VAR from coarsening (too noisy). Only coarsen RETURN, IF, EXCEPTION, LOOP.
|
||
|
||
**Result:**
|
||
|
||
| Codebase | Raw coverage | Coarsened coverage | Cross-pkg contexts |
|
||
|----------|-------------|-------------------|-------------------|
|
||
| RAGSAK k=3 | 18.5% | 6.0% | 87 (was 69) |
|
||
| Flask k=2 | 4.5% | 15.8% | 35 (was 54) |
|
||
| Flask k=3 | 17.1% | 17.3% | 28 (was 40) |
|
||
|
||
Key cross-package patterns (coarsened):
|
||
- `('IF', 'isEmpty', 'isEmpty')` — 8 RAGSAK packages (null-check convention)
|
||
- `('IF', 'isNullOrBlank', 'isNullOrBlank')` — 6 RAGSAK packages
|
||
- `('RETURN', 'render_template', 'render_template')` — 5 Flask packages
|
||
- `('IF', 'RETURN')` — 8 RAGSAK packages (guard clause pattern)
|
||
|
||
**Verdict:** The 4 high-signal categories (RETURN, IF, EXCEPTION, LOOP) DO
|
||
reveal cross-package conventions. Coarsening trades per-package coverage for
|
||
cross-package reach. Whether this is useful depends on the use case:
|
||
- For code completion: raw is better (specific method names)
|
||
- For documentation: coarsened is better (structural conventions)
|
||
|
||
---
|
||
|
||
## Round 7: Four-Codebase Evaluation
|
||
|
||
**Method:** Run raw vs coarsened on RAGSAK (Kotlin), Flask (Python),
|
||
kotlinx.coroutines (Kotlin), FastAPI (Python).
|
||
|
||
**Results (k=3):**
|
||
|
||
| Codebase | Raw cov | Coarse cov | Raw xpkg | Coarse xpkg | Δ |
|
||
|----------|---------|------------|----------|-------------|---|
|
||
| RAGSAK | 18.5% | 6.0% | 69 | 87 | +18 |
|
||
| Flask | 17.1% | 17.3% | 40 | 27 | -13 |
|
||
| Coroutines | 30.7% | 19.1% | 524 | 539 | +15 |
|
||
| FastAPI | 12.4% | 20.7% | 82 | 97 | +15 |
|
||
|
||
Cross-package patterns discovered:
|
||
- FastAPI: `('response', 'client', 'get')` — 72 packages (HTTP request pattern)
|
||
- Coroutines: `('RETURN', 'EXCEPTION', 'UnsupportedOperationException')` — 13 packages
|
||
- RAGSAK: `('IF', 'isEmpty', 'isEmpty')` — 8 packages (null-check)
|
||
- Flask: `('RETURN', 'render_template', 'render_template')` — 5 packages
|
||
|
||
**Verdict:** The conventions vs completions tradeoff is real and measurable.
|
||
Coarsening consistently trades per-package coverage for cross-package reach.
|
||
FastAPI is the exception: coverage improves (12.4% → 20.7%) because its
|
||
structural tokens (response/client patterns) are highly repetitive.
|
||
|
||
---
|
||
|
||
## Round 8: Frequency Threshold Sweep
|
||
|
||
**Hypothesis:** The fixed `min_coverage=0.2` is too aggressive. Lower thresholds
|
||
reveal more patterns while still filtering noise.
|
||
|
||
**Method:** Test thresholds 0.00–0.20 on all 4 codebases. Measure symbol count,
|
||
surviving sequences, SORE success, coverage.
|
||
|
||
**Results:**
|
||
|
||
| Codebase | Thresh | Syms | Seqs | SOREs | Coverage |
|
||
|----------|--------|------|------|-------|----------|
|
||
| RAGSAK | 0.01 | 130 | 1270 | 61 | 26.1% |
|
||
| RAGSAK | 0.05 | 16 | 919 | 40 | 45.4% |
|
||
| RAGSAK | 0.10 | 5 | 657 | 19 | 69.1% |
|
||
| Flask | 0.01 | 47 | 784 | 28 | 35.2% |
|
||
| Flask | 0.05 | 5 | 414 | 7 | 44.4% |
|
||
| Flask | 0.10 | 2 | 237 | 5 | 100.0% |
|
||
| Coroutines | 0.01 | 76 | 4802 | 175 | 40.4% |
|
||
| FastAPI | 0.01 | 23 | 2678 | 14 | 17.0% |
|
||
|
||
Key findings:
|
||
- Coverage increases with threshold (trivial: 1 symbol = 100% coverage)
|
||
- Sweet spot: 0.01–0.05. Enough symbols for meaningful patterns, enough
|
||
filtering to remove noise.
|
||
- At 0.01: RAGSAK gets `warn.status.body.ErrorResponse` (real convention)
|
||
- At 0.05: that pattern disappears (too aggressive)
|
||
- Flask dies at 0.15+ (0 symbols survive)
|
||
|
||
---
|
||
|
||
## What We Learned (Summary)
|
||
|
||
1. **Per-package grouping is too sparse.** 1-3 sequences per package isn't
|
||
enough for any inference method to produce general patterns.
|
||
|
||
2. **Cross-package exact matches are rare.** Only 38 in RAGSAK, mostly trivial
|
||
single-call sequences.
|
||
|
||
3. **Reduce doesn't help at our abstraction level.** It merges states within
|
||
one automaton; we need to merge patterns across packages.
|
||
|
||
4. **Scoring/selection isn't the bottleneck.** MDL vs Language Size doesn't
|
||
change what patterns are found.
|
||
|
||
5. **The calling context prefix is the right signal** but grouping by it
|
||
produces groups that are either too large, too diverse, or trivial.
|
||
|
||
6. **GBNF converter works correctly** but the input patterns are too specific.
|
||
|
||
---
|
||
|
||
## Next: Structural Coarsening + Cross-Package Detection
|
||
|
||
### Idea
|
||
|
||
Collapse method names → categories using tree-sitter capture names. This
|
||
converts textual sequences into structural shapes:
|
||
|
||
```
|
||
('isNullOrBlank', 'error', 'error') → (CALL, ERROR, ERROR)
|
||
('raise', 'ValueError', 'ValueError') → (CALL, ERROR, ERROR)
|
||
```
|
||
|
||
Same structural shape, different packages → cross-package convention.
|
||
|
||
### Why This Might Work
|
||
|
||
- We already extract tree-sitter capture names in `code.py:56-69`
|
||
- We already classify nodes into categories in `code.py:72-100`
|
||
(`lit`, `call`, `var`, `lambda`, `kwarg`, `expr`, `template`, `other`)
|
||
- The behavioral prefix filter (`CALL_PREFIXES`) keeps raw text; we need a
|
||
parallel path that keeps the category instead
|
||
- Coarsened sequences have smaller alphabets → more methods per group →
|
||
better inference
|
||
- Patterns like `(CALL, ERROR, ERROR)` are meaningful conventions that
|
||
repeat across packages
|
||
|
||
### What We Need
|
||
|
||
1. **Coarsening map:** `capture_name → category` using the existing
|
||
`CALL_PREFIXES`, `ARG_LITERAL_TYPES`, `LAMBDA_TYPES` classifications
|
||
plus a new `ERROR_TYPES` set
|
||
|
||
2. **Coarsened sequence extraction:** Same pipeline as now, but output
|
||
category labels instead of method names
|
||
|
||
3. **Cross-package grouping:** Group by coarsened shape (first k categories),
|
||
find shapes that appear in ≥2 packages
|
||
|
||
4. **SORE inference on coarsened sequences:** Smaller alphabet, more examples
|
||
per group → better patterns
|
||
|
||
5. **Evaluation:** Compare coarsened patterns vs raw patterns on:
|
||
- Coverage (% of methods in learned groups)
|
||
- Cross-package reach (# of packages per pattern)
|
||
- Usefulness for constrained generation (GBNF quality)
|
||
|
||
### Open Questions
|
||
|
||
- Does coarsening lose too much specificity? `(CALL, ERROR, ERROR)` is less
|
||
informative than `(raise, ValueError, ValueError)` — is the tradeoff worth it?
|
||
- What categories to use? The existing classifications in `code.py` are a
|
||
starting point but may need refinement (e.g., separating ERROR from CALL)
|
||
- How to handle the long tail? Most sequences are 1-2 symbols — coarsening
|
||
doesn't help much for those
|
||
- Is the GBNF output useful at all? Maybe the output should be a conditional
|
||
frequency table instead of a grammar
|
||
|
||
### Experiment Design
|
||
|
||
**Phase 1: Coarsening Proof of Concept**
|
||
- Implement coarsening map in `code.py`
|
||
- Add `--coarsen` flag to CLI
|
||
- Run on RAGSAK + Flask, compare raw vs coarsened patterns
|
||
- Measure: alphabet size reduction, group size increase, pattern count
|
||
|
||
**Phase 2: Cross-Package Detection**
|
||
- Group coarsened sequences by shape (first k categories)
|
||
- Find shapes appearing in ≥2 packages
|
||
- For each shape, infer SORE on coarsened sequences
|
||
- Measure: cross-package patterns found, coverage improvement
|
||
|
||
**Phase 3: Output Quality**
|
||
- Convert coarsened SOREs to GBNF
|
||
- Evaluate: are the GBNF rules more general/useful than raw SOREs?
|
||
- Compare: coarsened GBNF vs raw GBNF vs conditional frequency table
|
||
|
||
---
|
||
|
||
## The "Redacted" Concept
|
||
|
||
When collapsing method names to categories, we lose the specific method name
|
||
but gain the structural pattern. This is a form of **abstraction** — moving
|
||
from concrete examples to general rules.
|
||
|
||
The question is whether the abstraction is at the right level:
|
||
- Too specific: `(raise, ValueError, ValueError)` — package-specific noise
|
||
- Right level: `(CALL, ERROR, ERROR)` — cross-package convention
|
||
- Too abstract: `(X, Y, Y)` — trivial, tells the LLM nothing
|
||
|
||
The middle ground depends on the category vocabulary. We need enough categories
|
||
to be informative (CALL, ERROR, LIT, VAR, TYPE, etc.) but not so many that
|
||
every sequence is unique.
|