Added kotlinx.coroutines (1040 .kt files) and FastAPI (1129 .py files). Key findings across 4 codebases: - Coarsening trades per-package coverage for cross-package reach - FastAPI: coverage IMPROVES (12.4% -> 20.7%), cross-pkg +15 - Coroutines: cross-pkg +15, real exception-handling patterns - RAGSAK: cross-pkg +18, null-check conventions across 8 packages - Flask: slight cross-pkg loss, rendering conventions still visible The conventions vs completions tradeoff is real and measurable.
13 KiB
Experiment Log — Grammar Inference Pipeline
Track what we tried, what worked, what failed, and what's next. Each experiment includes: hypothesis, method, result, verdict.
Round 1: Context Strategies (commit bbdfe93)
Hypothesis: The calling context (prefix before the method body) determines which methods share a convention. Better context grouping → better grammars.
Method: Tested 4 strategies on RAGSAK (Kotlin) and Flask (Python):
- Baseline: Group by package directory
- Option A: Group by last k components of file path (
file_path_k{k}) - Option B: Group by first k symbols of call sequence (
first_k_sym_{k}) - Option C: Two-dimensional: (path_k, first_k_symbols)
Result:
| Strategy | RAGSAK patterns | RAGSAK coverage | Flask patterns | Flask coverage |
|---|---|---|---|---|
| Package baseline | 73 | 12.0% | 21 | 10.7% |
| File path k=1 | 73 | 12.0% | 21 | 10.7% |
| First k=1 | 20 | 4.6% | 2 | 1.4% |
| First k=3 | 47 | 12.0% | 21 | 10.7% |
Verdict: Package grouping and first_k_sym_3 produce similar results. Cross-package grouping by first symbol is too sparse — most groups are either too large (skipped) or too diverse (skipped). The useful patterns are package-specific, not cross-package.
Round 2: Reduce Algorithm (commit b516b29)
Hypothesis: Reduce (Algorithm 4, TODS 2010) merges structurally similar contexts, revealing cross-package patterns by unifying equivalent states.
Method: Implemented faithful Reduce with support-weighted SOA edit distance, adjunction, iterative merging, and minimize. Tested at ε=0.05 to 0.4.
Result:
- RAGSAK at ε=0.3: 1 merge (
JobStatus.every.getJobStatus↔JobStatus.now.minusMinutes) - Flask at ε=0.3: 1 merge (
def.boolean↔def.is_boolean) - Coverage improvement: negligible (< 1%)
Verdict: Reduce doesn't help. The contexts we produce are already too specific (unique per package) for the distance metric to find meaningful merges. Reduce works when you have a large SOA with many equivalent states — we have one SOA per package with few states. Wrong abstraction level.
Why it failed: Reduce merges states in a single automaton. We're producing one automaton per package group. There's nothing to merge across packages because each package gets its own inference run. Reduce would need to operate on a cross-package SOA, which we don't build.
Round 3: Language Size Scoring (commit dfb56a0)
Hypothesis: Bex et al.'s Language Size measure (arXiv:1004.2372, Section 4.3.1) is a better scoring function than MDL for our use case.
Method: Implemented lang_size_score() as default scoring method. Added
diversity threshold: skip groups with unique_ratio > 0.9 or methods < 5.
Result: 39 new tests. Pipeline runs correctly with new scoring. Coverage numbers similar to before — scoring method doesn't change which patterns are found, just which grammar is selected per group.
Verdict: Scoring method is not the bottleneck. The problem is upstream (pattern extraction), not downstream (pattern selection).
Round 4: GBNF Output (commit 011df39)
Hypothesis: SORE → GBNF conversion enables constrained LLM generation. SORE operators map directly to GBNF syntax.
Method: Implemented recursive descent parser for SORE, AST intermediate representation, and GBNF renderer. 15 tests.
Result: All tests pass. to_gbnf('raise.(ValueError)+') →
"raise" "ValueError"+. Correct mapping of +, ?, *, |, ., parens.
Verdict: Implementation works. But the input SOREs are too specific to individual packages to be useful for constrained generation. The converter is correct; the patterns it converts are the problem.
Round 5: Cross-Package Exact Matches
Hypothesis: Some call sequences appear verbatim in multiple packages. These are the real cross-package conventions.
Method: Grouped all sequences by exact tuple match across packages.
Result: 38 exact cross-package sequences in RAGSAK. Most are trivial:
('clearAllMocks',)— 4 packages (test teardown)('Builder',)— 4 packages (builder pattern)('Any',)— 4 packages (Kotlin type)('get',)— 4 packages (getter)
Interesting ones:
('assumeTrue', 'isDockerAvailable', 'start', 'pullAndWarmup')— 4 packages (Docker test setup)('isNullOrBlank', 'error', 'error')— 3 packages (null check → error)('sortedBy', 'map', 'toDescriptor')— 3 packages (data pipeline)('ObjectMapper', 'findAndRegisterModules')— 2 packages (Jackson config)
Verdict: Exact matches are too rare and mostly trivial. The real cross-package patterns are structural, not textual — "null check → error" appears with different method names in different packages.
Round 6: Structural Coarsening (Experiment coarsen_eval.py)
Hypothesis: Coarsening structural tokens (keywords, types) while keeping function names raw reveals cross-package patterns that pure-text misses.
Method: coarsen_token() maps tree-sitter capture names to categories.
Function calls kept as raw text (they ARE the content). Only structural tokens
coarsened: RETURN, RAISE, IF, LOOP, EXCEPTION, KW, TYPE, FUN, ATTR.
Result:
| Codebase | Structural % | Raw coverage | Coarsened coverage | Δ |
|---|---|---|---|---|
| RAGSAK | 4.7% | 18.6% | 18.3% | -0.3% |
| Flask | 36.1% | 17.5% | 28.8% | +11.3% |
Key findings:
- RAGSAK: 95.3% function calls → coarsening has nothing to work with
- Flask: 36.1% structural → coarsening significantly improves grouping
- Flask k=3: cross-package contexts increase 40 → 47
- Coarsened cross-package shapes:
('IF', 'KW', 'RETURN')in 4 packages,('RETURN', 'render_template', 'render_template')in 6 packages
Verdict: Coarsening helps codebases with rich structural tokens (Python: IF, LOOP, EXCEPTION, KW). Doesn't help codebases dominated by function calls (Kotlin: 95% calls). The approach is sound but language-dependent in practice — depends on how rich the highlights.scm is.
What we learned:
- The 5% structural tokens DO carry signal when they exist
- Python highlights.scm is much richer than Kotlin's
- Coarsening is not dead — it's a tool for languages with rich captures
- The real question is whether the coarsened patterns are USEFUL, not just whether they exist
Round 6b: Kotlin Capture Fix + Minimal Coarsening
Hypothesis: Kotlin's highlights.scm uses bare captures (conditional,
exception, repeat) while BEHAVIORAL_PREFIXES expected dotted forms
(keyword.conditional, etc.). All Kotlin structural context was being dropped.
Method: Added bare captures to BEHAVIORAL_PREFIXES. Dropped ATTR, TYPE, VAR from coarsening (too noisy). Only coarsen RETURN, IF, EXCEPTION, LOOP.
Result:
| Codebase | Raw coverage | Coarsened coverage | Cross-pkg contexts |
|---|---|---|---|
| RAGSAK k=3 | 18.5% | 6.0% | 87 (was 69) |
| Flask k=2 | 4.5% | 15.8% | 35 (was 54) |
| Flask k=3 | 17.1% | 17.3% | 28 (was 40) |
Key cross-package patterns (coarsened):
('IF', 'isEmpty', 'isEmpty')— 8 RAGSAK packages (null-check convention)('IF', 'isNullOrBlank', 'isNullOrBlank')— 6 RAGSAK packages('RETURN', 'render_template', 'render_template')— 5 Flask packages('IF', 'RETURN')— 8 RAGSAK packages (guard clause pattern)
Verdict: The 4 high-signal categories (RETURN, IF, EXCEPTION, LOOP) DO reveal cross-package conventions. Coarsening trades per-package coverage for cross-package reach. Whether this is useful depends on the use case:
- For code completion: raw is better (specific method names)
- For documentation: coarsened is better (structural conventions)
Round 7: Four-Codebase Evaluation
Method: Run raw vs coarsened on RAGSAK (Kotlin), Flask (Python), kotlinx.coroutines (Kotlin), FastAPI (Python).
Results (k=3):
| Codebase | Raw cov | Coarse cov | Raw xpkg | Coarse xpkg | Δ |
|---|---|---|---|---|---|
| RAGSAK | 18.5% | 6.0% | 69 | 87 | +18 |
| Flask | 17.1% | 17.3% | 40 | 27 | -13 |
| Coroutines | 30.7% | 19.1% | 524 | 539 | +15 |
| FastAPI | 12.4% | 20.7% | 82 | 97 | +15 |
Cross-package patterns discovered:
- FastAPI:
('response', 'client', 'get')— 72 packages (HTTP request pattern) - Coroutines:
('RETURN', 'EXCEPTION', 'UnsupportedOperationException')— 13 packages - RAGSAK:
('IF', 'isEmpty', 'isEmpty')— 8 packages (null-check) - Flask:
('RETURN', 'render_template', 'render_template')— 5 packages
Verdict: The conventions vs completions tradeoff is real and measurable. Coarsening consistently trades per-package coverage for cross-package reach. FastAPI is the exception: coverage improves (12.4% → 20.7%) because its structural tokens (response/client patterns) are highly repetitive.
What We Learned (Summary)
-
Per-package grouping is too sparse. 1-3 sequences per package isn't enough for any inference method to produce general patterns.
-
Cross-package exact matches are rare. Only 38 in RAGSAK, mostly trivial single-call sequences.
-
Reduce doesn't help at our abstraction level. It merges states within one automaton; we need to merge patterns across packages.
-
Scoring/selection isn't the bottleneck. MDL vs Language Size doesn't change what patterns are found.
-
The calling context prefix is the right signal but grouping by it produces groups that are either too large, too diverse, or trivial.
-
GBNF converter works correctly but the input patterns are too specific.
Next: Structural Coarsening + Cross-Package Detection
Idea
Collapse method names → categories using tree-sitter capture names. This converts textual sequences into structural shapes:
('isNullOrBlank', 'error', 'error') → (CALL, ERROR, ERROR)
('raise', 'ValueError', 'ValueError') → (CALL, ERROR, ERROR)
Same structural shape, different packages → cross-package convention.
Why This Might Work
- We already extract tree-sitter capture names in
code.py:56-69 - We already classify nodes into categories in
code.py:72-100(lit,call,var,lambda,kwarg,expr,template,other) - The behavioral prefix filter (
CALL_PREFIXES) keeps raw text; we need a parallel path that keeps the category instead - Coarsened sequences have smaller alphabets → more methods per group → better inference
- Patterns like
(CALL, ERROR, ERROR)are meaningful conventions that repeat across packages
What We Need
-
Coarsening map:
capture_name → categoryusing the existingCALL_PREFIXES,ARG_LITERAL_TYPES,LAMBDA_TYPESclassifications plus a newERROR_TYPESset -
Coarsened sequence extraction: Same pipeline as now, but output category labels instead of method names
-
Cross-package grouping: Group by coarsened shape (first k categories), find shapes that appear in ≥2 packages
-
SORE inference on coarsened sequences: Smaller alphabet, more examples per group → better patterns
-
Evaluation: Compare coarsened patterns vs raw patterns on:
- Coverage (% of methods in learned groups)
- Cross-package reach (# of packages per pattern)
- Usefulness for constrained generation (GBNF quality)
Open Questions
- Does coarsening lose too much specificity?
(CALL, ERROR, ERROR)is less informative than(raise, ValueError, ValueError)— is the tradeoff worth it? - What categories to use? The existing classifications in
code.pyare a starting point but may need refinement (e.g., separating ERROR from CALL) - How to handle the long tail? Most sequences are 1-2 symbols — coarsening doesn't help much for those
- Is the GBNF output useful at all? Maybe the output should be a conditional frequency table instead of a grammar
Experiment Design
Phase 1: Coarsening Proof of Concept
- Implement coarsening map in
code.py - Add
--coarsenflag to CLI - Run on RAGSAK + Flask, compare raw vs coarsened patterns
- Measure: alphabet size reduction, group size increase, pattern count
Phase 2: Cross-Package Detection
- Group coarsened sequences by shape (first k categories)
- Find shapes appearing in ≥2 packages
- For each shape, infer SORE on coarsened sequences
- Measure: cross-package patterns found, coverage improvement
Phase 3: Output Quality
- Convert coarsened SOREs to GBNF
- Evaluate: are the GBNF rules more general/useful than raw SOREs?
- Compare: coarsened GBNF vs raw GBNF vs conditional frequency table
The "Redacted" Concept
When collapsing method names to categories, we lose the specific method name but gain the structural pattern. This is a form of abstraction — moving from concrete examples to general rules.
The question is whether the abstraction is at the right level:
- Too specific:
(raise, ValueError, ValueError)— package-specific noise - Right level:
(CALL, ERROR, ERROR)— cross-package convention - Too abstract:
(X, Y, Y)— trivial, tells the LLM nothing
The middle ground depends on the category vocabulary. We need enough categories to be informative (CALL, ERROR, LIT, VAR, TYPE, etc.) but not so many that every sequence is unique.