9.8 KiB
Experiment Results — Context Strategies + Reduce
Date: 2026-07-12
Codebases
| Codebase | Files | Methods | Language |
|---|---|---|---|
| RAGSAK | 462 | 1594 | Kotlin |
| Flask | 24 | 1391 | Python |
RAGSAK Results
| Strategy | Contexts | SOREs | Coverage | Notes |
|---|---|---|---|---|
| Baseline (package) | 116 | 2 | 0.6% | Structural — useless |
| File path k=1 | 43 | 2 | 1.1% | Structural — no better |
| File path k=2 | 54 | 2 | 1.1% | Same |
| File path k=3 | 54 | 2 | 1.1% | Same |
| First 1 symbol | 134 | 20 | 4.6% | Behavioral — 10× better |
| First 2 symbols | 141 | 39 | 9.8% | Behavioral — 20× better |
| First 3 symbols | 117 | 47 | 12.0% | Winner |
| Two-dim (p1,s1) | 151 | 24 | 6.3% | Hybrid — worse |
| Two-dim (p1,s2) | 119 | 33 | 8.7% | Hybrid — worse |
| Two-dim (p2,s1) | 146 | 24 | 6.3% | Hybrid — worse |
| Two-dim (p2,s2) | 117 | 33 | 8.7% | Hybrid — worse |
| Return type heuristic | 5 | 0 | 0.0% | Useless |
Reduce Results (RAGSAK)
| Base | Threshold | Merges | Coverage | Notes |
|---|---|---|---|---|
| k=1 | 0.05-0.4 | 0 | 4.6% | No merges — contexts too distinct |
| k=2 | 0.05-0.2 | 0 | 9.8% | No merges |
| k=2 | 0.3 | 1 | 9.6% | Merged (JobStatus, every) with (JobStatus, now) |
| k=2 | 0.4 | 2 | 9.5% | Merged 2 pairs |
| k=3 | 0.05-0.2 | 0 | 12.0% | No merges |
| k=3 | 0.3 | 1 | 11.7% | Merged (JobStatus, every, getJobStatus) with (JobStatus, now, minusMinutes) |
| k=3 | 0.4 | 1 | 11.7% | Same merge |
Flask Results
| Strategy | Contexts | SOREs | Coverage | Notes |
|---|---|---|---|---|
| Baseline (package) | 9 | 0 | 0.0% | Structural — useless |
| File path k=1 | 7 | 0 | 0.0% | Same |
| File path k=2 | 9 | 0 | 0.0% | Same |
| File path k=3 | 9 | 0 | 0.0% | Same |
| First 1 symbol | 53 | 2 | 1.4% | Behavioral — only 2 SOREs |
| First 2 symbols | 96 | 18 | 7.3% | Behavioral — 15× better |
| First 3 symbols | 85 | 21 | 10.7% | Winner |
| Two-dim (p1,s1) | 68 | 5 | 2.2% | Hybrid — worse |
| Two-dim (p1,s2) | 96 | 22 | 10.1% | Close to behavioral |
| Two-dim (p2,s1) | 69 | 5 | 2.2% | Worse |
| Two-dim (p2,s2) | 93 | 19 | 9.5% | Close to behavioral |
| Return type heuristic | 3 | 0 | 0.0% | Useless |
Reduce Results (Flask)
| Base | Threshold | Merges | Coverage | Notes |
|---|---|---|---|---|
| k=1 | 0.05-0.4 | 0 | 1.4% | No merges |
| k=2 | 0.05-0.2 | 0 | 7.3% | No merges |
| k=2 | 0.3-0.4 | 1 | 7.3% | Merged (def, boolean) with (def, is_boolean) |
| k=3 | 0.05-0.2 | 0 | 10.7% | No merges |
| k=3 | 0.3-0.4 | 1 | 10.7% | Merged (def, boolean, return) with (def, is_boolean, return) |
Cross-Codebase Comparison
| Metric | RAGSAK | Flask |
|---|---|---|
| Best strategy | First 3 symbols | First 3 symbols |
| Best coverage | 12.0% | 10.7% |
| SOREs (best) | 47 | 21 |
| Reduce merges (ε=0.3) | 1 | 1 |
| Reduce impact on coverage | -0.3% | 0% |
Key Findings
1. Behavioral grouping wins consistently
- Both codebases: First-k-symbols beats all other strategies
- Structural (file path) is useless — doesn't predict behavioral similarity
- Hybrid (2D) is worse than pure behavioral
2. Coverage ceiling is ~10-12%
- RAGSAK: 12.0% with 47 SOREs
- Flask: 10.7% with 21 SOREs
- Most methods have unique call patterns — they don't share prefixes
3. Reduce has minimal impact
- Very few merges at any threshold (0-2 per codebase)
- Support-weighted distance is conservative — requires very similar SOAs
- Merges that do happen don't improve coverage
- Reason: contexts created by first-k-symbols are already quite distinct
4. The few Reduce merges are meaningful
- RAGSAK:
(JobStatus, every, getJobStatus)↔(JobStatus, now, minusMinutes)— same polling pattern, different initial call - Flask:
(def, boolean)↔(def, is_boolean)— same type-checking pattern, different method name
5. Flask is harder than RAGSAK
- Flask has fewer files (24 vs 462) but similar methods (1391 vs 1594)
- Flask has more unique methods per package — less repetition
- Flask SOREs are shorter/simpler — less compressible patterns
Round 5: Cross-Package Exact Matches
Hypothesis: Some call sequences appear verbatim in multiple packages. These are the real cross-package conventions.
Method: Group all sequences by exact tuple match, count packages per sequence.
Result: 38 exact cross-package sequences in RAGSAK. Most trivial:
('clearAllMocks',)— 4 packages (test teardown)('Builder',)— 4 packages (builder pattern)('Any',)— 4 packages (Kotlin type)
Interesting ones:
('assumeTrue', 'isDockerAvailable', 'start', 'pullAndWarmup')— 4 pkgs (Docker setup)('isNullOrBlank', 'error', 'error')— 3 pkgs (null check → error)('sortedBy', 'map', 'toDescriptor')— 3 pkgs (data pipeline)('ObjectMapper', 'findAndRegisterModules')— 2 pkgs (Jackson config)
Verdict: Exact matches too rare and mostly trivial. The real cross-package patterns are structural, not textual — "null check → error" appears with different method names in different packages.
Summary of Failed/Dismissed Approaches
| Approach | Why it failed |
|---|---|
| Per-package inference | Too sparse (1-3 seqs/package) |
| Reduce algorithm | Wrong abstraction level — merges states within one automaton, not across packages |
| Cross-package grouping by first symbol | 4.6% / 1.4% coverage — most groups skipped |
| Exact cross-package matches | 38 sequences, mostly trivial single-call |
| MDL vs Language Size scoring | Scoring isn't the bottleneck — pattern extraction is |
What Actually Works
- Behavioral grouping (first 3 symbols) — 12% / 10.7% coverage, consistent across codebases
- Calling context as prefix — the right signal, but package-specific
- GBNF conversion — correct implementation, but input patterns too specific
Next: Structural Coarsening + Cross-Package Detection
See EXPERIMENT_LOG.md for full reasoning and experiment design.
Core idea: collapse method names → categories using tree-sitter capture names. Converts textual sequences into structural shapes that repeat across packages.
('isNullOrBlank', 'error', 'error') → (CALL, ERROR, ERROR)
('raise', 'ValueError', 'ValueError') → (CALL, ERROR, ERROR)
Files Generated
experiments/results/ragsak_summary.json— RAGSAK metricsexperiments/results/flask_summary.json— Flask metricsexperiments/context_eval.py— experiment runner (supports multiple codebases)experiments/EXPERIMENT_LOG.md— full experiment history and next stepsbex/reduce.py— Algorithm 4 (TODS 2010) implementationbex/gbnf.py— SORE → GBNF convertertests/test_reduce.py— 24 tests for Reducetests/test_gbnf.py— 15 tests for GBNF converter
Round 20: AST Pipeline + Scoring Fixes (2026-07-13)
Commit range: ea6cac5 → 8b3a454
Codebases
| Codebase | Files | Methods | Language |
|---|---|---|---|
| RAGSAK | 462 | 1609 | Kotlin |
| FastAPI | — | — | Python |
| Zod | — | — | TypeScript |
Scoring: Language Size over MDL (ADR-13)
Abandoned MDL scoring — it rewards short expressions, so generic info+ beat specific
a.b.c.d.e+ (21% vs 98% success in Bex paper). Language Size (lang_size_score) chosen.
| Metric | Bag grammar | Structured grammar |
|---|---|---|
lang_size_score |
9975 | 20 |
mdl_score |
10^12 (clamped) | 10^12 (clamped) |
Final Defaults
| Parameter | Before | After |
|---|---|---|
decompose |
False | True |
max_seq_length |
5 | 4 |
idregex_refine |
False | True |
_COUNT_CAP |
10^12 | 10^30 |
Results
| Codebase | Grammars | Pure Bags | Structured | Bag % |
|---|---|---|---|---|
| RAGSAK (v4) | 126 | 6 | 120 | 4.8% |
| FastAPI (v3) | 143 | 26 | 117 | 18.2% |
| Zod (v3) | 23 | 5 | 18 | 21.7% |
Quality breakdown: ~85% of grammars across codebases remain orderless bags (A|B|C)+.
The ~15% that are structured represent real sequential flows:
- Web controller tests:
post→jsonPath→isEqualTo→exchange→expectStatus - API client patterns:
request→header→send→statusCode→jsonPath - Builder chains:
builder→field→value→build→validate
iDRegEx Findings
iDRegEx does NOT help at small scale. On 3-method groups, iDRegEx achieves only 3.8× tighter (below the 10× gate threshold). The gate correctly rejects it.
| Group size | CRX lang_size | iDRegEx lang_size | Ratio |
|---|---|---|---|
| 3 methods | 15 | 4 | 3.8× |
4 methods (storage) |
— | — | 91× (outlier) |
Bags survive because:
- CRX emits one grammar deterministically (no alternative to compare)
lang_size_scoreonly ranks between algorithms, not within CRX's own output- iDRegEx is too slow for large groups (200s+ timeout on 2036m FastAPI tests)
Key Insight
The grammar inference pipeline is fundamentally limited by the input: if methods in a package don't share a sequential calling pattern, no algorithm can find one. The ~15% structured grammars represent genuinely reusable patterns; the ~85% bags represent packages with diverse, unrelated methods grouped only by directory proximity.
Files
bex/grammar.py: AST nodes,_count_concatmemoization,_COUNT_CAP = 10^30bex/crx.py: CRX algorithm (AST-based)bex/mdl.py:lang_size_score,model_cost,data_costbex/idregex.py: iDRegEx algorithmbex/decompose.py: sequence decompositionbex/tag_preprocessor/analyze.py: pipeline orchestration, all defaultsexperiments/results/round20_ast_verify/: full experiment data (v2/v3/v4)