Key findings: - File path grouping (structural) = baseline (0.6% coverage) — no improvement - First-k-symbols grouping (behavioral) = 20× improvement (12% coverage at k=3) - Two-dimensional (path + symbols) = worse than behavioral alone - Winner: Option B k=3 — 47 SOREs, 12% coverage Smart filters make experiments fast (<1s vs minutes): - max_unique_ratio=0.85 - max_alphabet=20 - max_soa_edges=100 Preserved in experiments/results/ for future reference.
99 lines
5 KiB
Markdown
99 lines
5 KiB
Markdown
# Experiment Results — RAGSAK Context Strategies
|
||
|
||
**Date**: 2026-07-11
|
||
**Codebase**: RAGSAK (462 .kt files, 1594 methods, 137 packages)
|
||
|
||
## Summary Table
|
||
|
||
| Strategy | Contexts | SORE OK | Coverage | Notes |
|
||
|----------|----------|---------|----------|-------|
|
||
| Baseline (package) | 116 | 2 | 0.6% | Current approach — terrible |
|
||
| File path k=1 | 43 | 2 | 1.1% | Structural — no better than baseline |
|
||
| File path k=2 | 54 | 2 | 1.1% | Same — structural doesn't help |
|
||
| File path k=3 | 54 | 2 | 1.1% | Same — structural doesn't help |
|
||
| **First 1 symbol** | 134 | **20** | **4.6%** | Behavioral — 10× better |
|
||
| **First 2 symbols** | 141 | **39** | **9.8%** | Behavioral — 20× better |
|
||
| **First 3 symbols** | 117 | **47** | **12.0%** | Behavioral — **winner** |
|
||
| Two-dim (p1,s1) | 151 | 24 | 6.3% | Hybrid — worse than behavioral alone |
|
||
| Two-dim (p1,s2) | 119 | 33 | 8.7% | Hybrid — worse than behavioral alone |
|
||
| Two-dim (p2,s1) | 146 | 24 | 6.3% | Hybrid — worse than behavioral alone |
|
||
| Two-dim (p2,s2) | 117 | 33 | 8.7% | Hybrid — worse than behavioral alone |
|
||
| Return type heuristic | 5 | 0 | 0.0% | Too coarse — useless |
|
||
|
||
## Key Findings
|
||
|
||
### 1. Behavioral grouping > Structural grouping
|
||
- File path strategies (A) produce same results as baseline — structural similarity doesn't predict behavioral similarity
|
||
- First-k-symbols strategies (B) produce 20× more SOREs — behavioral prefix is the right signal
|
||
- **Winner: Option B, k=3** — 47 SOREs, 12% coverage
|
||
|
||
### 2. Two-dimensional doesn't help
|
||
- Adding file path to behavioral grouping **reduces** performance
|
||
- The structural component drags down the behavioral signal
|
||
- Conclusion: **drop the structural dimension entirely**
|
||
|
||
### 3. The successful SOREs reveal real patterns
|
||
|
||
**Test patterns** (most common):
|
||
- `post.uri.exchange.expectStatus.(expectBody.jsonPath.isEqualTo)?` — HTTP test flow
|
||
- `runTest.post.uri.bodyValue.(ChatRequest|SessionChatRequest).exchange.expectStatus` — test setup
|
||
- `every.getJobStatus.get.uri.exchange.expectStatus.(expectBody.jsonPath.isEqualTo)?` — polling pattern
|
||
|
||
**Assertion patterns**:
|
||
- `assertEquals.(of.(assertFailsWith)?)+` — assertion chain
|
||
- `filesIn.assertTrue.(hasImport)+` — file validation
|
||
- `trim.lowercase.(warn)+` — string processing
|
||
|
||
**Mock patterns**:
|
||
- `mockk.every....` — mock setup
|
||
- `clearAllMocks` — cleanup
|
||
|
||
**CRUD patterns**:
|
||
- `info.deleteByJobId` — delete operation
|
||
- `info.deleteByKnowledgeBaseId` — delete operation
|
||
- `update.trimIndent.now.insertRow` — update operation
|
||
|
||
### 4. Why coverage is still low (12%)
|
||
- Most methods have **unique call patterns** — they don't share a 3-symbol prefix with any other method
|
||
- The 3-symbol prefix is too specific for diverse codebases
|
||
- Need either:
|
||
- More data (more codebases to train on)
|
||
- Looser grouping (k=1 or k=2, but with merging)
|
||
- Symbol coarsening (collapse specific symbols into categories)
|
||
|
||
### 5. Smart filters work well
|
||
- `max_unique_ratio=0.85`: Skips 60-80% of groups (too diverse)
|
||
- `max_alphabet=20`: Skips groups with too many unique symbols
|
||
- `max_soa_edges=100`: Skips complex SOAs that would be slow
|
||
- **Result**: Experiment runs in <1s instead of minutes
|
||
|
||
## Successful SORE Examples (Top 10 by group size)
|
||
|
||
| Context | Methods | Unique Ratio | SORE |
|
||
|---------|---------|--------------|------|
|
||
| `ery {.stKnowledgeBases().gRequest(m` | 12 | 0.333 | `ery {.stKnowledgeBases().gRequest(m.(ckKnowledgeBase(re.ertEquals(Kn."k\|eckKnowledgeBase(r.(sertEquals(K.("\|(sertEquals(` |
|
||
| `mockk` | 11 | 0.091 | `mockk` |
|
||
| `filesIn.assertTrue.hasImport` | 7 | 0.286 | `filesIn.assertTrue.(hasImport)+` |
|
||
| `of` | 7 | 0.143 | `of` |
|
||
| `assertEquals.of.assertFailsWith` | 7 | 0.429 | `assertEquals.(of.(assertFailsWith)?)+` |
|
||
| `filesIn.filter.contains` | 5 | 0.400 | `filesIn.filter.contains.assertTrue.(hasImport)+` |
|
||
| `trim.lowercase.warn` | 5 | 0.400 | `trim.lowercase.(warn)+` |
|
||
| `DoclingConfig.assertThatThrownBy.validateCriticalSettings` | 5 | 0.200 | `DoclingConfig.assertThatThrownBy.validateCriticalSettings.isInstanceOf.hasMessageContaining` |
|
||
| `defaultCapabilityId` | 4 | 0.250 | `defaultCapabilityId` |
|
||
| `clearAllMocks` | 4 | 0.250 | `clearAllMocks` |
|
||
|
||
## Next Steps
|
||
|
||
1. **Symbol coarsening**: Collapse specific symbols into categories (GETTER, SETTER, BUILDER, etc.) to reduce fragmentation
|
||
2. **Hierarchical grouping**: Group by k=1 first, then sub-group by k=2 within each group
|
||
3. **Merge similar contexts**: After initial grouping, merge contexts with similar SOREs (Bex's Reduce algorithm)
|
||
4. **Test on more codebases**: RAGSAK may not be representative — need more data
|
||
5. **Accept the limit**: Some methods are genuinely unique — 12% coverage may be the ceiling for this approach
|
||
|
||
## Files Generated
|
||
|
||
- `experiments/results/summary.json` — metrics for all strategies
|
||
- `experiments/results/*.json` — detailed results per strategy
|
||
- `experiments/context_eval.py` — experiment runner script
|
||
- `Papers/02-06-2026/context-design-analysis.md` — analysis of all options
|
||
- `Papers/02-06-2026/bex-research-summary.md` — summary of Bex's papers
|