docs: handover document for project continuity

This commit is contained in:
tobjend 2026-07-12 19:54:10 +02:00
parent ee60b62854
commit 44dd34a46f

124
experiments/HANDOVER.md Normal file
View file

@ -0,0 +1,124 @@
# Handover — Behavioral Grammar Inference Project
## Current Status
**Branch**: `feature/treesitter-tag-queries` (PR #2)
**Last commit**: `ee60b62` — fix: YAML expansion bug and add GBNF to output
**Tests**: 269 passed, 8 warnings, 0 failures
## What We Built
### Source Code Analysis Pipeline
A **language-agnostic pipeline** that infers per-package calling conventions from any codebase:
1. **Tree-sitter AST** → extract method-level behavioral sequences (call chains, control flow)
2. **Algorithm 7 (CRX)** — generalized regular expression inference from examples
3. **YAML/GBNF output** — structured grammars grouped by package, ready for constrained decoding
### Key Features
- **Zero per-language code** — tree-sitter highlights.scm + behavioral prefix filter
- **10 supported languages**: Kotlin, Python, JavaScript, TypeScript, Go, Rust, Java, C++, Ruby, Swift
- **`--slice package`** — per-package grammars (not per-file)
- **`--split-mixed`** — separates interleaved calling conventions
- **`--decompose`** — decomposition forest for complex sequences (7× more grammars on RAGSAK)
- **`min_structure=0.5`** — filters out "flat bag" noise patterns
- **GBNF output** — llama-compatible constrained decoding grammars
### Algorithm Choice
**CRX (Algorithm 7)** is the default — fast (2ms), always produces something. Refined CRX (cluster-then-infer) is available via `--crx-method refined`. kORE and iDRegEx are opt-in via `--kore` and `--idregex` flags.
### GBNF Grammar Format
Each YAML entry includes a `gbnf` field with the llama-compatible GBNF grammar. This can be passed directly to llama.cpp server API for constrained decoding.
## Files to Know
| File | Purpose |
|------|---------|
| `bex/tag_preprocessor/analyze.py` | Full pipeline: `analyze_directory()``analyze_by_package()``_infer_group()``_build_yaml_output()` |
| `bex/tag_preprocessor/code.py` | `preprocess_by_method()` — AST to behavioral sequences |
| `bex/crx.py` | Standard CRX (Algorithm 7) |
| `bex/crx_refined.py` | Cluster-then-infer CRX |
| `bex/gbnf.py` | SORE→GBNF converter, `validate_sore()`, `grammar_structure_score()` |
| `bex/decompose.py` | Decomposition forest |
| `bex/distributional.py` | Crucio-inspired distributional clustering |
| `bex/grammar_index.py` | `GrammarIndex` class for resolving files to grammars |
| `bex/mcp_server.py` | MCP server with `analyze_directory`, `get_grammar`, `get_package_grammars` |
| `bex/mdl.py` | `lang_size_score()` (default), `mdl_score()` (fallback) |
| `bex/reduce.py` | Algorithm 4 (TODS 2010) for grammar reduction |
| `bex/ensemble.py` | `infer_ensemble()` — combine multiple algorithms |
## Experiments
### Key Findings
1. **CRX wins on simplicity** — no post-processing needed, flat chains are interpretable
2. **`DEFAULT_COVERAGE=0.05`** — was 0.8, almost filtered everything out
3. **`min_methods=3`** — sweet spot (was 5, lost 9 FastAPI grammars)
4. **Decomposition helps diverse codebases** — RAGSAK 4→27, FastAPI 16→29 high-structure grammars
5. **Decomposition hurts structured codebases** — kotlinx.coroutines 24→15
### Decision Matrix (CRX vs Refined CRX)
- CRX struct ≥ 0.2 → use CRX (already good)
- CRX struct < 0.05 use refined (flat bag)
- Group size ≤ 50 → use refined (safe to cluster)
- Group size > 50 → use CRX (refined likely trivial)
### Research Positioning
We are **unique**: first to infer behavioral grammars from source code execution patterns. Related work:
- Panini (white-box CFG from parsers)
- Crucio (black-box CFG from examples)
- XGrammar/DOMINO (constrained decoding)
- Typify/REST (type inference)
Our niche: discover patterns that should be inferred/enforced/typed.
## Open Questions
### 1. Grammar Usefulness for LLM Code Generation
MCP tools are ready but haven't validated if grammars help an LLM during generation. Need to test:
- Does constrained decoding with GBNF improve code quality?
- Do grammars reduce hallucination in call chains?
- Can grammars be used for code completion suggestions?
### 2. Decomposition Trade-off
Decomposition helps diverse codebases but hurts already-structured ones. Need auto-detection:
- If codebase already has good structure → skip decomposition
- If codebase is diverse → apply decomposition
### 3. Cross-Codebase Grammar Reuse
Can grammars from one project inform another? (e.g., "Spring Boot service patterns")
## How to Run
```bash
# Basic analysis
bex --include "*.py" --slice package --main-only --format yaml --output grammars.yml
# With decomposition + structure filtering
bex --include "*.py" --slice package --main-only --decompose --min-structure 0.5 --format yaml
# Full pipeline (what we tested)
bex --include "*.kt" --slice package --main-only --split-mixed --decompose --min-structure 0.5 --format yaml
# MCP server
bex serve --port 8080
```
## Test Coverage
- `tests/test_distributional.py`: 23 tests (distributional clustering)
- `tests/test_decompose.py`: 12 tests (decomposition forest)
- `tests/test_gbnf.py`: 28 tests (GBNF conversion)
- `tests/test_crx_refined.py`: 20 tests (refined CRX)
- `tests/test_grammar_index.py`: 14 tests (grammar index)
- `tests/test_analyze.py`: Pipeline tests
- `tests/test_reduce.py`: Algorithm 4 tests
- `tests/test_mdl.py`: MDL scoring tests
Total: 269 tests passing
## Next Steps
1. **Validate grammar usefulness** — test constrained decoding with llama.cpp
2. **Auto-detect decomposition** — skip if codebase already structured
3. **Cross-project grammar reuse** — share patterns across codebases
4. **IDE integration** — grammar-aware code completion