From 44dd34a46f96f84c964c047403ded5527ccf16c2 Mon Sep 17 00:00:00 2001 From: tobjend Date: Sun, 12 Jul 2026 19:54:10 +0200 Subject: [PATCH] docs: handover document for project continuity --- experiments/HANDOVER.md | 124 ++++++++++++++++++++++++++++++++++++++++ 1 file changed, 124 insertions(+) create mode 100644 experiments/HANDOVER.md diff --git a/experiments/HANDOVER.md b/experiments/HANDOVER.md new file mode 100644 index 0000000..0711579 --- /dev/null +++ b/experiments/HANDOVER.md @@ -0,0 +1,124 @@ +# Handover — Behavioral Grammar Inference Project + +## Current Status + +**Branch**: `feature/treesitter-tag-queries` (PR #2) +**Last commit**: `ee60b62` — fix: YAML expansion bug and add GBNF to output +**Tests**: 269 passed, 8 warnings, 0 failures + +## What We Built + +### Source Code Analysis Pipeline +A **language-agnostic pipeline** that infers per-package calling conventions from any codebase: + +1. **Tree-sitter AST** → extract method-level behavioral sequences (call chains, control flow) +2. **Algorithm 7 (CRX)** — generalized regular expression inference from examples +3. **YAML/GBNF output** — structured grammars grouped by package, ready for constrained decoding + +### Key Features +- **Zero per-language code** — tree-sitter highlights.scm + behavioral prefix filter +- **10 supported languages**: Kotlin, Python, JavaScript, TypeScript, Go, Rust, Java, C++, Ruby, Swift +- **`--slice package`** — per-package grammars (not per-file) +- **`--split-mixed`** — separates interleaved calling conventions +- **`--decompose`** — decomposition forest for complex sequences (7× more grammars on RAGSAK) +- **`min_structure=0.5`** — filters out "flat bag" noise patterns +- **GBNF output** — llama-compatible constrained decoding grammars + +### Algorithm Choice +**CRX (Algorithm 7)** is the default — fast (2ms), always produces something. Refined CRX (cluster-then-infer) is available via `--crx-method refined`. kORE and iDRegEx are opt-in via `--kore` and `--idregex` flags. + +### GBNF Grammar Format +Each YAML entry includes a `gbnf` field with the llama-compatible GBNF grammar. This can be passed directly to llama.cpp server API for constrained decoding. + +## Files to Know + +| File | Purpose | +|------|---------| +| `bex/tag_preprocessor/analyze.py` | Full pipeline: `analyze_directory()` → `analyze_by_package()` → `_infer_group()` → `_build_yaml_output()` | +| `bex/tag_preprocessor/code.py` | `preprocess_by_method()` — AST to behavioral sequences | +| `bex/crx.py` | Standard CRX (Algorithm 7) | +| `bex/crx_refined.py` | Cluster-then-infer CRX | +| `bex/gbnf.py` | SORE→GBNF converter, `validate_sore()`, `grammar_structure_score()` | +| `bex/decompose.py` | Decomposition forest | +| `bex/distributional.py` | Crucio-inspired distributional clustering | +| `bex/grammar_index.py` | `GrammarIndex` class for resolving files to grammars | +| `bex/mcp_server.py` | MCP server with `analyze_directory`, `get_grammar`, `get_package_grammars` | +| `bex/mdl.py` | `lang_size_score()` (default), `mdl_score()` (fallback) | +| `bex/reduce.py` | Algorithm 4 (TODS 2010) for grammar reduction | +| `bex/ensemble.py` | `infer_ensemble()` — combine multiple algorithms | + +## Experiments + +### Key Findings +1. **CRX wins on simplicity** — no post-processing needed, flat chains are interpretable +2. **`DEFAULT_COVERAGE=0.05`** — was 0.8, almost filtered everything out +3. **`min_methods=3`** — sweet spot (was 5, lost 9 FastAPI grammars) +4. **Decomposition helps diverse codebases** — RAGSAK 4→27, FastAPI 16→29 high-structure grammars +5. **Decomposition hurts structured codebases** — kotlinx.coroutines 24→15 + +### Decision Matrix (CRX vs Refined CRX) +- CRX struct ≥ 0.2 → use CRX (already good) +- CRX struct < 0.05 → use refined (flat bag) +- Group size ≤ 50 → use refined (safe to cluster) +- Group size > 50 → use CRX (refined likely trivial) + +### Research Positioning +We are **unique**: first to infer behavioral grammars from source code execution patterns. Related work: +- Panini (white-box CFG from parsers) +- Crucio (black-box CFG from examples) +- XGrammar/DOMINO (constrained decoding) +- Typify/REST (type inference) + +Our niche: discover patterns that should be inferred/enforced/typed. + +## Open Questions + +### 1. Grammar Usefulness for LLM Code Generation +MCP tools are ready but haven't validated if grammars help an LLM during generation. Need to test: +- Does constrained decoding with GBNF improve code quality? +- Do grammars reduce hallucination in call chains? +- Can grammars be used for code completion suggestions? + +### 2. Decomposition Trade-off +Decomposition helps diverse codebases but hurts already-structured ones. Need auto-detection: +- If codebase already has good structure → skip decomposition +- If codebase is diverse → apply decomposition + +### 3. Cross-Codebase Grammar Reuse +Can grammars from one project inform another? (e.g., "Spring Boot service patterns") + +## How to Run + +```bash +# Basic analysis +bex --include "*.py" --slice package --main-only --format yaml --output grammars.yml + +# With decomposition + structure filtering +bex --include "*.py" --slice package --main-only --decompose --min-structure 0.5 --format yaml + +# Full pipeline (what we tested) +bex --include "*.kt" --slice package --main-only --split-mixed --decompose --min-structure 0.5 --format yaml + +# MCP server +bex serve --port 8080 +``` + +## Test Coverage + +- `tests/test_distributional.py`: 23 tests (distributional clustering) +- `tests/test_decompose.py`: 12 tests (decomposition forest) +- `tests/test_gbnf.py`: 28 tests (GBNF conversion) +- `tests/test_crx_refined.py`: 20 tests (refined CRX) +- `tests/test_grammar_index.py`: 14 tests (grammar index) +- `tests/test_analyze.py`: Pipeline tests +- `tests/test_reduce.py`: Algorithm 4 tests +- `tests/test_mdl.py`: MDL scoring tests + +Total: 269 tests passing + +## Next Steps + +1. **Validate grammar usefulness** — test constrained decoding with llama.cpp +2. **Auto-detect decomposition** — skip if codebase already structured +3. **Cross-project grammar reuse** — share patterns across codebases +4. **IDE integration** — grammar-aware code completion