grammar-inference-engine/experiments/HANDOVER.md

5.5 KiB
Raw Blame History

Handover — Behavioral Grammar Inference Project

Current Status

Branch: feature/treesitter-tag-queries (PR #2) Last commit: ee60b62 — fix: YAML expansion bug and add GBNF to output Tests: 269 passed, 8 warnings, 0 failures

What We Built

Source Code Analysis Pipeline

A language-agnostic pipeline that infers per-package calling conventions from any codebase:

  1. Tree-sitter AST → extract method-level behavioral sequences (call chains, control flow)
  2. Algorithm 7 (CRX) — generalized regular expression inference from examples
  3. YAML/GBNF output — structured grammars grouped by package, ready for constrained decoding

Key Features

  • Zero per-language code — tree-sitter highlights.scm + behavioral prefix filter
  • 10 supported languages: Kotlin, Python, JavaScript, TypeScript, Go, Rust, Java, C++, Ruby, Swift
  • --slice package — per-package grammars (not per-file)
  • --split-mixed — separates interleaved calling conventions
  • --decompose — decomposition forest for complex sequences (7× more grammars on RAGSAK)
  • min_structure=0.5 — filters out "flat bag" noise patterns
  • GBNF output — llama-compatible constrained decoding grammars

Algorithm Choice

CRX (Algorithm 7) is the default — fast (2ms), always produces something. Refined CRX (cluster-then-infer) is available via --crx-method refined. kORE and iDRegEx are opt-in via --kore and --idregex flags.

GBNF Grammar Format

Each YAML entry includes a gbnf field with the llama-compatible GBNF grammar. This can be passed directly to llama.cpp server API for constrained decoding.

Files to Know

File Purpose
bex/tag_preprocessor/analyze.py Full pipeline: analyze_directory()analyze_by_package()_infer_group()_build_yaml_output()
bex/tag_preprocessor/code.py preprocess_by_method() — AST to behavioral sequences
bex/crx.py Standard CRX (Algorithm 7)
bex/crx_refined.py Cluster-then-infer CRX
bex/gbnf.py SORE→GBNF converter, validate_sore(), grammar_structure_score()
bex/decompose.py Decomposition forest
bex/distributional.py Crucio-inspired distributional clustering
bex/grammar_index.py GrammarIndex class for resolving files to grammars
bex/mcp_server.py MCP server with analyze_directory, get_grammar, get_package_grammars
bex/mdl.py lang_size_score() (default), mdl_score() (fallback)
bex/reduce.py Algorithm 4 (TODS 2010) for grammar reduction
bex/ensemble.py infer_ensemble() — combine multiple algorithms

Experiments

Key Findings

  1. CRX wins on simplicity — no post-processing needed, flat chains are interpretable
  2. DEFAULT_COVERAGE=0.05 — was 0.8, almost filtered everything out
  3. min_methods=3 — sweet spot (was 5, lost 9 FastAPI grammars)
  4. Decomposition helps diverse codebases — RAGSAK 4→27, FastAPI 16→29 high-structure grammars
  5. Decomposition hurts structured codebases — kotlinx.coroutines 24→15

Decision Matrix (CRX vs Refined CRX)

  • CRX struct ≥ 0.2 → use CRX (already good)
  • CRX struct < 0.05 → use refined (flat bag)
  • Group size ≤ 50 → use refined (safe to cluster)
  • Group size > 50 → use CRX (refined likely trivial)

Research Positioning

We are unique: first to infer behavioral grammars from source code execution patterns. Related work:

  • Panini (white-box CFG from parsers)
  • Crucio (black-box CFG from examples)
  • XGrammar/DOMINO (constrained decoding)
  • Typify/REST (type inference)

Our niche: discover patterns that should be inferred/enforced/typed.

Open Questions

1. Grammar Usefulness for LLM Code Generation

MCP tools are ready but haven't validated if grammars help an LLM during generation. Need to test:

  • Does constrained decoding with GBNF improve code quality?
  • Do grammars reduce hallucination in call chains?
  • Can grammars be used for code completion suggestions?

2. Decomposition Trade-off

Decomposition helps diverse codebases but hurts already-structured ones. Need auto-detection:

  • If codebase already has good structure → skip decomposition
  • If codebase is diverse → apply decomposition

3. Cross-Codebase Grammar Reuse

Can grammars from one project inform another? (e.g., "Spring Boot service patterns")

How to Run

# Basic analysis
bex --include "*.py" --slice package --main-only --format yaml --output grammars.yml

# With decomposition + structure filtering
bex --include "*.py" --slice package --main-only --decompose --min-structure 0.5 --format yaml

# Full pipeline (what we tested)
bex --include "*.kt" --slice package --main-only --split-mixed --decompose --min-structure 0.5 --format yaml

# MCP server
bex serve --port 8080

Test Coverage

  • tests/test_distributional.py: 23 tests (distributional clustering)
  • tests/test_decompose.py: 12 tests (decomposition forest)
  • tests/test_gbnf.py: 28 tests (GBNF conversion)
  • tests/test_crx_refined.py: 20 tests (refined CRX)
  • tests/test_grammar_index.py: 14 tests (grammar index)
  • tests/test_analyze.py: Pipeline tests
  • tests/test_reduce.py: Algorithm 4 tests
  • tests/test_mdl.py: MDL scoring tests

Total: 269 tests passing

Next Steps

  1. Validate grammar usefulness — test constrained decoding with llama.cpp
  2. Auto-detect decomposition — skip if codebase already structured
  3. Cross-project grammar reuse — share patterns across codebases
  4. IDE integration — grammar-aware code completion