# Where We Fit in the Science ## The Landscape (2024-2026) Grammar inference for code is active across 4 distinct research areas: ``` ┌─────────────────────────────────────────────────────────────────────┐ │ Grammar Inference Landscape │ ├─────────────────┬─────────────────┬─────────────────┬───────────────┤ │ Black-Box CFG │ White-Box CFG │ Constrained │ Behavioral │ │ Inference │ Inference │ Decoding │ Type Inference│ ├─────────────────┼─────────────────┼─────────────────┼───────────────┤ │ Crucio (ICSE'26)│ Panini (OOPSLA'25)│ XGrammar │ REST (OOPSLA'25)│ │ Kedavra (ASE'24)│ Leon (ICSE'25) │ DOMINO │ Typify (ICPC'26)│ │ Arvada │ │ ASAp (NeurIPS'24)│ RightTyper │ │ Treevada │ │ CRANE (ICML'25) │ DAInfer+ │ │ │ │ TreeCoder │ │ └─────────────────┴─────────────────┴─────────────────┴───────────────┘ ``` ## What Others Do ### Black-Box CFG Inference (Crucio, Kedavra, Arvada) - **Input**: Example strings + oracle (accept/reject) - **Output**: Context-free grammar - **Method**: Decompose strings, generalize via distributional analysis - **Target**: Programming language grammars (JSON, XML, C, Java) - **Limitation**: Needs oracle, assumes regular/context-free languages ### White-Box CFG Inference (Panini, Leon) - **Input**: Parser source code (ad hoc parsers) - **Output**: Regular grammar - **Method**: Refinement type inference + abstract interpretation - **Target**: String parsing functions (split, regex, format) - **Limitation**: Only works on parser functions, not behavioral patterns ### Constrained Decoding (XGrammar, DOMINO, ASAp, CRANE) - **Input**: Grammar + LLM - **Output**: LLM constrained to grammar - **Method**: Mask invalid tokens during generation - **Target**: JSON, SQL, code with strict syntax - **Limitation**: Requires pre-existing grammar, only enforces syntax ### Behavioral Type Inference (REST, Typify, RightTyper) - **Input**: Source code - **Output**: Type annotations - **Method**: Static/dynamic analysis + type inference - **Target**: Function signatures, return types, parameter types - **Limitation**: Types only, not behavioral patterns ## Where We Are Different **We are the only approach that infers behavioral grammars from execution patterns.** ``` Others: Parser source code → Grammar (for input validation) Example strings → Grammar (for language definition) Grammar → LLM (for output constraint) Us: Source code → Behavioral sequences → Grammar (for usage patterns) Grammar → LLM (for context/constraint) ``` ### Key Distinctions | Aspect | Others | Us | |--------|--------|-----| | **Input** | Parser code or example strings | Any codebase (behavioral sequences) | | **Output** | Grammar for input validation | Grammar for usage patterns | | **Target** | String parsing functions | API call sequences | | **Granularity** | Per-function or per-language | Per-package/module | | **Language support** | Usually 1 language | Any tree-sitter supported language | | **Use case** | Formal verification, testing | LLM code generation guidance | ### The Gap We Fill 1. **Panini** infers grammars for individual ad hoc parsers (e.g., `json.loads`). We infer grammars for *how packages are used* (e.g., Flask route patterns). 2. **Crucio/Kedavra** infer grammars from input/output examples. We infer grammars from *observed execution patterns* — no oracle needed, we have the source. 3. **XGrammar/DOMINO** enforce pre-existing grammars during LLM generation. We *discover* the grammars that should be enforced. 4. **REST/Typify** infer types (what something is). We infer *behavioral patterns* (how something is used). ## Our Contribution ### The Behavioral Grammar Concept **Definition**: A behavioral grammar captures the valid sequences of API calls within a package or module, expressed as a regular expression. ```python # Example: Flask route handler Grammar: GET_RETURN (POST_RETURN)* RETURN # Means: Routes often start with GET, sometimes POST, always return ``` ### The Pipeline ``` Source Code → tree-sitter AST → Behavioral Sequences → BEX Algorithms → YAML/GBNF ↑ Our innovation (language-agnostic preprocessing) ``` ### What Makes It Work 1. **Behavioral prefix extraction**: We capture the *intent* of code, not the implementation 2. **Token coarsening**: RETURN, IF, EXCEPTION, LOOP — abstract enough to generalize 3. **Package slicing**: Context matters (Flask vs FastAPI vs Django) 4. **Recursive splitting**: Separate mixed groups by first symbol 5. **Multiple algorithms**: CRX (fast), refined CRX (tighter), iDRegEx (rare) ## Where We Don't Fit (Yet) ### Limitations Compared to Others 1. **Not formal verification**: Our grammars are approximate, not proven correct 2. **Not parser inference**: We don't infer grammars for string parsing 3. **Not constraint enforcement**: We don't yet integrate with XGrammar/DOMINO 4. **Not type inference**: We complement types, not replace them 5. **Not language-specific**: We don't leverage language-specific type systems ### What We Could Become 1. **Grammar + Type hybrid**: Combine our behavioral grammars with Typify's type inference 2. **Constrained decoding integration**: Feed our grammars to XGrammar for LLM guidance 3. **API specification inference**: Combine with DAInfer+ for full API contracts 4. **Testing**: Use behavioral grammars for test case generation 5. **Documentation**: Auto-generate usage patterns from code ## Research Positioning ### Our Niche ``` ┌─────────────────────────────┐ │ Behavioral Grammar │ │ Inference (Us) │ └──────────┬──────────────────┘ │ ┌──────────────────────┼──────────────────────┐ │ │ │ ▼ ▼ ▼ ┌───────────────┐ ┌───────────────┐ ┌───────────────┐ │ Tree-sitter │ │ BEX Algorithm │ │ LLM Context │ │ Preprocessing │ │ Adaptation │ │ Generation │ │ (Language- │ │ (XML → Code │ │ (Grammar → │ │ agnostic) │ │ patterns) │ │ Prompt/ │ │ │ │ │ │ Constraint) │ └───────────────┘ └───────────────┘ └───────────────┘ ``` ### Related Work (Cite) 1. **Bex et al. (2010)**: "Inference of concise regular expressions and DTDs" — CRX algorithm 2. **Schröder & Cito (2025)**: "Static inference of regular grammars for ad hoc parsers" — Panini 3. **Li et al. (2024)**: "Incremental context-free grammar inference in black box settings" — Kedavra 4. **Li et al. (2026)**: "Context-free grammar inference for complex programming languages" — Crucio 5. **Dong et al. (2024)**: "XGrammar: Flexible and efficient structured generation engine" — Constrained decoding 6. **Park et al. (2025)**: "Flexible and efficient grammar-constrained decoding" — GCD 7. **Tam et al. (2025)**: "Grammar-constrained decoding makes LLMs better logical parsers" — GCD + reasoning 8. **Masoudian et al. (2026)**: "DAInfer+: Neurosymbolic inference of API specifications" — API contracts 9. **Typify (2026)**: "Usage-driven static analyzer for precise Python type inference" — Type inference ### Our Novelty 1. **First to apply BEX to behavioral sequences** (not XML/input data) 2. **Language-agnostic preprocessing** via tree-sitter (not language-specific) 3. **Package-level behavioral patterns** (not per-function or per-language) 4. **Grammar as LLM context** (not formal verification or testing) 5. **Hybrid approach** combining BEX algorithms with modern preprocessing ## Summary We occupy a unique position: **behavioral grammar inference from source code**. Others do: - Grammar inference for parsers (Panini, Crucio) - Grammar enforcement for LLMs (XGrammar, DOMINO) - Type inference for code (Typify, REST) We do: **Discover the patterns that should be inferred, enforced, or typed**.