4.5 KiB
Dervish: Automated Behavioral Grammar Inference for LLM Constrained Decoding
The Breakthrough
We built a system that automatically discovers API usage patterns from source code and converts them into GBNF grammars that constrain LLM output during code generation.
No manual grammar writing. No prompt engineering. Pure automated inference.
What We Achieved
Scale
- 233 grammars inferred across 3 production codebases
- RAGSAK (Kotlin): 102 grammars from 1,742 methods
- FastAPI (Python): 121 grammars from 4,861 methods
- Zod (TypeScript): 10 grammars from 6,203 methods
- Total: 12,806 methods analyzed automatically
Quality
- 89/102 RAGSAK grammars pass quality gate (87%)
- 116/121 FastAPI grammars pass quality gate (96%)
- 9/10 Zod grammars pass quality gate (90%)
- Top 15-20 grammars have genuine domain-specific patterns with ordering
Precision
- 84-94% precision after noise filtering (up from 28-55%)
- Automated noise detection removes test/stdlib tokens
- Quality gate filters useless single-token and bag grammars
Real-World Grammars We Discovered
RAG Pipeline Pattern
listKnowledgeBases → RagRequest → checkKnowledgeBase → request → knowledgeBaseId
36 methods follow this exact sequence. The grammar tells the LLM: "When calling RAG APIs, list KBs first, then build request, then check KB exists."
Document Processing Pipeline
DocumentParsingRequest → GraphDocument → ParsedDocument → asDocumentId → asFilename
6 methods in the doc-parser module follow this pattern. The grammar constrains the LLM to use the correct document processing sequence.
ID Conversion Pattern
asJobId → asDocumentId → asLogicalDocumentId → findOrCreateInactive → asFilename
30 methods convert IDs this way. The grammar ensures the LLM uses the right conversion function for each ID type.
CRUD Operations
findById → saveAll → (parse|runBlocking) → orElseThrow
23 methods follow this Spring Data pattern. The grammar constrains the LLM to proper CRUD sequencing.
Health Check Pattern
HealthCheckReply → healthCheckAsync → collectionExistsAsync → verifyConnectivityAsync
6 methods perform health checks. The grammar ensures the LLM calls all required health check endpoints.
How It Works
1. Tree-Sitter Extraction
Extracts behavioral sequences from source code using tree-sitter AST parsing. Language-agnostic — works with Kotlin, Python, TypeScript, and more.
2. Package-Level Inference
Groups methods by package/directory, then infers regular expression grammars using BEX family algorithms (CRX, iDRegEx). Finds patterns across hundreds of methods automatically.
3. Noise Filtering
Automatically removes test framework calls (assertEquals, mockk, verify) and stdlib calls (listOf, mapOf, filter) from grammars. Precision jumps from 28-55% to 84-94%.
4. Quality Gate
Scores each grammar on structure (ordering, alternation groups, symbol count). Filters out useless single-token and bag grammars. Keeps only grammars that genuinely constrain LLM output.
5. GBNF Output
Converts AST grammars to GBNF format for llama.cpp constrained decoding. Ready to plug into any LLM inference pipeline.
The Impact
Before Dervish
- LLMs generate code without knowing your codebase's conventions
- API calls follow patterns but LLMs don't learn them
- Code review catches convention violations after the fact
After Dervish
- LLMs constrained to your codebase's actual usage patterns
- API calls follow discovered sequences automatically
- Convention violations prevented at generation time
What's Next
- MCP Integration — Expose grammars via MCP tool for opencode
- Dynamic Regeneration — Re-infer grammars when codebase changes
- Multi-Language Expansion — Apply to Go, Rust, Java, C++
- Cross-Codebase Learning — Transfer patterns between projects
The Numbers That Matter
| Metric | Value |
|---|---|
| Codebases analyzed | 3 |
| Total methods | 12,806 |
| Grammars inferred | 233 |
| Grammars passing quality gate | 214 |
| Domain-specific patterns discovered | 15-20 |
| Precision after filtering | 84-94% |
| Manual grammar writing required | 0 |
Bottom Line
We turned 12,806 methods across 3 codebases into 233 behavioral grammars that constrain LLM output to your actual API usage patterns.
No prompt engineering. No manual rules. Pure automated inference from your source code.
The LLMs now know your codebase's conventions — because we taught them.