Commit graph

11 commits

Author SHA1 Message Date
tobjend
ea6cac53e3 wip: AST foundation — grammar.py, expr.py→AST, soa.py→AST labels 2026-07-13 01:13:48 +02:00
tobjend
9ca56e2c69 MCP server: add decomposition + quality filtering params
- analyze_directory tool: add decompose, max_seq_length, cluster_method,
  crx_method params for Crucio-inspired improvements
- min_structure default bumped to 0.5 for MCP (only high-structure
  grammars returned to agents)
- _build_yaml_output: filter entries below min_structure threshold
- README: update source code analysis section, fix kORE/iDRegEx mentions
- 269 tests pass

Co-authored-by: OpenCode <opencode@corentic.eu>
2026-07-12 18:56:35 +02:00
tobjend
44415c5b42 feat: runtime grammar lookup via MCP + fix YAML output
MCP server:
- analyze_directory: updated signature to match CLI (split_mixed,
  min_structure, min_methods, min_coverage=0.05)
- get_grammar(directory, file_path, context_symbol): returns the right
  GBNF for constrained generation at code generation time
- get_package_grammars(directory, file_path): lists all grammars for
  a file's package, ranked by quality

YAML output:
- _build_yaml_output now includes leaf grammars from recursive split
  (all_grammars in meta), so each calling context gets its own entry

Workflow for agents:
  1. analyze_directory → grammars.yml persisted
  2. get_grammar(file, 'return') → GBNF for constrained generation
  3. LLM generates code following the package's convention
2026-07-12 14:08:12 +02:00
tobjend
dfb56a083a WIP: language size scoring + diversity threshold (step 1 pending) 2026-07-11 22:56:42 +02:00
tobjend
830104b399 feat: add analyze_directory MCP tool — scan codebase, infer conventions, persist to .dervish/
Some checks failed
ci/woodpecker/push/woodpecker Pipeline failed
ci/woodpecker/pr/woodpecker Pipeline failed
2026-07-11 21:28:35 +02:00
tobjend
e94c52b71a docs: update stale docs — remove kORE from default ensemble, add tag preprocessor CLI
Some checks failed
ci/woodpecker/push/woodpecker Pipeline failed
ci/woodpecker/pr/woodpecker Pipeline failed
2026-07-11 20:51:10 +02:00
tobjend
036a84cc76 docs: add min_coverage to MCP tool + README, include core in output
All checks were successful
ci/woodpecker/push/woodpecker Pipeline was successful
ci/woodpecker/pr/woodpecker Pipeline was successful
2026-07-01 15:16:24 +02:00
tobjend
b8cc40177c remove redundant infer_grammar tool; update docs to single-tool MCP 2026-07-01 13:15:19 +02:00
tobjend
d7477344a6 move format-specific adapters to examples/, purge format-specific MCP tools 2026-07-01 10:36:14 +02:00
tobjend
0e2aec582b Grammar inference engine: CRX + iDRegEx ensemble with MDL scoring, MCP server, showcase, and blog post
- Ensemble inference (infer_ensemble) runs both CRX and iDRegEx, picks best by MDL
- CRX: CRX algorithm for wide coverage (accepts all sequences, large vocabulary)
- iDRegEx: iDRegEx for minimal core grammar (tightest common pattern)
- MDL scoring: fixed model_cost to count alphabet symbol occurrences, fixed dispatch order in _count_words_fast
- Fixed _match_tokens: rewritten as _match_possible with proper backtracking
- Fixed _parse_parts disjunction: children use _parse_flat_symbol to avoid dot-splitting
- MCP server: infer_best_grammar and infer_grammar tools
- Added prefer parameter (crx/idregex) to skip ensemble
- 28 passing tests
- SHOWCASE.md with Geerlingguy Galaxy demonstration
- blog_post.md with full technical deep-dive
2026-07-01 09:51:41 +02:00
tobjend
adc52c99ec Add MCP server: grammar inference via FastMCP
- bex/mcp_server.py: FastMCP server with 3 tools:
  * infer_grammar(sequences, method='crx'|'idregex')
  * infer_yaml_grammar(yaml_dir, pattern, method)
  * infer_ansible_role_grammar(roles_dir)
- pyproject.toml: add bex-mcp console_scripts entry point
2026-07-01 08:03:10 +02:00