grammar-inference-engine/experiments
tobjend 93d53f164a fix: byte/char offset mismatch in tree-sitter text extraction
Root cause: parser.parse(code.encode()) returns byte offsets, but
code[node.start_byte:node.end_byte] indexed into a Python string
(character offsets). Non-ASCII chars caused cumulative drift → truncated symbols.

Fix: store code_bytes = code.encode(), index into that, decode only final text.

Impact:
- Zod: 653 truncated symbols → 0
- RAGSAK: 653 truncated symbols → 0
- RAGSAK grammars: 8 → 27 (3.4×)
- FastAPI grammars: 33 → 111 (3.4×)
- Malformed: 6 → 2 (RAGSAK), 0 (FastAPI/Flask)

Also removed sanitize_symbol() (dead code after fix) and call_only parameter.
2026-07-12 15:57:27 +02:00
..
results feat: grammar_structure_score + min_structure filter 2026-07-12 03:02:50 +02:00
coarsen_eval.py feat: 4-codebase evaluation — conventions vs completions tradeoff 2026-07-12 01:14:09 +02:00
context_eval.py feat: implement Reduce algorithm (Algorithm 4, TODS 2010) 2026-07-12 00:11:38 +02:00
EXPERIMENT_LOG.md fix: byte/char offset mismatch in tree-sitter text extraction 2026-07-12 15:57:27 +02:00
freq_eval.py feat: frequency threshold sweep — 0.01-0.20 across 4 codebases 2026-07-12 01:26:45 +02:00
gbnf_eval.py perf: iDRegEx opt-in, GBNF newline fix, OverflowError fix 2026-07-12 02:52:05 +02:00
RESULTS.md docs: experiment log + cross-package analysis 2026-07-12 00:45:21 +02:00