Refined wins 7/9 when producing useful grammars (model_cost >= 2). Trivial output (single symbol) in 5/13 cases on large groups. CRX wins only once (FastAPI tests, 3618 methods). Key finding: refined CRX is better ~78% of the time but needs triviality check (model_cost >= 2) to avoid single-symbol grammars. |
||
|---|---|---|
| .. | ||
| results | ||
| coarsen_eval.py | ||
| context_eval.py | ||
| EXPERIMENT_LOG.md | ||
| freq_eval.py | ||
| gbnf_eval.py | ||
| RESULTS.md | ||