grammar-inference-engine/tests
tobjend 8b2899d16e Phase 2: decomposition forest for complex sequences
Decomposition breaks long sequences into shorter fragments before
inference. This helps when sequences are too long for CRX to handle
(>5 symbols → flat bags).

Results:
- RAGSAK: 21 → 80 grammars (3.8× increase)
- FastAPI: 111 → 118 grammars (small increase)

Changes:
- bex/decompose.py: decompose_sequence(), decompose_all(), decompose_with_coverage()
- bex/tag_preprocessor/analyze.py: --decompose, --max-seq-length flags
- Skip diversity check when decomposing (decomposition creates diverse fragments)
- 12 new tests in tests/test_decompose.py

Co-authored-by: OpenCode <opencode@corentic.eu>
2026-07-12 17:49:35 +02:00
..
test_analyze.py feat: iDRegEx refinement for CRX flat bags (Round 16) 2026-07-12 16:54:40 +02:00
test_bex.py clean up agent-betraying comments; fix stale test names 2026-07-01 13:26:03 +02:00
test_crx_refined.py feat: CRX refined — cluster-then-infer for tighter grammars 2026-07-12 01:50:40 +02:00
test_decompose.py Phase 2: decomposition forest for complex sequences 2026-07-12 17:49:35 +02:00
test_distributional.py feat: distributional clustering for better grouping (Phase 1) 2026-07-12 17:37:36 +02:00
test_ensemble.py drop kORE from default ensemble, add --kore flag to opt in 2026-07-04 02:58:07 +02:00
test_gbnf.py fix(gbnf): handle disjunction inside parens and compound repetition 2026-07-12 02:20:58 +02:00
test_grammar_index.py feat: grammar index for runtime lookup 2026-07-12 14:08:05 +02:00
test_kore.py WIP: language size scoring + diversity threshold (step 1 pending) 2026-07-11 22:56:42 +02:00
test_reduce.py feat: implement Reduce algorithm (Algorithm 4, TODS 2010) 2026-07-12 00:11:38 +02:00
test_scoring.py WIP: language size scoring + diversity threshold (step 1 pending) 2026-07-11 22:56:42 +02:00