Decomposition breaks long sequences into shorter fragments before inference. This helps when sequences are too long for CRX to handle (>5 symbols → flat bags). Results: - RAGSAK: 21 → 80 grammars (3.8× increase) - FastAPI: 111 → 118 grammars (small increase) Changes: - bex/decompose.py: decompose_sequence(), decompose_all(), decompose_with_coverage() - bex/tag_preprocessor/analyze.py: --decompose, --max-seq-length flags - Skip diversity check when decomposing (decomposition creates diverse fragments) - 12 new tests in tests/test_decompose.py Co-authored-by: OpenCode <opencode@corentic.eu>
4.4 KiB
4.4 KiB
Phase 2: Decomposition Forest
Goal
Break down complex/long behavioral sequences into shorter ones that still capture the pattern. This helps when sequences are long and diverse, making CRX produce flat bags.
Current Problem
Long sequences like:
["if", "return", "if", "return", "if", "return"]
CRX sees 6 symbols, tries to find a pattern, often produces flat bags like (if|return)*.
If we decompose into shorter examples:
["if", "return"]
["if", "return"]
["if", "return"]
CRX sees a clear pattern: if.return (repeated).
Crucio's Approach
Crucio uses three decomposition strategies:
- Binary maximum subsequence deletion: Split in half, delete max from each half
- Maximum subsequence deletion: Delete largest contiguous chunk
- Subsequence replacement: Replace a chunk with a shorter version
Key insight: Decomposed sequences must preserve grammar coverage (be valid under the same grammar).
Our Adaptation
For behavioral sequences, we need simpler decomposition:
- Prefix extraction: Take first N symbols
- Suffix extraction: Take last N symbols
- Window extraction: Take middle N symbols
- Pattern extraction: Find repeated patterns and extract one instance
Implementation Plan
File: bex/decompose.py (new)
"""Decomposition forest for behavioral sequences.
Inspired by Crucio's decomposition forest (ICSE 2026).
Breaks down long sequences into shorter ones that preserve patterns.
"""
def decompose_sequence(seq, max_length=5):
"""Decompose a sequence into shorter fragments.
Strategies:
1. If seq <= max_length, return as-is
2. Extract prefixes of length 1..max_length
3. Extract suffixes of length 1..max_length
4. Extract windows of length max_length
Args:
seq: List of symbols
max_length: Maximum fragment length
Returns:
List of fragments (shorter sequences)
"""
if len(seq) <= max_length:
return [seq]
fragments = []
# Prefixes
for i in range(1, min(max_length + 1, len(seq))):
fragments.append(seq[:i])
# Suffixes
for i in range(1, min(max_length + 1, len(seq))):
fragments.append(seq[-i:])
# Windows
for start in range(0, len(seq) - max_length + 1):
fragments.append(seq[start:start + max_length])
return fragments
def decompose_all(sequences, max_length=5):
"""Decompose all sequences in a list.
Args:
sequences: List of lists of symbols
max_length: Maximum fragment length
Returns:
List of fragments (shorter sequences)
"""
all_fragments = []
for seq in sequences:
all_fragments.extend(decompose_sequence(seq, max_length))
return all_fragments
def filter_by_coverage(fragments, min_coverage=0.5):
"""Keep only fragments that appear in at least min_coverage of original sequences.
This ensures we keep patterns that are common, not rare.
"""
from collections import Counter
# Count how many original sequences each fragment appears in
fragment_counts = Counter()
for frag in fragments:
fragment_counts[tuple(frag)] += 1
# Keep fragments that appear frequently enough
min_count = int(len(fragments) * min_coverage)
return [list(frag) for frag, count in fragment_counts.items()
if count >= min_count]
Integration with Pipeline
Add --decompose flag:
parser.add_argument('--decompose', action='store_true',
help='Decompose long sequences before inference')
parser.add_argument('--max-seq-length', type=int, default=5,
help='Maximum sequence length after decomposition')
In _infer_group:
if decompose:
symbol_seqs = decompose_all(symbol_seqs, max_length=max_seq_length)
Expected Benefits
- Shorter sequences: CRX works better on shorter inputs
- Clearer patterns: Decomposition reveals underlying structure
- Fewer flat bags: Long diverse sequences become short uniform ones
Test Plan
- Unit tests:
tests/test_decompose.py - Integration: Compare grammar count with/without decomposition
- Metric:
grammar_structure_score()should improve
Questions to Answer
- Does decomposition actually improve grammar quality?
- What max_length works best?
- How much slower is it?
- Does it help on flat bags specifically?