grammar-inference-engine/experiments/PHASE2_PLAN.md
tobjend 8b2899d16e Phase 2: decomposition forest for complex sequences
Decomposition breaks long sequences into shorter fragments before
inference. This helps when sequences are too long for CRX to handle
(>5 symbols → flat bags).

Results:
- RAGSAK: 21 → 80 grammars (3.8× increase)
- FastAPI: 111 → 118 grammars (small increase)

Changes:
- bex/decompose.py: decompose_sequence(), decompose_all(), decompose_with_coverage()
- bex/tag_preprocessor/analyze.py: --decompose, --max-seq-length flags
- Skip diversity check when decomposing (decomposition creates diverse fragments)
- 12 new tests in tests/test_decompose.py

Co-authored-by: OpenCode <opencode@corentic.eu>
2026-07-12 17:49:35 +02:00

4.4 KiB

Phase 2: Decomposition Forest

Goal

Break down complex/long behavioral sequences into shorter ones that still capture the pattern. This helps when sequences are long and diverse, making CRX produce flat bags.

Current Problem

Long sequences like:

["if", "return", "if", "return", "if", "return"]

CRX sees 6 symbols, tries to find a pattern, often produces flat bags like (if|return)*.

If we decompose into shorter examples:

["if", "return"]
["if", "return"]
["if", "return"]

CRX sees a clear pattern: if.return (repeated).

Crucio's Approach

Crucio uses three decomposition strategies:

  1. Binary maximum subsequence deletion: Split in half, delete max from each half
  2. Maximum subsequence deletion: Delete largest contiguous chunk
  3. Subsequence replacement: Replace a chunk with a shorter version

Key insight: Decomposed sequences must preserve grammar coverage (be valid under the same grammar).

Our Adaptation

For behavioral sequences, we need simpler decomposition:

  1. Prefix extraction: Take first N symbols
  2. Suffix extraction: Take last N symbols
  3. Window extraction: Take middle N symbols
  4. Pattern extraction: Find repeated patterns and extract one instance

Implementation Plan

File: bex/decompose.py (new)

"""Decomposition forest for behavioral sequences.

Inspired by Crucio's decomposition forest (ICSE 2026).
Breaks down long sequences into shorter ones that preserve patterns.
"""

def decompose_sequence(seq, max_length=5):
    """Decompose a sequence into shorter fragments.
    
    Strategies:
    1. If seq <= max_length, return as-is
    2. Extract prefixes of length 1..max_length
    3. Extract suffixes of length 1..max_length
    4. Extract windows of length max_length
    
    Args:
        seq: List of symbols
        max_length: Maximum fragment length
        
    Returns:
        List of fragments (shorter sequences)
    """
    if len(seq) <= max_length:
        return [seq]
    
    fragments = []
    
    # Prefixes
    for i in range(1, min(max_length + 1, len(seq))):
        fragments.append(seq[:i])
    
    # Suffixes
    for i in range(1, min(max_length + 1, len(seq))):
        fragments.append(seq[-i:])
    
    # Windows
    for start in range(0, len(seq) - max_length + 1):
        fragments.append(seq[start:start + max_length])
    
    return fragments


def decompose_all(sequences, max_length=5):
    """Decompose all sequences in a list.
    
    Args:
        sequences: List of lists of symbols
        max_length: Maximum fragment length
        
    Returns:
        List of fragments (shorter sequences)
    """
    all_fragments = []
    for seq in sequences:
        all_fragments.extend(decompose_sequence(seq, max_length))
    return all_fragments


def filter_by_coverage(fragments, min_coverage=0.5):
    """Keep only fragments that appear in at least min_coverage of original sequences.
    
    This ensures we keep patterns that are common, not rare.
    """
    from collections import Counter
    
    # Count how many original sequences each fragment appears in
    fragment_counts = Counter()
    for frag in fragments:
        fragment_counts[tuple(frag)] += 1
    
    # Keep fragments that appear frequently enough
    min_count = int(len(fragments) * min_coverage)
    return [list(frag) for frag, count in fragment_counts.items() 
            if count >= min_count]

Integration with Pipeline

Add --decompose flag:

parser.add_argument('--decompose', action='store_true',
                    help='Decompose long sequences before inference')
parser.add_argument('--max-seq-length', type=int, default=5,
                    help='Maximum sequence length after decomposition')

In _infer_group:

if decompose:
    symbol_seqs = decompose_all(symbol_seqs, max_length=max_seq_length)

Expected Benefits

  1. Shorter sequences: CRX works better on shorter inputs
  2. Clearer patterns: Decomposition reveals underlying structure
  3. Fewer flat bags: Long diverse sequences become short uniform ones

Test Plan

  1. Unit tests: tests/test_decompose.py
  2. Integration: Compare grammar count with/without decomposition
  3. Metric: grammar_structure_score() should improve

Questions to Answer

  1. Does decomposition actually improve grammar quality?
  2. What max_length works best?
  3. How much slower is it?
  4. Does it help on flat bags specifically?