grammar-inference-engine/experiments/PHASE1_PLAN.md
tobjend 841f5efcf5 feat: distributional clustering for better grouping (Phase 1)
Inspired by Crucio's distributional matrix (ICSE 2026).
Groups symbols by context similarity instead of first-symbol.

New module: bex/distributional.py
- extract_contexts(): get (left, right) context pairs
- build_distribution_matrix(): symbol × context counts
- cluster_symbols(): agglomerative clustering by similarity
- split_by_cluster(): split sequences by cluster membership
- distributional_split(): drop-in replacement for _split_by_first_symbol

CLI: --cluster-method distributional (opt-in, first-symbol is default)

Example: 'if' and 'while' both appear before 'return'
→ first-symbol: 3 groups (if, while, return)
→ distributional: 2 groups (if/while merged, return)

23 tests pass. Full suite: 257 tests pass.
2026-07-12 17:37:36 +02:00

169 lines
4.6 KiB
Markdown

# Phase 1: Distributional Clustering
## Goal
Replace `_split_by_first_symbol()` with smarter clustering based on context similarity.
This is inspired by Crucio's distributional matrix but adapted for behavioral sequences.
## Current Problem
Our current `_split_by_first_symbol()` is too crude:
```python
# Current: split by first symbol
groups[seq[0]].append(seq)
# Problem: "if return" and "return if" look different but might be equivalent
# if they appear in similar contexts
```
## Crucio's Insight
Symbols that appear in the same contexts are distributionally equivalent:
```
Symbol "a" appears in contexts: [_, b], [c, _], [_, _]
Symbol "b" appears in contexts: [a, _], [_, c], [_, _]
→ If "a" and "b" have same context distribution, they're equivalent
```
## Our Adaptation
### Step 1: Context Extraction
For each symbol in all sequences, extract context pairs:
```python
# Example sequences:
# ["if", "return", "if", "return"]
# ["return", "if", "return"]
# Contexts for "return":
# - ("if", "if") at position 1
# - ("if", None) at position 3
# - (None, "if") at position 0
# Contexts for "if":
# - (None, "return") at position 0
# - ("return", "return") at position 1
# - ("return", None) at position 2
```
### Step 2: Distribution Vector
For each symbol, create a distribution vector:
```python
# Symbol: "return"
# Context distribution: {("if", "if"): 1, ("if", None): 1, (None, "if"): 1}
# Symbol: "if"
# Context distribution: {(None, "return"): 1, ("return", "return"): 1, ("return", None): 1}
```
### Step 3: Similarity Measure
Compare context distributions using cosine similarity or Jaccard:
```python
def context_similarity(sym1_contexts, sym2_contexts):
# Compare the sets of contexts
# High similarity → symbols are distributionally equivalent
```
### Step 4: Clustering
Group symbols with high similarity:
```python
# Cluster 1: ["if", "while", "for"] (conditional contexts)
# Cluster 2: ["return", "yield"] (return contexts)
# Cluster 3: ["class", "def"] (definition contexts)
```
### Step 5: Split Sequences by Cluster
Replace first-symbol split with cluster-based split:
```python
# Instead of:
# groups[seq[0]].append(seq)
# Do:
# cluster = symbol_to_cluster[seq[0]]
# groups[cluster].append(seq)
```
## Implementation Plan
### File: `bex/distributional.py` (new)
```python
"""Distributional clustering for behavioral sequences.
Inspired by Crucio's distributional matrix (ICSE 2026).
Groups symbols by context similarity for better sequence classification.
"""
def extract_contexts(sequences):
"""Extract context pairs for each symbol."""
# For each symbol, collect (left_context, right_context) pairs
pass
def build_distribution_matrix(contexts):
"""Build distribution matrix from contexts."""
# Rows = symbols, Columns = unique contexts
# M[i,j] = count of symbol i in context j
pass
def cluster_symbols(distribution_matrix, threshold=0.7):
"""Cluster symbols by distribution similarity."""
# Use cosine similarity or Jaccard
# Return dict: symbol → cluster_id
pass
def split_by_cluster(sequences, clusters):
"""Split sequences by first symbol's cluster."""
# Instead of first-symbol, use cluster membership
pass
```
### File: `bex/tag_preprocessor/analyze.py` (modify)
Add `--cluster-method` flag:
```python
# New flag
parser.add_argument('--cluster-method',
choices=['first-symbol', 'distributional'],
default='first-symbol',
help='Method to split mixed groups')
# In _recursive_split():
if cluster_method == 'distributional':
clusters = cluster_symbols(sequences)
return split_by_cluster(sequences, clusters)
else:
return _split_by_first_symbol(sequences)
```
## Expected Benefits
1. **Better grouping**: Symbols with same context → same group
2. **More general**: Handles cases where first symbol varies
3. **Still fast**: Distributional clustering is O(n * k) where n = symbols, k = contexts
## Test Plan
1. **Unit tests**: `tests/test_distributional.py`
- Test context extraction
- Test distribution matrix
- Test clustering
- Test sequence splitting
2. **Integration test**: Run on RAGSAK with `--cluster-method distributional`
- Compare grammar count and quality vs first-symbol
3. **Evaluation metric**: `grammar_structure_score()` and `lang_size_score()`
- Higher structure = better grouping
- Tighter grammars = better patterns captured
## Questions to Answer
1. Does distributional clustering actually improve grammar quality?
2. What similarity threshold works best?
3. How much slower is it than first-symbol?
4. Does it help on flat bags specifically?