grammar-inference-engine/docs/adr/0004-frequency-filter-with-min-coverage.md

52 lines
2.3 KiB
Markdown
Raw Normal View History

# 4. Frequency filter with `min_coverage` threshold
**Date:** 2026-07-03
**Status:** Accepted
## Context
A raw method sequence can contain hundreds of unique call tokens, many of which appear in only 1-2 methods. These rare symbols are noise — they inflate grammar size, confuse the inference algorithm, and dilute the signal of common conventions.
We need a principled way to discard rare symbols while keeping the behavioral patterns that define the codebase.
## Decision
Apply a frequency filter *before* clustering: remove any symbol that appears in fewer than `min_coverage` fraction of method sequences.
Default threshold: `0.2` (20% of methods must contain the symbol).
Filtering is done by `frequency_filter()` in `analyze.py`:
```python
n_files = len(sequences)
threshold = max(1, int(n_files * min_coverage))
symbol_file_count = Counter()
for seq in sequences:
seen = set()
for _, text, _ in seq:
symbol_file_count[text] += 1 if text not in seen else 0
seen.add(text)
keep = {text for text, count in symbol_file_count.items() if count >= threshold}
```
A symbol is counted once per file (not once per occurrence) to avoid skew from files that repeat the same symbol many times.
## Consequences
**Positive:**
- Removes noise before clustering, improving cluster quality.
- Prevents rare one-off function calls from creating spurious n-gram matches.
- The `common` vs `other` cluster distinction is sharper because the filter removes tokens that would appear in neither.
**Negative:**
- With large codebases (600+ methods), 20% threshold may be too aggressive — a symbol needs 120+ occurrences to survive.
- Mitigation: `--min-coverage` flag lets users tune per codebase.
- Production code often needs lower values (`0.05`) because methods are more diverse than tests.
- The threshold is relative, not absolute. A 3-file project keeps anything in 1+ files (`max(1, 3*0.2)` = 1).
## Alternatives Considered
- **No filter**: CRX produces `(a+b+c+d+e+f+g+h+i+j+...)+` — the vocabulary is too large to be informative.
- **Absolute threshold**: `min_occurrences=5`. Doesn't scale — works for small projects, wrong for large ones.
- **TF-IDF style weighting**: More sophisticated but adds complexity. The simple coverage filter works well in practice.