4.1 KiB
13. Replace MDL scoring with Language Size measure
Date: 2026-07-11
Status: Accepted
Context
Our ensemble grammar inference uses a scoring function to select the best grammar when multiple algorithms (CRX, iDRegEx, kORE) produce candidates. The scoring function determines which grammar wins — so a bad scoring function means the ensemble picks bad grammars, even when the algorithms produce good ones.
The concrete problem
On real codebases, the ensemble would pick overly generic grammars like info+
over specific ones like info.file.template.shell.service+. The generic grammar
accepts an astronomically large language (every permutation of info at every
length), while the specific grammar accepts exactly one word of each length.
Any human would pick the specific one. Our scorer picked the generic one.
Why MDL failed
We used Minimum Description Length (MDL): score = model_cost + data_cost,
where model_cost = len(expr) and data_cost = Σ log₂(|L(r)| at seq length).
The problem: model_cost('info+') = 1 (one symbol occurrence), while
model_cost('info.file.template.shell.service+') = 5. MDL rewards short
expressions. When model_cost is small relative to data_cost, short generic
patterns win even though they accept far more spurious words.
The paper that fixed it
Bex, Goethals, Penninckx, Van Gucht, and Van den Bussche published "Learning Deterministic Regular Expressions for the Inference of Schemas from XML Data" (arXiv:1004.2372, also VLDB 2007). They proposed the iDRegEx algorithm and evaluated two scoring measures:
- Language Size (Section 4.3.1): select the expression that accepts the
fewest words up to length
n = 2m + 1. This directly measures specificity. - MDL (Section 4.3.2): model cost + data cost, based on Adriaans & Vitányi (2006). This rewards short expressions.
Their results (Section 5, line 1470): on a corpus of synthetic regular expressions with alphabet size 5, Language Size achieved 98% success rate while MDL achieved only 21%. They explicitly abandoned MDL: "Therefore in the remainder of this section we only consider iDRegEx with the language size criterion."
Decision
Replace MDL with Language Size as the default scoring function. Keep MDL as a
fallback enabled via scoring_method='mdl' parameter.
Implementation
def lang_size_score(expr, sequences):
"""Language Size: Σ |L(r)|_len(seq) — words at each sequence length.
From Bex et al. (arXiv:1004.2372), Section 4.3.1.
Lower is better — the grammar that accepts the fewest words wins.
"""
if not sequences:
return lang_size(expr, 2 * model_cost(expr) + 1)
total = 0
for seq in sequences:
total += _count_words_fast(expr, len(seq))
return total
Counts words at exactly the lengths present in the input sequences, not all lengths 0..n. This is a practical adaptation of the paper's measure: the paper evaluates all candidates at the same fixed n, but our candidates have different n values. Counting at observed sequence lengths gives the same result — the grammar accepting the fewest words at the relevant lengths wins.
The existing _count_words_fast function already computes the exact word count
needed by this measure. No new algorithms required — just a different aggregation.
Consequences
Positive:
- Ensemble picks specific grammars over generic ones (98% vs 21% on Bex's corpus)
_count_words_fastalready exists and is LRU-cached — zero new code for the hard partmodel_coststill used for tie-breaking (shortest expression when language sizes equal)
Negative:
lang_sizecan be expensive for expressions with large alphabets (exponential in worst case), butn = 2*model_cost + 1keeps it bounded in practice- Old MDL results stored in YAML will have different scores than new runs — not a compatibility issue since scores are internal selection criteria, not persisted
Migration:
mdl_scorepreserved as fallback:scoring_method='mdl'mdl_score_simplein ensemble.py updated to uselang_size_scoreby default- CLI flag
--scoring-methodcontrols which is used (default:langsize)