POST-TRAINING · REWARD-MODEL EVALUATION

Cross-View Contamination Inflates Reward-Model Validation Accuracy

Can a validation split be contaminated when no complete rows match?

Amir Reza Peimani2026Submitted / under review

Related views of the same interaction can cross a train–validation split undetected. Controlled ancestor-inclusion experiments isolate how those relationships inflate reward-model evaluation.

Post-training datasets contain preferences, critiques, edits, and comparisons derived from shared interactions. Checking each view for duplicate rows misses relationships that a learner can exploit.

Approach

Recover field-level ancestry, then manipulate whether a target’s exact edit ancestor appears in training. Compare the resulting gain on exposed targets with ancestry-clean validation and untouched evaluation.

Main result

+14.3 ppWord TF–IDF target-minus-clean effect

In the randomized nested-prefix experiment, ancestor inclusion raised Word TF–IDF exposed-target accuracy from 49.80% to 64.45%. The target-minus-clean effect was 14.26 percentage points (95% CI 10.16–18.36). A separate matched intervention found smaller inflation for character models and no comparable effect for Qwen2.5-1.5B.

Validation inflation as ancestors enter training. Original Figure 2: training-condition comparisons, nested prefixes, controlled effects, and ancestor-arrival timing. Word TF–IDF and frozen MiniLM are shown here; Qwen is evaluated separately in the manuscript.
Validation inflation as ancestors enter training. Original Figure 2: training-condition comparisons, nested prefixes, controlled effects, and ancestor-arrival timing. Word TF–IDF and frozen MiniLM are shown here; Qwen is evaluated separately in the manuscript. Open full resolution ↗
Contamination changes which model validation selects
Figure 4. Exposed-target and ancestry-clean validation rank six reward-model candidates differently and select different winners. The right panel compares their untouched accuracy and discordant contexts.
Figure 4. Exposed-target and ancestry-clean validation rank six reward-model candidates differently and select different winners. The right panel compares their untouched accuracy and discordant contexts. Open full resolution ↗