POST-TRAINING · REWARD-MODEL EVALUATION
Cross-View Contamination Inflates Reward-Model Validation Accuracy
Can a validation split be contaminated when no complete rows match?
Related views of the same interaction can cross a train–validation split undetected. Controlled ancestor-inclusion experiments isolate how those relationships inflate reward-model evaluation.
Post-training datasets contain preferences, critiques, edits, and comparisons derived from shared interactions. Checking each view for duplicate rows misses relationships that a learner can exploit.
Approach
Recover field-level ancestry, then manipulate whether a target’s exact edit ancestor appears in training. Compare the resulting gain on exposed targets with ancestry-clean validation and untouched evaluation.
Main result
In the randomized nested-prefix experiment, ancestor inclusion raised Word TF–IDF exposed-target accuracy from 49.80% to 64.45%. The target-minus-clean effect was 14.26 percentage points (95% CI 10.16–18.36). A separate matched intervention found smaller inflation for character models and no comparable effect for Qwen2.5-1.5B.

Contamination changes which model validation selects
