Aleph Alpha
Abstract pink gradient with translucent stepped blocks on the left and a curved orange wave sweeping through the right
Research

Fabian Gebhart

Generalization or Memorization? What Contaminated Training Data Does to Benchmark Scores

TL;DR


Introduction

How to detect contamination?

Evaluation instance

A bakery sells croissants in boxes of 6. On Monday it sold 14 boxes, and on Tuesday it sold 9 more boxes than on Monday. How many croissants did the bakery sell on Tuesday? Show your work and answer with a single number.

Answer 138

inside a matched 5-gram answer within 50 tokens

Question overlap1.00 × 0.75 = 0.75
39 of 39 unique 5-grams matched, weighted by IDF; 5 boilerplate 5-grams count zero
Answer nearby1.00 × 0.25 = 0.25
exact answer found within 50 tokens of the match
Weighted sum1.00
Length rule: 44 instance tokens need 0.84, so the score is scaled by 0.951.00
1.00Flagged: document dropped

Every question 5-gram matches and the answer sits a few tokens after the question. The question part contributes 0.75, the answer 0.25: a perfect 1.0, and the document is dropped.

Benchmark text is everywhere

Two stacked horizontal bars. The top bar shows the scanned pool: about a third instruction and QA data, a third source code, a third web crawl, with small slices of mathematics and other sources. The bottom bar shows the removed documents: two thirds instruction and QA data, then source code, mathematics and web crawl at about a tenth each.
Horizontal bar chart of ten evaluation suites by contamination matches. GSM8K and MATH in bits-per-byte format lead, followed by MMLU and MMLU-Pro. ARC and HellaSwag follow at a distance, then MMLU-Pro with chain of thought, PIQA, TriviaQA and Global MMLU.

Benchmark ARC_OLMES (train split)Score 1.00Answer found yes

Evaluation instance
Question: A single prokaryotic cell can divide several times in an hour. Few eukaryotic cells can divide as quickly. Which of the following statements best explains this difference?
Expected answer
Eukaryotic cells are more structurally complex than prokaryotic cells.
Training document (1,434 characters)
A single prokaryotic cell can divide several times in an hour. Few eukaryotic cells can divide as quickly. Which of the following statements best explains this difference? A. Eukaryotic cells are smaller than prokaryotic cells. B. Eukaryotic cells have less DNA than prokaryotic cells. C. Eukaryotic cells perform more specialized functions than prokaryotic cells. D. Eukaryotic cells are more structurally complex than prokaryotic cells. Answer: D **Explanation:** The correct answer is D because eukaryotic cells are more structurally complex than prokaryotic cells, which contributes to their slower division rate. Prokaryotic cells (e.g., bacteria) have a simpler structure, lack a nucleus and membrane-bound organelles, and replicate their single circular chromosome quickly. In contrast,[… 631 characters …]

shared passage matched answer [… n characters …] text without a match

Does it matter?

Expected appearances of a removed document, by size of the mid-training mix

Appearances per removed document
Arm Data Tokens in the mix: total / contaminated Appearances per removed document
DecontaminatedThe decontaminated proxy mix90B / 00
Naturally contaminatedThe same data mix before decontamination1T / about 0.04%0.007
Target-scale exposureThe decontaminated mix plus the removed documents, each as often as the production run would have seen it90B / 1.5B (1.6%)0.79 on average, as in the production run
Fully contaminatedThe decontaminated mix plus every removed document, once90B / 7.4B (8.2%)up to 1, since the runs stop short of the full mix
Four line charts of accuracy over training steps 500 to 5,000 for the four arms: HumanEval, MMLU, TriviaQA and GSM8K. The decontaminated and naturally contaminated lines track each other, apart from single-checkpoint swings. The fully contaminated line sits far above from the first evaluation on HumanEval and TriviaQA and climbs on MMLU. The target-scale exposure line sits between, close to the fully contaminated one on GSM8K.
Gain over the decontaminated arm at the end of the proxy runs, in points, on the questions the arms have seen (flagged by the detector or present verbatim in the removed documents) and on the questions they have never seen. Gains are averaged over questions. The MMLU gains in the text average over its 57 subjects and over the last five evaluations instead; the subject average weighs small subjects more. ± is the standard error of the mean gain.
Benchmark Questions seen / never seen Target-scale exposure Fully contaminated
MMLU13,052 / 815+6.1 ±0.4 / +3.2 ±1.6+24.4 ±0.4 / +20.2 ±1.6
TriviaQA6,963 / 1,030+16.5 ±0.6 / +13.1 ±1.5+28.1 ±0.6 / +25.6 ±1.5

What decontamination did not catch

Line chart over tokens seen, from 12 to 23.4 trillion. pass@1 (solid) and the share of completions that recite the reference solution (dashed) for pre-training in cream and mid-training in red. In pre-training both lines swing together between 0.3 and 0.95. At the switch to mid-training the recitation share drops to about 0.1 and pass@1 settles at 0.7. The share then spikes twice to almost 0.5 with pass@1 following to 0.85. The two pass@1 peaks are ringed and labeled spike 1 and spike 2.

Conclusion

Acknowledgements