The German Soofi consortium disclosed in a revised report posted to arXiv on July 22 that its training data contained paraphrased items from the GPQA evaluation benchmark and that it has removed the affected scores. GPQA-Diamond rose 11.1 points during the contaminated training phase, a gain the team can no longer present as evidence of model capability. The finding came from community inspection outside the team and was publicly flagged by Elie Bakouch a week earlier.
The initial release covered on July 19 placed Soofi S at the top of the fully open group on English and German aggregate scores. The model was trained on about 27 trillion tokens using roughly 253,000 B200 GPU-hours at Deutsche Telekom's Munich facility, according to the paper.
What Changed
- Germany's Soofi consortium disclosed in a July 22 arXiv revision that paraphrased GPQA evaluation items entered its training data, and it removed GPQA-Diamond and GPQA-Diamond-DE from every table while withdrawing the report's held-out suites.
- Researcher Elie Bakouch flagged the eval-set training on X on July 15; the team verified the finding against its construction pipeline and confirmed it seven days later.
- The root cause was a mislabeled split: GPQA publishes its entire evaluation set in a single Hugging Face split labeled train, and Soofi's pipeline selected splits by name rather than meaning.
- An audit found the same failure in four benchmarks (GPQA, TruthfulQA, BLiMP, Inverse Scaling); a corrected QA-base dataset is live, and future mixtures get allowlist validation plus n-gram screening.
AI-generated summary, reviewed by an editor. More on our AI guidelines.
Bakouch's July 15 warning
Bakouch posted his finding on X on July 15 after taking a second look at the release. In the post, Bakouch wrote “they literally train on benchmark eval set: gpqa diamond alone”.
Section 4.3 of the revised paper opens with the team's disclosure, “We regret to disclose a benchmark-contamination incident discovered after the initial version of this report.” The Soofi authors confirmed that community members had found paraphrased GPQA evaluation items, including GPQA-Diamond, in both English and machine-translated German. In the revision posted seven days after Bakouch's message, they stated, “We subsequently verified this finding against our construction pipeline and confirmed it.”
How the train label entered QA-base
GPQA has no training split, the report explains. Its complete evaluation set is published as a single split carrying Hugging Face's default train label. Soofi's QA-base dataset was designed to contain paraphrases of training splits from 25 standard benchmarks in English and German. The pipeline selected material by the label's name, and the resulting mixture included the GPQA evaluation items.
The authors acknowledged responsibility in the report, writing, “We state this as an explanation of the failure mode, not as an excuse: verifying that a split labeled train is in fact a training split was our responsibility.”
Get Implicator.ai in your inbox
Strategic AI news from San Francisco. No hype, no "AI will change everything" throat clearing. Just what moved, who won, and why it matters. Daily at 6am PST.
No spam. Unsubscribe anytime.
GPQA-Diamond rises from 32.3 to 43.4
Soofi's checkpoint monitoring showed no abrupt jump when the contaminated data entered the mixture. GPQA-Diamond moved from 32.3 to 43.4 across annealing checkpoints, while HumanEval improved by 18.3 points under the same protocol, a comparison that placed the GPQA rise within the range of legitimate gains on other benchmarks and led the team to mistake it for a genuine capability gain. The authors described the pattern as “slow inflation that is indistinguishable, at the trajectory level, from genuine capability gains.” The team never treated trajectory monitoring as a contamination defense and credited community inspection of the data with exposing the problem.
Know someone who'd find this useful? ✉️ Email it to a friend in one click, or they can subscribe free here.
Four benchmarks in the revised audit
The team's audit of source loaders and data files found the same split-label failure in GPQA, TruthfulQA, BLiMP, and the Inverse Scaling tasks. Each was an evaluation-only benchmark whose sole published split carried a train or validation label. Only GPQA appeared in Soofi's evaluation suite; the other three were disclosed so independently produced scores would not be read as genuine model performance.
Version 3 removed GPQA-Diamond and GPQA-Diamond-DE from all tables and figures and withdrew the initial report's held-out suites. The English and German suite means were recomputed without GPQA and the withdrawn held-out group for all 16 models on the same basis; the revision records that the relative rankings elsewhere in Section 4 were unchanged.
The revision notes that the corrected QA-base dataset is now available on Hugging Face with items derived from all four affected benchmarks removed and that the dataset and model cards document the incident. For future training runs, each source dataset is validated against a per-dataset allowlist derived from its original publication. Final mixtures are screened against the complete evaluation suite through n-gram overlap before training.
Frequently Asked Questions
What did the Soofi consortium disclose?
In a revised arXiv report posted July 22, the consortium disclosed that its training data contained paraphrased items from the GPQA evaluation benchmark in English and machine-translated German. It removed GPQA-Diamond and GPQA-Diamond-DE from all tables and figures and withdrew the initial version's held-out suites.
Who found the contamination?
Community members inspecting the published QA-base dataset. Elie Bakouch posted on X on July 15 that the model had trained on the GPQA-Diamond evaluation set; the Soofi team then verified the finding against its construction pipeline and confirmed it.
How did test data end up in the training mix?
GPQA ships its entire evaluation set as a single Hugging Face split carrying the default label train. Soofi's pipeline selected splits by name rather than semantics, so the evaluation items were ingested as if they were training data.
Did the contamination change Soofi's rankings?
The revised report says no. The English and German suite means were recomputed without GPQA and the withdrawn held-out group for all 16 models on the same basis, and it records the relative rankings as unchanged.
What happens to the dataset now?
A corrected QA-base release is live on Hugging Face with items from all four affected benchmarks removed. The dataset and model cards document the incident, and future training mixtures are validated against per-dataset allowlists and screened via n-gram overlap.
AI-generated summary, reviewed by an editor. More on our AI guidelines.



IMPLICATOR