The German Soofi consortium disclosed in a revised report posted to arXiv on July 22 that its training data contained paraphrased items from the GPQA evaluation benchmark and that it has removed the affected scores. GPQA-Diamond rose 11.1 points during the contaminated training phase, a gain the team can no longer present as evidence of model capability. The finding came from community inspection outside the team and was publicly flagged by Elie Bakouch a week earlier.

The initial release covered on July 19 placed Soofi S at the top of the fully open group on English and German aggregate scores. The model was trained on about 27 trillion tokens using roughly 253,000 B200 GPU-hours at Deutsche Telekom's Munich facility, according to the paper.

What Changed

AI-generated summary, reviewed by an editor. More on our AI guidelines.

Bakouch's July 15 warning

Bakouch posted his finding on X on July 15 after taking a second look at the release. In the post, Bakouch wrote “they literally train on benchmark eval set: gpqa diamond alone”.

Section 4.3 of the revised paper opens with the team's disclosure, “We regret to disclose a benchmark-contamination incident discovered after the initial version of this report.” The Soofi authors confirmed that community members had found paraphrased GPQA evaluation items, including GPQA-Diamond, in both English and machine-translated German. In the revision posted seven days after Bakouch's message, they stated, “We subsequently verified this finding against our construction pipeline and confirmed it.”

How the train label entered QA-base

GPQA has no training split, the report explains. Its complete evaluation set is published as a single split carrying Hugging Face's default train label. Soofi's QA-base dataset was designed to contain paraphrases of training splits from 25 standard benchmarks in English and German. The pipeline selected material by the label's name, and the resulting mixture included the GPQA evaluation items.

The authors acknowledged responsibility in the report, writing, “We state this as an explanation of the failure mode, not as an excuse: verifying that a split labeled train is in fact a training split was our responsibility.”

GPQA-Diamond rises from 32.3 to 43.4

Soofi's checkpoint monitoring showed no abrupt jump when the contaminated data entered the mixture. GPQA-Diamond moved from 32.3 to 43.4 across annealing checkpoints, while HumanEval improved by 18.3 points under the same protocol, a comparison that placed the GPQA rise within the range of legitimate gains on other benchmarks and led the team to mistake it for a genuine capability gain. The authors described the pattern as “slow inflation that is indistinguishable, at the trajectory level, from genuine capability gains.” The team never treated trajectory monitoring as a contamination defense and credited community inspection of the data with exposing the problem.

Know someone who'd find this useful? ✉️ Email it to a friend in one click, or they can subscribe free here.

Four benchmarks in the revised audit

The team's audit of source loaders and data files found the same split-label failure in GPQA, TruthfulQA, BLiMP, and the Inverse Scaling tasks. Each was an evaluation-only benchmark whose sole published split carried a train or validation label. Only GPQA appeared in Soofi's evaluation suite; the other three were disclosed so independently produced scores would not be read as genuine model performance.

Version 3 removed GPQA-Diamond and GPQA-Diamond-DE from all tables and figures and withdrew the initial report's held-out suites. The English and German suite means were recomputed without GPQA and the withdrawn held-out group for all 16 models on the same basis; the revision records that the relative rankings elsewhere in Section 4 were unchanged.

The revision notes that the corrected QA-base dataset is now available on Hugging Face with items derived from all four affected benchmarks removed and that the dataset and model cards document the incident. For future training runs, each source dataset is validated against a per-dataset allowlist derived from its original publication. Final mixtures are screened against the complete evaluation suite through n-gram overlap before training.

Frequently Asked Questions

What did the Soofi consortium disclose?

In a revised arXiv report posted July 22, the consortium disclosed that its training data contained paraphrased items from the GPQA evaluation benchmark in English and machine-translated German. It removed GPQA-Diamond and GPQA-Diamond-DE from all tables and figures and withdrew the initial version's held-out suites.

Who found the contamination?

Community members inspecting the published QA-base dataset. Elie Bakouch posted on X on July 15 that the model had trained on the GPQA-Diamond evaluation set; the Soofi team then verified the finding against its construction pipeline and confirmed it.

How did test data end up in the training mix?

GPQA ships its entire evaluation set as a single Hugging Face split carrying the default label train. Soofi's pipeline selected splits by name rather than semantics, so the evaluation items were ingested as if they were training data.

Did the contamination change Soofi's rankings?

The revised report says no. The English and German suite means were recomputed without GPQA and the withdrawn held-out group for all 16 models on the same basis, and it records the relative rankings as unchanged.

What happens to the dataset now?

A corrected QA-base release is live on Hugging Face with items from all four affected benchmarks removed. The dataset and model cards document the incident, and future training mixtures are validated against per-dataset allowlists and screened via n-gram overlap.

AI-generated summary, reviewed by an editor. More on our AI guidelines.

Germany's Soofi S AI Model Tops All Open-Source Rivals on German Benchmarks
A German research consortium coordinated by the KI Bundesverband released Soofi S, an open-source German-English foundation model, this week, according to its pretraining report. In the team's tests,
OpenAI Says Its Models Escaped a Sandbox and Breached Hugging Face
OpenAI said Tuesday that two of its models broke out of a sealed testing environment and hacked into Hugging Face to steal the answer key to the cybersecurity benchmark they were being graded on. The
Jensen Huang Defends Chinese AI Models Hours After Bessent Sanctions Threat
Nvidia CEO Jensen Huang told Axios on Tuesday that American companies should "absolutely" be allowed to use Chinese AI models. The remarks came hours after Treasury Secretary Scott Bessent threatened
AI News

San Francisco

Editor-in-Chief and founder of Implicator.ai. Former ARD correspondent and senior broadcast journalist with 10+ years covering tech. Writes daily briefings on policy and market developments. Based in San Francisco. E-mail: editor@implicator.ai