CrossFit reduces false agreement in self-evolving search agents by changing which solver rewards the question proposer. In experiments with Qwen3.5-4B and Qwen3.5-9B, it lowers final-round false-agreement mass from 6.1% to 3.0% and from 8.8% to 3.7%, respectively, while improving average Cover-EM across seven search benchmarks.
The failure the authors call co-cheating arises when a proposer creates questions and pseudo-labels, then a solver trained on those labels evaluates later proposals from the same sources. The two agents can agree on an incorrect answer, raising the internal reward without improving external correctness. A post-hoc audit finds this false agreement grows over successive rounds.
How the audit distinguishes agreement from correctness
The audit checks saved source documents, adopted pseudo-labels, and five solver responses used to calculate proposer reward. An external judge constructs evidence-backed references from the source and assesses those saved outputs; unresolved cases are left unresolved rather than forced into a correct-or-incorrect label. Because the auditor does not affect training or reward, it measures the loop without changing it.
A quick word from our friends at Comparitech…
Comparitech – Cybersecurity and privacy without the noise.
Comparitech covers the latest in cybersecurity, privacy, online threats, and the technology affecting how people and businesses protect themselves online.
The key diagnostic is false-agreement mass: the fraction of evaluated label-response pairs that match each other but are wrong according to the source-based audit. The authors distinguish this from lost credit, where a wrong label disagrees with a correct solver response. In round one, false-agreement mass is 0.004 for Qwen3.5-4B and 0.003 for Qwen3.5-9B; by round three it reaches 0.061 and 0.088.
The audit suggests that rising agreement alone is misleading: agreement becomes more optimistic while label and solver correctness fail to rise in step. The authors describe this as shared mistakes replacing disagreement, rather than intentional coordination. The auditor is an LLM judge, and the paper notes that automated judges can have systematic biases; its evidence-based procedure leaves unsupported cases unresolved.
MSV checks labels; CrossFit changes feedback ancestry
Multi-sample verification (MSV) tests each proposed question with three answers generated using its source and three generated without the source. If both sets produce compatible majorities, MSV admits the task and replaces the proposer’s draft label with the consensus; otherwise it rejects the task. The two views test source support and answer stability, but six samples from one model can still share errors.
CrossFit instead divides source documents into two folds and keeps every question from a document in its assigned fold. An auxiliary solver trained on fold A evaluates questions from fold B, while a solver trained on B evaluates questions from A. This source-level exclusion prevents the feedback solver from having trained on pseudo-labels derived from the source it scores; the main solver still trains on all admitted questions.
The distinction from Dr. Zero is the provenance of proposer feedback, not the reward rule or the main solver’s training set. MSV changes which question-label pairs enter training, while CrossFit changes which solver scores proposals. The method borrows the exclusion principle of statistical cross-fitting, but the authors explicitly do not claim its asymptotic guarantees for this adaptive curriculum. Unlike methods that train a proposer to generate counterfactual evidence edits, CrossFit keeps the source documents and changes which solver evaluates each source fold.
Benchmark gains and what the controls establish
On a fixed suite of 1,325 questions from Natural Questions, TriviaQA, PopQA, HotpotQA, 2WikiMultiHopQA, MuSiQue, and Bamboogle, CrossFit improves average Cover-EM over coupled self-evolution by 8.8 points at 4B and 8.4 at 9B. Gains are larger on the four multi-hop benchmarks, averaging 10.0 and 10.9 points, than on the single-hop datasets, at 7.3 and 5.2 points.
MSV alone reduces false-agreement mass from 6.1% to 5.7% at 4B and from 8.8% to 7.2% at 9B, but leaves substantial residual co-cheating. CrossFit lowers it further to 3.0% and 3.7%; combining both interventions reaches 2.0% and 1.7%. The combined method adds only 0.3 Cover-EM points over CrossFit alone at each scale. The modest gain from MSV alone contrasts with oracle-grounded self-evolution, which uses an external oracle to check causal reasoning steps rather than excluding source ancestry in search feedback.
Controls test whether source exclusion, rather than simply adding an evaluator, explains the effect. Same-source and full-data auxiliary solvers remain close to coupled feedback, while randomly splitting questions is less effective than splitting by source ID. In fixed-bank replay, source-ID feedback cuts false agreement from 0.058 and 0.073 to 0.004 and 0.001 at 4B and 9B, holding proposals and labels fixed.
Limits and implications for self-evolution
CrossFit blocks direct reuse of source-derived pseudo-labels in the feedback solver, but it is not a truth oracle. The authors acknowledge that auxiliary solvers can still share errors from pretraining or overlapping evidence. MSV also has a concrete cost: six additional labeler generations per candidate, including search and coordination, while only partially reducing false agreement. A separate approach uses verifier-grounded benchmark coevolution to evolve Lean proof workflows; CrossFit instead addresses false agreement in search-agent curricula.
The results position CrossFit as a targeted change to the feedback path in self-generated curricula, rather than a replacement for search or answer verification. Compared with Dr. Zero’s coupled feedback, it improves both the audit’s correctness measures and downstream search performance; compared with MSV, it addresses the training ancestry of the evaluator. The authors identify end-to-end efficiency and exclusion across connected sources as open tests of the approach. A related line of work, evidence-verifiable self-evolution, focuses on whether answers are supported by evidence, whereas CrossFit targets the provenance of proposer feedback. The broader concern that internally generated supervision can plateau is also examined through a closed-loop generalization gap, though this paper measures shared errors in search-agent feedback.


