A paper by five Sakana AI researchers reports that its Multi-Layered Review system, run as an ensemble of four reviews, caught 73.43% of contradictions planted in a paper’s main-claim statements. The paper was posted to arXiv on 8 October 2026 and is marked as published in Transactions on Machine Learning Research in July 2026; MarkTechPost covered it on 10 October.
The errors were synthetic: the authors inserted contradictions into 257 published conference papers and measured which review systems spotted them. On real retracted papers the same system scored lower, and the paper says its lead there is “less pronounced”.
How the Benchmark Works
The authors built what they call a Contradiction Benchmark. They collected 257 papers published at ACL, AISTATS, CVPR and ICML in 2025 and NeurIPS in 2024, restricted to papers under permissive Creative Commons licences.
For each paper an LLM built a knowledge graph of its claims and evidence, and the pipeline inserted a contradiction at different distances from a “main claim” node. The paper defines distance 0 as a main-claim node, and says distances “are defined to be inversely related to contradiction severity”. Table 1 lists between 209 and 242 contradictions per venue.
The paper describes the motivation as a “verification-centric perspective on LLM-assisted peer review, emphasizing error detection as a critical and resource-intensive task.”

Business Pill · VERIFICATION COST
A one-minute explainer of verification cost: why output that is cheap to make can still be expensive to check. It teaches the general idea only and says nothing about any company in this story.
The key insight: As we read it, the paper measures review systems on what a reviewer has to catch rather than on how closely they imitate human reviews: on planted main-claim contradictions, its four-review system caught 73.43%, and on real retracted papers 26.07% and 16.11% under two judging settings.
What Sakana Built
The Multi-Layered Review system has three roles: an Appendix Agent, a Literature Review Agent and a Review Agent. The paper says the appendix is summarised by Claude Haiku 3.5, the literature review runs on Claude Sonnet 4 with web search, and the review is produced through a three-pass prompt chain: an outline, a detailed critique, then a synthesis.
For the benchmark, MLR used only the Review Agent. The comparison systems were LLM-Review, the AI Reviewer and AgentReview, which the paper says were chosen on “code availability, system diversity, and cost considerations.”
The Results on Planted Errors
Table 14 gives the main-claim figures. MLR caught 73.43% of distance-0 contradictions with four reviews and 60.79% with one. The baselines caught 14.81% (AgentReview), 14.56% (LLM-Review) and 11.17% (AI Reviewer).
Across all distances the paper reports 40.95% for MLR with four reviews, against 5.95% to 6.50% for the three baselines. In its text the paper says MLR “detects more than 70% of distance 0 contradictions—corresponding to a fourfold improvement over the next strongest baseline.”
The paper flags a confound itself: the baselines use OpenAI GPT models, while MLR uses Claude. In an ablation that gave LLM-Review the same Claude model, distance-0 accuracy rose to 35.40%. The authors conclude that the model choice and the system design both contribute meaningfully.

Real Retracted Papers
The authors also tested the systems on the WithdrarXiv-Check dataset of papers withdrawn from arXiv, with their retraction comments. With a judge counting a review as a hit when it raised a similar problem, MLR scored 26.07% against 18.48% for AgentReview, 13.74% for the AI Reviewer and 5.21% for LLM-Review. Under an exact-match judge MLR scored 16.11%, against 9.00% for the AI Reviewer.
The paper notes that “all review systems, including MLR, perform worse on real mistakes than on our benchmark”, and calls the Contradiction Benchmark “an initial step toward evaluating the error detection capabilities of LLM reviewers.”
Manipulation, Cost and Users
The authors injected text designed to sway LLM reviewers into 50 rejected papers. They report that “all systems are strongly influenced by the injected text”, and that in 8 of the 50 cases MLR detected the manipulative text.
On cost, the paper puts MLR at about $0.42 in input tokens and $0.05 in output tokens per review, and says MLR and the AI Reviewer “both of which incur about USD 0.50 per review.”
In a user study with 38 review sessions, participants marked 305 of the 378 review comments they gave feedback on as correct, which the paper reports as 81%.
The Structural Read
The paper reports two sources of its gain. Swapping the baseline LLM-Review onto the same Claude model lifted its main-claim accuracy from 14.56% to 35.40%, and Sakana’s multi-pass design took it to 60.79% with a single review.
The benchmark is synthetic by design. The paper says the inserted errors give “unambiguous evaluation targets”, and its own audit found some of them read less naturally than genuine mistakes.
As we read it, the comparison that matters for anyone checking AI-assisted work is the gap between the planted-error and real-paper results, the same gap seen when AI agents produce scientific output that someone still has to verify.
Teo, Yamada, Kotyan, Imajuku and Clanuwat, Sakana AI, arXiv 2610.11087
“these systems should be used strictly as assistive tools to complement careful human review.”
Three Implications
CONFERENCE ORGANISERS The paper reports that all systems tested were strongly influenced by text injected to sway LLM reviewers, with MLR detecting it in 8 of 50 papers.
RESEARCH TEAMS The paper puts a review from MLR or the AI Reviewer at about USD 0.50, and reports that participants in its user study agreed with 305 of 378 comments they gave feedback on.
ANYONE READING BENCHMARK CLAIMS The 73.43% is a main-claim figure on synthetic errors with four reviews; the full-benchmark figure is 40.95% and the real-paper figures are 26.07% and 16.11%.
The Business Engineer Lens
This story maps onto the Business Engineer framework AI & The Harness Theory.
The analysis describes the next step as “an operating system that wraps your AI”, a harness that “compresses your thinking, your team’s feedback, and the model’s raw capability into something replicable.”
As we read it, the paper is a measured case of that idea: the same Claude model scored 35.40% on main-claim errors inside a simple reviewer and 60.79% inside Sakana’s three-pass design, before ensembling.
What Is Not Established
The paper says its analysis is “primarily limited to the machine learning domain” and that it could not compare against closed-source or proprietary review systems. It says data and code are “available upon request.”
In a human audit of 50 planted contradictions, the paper reports that 34% were rated as lacking coherence across adjacent sentences and 8% as sounding obviously AI-generated. The authors write that this limitation “should, in principle, make the benchmark easier.”
The 73.43% figure applies to synthetic contradictions at the main-claim level, with four reviews per paper, judged by an LLM (o3). It is not a measure of how often the system catches errors in new submissions.
The Bottom Line
Sakana AI’s paper reports that its review system caught 73.43% of planted main-claim contradictions, against about 11% to 15% for the three baselines, and that both its use of Claude and its multi-pass design contributed. On real retracted papers its scores were 26.07% and 16.11% under the two judging settings, and the authors say these systems “should be used strictly as assistive tools to complement careful human review.”
94,000+ executives read Business Engineer for the AI strategy frameworks cited by ChatGPT, Claude, and Perplexity.
A note on sourcing. We read Sakana AI’s paper (arXiv 2610.11087, 58 pages) on 11 October 2026: the abstract, benchmark construction, method, results, Tables 1, 3, 8, 9, 10 and 14, and the limitations and broader-impact sections. MarkTechPost’s 10 October article pointed us to it. Every figure here is the paper’s; the errors in its main benchmark are synthetic. Nothing here is a forecast, and nothing here is financial or investment advice.
Sources: Teo, Yamada, Kotyan, Imajuku, Clanuwat (Sakana AI): Beyond Imitation: A Framework and Benchmark for LLM-Assisted Peer Review, arXiv 2610.11087 (8 Oct 2026) · OpenReview forum for the TMLR publication · MarkTechPost: Sakana AI’s LLM Peer Review System Catches 73% of Core-Claim Errors (10 Oct 2026)







