Meta Researchers’ 27B Agent Tops Claude Code on Novelty

A research agent running on a 27-billion-parameter open model scored higher overall than Claude Code and Codex at proposing research directions, on a benchmark built by its authors, according to the paper “IdeaScientist: Orchestrating Agents for Grounded Scientific Ideation”, posted to arXiv on 2 October by researchers at Meta, Carnegie Mellon University and other universities. On a score from an LLM judge, the trained system reached 74.6, against 68.7 for Claude Code SDK on Claude-4.8-Opus and 69.5 for Codex SDK on GPT-5.4. The gap comes from novelty; on proposal quality, both of those agents scored higher.

The lead author, Jiarui Liu of Carnegie Mellon, did the work during an internship at Meta; co-authors are listed at Meta Reality Labs, FAIR at Meta, the University of Illinois Urbana-Champaign, the University of Toronto, Princeton, the University of Tokyo and RIKEN AIP. DAIR.AI highlighted the paper on X on 10 October. The authors have published the code on GitHub and the dataset on Hugging Face.

How IdeaScientist Works

The paper treats research ideation as its own task: given a problem and a challenge that current methods fail to meet, produce a grounded research proposal. It builds on the intuition that “a challenge in one field can often be addressed by a mechanism that solved an analogous challenge in another”.

An orchestrator keeps the research plan and calls three specialised sub-agents. A gap finder examines closely related work for methodological limitations, an innovator draws on work from other problem settings to develop transferable solution intuitions, and a report writer turns the surviving intuition into a full proposal with a fixed schema. A paper reader and a reviewer work in separate contexts.

Each of the three roles is trained separately with reinforcement learning, using a separate LoRA adapter per role and Group Relative Policy Optimization. The reward combines an LLM judge’s rating of the artifact, its agreement with a reference reconstructed from a human-written paper, and the overlap between the papers each one cites.

Overall score on the paper’s 277 held-out test problems, the mean of the novelty and proposal-quality av
Overall score on the paper’s 277 held-out test problems, the mean of the novelty and proposal-quality averages, scored by a Qwen3.6-27B judge (Table 1 of the IdeaScientist paper). IdeaScientist-Trained on Qwen3.6-27B: 74.6; Codex SDK on GPT-5.4: 69.5; Claude Code SDK on Claude-4.8-Opus: 68.7; DeepScientist on Qwen3.6-27B: 60.6. The untrained harness on Claude-4.8-Opus scored 77.7.

Business Pill · PROXY MEASURE

A one-minute explainer of proxy measures: how a quick score stands in for the thing we actually want to know. It teaches the general idea only and says nothing about any person or organisation in this story.

The key insight: As we read it, the paper’s result is about structure more than model size: an open 27-billion-parameter backbone, split into trained roles, scored higher on novelty than two commercial coding agents running their own loops. The same scores show the commercial agents ahead on proposal quality, and the highest overall score in the table came from the authors’ harness running on Claude-4.8-Opus.

The Idea Vault and the Test

The authors built what they call the Svalbard Idea Vault from the arXiv API: 4.7 million papers published from 1986 through 18 June 2026, 2.8 million of them with full text. From these they extracted 2.77 million tuples of problem definition, challenge, research gap, intuition and proposal.

They set a cutoff date of 1 January 2026. From machine-learning papers published on or after that date, filters left 14,977 methodological target papers, split into 14,500 for training, 200 for validation and 277 held out for testing. During evaluation the agent can search only literature from before the cutoff, so the test asks whether it can arrive at a direction before human researchers published it. The targets are rewritten into the proposal schema with measured experimental outcomes removed.

Every system receives the same problem definition and output schema, uses the same tools on the same offline paper database, and has internet access disabled. External systems, including Claude Code and Codex, keep their native agent loops.

The Scores

Each proposal is scored on novelty, measured against all literature, pre-cutoff literature and same-domain literature, and on six proposal-quality criteria commonly used in peer review: relevance, clarity, specificity, actionability, soundness and impact. The overall score is the mean of the two averages. All scores in the main table come from a Qwen3.6-27B judge.

On the Qwen3.6-27B backbone, IdeaScientist-Trained scored 74.6 overall. The strongest other open-source autoresearch system on the same backbone, DeepScientist, scored 60.6, the gap the paper reports as 14.0%. Claude Code SDK on Claude-4.8-Opus scored 68.7 and Codex SDK on GPT-5.4 scored 69.5, which the paper reports as gaps of 5.9% and 5.1%.

The split matters. On average novelty, IdeaScientist-Trained scored 68.8, against 51.9 for Claude Code SDK and 54.7 for Codex SDK. On average proposal quality, Claude Code SDK scored 85.5 and Codex SDK 84.4, against 80.4 for IdeaScientist-Trained.

The highest overall score in the table, 77.7, came from the untrained IdeaScientist harness running on Claude-4.8-Opus itself.

Average novelty and proposal quality: Claude Code SDK 51.9 and 85.5, Codex SDK 54.7 and 84.4, IdeaScientist-Trained 68.8 and 80.4
Average novelty and average proposal quality on the 277 held-out test papers, scored by a Qwen3.6-27B judge (Table 1 of the IdeaScientist paper): Claude Code SDK on Claude-4.8-Opus 51.9 and 85.5, Codex SDK on GPT-5.4 54.7 and 84.4, IdeaScientist-Trained on Qwen3.6-27B 68.8 and 80.4.

Which Parts Mattered

In the authors’ ablation on Qwen3.6-27B, novelty against all literature rose from 0.595 with the report writer alone to 0.634 with the gap finder added and 0.660 with the innovator added. Letting the innovator search other fields, rather than only the same domain, raised same-domain novelty from 36.3% to 67.0%.

Training improved each role in isolated held-out evaluation: the gap finder’s reward rose from 56.6% to 66.7%, the innovator’s from 58.8% to 69.5% and the report writer’s from 78.6% to 89.1%. Recomposed, the trained system improved on the untrained one by 1.9% on average novelty and 0.4% on average proposal quality.

How the Scores Were Judged

The authors compared LLM judges with two human experts who scored 50 proposals generated by 9 systems. They write that “scoring research proposals is inherently subjective”: the two experts correlated at r = 0.56, and the top three LLM judges correlated with the experts’ average at r = 0.51 to 0.59. They chose Qwen3.6-27B as the main judge, reporting that its self-preference bias was negligible and lower than Claude-4.8’s.

In a blind pairwise human evaluation on 50 test papers, two annotators compared IdeaScientist-Trained with other systems on the same backbone. The paper reports that it won by a statistically significant margin against every system except DeepScientist, and that its edge over the untrained version was not statistically significant.

The Structural Read

The design separates three jobs that a single agent loop usually blends: finding what is missing in close prior work, looking for a mechanism in another field, and writing the proposal up. The authors train each job against its own reference artifact.

As we read it, the ablations point the same way. Adding the gap finder and then the innovator raised novelty at each step, and letting the innovator search other fields nearly doubled novelty measured against the proposal’s own domain, from 36.3% to 67.0%.

The measurement is the open question the paper itself raises. Its two human experts agreed with each other at r = 0.56, about as well as the best LLM judges agreed with them, and every score is taken before any experiment tests whether a proposal works.

The IdeaScientist paper (Liu et al., arXiv 2610.04074, 2026), on evaluating research ideas

“Pre-execution evaluation must substitute observable proxies for eventual scientific value, but different proxies answer different questions.”

Three Implications

RESEARCH TEAMS The paper reports that a structured harness of trained roles on an open 27B model scored higher on novelty than Claude Code SDK and Codex SDK in its benchmark, while scoring lower on proposal quality.

AGENT BUILDERS In the authors’ ablation, each added sub-agent raised novelty, and cross-domain retrieval raised same-domain novelty from 36.3% to 67.0%.

ANYONE READING THE SCORES The paper notes that human experts disagree to a notable degree on proposal quality (r = 0.56), and that its main scores come from a Qwen3.6-27B judge before any experiment is run.

The Business Engineer Lens

This story maps onto the Business Engineer framework The Agentic Harness War.

The analysis puts it this way: “Value capture migrates up from the model to the harness.”

As we read it, the IdeaScientist paper is a research-lab example of that claim: the same untrained harness scored highest of all on Claude-4.8-Opus, and on a 27B open backbone it still scored above two commercial agents on novelty.

What Is Not Established

This is an arXiv preprint, and the document does not indicate journal peer review. The benchmark, the test split and the scoring were designed by the authors, and the main scores come from an LLM judge. The human pairwise evaluation compared systems on the same backbone, so the comparison with Claude Code SDK and Codex SDK rests on the judge’s scores.

The proposals are scored before any experiment is run. The paper notes that “Pre-execution evaluation must substitute observable proxies for eventual scientific value”, and cites expert studies finding that “greater novelty may come at the cost of feasibility in actual execution”.

Business Engineer Framework

The Agentic Harness War

A Business Engineer analysis of the race to own the harness around AI agents, and of why the layer that wraps the model, not the model alone, shapes what agents deliver.

Read the Map of AI →

The Bottom Line

In the authors’ benchmark, a structured harness of trained sub-agents on a 27-billion-parameter open model scored 74.6 against 68.7 for Claude Code SDK, with the gap coming from novelty while the commercial agents scored higher on proposal quality. The scores come from an LLM judge, on 277 held-out papers, before any experiment was run.

95,000+ executives read Business Engineer for the AI strategy frameworks cited by ChatGPT, Claude, and Perplexity.

A note on sourcing. On 11 October 2026 we read the full 36-page PDF of “IdeaScientist: Orchestrating Agents for Grounded Scientific Ideation” (arXiv 2610.04074v1, submitted 2 October 2026), from which every figure and quotation here is taken. A DAIR.AI post on X on 10 October brought it to our attention. The scores are the authors’, from their own benchmark and LLM judge. Nothing here is a forecast or investment advice.

Sources: Liu et al., ‘IdeaScientist: Orchestrating Agents for Grounded Scientific Ideation’, arXiv 2610.04074 (2 Oct 2026) · DAIR.AI on X (10 Oct 2026)

Scroll to Top

Discover more from FourWeekMBA

Subscribe now to keep reading and get access to the full archive.

Continue reading

FourWeekMBA