Google Research Cogentic Reports Five Open Math Results

This piece rests on one arXiv preprint by the system’s own authors. This publication has not read the five companion papers and has not checked any proof. The differences and percentages are this publication’s own arithmetic.

A multi-agent harness built on Gemini reports novel results on five open problems in online learning, auction theory, and mechanism design — according to a single preprint by its own authors.

What Happened

This piece is built on one source: arXiv preprint 2609.40324, version 1, dated 30 September 2026, by authors affiliated with Google Research. This publication read the paper in full. It is the authors’ own account of their own system. This publication has not read the five companion papers and has not checked any proof or any prior-state claim.

The paper describes Cogentic, a multi-agent harness that uses Gemini as its base model. According to the paper’s abstract, Cogentic produced novel results on five open problems across online learning, auction theory, and mechanism design. The paper states that each result was independently verified by domain experts and is developed in full in companion papers. That verification claim belongs to the paper; this publication makes no independent assessment of it.

The problems, the paper says, sit at the hardness level of open questions in flagship theoretical computer science conferences. The authors focused on problems within their own areas of expertise. None of the results here are machine-checked in a proof assistant such as Lean; the paper raises that possibility but the work reported does not include it.

Lower is better in both pairs. The paper's Table 1 gives the prior figure and the Cogentic figure for each; th
Lower is better in both pairs. The paper’s Table 1 gives the prior figure and the Cogentic figure for each; this publication has not checked either, and the revenue and PoA ratios are not comparable with each other.

How the Harness Works

The paper describes a division of labour modelled on a research group. An orchestrator decides what gets worked on. It does not perform mathematical derivations itself and is not permitted a mathematical opinion.

Provers draft proofs in parallel. After each prover completes a draft, a verifier checks it. A separate verifier then reads all of the round’s drafts side by side. A draft is accepted only when it passes both checks. Both verifiers start from the assumption that every step is wrong until justified.

Literature reviewers retrieve related work. Each round leaves a record of what was tried and a ledger of verified results. An auditor re-verifies intermediate lemmas in isolation before later provers may reuse them.

A process advisor reviews the logs and adjusts instructions. Like the orchestrator, it is barred from offering a mathematical opinion. At the end, a comparator selects the strongest proof, a formal writer expands it into a manuscript, and a final audit checks the document against the accepted proof.

The paper says the system needs only a problem statement, without expert hints, and works autonomously until it produces a result in the form of a paper.

The key insight: The harness’s design enforces adversarial skepticism by rule — verifiers are instructed to assume every proof step is wrong until justified, and both the orchestrator and process advisor are structurally prohibited from offering a mathematical opinion.

The Inference Budget

The paper states the results were based on Cogentic runs using on the order of 100 calls to Gemini for most problems and on the order of 1,000 for the hardest. Those are order-of-magnitude figures as the paper gives them.

The paper calls this a relatively low inference budget. It adds that the system has headroom for further optimization of the number of calls and that the design allows for naturally scaling up to solve harder problems. The paper gives no dollar cost, no token count, and no wall-clock time. Neither does this article.

The Five Results, in Prose

The prior-state figures below are the paper’s account and were not checked by this publication. The differences and percentages marked are this publication’s own arithmetic.

Online inverse linear optimization. The prior state, per the paper, was an efficient bound that grew with the number of rounds, with a bound that did not grow achievable only via a rule that is not efficient. The paper reports the first efficient and first proper bound that no longer grows with the number of rounds, achieved efficiently, at a cost of order-d-squared arithmetic per round.

On this result, the paper itself notes that after the companion paper appeared, another researcher obtained a tighter bound, but that the new algorithm is inefficient and an efficient algorithm with that regret remains open.

Two-sided market competition complexity. The prior constant required at least 20,000 agents per side. The paper reports that adding just two agents on the smaller side suffices, and, per Table 1, that adding one does not suffice for any mechanism that is dominant-strategy incentive compatible, individually rational and weakly budget-balanced. That is a reduction from 20,000 to 2 — this publication’s arithmetic, labelled.

Anytime regret with many experts. The best known anytime guarantee was a factor of the square root of two worse than the fixed-horizon guarantee, and whether that gap was necessary was open. The paper reports that the leading constants coincide as the number of experts grows.

Simple versus optimal revenue for a single additive buyer. The prior bound, per the paper, was that 5.2 times the larger of selling separately or bundling is at least the optimal revenue. The paper reports 3.52. That is a drop of 1.68, about 32 percent of 5.2 — this publication’s arithmetic, labelled.

Price of anarchy for autobidding auctions. The prior bound for two bidders was 1.8. The paper reports an optimal 1.5 for two bidders for anonymous, monotone mechanisms, and a bound of 2 minus 1/(4n+1) for n bidders, with a general tight mechanism for n bidders previously open. The drop from 1.8 to 1.5 is 0.3, about 17 percent of 1.8 — this publication’s arithmetic, labelled. The revenue ratio and the price-of-anarchy ratio measure different things and are not compared with each other.

What the Paper Says About Its Limits

The paper’s limits section is worth reading carefully. The results are natural-language proofs, not machine-checked proofs. The paper says this was possible because the problems were chosen from areas the authors work in.

Some companion papers include coauthors who had already been working on the corresponding problems. The authors say they checked the argument, wrote the exposition around it, and in some cases carried it further than the harness had.

Cogentic paper — verbatim

“A system like this can produce candidate results faster than they can be read, and the gap widens as the compute budget grows.”

The paper raises formalizing results in a proof assistant such as Lean so that correctness is settled mechanically. It adds that human understanding of the solution might lag behind. The results reported here are not machine-checked.

What Is Not Established

Not established, and therefore absent from this analysis: whether any proof is correct beyond the paper’s statement that experts verified them; the content of the companion papers; whether the prior-state figures in Table 1 are complete; any dollar cost, token count or running time; how many runs were attempted, or on how many problems the system failed, since the paper reports only the five results; whether any result will be accepted at a venue; and any comment from outside the author team.

None of the above is investment advice. It reports what one preprint says about a research system.

Business Engineer Framework

Founders, Distributors, Enablers

The Business Engineer framework maps how value is created across an AI ecosystem into Founders, Distributors and Enablers.

Explore the Map of AI →

The Bottom Line

A single preprint by Google Research authors reports that Cogentic, a multi-agent harness running on Gemini, produced novel results on five open problems, with a Gemini call budget the paper calls relatively low and verification built into the workflow. The proofs are not machine-checked and the companion papers were not read here. Everything above rests on the authors’ own account.


Source: Cogentic, arXiv preprint 2609.40324 v1, 30 September 2026, Google Research authors. This publication read the preprint in full. The companion papers were not read. No proof was independently checked. Prior-state figures are the paper’s account. Differences and percentages are this publication’s own arithmetic, labelled as such.

91,000+ executives read Business Engineer for the AI strategy frameworks cited by ChatGPT, Claude, and Perplexity.

Everything above comes from one arXiv preprint, 2609.40324 version 1, dated 30 September 2026, written by authors affiliated with Google Research and read in full by this publication. This publication has not read the five companion papers and has not checked any proof or any prior-state figure in Table 1. The statement that domain experts verified each result is the paper’s. The paper gives the number of Gemini calls as on the order of 100 for most problems and on the order of 1,000 for the hardest, and states no dollar cost, token count or running time.

The proofs are in natural language and are not machine-checked. The differences and percentages are this publication’s own arithmetic. The revenue and price-of-anarchy ratios measure different things and are not compared with each other. Nothing above predicts anything and nothing here is investment advice.

Scroll to Top

Discover more from FourWeekMBA

Subscribe now to keep reading and get access to the full archive.

Continue reading

FourWeekMBA