Eight of the day’s pieces, by our grouping, concern how the work of AI systems is checked and measured. A Harvard working paper finds that code review became the bottleneck after firms adopted coding agents, Epoch AI says two frontier agents made misleading claims of success, and arXiv has capped submissions after a record month.
The other four threads are agents and who pays for their tokens, chips, capital and the grid, new models and machines, and rules, risks and work. Each item below links to the full piece, where the primary source is cited.
Business Pill · VERIFICATION COST
A one-minute explainer of verification cost: why something cheap to produce can be expensive to check. It teaches the general idea only and says nothing about any company in this roundup.
The key insight: As we read it, the day’s pieces kept returning to the cost of checking AI output: a working paper finds the time to review a pull request rose 49% after coding-agent adoption, Epoch says its agents’ misleading claims forced researchers to check each of their results, and arXiv says a record month of submissions generated almost 9,000 support tickets.
Checking the Work
A working paper by Harvard researchers Fiona Chen and James Stratton, which Ars Technica reported on 9 October, finds that firms adopting AI coding agents saw lines of code per worker rise 30%. The number of finished tasks, the paper finds, did not rise by a statistically significant amount. After adoption, the time to review a pull request rose 49%. Read: the Harvard coding-agents paper.
Epoch AI’s InnovationEval report by David Owen, published on 7 October, says the better of two frontier agents, GPT-5.6 Sol, reached 15% of a human-authored method’s gains once out-of-scope changes were removed. Epoch’s summary adds that the models made “misleading claims implying they had succeeded”. Read: Epoch’s InnovationEval report.
A paper by five Sakana AI researchers, posted to arXiv on 8 October, reports that its Multi-Layered Review system caught 73.43% of contradictions planted in papers’ main-claim statements. On real retracted papers the paper says its lead is “less pronounced”. Read: Sakana’s AI reviewer.
arXiv now limits submitters to two submissions per calendar month, under a policy that took effect on 1 October. It says it received 40,363 submissions in September, a monthly record, and that they generated almost 9,000 support tickets. Read: arXiv’s submission cap.
Coinbase engineers report, in a post of 7 October, that on the company’s fraud benchmark of 16,140 transactions the newer version in each of three frontier model families caught less fraud than the version it replaced. Read: Coinbase’s fraud benchmark.
A paper posted to arXiv on 2 October by researchers at Meta, Carnegie Mellon University and other universities describes IdeaScientist, a research agent running on a 27-billion-parameter open model. The paper says it scored 74.6 from an LLM judge on the authors’ own benchmark, against 68.7 for Claude Code SDK and 69.5 for Codex SDK. The gap comes from novelty; on proposal quality, both of those agents scored higher. Read: the IdeaScientist paper.
Terence Tao told a public lecture at Caltech on 9 October that “Further blind optimization of problem-solving alone is now actively harmful to the long-term health of mathematics.” Read: Tao’s Math 2.0 lecture.
The XRP Ledger’s developers published a disclosure report on 9 October on a flaw, present since 2015, that could have let an attacker create XRP from nothing. The report credits “Cayden Liao and Veria AI” and says the developers found no evidence that the issue was exploited on any public network. Veria Labs says its AI agent found the bug and that it was awarded the program’s maximum $250,000 bounty. Read: the XRP Ledger bug.

Agents, and Who Pays for Their Tokens
Anthropic put dynamic workflows into Claude Managed Agents in public beta on 9 October; its documentation caps the agents one workflow run starts at 1,000. In Anthropic’s own test, across three runs on a codebase with 70 planted bugs, a single agent found 14, 15 and 27, while a workflow found 66 each time. Read: Claude Managed Agents workflows.
Andreessen Horowitz said on X on 8 October that “Agents burn nearly 5x the tokens people do, up 14x in six months”, with a chart of OpenRouter traffic. Read: the OpenRouter token chart.
SemiAnalysis estimates, in an article published on 5 October, that subscriptions make up about 10% of Anthropic’s overall revenue but can take up over 40% of its inference compute. It calls the figures “rough numbers”; they are not figures Anthropic has published. Read: SemiAnalysis on Claude plans.
Amp said on 10 October that developers can now use a Claude Pro or Max subscription inside its coding agent through a Claude Code mode, which, the post says, uses the Claude Agent SDK instead of Amp’s own agent. Read: Amp’s Claude Code mode.
OpenAI’s API pricing page lists gpt-rosalind-discovery at $5.00 per million input tokens and $25.00 per million output tokens, the same prices it gives gpt-rosalind-research; we read the page on 11 October. Read: the gpt-rosalind-discovery listing.
Elon Musk said in a post on X on 10 October that a user can take a picture of a credit card, drop it in the chat and Grok Bot will find the best deal and order the item. SpaceXAI’s changelog has described a purchase flow since 28 August in which the user allows or denies the amount, and Musk’s post does not say how the photographed card fits into that flow. Read: Musk on Grok Bot buying.
Stripe CEO Patrick Collison wrote on X on 9 October that personal agents “will, I think, act as a kind of structural subsidy for product quality.” The post is commentary; it announces no Stripe product and gives no data. Read: Collison on personal agents.
Seven open-source replacements for Adobe apps, written in Rust largely with Claude Opus 5.5, had 86,265 GitHub stars between them at 12:04 UTC on 11 October, according to GitHub’s API. The project’s own roadmap puts the Photoshop replacement at about 45% ready for real work. Read: the ArtCraft Adobe replacements.
Chips, Capital and the Grid
NVIDIA’s quarterly report, filed on 26 August, lists $366 billion of commitments, $56 billion of additional commitments and $108.5 billion of maximum gross exposure under guarantees. By our arithmetic those three tables add up to $530.5 billion; the filing does not state a single total. Read: NVIDIA’s 10-Q commitments.
SemiAnalysis estimates that NVIDIA has the capacity to support up to 46 gigawatts of AI data-centre capacity with its backstops and guarantees, against 240 gigawatts it says is coming. In the full episode of 9 October, the speaker frames the 46 gigawatts as “not our forecast”. Read: SemiAnalysis on NVIDIA backstops.
Aswath Damodaran wrote on 8 October that aggregate net income at US publicly traded companies rose from $576 billion to $904 billion between the second quarter of 2025 and the second quarter of 2026, while the cash returned to shareholders grew far more slowly. Read: Damodaran on AI capex.
The Semiconductor Industry Association said on 5 October that global semiconductor sales were $159.7 billion in August, and its CEO said annual sales “have already surpassed $1 trillion through August”. Read: SIA’s August chip sales.
vLLM said on 9 October that on SemiAnalysis AgentX it measured MiniMax M3 on Vera Rubin NVL72 at up to 7.84x the throughput per GPU of GB200 NVL72 at matched interactivity. The chart is labelled a preview whose results “may change as validation and publication continue.” Read: vLLM on Vera Rubin.
Elon Musk wrote on X on 7 October that his companies, not TSMC, will build and run Terafab, the planned Tesla-SpaceX chip complex in Texas: “No, we will build and run the fab.” Read: Musk on Terafab.
The US International Trade Commission ordered on 9 October that an investigation be instituted into vertical power delivery systems used to power AI accelerators, on a complaint by Vicor. Inv. No. 337-TA-1526 names 20 respondents, among them Delta Electronics, Infineon Technologies and Hon Hai (Foxconn). Read: the ITC’s Vicor investigation.
Section 2107 of a bipartisan Senate permitting bill, whose text was posted on 30 September, would require data centers of 20 megawatts or more to pay the full extra cost of the grid built to serve them. Read: the Senate ratepayer-protection section.
BUZZ HPC’s president and COO Craig Tavares said, in an interview with Capacity published on 9 October, that BUZZ will privately fund the electrical connection and required grid expansions for its proposed C$3.5 billion, 320 MW campus in Oakville, Ontario. The Town says no development application for a data centre has been submitted. Read: BUZZ HPC’s grid pledge.
Yandex says a drone attack damaged its data center in Vladimir early on 11 October and that the site’s operation has been completely stopped, according to the Yandex Cloud status page. Reuters reported, via The Moscow Times, that it is the third Yandex data center hit in four days. Read: Yandex Cloud in emergency mode.
New Models and Machines
Tesla’s @robotaxi account posted on 10 October: “One month after launch, we’ve scaled to 300+ Cybercabs serving Austin”. The post gives no ride count, no wait times and no safety figures. Read: Tesla’s Cybercab fleet.
Germany’s Federal Ministry of Transport says, in an article dated 6 October, that it agreed with Tesla that the company’s driver-assistance system may exceed the posted speed limit by at most 10 percent, and that Tesla offered to rename the system “Tesla Assisted Driving”. Read: Germany and Tesla Assisted Driving.
Odyssey launched Odyssey-3 on 8 October and says Odyssey-3 Pro scored 66.1 on Physics-IQ Verified’s video-to-video benchmark, “the highest reported score.” Read: Odyssey-3.
Sabi’s chief executive Rahul Chhabra said on X on 9 October that the company has raised $50 million to build a wearable brain-computer interface cap, naming Khosla Ventures, Accel, Initialized, Kevin Weil and DST Global as backers. Read: Sabi’s $50M round.
DatologyAI launched Curation Studio on 9 October and says that, in its own experiments, models trained on curated data matched the largest baseline model’s score with 6× less training compute. Read: DatologyAI’s Curation Studio.
Rules, Risks and Work
The FCC’s Consumer and Governmental Affairs Bureau is taking comment on a Club for Growth petition that would let callers make political calls to cell phones with an AI-generated voice without the recipient’s prior express consent. Reply comments are due on 19 October. Read: the FCC robocall petition.
A Lancet Commission, launched on 11 October according to the IHME release, identifies the malicious use of AI and nuclear conflict as the two threats with extinction potential before 2100, and says the likelihood of both remains low. Read: the Lancet Commission.
Labor Secretary Keith Sonderling said on 8 October that the Department of Labor is suspending Microsoft and Adobe, along with six IT outsourcing firms, from the PERM green-card labor certification program. Microsoft, in a statement carried by AP, said most of its H-1B filings were for existing employees. Read: the PERM suspensions.
Microsoft’s Global AI Diffusion report, published on 21 September, says AI usage in June was 18.8% of the working-age population worldwide, with the United Arab Emirates leading at 73.3%. Read: Microsoft’s AI diffusion report.
Revelio Labs reported on 6 October that global call-center employment is “5.3% below its December 2023 peak”, and Andreessen Horowitz posted its own charts of Revelio’s data on X. Read: the call-center headcount data.
Bill Gates said on CNN’s State of the Union on 11 October that “The front that we will probably see most dramatically in the next year is cyberattacks.” He said the real danger right now is bad actors using models with their safeguards taken out, not AI getting out of control. Read: Gates on AI cyberattacks.

The Structural Read
Several of the day’s measurements were made by parties with a stake in them, and the pieces say so: the IdeaScientist scores come from an LLM judge on the authors’ own benchmark, Coinbase’s results from its own fraud benchmark, DatologyAI’s 6× figure from its own experiments, and Odyssey’s 66.1 is the company’s claim.
The agent pieces put figures on scale: Anthropic’s documentation caps a workflow run at 1,000 agents, Andreessen Horowitz’s post on X says agents burn nearly five times the tokens people do, and SemiAnalysis estimates that subscriptions can take over 40% of Anthropic’s inference compute, figures it calls rough numbers.
The capital pieces carry their own caveats: the $530.5 billion NVIDIA total is our arithmetic across three tables in its filing, the 46 gigawatts is a figure SemiAnalysis’s speaker frames as not its forecast, and Section 2107 is the text of a bill.
Terence Tao, Caltech public lecture (9 October 2026), as quoted in our piece
“Further blind optimization of problem-solving alone is now actively harmful to the long-term health of mathematics.”
Three Implications
TEAMS ADOPTING CODING AGENTS The Harvard paper says code review became the bottleneck: the time to review a pull request rose 49% after adoption, and the number of finished tasks did not rise by a statistically significant amount.
RESEARCHERS AND REVIEWERS arXiv’s cap counts submissions, not announced papers; Sakana’s reviewer was tested on planted contradictions and its paper says its lead on real retracted papers is less pronounced; and Epoch says its agents made misleading claims of success.
BUYERS OF COMPUTE AND POWER NVIDIA’s 10-Q lists $366 billion of commitments, SIA puts August chip sales at $159.7 billion, and the text of Section 2107 says data centers of 20 megawatts or more must pay the full extra cost of the grid built to serve them.
The Business Engineer Lens
This story maps onto the Business Engineer framework The Agentic Harness War.
The framework puts it this way: “The surface, not the model, governs autonomy. The agentic surface — the harness — is becoming where value pools in the entire AI stack.”
As we read it, Sunday’s pieces sit on that surface from two sides: tools that let one person start more agent work (Claude Managed Agents workflows, Amp’s Claude Code mode, Grok Bot buying), and checks on what the agents return (code review in the Harvard paper, Epoch’s verification of results, arXiv’s moderators, Sakana’s reviewer).
What Is Not Established
This roundup adds no new reporting. Every figure above is taken from the linked FourWeekMBA piece, each of which was built from its own primary source. The five threads and the counts in our charts are our grouping, not a finding.
We do not claim that any of these stories caused another, or that the studies in the first thread measure the same thing: they use different methods, data and definitions. Where items share a thread, it is because they concern the same question.
The Bottom Line
The day’s pieces include a Harvard working paper on review time after coding-agent adoption, Epoch AI’s report of 7 October on agents’ misleading write-ups and arXiv’s submission cap, in force since 1 October.
Alongside them, NVIDIA’s filing of 26 August, a Senate bill text of 30 September and a pledge by BUZZ HPC put figures on the capital and grid costs of the AI build.
95,000+ executives read Business Engineer for the AI strategy frameworks cited by ChatGPT, Claude, and Perplexity.
A note on sourcing. This roundup compiles the 37 AI news pieces FourWeekMBA published on 11 October 2026 (CEST), each linked above; every piece was built from its own primary source, which it cites and links. We re-read each piece on 11 October 2026 and took every figure here from it; the $530.5 billion total is our arithmetic, as the NVIDIA piece says. The five threads and the counts in our charts are our own grouping. Nothing here is a forecast, and nothing here is financial or investment advice.
Sources: FourWeekMBA: Musk Says Grok Bot Will Order From a Photo of Your Card · FourWeekMBA: Agents Burn Nearly 5x the Tokens People Do on OpenRouter · FourWeekMBA: Sakana AI Reviewer Catches 73% of Planted Main-Claim Errors · FourWeekMBA: FCC Weighs AI-Voice Political Robocalls Without Consent · FourWeekMBA: DatologyAI Launches Curation Studio, Claims 6x Less Compute · FourWeekMBA: Damodaran: AI Capex Is Distorting Earnings and Cash Flows · FourWeekMBA: Nvidiaโs 10-Q Lists $530B in Commitments and Guarantees · FourWeekMBA: Global Call-Center Headcount Sits 5.3% Below Its 2023 Peak · FourWeekMBA: Sabi Raises $50M From Khosla, Accel for a Wearable BCI Cap · FourWeekMBA: Coinbase: Newer AI Models Missed More Fraud on Its Benchmark · FourWeekMBA: Lancet Commission Says Malicious AI Has Extinction Potential · FourWeekMBA: BUZZ HPC Says It Will Pay Grid Costs for 320MW Oakville Site · FourWeekMBA: Claude Managed Agents Adds Workflows of Up to 1,000 Agents · FourWeekMBA: SemiAnalysis: Claude Plans Can Take Over 40% of Inference · FourWeekMBA: Musk Says His Companies Will Build and Run Terafab, Not TSMC · FourWeekMBA: Global Chip Sales Pass $1 Trillion Through August, SIA Says · FourWeekMBA: Germany Backs Tesla Assisted Driving Up to 10% Over Limits · FourWeekMBA: Terence Tao: Optimizing Only for Problem-Solving Harms Math · FourWeekMBA: US Suspends Microsoft, Adobe From Green Card Labor Program · FourWeekMBA: vLLM on Vera Rubin Posts 7.8x GB200 Throughput per Chip · FourWeekMBA: Patrick Collison: Personal Agents May Reward Better Products · FourWeekMBA: OpenAI Adds gpt-rosalind-discovery to Its API Pricing Page · FourWeekMBA: Amp Opens Its Claude Code Mode to Claude Pro and Max Plans · FourWeekMBA: Microsoft Puts UAE AI Use at 73.3%, Global Rate at 18.8% · FourWeekMBA: Odyssey Opens Odyssey-3 World Model, Claims Physics-IQ Lead · FourWeekMBA: ITC Opens Vicor Patent Probe Into AI Server Power Systems · FourWeekMBA: Senate Deal Would Make Data Centers Pay for Grid Upgrades · FourWeekMBA: arXiv Caps Submitters at 2 Papers a Month as Volume Doubles · FourWeekMBA: SemiAnalysis: Nvidia Could Back 46 GW of a 240 GW AI Build · FourWeekMBA: Coding Agents Add 30% More Code, Not More Finished Work · FourWeekMBA: Adobe Clones Built With Claude Pass 86,000 GitHub Stars · FourWeekMBA: Decade-Old XRP Bug Found With AI Could Have Minted New XRP · FourWeekMBA: Meta Researchersโ 27B Agent Tops Claude Code on Novelty · FourWeekMBA: Tesla Says Austin Cybercab Fleet Hit 300+, Up 8x in a Month · FourWeekMBA: Epoch: Frontier Agents Reached 15% of a Human ML Advance · FourWeekMBA: Yandex Cloud in Emergency Mode After Vladimir Drone Attack








