Cognition and Factory Bet Opposite Sides of the AI Stack

Two well-funded companies are selling autonomous software engineering to enterprises — and they have drawn opposite conclusions about where the intelligence should sit.

Factory’s figures below come from its own newsroom; Cognition’s from CoreWeave’s engineering blog. Neither set has been independently reproduced. The three cost percentages have different scopes: a June benchmark range, an aggregate production figure, and a per-user median at one customer. Nothing below declares a winner between the two approaches. Nothing here is investment advice.

What Happened

Cognition, the company behind the Devin autonomous coding agent, this week became the first customer to run production agentic workloads on Nvidia Vera Rubin NVL72 racks at CoreWeave. It scaled to thousands of GPUs there. The benchmarks Cognition published were run on SWE-2, its own model — trained in-house through pre-training, fine-tuning and reinforcement-learning post-training. The claimed 4.8 times token throughput per GPU at matched interactivity is a statement about its own vertically integrated stack.

Factory went public with a different story on September 15. It raised $200 million at a $5 billion valuation, bringing total funding past $400 million. Investors include Blackstone, Khosla Ventures, Sequoia Capital, Insight Partners, NEA, Evantic Capital, Sound Ventures, Mantis VC and Clearlake. The company was founded in 2023 and names Nvidia, Blackstone, Royal Bank of Canada, Palo Alto Networks, Adobe and T-Mobile as customers.

Factory’s pitch is the inverse of Cognition’s. Its flagship product, the Factory Router, picks a model per task from whichever frontier models are available. Factory says the Router is now available in Factory Private, in private preview, which it calls the first and only on-premises model router built for software development agents. Enterprises can control which models run, and where — cloud, on-premises or fully air-gapped.

The sequence on Factory’s side, from its own newsroom: available on the Claude Marketplace from 9 September; $200 million at a $5 billion valuation announced 15 September by co-founders Matan Grinberg and Eno Reyes; the Router cost results published 23 September by Abhay Singhal; the OpenAI B2B marketplace listing on 29 September; and Automations, which lets a Droid run a recurring workflow on a schedule or trigger, generally available on 30 September.

The key insight: Both companies are selling autonomous software engineering. One is betting that owning the model is the durable advantage. The other is betting that being indifferent to the model is. Those are genuinely opposite structural positions — and both are well-capitalised enough to find out which is right.

Factory publishes the quality side alongside the cost side: routed runs at 99 per cent of Claude Opus 4.7's pa
Factory publishes the quality side alongside the cost side: routed runs at 99 per cent of Claude Opus 4.7’s pass rate on Terminal-Bench 2 and 96 per cent on Legacy-Bench.

The Structural Read

The FDE Framework — Founders, Distributors, Enablers — is useful here. Both companies look like Founders from the outside. But they are integrating in different directions.

Cognition is integrating downward into the stack. It trains SWE-2 itself. It queued first for Vera Rubin NVL72 hardware at CoreWeave. Its throughput claim is meaningful precisely because it controls every layer from model weights to inference infrastructure. The moat, if it holds, is compound: better models plus cheaper compute per token, accruing together.

Factory is integrating sideways — across model providers, deployment environments and enterprise procurement channels. The Router does not care which frontier model wins. Factory Private does not care whether the enterprise runs in the cloud or behind an air gap. The moat, if it holds, is the opposite: it gets stronger the more heterogeneous the model landscape becomes.

The marketplace positioning is the sharpest single data point in Factory’s story. Since September 9, Factory has been available on the Claude Marketplace, where enterprise customers can apply committed Anthropic spend toward Factory. Since September 29, eligible OpenAI enterprise customers can do the same through the OpenAI B2B marketplace.

So an enterprise can use either frontier lab’s committed budget to buy a product whose headline claim is 63 percent lower inference cost against published frontier rates. Both Anthropic and OpenAI evidently judge the distribution worth it. Matan Grinberg, Factory’s CEO, frames the Anthropic relationship as complementary: “Claude gives engineering teams frontier reasoning; Factory turns that reasoning into governed, production-scale software delivery.”

Factory monetises the committed spend of both labs while selling a layer designed to reduce dependence on their per-token inference pricing. That is an unusual position in the stack. It is reported here neutrally — both labs made a deliberate distribution decision, and neither appears to be disadvantaged by the arrangement as described.

The Cost and Quality Numbers

Factory’s inference cost claims come in three distinct scopes, and they should not be read as a single rising number.

When the Router launched in June, benchmarks showed 20 to 25 percent lower inference costs. That is a benchmark figure from a specific point in time.

In aggregate production, sessions using the Router now cost 63 percent less than the same workload at published frontier-model rates. That is a production aggregate across Factory’s customer base, not a benchmark.

A third figure is narrower still. Micaela Stump, a team lead product manager on Adyen’s internal developer platform, reports that router users there saved a median of 72 percent, and that the Router saved thousands of dollars in a single week at Adyen specifically. That is a per-user median at one named customer.

A cost claim without a quality measure is empty. Factory publishes both. Routed runs achieve 99 percent of Claude Opus 4.7’s pass rate on Terminal-Bench 2, and 96 percent on Legacy-Bench. Those figures matter because they establish what, if anything, is given up to reach the cost reduction.

What is not established, and therefore absent here: Factory’s revenue, Cognition’s valuation or total funding, the Router’s specific model mix, and any independent reproduction of either company’s benchmarks.

Set the Router’s published figures out in order, because they are not one number. A June benchmark showed 20 to 25 per cent lower inference cost against published frontier rates. In production, sessions now cost 63 per cent less in aggregate. At Adyen specifically, router users saved a median of 72 per cent.

Factory reports the quality side against the same comparison. Routed runs reached 99 per cent of Claude Opus 4.7’s pass rate on Terminal-Bench 2 and 96 per cent on Legacy-Bench.

Three Implications

FOR ENTERPRISES BUYING AI ENGINEERING TOOLS

The architecture choice is upstream of the vendor decision. An enterprise that wants full control over model selection and deployment environment has a different set of vendors than one that wants a single optimised stack.

91,000+ executives read Business Engineer for the AI strategy frameworks cited by ChatGPT, Claude, and Perplexity.

The Factory figures above come from three posts on the company’s own newsroom at factory.ai/news, dated 9, 15 and 23 September 2026, plus its news index entries of 29 and 30 September, all read directly. The Cognition figures come from CoreWeave’s engineering blog of 30 September 2026, also read directly. Nothing has been independently verified and no benchmark has been reproduced here. The three cost figures describe different things and are not one number growing.

The 20 to 25 per cent range is from a June benchmark. The 63 per cent is an aggregate across production sessions. The 72 per cent is a median for router users at one named customer, Adyen, quoted by an Adyen employee. Factory publishes quality figures alongside them, and those are reported above. Factory’s availability on both the Claude Marketplace and the OpenAI B2B marketplace is reported as the companies describe it.

Both labs list Factory voluntarily, and nothing above suggests Factory is exploiting either of them or that either lab has been disadvantaged. Nothing above declares a winner between the two architectures, and nothing above predicts how either company will fare. They are described as two different answers to the same question. Not established and therefore absent: Factory revenue or recurring revenue, Cognition’s valuation or total funding, the Factory Router’s model mix, either company’s customer count beyond what each states, and any independent reproduction of either set of benchmarks. Nothing here is investment advice.

Sources: factory.ai · factory.ai · fourweekmba.com · factory.ai · coreweave.com

Scroll to Top

Discover more from FourWeekMBA

Subscribe now to keep reading and get access to the full archive.

Continue reading

FourWeekMBA