Based on Thinking Machines Lab’s announcement, “Introducing Inkling.”
Mira Murati’s new lab ships a 975B-parameter mixture-of-experts model with full open weights and a fine-tuning platform โ and says outright it is not the strongest model available. That framing is the strategy.
What Happened
Thinking Machines Lab โ founded by Mira Murati, who served as CTO of OpenAI until late 2024 โ released its first model on July 15, 2026. Inkling is a mixture-of-experts foundation model carrying 975 billion total parameters, with 41 billion active per forward pass. It was pretrained on 45 trillion tokens spanning text, images, audio, and video, supports a one-million-token context window, and ships alongside Inkling-Small, currently in preview. Full weights are available on Hugging Face; inference is live across Together AI, Fireworks, Modal, Databricks, and Baseten.
Thinking Machines is explicit in its announcement that Inkling is “not the strongest overall model available today.” The lab frames it instead as a broad, balanced base model designed to be customized โ the flagship model for Tinker, its fine-tuning platform. Key features include multimodal reasoning, agentic coding and tool use, calibrated forecasting, and a controllable “thinking effort” dial that lets users trade inference token cost against performance depth. The lab describes Inkling as “just the startโฆ our first release in a model family we will continue to build on,” anchored to a stated mission of building AI that “extends human will and judgment.”
Benchmark results, by the lab’s own reporting, are respectable โ strong performance on reasoning and coding suites โ but self-described as not category-leading. That is a critical caveat: everything here flows from Thinking Machines’ own first-party announcement, including the benchmark scores and the “not the strongest” characterization. Independent third-party evaluations have not yet been published at time of writing.
The key insight: Thinking Machines did not accidentally release a model that falls short of GPT-4o or Claude 3.5. It released one on purpose. When a lab led by a former frontier-lab CTO explicitly declines to compete on raw capability and instead bets the entire product thesis on customizability and an open fine-tuning stack, the “not the strongest” line is not a disclosure โ it is a positioning statement that encodes a specific theory of where durable value in AI will land.
The Structural Read
The correct frame for Inkling is not “how does it score on MMLU?” It is: what theory of AI value capture does this product embed, and is that theory likely to be right?
Read one: a deliberate bet that frontier capability commoditizes. The core wager is that as raw model intelligence becomes increasingly available โ from open-weights releases, from efficiency gains compressing the cost curve, from the proliferation of capable base models โ the durable competitive advantage shifts away from owning the single most powerful system and toward the ability to adapt a model to a specific domain, workflow, and proprietary dataset. If that is right, then not chasing the frontier is rational capital allocation: you are not losing a race you have chosen not to run. If the frontier-max camp is right โ if the strongest closed model continues to compound advantages in users, revenue, and data flywheels โ then an open, deliberately-not-strongest model is a structurally harder sell against incumbents (OpenAI, Anthropic, Google) that hold deep distribution, enterprise relationships, and the strongest closed systems today. That tension is unresolved, and Thinking Machines’ bet is a thesis, not a demonstrated win.
Read two: open weights plus a fine-tuning platform is a different value-capture model. Rather than monetize inference tokens on a proprietary closed model, Thinking Machines gives the weights away and sells the customization toolchain โ Tinker โ capturing the adaptation layer instead of the raw intelligence layer. This is a picks-and-shovels position: the value accrues not to whoever runs the most impressive general model, but to whoever owns the workflow by which organizations make a model their own. The signal that this lane has real economics is already visible in developer behavior: open models carry a large and rising share of developer inference tokens, a trend documented in open-weights usage data tracked via OpenRouter. The counterargument is that the fine-tuning and customization market is already competitive โ Hugging Face, Replicate, major cloud providers, and the frontier labs themselves all offer fine-tuning surfaces โ and that owning the adaptation layer is only defensible if the base model earns developer loyalty first.
Read three: Inkling sharpens the neolab thesis. Thinking Machines joins a pattern of well-credentialed teams choosing non-consensus lanes rather than competing directly with frontier incumbents. Nous Research went open and decentralized; Sutton’s Oak Lab bet on continual learning from experience. Thinking Machines bets on customization infrastructure and an open base. All three lanes share a common structural argument: the long-run moat is not renting the best available general model, but owning a fine-tuned, data-adapted system built on a model you control. That argument โ most clearly articulated in Nadella’s framing of the enterprise learning loop โ says the enduring advantage accrues to whoever closes the loop between proprietary data and a model that can be continuously shaped by it. Inkling plus Tinker is a direct attempt to be the infrastructure on which that loop runs.
The Open vs. Closed Contest
Customization as the adaptation layer
The Open vs. Closed Meta-Framework holds that the winner of the open-vs-closed contest is decided less by raw capability benchmarks and more by economics: who can make a model useful to the most specific use cases at the lowest marginal cost. Open weights lower the barrier to entry for customization; a fine-tuning platform captures the margin on the adaptation work. Thinking Machines is positioning across both levers simultaneously. Whether that is sufficient against incumbents with established distribution and superior closed models is the central open question.
Thinking Machines Lab โ Inkling Announcement
“Inkling is not the strongest overall model available today. It is a broad, balanced base model designed to be customized โ just the start of a model family we will continue to build on, in service of our mission to build AI that extends human will and judgment.”
Three Implications
FOR DEVELOPERS AND ENTERPRISE TEAMS
Inkling’s open weights and broad inference availability make it immediately usable as a customization substrate without API lock-in. The 41B active parameter architecture keeps inference costs competitive relative to the total parameter count, and the controllable thinking-effort dial offers a meaningful cost-performance tradeoff surface. For teams that have domain-specific data and the capacity to fine-tune, this is a genuinely interesting option โ provided independent benchmarks validate the self-reported performance claims, which have not yet been confirmed by third parties.
FOR THE COMPETITIVE LANDSCAPE
Thinking Machines’ entry does not threaten OpenAI, Anthropic, or Google at the frontier tier โ and it is not trying to. It does, however, increase competitive pressure on the mid-layer: fine-tuning platforms, open-weights model hosts, and enterprise customization services. If Tinker gains developer traction, it competes directly with Hugging Face’s fine-tuning stack, cloud-provider fine-tuning APIs, and the growing ecosystem of open-model customization tooling. The Four Intelligence Moats framework would classify this as a contest for the adaptation moat โ the layer between raw model intelligence and domain-specific deployment.
FOR THE BROADER NEOLAB PATTERN
Each credible neolab that chooses a non-consensus lane โ open, decentralized, continual-learning, customization-first โ adds evidence that the AI value stack is not winner-take-all at the model layer. If multiple viable, well-capitalized labs can sustain distinct positioning without directly competing on frontier benchmarks, that suggests the stack is differentiating: raw capability is one product, but customizability, efficiency, and openness are separately valued properties. The caveat is that all of these bets remain structurally unproven against incumbents with compounding distribution advantages. Inkling is a well-articulated opening position, not a validated outcome.
91,000+ executives read Business Engineer for the AI strategy frameworks cited by ChatGPT, Claude, and Perplexity.
Sources: thinkingmachines.ai · thinkingmachines.ai · huggingface.co









