Grok 4.7 Holds Its Token Price While Capability Rises — and That Gap Is What Standard Metrics Miss

SpaceXAI’s Grok 4.7 keeps posted API pricing identical to its predecessor — and that unchanged number conceals the only price move that actually matters to buyers.

GROK 4.7 — KEY NUMBERS (VENDOR-REPORTED; NOT INDEPENDENTLY VERIFIED)

46.3%

CursorBench 4.0 — Grok 4.7
vs 40.4% for Grok 4.6 (SpaceXAI own figures)

38.0%

Terminal-Bench 4.0 — Grok 4.7
vs 20.3% for Grok 4.6 (SpaceXAI own figures)

$2 / $6

Input / Output per 1M tokens
Unchanged from Grok 4.6

Price for fast variant
2× speed, 2× posted rate

DISCLOSURE

Benchmark figures throughout this article are SpaceXAI’s own vendor-reported numbers, not independent evaluations. The quality-adjusted deflation argument depends entirely on those benchmarks being representative of real work — which is not established. This is not investment advice.

What Happened

SpaceXAI announced Grok 4.7 on September 21, 2026, describing it as its most capable model for coding and knowledge work. The model ships with a larger base than Grok 4.6, longer reinforcement learning runs targeting multi-hour tasks, stronger self-checking, and improved long-context management. It is available through Cursor, Grok Build, the Grok API, and a range of third-party coding harnesses, model routers, and cloud platforms.

The number that anchors the business story is the one that did not change. Headline API pricing holds at $2 per million input tokens and $6 per million output tokens — identical to Grok 4.6. A fast variant runs at twice the speed for twice the posted rate. Against that flat price, SpaceXAI’s own benchmark table reports CursorBench 4.0 moving from 40.4% to 46.3%, and Terminal-Bench 4.0 moving from 20.3% to 38.0%. Those are vendor figures, not third-party evaluations, and that qualifier applies every time they appear below.

Separately, as relayed through an onward amplification of The Information rather than from the original source, OpenAI is reported to have largely automated the training of new experimental models — internal models writing GPU kernels and optimizing runtime code, with researchers supplying a single optimization example and letting the system run for weeks. That is The Information’s reporting as relayed onward, not an OpenAI announcement, and it carries a weaker evidential chain than the SpaceXAI release.

The key insight: Nobody buys tokens. Buyers purchase completed work — a function that runs, a document that holds up, a task that closes without a retry. If the same spend completes more of that work, the effective price per unit has fallen even though the invoice is identical. That form of deflation is entirely invisible to any metric built on posted rates or blended token cost — and it rests entirely on vendor benchmarks being representative of real work, which is not established here.

VENDOR-REPORTED BENCHMARK MOVEMENT

CursorBench 4.0

Grok 4.6 40.4%
Grok 4.7 46.3%

Terminal-Bench 4.0

Grok 4.6 20.3%
Grok 4.7 38.0%

All figures are SpaceXAI’s own vendor-reported benchmarks. Not independently verified.

If the same spend completes more work, the effective price fell — and no price index recorded it.
If the same spend completes more work, the effective price fell — and no price index recorded it.

The Structural Read

The Map of AI framework identifies nine layers in the AI stack, from raw compute through to the harness layer where model output gets turned into organizational output. The pricing move here — or more precisely, the non-move — is primarily a harness-layer event. It changes what buyers can accomplish at the harness without changing what they pay at the model.

A price index built on posted token rates or blended cost per token cannot see quality-adjusted deflation. The denominator that changed — work completed per unit of spend — is not in the index. This is the standard quality-adjustment problem, familiar from decades of argument about how to measure inflation in computing hardware, and it is unusually acute in AI because the unit of useful output (a working function, a completed task, a document that survives review) is far harder to define than a transistor count or a clock speed. All of this rests on SpaceXAI’s own benchmark figures being representative of real-world work, which is not established by the vendor announcement and not verified here.

Earlier data published here showed the frontier share of usage falling alongside a declining blended token price, and the reading that comes most easily is that work is moving down the price ladder. A frontier model holding its posted price while its vendor-reported capability rises is also getting cheaper per unit of work — just not in a way any posted-price series can record. Both movements can be present at once — a falling posted rate at one end and quality-adjusted deflation at the other — and a blended index conflates them, because it observes only one of the two dimensions. Nothing here establishes which end of the market is gaining. That is a limit on what that measurement can tell anyone, not a correction to the underlying data.

Map of AI — Harness Layer

“The harness layer is where model capability converts into organizational throughput. When posted price holds and vendor-reported capability rises, the value transfer happens entirely at that layer — and it is invisible to every metric that stops at the model boundary.”

Self-checking deserves its own paragraph because it sits on a different economic thread. SpaceXAI describes stronger self-checking alongside longer reinforcement learning on multi-hour tasks. A system that checks its own output is taking on a step that previously belonged to the person receiving that output. The economic value of that is not in the quality of the result — it is in the review time removed from the human side of the workflow.

Review is the cost that scales with volume in any arrangement where a person adjudicates machine output, and it is the cost that quietly caps how much such output an organization can actually absorb. Nothing here claims that this model, or any model, reliably checks its own work, is calibrated, or can be trusted without review. SpaceXAI describes a capability; nothing here verifies it. The distinction between a described capability and a demonstrated one is the entire difference between a marketing claim and a finding.

The OpenAI item reported by The Information — as relayed onward, not from the original source, and not an OpenAI announcement — sits one level further back in the same direction. If accurate, it describes automation applied to the production of new models rather than to their outputs. That is a different kind of claim, and it should be held to a higher evidential standard precisely because it is the more structurally significant of the two stories. Nothing here describes it as recursive self-improvement, an intelligence explosion, or a step toward superintelligence. It is a reported process claim, at two removes, about a company that has not announced it.

Three Implications

IMPLICATION 1 — THE PRICE SERIES PROBLEM

Any organization using posted token rates or blended cost-per-token as its primary AI spend metric is observing only one dimension of what the market is doing. Quality-adjusted deflation — contingent entirely on vendor benchmarks reflecting real work, which is not established — does not appear in that measurement at all. The metric is not wrong; it is incomplete in a way that matters more as capability variance between versions widens.

IMPLICATION 2 — REVIEW AS THE BINDING CONSTRAINT

If self-checking capabilities develop in the direction SpaceXAI describes — and the company describes a capability, not a verified finding — the binding constraint on AI output volume shifts from generation cost to review capacity. Organizations that have not yet thought about review as a workflow design problem rather than a quality-assurance afterthought will encounter that constraint before they expect to.

IMPLICATION 3 — SOURCING DISCIPLINE AS ANALYTICAL INFRASTRUCTURE

The two stories covered here carry meaningfully different evidential weights: a vendor announcement with named benchmark figures, and a process claim relayed at two removes from the original reporting. Treating them as equivalent — because both sound significant — is an analytical error that compounds as the pace of AI news accelerates. The more structurally consequential a claim, the stronger the sourcing bar it requires, not the weaker one it tends to receive.

91,000+ executives read Business Engineer for the AI strategy frameworks cited by ChatGPT, Claude, and Perplexity.

This is not investment advice. The benchmark figures above are SpaceXAI’s own reported numbers, not independent evaluations. The argument that effective prices fell depends entirely on those benchmarks being representative of real work, which is not established here. The company describes stronger self-checking as a capability; nothing above verifies it, claims any model reliably checks its own work, or suggests output can be trusted without review. The report that OpenAI has largely automated the training of experimental models reached this piece through an onward amplification of The Information rather than the original, and is not an OpenAI announcement; nothing above characterises it as recursive self-improvement or as a step toward any broader capability threshold, and nothing above speculates about what would follow if it held. No usage, revenue, adoption, compute, cost or headcount figure appears above, no other company’s model, price or score is compared, and no partner’s commercial terms are described. Nothing is predicted.

Sources: x.ai · x.com

Scroll to Top

Discover more from FourWeekMBA

Subscribe now to keep reading and get access to the full archive.

Continue reading

FourWeekMBA