GPT-6 Astra configured for legal research tells you more about its own limits in the go-to-market design than in any benchmark number — and the benchmark number is already doing a lot of work.
What Happened
OpenAI introduced Astra for Law on September 17, 2026 — GPT-6 Astra configured with a purpose-built legal search capability, instructions tuned for legal analysis and writing, and firm-facing controls. Behind it sits a legal search index spanning more than 230 million URLs of U.S. case law, statutes, regulations, court rules, and administrative decisions, with sources added on a daily basis. Access runs through a Trusted Access programme inside ChatGPT and Codex, offered first to selected firms, with Am Law 200 as the stated target.
On the published benchmark — 200 questions drawn from the private validation set of Vals AI’s Legal Research Bench, both systems run at their highest reasoning effort — Astra for Law cleared an overall correctness check on 54% of questions. GPT-6 Astra with web search alone reached 38.7% on the same questions. Those are the published figures on that specific benchmark; 200 questions is a small sample drawn from a third party’s private validation set, and they are not a general measure of legal accuracy.
Three early deployments are named. At Sullivan & Cromwell, an agreement analyzer pulls the firm’s own negotiating playbooks and selected precedents into the review of a new deal and turns what it finds into proposed redlines. At Ropes & Gray, the work went into deal diligence. Cooley’s tool, called GO Public, handles IPO preparation — drafting the filing included. No price, contract value, number of firms, number of seats, or general-availability date is given in the announcement.
The key insight: The distribution model — Trusted Access, selected firms first, Am Law 200, delivered inside existing professional tools rather than as an open product — is itself a design document. It encodes the assumption that a qualified reviewer is a precondition of the product, not an optional upgrade. The benchmark number and the go-to-market structure are saying the same thing from two different directions.

The Structural Read
Start with the benchmark, then read the distribution model, and notice that they are consistent with each other in a way that is actually informative. A system that clears an overall correctness check on roughly half of 200 research questions drawn from a private third-party validation set is not one you hand over and walk away from. That is not a criticism of the system — it is a description of what the number means as a design input.
The go-to-market structure reads the same way: Trusted Access, selected firms first, Am Law 200, delivered inside ChatGPT and Codex rather than as a standalone open product. That is the shape a distribution model takes when the buyer’s existing professional judgment is load-bearing — when the reviewer is not an optional extra but the mechanism by which the product functions.
Business Engineer — Harness Theory
The Firm That Harnessed Its Own History
Harness Theory holds that the companies best positioned to capture AI value are those that load a general model with something the model’s vendor cannot supply. All three named deployments are exactly that. Sullivan & Cromwell’s agreement analyzer is GPT-6 Astra plus Sullivan & Cromwell’s own accumulated record of what it concedes and what it never concedes — an asset that took decades of transactions to produce. No vendor can ship that. No competitor can copy it. The model is available to anyone on the Trusted Access programme. The playbook is not.
The delta and the level are two distinct findings, and both belong in view. Roughly fifteen points separate Astra for Law from the same underlying model with ordinary web search — on the same 200 questions, same benchmark, same reasoning effort. That is a substantial demonstration that a curated index of primary law outperforms general retrieval at this task. It is the strongest evidence in the announcement that the expensive, unglamorous infrastructure work is the part that produced a measurable result.
But a delta is not a level. Moving from 38.7% to 54% changes how much verification is needed. It does not change whether verification is needed. Those are different questions, and conflating them would be a misreading of what the number actually shows.
The index itself deserves a separate observation, though not the one it usually gets. The headline number — 230 million URLs — is less operationally significant than the three words that follow it in the announcement: sources added daily. Primary law changes continuously. A research corpus that has fallen behind is, in a specific sense, worse than one whose gaps are obvious: a confident answer drawn from superseded material is harder to catch than no answer at all. That is a general property of legal research infrastructure. It is not a claim about this index’s freshness, completeness, or quality — none of which is established here. It is simply the reason daily addition is the hard operational part, not the scalable part.
Benchmark Comparison — Vals AI Legal Research Bench
200 questions · private validation set · highest reasoning effort · not a general measure of legal accuracy
~15-point gap shows a curated primary-law index outperforms general retrieval at this task. The delta is not the level: both figures require verification by a qualified reviewer.
Three Implications
IMPLICATION 1 — THE CORRECTNESS LEVEL IS THE DISTRIBUTION DESIGN CONSTRAINT
When a correctness figure on a published benchmark is the input that shapes how a product reaches the market, the go-to-market structure becomes a readable artifact. Trusted Access and selective rollout targeting Am Law 200 — firms with in-house expertise to review AI-generated research output — is the architecture of a product whose usefulness depends on a qualified reviewer being present. That is not a limitation being worked around. It is the product as designed, and the design is internally consistent.
IMPLICATION 2 — FIRM-SPECIFIC JUDGEMENT IS THE UNCOPIABLE INPUT
All three named deployments share a structure: a general model loaded with a particular firm’s accumulated record of prior decisions — what it has conceded, what it has not, how it has approached diligence, what its IPO preparation process looks like. That accumulated judgement is the product the firm is deploying, not the model alone. The model is available to any firm on the programme. The playbook, the precedents, the negotiating history — those were produced by doing the work over years and are not replicable by any outside party. The saleable configuration is inseparable from the firm’s own institutional record.
IMPLICATION 3 — DAILY FRESHNESS IS THE OPERATIONALLY HARD PART, NOT THE INDEXING SCALE
230 million URLs is a large number and a meaningful one, but it is not the operationally demanding commitment in this announcement. Primary law changes continuously — statutes are amended, regulations are updated, decisions are issued. A research corpus that has fallen behind is, in a specific sense, more dangerous than a corpus with obvious gaps, because a confident answer drawn from superseded material is harder to identify and correct than the absence of an answer. Daily addition of sources is the commitment that addresses that asymmetry. Building and sustaining that operational cadence across U.S. case law, statutes, regulations, court rules, and administrative decisions is the expensive, unglamorous infrastructure work — and the ~15-point delta on the published benchmark is the evidence that it produced a measurable result at this task.









