OpenAI’s Astra for Law and the Architecture of a Product That Requires Its Buyer to Check the Work

GPT-6 Astra configured for legal research tells you more about its own limits in the go-to-market design than in any benchmark number — and the benchmark number is already doing a lot of work.

ASTRA FOR LAW — KEY FIGURES · SEPT 17, 2026

54%

Astra for Law overall correctness — 200 questions, Vals AI Legal Research Bench private validation set, highest reasoning effort

38.7%

GPT-6 Astra with web search only — same 200 questions, same benchmark, same reasoning effort

230M+

URLs in the legal search index — U.S. case law, statutes, regulations, court rules, administrative decisions

3

Named early deployments — Sullivan & Cromwell, Ropes & Gray, Cooley — all Am Law 200 targets

What Happened

OpenAI introduced Astra for Law on September 17, 2026 — GPT-6 Astra configured with a purpose-built legal search capability, instructions tuned for legal analysis and writing, and firm-facing controls. Behind it sits a legal search index spanning more than 230 million URLs of U.S. case law, statutes, regulations, court rules, and administrative decisions, with sources added on a daily basis. Access runs through a Trusted Access programme inside ChatGPT and Codex, offered first to selected firms, with Am Law 200 as the stated target.

On the published benchmark — 200 questions drawn from the private validation set of Vals AI’s Legal Research Bench, both systems run at their highest reasoning effort — Astra for Law cleared an overall correctness check on 54% of questions. GPT-6 Astra with web search alone reached 38.7% on the same questions. Those are the published figures on that specific benchmark; 200 questions is a small sample drawn from a third party’s private validation set, and they are not a general measure of legal accuracy.

Three early deployments are named. At Sullivan & Cromwell, an agreement analyzer pulls the firm’s own negotiating playbooks and selected precedents into the review of a new deal and turns what it finds into proposed redlines. At Ropes & Gray, the work went into deal diligence. Cooley’s tool, called GO Public, handles IPO preparation — drafting the filing included. No price, contract value, number of firms, number of seats, or general-availability date is given in the announcement.

LAUNCH TIMELINE

Sept 17, 2026

OpenAI introduces Astra for Law — GPT-6 Astra configured with legal search, a 230M+ URL primary-law index updated daily, and firm-level controls. Access via Trusted Access in ChatGPT and Codex.

Sept 17, 2026 — benchmark published

Vals AI Legal Research Bench (private validation set, 200 questions, highest reasoning effort): Astra for Law 54%, GPT-6 Astra with web search 38.7%. Sample size and private third-party provenance apply to both figures.

Early access — selected firms

Sullivan & Cromwell (agreement analyzer with firm playbooks and precedents, producing proposed redlines), Ropes & Gray (deal diligence), and Cooley (GO Public — IPO preparation including filing drafts) named as early deployments.

The key insight: The distribution model — Trusted Access, selected firms first, Am Law 200, delivered inside existing professional tools rather than as an open product — is itself a design document. It encodes the assumption that a qualified reviewer is a precondition of the product, not an optional upgrade. The benchmark number and the go-to-market structure are saying the same thing from two different directions.

A fifteen-point gain changes how much checking is needed. It does not change whether checking is needed, and t
A fifteen-point gain changes how much checking is needed. It does not change whether checking is needed, and the distribution model is built around that.

The Structural Read

Start with the benchmark, then read the distribution model, and notice that they are consistent with each other in a way that is actually informative. A system that clears an overall correctness check on roughly half of 200 research questions drawn from a private third-party validation set is not one you hand over and walk away from. That is not a criticism of the system — it is a description of what the number means as a design input.

The go-to-market structure reads the same way: Trusted Access, selected firms first, Am Law 200, delivered inside ChatGPT and Codex rather than as a standalone open product. That is the shape a distribution model takes when the buyer’s existing professional judgment is load-bearing — when the reviewer is not an optional extra but the mechanism by which the product functions.

Business Engineer — Harness Theory

The Firm That Harnessed Its Own History

Harness Theory holds that the companies best positioned to capture AI value are those that load a general model with something the model’s vendor cannot supply. All three named deployments are exactly that. Sullivan & Cromwell’s agreement analyzer is GPT-6 Astra plus Sullivan & Cromwell’s own accumulated record of what it concedes and what it never concedes — an asset that took decades of transactions to produce. No vendor can ship that. No competitor can copy it. The model is available to anyone on the Trusted Access programme. The playbook is not.

The delta and the level are two distinct findings, and both belong in view. Roughly fifteen points separate Astra for Law from the same underlying model with ordinary web search — on the same 200 questions, same benchmark, same reasoning effort. That is a substantial demonstration that a curated index of primary law outperforms general retrieval at this task. It is the strongest evidence in the announcement that the expensive, unglamorous infrastructure work is the part that produced a measurable result.

But a delta is not a level. Moving from 38.7% to 54% changes how much verification is needed. It does not change whether verification is needed. Those are different questions, and conflating them would be a misreading of what the number actually shows.

The index itself deserves a separate observation, though not the one it usually gets. The headline number — 230 million URLs — is less operationally significant than the three words that follow it in the announcement: sources added daily. Primary law changes continuously. A research corpus that has fallen behind is, in a specific sense, worse than one whose gaps are obvious: a confident answer drawn from superseded material is harder to catch than no answer at all. That is a general property of legal research infrastructure. It is not a claim about this index’s freshness, completeness, or quality — none of which is established here. It is simply the reason daily addition is the hard operational part, not the scalable part.

Benchmark Comparison — Vals AI Legal Research Bench

200 questions · private validation set · highest reasoning effort · not a general measure of legal accuracy

Astra for Law (configured, curated index) 54%
GPT-6 Astra with web search only 38.7%

~15-point gap shows a curated primary-law index outperforms general retrieval at this task. The delta is not the level: both figures require verification by a qualified reviewer.

Three Implications

IMPLICATION 1 — THE CORRECTNESS LEVEL IS THE DISTRIBUTION DESIGN CONSTRAINT

When a correctness figure on a published benchmark is the input that shapes how a product reaches the market, the go-to-market structure becomes a readable artifact. Trusted Access and selective rollout targeting Am Law 200 — firms with in-house expertise to review AI-generated research output — is the architecture of a product whose usefulness depends on a qualified reviewer being present. That is not a limitation being worked around. It is the product as designed, and the design is internally consistent.

IMPLICATION 2 — FIRM-SPECIFIC JUDGEMENT IS THE UNCOPIABLE INPUT

All three named deployments share a structure: a general model loaded with a particular firm’s accumulated record of prior decisions — what it has conceded, what it has not, how it has approached diligence, what its IPO preparation process looks like. That accumulated judgement is the product the firm is deploying, not the model alone. The model is available to any firm on the programme. The playbook, the precedents, the negotiating history — those were produced by doing the work over years and are not replicable by any outside party. The saleable configuration is inseparable from the firm’s own institutional record.

IMPLICATION 3 — DAILY FRESHNESS IS THE OPERATIONALLY HARD PART, NOT THE INDEXING SCALE

230 million URLs is a large number and a meaningful one, but it is not the operationally demanding commitment in this announcement. Primary law changes continuously — statutes are amended, regulations are updated, decisions are issued. A research corpus that has fallen behind is, in a specific sense, more dangerous than a corpus with obvious gaps, because a confident answer drawn from superseded material is harder to identify and correct than the absence of an answer. Daily addition of sources is the commitment that addresses that asymmetry. Building and sustaining that operational cadence across U.S. case law, statutes, regulations, court rules, and administrative decisions is the expensive, unglamorous infrastructure work — and the ~15-point delta on the published benchmark is the evidence that it produced a measurable result at this task.

Business Engineer Framework

91,000+ executives read Business Engineer for the AI strategy frameworks cited by ChatGPT, Claude, and Perplexity.

The figures of 54% and 38.7% are overall correctness results on 200 questions drawn from the private validation set of a third party’s benchmark, Vals AI’s Legal Research Bench, with both systems run at their highest reasoning effort. Two hundred questions is a small sample and the set is private. These are the published figures for that benchmark and are not a general measure of legal accuracy. Nothing here claims any rate of error, hallucination or fabricated citation, and nothing here references any court sanction, disciplinary matter or incident involving AI in legal practice. Nothing here describes the system as reliable, unreliable, safe, unsafe, fit or unfit for any purpose, and nothing extrapolates the benchmark or predicts future scores. The three named deployments are described only as reported. No outcome, time saving, cost saving or quality effect is claimed, and no firm is characterised. No claim is made about the index’s freshness, completeness or quality; it is not described as a moat, no competitive advantage is claimed, it is not compared to any other corpus, no competitor is named as advantaged or disadvantaged, and no competitor’s benchmark score is stated. No price, contract value, number of firms or seats, or general-availability date is reported, and none appears above. Nothing here says anything about legal employment, junior lawyers, headcount or billing models, and no market size, growth rate or share is stated. OpenAI is a private company and the named firms are partnerships. No valuation, share-price or market-capitalisation claim is made. This is business analysis. It offers no legal advice and no guidance on professional responsibility, supervision duties or malpractice; it is not investment advice, no view is expressed on any security, and no recommendation is made.

Sources: openai.com · lawnext.com · siliconangle.com

Scroll to Top

Discover more from FourWeekMBA

Subscribe now to keep reading and get access to the full archive.

Continue reading

FourWeekMBA