OpenAI Says GPT-6 Astra Scored 55.0% on 11 Ironclad Tasks

OpenAI published a post on 6 October 2026 titled “Advancing computer use with Ironclad.” It says Ironclad, “a leader in AI contracting,” is the first of a small number of software companies it is working with to train AI agents on specialized business software, and that GPT-6 Astra is “our first frontier model trained on Ironclad tasks.”

OpenAI reports that, on its research evaluation of 11 tasks, GPT-6 Astra’s average score was 55.0%, against 41.6% for GPT-5.6 Sol, and that estimated time per attempt fell from 37.0 minutes to 19.2 minutes. These are OpenAI’s own figures. As we read its footnotes, the times are simulated and the 11 tasks are not all of Ironclad’s workflows.

Business Pill · PRACTICE MADE ON PURPOSE

A one-minute explainer of the idea behind part of this story: practice data made on purpose, known as synthetic data. It teaches the concept, not this story’s figures.

The key insight: The post reports a score and a training method together, and the two share a source. GPT-6 Astra is, in OpenAI’s words, “our first frontier model trained on Ironclad tasks,” and it is scored on 11 research tasks that Ironclad employees and OpenAI users of Ironclad helped identify. As we read the post, the figures show how that model performs on 11 research tasks chosen with the same partner. The post reports no test of whether the gain carries to other contracts or other software.

What OpenAI Says It Set Out to Do

The post says: “We’re exploring how to train models to understand a company’s business rules, execute multi-step workflows, and verify that their work meets the original requirements.” It says OpenAI is “partnering directly with a small number of software companies that understand these workflows best,” and that it is identifying challenging, high-value tasks with them and turning them into research problems for training and evaluating its models.

The post says OpenAI worked with Ironclad’s team to develop tasks that “require agents to configure agreements, approvals, and reusable legal terms.” It says Ironclad’s expertise was “instrumental in defining what success looks like.”

OpenAI's reported mean rubric scores on its 11 Ironclad research tasks, and its simulated time per attempt. Al
OpenAI’s reported mean rubric scores on its 11 Ironclad research tasks, and its simulated time per attempt. All figures are OpenAI’s, from its 6 October 2026 post; the times are simulated estimates, not measured customer time savings, and the results do not cover all Ironclad workflows.

The 11 Tasks and How They Were Graded

According to the post, Ironclad employees and people who use Ironclad at OpenAI helped OpenAI’s researchers identify 11 tasks across legal, commercial and procurement work. It gives three examples: setting up nondisclosure agreements, creating procurement approval processes, and updating a reusable legal clause so that it reflects the jurisdiction a requester selects.

OpenAI says it estimates that an experienced user would take about 30 to 40 minutes per task, on average. It says each task was evaluated against 8 to 50 criteria, depending on its complexity, and the post’s results table labels the score “Mean rubric score.”

The post gives a worked example of a legal operations team setting up a process for buying software, in which Finance may need to approve purchases above a certain amount, Security may need to review certain requests and Legal may need to review nonstandard terms. It says an agent doing the same work must configure Finance approval above the spending threshold and check that requests above and below it follow the right paths, and that “the finished process must work across the situations it was designed to handle.”

How the Models Were Trained

The post says Ironclad “provided hosted software environments of their product where the models could practice these tasks.” It says OpenAI’s researchers developed synthetic training tasks “around representative workflows” and used reinforcement learning to help the models improve through practice and feedback.

A footnote says OpenAI created the simulated tasks from contracts publicly available in the SEC’s EDGAR database, after applying filters designed to remove personal information. It says OpenAI did not use OpenAI customer data, OpenAI’s internal contracts, or nonpublic Ironclad customer data or contracts for training or evaluation. We have not checked this statement.

What OpenAI Reports

OpenAI says it compared Astra and GPT-5.6 Sol using Max reasoning for Astra and High reasoning for Sol, “the settings where each model scored highest.” Across the 11 research tasks, the post says Astra’s average score was 55.0%, compared with 41.6% for Sol, and estimated average time per attempt fell from 37.0 minutes for Sol to 19.2 minutes for Astra.

The post states these as a score 32% higher and a time per attempt 48% lower. Our arithmetic: 55.0 divided by 41.6 is about 1.32, and 19.2 divided by 37.0 is about 0.52, which agree with those two statements. The gap in the scores is 13.4 percentage points.

The post also says: “An internal model used in the development of Astra achieved an even stronger 63.7% on these tasks.” That is a third figure on the same tasks, from a model that is not the one the post introduces as GPT-6 Astra.

Two video clips in the post show one task. The post says Astra met about 94% of the task criteria in an estimated 20 minutes, and GPT-5.6 Sol met about 85% in an estimated 32 minutes. It does not say whether this task is typical of the 11.

What the Footnotes Limit

The first footnote says: “These results cover the 11 research tasks, not all Ironclad workflows.” It adds that the times are “simulated estimates based on assumed model processing and generation speeds, not measured customer time savings.”

The post’s section on Ironclad says that a contracting process must handle exceptions while preserving the rules a business depends on, and that human oversight still matters as agents improve at complex contracting tasks. It closes with a quote it attributes to Sunita Verma, Ironclad’s Chief Technology Officer, who is quoted as saying that agents “need to understand the full contracting lifecycle, including how business workflows connect while preserving the controls teams rely on.” That quote is as the post reports it; we have not verified it with Ironclad.

OpenAI’s Invitation to Other Software Companies

The post closes by inviting “a small number of software companies to work directly with our research and engineering teams on important professional tasks that today’s agents still can’t reliably complete.” It asks for a concrete example of what the agent is being asked to do, evidence of where it fails, and how a successful result would be judged.

It says partners should also be able to bring people who know the work deeply, a secure environment for testing, and data that can be safely used for research.

The Structural Read

We read the post as describing a loop with the partner at both ends. Ironclad employees and OpenAI users of Ironclad help pick the tasks, rubrics of 8 to 50 criteria define success, hosted copies of Ironclad’s product give the models somewhere to practise, and a set of research tasks measures the result. The post itself says this combination “ensures we are improving on real-world tasks that are most important to our customers.”

The time figure is a model of time rather than a clock reading. OpenAI’s footnote calls the times “simulated estimates based on assumed model processing and generation speeds,” and the post’s own estimate for an experienced person is about 30 to 40 minutes per task. The post reports no measured time for a person or a customer using either model.

The post sets Astra’s 55.0% beside two other figures on the same tasks: 41.6% for GPT-5.6 Sol and 63.7% for an internal model used in developing Astra. Our arithmetic: Astra sits 13.4 percentage points above Sol and 8.7 below the internal model. The post says OpenAI aims to bring the internal model’s further gains to future models.

OpenAI, in its 6 October 2026 post on Ironclad

“Getting individual steps right is not enough: the finished process must work across the situations it was designed to handle.”

Three Implications

THE SCORE IS OPENAI’S, ON TASKS CHOSEN WITH IRONCLAD The post reports a mean rubric score on 11 research tasks, each graded against 8 to 50 criteria, and gives three example tasks rather than all eleven. Its first footnote says the results do not cover all Ironclad workflows.

THE TIMES ARE SIMULATED OpenAI’s footnote says the times are estimates based on assumed model processing and generation speeds, not measured customer time savings. The post’s 37.0 and 19.2 minutes are described by its footnote in those terms.

THE TRAINING DATA IS DESCRIBED AS PUBLIC CONTRACTS OpenAI’s footnote says the simulated tasks were created from contracts publicly available in the SEC’s EDGAR database, with filters designed to remove personal information, and that OpenAI customer data, its internal contracts and nonpublic Ironclad customer data or contracts were not used for training or evaluation. We have not checked this statement.

What Is Not Established

In the post as we read it, OpenAI states no commercial terms for the collaboration: no payment, price, term or exclusivity. It gives three example tasks, not all 11, and it does not say how many attempts were run per task or how much the scores varied between attempts. It names no model from another company; the comparison is with GPT-5.6 Sol and an internal OpenAI model.

The post says Astra is OpenAI’s first frontier model trained on Ironclad tasks; it does not say how much the training tasks overlap with the 11 evaluation tasks. We did not see who or what applied the rubric criteria, or any evaluation of these tasks by anyone other than OpenAI.

We read OpenAI’s post through a text copy, because the site did not return the page to our direct request. We looked at the first pages of Ironclad’s press, blog and journal listings through the same kind of service on 6 October 2026 and saw no mention of OpenAI there; those copies may lag. We did not contact OpenAI or Ironclad, and we have not tested any model on these tasks.

Readers can also see our report on OpenAI’s and Atlassian’s posts about GPT-6 models in Atlassian’s products; this piece does not compare the two.

Business Engineer Framework

The Map of AI — Where Model Labs Meet Business Software

The Map of AI places more than 200 companies across nine layers of the stack, from silicon to application. It shows where model labs sit relative to the software companies they work with.

Read the Map of AI →

The Bottom Line

On 6 October 2026 OpenAI said it had trained GPT-6 Astra on tasks built with Ironclad, and reported a mean rubric score of 55.0% against 41.6% for GPT-5.6 Sol on 11 research tasks, with simulated times of 19.2 and 37.0 minutes per attempt. The figures are OpenAI’s, on tasks OpenAI chose with Ironclad, and its footnotes say the times are not measured customer savings and the results do not cover all Ironclad workflows.

91,000+ executives read Business Engineer for the AI strategy frameworks cited by ChatGPT, Claude, and Perplexity.

A note on sourcing. This piece rests on OpenAI’s post of 6 October 2026, “Advancing computer use with Ironclad,” read in full, including its footnotes, as a text copy because the site did not return the page to our direct request. We did not obtain any statement from Ironclad itself; the quote the post attributes to Ironclad’s Chief Technology Officer is as the post reports it, and we haven’t verified it with Ironclad. We haven’t checked OpenAI’s figures independently. We did not contact OpenAI or Ironclad, and we have not tested any model on these tasks. Nothing here is a forecast, and nothing here is financial or investment advice.

Sources: OpenAI: Advancing computer use with Ironclad (6 Oct 2026), read as a text copy including footnotes

Scroll to Top

Discover more from FourWeekMBA

Subscribe now to keep reading and get access to the full archive.

Continue reading

FourWeekMBA