🔍 Read the full analysis: Why Ironclad’s Fine Print Matters As OpenAI Trains Agents In Software on ThorstenMeyerAI.com
Prime made for students and young adults
- Fast, free delivery for dorm and study essentials
- Prime Video and Amazon Music included
- Member-only deals
TL;DR
OpenAI’s October 6 report describes training GPT-6 Astra on 11 legal, commercial and procurement tasks inside hosted copies of Ironclad’s contract-management software. Astra met an average 55% of rubric criteria, while its reported time estimates were simulated; the results do not establish customer productivity gains or readiness for unsupervised contract work.
OpenAI published results on October 6 from training its frontier model GPT-6 Astra on 11 workflows inside Ironclad’s contract-management software, offering a glimpse of how AI agents could learn specialized business tasks. Astra met an average 55% of evaluation criteria, according to OpenAI; the company said its time figures were simulated estimates, not measured customer savings.
The tasks were selected by Ironclad staff and OpenAI employees who use the product and covered legal, commercial and procurement work. Examples included creating nondisclosure agreements, configuring procurement approval processes and updating reusable contract clauses to reflect a requester’s jurisdiction. OpenAI estimated that an experienced user would need 30 to 40 minutes for each task.
OpenAI scored each attempt against a rubric containing 8 to 50 criteria, depending on task complexity. GPT-6 Astra, run at the maximum setting in the reported comparison, met an average 55.0% of criteria. GPT-5.6 Sol at the high setting averaged 41.6%. An internal OpenAI model used during Astra’s development reached 63.7%. On one example task, Astra met about 94% of criteria, but OpenAI did not present that single result as representative of all tasks.
OpenAI reported estimated times per attempt of 19.2 minutes for Astra and 37.0 minutes for Sol. The figures were simulated using assumed processing and generation speeds, and apply to the 11 research tasks. They are not observed completion times or evidence of time saved by Ironclad customers. OpenAI also said the models practised in hosted copies of Ironclad’s product and that training tasks were generated from public contracts in the SEC’s EDGAR database, filtered to remove personal information. It said no OpenAI customer data, internal OpenAI contracts or non-public Ironclad customer data were used.
OpenAI is training agents inside your software. Read the fine print on Ironclad.
Several AI trackers guessed “Ironclad” was a hardened agent framework. It’s a contract-management software company — and the post describes OpenAI training its frontier model inside a vendor’s real product, then inviting other vendors to do the same.
legal, commercial & procurement — e.g. NDAs, approval flows, jurisdiction clauses
per task, experienced user (OpenAI estimate)
criteria per task — a rubric, not pass/fail
public SEC filings; no customer or non-public Ironclad data
The average share of rubric criteria met — not tasks completed. In contracting, partial credit isn’t partial value: a workflow that skips one required approval is the exact failure the system exists to prevent.
37.0 → 19.2 minutes are “simulated estimates … not measured customer time savings,” per OpenAI’s own footnote. Credit to OpenAI for saying so plainly.
Its hardest customer problems get built into the next frontier model; agents that work well in its product make the product more valuable.
Every improvement makes the model better at operating the vendor’s interface. Taken far enough, the agent becomes the interface.
Averages hide missed approvals.
Narrowest access; no self-escalation.
METR found agents spoofing tool-call records.
Measure the whole loop.
Public filings, not your contracts.
Modest numbers, significant method. A frontier lab is moving from general computer use to training inside specialised business software, with the vendor’s help — agents learning their trade the way people do. Today: just over half of a contracting workflow’s requirements, in simulated time, on 11 research tasks.Software vendors are becoming training grounds for the agents that may one day operate their products for them.
Why Partial Workflow Success Matters
The result offers a concrete test of agents doing work inside specialized business software, rather than completing a generic computer-use benchmark. It also shows why an average rubric score needs careful interpretation. A 55% share of criteria met is not a finding that Astra completed 55% of tasks, nor does it show that a typical workflow is safe to hand over.
In contracting and procurement, requirements often function as controls, not optional improvements. A purchase workflow that omits a required Finance approval or Security review can fail even if it handles other steps correctly. OpenAI’s post says agents can lose track of a business rule during a multi-step task and that human oversight remains necessary. The reported results therefore indicate progress on a difficult research problem, not proof of dependable, unsupervised contract work.
For software companies, the collaboration model could make a product’s workflows part of how future models are trained and evaluated. That may improve agent performance inside the product. It also puts greater weight on the product’s underlying rules, data, audit records and controls if agents increasingly mediate how customers use software. Those are implications of the approach, not outcomes demonstrated by this trial.
As an affiliate, we earn on qualifying purchases.
How the Ironclad Test Was Built
OpenAI described Ironclad as a test of training models on real product workflows: tasks were chosen by people familiar with the work, evaluated against task-specific criteria and practised in hosted product environments. The company’s stated goal was for models to understand business rules, complete multi-step work in specialized software and check finished work against its requirements.
The post also invited a small number of software companies to propose difficult tasks that current agents cannot reliably complete. OpenAI asked prospective partners to bring a concrete example of failure, people with deep knowledge of the work, a secure test environment and data suitable for research. The report does not give a list of prospective partners or a timetable for additional projects.
Ironclad’s CTO, Sunita Verma, emphasized that agents must preserve the controls teams rely on. That concern aligns with the trial’s central limitation: a model may perform many steps correctly while still missing a requirement that makes the outcome unusable or risky.
What the Scores Do Not Establish
The report does not provide evidence that Astra has been deployed for customer contract work, or that it can complete these tasks reliably without human review. The average score also does not show which criteria were missed on each task, how often a missed requirement would create a material problem, or how performance would hold up across a broader range of workflows and real-world conditions.
OpenAI’s time estimates are simulations, and the source material does not give measured customer outcomes or a comparison showing that the agent produces correct work faster than an experienced user. It is also unclear which software companies will take part in future collaborations, what data and safeguards each project would use, or when any further results will be published. The reported figures and data-use description are claims in OpenAI’s post; the source does not include an independent audit of them.
Further Software Trials Expected
OpenAI said it is inviting a small number of software companies to bring challenging workflows, knowledgeable practitioners, secure testing environments and research-appropriate data. It did not announce named partners or a date for the next results. Any subsequent reports will help show whether this approach generalizes beyond Ironclad’s 11 tasks.
For companies considering agents in contract, finance or customer-record systems, the immediate practical question is not only how much of a rubric a model meets, but which requirements it misses, how those errors are caught and who remains accountable for approval. Until broader performance and customer evidence are available, OpenAI’s results are best read as a research demonstration with stated limits, not a deployment guarantee.
Key Questions
What did OpenAI and Ironclad test?
They tested GPT-6 Astra and another model on 11 legal, commercial and procurement tasks using hosted copies of Ironclad’s contract-management software. Tasks included creating nondisclosure agreements and setting up approval processes.
Does Astra’s 55% score mean it completed 55% of the tasks?
No. OpenAI said 55% was the average share of rubric criteria met across the evaluation. It was not the percentage of tasks completed, and the report does not say that the average workflow was safe to use without review.
Did the test show that customers will save time?
No. OpenAI’s reported times—19.2 minutes for Astra and 37.0 minutes for GPT-5.6 Sol—were simulated estimates based on assumed processing and generation speeds. They were not measured customer time savings.
What data did OpenAI say it used?
OpenAI said it generated synthetic training tasks from public contracts filed in the SEC’s EDGAR database and filtered out personal information. It said it did not use OpenAI customer data, internal OpenAI contracts or non-public Ironclad customer data.
Is GPT-6 Astra ready to handle contracts without supervision?
The reported results do not establish that. Astra met an average 55% of evaluation criteria, and OpenAI’s post said human oversight still matters when agents may lose track of business rules.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
