AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Why Ironclad’s Fine Print Matters As OpenAI Trains Agents In Software on ThorstenMeyerAI.com

Age 18–24?Offer from Amazon

Prime made for students and young adults

  • Fast, free delivery for dorm and study essentials
  • Prime Video and Amazon Music included
  • Member-only deals
Try Prime for Young Adults Free trial for eligible 18–24 year olds
As an affiliate, we earn on qualifying purchases.

TL;DR

OpenAI’s October 6 report describes training GPT-6 Astra on 11 legal, commercial and procurement tasks inside hosted copies of Ironclad’s contract-management software. Astra met an average 55% of rubric criteria, while its reported time estimates were simulated; the results do not establish customer productivity gains or readiness for unsupervised contract work.

OpenAI published results on October 6 from training its frontier model GPT-6 Astra on 11 workflows inside Ironclad’s contract-management software, offering a glimpse of how AI agents could learn specialized business tasks. Astra met an average 55% of evaluation criteria, according to OpenAI; the company said its time figures were simulated estimates, not measured customer savings.

The tasks were selected by Ironclad staff and OpenAI employees who use the product and covered legal, commercial and procurement work. Examples included creating nondisclosure agreements, configuring procurement approval processes and updating reusable contract clauses to reflect a requester’s jurisdiction. OpenAI estimated that an experienced user would need 30 to 40 minutes for each task.

OpenAI scored each attempt against a rubric containing 8 to 50 criteria, depending on task complexity. GPT-6 Astra, run at the maximum setting in the reported comparison, met an average 55.0% of criteria. GPT-5.6 Sol at the high setting averaged 41.6%. An internal OpenAI model used during Astra’s development reached 63.7%. On one example task, Astra met about 94% of criteria, but OpenAI did not present that single result as representative of all tasks.

OpenAI reported estimated times per attempt of 19.2 minutes for Astra and 37.0 minutes for Sol. The figures were simulated using assumed processing and generation speeds, and apply to the 11 research tasks. They are not observed completion times or evidence of time saved by Ironclad customers. OpenAI also said the models practised in hosted copies of Ironclad’s product and that training tasks were generated from public contracts in the SEC’s EDGAR database, filtered to remove personal information. It said no OpenAI customer data, internal OpenAI contracts or non-public Ironclad customer data were used.

At a glance
reportWhen: Published October 6; further software-c…
The developmentOpenAI published results from a collaboration with Ironclad that trained a frontier model on professional workflows inside the contract-management product and invited other software companies to propose similar research tasks.
OpenAI × Ironclad — Insights
AI Dispatch · Insights · 7 October 2026

OpenAI is training agents inside your software. Read the fine print on Ironclad.

Several AI trackers guessed “Ironclad” was a hardened agent framework. It’s a contract-management software company — and the post describes OpenAI training its frontier model inside a vendor’s real product, then inviting other vendors to do the same.

What they did
Tasks
11

legal, commercial & procurement — e.g. NDAs, approval flows, jurisdiction clauses

Human time
30–40m

per task, experienced user (OpenAI estimate)

Grading
8–50

criteria per task — a rubric, not pass/fail

Training data
EDGAR

public SEC filings; no customer or non-public Ironclad data

The results — and what the footnotes say
GPT-5.6 Sol (high) · criteria met41.6%
GPT-6 Astra (max) · criteria met55.0%
Internal model · criteria met63.7%
What “55%” means

The average share of rubric criteria met — not tasks completed. In contracting, partial credit isn’t partial value: a workflow that skips one required approval is the exact failure the system exists to prevent.

The time numbers are simulated

37.0 → 19.2 minutes are “simulated estimates … not measured customer time savings,” per OpenAI’s own footnote. Credit to OpenAI for saying so plainly.

~20 simulated minutes, ~half the criteria, and a human checks every requirement — vs 30–40 minutes for an expert done right. For now, the human is still the faster route to a correct workflow. The trend is the story.
The bigger story: software vendors as training grounds
Upside for the vendor

Its hardest customer problems get built into the next frontier model; agents that work well in its product make the product more valuable.

Risk for the vendor

Every improvement makes the model better at operating the vendor’s interface. Taken far enough, the agent becomes the interface.

The post frames it as showing why “a full contracting platform remains essential.” Winners will be vendors whose value is in rules, records and controls — not the screens an agent learns to click.
Five questions before letting agents into your systems of record
Which criteria failed?

Averages hide missed approvals.

What permissions?

Narrowest access; no self-escalation.

Tamper-proof logs?

METR found agents spoofing tool-call records.

Who checks, how long?

Measure the whole loop.

Whose training data?

Public filings, not your contracts.

The take

Modest numbers, significant method. A frontier lab is moving from general computer use to training inside specialised business software, with the vendor’s help — agents learning their trade the way people do. Today: just over half of a contracting workflow’s requirements, in simulated time, on 11 research tasks.Software vendors are becoming training grounds for the agents that may one day operate their products for them.

Source: OpenAI, “Advancing computer use with Ironclad” (6 Oct 2026) — tasks, criteria, EDGAR training data, 55.0% vs 41.6%, 19.2 vs 37.0 simulated minutes, 63.7% internal model, simulation footnote, collaboration invitation. Mischaracterisations of “Ironclad” in automated AI-news trackers (7 Oct 2026). METR investigation as covered here. Analysis is the author’s.
thorstenmeyerai.com

Why Partial Workflow Success Matters

The result offers a concrete test of agents doing work inside specialized business software, rather than completing a generic computer-use benchmark. It also shows why an average rubric score needs careful interpretation. A 55% share of criteria met is not a finding that Astra completed 55% of tasks, nor does it show that a typical workflow is safe to hand over.

In contracting and procurement, requirements often function as controls, not optional improvements. A purchase workflow that omits a required Finance approval or Security review can fail even if it handles other steps correctly. OpenAI’s post says agents can lose track of a business rule during a multi-step task and that human oversight remains necessary. The reported results therefore indicate progress on a difficult research problem, not proof of dependable, unsupervised contract work.

For software companies, the collaboration model could make a product’s workflows part of how future models are trained and evaluated. That may improve agent performance inside the product. It also puts greater weight on the product’s underlying rules, data, audit records and controls if agents increasingly mediate how customers use software. Those are implications of the approach, not outcomes demonstrated by this trial.

Amazon

contract management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How the Ironclad Test Was Built

OpenAI described Ironclad as a test of training models on real product workflows: tasks were chosen by people familiar with the work, evaluated against task-specific criteria and practised in hosted product environments. The company’s stated goal was for models to understand business rules, complete multi-step work in specialized software and check finished work against its requirements.

The post also invited a small number of software companies to propose difficult tasks that current agents cannot reliably complete. OpenAI asked prospective partners to bring a concrete example of failure, people with deep knowledge of the work, a secure test environment and data suitable for research. The report does not give a list of prospective partners or a timetable for additional projects.

Ironclad’s CTO, Sunita Verma, emphasized that agents must preserve the controls teams rely on. That concern aligns with the trial’s central limitation: a model may perform many steps correctly while still missing a requirement that makes the outcome unusable or risky.

What the Scores Do Not Establish

The report does not provide evidence that Astra has been deployed for customer contract work, or that it can complete these tasks reliably without human review. The average score also does not show which criteria were missed on each task, how often a missed requirement would create a material problem, or how performance would hold up across a broader range of workflows and real-world conditions.

OpenAI’s time estimates are simulations, and the source material does not give measured customer outcomes or a comparison showing that the agent produces correct work faster than an experienced user. It is also unclear which software companies will take part in future collaborations, what data and safeguards each project would use, or when any further results will be published. The reported figures and data-use description are claims in OpenAI’s post; the source does not include an independent audit of them.

Further Software Trials Expected

OpenAI said it is inviting a small number of software companies to bring challenging workflows, knowledgeable practitioners, secure testing environments and research-appropriate data. It did not announce named partners or a date for the next results. Any subsequent reports will help show whether this approach generalizes beyond Ironclad’s 11 tasks.

For companies considering agents in contract, finance or customer-record systems, the immediate practical question is not only how much of a rubric a model meets, but which requirements it misses, how those errors are caught and who remains accountable for approval. Until broader performance and customer evidence are available, OpenAI’s results are best read as a research demonstration with stated limits, not a deployment guarantee.

Key Questions

What did OpenAI and Ironclad test?

They tested GPT-6 Astra and another model on 11 legal, commercial and procurement tasks using hosted copies of Ironclad’s contract-management software. Tasks included creating nondisclosure agreements and setting up approval processes.

Does Astra’s 55% score mean it completed 55% of the tasks?

No. OpenAI said 55% was the average share of rubric criteria met across the evaluation. It was not the percentage of tasks completed, and the report does not say that the average workflow was safe to use without review.

Did the test show that customers will save time?

No. OpenAI’s reported times—19.2 minutes for Astra and 37.0 minutes for GPT-5.6 Sol—were simulated estimates based on assumed processing and generation speeds. They were not measured customer time savings.

What data did OpenAI say it used?

OpenAI said it generated synthetic training tasks from public contracts filed in the SEC’s EDGAR database and filtered out personal information. It said it did not use OpenAI customer data, internal OpenAI contracts or non-public Ironclad customer data.

Is GPT-6 Astra ready to handle contracts without supervision?

The reported results do not establish that. Astra met an average 55% of evaluation criteria, and OpenAI’s post said human oversight still matters when agents may lose track of business rules.

Source: ThorstenMeyerAI.com

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Access To Justice Tech: Grammarly As A Self-Help Legal Tool

A new AI-powered drafting tool, inspired by Grammarly, is being developed to assist non-lawyer litigants and small businesses in creating court documents accurately.

Loan covenant calendar for bootstrapped companies

A new workflow tool is being tested to help small, bootstrapped companies manage loan covenant obligations more effectively, addressing common reporting challenges.

Defense Cybersecurity Compliance: From Readiness To Certification

A proposed readiness workflow targets small DoD contractors facing CMMC Level 2 requirements, but demand and compliance claims need validation.

Improving B2B SaaS Procurement With Seamless Vendor Approval Processes

A new vendor approval process for mid-market companies streamlines onboarding through a staged, email-based system, reducing cycle times and increasing security.