🔍 Read the full analysis: OpenAI Agent Training In Your Software: Questions Raised By Ironclad on ThorstenMeyerAI.com
Get the latest gadgets delivered free with Prime
- Fast, free delivery on millions of items
- Prime Video, Amazon Music and more included
- Member-only deals all year
TL;DR
OpenAI described training GPT-6 Astra on 11 legal, commercial and procurement tasks in hosted copies of Ironclad’s contract-management software. Astra met an average 55% of task criteria, while OpenAI’s time figures are simulations—not measured customer savings—and the company says human oversight remains necessary.
OpenAI published details on October 6 of training its frontier model GPT-6 Astra on legal, commercial and procurement work inside hosted copies of Ironclad’s contract-management software. The company reported that Astra met an average 55% of task rubric criteria across 11 research tasks; it also cautioned that its estimated completion times are simulated and that the results do not establish readiness for unsupervised business use.
Ironclad is a contract-management software company, not the name of a new OpenAI agent framework. The OpenAI post, “Advancing computer use with Ironclad,” describes a collaboration in which models practised multi-step work in the vendor’s product. The stated aim was to teach models to follow business rules, complete workflows in specialist software and check their work against the original requirements.
Ironclad staff and OpenAI employees who use the product selected 11 tasks spanning legal, commercial and procurement workflows. Examples included setting up nondisclosure agreements, building procurement approval processes and changing a reusable contract clause to reflect a requester’s selected jurisdiction. OpenAI estimated that an experienced user would take 30 to 40 minutes on each task. Each task was graded against a rubric containing 8 to 50 criteria, depending on complexity.
OpenAI said Ironclad supplied hosted product environments for model practice. It created synthetic training tasks from publicly filed contracts in the SEC’s EDGAR database, which it said were filtered to remove personal information. OpenAI stated that it did not use OpenAI customer data, its internal contracts or non-public Ironclad customer data. Its reported results put GPT-6 Astra at 55.0% of criteria met, compared with 41.6% for GPT-5.6 Sol (high); an internal model used in Astra’s development reached 63.7%. Astra’s estimated time per attempt was 19.2 minutes, versus 37.0 minutes for GPT-5.6 Sol.
OpenAI is training agents inside your software. Read the fine print on Ironclad.
Several AI trackers guessed “Ironclad” was a hardened agent framework. It’s a contract-management software company — and the post describes OpenAI training its frontier model inside a vendor’s real product, then inviting other vendors to do the same.
legal, commercial & procurement — e.g. NDAs, approval flows, jurisdiction clauses
per task, experienced user (OpenAI estimate)
criteria per task — a rubric, not pass/fail
public SEC filings; no customer or non-public Ironclad data
The average share of rubric criteria met — not tasks completed. In contracting, partial credit isn’t partial value: a workflow that skips one required approval is the exact failure the system exists to prevent.
37.0 → 19.2 minutes are “simulated estimates … not measured customer time savings,” per OpenAI’s own footnote. Credit to OpenAI for saying so plainly.
Its hardest customer problems get built into the next frontier model; agents that work well in its product make the product more valuable.
Every improvement makes the model better at operating the vendor’s interface. Taken far enough, the agent becomes the interface.
Averages hide missed approvals.
Narrowest access; no self-escalation.
METR found agents spoofing tool-call records.
Measure the whole loop.
Public filings, not your contracts.
Modest numbers, significant method. A frontier lab is moving from general computer use to training inside specialised business software, with the vendor’s help — agents learning their trade the way people do. Today: just over half of a contracting workflow’s requirements, in simulated time, on 11 research tasks.Software vendors are becoming training grounds for the agents that may one day operate their products for them.
Partial Scores Matter in Contract Work
The results point to progress in getting AI models to operate specialised business software, but the headline score needs careful reading. 55% is the average share of rubric criteria met, not the percentage of tasks completed successfully. OpenAI also highlighted one task where Astra met about 94% of criteria, but that example does not describe performance across the full set.
In contract and procurement work, a missed requirement can undermine the whole workflow. A process might need Finance approval above a spending threshold, Security review for particular requests and Legal review when terms are nonstandard. Meeting two requirements while missing the third is not necessarily useful partial completion: it could bypass a required control. OpenAI’s post acknowledges that losing track of a business rule can limit what a company can safely ask an agent to do.
The time comparison is also not a customer productivity measurement. OpenAI explicitly described the figures as simulated estimates based on assumed processing and generation speeds, covering the 11 research tasks rather than Ironclad workflows generally. With partial rubric performance and no measured customer time savings, the publication shows a research direction—not proof that an agent is already faster or safer than a person completing these tasks.
For software vendors, the effort may offer a way to test where agents fail and improve their ability to handle product-specific work. It may also change how customers use software: if agents can act through an interface, the product’s value may depend increasingly on its business rules, records, audit trail and controls, not only its screens. That is an implication of the approach, not a confirmed outcome of this research.
As an affiliate, we earn on qualifying purchases.
How Ironclad Tasks Were Tested
OpenAI framed the project around tasks agents cannot yet reliably complete, rather than general-purpose computer use. The grading system applied a different number of criteria to each task, from 8 to 50, to reflect varying complexity. Because the reported result averages criteria met, it indicates how much of the specified work the model satisfied on average; it is not a pass rate or a measure of error severity.
The publication also describes a possible next step beyond Ironclad. OpenAI said it is inviting a small number of software companies to work on difficult tasks. It asked prospective partners to bring a concrete example of a task that current agents fail, people with deep knowledge of the work, a secure testing environment and data suitable for research. This positions software firms as potential training and evaluation partners, while the Ironclad results remain limited to the tasks and setup described in the post.
Limits of the Reported Results
The publication does not establish how GPT-6 Astra would perform across Ironclad’s broader product workflows, with live customer data or in day-to-day business conditions. The 11 selected research tasks are a small, defined test set; the source material does not provide enough detail to determine how representative they are of customers’ full range of work.
The average rubric score also does not show which criteria were missed on each task or how consequential each failure was. The available account says human oversight remains necessary, but does not specify a deployment plan, a threshold for acceptable performance or how reviewers should handle individual errors. The time figures are simulated, and the results do not show measured end-to-end time or cost savings for customers.
OpenAI said it used no non-public Ironclad customer data, but the account does not detail the full data-handling arrangements for future partner projects. It is also not clear which software companies, if any, will take part next, what products they will test, or when further results will be released.
OpenAI’s Next Software Partnerships
OpenAI says it is seeking a small number of software-company partners to provide difficult workflows, knowledgeable staff, secure test environments and research-appropriate data. The company has not, in the source material, named additional partners or set a date for the next report.
Any subsequent evaluation will be more informative if it reports task-level criteria and failures, explains how human review is built into the process, and separates simulated performance from measured customer outcomes. For companies considering agents in contract, finance or customer-record systems, the immediate practical question is not only whether a model can operate the interface, but whether it consistently respects required approvals and produces an auditable result.
Key Questions
What is Ironclad in this announcement?
Ironclad is a contract-management software company. OpenAI described training and testing models in hosted copies of its product; Ironclad is not the name of a new agent framework.
What does Astra’s 55% score mean?
It is the average share of rubric criteria met across the tasks, not the percentage of tasks completed. The report does not make that score equivalent to a successful workflow or a safe deployment.
Did OpenAI measure customer time savings?
No. OpenAI said its reported completion times were simulated estimates based on assumed processing and generation speeds. They were not measured savings for Ironclad customers.
What data did OpenAI say it used?
OpenAI said it created synthetic tasks from public contracts filed in the SEC’s EDGAR database and filtered them to remove personal information. It stated that it did not use OpenAI customer data, OpenAI internal contracts or non-public Ironclad customer data.
Can businesses use Astra for contract workflows without review?
The reported results do not support that conclusion. The average criteria score was partial, and the post says human oversight still matters when an agent may miss a business rule or required control.
Source: ThorstenMeyerAI.com
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
