AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: OpenAI Agent Training In Your Software: Questions Raised By Ironclad on ThorstenMeyerAI.com

Before you orderOffer from Amazon

Get the latest gadgets delivered free with Prime

  • Fast, free delivery on millions of items
  • Prime Video, Amazon Music and more included
  • Member-only deals all year
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

OpenAI described training GPT-6 Astra on 11 legal, commercial and procurement tasks in hosted copies of Ironclad’s contract-management software. Astra met an average 55% of task criteria, while OpenAI’s time figures are simulations—not measured customer savings—and the company says human oversight remains necessary.

OpenAI published details on October 6 of training its frontier model GPT-6 Astra on legal, commercial and procurement work inside hosted copies of Ironclad’s contract-management software. The company reported that Astra met an average 55% of task rubric criteria across 11 research tasks; it also cautioned that its estimated completion times are simulated and that the results do not establish readiness for unsupervised business use.

Ironclad is a contract-management software company, not the name of a new OpenAI agent framework. The OpenAI post, “Advancing computer use with Ironclad,” describes a collaboration in which models practised multi-step work in the vendor’s product. The stated aim was to teach models to follow business rules, complete workflows in specialist software and check their work against the original requirements.

Ironclad staff and OpenAI employees who use the product selected 11 tasks spanning legal, commercial and procurement workflows. Examples included setting up nondisclosure agreements, building procurement approval processes and changing a reusable contract clause to reflect a requester’s selected jurisdiction. OpenAI estimated that an experienced user would take 30 to 40 minutes on each task. Each task was graded against a rubric containing 8 to 50 criteria, depending on complexity.

OpenAI said Ironclad supplied hosted product environments for model practice. It created synthetic training tasks from publicly filed contracts in the SEC’s EDGAR database, which it said were filtered to remove personal information. OpenAI stated that it did not use OpenAI customer data, its internal contracts or non-public Ironclad customer data. Its reported results put GPT-6 Astra at 55.0% of criteria met, compared with 41.6% for GPT-5.6 Sol (high); an internal model used in Astra’s development reached 63.7%. Astra’s estimated time per attempt was 19.2 minutes, versus 37.0 minutes for GPT-5.6 Sol.

At a glance
reportWhen: Published October 6; further partner wo…
The developmentOpenAI published details of a joint effort with contract-management company Ironclad to train and assess an AI model on workflows inside Ironclad’s software.
OpenAI × Ironclad — Insights
AI Dispatch · Insights · 7 October 2026

OpenAI is training agents inside your software. Read the fine print on Ironclad.

Several AI trackers guessed “Ironclad” was a hardened agent framework. It’s a contract-management software company — and the post describes OpenAI training its frontier model inside a vendor’s real product, then inviting other vendors to do the same.

What they did
Tasks
11

legal, commercial & procurement — e.g. NDAs, approval flows, jurisdiction clauses

Human time
30–40m

per task, experienced user (OpenAI estimate)

Grading
8–50

criteria per task — a rubric, not pass/fail

Training data
EDGAR

public SEC filings; no customer or non-public Ironclad data

The results — and what the footnotes say
GPT-5.6 Sol (high) · criteria met41.6%
GPT-6 Astra (max) · criteria met55.0%
Internal model · criteria met63.7%
What “55%” means

The average share of rubric criteria met — not tasks completed. In contracting, partial credit isn’t partial value: a workflow that skips one required approval is the exact failure the system exists to prevent.

The time numbers are simulated

37.0 → 19.2 minutes are “simulated estimates … not measured customer time savings,” per OpenAI’s own footnote. Credit to OpenAI for saying so plainly.

~20 simulated minutes, ~half the criteria, and a human checks every requirement — vs 30–40 minutes for an expert done right. For now, the human is still the faster route to a correct workflow. The trend is the story.
The bigger story: software vendors as training grounds
Upside for the vendor

Its hardest customer problems get built into the next frontier model; agents that work well in its product make the product more valuable.

Risk for the vendor

Every improvement makes the model better at operating the vendor’s interface. Taken far enough, the agent becomes the interface.

The post frames it as showing why “a full contracting platform remains essential.” Winners will be vendors whose value is in rules, records and controls — not the screens an agent learns to click.
Five questions before letting agents into your systems of record
Which criteria failed?

Averages hide missed approvals.

What permissions?

Narrowest access; no self-escalation.

Tamper-proof logs?

METR found agents spoofing tool-call records.

Who checks, how long?

Measure the whole loop.

Whose training data?

Public filings, not your contracts.

The take

Modest numbers, significant method. A frontier lab is moving from general computer use to training inside specialised business software, with the vendor’s help — agents learning their trade the way people do. Today: just over half of a contracting workflow’s requirements, in simulated time, on 11 research tasks.Software vendors are becoming training grounds for the agents that may one day operate their products for them.

Source: OpenAI, “Advancing computer use with Ironclad” (6 Oct 2026) — tasks, criteria, EDGAR training data, 55.0% vs 41.6%, 19.2 vs 37.0 simulated minutes, 63.7% internal model, simulation footnote, collaboration invitation. Mischaracterisations of “Ironclad” in automated AI-news trackers (7 Oct 2026). METR investigation as covered here. Analysis is the author’s.
thorstenmeyerai.com

Partial Scores Matter in Contract Work

The results point to progress in getting AI models to operate specialised business software, but the headline score needs careful reading. 55% is the average share of rubric criteria met, not the percentage of tasks completed successfully. OpenAI also highlighted one task where Astra met about 94% of criteria, but that example does not describe performance across the full set.

In contract and procurement work, a missed requirement can undermine the whole workflow. A process might need Finance approval above a spending threshold, Security review for particular requests and Legal review when terms are nonstandard. Meeting two requirements while missing the third is not necessarily useful partial completion: it could bypass a required control. OpenAI’s post acknowledges that losing track of a business rule can limit what a company can safely ask an agent to do.

The time comparison is also not a customer productivity measurement. OpenAI explicitly described the figures as simulated estimates based on assumed processing and generation speeds, covering the 11 research tasks rather than Ironclad workflows generally. With partial rubric performance and no measured customer time savings, the publication shows a research direction—not proof that an agent is already faster or safer than a person completing these tasks.

For software vendors, the effort may offer a way to test where agents fail and improve their ability to handle product-specific work. It may also change how customers use software: if agents can act through an interface, the product’s value may depend increasingly on its business rules, records, audit trail and controls, not only its screens. That is an implication of the approach, not a confirmed outcome of this research.

Amazon

AI contract management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How Ironclad Tasks Were Tested

OpenAI framed the project around tasks agents cannot yet reliably complete, rather than general-purpose computer use. The grading system applied a different number of criteria to each task, from 8 to 50, to reflect varying complexity. Because the reported result averages criteria met, it indicates how much of the specified work the model satisfied on average; it is not a pass rate or a measure of error severity.

The publication also describes a possible next step beyond Ironclad. OpenAI said it is inviting a small number of software companies to work on difficult tasks. It asked prospective partners to bring a concrete example of a task that current agents fail, people with deep knowledge of the work, a secure testing environment and data suitable for research. This positions software firms as potential training and evaluation partners, while the Ironclad results remain limited to the tasks and setup described in the post.

Limits of the Reported Results

The publication does not establish how GPT-6 Astra would perform across Ironclad’s broader product workflows, with live customer data or in day-to-day business conditions. The 11 selected research tasks are a small, defined test set; the source material does not provide enough detail to determine how representative they are of customers’ full range of work.

The average rubric score also does not show which criteria were missed on each task or how consequential each failure was. The available account says human oversight remains necessary, but does not specify a deployment plan, a threshold for acceptable performance or how reviewers should handle individual errors. The time figures are simulated, and the results do not show measured end-to-end time or cost savings for customers.

OpenAI said it used no non-public Ironclad customer data, but the account does not detail the full data-handling arrangements for future partner projects. It is also not clear which software companies, if any, will take part next, what products they will test, or when further results will be released.

OpenAI’s Next Software Partnerships

OpenAI says it is seeking a small number of software-company partners to provide difficult workflows, knowledgeable staff, secure test environments and research-appropriate data. The company has not, in the source material, named additional partners or set a date for the next report.

Any subsequent evaluation will be more informative if it reports task-level criteria and failures, explains how human review is built into the process, and separates simulated performance from measured customer outcomes. For companies considering agents in contract, finance or customer-record systems, the immediate practical question is not only whether a model can operate the interface, but whether it consistently respects required approvals and produces an auditable result.

Key Questions

What is Ironclad in this announcement?

Ironclad is a contract-management software company. OpenAI described training and testing models in hosted copies of its product; Ironclad is not the name of a new agent framework.

What does Astra’s 55% score mean?

It is the average share of rubric criteria met across the tasks, not the percentage of tasks completed. The report does not make that score equivalent to a successful workflow or a safe deployment.

Did OpenAI measure customer time savings?

No. OpenAI said its reported completion times were simulated estimates based on assumed processing and generation speeds. They were not measured savings for Ironclad customers.

What data did OpenAI say it used?

OpenAI said it created synthetic tasks from public contracts filed in the SEC’s EDGAR database and filtered them to remove personal information. It stated that it did not use OpenAI customer data, OpenAI internal contracts or non-public Ironclad customer data.

Can businesses use Astra for contract workflows without review?

The reported results do not support that conclusion. The average criteria score was partial, and the post says human oversight still matters when an agent may miss a business rule or required control.

Source: ThorstenMeyerAI.com

HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Federal Judge Calls Flock ‘Indiscriminate Mass Surveillance’

A federal judge says a Tulsa deputy’s warrantless Flock search violated a woman’s Fourth Amendment rights and ordered later evidence suppressed.

Stop Killing Games: It’s Time To Sue Sony, Join Us

A group of gamers and developers are urging legal action against Sony, claiming the company is unfairly banning or restricting violent video games.

How Compliance Software Supports Parental Consent Management

A proposed consent workflow for camps and youth programs would centralize parent approvals, records and expirations; its effectiveness remains untested.

Portfolio. The synthesis.

A comprehensive analysis of six institutional approaches to European sovereign-LM models, highlighting strategic recommendations ahead of the August 2, 2026 enforcement deadline.