🔍 Read the full analysis: OpenAI Is Training Agents In Your Software. Ironclad’s Terms Matter on ThorstenMeyerAI.com
Prime made for students and young adults
- Fast, free delivery for dorm and study essentials
- Prime Video and Amazon Music included
- Member-only deals
TL;DR
OpenAI reported training GPT-6 Astra on 11 legal, commercial and procurement tasks in hosted copies of Ironclad’s contract-management software. Astra met an average 55% of task criteria, while estimated completion times were simulations, not measured customer savings; OpenAI says human oversight remains important.
OpenAI said on October 6 that it trained its frontier model GPT-6 Astra on 11 legal, commercial and procurement tasks in hosted copies of Ironclad’s contract-management software. The report describes a research effort to teach agents to follow business rules inside specialized products; Astra met an average 55% of rubric criteria, and OpenAI said human review remains necessary.
OpenAI and Ironclad selected tasks such as preparing nondisclosure agreements, configuring procurement approvals and updating a reusable contract clause to reflect a requester’s jurisdiction. OpenAI estimated that an experienced user would take 30 to 40 minutes per task. Each task was scored against between 8 and 50 criteria, depending on its complexity.
Ironclad provided hosted copies of its product for model practice. OpenAI said it generated synthetic training tasks from publicly filed contracts in the U.S. Securities and Exchange Commission’s EDGAR database, with personal information filtered out. The company said it did not use OpenAI customer data, its internal contracts or non-public Ironclad customer data.
In OpenAI’s comparison, GPT-6 Astra met an average 55.0% of criteria, compared with 41.6% for GPT-5.6 Sol at a high setting. An internal model used during Astra’s development reached 63.7%. OpenAI also reported a 19.2-minute estimated time per Astra attempt, against 37.0 minutes for GPT-5.6 Sol. The figures measure rubric coverage and simulated task time, not the percentage of tasks completed or verified customer productivity gains.
OpenAI is training agents inside your software. Read the fine print on Ironclad.
Several AI trackers guessed “Ironclad” was a hardened agent framework. It’s a contract-management software company — and the post describes OpenAI training its frontier model inside a vendor’s real product, then inviting other vendors to do the same.
legal, commercial & procurement — e.g. NDAs, approval flows, jurisdiction clauses
per task, experienced user (OpenAI estimate)
criteria per task — a rubric, not pass/fail
public SEC filings; no customer or non-public Ironclad data
The average share of rubric criteria met — not tasks completed. In contracting, partial credit isn’t partial value: a workflow that skips one required approval is the exact failure the system exists to prevent.
37.0 → 19.2 minutes are “simulated estimates … not measured customer time savings,” per OpenAI’s own footnote. Credit to OpenAI for saying so plainly.
Its hardest customer problems get built into the next frontier model; agents that work well in its product make the product more valuable.
Every improvement makes the model better at operating the vendor’s interface. Taken far enough, the agent becomes the interface.
Averages hide missed approvals.
Narrowest access; no self-escalation.
METR found agents spoofing tool-call records.
Measure the whole loop.
Public filings, not your contracts.
Modest numbers, significant method. A frontier lab is moving from general computer use to training inside specialised business software, with the vendor’s help — agents learning their trade the way people do. Today: just over half of a contracting workflow’s requirements, in simulated time, on 11 research tasks.Software vendors are becoming training grounds for the agents that may one day operate their products for them.
Why Contract Workflows Need Oversight
The results point to a possible change in how AI models are trained for work: rather than learning only from general examples, agents can practise tasks inside the software and workflows where businesses actually carry them out. OpenAI says the aim is for models to understand business rules, complete multi-step work and check results against requirements. It is also asking other software companies to bring difficult tasks for similar research.
But an average score of 55% of criteria met does not establish that a contract workflow is safe to run without review. A missed requirement could be a required Finance approval, a Security review or escalation of nonstandard language to Legal. In such cases, partial completion may still leave a process defective. OpenAI’s results are research measurements, not evidence that the agent can independently manage consequential business decisions.
For software vendors, there is a strategic trade-off. A model that works well inside a vendor’s product could make the product more useful and expose where its workflows need improvement. At the same time, customers may increasingly interact with the agent rather than the product’s screens. The long-term value of a platform may depend more on its business rules, records, audit trail and controls than on its interface alone. That is an implication of the partnership model, not a reported outcome of this test.
contract management software for legal teams
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
How the Ironclad Test Was Set Up
OpenAI’s October 6 post described the Ironclad work as training models to handle specialized software tasks and verify whether the resulting work met the original requirements. Ironclad is a contract-management software company, not the name of a new agent framework. The project used a defined set of 11 tasks selected by Ironclad staff and OpenAI employees who use the product.
The evaluation used task-specific criteria rather than a simple pass-or-fail measure. That makes the reported average useful as a measure of how many requirements the model satisfied, but it does not show which requirements were missed in every task or whether any particular workflow would pass a company’s deployment threshold. OpenAI also described a showcase task on which Astra met about 94% of criteria; that result is one task, not the overall average.
OpenAI’s time estimates need separate treatment from the rubric scores. The company said they were simulated using assumed processing and generation speeds, rather than measured time savings for customers. They cover the 11 research tasks and do not establish how quickly the model would perform across Ironclad’s broader range of workflows.
As an affiliate, we earn on qualifying purchases.
What the Scores Do Not Show
The report does not establish how Astra would perform across Ironclad’s full product or on live customer work. It also does not provide a breakdown here of the specific criteria the model missed on each task, or a threshold at which OpenAI or Ironclad would consider a workflow ready for deployment. The 55% average should not be read as a 55% task-completion rate.
OpenAI’s estimated task times are simulations, not observed customer outcomes, and the research does not show net time saved after a person checks and corrects the work. OpenAI said human oversight still matters. The source material also does not specify a commercial deployment schedule, the terms of prospective partnerships, or what additional data other software partners might contribute. OpenAI’s stated data limits apply to this Ironclad project and should not be assumed to define every future partnership.
AI legal document automation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What OpenAI’s Partner Call Requires
OpenAI said it is inviting a small number of software companies to work on tasks current agents cannot reliably complete. It asks prospective partners to bring a specific failure example, people with deep knowledge of the work, a secure test environment and data that can safely be used for research. The company has not identified the partners it may select or announced a timetable for further projects in the source material.
For businesses considering agents in contract, finance or customer-record systems, the next practical step is to ask vendors how performance is evaluated: which requirements were missed, how exceptions are handled, what audit records are kept and who approves consequential actions. Further results will need to show whether agents can satisfy complete workflows reliably, not only improve an average rubric score.
As an affiliate, we earn on qualifying purchases.
Key Questions
What did OpenAI test with Ironclad?
OpenAI tested a model on 11 legal, commercial and procurement tasks in hosted copies of Ironclad’s contract-management software, using criteria to score each task.
Did GPT-6 Astra complete 55% of the tasks?
No. OpenAI reported that Astra met an average 55% of the rubric criteria. That is not a task-completion rate, and the report does not say that the model completed 55% of tasks successfully.
Did the test show customers would save time?
No measured customer savings were reported. OpenAI said its 19.2-minute estimate for Astra was simulated using assumed processing and generation speeds and applied to the 11 research tasks.
What data did OpenAI say it used?
OpenAI said it created synthetic tasks from publicly filed SEC EDGAR contracts, filtered personal information, and did not use OpenAI customer data, OpenAI internal contracts or non-public Ironclad customer data.
Can businesses use the agent for contract work without review?
The reported results do not establish that. OpenAI said human oversight still matters, and the average criteria score does not show that every required approval or control was followed.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
