OpenAI Is Training Agents In Your Software. Ironclad’s Terms Matter
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: OpenAI Is Training Agents In Your Software. Ironclad’s Terms Matter on ThorstenMeyerAI.com

Age 18–24?Offer from Amazon

Prime made for students and young adults

  • Fast, free delivery for dorm and study essentials
  • Prime Video and Amazon Music included
  • Member-only deals
Try Prime for Young Adults Free trial for eligible 18–24 year olds
As an affiliate, we earn on qualifying purchases.

TL;DR

OpenAI reported training GPT-6 Astra on 11 legal, commercial and procurement tasks in hosted copies of Ironclad’s contract-management software. Astra met an average 55% of task criteria, while estimated completion times were simulations, not measured customer savings; OpenAI says human oversight remains important.

OpenAI said on October 6 that it trained its frontier model GPT-6 Astra on 11 legal, commercial and procurement tasks in hosted copies of Ironclad’s contract-management software. The report describes a research effort to teach agents to follow business rules inside specialized products; Astra met an average 55% of rubric criteria, and OpenAI said human review remains necessary.

OpenAI and Ironclad selected tasks such as preparing nondisclosure agreements, configuring procurement approvals and updating a reusable contract clause to reflect a requester’s jurisdiction. OpenAI estimated that an experienced user would take 30 to 40 minutes per task. Each task was scored against between 8 and 50 criteria, depending on its complexity.

Ironclad provided hosted copies of its product for model practice. OpenAI said it generated synthetic training tasks from publicly filed contracts in the U.S. Securities and Exchange Commission’s EDGAR database, with personal information filtered out. The company said it did not use OpenAI customer data, its internal contracts or non-public Ironclad customer data.

In OpenAI’s comparison, GPT-6 Astra met an average 55.0% of criteria, compared with 41.6% for GPT-5.6 Sol at a high setting. An internal model used during Astra’s development reached 63.7%. OpenAI also reported a 19.2-minute estimated time per Astra attempt, against 37.0 minutes for GPT-5.6 Sol. The figures measure rubric coverage and simulated task time, not the percentage of tasks completed or verified customer productivity gains.

At a glance
reportWhen: Published October 6; ongoing research
The developmentOpenAI described training a frontier model on workflows inside Ironclad’s contract-management product and invited other software companies to propose similar research partnerships.
OpenAI × Ironclad — Insights
AI Dispatch · Insights · 7 October 2026

OpenAI is training agents inside your software. Read the fine print on Ironclad.

Several AI trackers guessed “Ironclad” was a hardened agent framework. It’s a contract-management software company — and the post describes OpenAI training its frontier model inside a vendor’s real product, then inviting other vendors to do the same.

What they did
Tasks
11

legal, commercial & procurement — e.g. NDAs, approval flows, jurisdiction clauses

Human time
30–40m

per task, experienced user (OpenAI estimate)

Grading
8–50

criteria per task — a rubric, not pass/fail

Training data
EDGAR

public SEC filings; no customer or non-public Ironclad data

The results — and what the footnotes say
GPT-5.6 Sol (high) · criteria met41.6%
GPT-6 Astra (max) · criteria met55.0%
Internal model · criteria met63.7%
What “55%” means

The average share of rubric criteria met — not tasks completed. In contracting, partial credit isn’t partial value: a workflow that skips one required approval is the exact failure the system exists to prevent.

The time numbers are simulated

37.0 → 19.2 minutes are “simulated estimates … not measured customer time savings,” per OpenAI’s own footnote. Credit to OpenAI for saying so plainly.

~20 simulated minutes, ~half the criteria, and a human checks every requirement — vs 30–40 minutes for an expert done right. For now, the human is still the faster route to a correct workflow. The trend is the story.
The bigger story: software vendors as training grounds
Upside for the vendor

Its hardest customer problems get built into the next frontier model; agents that work well in its product make the product more valuable.

Risk for the vendor

Every improvement makes the model better at operating the vendor’s interface. Taken far enough, the agent becomes the interface.

The post frames it as showing why “a full contracting platform remains essential.” Winners will be vendors whose value is in rules, records and controls — not the screens an agent learns to click.
Five questions before letting agents into your systems of record
Which criteria failed?

Averages hide missed approvals.

What permissions?

Narrowest access; no self-escalation.

Tamper-proof logs?

METR found agents spoofing tool-call records.

Who checks, how long?

Measure the whole loop.

Whose training data?

Public filings, not your contracts.

The take

Modest numbers, significant method. A frontier lab is moving from general computer use to training inside specialised business software, with the vendor’s help — agents learning their trade the way people do. Today: just over half of a contracting workflow’s requirements, in simulated time, on 11 research tasks.Software vendors are becoming training grounds for the agents that may one day operate their products for them.

Source: OpenAI, “Advancing computer use with Ironclad” (6 Oct 2026) — tasks, criteria, EDGAR training data, 55.0% vs 41.6%, 19.2 vs 37.0 simulated minutes, 63.7% internal model, simulation footnote, collaboration invitation. Mischaracterisations of “Ironclad” in automated AI-news trackers (7 Oct 2026). METR investigation as covered here. Analysis is the author’s.
thorstenmeyerai.com

Why Contract Workflows Need Oversight

The results point to a possible change in how AI models are trained for work: rather than learning only from general examples, agents can practise tasks inside the software and workflows where businesses actually carry them out. OpenAI says the aim is for models to understand business rules, complete multi-step work and check results against requirements. It is also asking other software companies to bring difficult tasks for similar research.

But an average score of 55% of criteria met does not establish that a contract workflow is safe to run without review. A missed requirement could be a required Finance approval, a Security review or escalation of nonstandard language to Legal. In such cases, partial completion may still leave a process defective. OpenAI’s results are research measurements, not evidence that the agent can independently manage consequential business decisions.

For software vendors, there is a strategic trade-off. A model that works well inside a vendor’s product could make the product more useful and expose where its workflows need improvement. At the same time, customers may increasingly interact with the agent rather than the product’s screens. The long-term value of a platform may depend more on its business rules, records, audit trail and controls than on its interface alone. That is an implication of the partnership model, not a reported outcome of this test.

Amazon

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How the Ironclad Test Was Set Up

OpenAI’s October 6 post described the Ironclad work as training models to handle specialized software tasks and verify whether the resulting work met the original requirements. Ironclad is a contract-management software company, not the name of a new agent framework. The project used a defined set of 11 tasks selected by Ironclad staff and OpenAI employees who use the product.

The evaluation used task-specific criteria rather than a simple pass-or-fail measure. That makes the reported average useful as a measure of how many requirements the model satisfied, but it does not show which requirements were missed in every task or whether any particular workflow would pass a company’s deployment threshold. OpenAI also described a showcase task on which Astra met about 94% of criteria; that result is one task, not the overall average.

OpenAI’s time estimates need separate treatment from the rubric scores. The company said they were simulated using assumed processing and generation speeds, rather than measured time savings for customers. They cover the 11 research tasks and do not establish how quickly the model would perform across Ironclad’s broader range of workflows.

Amazon

AI-powered contract review tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What the Scores Do Not Show

The report does not establish how Astra would perform across Ironclad’s full product or on live customer work. It also does not provide a breakdown here of the specific criteria the model missed on each task, or a threshold at which OpenAI or Ironclad would consider a workflow ready for deployment. The 55% average should not be read as a 55% task-completion rate.

OpenAI’s estimated task times are simulations, not observed customer outcomes, and the research does not show net time saved after a person checks and corrects the work. OpenAI said human oversight still matters. The source material also does not specify a commercial deployment schedule, the terms of prospective partnerships, or what additional data other software partners might contribute. OpenAI’s stated data limits apply to this Ironclad project and should not be assumed to define every future partnership.

Amazon

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What OpenAI’s Partner Call Requires

OpenAI said it is inviting a small number of software companies to work on tasks current agents cannot reliably complete. It asks prospective partners to bring a specific failure example, people with deep knowledge of the work, a secure test environment and data that can safely be used for research. The company has not identified the partners it may select or announced a timetable for further projects in the source material.

For businesses considering agents in contract, finance or customer-record systems, the next practical step is to ask vendors how performance is evaluated: which requirements were missed, how exceptions are handled, what audit records are kept and who approves consequential actions. Further results will need to show whether agents can satisfy complete workflows reliably, not only improve an average rubric score.

Amazon

AI contract analysis tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What did OpenAI test with Ironclad?

OpenAI tested a model on 11 legal, commercial and procurement tasks in hosted copies of Ironclad’s contract-management software, using criteria to score each task.

Did GPT-6 Astra complete 55% of the tasks?

No. OpenAI reported that Astra met an average 55% of the rubric criteria. That is not a task-completion rate, and the report does not say that the model completed 55% of tasks successfully.

Did the test show customers would save time?

No measured customer savings were reported. OpenAI said its 19.2-minute estimate for Astra was simulated using assumed processing and generation speeds and applied to the 11 research tasks.

What data did OpenAI say it used?

OpenAI said it created synthetic tasks from publicly filed SEC EDGAR contracts, filtered personal information, and did not use OpenAI customer data, OpenAI internal contracts or non-public Ironclad customer data.

Can businesses use the agent for contract work without review?

The reported results do not establish that. OpenAI said human oversight still matters, and the average criteria score does not show that every required approval or control was followed.

Source: ThorstenMeyerAI.com

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Thrymvault: A System Around Your Content

Thrymvault launches as a private, self-hosted workspace integrating documents, databases, AI prompts, and client portals to streamline content creation.

Capital: The Lever Beneath the Levers

Analysis of how private and public funding shapes AI infrastructure, revealing the fragile cycle of capital fueling AI giants and risks ahead.

AI Is the Alibi. The Reorg Is the Signal.

Coinbase’s recent layoffs are officially linked to AI-driven restructuring, but underlying market pressures suggest the true driver is crypto downturn and cost-cutting.

Home signal monitor: Mortgage Rates Inch to Another 6-Week Low

Mortgage rates have declined to their lowest point in six weeks, signaling potential shifts in the housing market and borrowing costs.