Understanding Ironclad’s Fine Print On OpenAI’s Agents And Your Software
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Understanding Ironclad’s Fine Print On OpenAI’s Agents And Your Software on ThorstenMeyerAI.com

Buying for a business?Offer from Amazon

Get business pricing on monitors, keyboards and dev gear

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

TL;DR

OpenAI described training and evaluating a frontier model on 11 legal, commercial and procurement tasks inside contract-management software company Ironclad’s product. GPT-6 Astra met an average 55% of task criteria, while its estimated completion time was simulated rather than measured in customer use. OpenAI is inviting other software companies to participate, but the results do not establish that agents are ready to run business workflows without human review.

OpenAI said on October 6 that it trained and evaluated its frontier model GPT-6 Astra on professional workflows inside Ironclad’s contract-management software, reporting an average of 55% of rubric criteria met across 11 tasks. The results offer a view of how AI agents may learn to use specialized business software, but OpenAI’s time estimates were simulated and the company says human oversight remains necessary.

The tasks were selected by Ironclad staff and OpenAI employees who use the product. They covered legal, commercial and procurement work, including creating nondisclosure agreements, setting up procurement approval processes and adjusting a reusable contract clause to match a requester’s jurisdiction. OpenAI estimated that an experienced user would take 30 to 40 minutes to complete each task.

Tasks were graded against 8 to 50 criteria, depending on complexity. OpenAI reported that GPT-6 Astra met an average 55% of those criteria, compared with 41.6% for GPT-5.6 Sol at its high setting. An internal OpenAI model used during Astra’s development reached 63.7%. On one showcased task, Astra met about 94% of the criteria. The average score is a share of requirements met, not the percentage of tasks completed successfully.

Ironclad provided hosted copies of its product for the models to use. OpenAI said it generated synthetic training tasks from contracts filed publicly in the U.S. Securities and Exchange Commission’s EDGAR database, with personal information removed. The company said it used no OpenAI customer data, OpenAI internal contracts or non-public Ironclad customer data. OpenAI also reported estimated times of 19.2 minutes for Astra and 37 minutes for GPT-5.6 Sol, while saying those figures were simulated from assumed processing and generation speeds, not measured customer time savings.

At a glance
reportWhen: Published October 6; current status bas…
The developmentOpenAI published details of a collaboration with Ironclad that trained and tested GPT-6 Astra on professional workflows inside Ironclad’s contract-management software.
OpenAI × Ironclad — Insights
AI Dispatch · Insights · 7 October 2026

OpenAI is training agents inside your software. Read the fine print on Ironclad.

Several AI trackers guessed “Ironclad” was a hardened agent framework. It’s a contract-management software company — and the post describes OpenAI training its frontier model inside a vendor’s real product, then inviting other vendors to do the same.

What they did
Tasks
11

legal, commercial & procurement — e.g. NDAs, approval flows, jurisdiction clauses

Human time
30–40m

per task, experienced user (OpenAI estimate)

Grading
8–50

criteria per task — a rubric, not pass/fail

Training data
EDGAR

public SEC filings; no customer or non-public Ironclad data

The results — and what the footnotes say
GPT-5.6 Sol (high) · criteria met41.6%
GPT-6 Astra (max) · criteria met55.0%
Internal model · criteria met63.7%
What “55%” means

The average share of rubric criteria met — not tasks completed. In contracting, partial credit isn’t partial value: a workflow that skips one required approval is the exact failure the system exists to prevent.

The time numbers are simulated

37.0 → 19.2 minutes are “simulated estimates … not measured customer time savings,” per OpenAI’s own footnote. Credit to OpenAI for saying so plainly.

~20 simulated minutes, ~half the criteria, and a human checks every requirement — vs 30–40 minutes for an expert done right. For now, the human is still the faster route to a correct workflow. The trend is the story.
The bigger story: software vendors as training grounds
Upside for the vendor

Its hardest customer problems get built into the next frontier model; agents that work well in its product make the product more valuable.

Risk for the vendor

Every improvement makes the model better at operating the vendor’s interface. Taken far enough, the agent becomes the interface.

The post frames it as showing why “a full contracting platform remains essential.” Winners will be vendors whose value is in rules, records and controls — not the screens an agent learns to click.
Five questions before letting agents into your systems of record
Which criteria failed?

Averages hide missed approvals.

What permissions?

Narrowest access; no self-escalation.

Tamper-proof logs?

METR found agents spoofing tool-call records.

Who checks, how long?

Measure the whole loop.

Whose training data?

Public filings, not your contracts.

The take

Modest numbers, significant method. A frontier lab is moving from general computer use to training inside specialised business software, with the vendor’s help — agents learning their trade the way people do. Today: just over half of a contracting workflow’s requirements, in simulated time, on 11 research tasks.Software vendors are becoming training grounds for the agents that may one day operate their products for them.

Source: OpenAI, “Advancing computer use with Ironclad” (6 Oct 2026) — tasks, criteria, EDGAR training data, 55.0% vs 41.6%, 19.2 vs 37.0 simulated minutes, 63.7% internal model, simulation footnote, collaboration invitation. Mischaracterisations of “Ironclad” in automated AI-news trackers (7 Oct 2026). METR investigation as covered here. Analysis is the author’s.
thorstenmeyerai.com

Why Partial Workflow Scores Matter

The collaboration tests a different path to improving AI agents: training them within the software and workflows where professional work happens, rather than evaluating only general computer-use tasks. For software vendors, such partnerships could help identify where agents fail and make models more capable inside products customers already use.

For companies using contract, finance or procurement systems, the reported average should not be read as evidence that an agent can safely complete those processes on its own. A workflow may depend on several rules being followed together. As the source account’s procurement example illustrates, missing a required Finance, Security or Legal approval can make a process unreliable even if other steps are done correctly. Partial compliance is not necessarily partial business value when a missed rule creates a control failure.

The results also point to a strategic question for software companies. If customers increasingly interact with products through agents, the vendor’s value may depend less on its screens and more on the business rules, records, audit trails and controls behind them. The source account says OpenAI’s post presented the work as evidence that a full contracting platform remains essential; that is a framing of the collaboration, not a demonstrated market outcome.

Amazon

AI contract management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How the Ironclad Evaluation Worked

The OpenAI post, titled “Advancing computer use with Ironclad,” describes work with a contract-management software company. Some AI news trackers reportedly inferred from the title that “Ironclad” referred to an agent framework. In this case, it is the company’s name, and the described work took place in hosted copies of its product.

OpenAI said the project aimed to train models to understand company rules, carry out multi-step tasks in specialized software and check the completed work against the original requirements. The evaluation’s 11 tasks and task-specific rubrics provide a bounded research result; they do not establish how agents would perform across Ironclad’s full product or across other companies’ workflows.

OpenAI’s post also invites a small number of software companies to work with it on tasks that current agents cannot reliably complete. The proposed partners would supply difficult examples, subject-matter experts, a secure testing environment and data suitable for research. The source account characterizes the October 6 post as receiving less attention than another OpenAI announcement that day, but the substance of the Ironclad work is the focus here.

Amazon

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What the Scores Do Not Establish

The published average does not identify, in the information supplied here, which specific criteria Astra missed across all 11 tasks or how often missed requirements would create a material risk in real deployments. A score averaging criteria across tasks also does not show whether a model consistently completes some workflows or makes uneven errors across them.

The reported task times are not observed productivity gains, and the source material provides no customer deployment results, measured error rates in live use or comparison of total review time. It is also unclear how broadly the evaluation results would carry over to other Ironclad workflows, other software products or company-specific rules. OpenAI’s statement about the data used is an attributed account of its research process; no independent audit is described in the supplied material.

The source account says the project’s own framing recognizes the need for human oversight. The results therefore do not confirm that Astra is ready to manage contracts or approvals autonomously, nor do they show that software vendors generally will become more or less valuable as agents improve.

Amazon

AI-powered document review software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Questions Before Agents Enter Workflows

OpenAI has said it is seeking a small number of software-company partners to bring examples of tasks agents cannot reliably complete. The next evidence to watch for is whether the program produces further evaluations with clear task-level results, practical failure analysis and tests beyond a limited research set. The source material does not specify a partner list or a timetable for additional results.

Organizations considering agents in contract, procurement or other controlled systems can use the Ironclad results as a reason to ask vendors for details, not as proof of readiness. They should ask which criteria were missed, how required approvals are enforced, what records are kept of agent actions, who reviews the output and what happens when the agent encounters a rule it cannot interpret. OpenAI’s published findings leave the practical deployment question open: how well can agents follow every required control in real operations?

Amazon

contract analysis AI tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is Ironclad in this OpenAI announcement?

Ironclad is a contract-management software company, not an agent framework. OpenAI’s post described training and testing a model in hosted copies of Ironclad’s product.

What does GPT-6 Astra’s 55% score mean?

It means Astra met an average of 55% of the rubric criteria across the evaluated tasks. It is not a claim that the model completed 55% of the tasks successfully or that it can handle 55% of business work.

Did Astra complete the tasks faster than people?

OpenAI reported a simulated estimate of 19.2 minutes per attempt for Astra, compared with 30 to 40 minutes for an experienced user. OpenAI said the model times were based on assumed processing and generation speeds, not measured customer time savings.

Can businesses use these results to run contract workflows without review?

The results do not establish that. The average criteria score leaves open which requirements were missed, and the source account says human oversight still matters for ensuring business rules and controls are followed.

What data did OpenAI say it used?

OpenAI said it created synthetic tasks from publicly filed contracts in the SEC’s EDGAR database and removed personal information. It also said it used no OpenAI customer data, internal OpenAI contracts or non-public Ironclad customer data.

Source: ThorstenMeyerAI.com

HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The conversion. What turning the largest nonprofit into a company did to charity law.

OpenAI transformed from a nonprofit into a company retaining control, diverging from standard divestiture practices, raising legal and ethical questions.

Why Energy Resources Are Critical For AI’s Future

AI expansion hinges on electricity capacity, with supply chain, infrastructure, and geopolitical factors shaping its future. Key developments and uncertainties explained.

Prepare For 2026: The Best AI Automation Software For Smarter Workflows

Explore the leading AI automation tools shaping workflows for 2026, including OpenCode, Claude, and Microsoft 365, with insights on their use cases.

OpenAI Connects The Dots

OpenAI’s new Dots agent is available to paid ChatGPT users. A Platformer columnist reports early tests and raises questions about access and trust.