🔍 Read the full analysis: Understanding Ironclad’s Fine Print On OpenAI’s Agents And Your Software on ThorstenMeyerAI.com
Get business pricing on monitors, keyboards and dev gear
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
TL;DR
OpenAI described training and evaluating a frontier model on 11 legal, commercial and procurement tasks inside contract-management software company Ironclad’s product. GPT-6 Astra met an average 55% of task criteria, while its estimated completion time was simulated rather than measured in customer use. OpenAI is inviting other software companies to participate, but the results do not establish that agents are ready to run business workflows without human review.
OpenAI said on October 6 that it trained and evaluated its frontier model GPT-6 Astra on professional workflows inside Ironclad’s contract-management software, reporting an average of 55% of rubric criteria met across 11 tasks. The results offer a view of how AI agents may learn to use specialized business software, but OpenAI’s time estimates were simulated and the company says human oversight remains necessary.
The tasks were selected by Ironclad staff and OpenAI employees who use the product. They covered legal, commercial and procurement work, including creating nondisclosure agreements, setting up procurement approval processes and adjusting a reusable contract clause to match a requester’s jurisdiction. OpenAI estimated that an experienced user would take 30 to 40 minutes to complete each task.
Tasks were graded against 8 to 50 criteria, depending on complexity. OpenAI reported that GPT-6 Astra met an average 55% of those criteria, compared with 41.6% for GPT-5.6 Sol at its high setting. An internal OpenAI model used during Astra’s development reached 63.7%. On one showcased task, Astra met about 94% of the criteria. The average score is a share of requirements met, not the percentage of tasks completed successfully.
Ironclad provided hosted copies of its product for the models to use. OpenAI said it generated synthetic training tasks from contracts filed publicly in the U.S. Securities and Exchange Commission’s EDGAR database, with personal information removed. The company said it used no OpenAI customer data, OpenAI internal contracts or non-public Ironclad customer data. OpenAI also reported estimated times of 19.2 minutes for Astra and 37 minutes for GPT-5.6 Sol, while saying those figures were simulated from assumed processing and generation speeds, not measured customer time savings.
OpenAI is training agents inside your software. Read the fine print on Ironclad.
Several AI trackers guessed “Ironclad” was a hardened agent framework. It’s a contract-management software company — and the post describes OpenAI training its frontier model inside a vendor’s real product, then inviting other vendors to do the same.
legal, commercial & procurement — e.g. NDAs, approval flows, jurisdiction clauses
per task, experienced user (OpenAI estimate)
criteria per task — a rubric, not pass/fail
public SEC filings; no customer or non-public Ironclad data
The average share of rubric criteria met — not tasks completed. In contracting, partial credit isn’t partial value: a workflow that skips one required approval is the exact failure the system exists to prevent.
37.0 → 19.2 minutes are “simulated estimates … not measured customer time savings,” per OpenAI’s own footnote. Credit to OpenAI for saying so plainly.
Its hardest customer problems get built into the next frontier model; agents that work well in its product make the product more valuable.
Every improvement makes the model better at operating the vendor’s interface. Taken far enough, the agent becomes the interface.
Averages hide missed approvals.
Narrowest access; no self-escalation.
METR found agents spoofing tool-call records.
Measure the whole loop.
Public filings, not your contracts.
Modest numbers, significant method. A frontier lab is moving from general computer use to training inside specialised business software, with the vendor’s help — agents learning their trade the way people do. Today: just over half of a contracting workflow’s requirements, in simulated time, on 11 research tasks.Software vendors are becoming training grounds for the agents that may one day operate their products for them.
Why Partial Workflow Scores Matter
The collaboration tests a different path to improving AI agents: training them within the software and workflows where professional work happens, rather than evaluating only general computer-use tasks. For software vendors, such partnerships could help identify where agents fail and make models more capable inside products customers already use.
For companies using contract, finance or procurement systems, the reported average should not be read as evidence that an agent can safely complete those processes on its own. A workflow may depend on several rules being followed together. As the source account’s procurement example illustrates, missing a required Finance, Security or Legal approval can make a process unreliable even if other steps are done correctly. Partial compliance is not necessarily partial business value when a missed rule creates a control failure.
The results also point to a strategic question for software companies. If customers increasingly interact with products through agents, the vendor’s value may depend less on its screens and more on the business rules, records, audit trails and controls behind them. The source account says OpenAI’s post presented the work as evidence that a full contracting platform remains essential; that is a framing of the collaboration, not a demonstrated market outcome.
As an affiliate, we earn on qualifying purchases.
How the Ironclad Evaluation Worked
The OpenAI post, titled “Advancing computer use with Ironclad,” describes work with a contract-management software company. Some AI news trackers reportedly inferred from the title that “Ironclad” referred to an agent framework. In this case, it is the company’s name, and the described work took place in hosted copies of its product.
OpenAI said the project aimed to train models to understand company rules, carry out multi-step tasks in specialized software and check the completed work against the original requirements. The evaluation’s 11 tasks and task-specific rubrics provide a bounded research result; they do not establish how agents would perform across Ironclad’s full product or across other companies’ workflows.
OpenAI’s post also invites a small number of software companies to work with it on tasks that current agents cannot reliably complete. The proposed partners would supply difficult examples, subject-matter experts, a secure testing environment and data suitable for research. The source account characterizes the October 6 post as receiving less attention than another OpenAI announcement that day, but the substance of the Ironclad work is the focus here.
As an affiliate, we earn on qualifying purchases.
What the Scores Do Not Establish
The published average does not identify, in the information supplied here, which specific criteria Astra missed across all 11 tasks or how often missed requirements would create a material risk in real deployments. A score averaging criteria across tasks also does not show whether a model consistently completes some workflows or makes uneven errors across them.
The reported task times are not observed productivity gains, and the source material provides no customer deployment results, measured error rates in live use or comparison of total review time. It is also unclear how broadly the evaluation results would carry over to other Ironclad workflows, other software products or company-specific rules. OpenAI’s statement about the data used is an attributed account of its research process; no independent audit is described in the supplied material.
The source account says the project’s own framing recognizes the need for human oversight. The results therefore do not confirm that Astra is ready to manage contracts or approvals autonomously, nor do they show that software vendors generally will become more or less valuable as agents improve.
AI-powered document review software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Questions Before Agents Enter Workflows
OpenAI has said it is seeking a small number of software-company partners to bring examples of tasks agents cannot reliably complete. The next evidence to watch for is whether the program produces further evaluations with clear task-level results, practical failure analysis and tests beyond a limited research set. The source material does not specify a partner list or a timetable for additional results.
Organizations considering agents in contract, procurement or other controlled systems can use the Ironclad results as a reason to ask vendors for details, not as proof of readiness. They should ask which criteria were missed, how required approvals are enforced, what records are kept of agent actions, who reviews the output and what happens when the agent encounters a rule it cannot interpret. OpenAI’s published findings leave the practical deployment question open: how well can agents follow every required control in real operations?
As an affiliate, we earn on qualifying purchases.
Key Questions
What is Ironclad in this OpenAI announcement?
Ironclad is a contract-management software company, not an agent framework. OpenAI’s post described training and testing a model in hosted copies of Ironclad’s product.
What does GPT-6 Astra’s 55% score mean?
It means Astra met an average of 55% of the rubric criteria across the evaluated tasks. It is not a claim that the model completed 55% of the tasks successfully or that it can handle 55% of business work.
Did Astra complete the tasks faster than people?
OpenAI reported a simulated estimate of 19.2 minutes per attempt for Astra, compared with 30 to 40 minutes for an experienced user. OpenAI said the model times were based on assumed processing and generation speeds, not measured customer time savings.
Can businesses use these results to run contract workflows without review?
The results do not establish that. The average criteria score leaves open which requirements were missed, and the source account says human oversight still matters for ensuring business rules and controls are followed.
What data did OpenAI say it used?
OpenAI said it created synthetic tasks from publicly filed contracts in the SEC’s EDGAR database and removed personal information. It also said it used no OpenAI customer data, internal OpenAI contracts or non-public Ironclad customer data.
Source: ThorstenMeyerAI.com
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
