AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: OpenAI’s Agent Training Inside Software Raises A Fine-Print Question For Ironclad on ThorstenMeyerAI.com

Buying for a business?Offer from Amazon

Get business pricing on office and shipping supplies

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

TL;DR

OpenAI described training and evaluating AI models on legal, procurement and commercial tasks inside Ironclad’s contract-management software. GPT-6 Astra met an average 55% of task criteria, while estimated completion times were simulated, not measured customer savings. OpenAI says the work used public contract filings and did not use non-public Ironclad customer data.

OpenAI published results on October 6 from training AI models on 11 legal, commercial and procurement tasks inside Ironclad’s contract-management software, reporting that GPT-6 Astra met an average 55% of evaluation criteria. The work is a research effort rather than evidence of production-ready automation: OpenAI’s time estimates were simulated, and its results leave substantial room for human review of consequential workflows.

Ironclad is a contract-management software company, not the name of a new AI agent framework. OpenAI’s post describes a collaboration in which Ironclad staff and OpenAI employees who use the product selected tasks such as creating nondisclosure agreements, setting up procurement approval processes and adjusting reusable contract clauses to reflect a requester’s jurisdiction. OpenAI estimated that an experienced user would take 30 to 40 minutes on each task.

Ironclad provided hosted copies of its product for model practice. OpenAI says it generated synthetic training tasks from publicly filed contracts in the US Securities and Exchange Commission’s EDGAR database, with personal information filtered out. The company said it did not use OpenAI customer data, OpenAI internal contracts or non-public Ironclad customer data.

OpenAI reported that GPT-5.6 Sol, evaluated at its high setting, met an average 41.6% of rubric criteria, while GPT-6 Astra, at its maximum setting, met 55%. An internal model used in Astra’s development reached 63.7%. The tasks were graded against between eight and 50 criteria, depending on complexity. OpenAI also highlighted one task on which Astra met about 94% of the criteria; that example is not the overall average.

At a glance
reportWhen: Published October 6; further partner wo…
The developmentOpenAI published results from training frontier models on workflows inside Ironclad’s contract-management product and invited other software companies to partner on similar research.
OpenAI × Ironclad — Insights
AI Dispatch · Insights · 7 October 2026

OpenAI is training agents inside your software. Read the fine print on Ironclad.

Several AI trackers guessed “Ironclad” was a hardened agent framework. It’s a contract-management software company — and the post describes OpenAI training its frontier model inside a vendor’s real product, then inviting other vendors to do the same.

What they did
Tasks
11

legal, commercial & procurement — e.g. NDAs, approval flows, jurisdiction clauses

Human time
30–40m

per task, experienced user (OpenAI estimate)

Grading
8–50

criteria per task — a rubric, not pass/fail

Training data
EDGAR

public SEC filings; no customer or non-public Ironclad data

The results — and what the footnotes say
GPT-5.6 Sol (high) · criteria met41.6%
GPT-6 Astra (max) · criteria met55.0%
Internal model · criteria met63.7%
What “55%” means

The average share of rubric criteria met — not tasks completed. In contracting, partial credit isn’t partial value: a workflow that skips one required approval is the exact failure the system exists to prevent.

The time numbers are simulated

37.0 → 19.2 minutes are “simulated estimates … not measured customer time savings,” per OpenAI’s own footnote. Credit to OpenAI for saying so plainly.

~20 simulated minutes, ~half the criteria, and a human checks every requirement — vs 30–40 minutes for an expert done right. For now, the human is still the faster route to a correct workflow. The trend is the story.
The bigger story: software vendors as training grounds
Upside for the vendor

Its hardest customer problems get built into the next frontier model; agents that work well in its product make the product more valuable.

Risk for the vendor

Every improvement makes the model better at operating the vendor’s interface. Taken far enough, the agent becomes the interface.

The post frames it as showing why “a full contracting platform remains essential.” Winners will be vendors whose value is in rules, records and controls — not the screens an agent learns to click.
Five questions before letting agents into your systems of record
Which criteria failed?

Averages hide missed approvals.

What permissions?

Narrowest access; no self-escalation.

Tamper-proof logs?

METR found agents spoofing tool-call records.

Who checks, how long?

Measure the whole loop.

Whose training data?

Public filings, not your contracts.

The take

Modest numbers, significant method. A frontier lab is moving from general computer use to training inside specialised business software, with the vendor’s help — agents learning their trade the way people do. Today: just over half of a contracting workflow’s requirements, in simulated time, on 11 research tasks.Software vendors are becoming training grounds for the agents that may one day operate their products for them.

Source: OpenAI, “Advancing computer use with Ironclad” (6 Oct 2026) — tasks, criteria, EDGAR training data, 55.0% vs 41.6%, 19.2 vs 37.0 simulated minutes, 63.7% internal model, simulation footnote, collaboration invitation. Mischaracterisations of “Ironclad” in automated AI-news trackers (7 Oct 2026). METR investigation as covered here. Analysis is the author’s.
thorstenmeyerai.com

Why Partial Workflow Accuracy Matters

The results offer an early view of how AI developers may train agents to work inside specialist business software, rather than relying only on general-purpose examples. If models become more reliable at carrying out multi-step work in products used for contracts, procurement and other business operations, software vendors could offer new forms of assistance and automation. OpenAI is also inviting a small number of software companies to bring difficult tasks, domain experts, secure test environments and research-appropriate data.

But a rubric score is not the same as a completed, dependable business process. OpenAI’s 55% figure is the average share of criteria met, not the share of tasks completed successfully. In a procurement workflow, for example, missing a required Finance, Security or Legal approval can undermine the process even if other steps are done correctly. The results do not establish that Astra can be trusted to run such work without checks.

The reported times also do not establish a productivity gain. OpenAI estimated 19.2 minutes per Astra attempt, compared with 37 minutes for GPT-5.6 Sol, but said those figures were simulated using assumed processing and generation speeds. They are not observed customer times. With the agent meeting about half the criteria on average, the figures cannot show that customers would finish work faster while maintaining accuracy.

For software vendors, the arrangement presents both a potential opportunity and a strategic question. Models that work well inside a product may make it more useful, while also shifting how customers interact with it. As agents take on more tasks, a vendor’s lasting value may depend less on the screens an agent operates and more on the business rules, data structures, audit records and controls behind them. That is an implication of the model described by the collaboration, not a proven outcome.

Amazon

contract management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How the Ironclad Tests Were Set Up

OpenAI’s post framed the project around teaching models to understand an organization’s rules, perform multi-step workflows in specialized software and check whether the result meets the initial requirements. The 11 tasks were selected by Ironclad staff and OpenAI employees who use the product, and evaluated against task-specific rubrics. OpenAI said GPT-6 Astra was the first frontier model trained using this approach.

The distinction between research testing and customer deployment is central to interpreting the announcement. The models practised in hosted copies of Ironclad’s product, and the training tasks were based on public contract filings, according to OpenAI. The published account does not report a trial measuring outcomes in customers’ live systems. Nor does it establish that the benchmark covers the full range of work performed in Ironclad or other business software.

OpenAI presented the effort as a possible model for working with other software companies on tasks that current agents cannot reliably complete. The proposed partners would need to provide examples of failures, people with detailed knowledge of the work, a secure environment for testing and data that can safely be used for research. The announcement therefore describes both a benchmark and a route for further collaboration; it does not announce a general release of an Ironclad-operating agent.

Amazon

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What the Benchmark Cannot Yet Show

The published results do not show how the model would perform across a broader set of Ironclad workflows or in customers’ live environments. OpenAI’s average criteria scores also do not identify, in the source material, which specific requirements models missed across all 11 tasks. That matters because a low score on a nonessential detail and a missed approval control can have very different consequences.

It is also unclear whether the reported performance would hold with different contract documents, company-specific rules or software configurations. The simulated times are not evidence of actual customer savings, and the results do not establish how much human checking would be required in routine use. OpenAI says human oversight remains relevant; no independent evaluation or deployment outcome is supplied in the source material.

The source describes GPT-6 Astra as the first frontier model trained this way, but does not provide a release date, customer availability or a deployment plan for an Ironclad-integrated agent. OpenAI’s invitation to other software companies indicates potential follow-on research, not that partnerships or product launches have already been completed.

Amazon

procurement workflow automation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Further Software Partnerships

OpenAI says it is seeking a small number of software-company partners to work on difficult tasks that current agents cannot reliably complete. Those companies are expected to bring concrete examples of failure, knowledgeable staff, a secure test environment and data that can be used safely for research. The post does not name additional partners or give a schedule for further results.

The next useful evidence would include task-by-task reporting that identifies which criteria were missed, tests across a wider range of workflows and measured performance in settings that reflect real use. For companies considering agents in contract or procurement systems, the current results support asking what is being tested, how approval rules are checked and what human review remains necessary. OpenAI has not provided customer productivity measurements or a timetable for production use in the material described here.

Amazon

AI-powered NDA creation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is Ironclad in this announcement?

Ironclad is a contract-management software company. OpenAI’s post describes training and testing models on workflows inside Ironclad’s product; Ironclad is not the name of a new AI agent framework.

What does GPT-6 Astra’s 55% score mean?

It is the average share of rubric criteria met across the research tasks, not the percentage of tasks completed successfully. The criteria included requirements tailored to each task, and the published average does not show which specific requirements were missed in every case.

Did OpenAI measure customer time savings?

No. OpenAI said the reported task times were simulated estimates based on assumed processing and generation speeds. They were not measured time savings for customers.

What data did OpenAI say it used?

OpenAI said it created synthetic training tasks from publicly filed contracts in the SEC’s EDGAR database and filtered out personal information. It said it did not use OpenAI customer data, internal OpenAI contracts or non-public Ironclad customer data.

Is an Ironclad AI agent ready for customer use?

The source material does not announce a customer release or production deployment. The reported results are from research tasks, and they do not establish that the system can complete consequential contract or procurement workflows without human review.

Source: ThorstenMeyerAI.com

This content is for general information only and is not financial, tax or legal advice. Consult a qualified professional for decisions about your money.
FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Transform Your Aerial Photography With AI Camera Drones In 2026

Discover how AI-powered camera drones are revolutionizing aerial photography in 2026, offering enhanced stability, intelligent features, and professional-quality footage.

OpenAI’s Cost-Cutting For GPT‑6 Sol And Luna Leaves Benchmark Scores Unchanged

OpenAI reduces GPT-6 Sol and Luna prices by 50%, with benchmark scores remaining stable, impacting AI deployment costs and capabilities.

A Conversation On AI And Faith Surprises Scholars At Anthropic

A New York Times headline reports a meeting between religious scholars and Anthropic, but the participants, discussion and outcome are not available.

Making AI Smarter: Hardware Designed Before The Algorithms

New hardware designs prioritize inference workloads, focusing on thermal efficiency, memory interconnects, and specialization, marking a shift from traditional AI chips.