AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Is AI Making Quality Control The Costliest Step? on ThorstenMeyerAI.com

Buying for a business?Offer from Amazon

Get business pricing on office and shipping supplies

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

TL;DR

AI systems are producing more mathematical manuscripts, code changes and contract work, but checking that output remains time-consuming. Industry reports and a peer-reviewed study point to pressure on software review, though the figures have limits and do not prove that quality control is now the costliest step across all industries. The scale of the problem, and who will be accountable for errors, remain unsettled.

AI is increasing the volume of work produced in mathematics, software and contract workflows, while the people and processes needed to check it remain limited, according to figures and examples compiled by ThorstenMeyerAI.com. The material suggests verification can constrain how much AI output organisations safely use, but it does not establish that quality control is universally the costliest step.

The source says OpenAI published 722 mathematical manuscripts this week, selected from work on roughly 4,000 problems, with an average result taking about three hours of compute. The manuscripts were grouped into 372 families. Some results were formally checked using Lean, while OpenAI cautioned that unformalized work could contain issues. The source contrasts that output with the careful checking of an earlier counterexample to an Erdős conjecture by five leading mathematicians; it does not provide the verification time or a like-for-like cost comparison.

Software metrics cited in the analysis point to a similar tension, though they come from different datasets and providers. Faros AI reported that teams merged 98% more pull requests during high-AI-adoption periods, while review time rose 91%. LinearB, analysing 8.1 million pull requests across 4,800 organisations, said AI-generated changes took 4.6 times longer to reach the start of review and were accepted 32.7% of the time, compared with 84.4% for human-written changes. A peer-reviewed 2026 study reportedly found that 61% of AI-agent pull requests received no human review before being merged or closed.

For contract work, the source says OpenAI partnered with contract-software company Ironclad to train GPT-6 Astra on real contracting workflows. On 11 tasks, the model met an average of 55% of evaluation criteria. That result is presented as an improvement over a previous model, but the source gives no prior score or detailed account of the evaluation. It indicates progress, not that the model can replace professional review.

At a glance
analysisWhen: Source material describes developments…
The developmentA source analysis draws on recent AI-generated mathematics, software-review metrics and a legal-workflow model evaluation to argue that verification is becoming a bottleneck as production speeds up.
The Referee Shortage — Post-Labor
AI Dispatch · Post-Labor · 7 October 2026

The referee shortage: AI made doing cheap and checking expensive

OpenAI’s model produced a maths result in about three hours of compute. Verifying one earlier result took five of the world’s leading mathematicians. That ratio is the next decade of work: producing is cheap and abundant; trusting is slow, human and scarce.

One pattern, three fields
Mathematics
722
manuscripts, ~3h compute each

Some Lean-checked; OpenAI warns unformalized ones “could have issues.” Verification abundance, adjudication scarcity.

Software
+98% / +91%
more PRs merged / longer review

Faros AI. LinearB (8.1M PRs): AI changes wait 4.6× longer, accepted 32.7% vs 84.4%.

Professional workflows
55%
of criteria met — Astra on Ironclad

Real progress. Someone still has to find the other 45% before the work can be used.

Generation collapsed. Verification didn’t. (conceptual, not to scale)
Cost to produce a resultdown
Cost to check a resultnot down
No author intent

Machine output arrives without reasoning a reviewer can interrogate. It looks locally clean and gives no clue where it’s wrong.

Checks the answer, not the question

A prover confirms the proof proves its statement; tests confirm what tests check. Neither confirms it’s what was needed.

Someone must be accountable

Contracts are signed, designs stamped, papers defended. Responsibility is institutional — you can’t hold a model to it.

Illustrative: $1 of model time + 4 minutes of review at $45/hour. Halving the model price saves 12.5%; one extra review minute erases it. In that example, review is three-quarters of the bill.
What happens when referees run out — already visible
Rubber-stamping
61%

of AI-agent pull requests got no human review at all (EASE 2026). Zero-review merges up 31.3% (Faros).

Triage by suspicion
38%

of reviewers deliberately deprioritise AI changes (LinearB). Good machine work waits behind bad.

Producer as filter
~4,000 → 372

OpenAI chose which maths families were significant. When referees can’t keep up, the producer’s filter becomes the review.

The apprenticeship paradox: reviewers are made by doing the work. The work AI absorbs — writing code, drafting contracts, proving lemmas — is exactly what trained the reviewers. Demand for judgement rises as its supply line shrinks.
What to do
Price verification

Budget review hours next to model spend.

Formalise checks

Provers, types, tests, policy engines.

Tier the review

Experts only where consequences are high.

Fund the referees

Who profits from generation pays for checking.

Protect apprenticeship

Keep some production human for learners.

The take

The first automation question was which jobs AI would do. The better one is which jobs AI makes more necessary: the ones that check, adjudicate and take responsibility. Expect a referee premium — senior engineers, auditors, specialist lawyers, reviewing scientists become the binding constraint on how much AI output anyone can use.Accountability — standing behind a result — may be the most durable form of human work there is.

Sources: OpenAI maths release & Erdős verification as covered here; arXiv:2608.28997; OpenAI × Ironclad (6 Oct 2026); Faros AI; LinearB 2026 (8.1M PRs); Duma et al., EASE 2026 — via secondary reporting. Several code-review sources sell review tools. Review-cost example illustrative. Analysis is the author’s.
thorstenmeyerai.com

Review Capacity May Limit AI Use

If output grows faster than review capacity, organisations face a practical constraint: they may be unable to use all the work AI produces without accepting greater risk or adding reviewers. In software, delays can slow releases; skipped checks can allow defects through. In contracts, a missed approval requirement or incorrect clause could have legal or commercial consequences. These are potential impacts, not outcomes quantified by the figures cited here.

The analysis also points to a workforce issue. Experienced reviewers usually develop judgment through years of producing and examining work. If AI takes over tasks that once trained junior staff, the pool of people qualified to evaluate high-stakes output could shrink even as demand for their approval increases. The source describes this as an apprenticeship risk; it does not provide workforce data establishing that the decline is already happening.

That makes the question broader than whether a model can generate a plausible answer. Institutions still need people who can decide whether the task was framed correctly, whether the answer fits its intended use and who is accountable if it fails. The people with that expertise may become a limiting resource, although the material does not quantify any resulting wage or employment premium.

Amazon

AI quality control software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Three Fields, One Verification Gap

The argument links examples from fields with different standards of review. In mathematics, formal proof systems such as Lean can check whether a proof follows from stated assumptions. They cannot by themselves determine whether the theorem addresses the right question or whether it is important. The source notes that some of OpenAI’s manuscripts were formalized and some were not; it does not specify how many fell into each group.

In software, automated tests and code-review processes can catch certain errors, but their usefulness depends on what they cover and how carefully people examine the changes. The figures cited come from Faros AI, LinearB and a peer-reviewed study, with distinct methods and measures. The source also warns that Faros and LinearB sell code-review tools, a commercial interest readers should keep in mind when interpreting their reports.

Contract workflows add a different kind of judgment: checking obligations, approvals and jurisdiction-specific language. The model evaluation described in the source covered 11 tasks and reported average criteria met, not a broad measure of legal accuracy or performance across all contract types. These examples support a question about review capacity, but they are not a single controlled comparison of the cost of producing versus checking work.

Amazon

software review automation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How Much Review Is Enough?

The evidence does not show that quality control is the costliest step across the economy. The examples use different kinds of work, measures and time periods, and the source supplies no comparable figures for the monetary cost of generation and review. The OpenAI manuscript count is an output figure, not a measure of how many results are correct or useful.

The software statistics also need careful interpretation. The source does not provide study methods, uncertainty ranges or enough detail to reconcile the different acceptance and review figures. Faros AI and LinearB sell review-related products, while the separate 2026 study is described as peer-reviewed but not identified by title or venue. The cited measures do not establish that AI authorship caused longer review times or lower acceptance rates.

It is also unclear how much checking can be automated, how many errors escape existing review, or whether organisations are already changing staffing and training in response. The source offers an argument about a possible reviewer shortage, not direct evidence that senior expertise or pay has risen.

Amazon

contract review AI tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Track Review Quality and Training

The immediate test for the claim will be whether future, methodologically transparent studies find the same pattern across organisations: more AI-assisted output, longer or less complete review, and measurable consequences for defects, rework or risk. Comparisons will need to account for differences in task difficulty, team size, adoption levels and review rules.

For the OpenAI mathematics work, useful follow-up information would include how many manuscripts receive formal verification, how results fare under independent review, and how the project handles corrections. For software and contracting, organisations would need to report not only review speed but also missed defects, downstream fixes and accountability arrangements.

The workforce question will take longer to assess. Employers and professional bodies can monitor whether junior staff still get opportunities to draft, test and revise work themselves, rather than only reviewing machine-generated drafts. The source identifies that training pathway as a concern, but whether AI adoption is weakening it remains to be established.

Amazon

mathematical verification software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

Does the evidence prove quality control is now the costliest step?

No. The cited examples suggest that review can be a bottleneck, but they do not compare generation and verification costs on a common basis or establish a universal ranking across industries.

Were all 722 OpenAI mathematics manuscripts verified?

The source says some results were formally checked in Lean and quotes OpenAI warning that unformalized results could have issues. It does not state how many manuscripts were formally verified or independently reviewed.

What do the software review figures show?

Faros AI reported more merged pull requests alongside longer review time in high-adoption periods. LinearB reported longer waits before review began and lower acceptance rates for AI-generated changes. These are separate findings, and the source does not establish that AI alone caused the differences.

Can AI check AI-generated work?

Automated tools can check specific properties, such as whether a proof follows formal rules or whether code passes tests. Those checks cannot automatically establish that the original task, assumptions or tests were appropriate; human judgment may still be needed.

Could AI reduce opportunities for junior workers to gain expertise?

The source raises this as a risk because experienced reviewers often develop judgment by producing and reviewing work. It does not provide data showing that junior training opportunities have already declined across these fields.

Source: ThorstenMeyerAI.com

This content is for general information only and is not financial, tax or legal advice. Consult a qualified professional for decisions about your money.
HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

How AI Helped Kimi K3 Surpass Expectations And End Price Competition

Kimi K3, a Chinese AI with 2.8 trillion parameters, debuts at a price matching Western models, signaling a shift in Chinese AI competitiveness.

Eucalyptus Resources Surges In Global Coverage

Eucalyptus Resources experiences a significant surge in international coverage, with 62 mentions in recent media monitoring reports, highlighting growing global interest.

FCC says it will move toward 2027 auction of mid-band wireless spectrum

FCC announces plans to move toward a 2027 auction of mid-band wireless spectrum, impacting 5G deployment and telecom industry strategies.

How Global Media Is Covering Aretha Franklin’s Music And Live Performances

An analysis of how international media is covering Aretha Franklin’s music and upcoming live shows, highlighting recent surges in coverage and implications for promoters.