Can AI’s Review Process Scale Alongside Its Output?
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Can AI’s Review Process Scale Alongside Its Output? on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get monitors, keyboards and dev gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

OpenAI said its model produced 722 mathematics manuscripts this week, while software-industry data cited in a recent analysis suggests AI-assisted code is also increasing review workloads. The figures point to a growing mismatch between the low cost of producing AI output and the human time needed to verify, judge and take responsibility for it.

OpenAI said this week its model produced 722 mathematical manuscripts from a programme that posed about 4,000 problems, intensifying debate over whether human review can keep pace with AI output. The development matters beyond mathematics: software-industry data cited by ThorstenMeyerAI.com also points to rising review time and uneven checks as AI-generated changes become more common.

The manuscripts were grouped into 372 families, and the source estimates that the average result took about three hours of compute to produce. OpenAI has said some results have formal checks in Lean, a proof-assistant system, while warning that unformalized results could have issues. The source does not provide a complete independent audit of all 722 manuscripts.

The difference between producing work and establishing its value is central to the concern. A formal system can test whether a proof follows from its stated assumptions; experts still need to judge whether the assumptions and theorem address the intended problem, whether the result is significant, and whether it withstands scrutiny. The source contrasts the manuscript volume with the careful verification by five leading mathematicians of an earlier result from the same programme, described there as a counterexample to an old Erdős conjecture.

Software metrics cited in the source offer a second example, though their methods and periods are not fully detailed in the supplied material. Faros AI reported that teams merged 98% more pull requests across low- and high-AI-adoption periods while review time rose 91%. LinearB said its analysis of 8.1 million pull requests across 4,800 organisations found AI-generated changes took 4.6 times longer for review to begin and were accepted 32.7% of the time, compared with 84.4% for human-written changes. These are reported findings, not a controlled demonstration that AI use alone caused the differences.

At a glance
reportWhen: Reported this week; the underlying stud…
The developmentOpenAI’s latest mathematics output, alongside software-review data, has sharpened concern that human verification capacity is not keeping pace with AI-generated work.
The Referee Shortage — Post-Labor
AI Dispatch · Post-Labor · 7 October 2026

The referee shortage: AI made doing cheap and checking expensive

OpenAI’s model produced a maths result in about three hours of compute. Verifying one earlier result took five of the world’s leading mathematicians. That ratio is the next decade of work: producing is cheap and abundant; trusting is slow, human and scarce.

One pattern, three fields
Mathematics
722
manuscripts, ~3h compute each

Some Lean-checked; OpenAI warns unformalized ones “could have issues.” Verification abundance, adjudication scarcity.

Software
+98% / +91%
more PRs merged / longer review

Faros AI. LinearB (8.1M PRs): AI changes wait 4.6× longer, accepted 32.7% vs 84.4%.

Professional workflows
55%
of criteria met — Astra on Ironclad

Real progress. Someone still has to find the other 45% before the work can be used.

Generation collapsed. Verification didn’t. (conceptual, not to scale)
Cost to produce a resultdown
Cost to check a resultnot down
No author intent

Machine output arrives without reasoning a reviewer can interrogate. It looks locally clean and gives no clue where it’s wrong.

Checks the answer, not the question

A prover confirms the proof proves its statement; tests confirm what tests check. Neither confirms it’s what was needed.

Someone must be accountable

Contracts are signed, designs stamped, papers defended. Responsibility is institutional — you can’t hold a model to it.

Illustrative: $1 of model time + 4 minutes of review at $45/hour. Halving the model price saves 12.5%; one extra review minute erases it. In that example, review is three-quarters of the bill.
What happens when referees run out — already visible
Rubber-stamping
61%

of AI-agent pull requests got no human review at all (EASE 2026). Zero-review merges up 31.3% (Faros).

Triage by suspicion
38%

of reviewers deliberately deprioritise AI changes (LinearB). Good machine work waits behind bad.

Producer as filter
~4,000 → 372

OpenAI chose which maths families were significant. When referees can’t keep up, the producer’s filter becomes the review.

The apprenticeship paradox: reviewers are made by doing the work. The work AI absorbs — writing code, drafting contracts, proving lemmas — is exactly what trained the reviewers. Demand for judgement rises as its supply line shrinks.
What to do
Price verification

Budget review hours next to model spend.

Formalise checks

Provers, types, tests, policy engines.

Tier the review

Experts only where consequences are high.

Fund the referees

Who profits from generation pays for checking.

Protect apprenticeship

Keep some production human for learners.

The take

The first automation question was which jobs AI would do. The better one is which jobs AI makes more necessary: the ones that check, adjudicate and take responsibility. Expect a referee premium — senior engineers, auditors, specialist lawyers, reviewing scientists become the binding constraint on how much AI output anyone can use.Accountability — standing behind a result — may be the most durable form of human work there is.

Sources: OpenAI maths release & Erdős verification as covered here; arXiv:2608.28997; OpenAI × Ironclad (6 Oct 2026); Faros AI; LinearB 2026 (8.1M PRs); Duma et al., EASE 2026 — via secondary reporting. Several code-review sources sell review tools. Review-cost example illustrative. Analysis is the author’s.
thorstenmeyerai.com

Review Time Is Becoming a Constraint

If generation grows faster than review capacity, organisations face a practical limit on how much AI-produced work they can safely use. More drafts, proofs or code changes do not automatically create more usable work; each still needs checks suited to its risks, and some decisions require a person able to explain and stand behind them.

The cited software figures describe several possible consequences: longer queues, lower review rates and reviewers giving AI-generated submissions extra scrutiny. The source also reports that a 2026 peer-reviewed study found 61% of AI-agent pull requests received no human review before merging or closing. Without details about the study’s sample and definition of review, that figure should be read as a finding within its reported scope, not a universal rate.

There is also a workforce implication. The source argues that experienced reviewers usually gain judgment through years of doing the underlying work. If AI changes reduce opportunities for junior staff to write code, draft contracts or develop proofs themselves, organisations may weaken the route by which future experts acquire that judgment. That is a risk raised by the analysis, not an established outcome across industries.

Amazon

AI code review tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Three Sectors, Similar Bottlenecks

Mathematics makes the distinction between checking and judging especially visible. A proof assistant can check a formal proof against a specified statement. It cannot by itself establish that the statement is the right one to pursue or explain what the result means. OpenAI’s warning about unformalized manuscripts also means the full collection should not be treated as uniformly verified.

In software, automated tests and review tools can catch some problems, but their coverage depends on what was tested and what reviewers inspect. The source cites Faros AI, LinearB and a peer-reviewed 2026 study, while cautioning that several of the data providers sell code-review tools. Their commercial interests do not invalidate the measurements, but readers need study methods and independent replication to assess how broadly the findings apply.

The source also describes an OpenAI partnership with contract-software company Ironclad involving a model called GPT-6 Astra, evaluated on 11 contracting tasks. It reportedly met 55% of evaluation criteria on average, an improvement over a previous model. That result is a reported benchmark, not proof that the model can safely handle contracts without legal review; the supplied material gives no task-by-task results or full evaluation methodology.

Amazon

AI manuscript verification software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How Much Work Was Verified?

The supplied source does not establish how many of the 722 manuscripts were independently checked, how many were formally verified, or how the published selection was made from the roughly 4,000 problems. It also does not provide enough detail to compare the mathematics programme’s review effort directly with the software studies.

The software statistics come from sources with differing methods and potential commercial interests. The material does not give their complete definitions, sampling procedures or measurement windows, so the numbers should not be combined into a single industry-wide estimate. The Ironclad evaluation likewise lacks enough methodological detail here to show how its criteria map to real contract outcomes.

More broadly, it remains unclear whether new review tools and workflows can expand expert capacity fast enough, or whether organisations will respond by accepting more unchecked work, delaying it, or narrowing what they use. The source presents these as possible patterns and risks, not settled outcomes.

Amazon

mathematics proof assistant tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

More Audits and Workflow Evidence

The next useful evidence will be independent audits of AI-generated mathematics, including clear counts of formally checked work, corrections and failures. For software, comparable studies should report how review time, defect rates and review coverage change under clearly defined levels of AI adoption, with methods that readers can examine.

For contract workflows, the reported 11-task result needs fuller evaluation details and evidence about how outputs perform when reviewed in actual practice. Across fields, organisations will also need to track who is accountable for approving work and whether junior staff still gain the experience required to become reliable reviewers.

For now, the confirmed development is the scale of OpenAI’s reported mathematics output, alongside separate studies and industry measurements pointing to review pressures in software. Whether verification can scale through better tools, more training or different institutional rules remains an open question.

Amazon

software review automation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What did OpenAI report this week?

OpenAI’s model produced 722 mathematical manuscripts in 372 families after being posed about 4,000 problems, according to the source. The source estimates an average of about three hours of compute per result.

Were all 722 manuscripts verified?

No. The source says some results were formally checked in Lean and quotes OpenAI warning that unformalized results could have issues. It does not state how many manuscripts were independently verified.

What do the software figures show?

The cited figures report rising review time and longer waits before review for AI-generated code changes. They come from different analyses with distinct methods; the supplied material does not establish that AI use alone caused the reported patterns.

Why can’t AI simply review AI output?

Automated checks can test specific properties, such as whether code passes tests or a formal proof follows stated rules. They may not establish whether the requirements, tests or original question were correct, and people or institutions often remain responsible for approving consequential work.

What remains unknown?

The total share of AI output receiving meaningful human review, the reliability of the cited findings across organisations, and whether review tools can increase expert capacity are all unclear from the available material.

Source: ThorstenMeyerAI.com

HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

What GPT-6 Means For Everyday Intelligent UI

A headline links GPT-6 with an intelligent interface for everyone, but provides no evidence of an announcement, features, access or timing.

The Evolution Of AI: Microsoft’s Signal Peak 2026 And Anthropic’s Models At The Forefront

Microsoft prepares to launch Project Perception, an AI security platform routing models from Microsoft, OpenAI, and Anthropic, challenging Anthropic’s Mythos.

Memory Stopped Being a Commodity

Micron’s new long-term contracts signal a fundamental change in memory markets, with buyers pre-funding capacity and memory no longer treated as a commodity.

Could Mistral Large 4 Be Your Pick Beyond The US And China?

Mistral Large 4 scores 38.4 on Artificial Analysis’ index, leading the stated non-US, non-China field but trailing major US and Chinese models.