AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Cheap AI Output, Costly Human Judgment on ThorstenMeyerAI.com

Age 18–24?Offer from Amazon

Prime made for students and young adults

  • Fast, free delivery for dorm and study essentials
  • Prime Video and Amazon Music included
  • Member-only deals
Try Prime for Young Adults Free trial for eligible 18–24 year olds
As an affiliate, we earn on qualifying purchases.

TL;DR

AI tools can produce large volumes of mathematical manuscripts, software changes and contract work, but the source material describes human verification as slower and harder to scale. Figures from industry reports and a peer-reviewed study point to review pressures in software, though some sources sell review tools and the evidence does not establish a single cause.

AI-generated work is becoming cheaper to produce, while verifying it remains a substantial human task, according to a report drawing on recent mathematics, software and contract-work examples. The mismatch matters because organisations may be able to generate more material than qualified reviewers can reliably assess, with consequences for quality, accountability and professional training.

The report says OpenAI published 722 mathematical manuscripts in a set of 372 families after posing about 4,000 problems to its model. It estimates that the average result took about three hours of compute to produce. Some results have been checked using Lean, a formal proof assistant; OpenAI cautioned that unformalized results could contain issues. The source contrasts this volume with the careful scrutiny given to an earlier result from the programme, described as a counterexample to an Erdős conjecture: five leading mathematicians were involved in its verification.

In software, the report cites several measurements that suggest review is not keeping pace uniformly with production. Faros AI reported that teams merged 98% more pull requests across low- and high-AI-adoption periods, while review time rose 91%. LinearB, analysing 8.1 million pull requests across 4,800 organisations, found AI-generated changes waited 4.6 times longer for review to begin and were accepted 32.7% of the time, compared with 84.4% for human-written changes. A peer-reviewed 2026 study found 61% of AI-agent pull requests received no human review before they were merged or closed. These figures come from different studies and should not be treated as directly comparable.

The source also points to OpenAI’s partnership with contract-software company Ironclad. In an evaluation across 11 contracting tasks, GPT-6 Astra met 55% of the criteria on average, according to the supplied material. That is described as an improvement over a previous model, but the remaining criteria still require attention before such work can be relied on. The source does not provide the evaluation protocol or a detailed breakdown of the criteria.

At a glance
reportWhen: Reported this week; software and contra…
The developmentA report argues that rising AI-generated output is increasing pressure on human review, drawing on OpenAI’s mathematics work, software-industry metrics and a contract-AI evaluation.
The Referee Shortage — Post-Labor
AI Dispatch · Post-Labor · 7 October 2026

The referee shortage: AI made doing cheap and checking expensive

OpenAI’s model produced a maths result in about three hours of compute. Verifying one earlier result took five of the world’s leading mathematicians. That ratio is the next decade of work: producing is cheap and abundant; trusting is slow, human and scarce.

One pattern, three fields
Mathematics
722
manuscripts, ~3h compute each

Some Lean-checked; OpenAI warns unformalized ones “could have issues.” Verification abundance, adjudication scarcity.

Software
+98% / +91%
more PRs merged / longer review

Faros AI. LinearB (8.1M PRs): AI changes wait 4.6× longer, accepted 32.7% vs 84.4%.

Professional workflows
55%
of criteria met — Astra on Ironclad

Real progress. Someone still has to find the other 45% before the work can be used.

Generation collapsed. Verification didn’t. (conceptual, not to scale)
Cost to produce a resultdown
Cost to check a resultnot down
No author intent

Machine output arrives without reasoning a reviewer can interrogate. It looks locally clean and gives no clue where it’s wrong.

Checks the answer, not the question

A prover confirms the proof proves its statement; tests confirm what tests check. Neither confirms it’s what was needed.

Someone must be accountable

Contracts are signed, designs stamped, papers defended. Responsibility is institutional — you can’t hold a model to it.

Illustrative: $1 of model time + 4 minutes of review at $45/hour. Halving the model price saves 12.5%; one extra review minute erases it. In that example, review is three-quarters of the bill.
What happens when referees run out — already visible
Rubber-stamping
61%

of AI-agent pull requests got no human review at all (EASE 2026). Zero-review merges up 31.3% (Faros).

Triage by suspicion
38%

of reviewers deliberately deprioritise AI changes (LinearB). Good machine work waits behind bad.

Producer as filter
~4,000 → 372

OpenAI chose which maths families were significant. When referees can’t keep up, the producer’s filter becomes the review.

The apprenticeship paradox: reviewers are made by doing the work. The work AI absorbs — writing code, drafting contracts, proving lemmas — is exactly what trained the reviewers. Demand for judgement rises as its supply line shrinks.
What to do
Price verification

Budget review hours next to model spend.

Formalise checks

Provers, types, tests, policy engines.

Tier the review

Experts only where consequences are high.

Fund the referees

Who profits from generation pays for checking.

Protect apprenticeship

Keep some production human for learners.

The take

The first automation question was which jobs AI would do. The better one is which jobs AI makes more necessary: the ones that check, adjudicate and take responsibility. Expect a referee premium — senior engineers, auditors, specialist lawyers, reviewing scientists become the binding constraint on how much AI output anyone can use.Accountability — standing behind a result — may be the most durable form of human work there is.

Sources: OpenAI maths release & Erdős verification as covered here; arXiv:2608.28997; OpenAI × Ironclad (6 Oct 2026); Faros AI; LinearB 2026 (8.1M PRs); Duma et al., EASE 2026 — via secondary reporting. Several code-review sources sell review tools. Review-cost example illustrative. Analysis is the author’s.
thorstenmeyerai.com

Review Capacity Sets the Pace

The practical constraint may shift from producing drafts to deciding which are safe, valid or useful. Review capacity affects how much AI output can be used: if qualified people cannot check it, organisations may delay work, accept it with less scrutiny or rely on the producer’s own judgment. Each choice carries different risks, from missed defects to bottlenecks that reduce the time savings AI was meant to provide.

The report’s examples also highlight that verification is not just a matter of checking whether an answer is internally consistent. A proof can be valid but address the wrong claim; software can pass tests that fail to capture the user’s need; a contract draft can miss an approval rule or jurisdiction-specific requirement. Human reviewers supply context and accountability that automated checks may not provide. The source argues that these responsibilities could make experienced engineers, lawyers, auditors and researchers more valuable, though it offers no labor-market data to quantify that effect.

There is a longer-term workforce question, too. Junior staff often gain expertise by doing the work that AI tools can now draft or generate. If those tasks disappear without replacement training, the future pool of people capable of reviewing complex work could shrink. That outcome is a concern raised by the report, not an established trend demonstrated by the cited figures.

Amazon

AI code review tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Three Fields, Similar Pressures

The examples span research, software and contracting, but they do not come from one coordinated study. In mathematics, formal proof systems such as Lean can check whether a proof follows from stated assumptions. That does not by itself establish that the theorem is important or that it answers the intended question. OpenAI’s warning about unformalized results is a qualification on the published set, not evidence that every such result is flawed.

Software metrics offer a larger data base, but the sources have limits. Faros AI and LinearB sell products related to engineering workflows or code review, and their findings should be read with that commercial interest in mind. The peer-reviewed 2026 study is a separate source, while the supplied material does not identify its sample or methods. Together, the figures indicate review-related concerns, but they do not prove that AI adoption alone caused every change.

The contract example is also an evaluation rather than evidence of routine deployment outcomes. The reported 55% average across 11 tasks indicates that the model met some, but not all, evaluation criteria. Without the task details, scoring method and comparison figures, readers cannot judge how the result translates to actual contract work.

““Verification abundance, adjudication scarcity.””

— The report’s summary of a recent paper

What the Figures Cannot Show

The cited software measurements use different samples and methods, and the source does not give enough detail to determine whether their time periods or definitions align. Several sources sell code-review tools, which is relevant when interpreting their findings. The numbers show reported associations, not proof that AI use alone caused longer waits, lower acceptance or missing reviews.

It is also unclear how many of the 722 mathematics manuscripts have since been independently checked, how often unformalized results contain errors, or how the Ironclad evaluation’s 11 tasks were selected and scored. The source’s broader claims about future job value and a shortage of experienced reviewers are interpretations; it provides no hiring, wage or workforce projections to confirm them.

More Evidence on Review Outcomes

The next useful evidence would be follow-up reporting on how many AI-generated results are independently verified, corrected or withdrawn, alongside transparent methods for the cited software and contract evaluations. Organisations adopting these systems will also need to track review time, defect rates and accountability—not just how quickly work is generated. The source material does not identify a scheduled next release or study, so the timing of further evidence is unknown.

Key Questions

What is the main development described?

The report says AI is generating large volumes of mathematical, software and contract work, while human verification remains slower and limited. It supports that argument with examples and studies, but does not establish that every field faces the same scale of pressure.

Did OpenAI publish 722 verified mathematical results?

The source says OpenAI published 722 mathematical manuscripts in 372 families after posing about 4,000 problems. Some were formally checked in Lean; OpenAI cautioned that unformalized results could have issues. The source does not say all 722 were formally verified.

What did the software studies report?

Faros AI reported more pull requests merged alongside longer review time, while LinearB reported longer waits for AI-generated changes to enter review and lower acceptance rates than for human-written changes. A separate peer-reviewed 2026 study reported that 61% of AI-agent pull requests received no human review before merging or closing. These are findings from distinct sources, not a single combined measure.

Does the evidence prove AI causes weaker review?

No. The cited figures suggest review pressures and differences associated with AI-generated work, but the supplied material does not establish causation. Some sources also sell code-review tools, and the studies’ methods and time periods are not fully described here.

What remains unknown about the contract evaluation?

The source reports that GPT-6 Astra met 55% of evaluation criteria on average across 11 tasks, but does not provide the task selection, scoring rules, detailed results or comparison figures. It is not clear how well that evaluation predicts performance in routine legal work.

Source: ThorstenMeyerAI.com

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

2027 Content Creator Laptop Picks: Top 10 With AI Power

Discover the best 2027 laptops for content creators, featuring AI-enhanced performance, high-quality displays, and portability. Top 10 picks explained.

The Real Cost Of A Local-Inference Rig In 2026

An in-depth analysis of the hardware costs for local AI inference in 2026, highlighting key factors, value strategies, and future implications.

OpenWrt One’s Role In Shaping Open Hardware Router Trends

Analysis of how OpenWrt One is shaping open hardware router development and influencing industry trends.

The Future Of AI: 10 Key Developments To Watch In 2026

An overview of the 10 most significant AI advancements to watch in 2026, based on current confirmed trends and emerging claims.