AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: When AI Makes Work Cheap, Careful Checking Gets Pricier on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the latest gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

A report from ThorstenMeyerAI.com describes a widening gap between cheap AI-generated work and the human effort needed to verify it. It cites OpenAI’s 722 mathematics manuscripts and software-workflow data showing longer review waits, while cautioning that several software sources sell code-review tools.

A report published this week argues that AI is lowering the cost of producing work faster than it is lowering the cost of checking it, a gap that could make expert review a bottleneck in mathematics, software and professional services. It points to OpenAI’s publication of 722 mathematical manuscripts and software data showing more code changes but longer review waits; the figures come from the report and cited studies, not a single independent evaluation.

According to the report, OpenAI’s model was given about 4,000 mathematical problems and generated 722 manuscripts grouped into 372 families. Some results were formally checked using Lean, a proof assistant. OpenAI cautioned that some results without formal verification could have issues. The report contrasts this output with the careful checking by five leading mathematicians of an earlier result from the same program, described as a counterexample to an old Erdős conjecture. The account does not provide the reviewers’ individual assessments or a complete independent audit of the manuscripts.

The report says software teams are also handling more AI-assisted work while facing review pressure. Faros AI found that teams merged 98% more pull requests in periods of high AI adoption, while review time rose 91%. LinearB, analysing 8.1 million pull requests from 4,800 organisations, reported that AI-generated changes waited 4.6 times longer for review to begin and were accepted 32.7% of the time, compared with 84.4% for human-written changes. These are findings as presented by the source; several cited companies sell code-review products, a potential interest readers should keep in mind.

A peer-reviewed 2026 study cited in the report found 61% of AI-agent pull requests received no human review before being merged or closed. Faros also reported a 31.3% rise in merges with no review during high-adoption periods. In a separate professional-work example, the source says OpenAI partnered with contract-software company Ironclad to train GPT-6 Astra on contracting workflows. On 11 tasks, Astra met an average of 55% of evaluation criteria. That result indicates improvement over the previous model, according to the report, but also leaves substantial criteria unmet; the source does not detail the full evaluation design.

At a glance
analysisWhen: Published this week, according to the s…
The developmentA report argues that growing AI output is making expert review a scarce and increasingly important part of mathematics, software development and professional workflows.
The Referee Shortage — Post-Labor
AI Dispatch · Post-Labor · 7 October 2026

The referee shortage: AI made doing cheap and checking expensive

OpenAI’s model produced a maths result in about three hours of compute. Verifying one earlier result took five of the world’s leading mathematicians. That ratio is the next decade of work: producing is cheap and abundant; trusting is slow, human and scarce.

One pattern, three fields
Mathematics
722
manuscripts, ~3h compute each

Some Lean-checked; OpenAI warns unformalized ones “could have issues.” Verification abundance, adjudication scarcity.

Software
+98% / +91%
more PRs merged / longer review

Faros AI. LinearB (8.1M PRs): AI changes wait 4.6× longer, accepted 32.7% vs 84.4%.

Professional workflows
55%
of criteria met — Astra on Ironclad

Real progress. Someone still has to find the other 45% before the work can be used.

Generation collapsed. Verification didn’t. (conceptual, not to scale)
Cost to produce a resultdown
Cost to check a resultnot down
No author intent

Machine output arrives without reasoning a reviewer can interrogate. It looks locally clean and gives no clue where it’s wrong.

Checks the answer, not the question

A prover confirms the proof proves its statement; tests confirm what tests check. Neither confirms it’s what was needed.

Someone must be accountable

Contracts are signed, designs stamped, papers defended. Responsibility is institutional — you can’t hold a model to it.

Illustrative: $1 of model time + 4 minutes of review at $45/hour. Halving the model price saves 12.5%; one extra review minute erases it. In that example, review is three-quarters of the bill.
What happens when referees run out — already visible
Rubber-stamping
61%

of AI-agent pull requests got no human review at all (EASE 2026). Zero-review merges up 31.3% (Faros).

Triage by suspicion
38%

of reviewers deliberately deprioritise AI changes (LinearB). Good machine work waits behind bad.

Producer as filter
~4,000 → 372

OpenAI chose which maths families were significant. When referees can’t keep up, the producer’s filter becomes the review.

The apprenticeship paradox: reviewers are made by doing the work. The work AI absorbs — writing code, drafting contracts, proving lemmas — is exactly what trained the reviewers. Demand for judgement rises as its supply line shrinks.
What to do
Price verification

Budget review hours next to model spend.

Formalise checks

Provers, types, tests, policy engines.

Tier the review

Experts only where consequences are high.

Fund the referees

Who profits from generation pays for checking.

Protect apprenticeship

Keep some production human for learners.

The take

The first automation question was which jobs AI would do. The better one is which jobs AI makes more necessary: the ones that check, adjudicate and take responsibility. Expect a referee premium — senior engineers, auditors, specialist lawyers, reviewing scientists become the binding constraint on how much AI output anyone can use.Accountability — standing behind a result — may be the most durable form of human work there is.

Sources: OpenAI maths release & Erdős verification as covered here; arXiv:2608.28997; OpenAI × Ironclad (6 Oct 2026); Faros AI; LinearB 2026 (8.1M PRs); Duma et al., EASE 2026 — via secondary reporting. Several code-review sources sell review tools. Review-cost example illustrative. Analysis is the author’s.
thorstenmeyerai.com

Review Capacity May Limit AI Use

The practical issue is not only whether AI can produce a plausible proof, code change or contract draft. Organisations need to determine whether the work answers the right question, follows relevant rules and can be responsibly used. Where expert review is limited, more output may mean longer queues, unchecked changes or selective review, rather than a matching increase in useful work.

The report describes this as a shift in the value of labour: when production becomes easier, people who can reliably judge its quality may become more important. Senior engineers, specialist lawyers, auditors and scientific reviewers could face greater demand for their time. That is an interpretation, not a measured forecast of wages or hiring. The cited figures show pressure in particular settings, but do not establish how widespread or permanent the trend is.

There is also a training concern. Experienced reviewers often develop judgement by doing the underlying work themselves: writing code, drafting agreements or proving results. If AI replaces too many early-career tasks, employers may reduce the opportunities through which future specialists gain that experience. Whether organisations can preserve those pathways while adopting AI remains an open question.

Amazon

software code review tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Three Fields, Similar Verification Gaps

The report brings together examples from mathematics, software and contracting rather than presenting one unified study. In mathematics, formal proof systems can check whether a proof follows from stated assumptions. In software, tests can confirm behaviour covered by those tests. Neither method alone establishes that the original theorem, requirements or tests capture the real problem. The report says the distinction between abundant verification and scarce adjudication is central: checking a defined object may be automated, while deciding whether it is the right object still calls for judgement.

The software statistics cited come from different organisations, datasets and measures. Faros compares team activity across periods of higher and lower AI adoption; LinearB examines pull requests across thousands of organisations; the 2026 study concerns AI-agent pull requests. They should not be treated as interchangeable measures or proof that AI alone caused each outcome. Taken together, they illustrate the report’s argument that production volume and review capacity can move in different directions.

The report also notes three possible responses when review capacity is stretched: changes may be merged without review, reviewers may deprioritise AI-generated work, or producers may decide which results are worth submitting. Each can affect quality or fairness in a different way. The cited material does not quantify how often all three occur across industries.

How Broad Is the Review Shortage?

The cited evidence does not establish that review bottlenecks are equally severe across industries or organisations. The report provides selected statistics, but the measures differ and some sources sell tools related to code review. The precise methods, adoption periods and comparison groups are not fully reproduced in the source material.

It is also unclear how much of the reported increase in review time or unreviewed work is caused by AI, rather than changes in team size, project complexity or other workflow factors. The Ironclad example gives an average score across 11 tasks but does not specify the criteria, task mix or independent validation. No wage, employment or productivity data in the material confirms a broader economic “referee premium.”

Finally, automated checking may improve. The source argues that tools cannot fully resolve questions of intent, relevance and accountability, but it does not measure how much human review can be reduced through better systems. The scale and timing of any lasting shift remain uncertain.

Track Review Outcomes and Training

The next useful evidence will show whether review queues and unreviewed AI-generated work persist as teams gain experience with these tools. Organisations and researchers would need comparable measures over time, including review duration, defect rates, rework, and whether reviewers had enough information to assess a change. For mathematics and contracting, independent evaluations could clarify how often generated work meets standards beyond a model’s own reported score.

Employers will also need to decide how junior staff can learn the work that supports later expert judgement. The source report raises that concern but does not identify a settled solution. For now, its central finding is a question of capacity: AI can raise the amount of work produced, while the supply of accountable human review may not grow at the same rate.

Key Questions

What is the main development described?

A report argues that AI is making it cheaper to produce work while leaving expert review comparatively slow and scarce. It uses examples from mathematics, software and contract workflows.

Did OpenAI publish 722 verified mathematical results?

The report says OpenAI published 722 manuscripts generated from about 4,000 problems. Some were formally checked in Lean, but OpenAI cautioned that unformalized results could have issues. The source does not say all 722 were independently verified.

What do the software figures show?

The report cites studies reporting more pull requests, longer waits for review and substantial numbers of AI-agent changes receiving no human review. The studies use different datasets and methods, and some cited companies sell code-review tools.

Does this prove AI is reducing jobs?

No. The material describes workflow and review patterns, not a demonstrated net effect on employment. It raises a concern that fewer entry-level drafting or coding tasks could weaken how future reviewers gain experience.

What remains uncertain?

It is unclear how widely the review bottleneck applies, how much AI causes the reported changes, and whether improved checking tools can reduce the need for human judgement. The source also does not establish a measured wage premium for reviewers.

Source: ThorstenMeyerAI.com

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Warhammer 40,000: Space Marine 2 Enters The Steam Most-played Chart

The latest Warhammer 40,000: Space Marine 2 has entered Steam’s most-played games chart, reaching rank 94 with a peak of over 16,500 players.

I Tried All The Best (And Worst) Doorbells

A comprehensive review of popular and unpopular doorbells, exploring their features, performance, and user experiences amid rising interest in home security tech.

2DWillNeverDie

Interest in the phrase 2DWillNeverDie is surging online. The trigger is unconfirmed, but the 2D-vs-3D animation debate is resurfacing.

The Rise Of ‘System One’ AI: What Jev’s Research Means For The Future

TypeSafe’s Jev introduces a new class of AI, System One models, focused on structured decisions, promising faster, cheaper automation without text generation.