The Rising Price Of Reviewing AI-Generated Work
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: The Rising Price Of Reviewing AI-Generated Work on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get office and shipping supplies delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

Recent examples and industry studies point to a growing gap between the volume of AI-generated work and the capacity to verify it. The reported figures vary by source, and some come from companies that sell review tools, but they raise questions about quality control, accountability and how future expert reviewers will be trained.

AI-generated work is arriving faster than many teams can review it, according to a source roundup spanning mathematics, software and contract work. The cited examples include 722 mathematical manuscripts published by OpenAI this week and software-industry data associating higher AI use with longer review waits, highlighting a potential bottleneck in deciding whether machine-produced work is reliable and fit for use.

The mathematics example illustrates the gap between producing work and validating it. OpenAI posed about 4,000 problems to a model and published 722 manuscripts grouped into 372 families. The source says the average result took about three hours of compute. Some results were formally checked using Lean, a proof assistant; OpenAI cautioned that some unformalized results “could have issues.” A separate result from the programme, described as a counterexample to an old Erdős conjecture, drew careful verification from five leading mathematicians, according to the source.

Software figures point in a similar direction, though they come with limits. Faros AI reported that teams merged 98% more pull requests across low- and high-AI-adoption periods while review time rose 91%. LinearB, analysing 8.1 million pull requests across 4,800 organisations, reported that AI-generated changes waited 4.6 times longer for review to begin and were accepted 32.7% of the time, compared with 84.4% for human-written changes. The source also cites a peer-reviewed 2026 study in which 61% of AI-agent pull requests received no human review before being merged or closed.

The figures do not establish that AI alone caused these outcomes. The source notes that several companies behind the software data sell code-review tools, a potential conflict readers should keep in mind. In contract work, OpenAI’s partnership with contract-software company Ironclad provides another example: the source says GPT-6 Astra met 55% of evaluation criteria on average across 11 tasks, an improvement over its predecessor. That result still leaves a substantial share of criteria unmet, though the source does not specify what each shortfall involved.

At a glance
reportWhen: Developing; the source cites recent stu…
The developmentA source roundup reports that AI systems are producing more mathematical manuscripts, code changes and contract work, while human verification remains a time-consuming constraint.
The Referee Shortage — Post-Labor
AI Dispatch · Post-Labor · 7 October 2026

The referee shortage: AI made doing cheap and checking expensive

OpenAI’s model produced a maths result in about three hours of compute. Verifying one earlier result took five of the world’s leading mathematicians. That ratio is the next decade of work: producing is cheap and abundant; trusting is slow, human and scarce.

One pattern, three fields
Mathematics
722
manuscripts, ~3h compute each

Some Lean-checked; OpenAI warns unformalized ones “could have issues.” Verification abundance, adjudication scarcity.

Software
+98% / +91%
more PRs merged / longer review

Faros AI. LinearB (8.1M PRs): AI changes wait 4.6× longer, accepted 32.7% vs 84.4%.

Professional workflows
55%
of criteria met — Astra on Ironclad

Real progress. Someone still has to find the other 45% before the work can be used.

Generation collapsed. Verification didn’t. (conceptual, not to scale)
Cost to produce a resultdown
Cost to check a resultnot down
No author intent

Machine output arrives without reasoning a reviewer can interrogate. It looks locally clean and gives no clue where it’s wrong.

Checks the answer, not the question

A prover confirms the proof proves its statement; tests confirm what tests check. Neither confirms it’s what was needed.

Someone must be accountable

Contracts are signed, designs stamped, papers defended. Responsibility is institutional — you can’t hold a model to it.

Illustrative: $1 of model time + 4 minutes of review at $45/hour. Halving the model price saves 12.5%; one extra review minute erases it. In that example, review is three-quarters of the bill.
What happens when referees run out — already visible
Rubber-stamping
61%

of AI-agent pull requests got no human review at all (EASE 2026). Zero-review merges up 31.3% (Faros).

Triage by suspicion
38%

of reviewers deliberately deprioritise AI changes (LinearB). Good machine work waits behind bad.

Producer as filter
~4,000 → 372

OpenAI chose which maths families were significant. When referees can’t keep up, the producer’s filter becomes the review.

The apprenticeship paradox: reviewers are made by doing the work. The work AI absorbs — writing code, drafting contracts, proving lemmas — is exactly what trained the reviewers. Demand for judgement rises as its supply line shrinks.
What to do
Price verification

Budget review hours next to model spend.

Formalise checks

Provers, types, tests, policy engines.

Tier the review

Experts only where consequences are high.

Fund the referees

Who profits from generation pays for checking.

Protect apprenticeship

Keep some production human for learners.

The take

The first automation question was which jobs AI would do. The better one is which jobs AI makes more necessary: the ones that check, adjudicate and take responsibility. Expect a referee premium — senior engineers, auditors, specialist lawyers, reviewing scientists become the binding constraint on how much AI output anyone can use.Accountability — standing behind a result — may be the most durable form of human work there is.

Sources: OpenAI maths release & Erdős verification as covered here; arXiv:2608.28997; OpenAI × Ironclad (6 Oct 2026); Faros AI; LinearB 2026 (8.1M PRs); Duma et al., EASE 2026 — via secondary reporting. Several code-review sources sell review tools. Review-cost example illustrative. Analysis is the author’s.
thorstenmeyerai.com

Human Review Becomes the Bottleneck

If organizations can generate drafts, code and research more quickly than experts can check them, the amount of usable AI output may be limited by review capacity, not generation speed. That affects whether teams can safely deploy code, rely on contract analysis or build research on machine-produced results. It also changes where organizations may need to invest: not only in tools that generate work, but in experienced people and processes that can evaluate it.

The source describes three possible responses when reviews fall behind: work may be merged without review, AI-generated changes may be put lower in the queue, or producers may decide which outputs deserve scrutiny. Each creates a different risk. Unreviewed work can carry errors forward; blanket suspicion can delay sound contributions; and relying on the producer’s own selection leaves less independent oversight. The cited data suggest these patterns are present in some settings, but do not show how widespread they are across industries.

There is also a workforce question. Experienced reviewers usually develop judgment through years of doing the underlying work: writing code, drafting contracts or producing mathematical results. If entry-level tasks shift toward prompting and editing AI output, organizations may need to create new ways for junior workers to gain the experience that makes later review credible. That is a concern raised by the source, not a measured forecast of future staffing shortages.

Amazon

AI review tools for code

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Evidence Across Three Workflows

The examples cover different kinds of work, so their numbers should not be treated as a single comparable measure. Mathematics relies on proof and expert interpretation; software teams track pull requests, review starts and acceptance; contract evaluations measure performance against task criteria. Together, they illustrate a shared distinction: producing an answer is not the same as establishing that it is correct or appropriate.

Automated checks can reduce some review work. A proof assistant can verify that a formal proof follows stated rules, and software tests can check specified behavior. But those checks cannot, by themselves, establish that the theorem addresses the intended question or that the tests cover the real requirements. Likewise, a contract system can meet many evaluation criteria while still requiring professional judgment before a document is used. Human review also matters because legal and institutional systems generally assign responsibility to people, such as the lawyer signing a contract or engineer approving a design.

Amazon

mathematical proof verification software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How Broad Is the Review Gap?

The cited figures do not provide a consistent measure of review quality across fields. The source does not give the underlying methods or full definitions for every statistic, and it flags that some software-data providers sell review products. It is also unclear how much of the longer wait for AI-generated code reflects its quality, how teams route that code, or differences in the kinds of changes being submitted.

The mathematics example also leaves open how many of the 722 manuscripts were independently checked, how many contained errors and how the published set was selected from the roughly 4,000 problems posed. The contract benchmark’s 55% average does not identify which criteria were missed or whether those omissions would prevent use in a real workflow. The evidence supports concern about a possible capacity gap, but it does not establish a universal shortage of reviewers or show that all AI-produced work requires more review than human work.

Amazon

AI-powered code review tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Tracking Review Capacity and Quality

The next useful evidence will show whether review times, defect rates and acceptance outcomes change as AI adoption grows, using transparent methods and comparisons that distinguish types of work. Organizations will also need to report what counts as a review: a quick approval, a substantive human check, or an automated test are not equivalent safeguards.

For employers, the immediate question is how to expand output without weakening accountability or removing the work through which junior staff learn. That may mean setting risk-based review requirements, recording who approved consequential work and preserving supervised opportunities to build professional judgment. The cited material does not announce a new industry standard or policy; it leaves the scale of the problem and the right response open.

Amazon

contract review software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is the main development?

The source describes a widening gap between the volume of AI-generated work and the time available for people to verify it, using examples from mathematics, software and contract tasks.

Does the evidence prove AI-generated work is less reliable?

No. The figures are drawn from different studies and company analyses. Some software data show lower acceptance rates for AI-generated changes, but the source does not establish that AI use alone caused the difference or that the results apply to every team.

Why can’t automated checks handle all verification?

Automated systems can check specified properties, such as whether a formal proof follows rules or code passes tests. They may not determine whether the right problem was addressed, whether tests reflect real needs, or who should take responsibility for the result.

What does the source say about training future reviewers?

It raises a concern that junior workers may get fewer chances to build expertise if AI takes over tasks that once served as training. The source does not provide evidence that a future reviewer shortage has already been measured.

Source: ThorstenMeyerAI.com

This content is for general information only and is not financial, tax or legal advice. Consult a qualified professional for decisions about your money.
HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Machine Economy — Capital-Heavy, Human-Light, Trading With Itself

Analysis of the emerging machine economy where AI-driven firms operate with minimal human input, reshaping markets and economic structures.

The Delegation Ladder: The Four Agentic Loops, and What Each One Lets You Stop Doing

An analysis of the four agentic loops in AI design, explaining what each allows you to stop doing and how they impact AI processes and management.

Optimize Your Work With These 12 AI Tools In 2026

Discover the 12 best AI productivity tools in 2026 that can transform your workflow, automate tasks, and enhance decision-making for professionals.

Agentic Loop Failure Modes: A Production Taxonomy at the End of Year One

A comprehensive taxonomy of failure modes in production agentic AI systems after one year of deployment, aiding debugging and architectural decisions.