🔍 Read the full analysis: The Rising Price Of Reviewing AI-Generated Work on ThorstenMeyerAI.com
Get office and shipping supplies delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
Recent examples and industry studies point to a growing gap between the volume of AI-generated work and the capacity to verify it. The reported figures vary by source, and some come from companies that sell review tools, but they raise questions about quality control, accountability and how future expert reviewers will be trained.
AI-generated work is arriving faster than many teams can review it, according to a source roundup spanning mathematics, software and contract work. The cited examples include 722 mathematical manuscripts published by OpenAI this week and software-industry data associating higher AI use with longer review waits, highlighting a potential bottleneck in deciding whether machine-produced work is reliable and fit for use.
The mathematics example illustrates the gap between producing work and validating it. OpenAI posed about 4,000 problems to a model and published 722 manuscripts grouped into 372 families. The source says the average result took about three hours of compute. Some results were formally checked using Lean, a proof assistant; OpenAI cautioned that some unformalized results “could have issues.” A separate result from the programme, described as a counterexample to an old Erdős conjecture, drew careful verification from five leading mathematicians, according to the source.
Software figures point in a similar direction, though they come with limits. Faros AI reported that teams merged 98% more pull requests across low- and high-AI-adoption periods while review time rose 91%. LinearB, analysing 8.1 million pull requests across 4,800 organisations, reported that AI-generated changes waited 4.6 times longer for review to begin and were accepted 32.7% of the time, compared with 84.4% for human-written changes. The source also cites a peer-reviewed 2026 study in which 61% of AI-agent pull requests received no human review before being merged or closed.
The figures do not establish that AI alone caused these outcomes. The source notes that several companies behind the software data sell code-review tools, a potential conflict readers should keep in mind. In contract work, OpenAI’s partnership with contract-software company Ironclad provides another example: the source says GPT-6 Astra met 55% of evaluation criteria on average across 11 tasks, an improvement over its predecessor. That result still leaves a substantial share of criteria unmet, though the source does not specify what each shortfall involved.
The referee shortage: AI made doing cheap and checking expensive
OpenAI’s model produced a maths result in about three hours of compute. Verifying one earlier result took five of the world’s leading mathematicians. That ratio is the next decade of work: producing is cheap and abundant; trusting is slow, human and scarce.
Some Lean-checked; OpenAI warns unformalized ones “could have issues.” Verification abundance, adjudication scarcity.
Faros AI. LinearB (8.1M PRs): AI changes wait 4.6× longer, accepted 32.7% vs 84.4%.
Real progress. Someone still has to find the other 45% before the work can be used.
Machine output arrives without reasoning a reviewer can interrogate. It looks locally clean and gives no clue where it’s wrong.
A prover confirms the proof proves its statement; tests confirm what tests check. Neither confirms it’s what was needed.
Contracts are signed, designs stamped, papers defended. Responsibility is institutional — you can’t hold a model to it.
of AI-agent pull requests got no human review at all (EASE 2026). Zero-review merges up 31.3% (Faros).
of reviewers deliberately deprioritise AI changes (LinearB). Good machine work waits behind bad.
OpenAI chose which maths families were significant. When referees can’t keep up, the producer’s filter becomes the review.
Budget review hours next to model spend.
Provers, types, tests, policy engines.
Experts only where consequences are high.
Who profits from generation pays for checking.
Keep some production human for learners.
The first automation question was which jobs AI would do. The better one is which jobs AI makes more necessary: the ones that check, adjudicate and take responsibility. Expect a referee premium — senior engineers, auditors, specialist lawyers, reviewing scientists become the binding constraint on how much AI output anyone can use.Accountability — standing behind a result — may be the most durable form of human work there is.
Human Review Becomes the Bottleneck
If organizations can generate drafts, code and research more quickly than experts can check them, the amount of usable AI output may be limited by review capacity, not generation speed. That affects whether teams can safely deploy code, rely on contract analysis or build research on machine-produced results. It also changes where organizations may need to invest: not only in tools that generate work, but in experienced people and processes that can evaluate it.
The source describes three possible responses when reviews fall behind: work may be merged without review, AI-generated changes may be put lower in the queue, or producers may decide which outputs deserve scrutiny. Each creates a different risk. Unreviewed work can carry errors forward; blanket suspicion can delay sound contributions; and relying on the producer’s own selection leaves less independent oversight. The cited data suggest these patterns are present in some settings, but do not show how widespread they are across industries.
There is also a workforce question. Experienced reviewers usually develop judgment through years of doing the underlying work: writing code, drafting contracts or producing mathematical results. If entry-level tasks shift toward prompting and editing AI output, organizations may need to create new ways for junior workers to gain the experience that makes later review credible. That is a concern raised by the source, not a measured forecast of future staffing shortages.
As an affiliate, we earn on qualifying purchases.
Evidence Across Three Workflows
The examples cover different kinds of work, so their numbers should not be treated as a single comparable measure. Mathematics relies on proof and expert interpretation; software teams track pull requests, review starts and acceptance; contract evaluations measure performance against task criteria. Together, they illustrate a shared distinction: producing an answer is not the same as establishing that it is correct or appropriate.
Automated checks can reduce some review work. A proof assistant can verify that a formal proof follows stated rules, and software tests can check specified behavior. But those checks cannot, by themselves, establish that the theorem addresses the intended question or that the tests cover the real requirements. Likewise, a contract system can meet many evaluation criteria while still requiring professional judgment before a document is used. Human review also matters because legal and institutional systems generally assign responsibility to people, such as the lawyer signing a contract or engineer approving a design.
mathematical proof verification software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
How Broad Is the Review Gap?
The cited figures do not provide a consistent measure of review quality across fields. The source does not give the underlying methods or full definitions for every statistic, and it flags that some software-data providers sell review products. It is also unclear how much of the longer wait for AI-generated code reflects its quality, how teams route that code, or differences in the kinds of changes being submitted.
The mathematics example also leaves open how many of the 722 manuscripts were independently checked, how many contained errors and how the published set was selected from the roughly 4,000 problems posed. The contract benchmark’s 55% average does not identify which criteria were missed or whether those omissions would prevent use in a real workflow. The evidence supports concern about a possible capacity gap, but it does not establish a universal shortage of reviewers or show that all AI-produced work requires more review than human work.
As an affiliate, we earn on qualifying purchases.
Tracking Review Capacity and Quality
The next useful evidence will show whether review times, defect rates and acceptance outcomes change as AI adoption grows, using transparent methods and comparisons that distinguish types of work. Organizations will also need to report what counts as a review: a quick approval, a substantive human check, or an automated test are not equivalent safeguards.
For employers, the immediate question is how to expand output without weakening accountability or removing the work through which junior staff learn. That may mean setting risk-based review requirements, recording who approved consequential work and preserving supervised opportunities to build professional judgment. The cited material does not announce a new industry standard or policy; it leaves the scale of the problem and the right response open.
As an affiliate, we earn on qualifying purchases.
Key Questions
What is the main development?
The source describes a widening gap between the volume of AI-generated work and the time available for people to verify it, using examples from mathematics, software and contract tasks.
Does the evidence prove AI-generated work is less reliable?
No. The figures are drawn from different studies and company analyses. Some software data show lower acceptance rates for AI-generated changes, but the source does not establish that AI use alone caused the difference or that the results apply to every team.
Why can’t automated checks handle all verification?
Automated systems can check specified properties, such as whether a formal proof follows rules or code passes tests. They may not determine whether the right problem was addressed, whether tests reflect real needs, or who should take responsibility for the result.
What does the source say about training future reviewers?
It raises a concern that junior workers may get fewer chances to build expertise if AI takes over tasks that once served as training. The source does not provide evidence that a future reviewer shortage has already been measured.
Source: ThorstenMeyerAI.com
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
