The AI Review Bottleneck: Cheap Output, Expensive Oversight
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: The AI Review Bottleneck: Cheap Output, Expensive Oversight on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get monitors, keyboards and dev gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

OpenAI reported producing 722 mathematical manuscripts from about 4,000 problems, while several software-industry analyses describe longer review queues and less consistent human checking of AI-generated code. The figures come from different sources and have limitations, but they point to a shared concern: producing work with AI can be faster than verifying it. How organizations preserve review quality and train future experts remains unresolved.

OpenAI reported producing 722 mathematical manuscripts from about 4,000 problems, an output that illustrates how quickly AI can generate work for experts to assess. A report published by ThorstenMeyerAI.com compares that volume with the human effort needed to verify earlier AI-assisted mathematical work and with software-industry data indicating that review can lag behind code production. The datasets do not establish one universal measure of an AI review bottleneck, but they raise a practical issue for employers and professional fields: more generated work does not automatically mean more work that can be trusted or used.

According to the source report, OpenAI’s model produced the manuscripts across 372 families of problems, after being given roughly 4,000 problems. Some results were checked using Lean, a system for formal proof verification. OpenAI cautioned that some results without formal verification could have issues. The report says the average result took about three hours of compute to produce, but it does not provide a full breakdown of that average or the human effort required to check all 722 manuscripts.

The report contrasts that output with the response to an earlier result from the same programme: a proposed counterexample to an Erdős conjecture that received careful scrutiny from five leading mathematicians. That comparison illustrates the difference between generating a candidate result and deciding whether it is correct, relevant and significant. It does not, by itself, establish that every manuscript requires five reviewers or the same level of scrutiny.

In software, the report cites several analyses with different samples and methods. Faros AI said teams in high-AI-adoption periods merged 98% more pull requests while review time rose 91%. LinearB, analysing 8.1 million pull requests from 4,800 organizations, reported that AI-generated changes waited 4.6 times longer for review to begin and were accepted 32.7% of the time, compared with 84.4% for human-written changes. The report also cites a peer-reviewed 2026 study in which 61% of AI-agent pull requests received no human review before being merged or closed. These findings are not interchangeable, and the source notes that several data providers sell code-review products.

At a glance
reportWhen: Reported this week; software and study…
The developmentA source report uses OpenAI’s mathematics output and software-review data to highlight the growing mismatch between AI-generated work and human capacity to verify it.
The Referee Shortage — Post-Labor
AI Dispatch · Post-Labor · 7 October 2026

The referee shortage: AI made doing cheap and checking expensive

OpenAI’s model produced a maths result in about three hours of compute. Verifying one earlier result took five of the world’s leading mathematicians. That ratio is the next decade of work: producing is cheap and abundant; trusting is slow, human and scarce.

One pattern, three fields
Mathematics
722
manuscripts, ~3h compute each

Some Lean-checked; OpenAI warns unformalized ones “could have issues.” Verification abundance, adjudication scarcity.

Software
+98% / +91%
more PRs merged / longer review

Faros AI. LinearB (8.1M PRs): AI changes wait 4.6× longer, accepted 32.7% vs 84.4%.

Professional workflows
55%
of criteria met — Astra on Ironclad

Real progress. Someone still has to find the other 45% before the work can be used.

Generation collapsed. Verification didn’t. (conceptual, not to scale)
Cost to produce a resultdown
Cost to check a resultnot down
No author intent

Machine output arrives without reasoning a reviewer can interrogate. It looks locally clean and gives no clue where it’s wrong.

Checks the answer, not the question

A prover confirms the proof proves its statement; tests confirm what tests check. Neither confirms it’s what was needed.

Someone must be accountable

Contracts are signed, designs stamped, papers defended. Responsibility is institutional — you can’t hold a model to it.

Illustrative: $1 of model time + 4 minutes of review at $45/hour. Halving the model price saves 12.5%; one extra review minute erases it. In that example, review is three-quarters of the bill.
What happens when referees run out — already visible
Rubber-stamping
61%

of AI-agent pull requests got no human review at all (EASE 2026). Zero-review merges up 31.3% (Faros).

Triage by suspicion
38%

of reviewers deliberately deprioritise AI changes (LinearB). Good machine work waits behind bad.

Producer as filter
~4,000 → 372

OpenAI chose which maths families were significant. When referees can’t keep up, the producer’s filter becomes the review.

The apprenticeship paradox: reviewers are made by doing the work. The work AI absorbs — writing code, drafting contracts, proving lemmas — is exactly what trained the reviewers. Demand for judgement rises as its supply line shrinks.
What to do
Price verification

Budget review hours next to model spend.

Formalise checks

Provers, types, tests, policy engines.

Tier the review

Experts only where consequences are high.

Fund the referees

Who profits from generation pays for checking.

Protect apprenticeship

Keep some production human for learners.

The take

The first automation question was which jobs AI would do. The better one is which jobs AI makes more necessary: the ones that check, adjudicate and take responsibility. Expect a referee premium — senior engineers, auditors, specialist lawyers, reviewing scientists become the binding constraint on how much AI output anyone can use.Accountability — standing behind a result — may be the most durable form of human work there is.

Sources: OpenAI maths release & Erdős verification as covered here; arXiv:2608.28997; OpenAI × Ironclad (6 Oct 2026); Faros AI; LinearB 2026 (8.1M PRs); Duma et al., EASE 2026 — via secondary reporting. Several code-review sources sell review tools. Review-cost example illustrative. Analysis is the author’s.
thorstenmeyerai.com

Review Capacity Shapes AI Adoption

If production grows faster than review capacity, an organization may have more drafts, proofs or code than qualified people can reliably assess. That can turn review time into a limit on usable AI output, even when the tools make initial production faster. It can also shift costs rather than remove them: time saved on drafting may be spent checking, correcting and documenting the result.

The software figures point to several possible responses, each with risks. Teams may merge work with little or no review, slow down AI-generated changes because they distrust them, or rely more heavily on the producer’s own quality filter. The figures cited do not show that every organization follows these patterns, but they make clear why adoption rates alone cannot describe productivity. Reliability and review practices matter alongside how much work a model generates.

The source report also identifies a workforce concern: junior employees often develop judgment through drafting, coding and solving problems themselves. If AI takes over much of that practice, organizations may need to deliberately create other ways for early-career workers to gain the experience required for senior review. That is a risk to consider, not an established outcome of the reported data.

Amazon

AI code review tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

From Generated Work to Verified Results

AI tools can generate candidate work at a pace that differs from the slower processes used to validate it. In mathematics, formal proof systems such as Lean can check whether a proof follows from its stated assumptions. That is valuable, but it does not decide whether the theorem addresses the right question or whether the result matters. In software, tests can check specified behavior without proving that the tests capture every requirement.

The same distinction applies in professional services. The source report describes OpenAI’s partnership with contract-software company Ironclad and says GPT-6 Astra met 55% of evaluation criteria on average across 11 tasks. The report characterizes this as an improvement over a previous model, but the supplied material does not identify the full evaluation method or establish how performance translates to routine legal work. A remaining portion of evaluation criteria is not automatically equivalent to a specific error rate; it signals that human assessment and task-specific validation still matter.

Across these examples, technical checks address only part of the problem. People also interpret what was requested, judge consequences and accept responsibility. The report’s central framing is that verification and adjudication remain scarce even when first drafts or candidate answers become easier to produce.

“Verification abundance, adjudication scarcity.”

— ThorstenMeyerAI.com report

Amazon

formal proof verification software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How Much Review Is Actually Needed

The available figures do not measure one common outcome. The mathematics output, software pull-request studies and contract-model evaluation involve different tasks, time periods and methods. The supplied material does not include enough detail to independently compare them or establish that AI adoption caused each reported change. The report also notes that some software-data sources sell review tools, a commercial interest readers should consider when interpreting their findings.

It remains unclear how many of the 722 manuscripts were correct, useful or independently assessed, and how much expert time a representative manuscript requires. The supplied information does not establish whether the cited software review patterns hold across industries or organizations. Nor does it show that junior employees will lose opportunities to develop judgment; that is a concern raised by the source, not a measured result in the cited figures.

There is also no single definition of successful review across the examples. A formally valid proof, accepted code change and contract meeting evaluation criteria are different standards. The scale of any net productivity gain depends on review quality, correction rates, consequences of errors and the time experts spend checking work—details the source material does not fully quantify.

Amazon

software review automation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Tracking Review and Training Practices

The next useful evidence will come from studies that track generated work and its review together: how long checking takes, what errors reviewers find, how often changes are revised or rejected, and what happens after work is approved. Comparisons will be more informative when they state the baseline, time window, task type and review standard, rather than treating output volume as a proxy for usable work.

Organizations adopting AI can also make review practices visible: identify who is accountable for approval, distinguish automated checks from human assessment, and record when work bypasses review. For fields that rely on experienced specialists, a parallel question is how junior workers can still practice the underlying skills—not only inspect machine-produced drafts—so that future reviewers have experience to draw on.

The source report offers no announced policy or next milestone from OpenAI, Faros AI, LinearB or Ironclad on this issue. For now, the development is best read as a warning to measure both sides of AI adoption: how quickly systems generate work and how well institutions can verify it.

Amazon

AI manuscript review software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What did OpenAI report producing?

The source report says OpenAI’s model produced 722 mathematical manuscripts across 372 problem families after receiving about 4,000 problems. Some results were formally checked in Lean; OpenAI said some unformalized results could have issues.

Do the figures prove that AI makes review slower?

No single figure proves that for all workplaces. The cited analyses report longer review times or different review outcomes in particular datasets, but they use different methods and samples. The source also notes that some providers have commercial interests in code-review tools.

Can automated tools verify AI-generated work?

They can check specified properties. A proof checker can verify a proof against its stated theorem, and software tests can check behavior covered by those tests. Those checks do not necessarily establish that the question, requirements or tests were appropriate; expert judgment may still be needed.

Why does the report raise concerns about junior workers?

It argues that people often gain the experience needed for senior review by doing the underlying work themselves. The source does not show that AI has already reduced training opportunities, but it raises the possibility that organizations will need deliberate ways to build those skills.

Source: ThorstenMeyerAI.com

EVERGREEN BESTSE

Evergreen bestsellers Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Entertainment signal monitor: Toy Story 5

Toy Story 5 is identified as a fast-moving development in entertainment, prompting early monitoring for industry operators. Details are still emerging.

A Beauty Trend Check On Pluralibacter Gergoviae And Shampoo

A business brief flags Pluralibacter gergoviae and shampoo as a trend signal, but offers no evidence of a contamination incident or recall.

Analog Comeback: Why Tactile, Human-Made Design Is Trending Again

AIThis post was created with the assistance of artificial intelligence (AI).You’re noticing…

Speaker Tracking Sounds Great – Until It Doesn’t

Absolutely impressive, speaker tracking can falter due to various issues, and understanding these can help you troubleshoot effectively.