AI’s Productivity Tradeoff: More Work To Review
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: AI’s Productivity Tradeoff: More Work To Review on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get school and study supplies delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

AI tools are producing more mathematical manuscripts, code changes and legal-work drafts, but reviewing those outputs remains slow and depends on human judgment. The supplied examples suggest a growing review bottleneck, though some figures come from vendors and the scale of the problem is not yet settled.

AI systems are producing more work in mathematics, software and contract workflows, while humans still need to check whether that work is correct and fit for use. Examples in source material from ThorstenMeyerAI.com include 722 mathematical manuscripts produced by an OpenAI programme and software-industry data reporting longer review waits for AI-generated changes. Together, they point to a potential bottleneck: output is scaling faster than the capacity to verify it.

According to the source, OpenAI posed about 4,000 mathematical problems to a model and published 722 manuscripts, grouped into 372 families. The source says some results were formally checked using Lean, a proof-assistant system, while OpenAI cautioned that some results without formal verification could have issues. A counterexample to an earlier Erdős conjecture from the same programme received careful verification from five leading mathematicians, according to the source. The materials do not give publication dates or identify the mathematicians.

Software metrics cited in the source also describe a gap between production and review. Faros AI reported that teams merged 98% more pull requests during high-AI-adoption periods, while review time rose 91%. LinearB, analysing 8.1 million pull requests across 4,800 organisations, reported that AI-generated changes waited 4.6 times longer for review to begin and were accepted 32.7% of the time, compared with 84.4% for human-written changes. The source also cites a peer-reviewed 2026 study finding that 61% of AI-agent pull requests received no human review before being merged or closed.

In contract work, the source describes a partnership between OpenAI and contract-software company Ironclad involving GPT-6 Astra, trained on real contracting workflows. On 11 tasks, the model met an average of 55% of the evaluation criteria, according to the source. That score indicates remaining review work, but the available material does not specify the evaluation method, task weighting or how the system performed on individual tasks.

At a glance
reportWhen: Source material describes developments…
The developmentA set of examples spanning AI-generated mathematics, software and contract work highlights a potential mismatch between faster production and limited human review capacity.
The Referee Shortage — Post-Labor
AI Dispatch · Post-Labor · 7 October 2026

The referee shortage: AI made doing cheap and checking expensive

OpenAI’s model produced a maths result in about three hours of compute. Verifying one earlier result took five of the world’s leading mathematicians. That ratio is the next decade of work: producing is cheap and abundant; trusting is slow, human and scarce.

One pattern, three fields
Mathematics
722
manuscripts, ~3h compute each

Some Lean-checked; OpenAI warns unformalized ones “could have issues.” Verification abundance, adjudication scarcity.

Software
+98% / +91%
more PRs merged / longer review

Faros AI. LinearB (8.1M PRs): AI changes wait 4.6× longer, accepted 32.7% vs 84.4%.

Professional workflows
55%
of criteria met — Astra on Ironclad

Real progress. Someone still has to find the other 45% before the work can be used.

Generation collapsed. Verification didn’t. (conceptual, not to scale)
Cost to produce a resultdown
Cost to check a resultnot down
No author intent

Machine output arrives without reasoning a reviewer can interrogate. It looks locally clean and gives no clue where it’s wrong.

Checks the answer, not the question

A prover confirms the proof proves its statement; tests confirm what tests check. Neither confirms it’s what was needed.

Someone must be accountable

Contracts are signed, designs stamped, papers defended. Responsibility is institutional — you can’t hold a model to it.

Illustrative: $1 of model time + 4 minutes of review at $45/hour. Halving the model price saves 12.5%; one extra review minute erases it. In that example, review is three-quarters of the bill.
What happens when referees run out — already visible
Rubber-stamping
61%

of AI-agent pull requests got no human review at all (EASE 2026). Zero-review merges up 31.3% (Faros).

Triage by suspicion
38%

of reviewers deliberately deprioritise AI changes (LinearB). Good machine work waits behind bad.

Producer as filter
~4,000 → 372

OpenAI chose which maths families were significant. When referees can’t keep up, the producer’s filter becomes the review.

The apprenticeship paradox: reviewers are made by doing the work. The work AI absorbs — writing code, drafting contracts, proving lemmas — is exactly what trained the reviewers. Demand for judgement rises as its supply line shrinks.
What to do
Price verification

Budget review hours next to model spend.

Formalise checks

Provers, types, tests, policy engines.

Tier the review

Experts only where consequences are high.

Fund the referees

Who profits from generation pays for checking.

Protect apprenticeship

Keep some production human for learners.

The take

The first automation question was which jobs AI would do. The better one is which jobs AI makes more necessary: the ones that check, adjudicate and take responsibility. Expect a referee premium — senior engineers, auditors, specialist lawyers, reviewing scientists become the binding constraint on how much AI output anyone can use.Accountability — standing behind a result — may be the most durable form of human work there is.

Sources: OpenAI maths release & Erdős verification as covered here; arXiv:2608.28997; OpenAI × Ironclad (6 Oct 2026); Faros AI; LinearB 2026 (8.1M PRs); Duma et al., EASE 2026 — via secondary reporting. Several code-review sources sell review tools. Review-cost example illustrative. Analysis is the author’s.
thorstenmeyerai.com

Review Capacity Becomes the Constraint

If AI makes drafting and coding faster without making review equally fast, the time saved at the production stage may be partly offset by longer queues, more checking and greater uncertainty about what can safely be used. The cited software figures suggest this can affect delivery speed and quality control, although they do not establish that AI adoption alone caused the changes.

The effect may reach beyond technology teams. Contracts and mathematical results depend on people who can judge whether an output answers the right question, not merely whether it is internally consistent. The source argues that experienced reviewers could become a scarce resource as organisations generate more material for them to assess. That is a plausible interpretation of the examples, not a measured economy-wide finding.

There is also a workforce concern. If junior staff use AI for tasks that once taught them how to write code, draft contracts or develop proofs, they may get fewer opportunities to build the expertise needed for later review roles. The source raises this as a risk; the provided data does not show whether training pathways are already shrinking or how quickly that might happen.

Amazon

AI review management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Three Fields, One Verification Gap

The examples cover different kinds of checking. In mathematics, software such as Lean can confirm that a formal proof follows from stated assumptions. It cannot, on its own, determine whether the theorem is the one researchers needed, whether the result matters or what it contributes. The source characterizes that distinction as a gap between formal verification and human adjudication.

In software, tests can check specified behaviours, but they cannot guarantee that the tests capture every requirement. A reviewer must still assess how a change fits the system and whether important cases have been missed. In contract work, a model may satisfy many evaluation criteria while still missing a jurisdictional requirement or an approval rule, examples raised in the source rather than confirmed incidents.

The source also notes that several cited software-data providers sell code-review tools. That commercial interest does not invalidate their measurements, but it is a reason to read the figures carefully and seek independent replication. The metrics use different samples and methods, so they should not be treated as directly comparable measures of one trend.

Amazon

code review tools for AI-generated code

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How Wide Is the Review Gap?

The available material does not establish that AI adoption caused the reported changes in software review time or acceptance rates. It does not provide the underlying study methods, dates for several datasets, or enough detail to compare organisations with different workloads and review policies. Some cited providers sell review tools, and the figures may reflect differences in how AI-generated changes are identified.

It is also unclear how the mathematics programme selected the 722 manuscripts from the roughly 4,000 problems, how many results have since been independently checked, and what proportion contain errors. For the Ironclad work, the source gives an average score across 11 tasks but not the evaluation criteria or a comparison with human performance. The examples therefore identify a potential pattern, not a definitive measure of the size of a cross-industry problem.

Amazon

formal verification software for mathematics

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Evidence to Watch as Use Expands

The next useful evidence will show whether review delays and unreviewed changes persist as organisations adjust their processes. Independent studies that publish their methods, compare similar teams and track outcomes over time could clarify whether AI use is driving the changes or coinciding with other shifts in workload and staffing.

For mathematics, the key follow-up is independent verification of results that have not been formally checked, alongside clearer reporting on which claims have passed that process. For contract systems, further detail on task-level performance, error types and human review requirements would help readers understand what a 55% average evaluation score means in practice. The source material does not identify a specific next release, review deadline or scheduled study.

Amazon

contract review automation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is the main development?

Examples cited by ThorstenMeyerAI.com show AI producing substantial volumes of mathematical and software work, while human review remains necessary and may be lagging behind output. The examples suggest a review bottleneck but do not establish its scale across all industries.

Were all 722 mathematical manuscripts verified?

No. The source says some results were formally checked in Lean and quotes OpenAI warning that unformalized results could have issues. It does not give a count of how many manuscripts have been independently verified.

What did the software figures measure?

Faros AI reported changes in pull-request volume and review time across different AI-adoption periods. LinearB reported review-start delays and acceptance rates for AI-generated versus human-written changes in its analysis of 8.1 million pull requests across 4,800 organisations. Those are separate datasets and should not be combined into one measure.

Does the evidence prove AI makes review slower?

No. The cited figures report associations and comparisons, but the supplied material does not provide enough methodological detail to establish that AI adoption caused the changes. Some cited software-data providers also sell review tools.

Why might human reviewers still be needed?

Automated checks can test whether work meets specified rules, but people may still need to decide whether the rules capture the real requirement, whether the result is appropriate and who is accountable for using it.

Source: ThorstenMeyerAI.com

HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

How AI And Computer Vision Are Changing Food Safety Checks

AI-driven computer vision is now reliably detecting food safety violations in restaurant walk-through photos, improving verification and compliance.

24 Ways To Make Jev Part Of Your AI Decision Workflow

Thorsten Meyer details 24 Jev use cases across publishing, commerce and operations; 3 live, 12 strong fits, with a four-condition test before wiring.

The Future Is Here: 14 AI Automation Tools To Optimize Workflows In 2026

A comprehensive roundup of 14 AI automation tools shaping workflows in 2026, highlighting key features, applications, and future implications.

The European Bet: How Mistral, Aleph Alpha, and Black Forest Labs Are Playing a Different Game

European AI vendors Mistral, Aleph Alpha, and Black Forest Labs are positioning for the EU AI Act’s enforcement, emphasizing compliance and sovereign deployment.