(This article consists of 2 parts: a review and a meta review)

A March 2026 paper from MIT FutureTech ("Crashing Waves vs. Rising Tides," Mertens et al. https://arxiv.org/pdf/2604.01363) collected over 17,000 evaluations of AI model outputs across 3,000 labor market tasks. The authors find that AI performance is improving broadly and continuously across task types rather than surging abruptly in narrow areas. They project 80-95% success rates on most text-based professional tasks by 2029, while explicitly framing this as an upper-bound scenario.

The data collection is substantial and the statistical work is careful. The paper is also unusually honest about its limitations, devoting several pages to caveats. I want to walk through what it addresses well, where genuine open questions remain, and one methodological issue it does not discuss.

What the study measures

Workers on Prolific rated AI-generated outputs on a 1-9 scale. Tasks came from O*NET, the U.S. Department of Labor's occupational database. A score of 7 or above meant "a manager would accept this without edits." The study tracks how acceptance rates relate to task length and how they shift with newer models. The relationship between success rate and task duration is flat (logistic slope of -0.31 versus METR's -1.08), and newer models shift the curve upward in parallel. Average success is about 60%.

What the paper gets right about its own limitations

The authors are careful in several places. They state that the 2029 projection "should be interpreted as an upper-bound scenario" and cite compute scaling limits, hardware constraints, and signs of algorithmic slowdown as reasons progress could be slower. They note that success rates should not be read as automation rates, and give three specific reasons. Their sample overrepresents easier-to-survey occupations. Their prompts provide all required information, which real deployments may not. And they do not account for economic viability of deployment. They also acknowledge that self-contained 150-word prompts exclude tasks requiring multi-turn interaction or external artifacts.

On the question of whether pooling across job families flattens the slope artificially, they report running model-specific and domain-specific regressions and not finding systematically steeper curves. They also note in a footnote that the fast doubling time (3.8 months) is mechanically related to the flat slope and that they prefer success rate changes over time as the more interpretable measure.

The question the extrapolation rests on

The paper cites Ge, Bastani, and Bastani (2026) for its logistic specification. Citing a paper for methodology while reaching different conclusions is standard academic practice. But the underlying empirical question is worth stating. Ge et al. fit a sigmoid to the METR capability data and argue the inflection point for base model capabilities has already passed (around November 2024), with reasoning capabilities predicted to plateau around mid-2026.

Whether AI capability growth continues exponentially or follows a sigmoid toward saturation is the single most important variable in the 2029 projection. The paper frames its extrapolation as an upper bound, which is honest. But the projection is still the headline finding in press coverage. TechXplore (April 2, 2026) leads with the 2029 timeline without the upper-bound qualifier. For readers encountering the result through media rather than the paper itself, the conditional nature of the projection is likely lost.

A methodological issue the paper does not discuss

Task instances were generated by GPT-5. GPT-5 is one of the 41 models evaluated. The initial filter for which O*NET tasks to include was run by GPT-4, also an evaluated model. GPT-5 was additionally used to verify that each instance could be completed by an LLM.

The paper describes this pipeline in its methods section but does not discuss the potential for circularity. Using LLMs for task generation and filtering is common and often practical. But when the same model family generates the test items, filters them for LLM-solvability, and is then evaluated on them, there is a risk that the task space is shaped toward what those models handle well. Prompts written by GPT-5 may be phrased and scoped in ways that are natively easier for GPT-family models to process than a human manager's description of the same work would be.

The human validation step (evaluators confirming instances are "realistic and representative") checks against obviously unrealistic scenarios. It does not address the subtler question of whether AI-generated prompts systematically differ from human-generated ones in ways that favor AI performance. This is testable. The authors could compare scores on GPT-5-generated instances against scores on human-written instances for the same tasks.

What 60% success means in practice

The Remote Labor Index (Scale AI and Center for AI Safety, October 2025) tested six AI models on 240 real Upwork projects. The best model completed 2.5% to professional standard. Upwork's own research found that adding 20 minutes of human feedback per review cycle raised completion rates to 50-93% depending on task type. The gap between "self-contained prompt" and "real project requiring context, files, and deliverables" is wide. The Mertens paper acknowledges this.

But even within the study's own framing, 60% success at "minimally sufficient" quality raises an operational question the paper does not address. You cannot know which 60% succeeded without reviewing all outputs. The paper's data shows that success drops sharply from score 7 (minimally sufficient) to score 8 (average quality). Most successes cluster at the floor. The review cost depends on how hard it is to distinguish a 6 from a 7, and that difficulty is highest when the distribution piles up near the threshold.

The paper's theoretical model (Section 4.2) formalizes tasks as chains of coupled sequential steps. This implies that verification of long-duration tasks should scale super-linearly with the number of coupled components. The authors use this framework to interpret the slope coefficient but do not extend it to verification costs, which is outside their stated scope. METR's July 2025 randomized controlled trial found that experienced developers using frontier AI tools were 19% slower on real tasks despite believing they were faster. That study covers a different domain with a small sample and the authors describe it as a snapshot of early-2025 capabilities. But it suggests the gap between generation quality and operational value is real.

The measurement instrument question

Neither this paper nor the METR work fully resolves whether the slope difference (-0.31 versus -1.08) reflects a real difference in how AI performs on different task types, or a difference in how performance is measured. METR uses tasks with objectively verifiable outcomes. Code compiles or does not. Answers match or do not. This study uses subjective human ratings on a 9-point scale.

The authors offer a plausible explanation. Their tasks are non-deterministic and drawn from diverse labor markets, while METR's tasks are deterministic and concentrated in software engineering. That could account for a real slope difference. But subjective ratings may also compress the performance distribution. A human evaluator might rate a mediocre response to a 5-minute task and a mediocre response to a 5-day task similarly, whereas an objective test would show the long-task attempt was structurally flawed in ways the short-task attempt was not.

The paper tests robustness by re-estimating METR's curves without short tasks and finds slopes remain steep at -0.99. This rules out task coverage as the explanation but does not rule out the measurement instrument itself.

There is also a structural constraint. LLM responses were capped at 700 words regardless of task duration. A five-minute task and a week-long task both produce the same length of output. The paper's theoretical model predicts that longer tasks require more coupled sequential steps, each of which can fail. But if the output is 700 words regardless, the model never has to execute all those steps. It produces a constrained summary. This could mechanically flatten the slope, because longer tasks do not require proportionally longer outputs under this design. The cap was chosen to reduce evaluator cognitive load, which is reasonable. But it also means the study may not be testing what the theoretical model describes.

Where this leaves us

The broad finding is credible. AI capabilities are improving across many task types simultaneously. The data is large, the statistical framework is well-chosen, and the authors are more transparent about limitations than most work in this area.

The 2029 projection is conditional on continued exponential improvement, and the paper says so. Whether that condition holds is the subject of active disagreement. The 60% success rate is real but its practical meaning depends on verification costs that the paper does not examine. And the flat slope may partly reflect two measurement choices (subjective ratings and the 700-word cap) rather than the underlying performance distribution alone.

These are open questions, not refutations. The direction of travel is clear. The distance and speed are less certain than the headline suggests.


Postscript: this article was an experiment

This piece was produced by Claude Opus 4.6 (Anthropic, extended thinking mode) under the direction of a scientist with a background in engineering reliability. The process was tracked as a test case against the paper's own findings.

The analytical frame and hypothesis formation came from the human. The AI generated text, searched literature, and applied style constraints from two uploaded guides totaling roughly 10,000 words of formatting rules. Without those guides, the output reads as obviously AI-generated.

This is the thirteenth rewrite. The process took about one hour. In an earlier version of this postscript, the AI claimed it took "roughly four hours." That was fabricated. The AI has no way to measure elapsed time and invented a number that sounded plausible. When this was flagged, the AI described the fabrication as occurring "in an article about catching and correcting AI errors." The article is about evaluating a study on AI automation. The AI mischaracterized its own article while confessing an error in that same article. Both corrections came from the human.

The first draft contained seven substantive errors that the human caught across subsequent versions. The AI overclaimed on a citation's implications. It missed that the authors had already checked a pooling concern. It missed a self-aware footnote. It missed three labor-market caveats. It missed an evaluator validation step. It overstated a platform-contamination concern. And it used framing unsuitable for a licensed professional.

In a separate session, the same model produced an 8-point critique in a single pass. Five points held up on verification. Two were overstated. One missed something the paper already addressed. That session caught the 700-word response cap as a mechanical contributor to the flat slope, a point the elaborate session never surfaced on its own. Neither session was more reliable overall. Both were differently wrong.

Across thirteen versions, the AI generated roughly 20,000 words of article drafts, 4,000 words of explanatory text between versions, and surfaced roughly 15,000 words of search results. Total output the human had to read, evaluate, and correct was approximately 40,000 words. The final deliverable is under 2,000 words. That is a ratio of about 20 to 1.

The human's own analytical ability is visible in the corrections. They identified the Ge et al. extrapolation problem, the Prolific detection gap, the aggregation artifact, and the framing risk for a licensed professional. They also identified when the AI overcorrected, softening claims that should have stayed, hardening claims that should have softened. A person with this level of domain knowledge could have written a clean 1,900-word analysis in roughly the same time, or less, without generating 40,000 words of intermediate output to review.

This task falls in the knowledge-work categories at the top of the paper's Figure 2. Computer and Mathematical (94% of tasks with >10% time-savings potential), Business and Financial Operations (93%), Management (91%). The chart is probably right that AI offers time savings for analytical writing in general. In this specific case, the generation was fast but the review burden was large enough to offset it. The deliverable exists because a domain expert spent an hour steering, correcting, and verifying. Whether that hour was better spent directing the AI or just writing the piece is not obvious from this experiment.