Ten times the output into the same approval queue is not a productivity forecast. It is a bottleneck scenario: generation accelerates, acceptance does not, and the gain becomes work waiting to be checked.
The number comes from my earlier essay on the three postures of AI work, where I wrote that people working in the Build posture could be "3x to 10x more productive on the right tasks." That was an author's estimate for one interaction posture on suitable tasks. Ten times was the governed illustrative upper bound of that range, not a measured enterprise result, and it is not an empirical productivity constant.
Nothing in this essay converts that number into people, capacity, or cost. The evidence does not permit any of those moves. The title uses the upper bound because it makes the mismatch visible: even dramatic task amplification can disappear inside an organization whose acceptance rate did not change.
The part I left out of Superworkers
I wrote Superworkers, Not Replacements as an argument for amplification. The core claim was that AI expands what a person can do rather than making the person irrelevant. I still believe that.
But the essay was incomplete in a consequential way. It stopped at the individual.
I described the faster research, the stronger draft, the better-timed information, and the human judgment that remained. I did not follow the work one step farther. I did not ask what happened when the amplified person handed the result to a reviewer, an approver, a control function, or another team whose operating cadence had not changed. That omission made the argument cleaner than the organization it was supposed to describe.
This is the correction.
A person's task productivity is the rate or quality at which that person completes a bounded task. Organizational productivity is the rate at which useful work clears the full system and becomes an accepted result. The first can improve while the second does not. That is not a paradox. It is what any flow does when one stage accelerates and the next stage stays fixed.
The research is unusually consistent on the first measure, provided we keep the bases attached.
Brynjolfsson, Li and Raymond studied 5,172 customer-support agents in one firm. Access to a generative assistant increased productivity by about 15 percent, measured specifically as issues resolved per hour, with the largest gains among less-experienced and lower-skilled workers. That is the favorable case. Customer support has fast, abundant, machine-readable outcome signals. A case is resolved or it is not; handle time and customer response provide additional checks. Verification is comparatively cheap, which is one reason the gain is clean.
Noy and Zhang studied 453 college-educated professionals completing short, occupation-specific writing tasks. ChatGPT access reduced time taken by about 40 percent and raised output quality, as judged by blinded evaluators, by about 18 percent. Again, the base matters. These were standalone writing tasks, not jobs, workflows, firms, or labor markets. The authors explicitly caution against turning task-level experimental results into general-equilibrium labor conclusions.
Both studies show task productivity. Neither establishes organizational productivity. The distinction is not semantic housekeeping. It is the difference between producing something and getting it accepted into use.
The queue starts with plausible output
Herbert Simon argued in 1971 that an information-rich environment consumes attention, making attention rather than information the scarce resource. Paraphrased into the present problem, a new information-processing subsystem helps only if it absorbs more demand on attention than it creates.
Generative AI plainly absorbs some demand. It retrieves, synthesizes, drafts, classifies, and acts. It also creates a much larger surface of plausible output to inspect. When that output is easy to check, the system clears it cheaply. When it is not, the work has merely moved.
I have already argued that verification is the pricing lever. This is what it does to structure.
That earlier essay derived the economics: different graders produce different cost curves. I will not repeat the pricing derivation here. The organizational consequence is narrower and, in my read, more uncomfortable. If verification remains expensive, the scarce role is no longer the person who can produce the first plausible answer. It is the person who can accept the answer without confusing plausibility for correctness.
The best evidence for the problem is Dell'Acqua and colleagues' published jagged-frontier experiment. The 2026 Organization Science study involved 758 knowledge workers at Boston Consulting Group. On 18 tasks deliberately selected to sit inside the model's capability frontier, people using AI completed 12.2 percent more tasks, completed them 25.1 percent faster, and produced work whose quality was approximately 32 percent higher.
Then the researchers moved one task just outside the frontier.
On that complex managerial problem, the AI-assisted participants were 19 percentage points less likely to reach the correct solution than the control. The underlying rates make the result harder to wave away: 84.5 percent of the control group reached the correct answer, compared with 60 percent and 70.6 percent in the two AI conditions.
Same broad population. Same tool. Adjacent knowledge work. Large gains on one side of a frontier that was not legible from the outside, material degradation on the other.
That last fact changes the operating model. If weak output looked weak, ordinary review would be enough. A junior checker could catch a broken structure, a missing section, or an answer that plainly did not make sense. The costly case is a wrong answer that arrives fluent, complete, and internally coherent. It does not advertise the place where it crossed the frontier.
Inspection asks whether the work looks acceptable. Verification asks whether it is correct. Cheap generation widens the gap between those two activities.
Human review is not a perfect answer either. Skitka, Mosier and Burdick showed that automated decision aids can induce both omission errors, where a person misses what the aid failed to flag, and commission errors, where a person follows an incorrect recommendation despite contradictory information. The point here is not to expand that taxonomy. It is simply that adding a human signature does not make a fallible verification system infallible.
Production rate against acceptance rate
Raising production without raising acceptance produces a queue, not output.
Producing faster without accepting faster converts the gain into a queue. Two panels are built identically, each with three horizontal bars: produced, accepted, and waiting. In the prior state produced and accepted are the same length and waiting is a stub. In the current state produced is much longer, accepted is exactly the same length as in the prior state, and waiting absorbs the entire difference. No axis, scale, or quantity is shown.
Prior state
- Produced
- Accepted
- Waiting
before / after amplification
Current state
- Produced
- Acceptedunchanged
- Waiting
Partial complementarity is still partial
The next Dell'Acqua study is easy to confuse with the first, so the distinction matters. It is a separate 2026 Organization Science article, a separate field experiment, and a separate question.
The Cybernetic Teammate study involved 791 professionals at Procter & Gamble working on new-product-development tasks. Individuals working with AI matched the performance of two-person teams working without it, and AI reduced functional siloing by helping people produce more balanced cross-functional solutions regardless of their professional background.
That is meaningful complementarity. It is also partial.
The published abstract's decomposition says AI primarily improved the quality of generated ideas while human judgment retained value in evaluative selection. I am using that finding carefully. Selecting the best among candidate ideas is narrower than verifying that a plausible output is correct. The study does not measure organizational verification cost. It does corroborate the direction of the asymmetry: the machine's contribution concentrates in generation, while evaluative judgment remains human.
This is where a loose amplification argument gets into trouble. Human and machine can be strongly complementary on one part of a task without being complementary across the entire workflow. Better candidate generation does not automatically produce better acceptance. It can produce more candidates requiring a judgment that did not get easier.
Humlum and Vestergaard's Danish labor evidence points the same way from much farther off, and Part 4 takes it up properly rather than here.
What this essay needs from the asymmetry is narrower, and it can be stated without leaving the task level. Generating a candidate answer is a search over material the model has seen a great deal of. Deciding whether a particular candidate is correct is a comparison against a standard that often exists nowhere in that material: the commitment already made, the exception granted last quarter, the constraint that is real and undocumented, the reason this case is not the case it resembles. Those are held by people, unevenly, and mostly in their heads. So the two halves of the workflow scale differently for a reason that is not going to be fixed by a better prompt. The supply of candidates rose. The standard they have to be judged against did not move, and neither did the number of people who hold it.
Why verification does not delegate cleanly
Part 1 of this series described Luis Garicano's model of hierarchy. Expertise is expensive, so routine problems are handled lower in the organization and only exceptions climb to people with scarcer knowledge. Layers economize on expert attention by filtering the flow.
Verification inverts that mechanism.
If the output can be confidently wrong while looking routine, you cannot identify the exception before checking it. The routine-looking case now requires expert attention precisely because appearance is no longer a reliable filter. The checker has to know enough to reconstruct the reasoning, test the evidence, notice what is missing, and recognize the edge condition the generator crossed.
That means verification cannot be delegated downward as readily as generation. A person with less domain knowledge than the producer may be able to check form, completeness, policy fields, and obvious contradictions. They cannot reliably check a substantive answer whose failure mode is that it looks right to anyone who does not already know why it is wrong.
This is an inference, not a measured law. No study in the evidence base measures how far verification can be delegated down an organizational hierarchy after AI adoption. I think the mechanism is strong because it follows directly from the jagged-frontier result and from what hierarchy is designed to do, but the inferential step should be visible.
It also explains why the new scarce role I called the domain practitioner-builder is structurally awkward. That person writes the domain's unit tests because the knowledge needed to define correct is the same knowledge needed to do the work. Expertise has to sit at the point of production, not wait several layers above it for an exception that nobody can classify until after it has been examined.
This is also why machine-checkable ground truth changes everything. I have made that case in Why Software Engineers Got Agents First, and it does not need another benchmark comparison here. Where the artifact can grade itself, verification can run with generation. Where the correct answer depends on contextual judgment, reconciliation across incomplete evidence, or a person accepting consequence, verification remains scarce attention.
So the queue is not merely a line of work waiting for a slow approver. That framing blames the person and misses the architecture. The queue is what the system produces when it buys generation without buying the ability to accept what was generated.
The strongest case against this argument
The serious objection is that this is transitional. Models will improve, error rates will fall, and verification demand will fall with them. What looks like a structural bottleneck in 2026 may look quaint once systems are reliably better than the people checking them.
That objection is plausible. Any honest version of this argument has to allow for it.
My response is that the cost is imposed by jaggedness, not only by the average error rate. A stronger model can be excellent across a wider field and still have a boundary that is difficult to see from the outside. Better average performance reduces how often verification finds an error. It does not necessarily reduce the expertise required to know when the answer crossed the frontier. In fact, more uniformly plausible output can make the remaining failures harder to detect.
But this is a contested position, not a settled finding. If verification cost falls at a rate comparable to generation cost in domains without machine-checkable ground truth, this essay is wrong. The queue would clear as the model improved, and the organizational consequence I am describing would be temporary.
There is a second objection: many approval queues are ceremony, not verification, and should simply be removed. Also plausible, and often correct. An approval that cannot change the work is not much of a control. But deleting a ceremonial gate and building a reliable acceptance mechanism are different operations. Part 5 will take that distinction directly. For now, the narrow claim is that where consequential work still requires substantive judgment, removing the signature does not remove the need.
A quantity-free example
Consider a specialist intake function. A substantive response once required an experienced practitioner to assemble context, interpret the request, and draft the answer. A generative assistant now produces a credible first draft without consuming the same sustained attention.
More drafts arrive. The senior reviewer who accepts those responses is unchanged, with the same calendar and the same reading speed. The reviewer also knows that plausibility is the failure mode, so faster skimming would defeat the purpose of the review.
The team reports two things at once: the tool is working, and the work is not moving faster. Both are true. They are measurements from different points in the same flow.
Nobody in that example is resisting the technology. Nobody is underperforming. The design accelerated one stage and left the acceptance apparatus alone. The queue is a property of that design.
This example is illustrative and intentionally sanitized. It contains no institutional fact, internal metric, system, or attributed experience.
The Monday test
Take one workflow where a drafting or analysis assistant is already in use. Do not begin with adoption, seats, prompts, or model quality. Measure two intervals separately:
- Time from request to first draft.
- Time from first draft to accepted.
If the first interval fell and the second did not, you bought task throughput that the organization cannot yet ship.
Then ask who could absorb the second interval. Could the work be checked mechanically? Could a policy constraint reject a bad submission before a person sees it? Could evidence be reconciled to source automatically? Where judgment remains, could someone with less domain knowledge than the producer perform it reliably?
If the answer to that last question is no, do not call the result a staffing problem. You have located a queue around scarce expertise. The sequence I would recommend, from operating experience rather than from any study cited here, is to build the checking apparatus before scaling the generating apparatus further.
What comes next
The queue is a symptom. It tells you that production and acceptance moved at different rates. It does not tell you which handoff should be redesigned, which approval is ceremonial, or where the work first began to break.
That requires a different question.
Ask employees where AI should be applied, as many organizations do, and the question sounds participatory while returning a catalogue of tools. The people doing the work hold something more valuable: direct evidence about where it waits, gets redone, gets re-explained, or leaves the formal process to survive.
Part 3 is about asking for that evidence without asking the witness to be the architect. The queue tells you to look downstream. The breakage tells you where to start.
Sources
- Brynjolfsson, Erik, Danielle Li, and Lindsey R. Raymond. "Generative AI at Work." Quarterly Journal of Economics 140, no. 2 (May 2025): 889–942. Publisher.
- Noy, Shakked, and Whitney Zhang. "Experimental Evidence on the Productivity Effects of Generative Artificial Intelligence." Science 381, no. 6654 (2023): 187–192. DOI.
- Dell'Acqua, Fabrizio, et al. "Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of Artificial Intelligence on Knowledge Worker Productivity and Quality." Organization Science 37, no. 2 (2026): 403–423. DOI.
- Dell'Acqua, Fabrizio, et al. "The Cybernetic Teammate: A Field Experiment on Generative AI and Teamwork." Organization Science 37, no. 4 (2026): 1217–1242. DOI.
- Humlum, Anders, and Emilie Vestergaard. "Still Waters, Rapid Currents: Early Labor Market Transformation under Generative AI." NBER Working Paper 33777, May 2025, revised March 2026. DOI.
- Simon, Herbert A. "Designing Organizations for an Information-Rich World." In Computers, Communications, and the Public Interest, edited by Martin Greenberger. Baltimore: Johns Hopkins Press, 1971. Paraphrased without a page locator.
- Garicano, Luis. "Hierarchies and the Organization of Knowledge in Production." Journal of Political Economy 108, no. 5 (2000): 874–904. DOI.
- Skitka, Linda J., Kathleen L. Mosier, and Mark Burdick. "Does Automation Bias Decision-Making?" International Journal of Human-Computer Studies 51, no. 5 (1999): 991–1006. DOI.
Operate. Publish. Teach.
