You did not transform the work. You installed a faster worker inside a job designed for scarcity.
That sentence is a synthesis across the studies below, not a measured rate. The pilots are real. The task gains are real. Those studies do not jointly estimate organization-level throughput, and individually none establishes it either. That outcome was generally not their measured endpoint. The model is not trapped by its context window. It is trapped by your job architecture.
The comfortable explanation is that models are still improving, data is not ready, context is not compiled, interfaces have not caught up. Each is partly true. None is the whole story, and reaching for any one keeps the argument where the executive team thinks it already has an answer.
The harder explanation is that the job architecture you dropped the model into was engineered against a very specific scarcity. Analyst time, manager attention, and reviewer bandwidth were scarce. Every role, span, approval, handoff, and career ladder in the modern enterprise is a priced response to those scarcities. The model changes what several of them cost. The architecture priced against them did not move.
I want to be careful how strong a claim this is, because a version of it is fashionable and wrong. I am not arguing that work design always dominates model quality or data quality. Model capability, training data, interface quality, and workflow design can each bind. The narrower claim is that good models and good data are not by themselves a sufficient operating model, and rollout is not a reliable mechanism for changing how coordinated work gets produced. That claim is inferred across the studies below. Nothing in any single paper measures a "job architecture ceiling." The name is mine.
The ceiling nobody drew on the roadmap
The Job Architecture Ceiling
A generative model accelerates the tasks, but the output hits the unchanged constraints of the job architecture.
A diagram showing a fixed horizontal ceiling representing seven components of job architecture: discretion, accountability, review, promotion, workload, visibility, and value capture. Below it, task capability expands upwards but hits the unchanged constraints.
An enterprise job, in the operational sense, is a bundle: a set of tasks, a discretion boundary, an accountability assignment, a rhythm of review, a promotion ladder, a workload norm, a visibility surface, and a share of what gets produced. When somebody says "job description," they are usually pointing at the first item and hoping the other seven come along. They do not.
A generative model enters that bundle through the first item. It changes what the tasks cost. The other seven are governed by roles, policies, calendars, systems, and ladders that a license agreement does not move on its own. The recovered capacity could become shorter queues, better service, worker time, margin, or simply more work. Job architecture helps shape that allocation; these studies do not predict it.
That is the ceiling, and it is observable in the operational signals executive teams already look at: queue depth, cycle time, rework rate, exception and override frequency, approval latency, downstream completion. When a pilot moves the first item in the bundle and none of those signals move, the ceiling is telling you where the pilot stopped and what the operating model was priced to accept. The three postures (chat, build, automate) describe how a person can use the model; the job architecture describes what they can do with the output.
The evidence, in the units it was measured in
The Evidence is Boundaried
Task gains are real, but broad rollout has not reliably produced organization-level transformation.
Three separate panels detailing distinct studies. Left: Procter & Gamble showing individual vs team performance. Middle: Copilot across 66 firms showing individual time savings. Right: Denmark administrative record showing minimal aggregate effect but task reallocation.
Procter & Gamble
791 professionals
Individuals with AI
Teams without AI
On bounded product-development challenges, AI reproduced some of the benefits collaboration usually supplies.
Copilot
7,137 knowledge workers across 66 firms
(ITT estimate for all assigned workers)
Individual access changed individually adjustable behavior; it did not reorganize the broader task bundle in six months.
Denmark
24,796 linked 2024 respondents
Average earnings/hours effect bound:
Of adopters reporting any time savings:
(non-exclusive response category)
Reallocation reported, but no detectable average effect on earnings or recorded hours.
Three studies are worth reading side by side. None measures a ceiling. Each measures a piece of what a ceiling would look like from the inside.
The Procter and Gamble experiment. Dell'Acqua and colleagues, in Organization Science 37, no. 4 (July–August 2026): 1217–1242, ran a preregistered field experiment with 791 professionals at Procter and Gamble on real new-product-development challenges, randomized into individual or two-person team, with or without AI (DOI 10.1287/orsc.2025.20702). Individuals with AI matched the performance of teams without AI on the studied challenges. The published abstract reports that AI produced more functionally balanced solutions, locates AI's contribution in the generation of ideas, and leaves evaluative selection with the human. On these bounded product-development challenges, AI reproduced some of the benefits collaboration usually supplies. That is not "AI replaced a team," it is a measured shift in what one person can produce.
The 66-firm Copilot study. Dillon, Jaffe, Immorlica and Stanton, in a forthcoming American Economic Review: Insights article (AEA article page), ran a six-month cross-industry randomized field experiment with 7,137 knowledge workers across 66 large firms. The forthcoming abstract reports that, in the second half of the experiment, the roughly 80 percent of treated workers who used the tool spent about two fewer hours per week on email and reduced work outside regular hours. For all assigned workers, the intent-to-treat estimate from the working paper is closer to 1.3 to 1.4 fewer email hours per week, flagged as PROVISIONAL (working paper). Beyond individual time savings, the authors did not detect a shift in the quantity or composition of workers' tasks from individual-level provision. Individual access changed individually adjustable behavior; it did not, in six months, reorganize the broader task bundle.
The Danish administrative record. Humlum and Vestergaard, in the March 2026 revision of NBER Working Paper 33777 titled Still Waters, Rapid Currents (DOI 10.3386/w33777), analyzed two survey rounds across eleven highly exposed occupations in Denmark. The 2024 round produced 25,241 complete responses, of which 24,796 linked to administrative registers; the 2023 round used a different structure, with 18,109 main-survey completions plus a separate 2,561-response follow-up-only questionnaire. Difference-in-differences estimates against matched earnings and hours through December 2024 ruled out average effects larger than about 2 percent, two years after ChatGPT's launch, at either worker or workplace level. Adopters reported saving about 3 percent of their work hours; 84.7 percent of adopters who reported any time savings selected "more time on other tasks," a conditional subgroup on a non-exclusive response category. This is measured; PROVISIONAL working-paper evidence, and the reallocation reading is the authors' own correlational interpretation. Early adoption and self-reported time savings coexisted with reorganization below the level of earnings and recorded hours.
One temptation the executive team should refuse before it forms. A per-task productivity percentage in one occupation, multiplied by seat count in another, is arithmetic. It is not an evidence-based workforce plan. None of the studies above licenses that multiplication: tasks are not seats, hours are not throughput, and the arrangement around the tool shapes where recovered time lands, if anywhere useful. If the diagnostic below finds no landing surface, the arithmetic has already lost its denominator.
"Work design" is a euphemism
"Work Design" is a Euphemism
The six concrete governance decisions that an operating envelope has to answer.
A two-column matrix. Left column lists the six governance decisions (Discretion, Accountability, Surveillance, Workload, Advancement, Who captures the gains). Right column translates them into operational realities without blaming the worker.
- Discretion
- How much of the work the role is allowed to decide without asking.
- Accountability
- Who is answerable when the AI-assisted work is wrong.
- Surveillance
- What the system watches, records, and reports upward.
- Workload
- Whether recovered time is invested in verification and learning, or absorbed as more throughput.
- Advancement
- How people become expert enough to hold the accountable role when the missing rung was automated.
- Who captures the gains
- Whether the surplus lands with the role as time, the customer as service, or the shareholder as margin.
Once the job architecture is on the table as one of the binding constraints, the phrase leaders reach for is "work design," and it is a polite name for six concrete decisions that any operating envelope has to answer.
Discretion. How much of the work the worker is allowed to decide without asking. A model that drafts fluently inside a role with narrow discretion produces drafts nobody has permission to send.
Accountability. Who is answerable when the AI-assisted work is wrong. Consequences do not deduplicate the way coordination does. Redistributing the work without naming the person who answers for it leaves an ownership gap.
Surveillance. What the system watches, records, and reports upward, whether or not the deployer set out to answer that question. Deployed systems can produce telemetry, and telemetry can be routed to the operator as feedback or to the principal as evidence. Those two routings push authority in opposite directions.
Workload. Whether the recovered time is invested in verification, learning, and better decisions, or absorbed as more of the same throughput. If throughput rises and reviewer capacity does not, the queue is the outcome.
Advancement. How people become expert enough to hold the accountable role. Skill formation runs through the tasks that got automated; if the ladder was priced against those tasks, it now has a rung missing.
Who captures the gains. Whether the productivity surplus lands with the worker as time, the customer as service, the shareholder as margin, or the executive team as ambition. The evidence does not predict who captures it. Leaving the allocation unspecified is still a distribution choice.
Each of the six is a governance decision, and none is un-encodable in your stack. Discretion, accountability, surveillance, and routing get expressed in RBAC, workflow engines, policy-as-code, audit logs, and eval gates every day, once someone has decided what they should be. The distinction is between license provisioning (access to the tool) and operating-envelope design, which is what the six decisions actually are. A program can do the first well and still treat the second as change management (a personal observation, not a measured rate). Naming the six on one page, with an owner beside each, is the shortest route out.
Skill formation is part of the operating design
There is a specific risk this series takes seriously, because it does not show up in short-window productivity data. Matthew Beane's comparative ethnography of robotic surgery, in Administrative Science Quarterly 64, no. 1 (March 2019): 87–123 (DOI 10.1177/0001839217751692), documented that the shift from open to robotic technique sharply reduced trainees' legitimate participation in operative work. Approved learning methods became ineffective; a minority developed competence through what Beane called shadow learning: premature specialization, abstract rehearsal, undersupervised struggle. The finding is qualitative and specific to a surgical setting, and does not establish that all AI deployments de-skill workers. What it establishes is a mechanism: a technology can improve how production is performed while silently breaking the pathway through which novices become experts.
Bastani and colleagues, in PNAS 122, no. 26 (June 2025): e2422633122 (DOI 10.1073/pnas.2422633122), ran a classroom randomized field experiment with 839 unique students across 2,848 student-session observations in a Turkish high school. During assisted practice, students using GPT Base and GPT Tutor scored 48 and 127 percent higher than control. On the subsequent unassisted exam, GPT Base students scored 17 percent below control. The safeguarded GPT Tutor arm, using teacher-supplied solutions, common mistakes, and hinting instructions, was statistically indistinguishable from control on the unassisted exam. The result is established for school mathematics and inferred when applied to workplace skill formation. Tool-enabled performance and retained skill are different outcomes, and the arrangement around the tool changes which one you get.
Training architecture is part of job architecture. Treating it as post-deployment change management is how the rung goes missing.
The evidence also shows where AI clearly works
The account above is not a case against AI. It is a case against the version of the strategy that stops at rollout.
Brynjolfsson, Li and Raymond, in the Quarterly Journal of Economics 140, no. 2 (May 2025): 889–942 (DOI 10.1093/qje/qjae044), studied a staggered deployment of a generative assistant to 5,172 customer-support agents at one firm. Access increased productivity, measured as issues resolved per hour, by about 15 percent on average, with larger gains among less experienced and lower-skilled agents. That is a clean task-level gain, consistent with a mechanism where verification is cheap because the ground truth arrives quickly and is machine-readable.
The MASAI trial, reported by Gommers and colleagues in The Lancet 407, no. 10527 (January 31, 2026): 505–514 (DOI 10.1016/S0140-6736(25)02464-X), randomized 105,934 participants in a Swedish population-based screening trial. AI-supported reading triaged low-risk examinations to one radiologist and high-risk examinations to two, meeting the prespecified non-inferiority criterion for interval cancer. The 2025 secondary-outcome paper (DOI 10.1016/S2589-7500(24)00267-X) reported an approximately 44.2 percent reduction in screen-reading workload. MASAI tested a bundled AI-supported reading protocol, not a detached model, and the trial cannot separate the model contribution from the workflow contribution. That is inferred reasoning about the mechanism. What it does establish is that a deliberately redesigned pathway achieved non-inferiority on interval cancer while reducing reader workload.
Read the two next to the first three and the pattern is boundaried but consistent. Where verification is cheap, or the arrangement around the model was deliberately designed, the surplus can land as a measured outcome. Extending that signal to settings where the arrangement was not designed is an inference these two studies do not license.
The strong version of the argument against this piece deserves stating. Technical architecture and bounded workflow redesign can produce real value without touching the enterprise operating model. QJE's customer-support gain arrived through disciplined tooling in a bounded flow. PNAS's safeguarded arm is a concrete bridge: teacher-written guardrails mitigated the unassisted-exam deficit seen in the unguarded GPT Base arm to statistical indistinguishability from control, a technical redesign doing the work. That version is right about the near term; it does not license the leap from a bounded win to an operating-model change.
The executive diagnostic
The version of this argument that does not translate is the one that ends with more strategy work. What this problem needs instead is one honest inspection of one job, at the level of specificity where it can be argued about.
Pick one role that has received AI tooling in the last twelve months. A real role, not an aspirational one. For one week, hold the tooling constant and examine the job. Seven lines fit on a page.
- What is inside the discretion boundary this week, and what remains outside.
- Who answers for the AI-assisted output when it is wrong, and how they were selected.
- What the tooling watches, records, and reports upward, and to whom.
- Whether the recovered time is going into verification, learning, and better decisions, or into more of the same work.
- Which rung of the current advancement ladder the AI has thinned, and what is replacing it.
- Where the productivity surplus is landing, and whether that landing was chosen.
- Baseline the six operational signals (queue depth, cycle time, rework rate, exception and override frequency, approval latency, downstream completion) and remeasure them at week's end. If none moved, name the one you expected to.
Then bring the seven lines back and name what would have to change for the surplus to land somewhere different. That artifact is not a strategy. It is the decision surface an operating-model conversation actually argues from.
A parallel move for the CHRO. Walk one advancement ladder in the same organization and ask, for each rung: which discretion the AI narrowed, which accountability it redistributed, and which pathway novices used to reach the top. If the rung the AI helped is the rung the ladder used to build, the ladder needs a new rung, not a new tool. That is a CHRO decision, and the AI program cannot make it on the CHRO's behalf.
If either exercise returns something you do not want to look at, the diagnostic is working. The rest of the series is written for the reader who reached that point.
Where this goes next
The Organization After Cheap Intelligence
Six parts taking one surface of the operating envelope each.
A navigation list of the six essays in the series, with Part 2 explicitly marked 'Start here if you read one'.
The next six essays take one part of that decision surface each and treat it seriously.
The Org Chart Was Built for Expensive Information reads the structure as a priced response to information that used to be expensive, and argues the price change made it mispriced rather than obsolete. Ten Times the Output Into the Same Approval Queue explains why individual amplification stops at the acceptance step. Don't Ask Employees Where to Put AI. Ask Where the Work Breaks. separates the witness role from the architect role.
AI Is Coming for the Coordination Work, Not Necessarily the Manager unbundles the manager's job and shows why the same capability supports at least three organizational shapes. Fewer Approvals, Stronger Controls treats accountability as an instrument rather than a compliance costume. Start With the Decision. Remove the Handoffs. takes the decision, not the task, as the unit of redesign.
If you only read one, start with Ten Times the Output Into the Same Approval Queue. It converts the ceiling above into the bottleneck most executive teams can feel this quarter.
The job architecture is not the only constraint, and this piece does not claim it always dominates. It is the constraint a rollout-first program is not equipped to see. That is where the series starts.
Sources
- Dell'Acqua, Fabrizio, Edward McFowland III, Ethan Mollick, Hila Lifshitz, Katherine C. Kellogg, Saran Rajendran, Lisa Krayer, François Candelon, and Karim R. Lakhani. "Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of Artificial Intelligence on Knowledge Worker Productivity and Quality." Organization Science 37, no. 2 (March–April 2026): 403–423. https://doi.org/10.1287/orsc.2025.21838. Referenced for the jagged-frontier framing that essay 2 develops.
- Dell'Acqua, Fabrizio, et al. "The Cybernetic Teammate: A Field Experiment on Generative AI and Teamwork." Organization Science 37, no. 4 (July–August 2026): 1217–1242. https://doi.org/10.1287/orsc.2025.20702. 791 professionals at Procter and Gamble (
established). - Dillon, Eleanor W., Sonia Jaffe, Nicole Immorlica, and Christopher T. Stanton. "Shifting Work Patterns with Generative AI." American Economic Review: Insights, forthcoming. AEA article page: https://www.aeaweb.org/articles?id=10.1257%2Faeri.20250275. NBER Working Paper 33795, revised November 2025: https://www.nber.org/system/files/working_papers/w33795/w33795.pdf. 7,137 workers across 66 firms; abstract-level findings
established, detailed estimatesPROVISIONAL. - Humlum, Anders, and Emilie Vestergaard. "Still Waters, Rapid Currents: Early Labor Market Transformation under Generative AI." NBER Working Paper 33777, May 2025, revised March 2026. https://doi.org/10.3386/w33777. Full text: https://www.rfberlin.com/wp-content/uploads/2026/03/26078.pdf. 2024 round: 25,241 complete responses, 24,796 linked to registers. 2023 round: 18,109 main-survey completions plus a separate 2,561-response follow-up-only questionnaire.
Measured; PROVISIONALworking-paper evidence. - Brynjolfsson, Erik, Danielle Li, and Lindsey R. Raymond. "Generative AI at Work." Quarterly Journal of Economics 140, no. 2 (May 2025): 889–942. https://doi.org/10.1093/qje/qjae044. 5,172 customer-support agents (
established). - Beane, Matthew. "Shadow Learning: Building Robotic Surgical Skill When Approved Means Fail." Administrative Science Quarterly 64, no. 1 (March 2019): 87–123. https://doi.org/10.1177/0001839217751692. Qualitative field study (
establishedas an organizational finding;inferredwhen applied to generative-AI settings). - Bastani, Hamsa, Osbert Bastani, Alp Sungu, Haosen Ge, Özge Kabakcı, and Rei Mariman. "Generative AI without guardrails can harm learning: Evidence from high school mathematics." PNAS 122, no. 26 (June 25, 2025): e2422633122. https://doi.org/10.1073/pnas.2422633122. 839 unique students / 2,848 observations (
established); the August 2025 correction adjusted an author affiliation only. - Gommers, Jessie, et al. "Interval cancer, sensitivity, and specificity comparing AI-supported mammography screening with standard double reading without AI in the MASAI study." The Lancet 407, no. 10527 (January 31, 2026): 505–514. https://doi.org/10.1016/S0140-6736(25)02464-X. Full-cohort secondary-outcome paper: Hernström et al., The Lancet Digital Health 7, no. 3 (March 2025): e175–e183, https://doi.org/10.1016/S2589-7500(24)00267-X. 105,934 participants randomized (
established).
Operate. Publish. Teach.
