brianletort.ai
← Library

GitHub Copilot research experiment

Verified evidenceCreator to judge

AI pair-programmed HTTP server

Inline generation can compress a bounded task but not prove lifecycle productivity.

Software engineering · Remote study

Collections: Human still decides · Embodied work

An editorial scene for GitHub Copilot research experiment contrasts reads the http-server specification and writes javascript without copilot. with offers inline code completions while the developer implements the same server task. in the ai pair-programmed http server workflow.

Executive brief

The operating-model shift, in one view.

Use for bounded acceleration, not universal delivery productivity.

AI value · Completion time

Verified

71.17 minutes treatment; 55.8% faster, P=.0017

Conditioned on completion; narrow task; author affiliations.

Before

Reads the HTTP-server specification and writes JavaScript without Copilot. → Runs tests, debugs, and submits the implementation.

After

Offers inline code completions while the developer implements the same server task. → Integrates suggestions, runs tests, debugs, and submits.

Human boundary

Developer owns code; tests check requirements.

Why it matters

Inline generation can compress a bounded task but not prove lifecycle productivity.

How the work changed

Before

How the work ran before the change.

  1. Step 1 of 2

    Randomized control developer

    Reads the HTTP-server specification and writes JavaScript without Copilot.

    ControlSame instructions and JavaScript familiarity requirements as treatment.

  2. Step 2 of 2

    Developer

    Runs tests, debugs, and submits the implementation.

    ControlCorrectness/completeness test suite; no production deployment.

What changed

Inline generation can compress a bounded task but not prove lifecycle productivity.

Decision rightHuman moves from creator to judge

After

How the same work runs now.

  1. Step 1 of 2

    GitHub Copilot

    Offers inline code completions while the developer implements the same server task.

    ControlSuggestion only; developer may accept, edit, or reject.

  2. Step 2 of 2

    Developer

    Integrates suggestions, runs tests, debugs, and submits.

    ControlDeveloper owns code; automated tests score correctness and completeness.

Process model built from the published workflow evidence for GitHub Copilot research experiment. Every step, actor, and control appears in full below.
Every step, actor, and control

Exception path

Incorrect suggestions are rejected or corrected.

Work removed

  • Some boilerplate typing and lookup

Decision authority

Developer owns code; tests check requirements.

Before

  1. 01

    Randomized control developer

    Reads the HTTP-server specification and writes JavaScript without Copilot.

    Control: Same instructions and JavaScript familiarity requirements as treatment.

  2. 02

    Developer

    Runs tests, debugs, and submits the implementation.

    Control: Correctness/completeness test suite; no production deployment.

After

  1. 01

    GitHub Copilot

    Offers inline code completions while the developer implements the same server task.

    Control: Suggestion only; developer may accept, edit, or reject.

  2. 02

    Developer

    Integrates suggestions, runs tests, debugs, and submits.

    Control: Developer owns code; automated tests score correctness and completeness.

Work that left the path

  • Some boilerplate typing and lookup

Human role before

Control-group developers implemented the HTTP-server specification, ran automated tests, debugged failures, and submitted the code to the study evaluator.

Human role after

Developers accept, edit or reject, debug and submit.

AI roleSuggests code inline during JavaScript implementation.

Outcomes

Completion time

Verified

160.89 minutes control71.17 minutes treatment; 55.8% faster, P=.0017

2022 · 95 randomized; 70 completed

Conditioned on completion; narrow task; author affiliations.

What leaders can reuse

Anti-pattern

Do not generalize to maintained systems.

Questions

  1. 01Where is the operating threshold set and who can override it?
  2. 02What measured result would trigger rollback or retraining?
  3. 03Which residual decisions must remain human-owned?

Portability conditions

  • Automated tests
  • Code review
  • Security scanning

Reputation risk

low

Evidence and authority

What the public record supports.

Current · updated

1 independent, 1 primary; publication outcomes are verified.

Bundle 1.0.0 · reviewed 2026-08-23 · stable ID 4900752c0088f2ed

Related transformations

More in Software engineering

Sources

Read the evidence, freshness, caveat, and version policy.