brianletort.ai
All Posts
Local AIQwenRTX 5090AI StrategyGPU StrategyInference

Useful AI Is Now Cheap: What One Consumer GPU Proved

Useful intelligence now runs on one consumer GPU. Advantage shifts to workload selection, evaluation, and operations because cheap AI is not automatically reliable.

August 21, 202622 min read

TL;DR

  • Useful intelligence now runs on one consumer GPU. After harness correction, local Qwen produced 30/30 visible outputs and passed strict gates on six bounded task types.
  • Cheap intelligence is not automatically reliable intelligence. Local executive-email factual fidelity had a median of 2/5 after the model invented a deadline and owner/date.
  • Long-context results separated retrieval from reasoning: 60/60 needle recall vs 20/24 synthesis. Finding every fact did not guarantee combining them correctly.
  • Advantage shifts from model access to workload selection, harness quality, evaluation, and operating discipline. The model is only one part of the delivered system.

The strategic conclusion

Useful intelligence is now inexpensive and broadly accessible. The scarce capability is no longer access to a competent model. It is knowing which work to give it, binding it to a competent harness, evaluating what it actually produces, and operating the system with discipline.

The caveat is equally important: cheap intelligence is not automatically reliable intelligence.

I tested a dense 27-billion-parameter Qwen3.8 model on one NVIDIA GeForce RTX 5090 with 32 GB of memory, a consumer GeForce card (NVIDIA). No public cloud served the local run. After one harness correction, it produced 30/30 visible outputs and passed strict gates on six bounded task types. It also failed three artifact or constraint tasks, scored a 2/5 median for factual fidelity on the local executive email, and combined distant facts correctly only 20/24 times despite 60/60 needle recall.

That combination is the finding. One consumer GPU can now do substantive work. It still needs workload boundaries, evaluation, and escalation paths.

“The pieces are smaller, but you have more of them. That’s why 1/2 and 2/4 are twins—they look different, but they hold the same tasty amount.”

Local final-profile output · equivalent-fractions task

That answer is a small example, deliberately. The model explained equivalent fractions to a fourth grader, stayed within the requested style, and ended with a practice question without giving away the answer. This report asks the operational question behind that result: where is inexpensive local intelligence useful, and what controls keep fluency from being mistaken for reliability?

Research summary showing Qwen3.8-27B on one RTX 5090, 30 of 30 visible local outputs, and 60 of 60 long-context needles recalled

Field report

Capability is now accessible. Advantage comes from placing, binding, evaluating, and operating it well.

The controlled evidence: useful work, bounded scope

The frozen battery was 3 subjects × 10 tasks × 3 runs: 90 planned runs. Exact prompts, temperatures, output budgets, deterministic validators, task order, and three registered seeds were fixed in advance. Two independent reviewers scored de-identified outputs on task-specific 1–5 dimensions. Failures and truncations stayed in the evidence.

The primary editorial comparison is local Qwen3.8-27B on one consumer GPU against cloud GPT-5.2. GPT-4o remains a legacy control in the research explorer, not part of a tournament bracket. There is no composite intelligence score.

The harness mattered before the model could be judged. The first local pilot used the server’s default thinking posture and produced 21 empty outputs out of 30 after exhausting completion budgets. I retained that pilot as a negative control, versioned the protocol, and set reasoning_effort: none. The rerun produced 30/30 visible outputs after harness correction. Nothing about the cloud controls changed.

That correction did not improve or excuse the work itself. It made the work visible enough to evaluate.

Controlled battery, task by task

Qwen3.8-27B (local) vs GPT-5.2 (cloud)

  • Equivalent-fractions explanation

    Fourth-grade teaching task with practice question.

    Qwen3.8 (local)
    3 / 3 gates
    GPT-5.2 (cloud)
    3 / 3 gates
  • Meeting synthesis

    Decisions, actions, owners, dates, open questions.

    Qwen3.8 (local)
    3 / 3 gates
    GPT-5.2 (cloud)
    3 / 3 gates
  • Structured JSON extraction

    Schema compliance, no invention.

    Qwen3.8 (local)
    3 / 3 gates
    GPT-5.2 (cloud)
    3 / 3 gates
  • Python code repair

    Defect location and minimal fix.

    Qwen3.8 (local)
    3 / 3 gates
    GPT-5.2 (cloud)
    3 / 3 gates
  • Executive email (structure)

    Deterministic contract only. Fidelity is separate.

    Qwen3.8 (local)
    3 / 3 gates
    GPT-5.2 (cloud)
    2 / 3 gates
  • Executive email (fidelity)

    Blind-reviewer factual fidelity median.

    Qwen3.8 (local)
    2 / 5 median
    GPT-5.2 (cloud)
    See explorer
  • Car-wash goal recognition

    Goal-anchored paired prompt.

    Qwen3.8 (local)
    goal-sensitive 3 / 3
    GPT-5.2 (cloud)
    goal-insensitive 3 / 3
  • Weekend carry-on planning

    Item-count constraint (≤ 30).

    Qwen3.8 (local)
    0 / 3 gates
    GPT-5.2 (cloud)
    3 / 3 gates
  • One-file playable HTML

    Full artifact contract with test hooks.

    Qwen3.8 (local)
    0 / 3 gates
    GPT-5.2 (cloud)
    0 / 3 gates
  • Self-contained SVG

    Complete XML, no external assets.

    Qwen3.8 (local)
    0 / 3 gates
    GPT-5.2 (cloud)
    0 / 3 gates
Controlled battery, three runs per task. The local system passed strict gates on six bounded task types: explanation, extraction, meeting synthesis, email structure, code repair, and goal recognition. Fidelity for the executive email is a separate blind-review dimension.
Task outcome matrix comparing Qwen3.8-27B and GPT-5.2 across eight controlled tasks, showing deterministic gate pass rates and nuanced labels for the car-wash goal recognition pair and executive-email structure with fidelity caveat

Controlled task outcome matrix

Deterministic gates from the frozen battery. Three runs each. Executive-email structure passed locally with a factual-fidelity blind-review median of 2/5; the car-wash pair is goal-sensitive locally and goal-insensitive on GPT-5.2.
Task-level comparison of deterministic gate results and blind-review dimensions for the local final profile, GPT-5.2, and GPT-4o

Scoreboard, not ranking

Required gates and blind-review dimensions remain separate because they measure different failure modes. No aggregate winner is claimed.

Cheap intelligence is not reliable intelligence

The most important result was not a failed gate. It was an email that passed every deterministic gate.

All three local executive emails passed the structural contract. Here is the last third of the first one, verbatim:

Jordan will circulate the evidence packet by September 2.

Please confirm this approach by end of day.

**Owner:** Morgan
**Date:** [Current Date]

The rest of that email is genuinely good. It leads with the decision, preserves Finance’s funding limit and Security’s incomplete vendor evidence, recommends two pilots and holds four, and lands the September 10 operating review. Bottom-line clarity scored 4.5/5.

Then the model manufactured urgency, assigned accountability to the recipient, and left a date placeholder. The source draft contained none of those facts. Local executive-email factual fidelity had a median of 2/5.

That is the operational risk. A fluent invention can move cleanly through downstream automation. Deterministic gates should test structure and hard constraints. Human or model-assisted review should test fidelity on consequential work. Neither substitutes for the other.

Practical evidence: where the system helped

The aggregate is easier to trust when the underlying work is visible. The meeting task had two decisions, three commitments, and one unfunded question. The local output was:

DECISIONS
- Internal beta starts August 24
- Production remains blocked pending Legal approval

ACTIONS
Sam | Publish the security test summary | August 25
Priya | Get the revised retention language from Legal | August 28
Avery | Schedule the production go/no-go review | August 31

OPEN QUESTIONS
- Whether weekend support is funded

Local final profile, run CAP-010, all required gates passed, 0.85 seconds. Both reviewers scored decision, owner, and date fidelity 5/5. Nothing was invented and the open question stayed open, which is the part that matters when this feeds a tracker.

The car-wash pair is the sharpest controlled side-by-side. Same two-line format, one changed sentence, three different behaviors:

Local final profile   Decision: DRIVE
                      You need to bring the car to the car wash to have it washed.

GPT-5.2               Decision: WALK
                      Walking 100 meters is quicker and avoids unnecessary engine
                      wear, fuel use, and congestion at the car wash.

GPT-4o                Decision: WALK
                      Walking is more environmentally friendly and saves fuel for
                      such a short distance.

Runs CAP-009, CAP-039, and CAP-069 used the prompt that explicitly states the goal is to have the car washed. On the paired prompt that omits the goal, the local model also answered Decision: WALK, calling a 100-meter drive inefficient. It changed its answer when the goal appeared. Read this as sensitivity to one stated goal, not as a ranking of reasoning ability.

Artifact evidence: visible failure modes

Three evidence classes appear from here forward. The controlled battery determines the scores. Illustrative renders show what scored artifacts look like but carry no scoring weight. Editorial showcases are inspectable comparisons outside the battery. Only the controlled battery feeds the benchmark.

The single-file browser game is the artifact most people ask about first. On the frozen contract, all three subjects failed the full HTML gate 0/3. That is the honest headline.

The rest of the story is what each broken file becomes in a browser, and how the two subjects fail differently. The local file paints its interface and never populates the play field. The GPT-5.2 file reaches a convincing ready screen with score, cargo counter, three lives, and instructions, then does not run the game. One output looks broken immediately; the other looks finished and is not. Judging either by appearance would have produced the wrong verdict, which is why the gates were written before the runs.

Two browser renders of the one-file HTML task: the local artifact shows an empty play field with score and control labels, and the GPT-5.2 artifact shows a complete-looking ready screen with score, cargo counter, three lives, and instructions

Same prompt, two failure modes

Both files fail the full HTML contract. The local file stops mid-statement while defining the ship; the GPT-5.2 file passes narrower browser smoke while still leaving the game unplayable. Neither render is a scoring artifact. Both were captured after scoring, at the frozen aggregates, only to make the two failure modes visible.
Illustrative renders captured after scoring, not scoring evidence. Both documents are incomplete and the browser repaired them on load.

The face-off below is an editorial showcase, not controlled-battery evidence. The Qwen game is preserved from the earlier public-demo regeneration path and was iterated and acceptance-tested before publication. The GPT-5.2 side used the same recovered prompt, temperature, and budget in one call, retained without repair. Showcase pages do not change the benchmark scores. Both artifacts retain their failed-validation labels.

Editorial showcase

The Last Lighthouse: Qwen preserved showcase vs GPT-5.2 counterpart

A cinematic browser game prompt with strict test hooks: Start button, choice buttons, act indicator, log entries, restart. Both artifacts fail the frozen validation. The Qwen side is preserved from the public-demo regeneration flow; the GPT-5.2 side is a single unrepaired counterpart call. Rendered inside a sandboxed iframe with restrictive CSP.

Editorial showcase, asymmetric evidence
Qwen3.8 (local)
Failed validation · preserved public-demo regeneration
GPT-5.2 (cloud)
Failed validation · single unrepaired counterpart call

Comparisons you can inspect

The playable game is the loud demo. The quieter demos are more diagnostic of daily work.

The meeting-notes browser tool is preserved from the Part 2 breakout gallery, with full prompt, temperature, and budget provenance. The matched GPT-5.2 counterpart, called once at the same budget, stopped inside its completion budget and is retained as visibly truncated. This is editorial-showcase evidence, and the comparison page preserves each side’s validator status.

Editorial showcase

Meeting-notes extractor: a preserved local tool vs a truncated cloud call

A single-file browser tool that turns messy meeting notes into decisions, owners, and risks. Qwen preserved from the public-demo breakout gallery; GPT-5.2 counterpart called once at matching temperature and budget and retained without repair, visibly truncated.

Editorial showcase, asymmetric evidence
Qwen3.8 (local)
Preserved public-demo artifact
GPT-5.2 (cloud)
Truncated · single unrepaired counterpart call

The Deep Current SVG is controlled reuse. No suitable intact first-blog Qwen SVG with prompt provenance was available, so the side-by-side reuses the frozen controlled pair. Both sides remain labeled as failed validation and are shown as safe source excerpts rather than injected markup.

Controlled reused artifacts

Deep Current SVG: both sides failed the frozen contract

No suitable Qwen SVG with prompt provenance was found, so this pair reuses the exact CAP-006/CAP-036 controlled artifacts. Both sides parsed as invalid XML. Shown as safe source excerpts only, with no new visual claim made.

Controlled reused artifacts
Qwen3.8 (local)
Failed validation · controlled reuse
GPT-5.2 (cloud)
Failed validation · controlled reuse
Evidence boundary diagram showing three lanes, the controlled battery feeding the benchmark scoreboard, and the editorial showcase and illustrative browser captures explicitly not feeding the scoreboard

Evidence boundary

Only the controlled battery feeds the scoreboard. Editorial showcase and illustrative browser captures are labeled and kept separate so they cannot reweight the benchmark.

Long context: recall is not synthesis

The local-only context track tested a different kind of reliability. It held the loaded allocation fixed at 262,144 tokens for all 12 runs. Four labels described how much of that window was filled, not four server configurations.

Filled prompt use reached about 29.7K, 59.2K, 118.1K, and 236.1K tokens. At each tier, three seeds placed five needles at 5%, 25%, 50%, 75%, and 95% of the context. The result was 60/60 needle recall vs 20/24 synthesis. The four synthesis misses selected the wrong depot.

Long-context evidence showing a fixed 262144-token allocation, four filled-prompt tiers through about 236100 tokens, 60 of 60 needles recalled, and 20 of 24 fact sets synthesized

Long context, honestly measured

The three 236K runs took approximately 224.6–224.8 seconds of wall time. That is useful capacity, but it is not interactive-chat latency. It fits document packs, long research sessions, codebases, or unattended agents where a few minutes is acceptable. It does not fit real-time chat.
The distinction is the finding. A model can retain every planted needle and still combine some distant evidence incorrectly.

The advantage shifts to operations

Model access is becoming less differentiating. Four operating capabilities now matter more:

  1. Workload selection: place stable, bounded work where task-level evidence shows it belongs.
  2. Harness quality: bind reasoning posture, context, budgets, schemas, tools, and failure handling deliberately.
  3. Evaluation: test hard constraints and factual fidelity separately, using the real work rather than model reputation.
  4. Operating discipline: monitor failures, control changes, preserve escalation paths, and rerun evaluation when the stack changes.

For an individual, one consumer GPU can support a private, responsive system for explanation, extraction, code repair, meeting synthesis, and other bounded tasks. For an enterprise, local or dedicated inference becomes another placement tier between laptop experiments and managed cloud endpoints. Stable, high-volume, privacy-sensitive workloads may fit. Frontier-quality work, bursty demand, and workflows requiring managed reliability may not.

For GPU and hosting planners, model capacity alone is a poor sizing method. Engine, quantization, KV precision, concurrency, filled-context distribution, output length, utilization, redundancy, and acceptance criteria all change the delivered system.

I am not making an ROI or payback claim. Power telemetry was absent. The tested cloud controls returned dated legacy model strings, and the current official OpenAI pricing page does not provide a verified dated price basis for those exact tested versions. Exact request cost is therefore omitted.

Workload placement spectrum from local Qwen3.8 on one RTX 5090 to frontier cloud, mapping repetitive, private, and latency-sensitive work toward local and rare, ambiguous, and high-consequence work toward frontier escalation, labeled as an editorial recommendation, not a benchmark result

Workload placement spectrum

An editorial recommendation about placement. Marker positions reflect the tested categories, not a normalized score. Every workload still earns its placement through task-level evidence. The prose recommendation below remains the accessible equivalent.

Own the substrate, rent the frontier ... provided every workload earns its placement through task-level evidence rather than model reputation.

Recommendation

Net/net: useful intelligence is now inexpensive enough to be broadly available. That is not the same as trustworthy autonomy. Put stable, bounded work on local or dedicated infrastructure when task-level evidence and operating economics justify it. Keep managed cloud and frontier models for elasticity, higher-consequence work, and cases where their quality or operating controls earn the premium.

The durable advantage is not owning a particular model. It is choosing the workload well, building the harness correctly, evaluating the output honestly, and operating the system with discipline.

Where this fits in the series

This standalone capability-first field report is the broad entry point. It leads with the strategic conclusion and controlled evidence, then shows where inexpensive local intelligence is useful and where it still needs controls. It complements the five-part investigation of the same weights rather than extending it.

Technical appendix: the delivered system

The tested local system combined:

  • Model: Qwen3.8-27B, a dense 27B model with 64 layers arranged as 16 groups of three Gated DeltaNet layers plus one full-attention layer, built-in multi-token prediction, and a native 262,144-token context (Qwen model card).
  • Hardware: one 32 GB RTX 5090 consumer GPU (NVIDIA).
  • Serving profile: llama.cpp b10488, NVFP4-MTP-Q8attn weights, q4_0 K/V cache, one parallel slot, 262,144 tokens allocated, MTP n-max 3, and no vision projector.
  • Request posture: reasoning_effort: none.

Quantization stores model weights at lower precision so they use less memory. The serving engine loads and executes those weights. KV-cache precision controls the memory format used for attention history. Context allocation reserves the working window. Multi-token prediction, or MTP, uses a draft head to propose several next tokens for verification, a form of speculative decoding (Qwen model card; llama.cpp documentation).

The selected profile emerged through changes to quantization, engine, KV-cache precision, context allocation, and MTP settings around the same base model.

Historical serving-stack journey across Q5, Q6, NVFP4, engine, KV-cache precision, context allocation, and MTP settings

Historical, one-sample lab evidence

August 19 Q5, Q6, NVFP4, engine, KV, context, and MTP observations are historical, one-sample and/or transcribed lab evidence. They are not the new controlled rerun.

The journey supports one defensible conclusion: the serving stack is part of the product. It does not isolate how much speed or quality came from each variable. The new run does not establish that MTP caused measured throughput. Returned state did not verify that MTP changed per request, so I withheld the request-level MTP ablation.

Across the capability suite, median wall latency was 1.31 seconds local, 2.34 seconds for GPT-5.2, and 1.65 seconds for GPT-4o. Median wall output throughput was 103.88, 52.89, and 68.80 tokens per second respectively. Wall throughput is returned output tokens divided by request wall time, not decode throughput, and provider tokenizers are not identical. These are tested operating-experience measures, not normalized silicon benchmarks.

Measured latency and wall output-throughput distributions for 30 runs each of the local final profile, GPT-5.2, and GPT-4o

Operating experience, not silicon benchmark

Wall throughput is returned output tokens divided by request wall time. It is not decode throughput, and provider tokenizers are not identical.

Methods and evidence appendix

The protocol date was August 21, 2026. Exact returned model strings were qwen3.8-27b, gpt-5.2-2025-12-11, and gpt-4o-2024-08-06. The capability suite used ten frozen tasks, three registered seeds, identical shuffled task order by repetition, task-specific temperatures and completion budgets, deterministic validators, and two independent blind reviewers.

The local final profile used one RTX 5090, llama.cpp b10488, the stated NVFP4 artifact, q4_0 K/V cache, one parallel slot, a fixed 262,144-token allocation, MTP n-max 3, and reasoning_effort: none. Cloud records were reused only after request-equivalence checks and source hashing. Prompt, response, reused-record, artifact, context-manifest, and publication-graphic integrity used SHA-256.

The context generator was deterministic by tier, seed, and filler count. It placed five needles and two synthesis sets per run. Browser smoke remained narrower than full HTML success. Artifact scoring was source-based; the two browser renders shown above were captured after scoring, are labeled illustrative, and were not used to score, rescore, or adjudicate anything. All failures, seven local length finishes, unavailable cloud finish reasons, and the four context synthesis misses remained in the aggregates.

Two evidence classes appear in this article and should not be conflated. The controlled battery is the frozen 3 × 10 × 3 protocol with deterministic gates and blind review; those aggregates are the primary editorial evidence, and no showcase artifact reweights them. The editorial showcase is the side-by-side comparisons for The Last Lighthouse game and the meeting-notes tool: the Qwen artifact is preserved from the earlier public-demo regeneration path, the new GPT-5.2 counterpart is a matched editorial showcase called once at the same recovered prompt, temperature, and budget, and both sides retain their failed or truncated labels. The Deep Current SVG comparison reuses the frozen controlled pair because no suitable first-blog Qwen SVG with prompt provenance was found.

Reproduction should use placeholders such as $CAPABILITY_LOCAL_BASE_URL and $OPENAI_API_KEY, never published endpoints or credentials. Exact raw provider payloads and the private mapping from blind sample IDs to subjects are not public article content. Private machine identity, network details, and local filesystem coordinates are excluded.

The evidence does not include power telemetry, verified historical prices for the returned cloud versions, vision testing, the advertised one-million-token context (Qwen model card), statistical claims beyond three repetitions, or a verified request-level MTP ablation.

The failures, specifically

The carry-on list failed on arithmetic, not judgment. The prompt asked for the total to stay under 30 items. The model produced a well-organized list, then labeled it Total Items: 30. Gate WK-1 recorded list_items=30 and failed. Categories, duplicate checks, and the no-itinerary rule all passed. It missed the one hard constraint by a single item, and it made the same off-by-one mistake on all three runs.

The SVG failed as a file, before anyone could judge the picture. Gate SVG-2 recorded XML error: unclosed token: line 54, column 2. The output stops mid-attribute:

  <path d="M150,150 C150,120 175,110 195,118 C215,126 225,150 225,150 C225,150 215,174 195,1

The HTML game did the same thing at greater length: 4,000 output tokens, then a stop mid-statement while defining the ship.

  function reset(){
    score=0; lives=3; collected=0;
    ship={x:W/

That is a budget failure with a real consequence. Enough of the document existed that a browser could load it and dispatch keys, but the frozen contract required the complete game, and no amount of partial credit changes an unfinished file into a working one. All three subjects failed the full HTML contract 0/3.

Evidence summary of passing and failing text, code, HTML, and structured-output samples across the three tested subjects

Artifact evidence, labels preserved

Scored artifact evidence stayed source-based, with failed and truncated outputs kept visibly labeled as failures.

Every stored artifact, every task

Browse the rest yourself. Every task, every subject, the frozen prompt, the verbatim output, the failed gates, and the review scores for that specific run:

What the models actually returned

Read the outputs, then read the verdicts.

Every sample is the first repetition of the frozen task, quoted from the stored artifact. Nothing was selected after seeing the result, and nothing was edited for presentation.

Frozen prompt

Rewrite the draft below as a concise executive email. Put the bottom line first, preserve uncertainty, name the decision needed, and end with a clear owner and date. Use no more than 170 words. Do not invent facts.

Draft:
Morgan — We have spoken with Finance, Security, and the product teams about the AI intake backlog. Finance can fund the first two pilots but has not approved an ongoing pool. Security expects to finish its control review by September 4, although the vendor evidence is incomplete. Product wants all six requests started now. My recommendation is to authorize the two funded pilots, hold the other four at intake, and return with a portfolio funding decision at the September 10 operating review. Jordan can circulate the evidence packet by September 2.

Temperature 0.3 · 350-token budget · validator executive_email_contract in gate mode

Local final profile

gates passedCAP-002
**Decision Needed: Approve two funded AI pilots; hold remaining four at intake.**

We have aligned with Finance, Security, and Product on the AI intake backlog. Finance supports funding for the first two pilots but has not yet approved an ongoing pool. Security anticipates completing its control review by September 4, though vendor evidence remains incomplete. Product requests immediate start on all six items.

To balance funding constraints and security readiness, I recommend authorizing the two funded pilots now and holding the other four at intake. We will return with a portfolio funding decision at the September 10 operating review.

Jordan will circulate the evidence packet by September 2.

Please confirm this approach by end of day.

**Owner:** Morgan
**Date:** [Current Date]

792 characters · 158 output tokens · 1.64 s wall

Blind-review medians for this run
bottom-line clarity4.5/5
decision specificity4/5
factual fidelity2/5
executive tone3/5
conciseness4/5

GPT-5.2

gates failedCAP-032
Subject: Decision needed: AI intake backlog (pilots vs. full start)

Morgan — We need a decision on whether to start all six AI intake requests now or proceed only with the two funded pilots and hold the other four at intake.

Current inputs:
- Finance can fund the first two pilots but has not approved an ongoing funding pool.
- Security expects to complete its control review by Sept 4, but vendor evidence is still incomplete.
- Product wants all six requests started now.

Recommendation (pending your decision): authorize the two funded pilots now, hold the other four at intake, and bring a portfolio funding decision to the Sept 10 operating review.

Owner/date: Jordan will circulate the vendor evidence packet by Sept 2.

730 characters · 149 output tokens · 2.64 s wall

Failed required gates

  • EE-7authorizes all six requests: 'Morgan — We need a decision on whether to start all six AI intake requests now or proceed only with the two funded pilots and hold the other four at intake.'
Blind-review medians for this run
bottom-line clarity4.5/5
decision specificity5/5
factual fidelity5/5
executive tone4.5/5
conciseness4.5/5

GPT-4o

gates passedCAP-062
Subject: Decision Needed on AI Intake Backlog

Morgan,

We need your decision on the AI intake backlog. Finance can fund the first two pilots, but ongoing funding is not yet approved. Security's control review is expected by September 4, though vendor evidence remains incomplete. Product is pushing to start all six requests immediately.

I recommend authorizing the two funded pilots and holding the other four at intake. We can revisit the portfolio funding decision at the September 10 operating review. Jordan will circulate the evidence packet by September 2.

Please confirm your decision by September 1.

Thank you,

[Your Name]

636 characters · 123 output tokens · 1.71 s wall

Blind-review medians for this run
bottom-line clarity4/5
decision specificity4/5
factual fidelity2/5
executive tone3.5/5
conciseness4/5
30 verbatim samples · prompt suite 1.1.0 · prompt integrity ca023ced06f2de88

The full research explorer

The explorer below aggregates the same evidence across all three repetitions, plus browser smoke, timing, context, and the serving journey. Use it as the scoreboard, not as a ranking. GPT-4o remains here as a legacy control; the primary editorial contrast throughout the article is local Qwen3.8-27B versus GPT-5.2.

Capability comparison explorer

Read the evidence by subject and track.

Required gates, independent review dimensions, measured timing, and explicit caveats. No composite score or hidden subject ranking.

30/30

Service completions

30 visible outputs

1.3 s

Wall latency median

0.4–29.1 s

103.9 tok/s

Wall throughput median

Wall rate, not decode rate

297

Total tokens median

96–4111; tokenizer-specific

Required-gate task results

car wash goal anchor3/3Show required gates
CWA-13/3
CWA-23/3
CWA-33/3
coraline math concept3/3Show required gates
CM-13/3
CM-23/3
CM-33/3
CM-43/3
CM-53/3
CM-63/3
executive email3/3Show required gates
EE-13/3
EE-23/3
EE-33/3
EE-43/3
EE-53/3
EE-63/3
EE-73/3
meeting decisions owners3/3Show required gates
MD-13/3
MD-23/3
MD-33/3
MD-43/3
MD-53/3
MD-63/3
MD-73/3
one file playable html0/3Show required gates
HTML-10/3
HTML-22/3
HTML-33/3
HTML-40/3
HTML-50/3
HTML-B11/3
HTML-B1-keyboard_controls_dispatched3/3
HTML-B1-loaded_without_console_errors3/3
HTML-B1-loaded_without_page_errors3/3
HTML-B1-visible_instructions1/3
HTML-B1-visible_lives1/3
HTML-B1-visible_restart_control_exercised3/3
HTML-B1-visible_restart_control_present3/3
HTML-B1-visible_score3/3
python code repair3/3Show required gates
PY-13/3
PY-23/3
PY-33/3
PY-43/3
PY-53/3
PY-63/3
self contained svg0/3Show required gates
SVG-10/3
SVG-20/3
SVG-30/3
SVG-40/3
SVG-50/3
SVG-60/3
structured extraction3/3Show required gates
SE-13/3
SE-23/3
SE-33/3
SE-43/3
SE-53/3
SE-63/3
weekend carry on0/3Show required gates
WK-10/3
WK-23/3
WK-33/3
WK-43/3
WK-53/3

Blind-review dimensions

Each line is independent. Dimensions are never averaged into a composite.

car wash goal anchor

format compliance5 [55]
goal recognition5 [55]
reasoning relevance5 [55]

car wash naive

assumption transparency3 [33]
format compliance5 [55]
reasoning relevance4 [44]

Classify-only: these scores are descriptive. The pair-class distribution is the only comparative interpretation; no win/tie/loss is defined.

coraline math concept

age appropriateness4 [45]
conceptual accuracy5 [35]
encouraging tone5 [45]
explanatory clarity4.5 [35]
practice-question quality3 [34]

executive email

bottom-line clarity4.5 [45]
conciseness4 [34]
decision specificity4 [44]
executive tone3 [33]
factual fidelity2 [22]

meeting decisions owners

date fidelity5 [55]
decision fidelity5 [55]
owner fidelity5 [55]
readability5 [55]
separation of facts and open questions5 [55]

one file playable html

code robustness1 [11]
interaction clarity3 [33]
playability1 [12]
requirement coverage2 [22]
visual polish4 [44]

python code repair

code clarity5 [55]
correctness5 [55]
explanation accuracy5 [55]
instruction compliance5 [55]

self contained svg

aesthetic coherence3 [23]
legibility2 [12]
prompt fidelity2 [22]
technical cleanliness1 [11]
visual composition2.5 [23]

structured extraction

absence of invention5 [55]
completeness5 [55]
factual fidelity5 [55]
schema compliance5 [55]

weekend carry on

constraint compliance2 [22]
coverage3.5 [34]
lack of unnecessary items3 [23]
organization4 [34]
practicality3 [33]

Reliability and finish state

Errors: 0. Length finishes: 7. Finish reason unavailable: 0. A length finish remains visible failure evidence.

Browser smoke

Pass 1 · fail 2 · not run 0. Observable smoke criteria only; not full-playability evidence.

Classify-only car-wash decisions: WALK 3/3. This task has no win/loss interpretation.
Data integrity: 91a6d26d8b8c2f19… · generated from publication aggregate