brianletort.ai
← All posts

Qwen3.8: Same Weights, Different Product

A five-part investigation into Qwen3.8-27B as a private daily driver... what breaks, what the harness fixes, what still belongs in the cloud, and how native 262K context fits on 32 GB.

New here? Start with Part 0. It explains the “thinking” trap in plain language, shows the harness delta (raw local 2/8 → harnessed 8/8), then moves through the car-wash trap, verified one-shot demos, rent-vs-own logic, and the serving-stack bakeoff.

Flagship field report · standalone

Useful AI Is Now Cheap: What One Consumer GPU Proved

Start with the flagship for the strategic conclusion and controlled evidence. Use Parts 0–4 for the investigation and technical depth behind it.

Read the flagship

Featured graphic

The harness made it usable. The serving stack made it fast.

The serving-stack capstone connects model architecture, native 262K context, and measured real-task throughput on one 32 GB consumer GPU.

Qwen3.8 serving-stack capstone: native 262K context on one 32 GB RTX 5090

Verified demos

Six verified one-shot demos (two regenerated + acceptance-tested), plus the full gallery.