Skip to content

Writing · Field notes from the operating layer

Ideas for operating AI, not merely adopting it.

Essays, field reports, frameworks, and technical guides for the people responsible for making AI work.

Current essay · Models & Engineering

The AI Most People Haven't Met Yet

Chatbots continue and coordinate symbolic information. A different class of model represents how a bounded environment may change, often conditioned on action. Both are useful. Treating them as the same thing is why enterprise AI conversations keep talking past each other.

Choose your path

Start with the work closest to the decisions you make.

Editorial pillars

Six editorial lenses. Each starts from a point of view, not a topic label.

Active series

Multi-part arguments, ordered as published.

Series / qwen38-local-wow

Qwen3.8: Same Weights, Different Product

A five-part investigation into how the harness and serving stack make Qwen3.8-27B useful: the reasoning dial, the harness delta (raw local 2/8 → harnessed 8/8), verified demos, portfolio economics, and native 262K context on 32 GB.

5 published parts5/5 planned
Browse in archive

Series / agent-native-work

Agent-Native Work

A four-part series on designing work for agents: why software got them first, every domain's AGENTS.md, the verification gap, and the practitioner-builder.

4 published parts4/4 planned
Browse in archive

Series / llm-os-modes

Modes of the LLM OS

A six-part series on the operating modes behind frontier AI: Chat, Agent, Deep Research, Cowork, and owned infrastructure.

6 published parts6/6 planned
Browse in archive

Series / token-economy

The Token Economy

A strategy and architecture series on token economics, model portfolios, and AI factory operations.

3 published parts3/3 planned
Browse in archive

Series / context-compilation

Context Compilation

The missing systems layer between retrieval and reasoning, from benchmark blind spots to measured evidence.

3 published parts3/3 planned
Browse in archive

Series / autonomous-stack

The Autonomous Stack

The architecture of intelligent systems, from the data substrate to agent runtimes and prescriptive intelligence.

4 published partsOngoing
Browse in archive

Series / agent-societies

Agent Societies

A field guide to what happens when agents interact at scale, from emergence to competence.

4 published partsOngoing
Browse in archive

Series / ai-native-computer

AI-Native Computer

A technical and operating-model series on what changes when AI becomes the computer, not just another app.

3 published partsOngoing
Browse in archive

Complete index

Archive

70 essays

Filters
Pillar
Audience
Format
Series
  1. Framework · September 10, 2026

    The AI Most People Haven't Met Yet

    Chatbots continue and coordinate symbolic information. A different class of model represents how a bounded environment may change, often conditioned on action. Both are useful. Treating them as the same thing is why enterprise AI conversations keep talking past each other.

  2. Research Note · September 10, 2026

    When a Model Can Rehearse the Future

    Rehearsal is where world models earn their keep — exploring possible futures inside a bounded scene before acting. But visual plausibility, controllability, physical executability, and downstream task utility are separate properties, and current systems succeed at some and struggle at others.

  3. Framework · September 10, 2026

    A Different Model Often Requires a Different Substrate

    Chatbots ride on documents. Making a world model useful in a specific enterprise setting tends to require a different substrate — a linked operational record of observations, conditions, actions, and outcomes, with rights and provenance to match. Document RAG is insufficient for that, not irrelevant.

  4. Research Note · August 21, 2026

    Useful AI Is Now Cheap: What One Consumer GPU Proved

    Useful intelligence now runs on one consumer GPU. Advantage shifts to workload selection, evaluation, and operations because cheap AI is not automatically reliable.

  5. Field Report · August 19, 2026

    262K on 32 GB: The Serving Stack That Changed the Desk

    Qwen3.8-27B can hold its native 262K context window on one RTX 5090. The reason is architectural, and the serving stack matters as much as the weights.

  6. Field Report · August 17, 2026

    Same Weights, Different Product: How the Harness Made Qwen3.8 Useful.

    An independent lab investigation. The same open-weight model, served two ways, produced either an empty afternoon or a daily driver. What changed was the harness, not the weights. A newcomer-friendly entry to a four-part series.

  7. Field Report · August 16, 2026

    The Trap That Walked: Qwen3.8 on a $6K Desk (Part 1)

    A free 27B open model on a home RTX 5090 drove the car-wash trap 5/5 where GPT-5.2 walked. The lesson is not local-beats-cloud. It is that default posture is not capability.

  8. Field Report · August 16, 2026

    The $6K Desk That Works: Near-Frontier Private AI (Part 2)

    An offline 8-task workday, synthetic privacy drills, and six verified one-file browser demos. Same weights, better harness, and the honest economics of owning versus renting.

  9. Framework · August 16, 2026

    What Still Rents: The Portfolio Case for Local AI

    A ~$6K desk changes the default. Here is what still belongs in the cloud, and how to route work between owning and renting without tribalism.

  10. Technical Guide · July 28, 2026

    What Memory Bandwidth Actually Buys You: LLM Inference Hardware in 2026

    A first-principles guide to inference hardware from desk to rack — M5 Max, DGX Spark clusters, prosumer PCs, RTX PRO 6000, Lenovo 8× H200/B200 nodes, AMD Instinct, plus prefill vs decode math, VRAM sizing, and API vs rent vs own TCO.

  11. Framework · July 10, 2026

    Most 'AI Agents' Aren't Agents

    Computer science has had rigorous definitions of 'agent' for 30 years. The industry took about two to break the word — and the cost isn't semantic. Agent-washing misprices risk in both directions: the safe systems get over-governed and the autonomous ones get under-governed. Here is the L0–L5 ladder I use to keep the term honest, and the governance that should follow each level.

  12. Framework · July 9, 2026

    Why Software Engineers Got Agents First

    85.2% versus 10.4%. Same tier of models, five weeks apart. That is not a domain gap — it is a grader gap. Software got agents first because its work came with a free compiler. One law, three multipliers, and the reason every other domain now has to build its own.

  13. Technical Guide · July 9, 2026

    Every Domain Needs Its AGENTS.md

    AGENTS.md is not a document. It is an interface stack — MCP tools, skill catalogs, policy packs, and audit streams — with a README on top. Five translations, one procurement boundary, and the difference between agent-legible and Potemkin agent-ready.

  14. Framework · July 9, 2026

    The Verification Gap

    Verification is not merely the reliability blocker — it is the pricing lever. Where a domain can verify cheaply, it prices on outcomes and tunes smaller models. Where it cannot, it stays hostage to frontier tokens. Margin follows verification.

  15. Framework · July 9, 2026

    The Rise of the Domain Practitioner-Builder

    The forward-deployed engineer is not a phase every domain passes through — it is a fork. Which side your domain walks down is decided by whether it owns its evals fast enough to outrun acquisition. The capstone of the four-part series on designing work for agents.

  16. Research Note · June 1, 2026

    The AI Platform Race Is Moving from Models to Execution

    The enterprise AI race is splitting into two models: integrated work systems that turn intent into completed work, and broad ecosystems that hand you powerful components and the integration bill. Three interactive positioning matrices and a quantitative 'when to use each' tool — across Anthropic, OpenAI, Microsoft, Google, AWS, Salesforce, ServiceNow, IBM, Databricks, and Snowflake.

  17. Technical Guide · May 25, 2026

    Running Your Own LLM OS: The Enterprise Build

    Your CEO asks whether you can build your own. The answer is yes. Here is what that actually means — four modes, four stacks from Frontier API to an 8x B200 chassis on your own silicon, the near-frontier OSS shift that changed the calculus, and the control spectrum that cost analysis keeps missing.

  18. Framework · May 18, 2026

    Cowork Mode: State Is the Coworker

    The difference between a chatbot and a coworker is state. Claude Code, Cursor, Operator, Codex, ChatGPT Projects. Persistent memory, skills, knowledge base, environment access. Session-long state — and the most dangerous un-governed surface in the enterprise today.

  19. Technical Guide · May 11, 2026

    Deep Research Mode: Planner, Swarm, Synthesizer

    Deep research is not a bigger chat. It is three sub-systems pretending to be one — a planner that decomposes the question, a swarm of agents that search in parallel, and a synthesizer that does a long-context reduce. 5 to 15 minutes. Hundreds of thousands of tokens. And the richest audit trail of any mode.

  20. Field Report · May 8, 2026

    I Stopped Using ChatGPT (and 10X'd My Work)

    A field report on how I actually use AI in May 2026 — a journey from Chat (3X) through Cowork (5X) and Build (10X) to Automate (30X), and what it means if you are not technical.

  21. Framework · May 5, 2026

    The Three Postures of AI Work: Chat, Build, Automate

    There are three Level-1 ways humans and AI work together — Chat (Human-to-GenAI), Build (Human-to-Agent), and Automate (Agent-to-Agent + Agent-to-Human). In 2026, chat is table stakes. The advantage lives in Build and Automate.

  22. Technical Guide · May 4, 2026

    Agent Mode: The Loop Is the Machine

    Agents are not a model. They are a loop. One Agent turn equals 5–50 Chat-mode calls, plus tools, plus state, plus a kill switch. Here is what you actually pay for when Cursor writes a PR — and what enterprise governance must cover that Chat-mode governance does not.

  23. Technical Guide · April 27, 2026

    Chat Mode: Single-Shot on Shared Silicon

    One prompt in. One response out. Fourteen infrastructure layers in between. Reasoning models are still Chat Mode — they just rent the GPU for longer. Here is what actually happens, and why it is still one machine.

  24. Framework · April 23, 2026

    From AI-Ready Infrastructure to AI Economics Platform

    Space, power, and cooling was the right product for the last era. It is not the right product for this one. A first-person argument — from inside Digital Realty — about where infrastructure platforms are actually going.

  25. Framework · April 20, 2026

    The Enterprise Token Scorecard

    Six numbers the CFO should read in thirty seconds. The metrics that separate mature AI operators from enthusiastic experimenters — and the trajectory that tells you, every quarter, whether the platform is actually being run.

  26. Framework · April 20, 2026

    Modes of the LLM OS: Why Frontier AI Runs in Four Modes, Not One

    When you hit enter in ChatGPT, Claude, or Cursor, you are not running one machine. You are running one of four operating modes of something that behaves like an operating system. Same GPUs. Five orders of magnitude in cost. Completely different governance surface.

  27. Technical Guide · April 19, 2026

    Designing the AI Control Plane

    Seventeen control planes, zero control. The architecture pattern that turns the CEO's token-economics argument and the Data Gravity placement argument into a single governed operating system for enterprise AI.

  28. Framework · April 19, 2026

    Operating Intelligence at Scale

    The economics of enterprise AI are now driven by routing, compression, caching, and infrastructure control. The AI factory pattern — dedicated GPU environments with federated routing — is becoming core enterprise infrastructure.

  29. Framework · April 18, 2026

    Data Gravity Meets Token Economics

    When 93% of enterprise data is created outside the public cloud, the AI question stops being 'which model' and starts being 'where does inference run'. The executive companion to The CEO's Guide to Token Economics.

  30. Framework · April 17, 2026

    The CEO's Guide to Token Economics

    Why boards should stop asking what AI costs and start asking what a verified outcome costs. A non-technical playbook for the operating discipline that will separate AI leaders from AI spenders.

  31. Framework · April 13, 2026

    What Context Engineering Actually Means

    RAG, MCP, memory systems, fine-tuning, prompt caching, AGENTS.md, knowledge graphs — everyone has a piece of the context puzzle. Nobody has the whole picture. Here's what's missing and why it matters.

  32. Framework · April 12, 2026

    The Enterprise Model Portfolio

    The answer to the token economics problem isn't one model — it's a portfolio of six specialized model types served as internal API services. Near-frontier open models now handle 80–90% of enterprise tasks at a fraction of the cost.

  33. Research Note · April 11, 2026

    The Benchmarks Are Lying to You

    The AI memory space has converged on benchmarks that measure retrieval — the easiest part of the problem. They don't test governance, safety, provenance, or compilation quality. Here's what's missing and why it matters.

  34. Technical Guide · April 11, 2026

    The Missing Layer

    Context Compilation Theory, Context IR, and the architecture between access and reasoning. How measuring benchmark gaps revealed a missing systems layer — and why it changes how we should build AI systems.

  35. Research Note · April 11, 2026

    The Evidence

    Eight metrics measured on a live system. The CRR journey from 48.6% to 100%. CompileBench: the benchmark that evaluates compilation decisions. And the open standard proposal.

  36. Framework · April 5, 2026

    The Stack That Thinks: Putting It All Together

    The Autonomous Stack is four layers: data substrate, agent runtime, proactive intelligence, and human interface. When all four work together, intelligence compounds.

  37. Framework · April 5, 2026

    The Token Bill Nobody's Ready For

    A single power user can generate 10-50 million AI tokens per day. Multiply that across an enterprise, and the math changes everything. Token economics is becoming the defining constraint of enterprise AI.

  38. Framework · March 29, 2026

    From Reactive to Prescriptive: The Proactive Agent Shift

    Today's agents wait to be asked. Tomorrow's will tell you what you're missing. The shift from reactive to prescriptive is where agents become genuinely valuable.

  39. Research Note · March 22, 2026

    The Runtime Wars: Agent Operating Systems Are Here

    Agent runtimes have crossed from frameworks to operating systems. ZeroClaw, OpenFang, and OpenClaw represent three competing philosophies for giving agents a durable lifecycle.

  40. Framework · March 15, 2026

    The Data Layer Nobody's Building

    Vector stores and RAG are table stakes. Real agent intelligence needs a continuous, multi-modal data substrate with episodic, semantic, relational, temporal, and contextual data.

  41. Research Note · February 1, 2026

    The Petri Dish: When Agents Build Societies

    I've been watching agents build a society. The emergent behaviors appearing when large numbers of agents interact without human orchestration point to something bigger than better chatbots.

  42. Technical Guide · January 26, 2026

    SemanticStudio: A Production-Ready Enterprise RAG Agent System

    Open-sourcing the multi-agent chat platform I built to test my AI-native architecture ideas. 28 domain agents, 5 configurable modes, 4-tier memory with Context Graph, GraphRAG-lite, and everything enterprises need to build production AI.

  43. Technical Guide · January 26, 2026

    Domain Agents: Specialization at Scale

    Why SemanticStudio uses specialized domain agents instead of one general-purpose assistant, and how to configure and manage them—from 12 to 50+ agents.

  44. Technical Guide · January 26, 2026

    RAG Chain Configuration: Models, Modes, and Fine-Tuning

    The power user's guide to configuring SemanticStudio's RAG chain—multi-provider LLM support, mode parameters, and full control over cost vs. quality.

  45. Technical Guide · January 26, 2026

    Memory as Infrastructure: The Complete 4-Tier System

    A deep dive into SemanticStudio's 4-tier memory architecture—working context, session memory, long-term memory, and the Context Graph. Progressive compression meets knowledge bridging.

  46. Technical Guide · January 26, 2026

    GraphRAG-lite: Beyond Vector Similarity

    How SemanticStudio's knowledge graph and entity resolution enable relationship discovery that pure vector RAG misses.

  47. Framework · January 3, 2026

    Results as a Service: Why 2026 Is the Year Outcomes Become the Product

    AI agents make outcome delivery feasible. Economic pressure makes it inevitable. Here's what RaaS actually is, where it's already working, and why the shift from 'pay for software' to 'pay for results' changes everything.

  48. Technical Guide · January 3, 2026

    RaaS Architecture: The Control Plane That Makes Outcomes Real

    RaaS isn't a pricing model—it's the commercialization of an execution loop. Here's what Result Contracts look like, how the Outcome Control Loop works, and what providers and consumers need to make outcome-based models real.

  49. Framework · December 22, 2025

    When AI Is the Front End: The Future of Software and SaaS

    If AI is the front end and the LLM is the CPU, what does that do to traditional software? Apps stop being destinations and become capability graphs.

  50. Framework · December 22, 2025

    Architecting the AI-Native Enterprise: A BDAT Playbook

    How should a leading organization design for an AI-native future? Using the BDAT lens—Business, Data, Application, Technology—we explore what's next.

  51. Technical Guide · December 5, 2025

    Context Engineering: Beyond Window Sizes

    How to architect RAG systems that overcome attention dilution and recency bias in large context windows.

  52. Technical Guide · December 1, 2025

    Agentic Architecture: Patterns That Scale

    Design patterns for multi-agent AI systems that actually work in production environments.

  53. Technical Guide · November 26, 2025

    Building RAG Systems at Enterprise Scale

    Lessons learned from implementing retrieval-augmented generation across hundreds of documents and thousands of users.

  54. Framework · November 22, 2025

    Data Products: The Foundation AI Needs

    Why treating data as a product is essential for AI success, and how to build the data infrastructure that makes AI work.

  55. Framework · November 16, 2025

    Teaching Machines, Teaching Humans

    What 5,000+ students and two decades of AI development have taught me about learning—both artificial and human.

  56. Framework · November 14, 2025

    Data Governance in the AI Era

    How traditional data governance practices must evolve to support AI initiatives while maintaining trust and compliance.

Newsletter · The Operating Layer

The dispatch for people accountable for making AI work.

A biweekly note on governed enterprise AI, written from inside the operating problem.

Read The Operating Layer
Writing | Brian Letort