Cover illustration

TheDaily Front

Issue No. #260824 Monday, August 24 2026 #260824 — MONDAY, AUGUST 24, 2026
The machines get cheaper, the rules get thicker, and the ocean keeps the ledger.
Monday, August 24, 2026 The Daily Front No. #260824 — Contents
30stories
10,071points
6,176comments
268kllm tokens
Assembled with 31 model calls — 173,124 tokens read, 94,953 written.

Highlights

How Europe is killing makers and micro-entrepreneurs

A marketplace operator argues that Europe’s compliance burden is squeezing the garage-scale hardware makers it hopes to protect.

MS Paint and Photos inivisibly watermark even locally generated output with GUID

Reverse engineering finds that local AI images made in familiar Windows apps may carry an invisible server-issued identifier.

Coding expertise is going to collapse from AI reliance

A warning that delegating the hard parts of programming to AI may also erode the path by which engineers acquire judgment.

Executable Is a SQLite Database

A bold systems experiment asks what happens when the executable file is redesigned as a SQLite database.

Oceans hit highest temperature on record

Record ocean temperatures bring the physical world’s largest warning to a front page crowded with digital ambitions.

From the Editor

The day’s papers find invention hemmed in from every side: by regulation, by hidden identifiers, by price wars, and by the old problem of whether convenience leaves us less capable. Meanwhile, the sea has posted a record of its own, which ought to make every other race for scale feel rather less abstract.

  1. How Europe is killing makers and micro-entrepreneurs3
  2. MS Paint and Photos inivisibly watermark even locally generated output with GUID4
  3. Coding expertise is going to collapse from AI reliance5
  4. Executable Is a SQLite Database6
  5. SeL4 security proofs now complete on AArch647
  6. Andreessen Horowitz is investing billions into a bleak future8
  7. Oceans hit highest temperature on record9
  8. FDA clears blood test to aid evaluation for Alzheimer's disease10
  9. OpenAI: GPT 5.6 Sol price reduction (until at least Nov 21)11
  10. I built a low-latency AI companion that plays Skyrim with me12
  11. Jabber/XMPP: 25 Years of Digital Independence13
  12. AI Chip Architectures14
  13. Where did all the public bathrooms go?15
  14. Hot Chips 2026: CUDA Targets RISC-V – By Chester Lam16
  15. Kodak DC50 now usable on the Apple II17
  16. LLMs could control their host machines by exploiting inference engines18
  17. AI and Infrastructure Engineering19
  18. OCR It – pull text out of un-copyable documents for your LLM20
  19. One corner of China’s internet is insisting that the Tang Dynasty never existed21
  20. Xiaomi: New CPU matches Apple cores single threaded, much faster multithreaded22
  21. IPFS Maintainers Winding Down23
  22. New EU-wide product repair rules come into force24
  23. iCloud+ Hide My Email addresses will remain on icloud.com25
  24. Show HN: GlassBox – what the browser reveals, and how identifiable you are26
  25. Show HN: PicoMQ – Durable Streams over HTTP, on object storage27
  26. Woman stranded in Spain after UK's eVisa system mistakes her for twin sister28
  27. Anthropic's best AI model struggles to attract users as cheaper tools thrive29
  28. I were 17, I'd learn how to build LLMs from scratch29
  29. The entire city of San Francisco as a video game29
  30. Show HN: A techno machine in one HTML file, with verifiable renders29
The Daily Front Page 2 of 30
Monday, August 24, 2026 The Daily Front No. #260824 — The Maker’s Burden
article

How Europe is killing makers and micro-entrepreneurs

by l-one-lone·▲ 1,147 points·682 comments·lectronz.com ↗
They are engineers, independent designers, and hardware enthusiasts working from spare rooms, garages, and tiny workshops.

Lectronz is a marketplace for open-source hardware makers and DIY electronics. Most of our sellers are not factories or well-funded start-ups. They are engineers, independent designers, and hardware enthusiasts working from spare rooms, garages, and tiny workshops.

Some earn a living from their products. Some sell only a handful of boards each year. Others build ten units simply because they created something useful and want to share it with the community. Occasionally, one of those experiments grows into a real business. Every Arduino begins somewhere.

But the European Union's new packaging rules now threaten to kill the world of makers and micro-entrepreneurs, putting jobs, livelihoods and an entire ecosystem of innovation at risk.

And this threat is not just limited to makers and engineers. It affects artists, craftspeople and other micro-entrepreneurs selling their work across the EU.

A good idea, a terrible implementation

The EU has required producers to take responsibility for packaging waste for many years through Extended Producer Responsibility (EPR) schemes. The new Packaging and Packaging Waste Regulation (PPWR), which generally applies from 12 August 2026, aims to harmonise packaging rules across the European Union and reduce waste.

The main idea of EPR is sensible: businesses that place packaging on the market should help finance its collection and recycling.

For makers, this means taking responsibility for the boxes, envelopes, plastic bags and other packaging used to deliver their products. This is an idea we can all get behind.

Unfortunately, instead of creating a single European system, the PPWR preserves a fragmented national model. A business selling directly to customers across the EU must register and fulfil its obligations separately in every Member State where its packaging becomes waste. For large companies, this is part of the cost of doing business; for micro-businesses selling only a handful of products into each country, the cost and administrative burden can be wildly disproportionate to the amount of packaging involved.


Imagine an engineer in Greece who designs a €25 open-source sensor board...

During the first year, he sells five to Germany, two to France, two to Austria and one to Belgium. Each ships in a small antistatic bag and a padded envelope. The amount of packaging generated for each sale is probably around 50 grams.

He has just become a packaging waste producer in four countries.

Based on indicative prices currently quoted by national schemes and compliance providers, the annual cost for France alone can look like this:

  • Registering for a packaging scheme, totalling €110 in fees per year.
  • Using the services of an Authorised representative, adding €190 to €300 in costs per year.
  • Spending time registering, documenting, and reporting waste created.

These indicative costs continue to add up for each country:

  • Belgium: €50 to €100 administrative fees per year, plus the services of an authorised representative (approx. €250 to €450).
  • Germany: registration is free, but packaging-scheme participation starts at approximately €10 per year, plus an authorised representative costing around €190 per year.
  • Austria: €250 administrative fees per year, plus the services of an authorised representative (approx. €100).

In short, the barrier to entry for these four countries totals €1150 per year in an optimistic scenario.

The weight-based environmental contribution associated with half a kilogram of packaging should be measured in cents. The bureaucracy required to account for it is measured in thousands of euros.


Now imagine you want to sell to all 27 Member States! To make it worthwhile, our Greek engineer needs to sell not 10 boards, not 100, but literally thousands of boards every year from the very start.

It simply isn’t worth it anymore.

Killing innovation softly

Often, innovation doesn’t come from large established corporations, but from small businesses that start from scratch with new ideas and little money. Before becoming successful and selling millions of products, many companies started selling 10, then 100, then 1000. Most businesses never make it there. But there has to be space where ideas can be tested. This is one of the reasons Lectronz exists.

In the past year, while some sellers on Lectronz sold hundreds of products, half of our registered sellers got fewer than 10 orders. This is not a bug, but the nature of a marketplace like Lectronz where makers are free to experiment with product ideas. Some ideas don’t work. Some creators on Lectronz only build 10 units and share them with the community without making a profit. But even products that “fail” have a value. When hardware creators share them with the community, they help others grow as well. One piece of hardware may unlock the creation of another, leading to new product ideas and innovation.

The EPR regulations threaten the existence of this innovative space in the EU.

EU policymakers keep sounding the alarm about Europe’s lack of innovation, but seem hell-bent on making it as hard as possible for innovation to emerge at all, with regulations that create a disproportionate barrier to entry for micro-enterprises and SMEs. It’s an environment where only big players like Amazon, Temu, or eBay can exist.

Lectronz is also a micro-enterprise

Lectronz collects a 5% fee on every transaction it processes. We waive this fee on the first five sales to encourage sellers to test our platform. After years of work, and with the recent surge of new sellers joining our platform in 2026, Lectronz now generates roughly the equivalent of one modest salary.

I did not build it to become the next Amazon. I built it because independent hardware creators deserve a marketplace designed for them.

If these rules force many of our sellers to withdraw from the European market, they could also make Lectronz itself unviable. After everything we have built together, that would be personally heartbreaking.

For now, Lectronz sellers should not expect any immediate disruption. It remains unclear how national authorities will enforce these rules against makers and micro-enterprises, and we will continue monitoring the situation closely.

What are the solutions?

If these regulations are applied strictly, the short-term solution for makers is simple: stop selling in the EU and ship exclusively to non-EU markets.

Yes, you read that right. For a French micro-entrepreneur, it makes more sense to ship products to the US than to ship to neighbouring Germany or Belgium, for example. This is true even with any US tariffs in place.

Of course, limiting sales to the US is not a viable solution for some sellers. It's also a loss for the European economy itself. I still hope that we can work out realistic solutions that can help restore the EU single market for micro-enterprises. Here are some ideas.

Solution #1: Introduce an EU-wide de minimis threshold.

Exempt small-volume sellers and micro-enterprises from cross-border packaging obligations. The threshold would apply only to producers that are below a specific volume of waste and/or a specific yearly turnover.

Solution #2: Create an EU EPR One Stop Shop.

Create a centralised EU portal where sellers can register, report waste, and pay truly reasonable fees at once, for all Member States where they ship products. This could mimic the mechanism that already exists for VAT with the One Stop Shop (OSS).

Ideally, since we are in 2026, most of this work should be done through a modern open RESTful API (not web forms) and open-source software, to be as automated as possible.

Solution #3: Allow marketplaces to represent and manage micro-enterprises collectively as if it were a single producer.

A mechanism should allow marketplaces like Lectronz or Tindie to register, report waste, and pay reasonable fees on behalf of all their sellers as if they were collectively one producer of waste.

This means that the marketplace would pay administrative fees and other EPR costs corresponding to a single producer, that would collectively represent all its sellers. For Lectronz, this would have a non-trivial impact in terms of cost and administrative work, but it might be achievable under the right conditions.

As stated above, using a common API standard for all countries would help automate things.

Make your voice heard

Again, to reiterate, we support the idea of reducing waste and promoting sustainability. But there’s got to be a better, simpler, and fairer way to do it.

This regulation is having a massive effect on the entire ecosystem of micro-businesses, not just makers. It affects artists who sell their creations online. Local traditional food producers who export their products across the EU. It also affects craftspeople who sell their work online through their own website or dedicated platforms like Etsy. Beyond the small world of makers and DIY electronics, this will have an impact on the livelihood of potentially hundreds of thousands of people in the EU.

And to be clear: these rules affect not only businesses in the EU, but any business that sells to buyers in the EU.

Jeanette Koňarčíková, an independent artist and micro-entrepreneur from Slovakia, launched an online petition to draw the attention of policymakers to this issue:

https://www.change.org/p/stop-destroying-eu-micro-businesses-immediate-moratorium-on-cross-border-epr-fees

The petition is thoughtful and well-written. I encourage you to read and sign it!

The European Commission also has an open public feedback page for this issue here:

https://ec.europa.eu/info/law/better-regulation/have-your-say/initiatives/15352-Packaging-and-packaging-waste-rules-on-national-registers-of-producers_en

Consider leaving feedback there as well.

Recently, the European Commission has begun to recognise part of the problem and has proposed suspending the requirement to appoint an authorised representative in every destination country until 2035. But this proposal has not yet been adopted. Unfortunately, this proposal may take time to be voted on and enter into force. By then, many small businesses may have closed. More importantly, removing the authorised-representative requirement would address only part of the problem. Rules like this risk undermining trust in the European project itself. What’s the point of the EU if the single market no longer exists for micro-enterprises?

Here at Lectronz, we will continue to move forward and hope for the best.

But make your voice heard now to make sure policymakers understand the urgency of this issue!

The Daily Front Page 3 of 30
Monday, August 24, 2026 The Daily Front No. #260824 — The Invisible Mark
article

MS Paint and Photos inivisibly watermark even locally generated output with GUID

by ComputerGuru·▲ 621 points·251 comments·xusheng.dev ↗
The GUID is embedded into the locally generated image as an invisible watermark.

Reverse engineering reveals how Paint and Photos embed a server-issued GUID into the pixels of locally generated AI images.

TL;DR

  • Microsoft Paint supports both local and cloud image generation
  • Paint and Photos also ship local AI models
  • The two apps send the prompt to a remote server for moderation
  • The server returns a GUID along with the moderated prompt
  • The GUID is embedded into the locally generated image as an invisible watermark
  • A separate visible-watermark setting does not control this invisible watermark
  • On Copilot+ PCs, image generation is local but prompt moderation remains remote
  • Microsoft discloses that Paint adds C2PA metadata to AI-generated images
  • AI-generated image saves limited to C2PA-preserving formats: PNG, JPEG, GIF, and .paint

Paint sends the user prompt to Microsoft’s moderation server, receives a moderated prompt and watermark GUID, generates the image locally, and embeds the GUID into the final image pixels

A curious look at Microsoft Paint

This research started with my curiosity about Paint. I recently had some success looking into less-explored Windows features like UCPD, WHESCVC, and I have long known that Microsoft added a bunch of AI features into the Paint app. I do not know if anyone actually uses Paint + AI to generate images, but I wanted to see how exactly the image generation works.

Before I started, I expected that it simply called a remote API to do the image generation. However, after I set up Binary Ninja MCP with Codex and started the analysis, I soon realized that Microsoft actually shipped local models in Windows as part of Copilot.

The Paint App is sitting in the following path (yes, they are all Windows Apps now):

C:\Program Files\WindowsApps\Microsoft.Paint_11.2605.71.0_x64__8wekyb3d8bbwe\PaintApp\

And there are four apparent model files with the .onnxe extension:

seg.onnxe          23.1 MB
inseg_enc.onnxe    28.0 MB
inseg_dec.onnxe    16.5 MB
mager.onnxe        302.4 MB

The format of seg.onnxe was previously known, i.e., when it is XORed with the string Microsoft_2023, it becomes a normal ONNX file. However, the format of the other three .onnxe files initially looked different.

It turned out that Microsoft had not changed the algorithm, only the key. segapi.dll contains a small key registry:

ps_enc_key.1.0.80-main -> "Microsoft_2023"
ps_enc_key.1.0.81-main -> a 4,096-byte alphanumeric string

After decryption, onnx.checker.check_model() works on all of them:

Model Graph seg.onnx 1,094 nodes, input input_image, output output inseg_enc.onnx 1,014 nodes, output image_embeddings inseg_dec.onnx 1,133 nodes, inputs for embeddings, points and masks; output masks mager.onnx 15,284 nodes, image/mask inputs; output output

A visible watermark

While walking through these files, I found a Watermarker.dll:

The properties of Watermarker.dll included with Microsoft Paint

This is not super surprising to me, because while I interacted with the Paint app, I already discovered that it has a setting to embed a visible watermark to the image that it produces:

Paint offers Never, Always, and Ask every time choices for its visible AI watermark

The visible watermark is just a small Copilot logo at the bottom right of the image, which is totally normal.

Then, out of nowhere, I decided to ask AI to analyze the DLL and see if it could also be embedding an invisible watermark. This is part of my intuition as a reverse engineer, because the file is 1.67 MB in size, which is unusually large for such trivial functionality (arguably, the visible watermark does not even require a separate DLL). Apparently, the recent Claude Code text-watermark announcement also played a role in prompting me to think about this possibility.

An invisible watermark

To begin with, the visible watermark is added by AddPerceptibleWatermark:

CPBDoc::Save(...)
  |
  `-- perceptible-watermark save helper(bitmap, WatermarkSetting)
        |
        +-- WatermarkSetting::Never
        |     `-- return the original bitmap
        |
        +-- WatermarkSetting::AskEveryTime
        |     `-- show the Yes / No confirmation popup
        |           +-- No: return the original bitmap
        |           `-- Yes: continue
        |
        `-- Always or confirmed Yes
              +-- Paint::AI::GetPerceptibleWatermarkSvg()
              `-- Paint::AI::AddPerceptibleWatermark(bitmap, SVG stream)
                    `-- composite the visible Copilot logo

Then there is also a different WmkWriteWatermark function:

Watermarker.dll!WmkWriteWatermark(
    output_pixels,
    payload,
    payload_length,
    width,
    height,
    stride,
    input_pixels,
    pixel_format);

Tracing the call tree, we can see WmkWriteWatermark is called after a local Stable Diffusion image generation. And if WmkWriteWatermark fails, Paint converts the entire generation into an error rather than returning the image without it:

CocreatorViewModel::GenerateImageAsync(...)
  |
  `-- Paint::AI::StableDiffusionHelpers::GenerateAsync(..., watermarkId, ...)
        |
        `-- Microsoft.ImageCreation.ImageGenerator
              |
              `-- NPU-generated image result
                    |
                    +-- output safety/moderation checks
                    |
                    +-- Paint::AI::AddWatermark(bitmap, watermarkId)
                    |     |
                    |     `-- Watermarker.dll!WmkWriteWatermark(...)
                    |           |
                    |           +-- success: return the watermarked bitmap
                    |           `-- failure: turn generation into an error
                    |
                    `-- construct successful StableDiffusionResult

Then it is natural to ask what the incoming payload actually is. It quickly becomes apparent that it must be 16 bytes:

if (payload_length < 16)
    return -6;

if (payload_length > 16)
    return -5;

It is funny to me that the code is using two different error codes when the payload is too short or too long. The function then ignores the length parameter and uses a hard-coded loop bound when it copies the payload:

for (size_t i = 0; i < 16; i++)
    message.push_back(payload[i]);

We do not yet know what the 16-byte payload is, but as we will see later, it is a GUID! WmkWriteWatermark does not embed the GUID directly. Its wrapper constructs the following 18-byte (144-bit) message:

0x4c || GUID[0..15] || (sum of the 16 GUID bytes modulo 256)

The core encoder rounds the usable image dimensions down to multiples of eight and keeps 144 counters, one for each bit. It requires every bit to be placed at least three times.

The encoder itself can be summarized as:

WmkWriteWatermark(output, guid, 16, width, height, stride, input, format)
  |
  +-- validate pointers, format, stride, and payload length
  +-- require width >= 192 and height >= 192
  +-- construct payload
  |     `-- 0x4c || GUID || byte-sum checksum
  +-- expand 18 bytes into 144 individual bits
  +-- round usable dimensions down to 8-pixel boundaries
  +-- scan/select suitable image blocks
  +-- quantize selected block/matrix values according to each bit
  +-- require at least three successful placements per bit
  |     |
  |     `-- insufficient capacity -> return -8
  `-- reconstruct RGB pixels into the output buffer

The embedding loop performs small quantized changes over selected image blocks. It contains 3-by-5 matrix operations and a matrix-decomposition routine, and it uses constants including 24.0, 0.25, 0.5, and 0.2. This looks like a content-adaptive block-domain, SVD-style watermark.

I am not an expert in image watermarking, but one thing should be clear – this is an invisible watermark! AI even wrote some code to call this function directly and tested it with a synthetic 512-by-512 BGRA image – 193,376 of the 262,144 pixels changed after adding the watermark.

That led to the next question. Where does the input of the watermark come from?

a GUID from remote prompt moderation

At the WmkWriteWatermark boundary, the payload is only a pointer and a length. Knowing that it must be 16 bytes was a clue, but many things can be 16 bytes. I therefore started walking backward through its callers. The immediate wrapper in PaintAIManager.dll has this symbolized signature:

Paint::AI::AddWatermark(
    Gdiplus::Bitmap& image,
    winrt::guid const& watermarkId);

winrt::guid, yikes! Now we know that the 16-byte watermark payload is indeed a GUID.

Further tracking the source, we find that the GUID actually comes from a network request. Before Paint runs the local image model, AIServices.dll sends the prompt and style to:

https://apsaiservices-a0fqcjc6bzbhgdcd.b02.azurefd.net/
v1/paint-cocreator/moderate-prompt

The request is JSON and contains at least these fields:

{
  "prompt": "...",
  "style": "...",
  "lastPromptGenerationId": "..."
}

The response parser expects:

{
  "revisedPrompt": "...",
  "promptGenerationId": "...",
  "watermarkId": "...",
  "containsHumanReference": false
}

Static analysis is nice, but at this point I wanted to see a real response from the server. I reused Paint’s own authenticated session and sent the following prompt through the moderation endpoint:

a cobalt blue circle above a tiny orange square

The server returned HTTP 200:

{
  "revisedPrompt": "a cobalt blue circle above a tiny orange square",
  "promptGenerationId": "74d9e06b-adea-43ce-85fe-186a26e2e34a",
  "watermarkId": "83424621-03cb-40e3-9808-a9fae837156d",
  "containsHumanReference": false
}

I also tried the prompt a portrait of a smiling person wearing a blue hat. This time the response contained a different pair of GUIDs and containsHumanReference was true. The field is therefore a server-side classification of whether the prompt refers to a human. Paint parses and stores it alongside the IDs, although I found no evidence that it controls the watermarking step itself.

ParseModerateResponse parses both ID strings as GUIDs and rejects zero values with InvalidPromptGenerationId or InvalidWatermarkId. The server’s watermarkId is what becomes part of the generated image:

PaintUI.dll
  `-- IPromptModerationService
        `-- PaintAIManager.dll
              `-- AIServices.dll!ModerateAsync(...)
                    |
                    +-- build JSON
                    |     +-- prompt
                    |     +-- style
                    |     `-- lastPromptGenerationId
                    |
                    +-- HTTPS POST /v1/paint-cocreator/moderate-prompt
                    |
                    `-- AIServices.dll!ParseModerateResponse(response)
                          +-- revisedPrompt
                          +-- promptGenerationId -> parse as GUID
                          +-- watermarkId        -> parse as GUID
                          `-- containsHumanReference
                                |
                                `-- PaintUI stores WatermarkId
                                      `-- StableDiffusionHelpers::GenerateAsync(..., watermarkId, ...)
                                            `-- local StableDiffusion result
                                                  `-- Paint::AI::AddWatermark(bitmap, winrt::guid const&)
                                                        `-- WmkWriteWatermark(..., guid, 16, ...)
                                                              `-- modified RGB pixels

In other words, “generated locally” does not mean that the complete operation is local. Microsoft receives and moderates the prompt, then issues the unique GUID that Paint embeds into the locally generated image. Paint also sends the previous promptGenerationId as lastPromptGenerationId with its next moderation request, allowing successive requests to be linked explicitly.

The same watermark GUID in C2PA metadata

There is another piece to this story. Paint does more than alter the pixels. It also attaches C2PA Content Credentials to the saved file. The code responsible for this lives in ProvenanceHelper.dll, backed by provenancesdk.dll.

For the local Stable Diffusion path, the flow looks like this:

local Stable Diffusion result
  |
  +-- Paint::AI::AddWatermark(bitmap, watermarkId)
  |     `-- Watermarker.dll!WmkWriteWatermark(..., watermarkId, 16, ...)
  |
  `-- AIServices.dll!SignIngredientOnlineAsync(..., promptGenerationId, image, ...)
        |
        +-- POST /v1/paint-cocreator/image-sign
        |     +-- imageMetadata
        |     |     +-- PromptGenerationId
        |     |     +-- GenerationSeed
        |     |     +-- CreativityLevel
        |     |     +-- AIFVersion
        |     |     `-- moderation scores
        |     `-- imageToSign.jpg
        |
        `-- ParseProvenanceResponse(...)
              `-- server-supplied C2PA manifest
                    `-- ProvenanceHelper::InsertManifestIngredient(...)
                          `-- AuthoringFinalizeOutputToBufferAsync(...)
                                `-- final image with C2PA metadata

Notice that the signing request sends PromptGenerationId, while the image already contains the separately returned watermarkId. The server assigned both values during moderation, so it can associate the signing request with the watermark already present in the submitted pixels.

I then saved a real image directly from Paint’s Image Creator and inspected its PNG chunks. Immediately after IHDR was an 18,979-byte caBX chunk containing a signed C2PA manifest. The interesting part was this:

{
  "c2pa.soft-binding": {
    "alg": "com.microsoft.invismark.1",
    "blocks": [
      {
        "scope": "the entire image",
        "value": "83424621-03cb-40e3-9808-a9fae837156d"
      }
    ]
  },
  "c2pa.actions.v2": {
    "actions": [
      {
        "action": "c2pa.watermarked",
        "description": "Content watermarked by Microsoft Responsible AI"
      }
    ]
  }
}

Decoded into something more readable, the manifest says:

  • Generator: Microsoft Responsible AI Provenance
  • AI system: Azure OpenAI ImageGen
  • Action: c2pa.watermarked
  • Algorithm: com.microsoft.invismark.1
  • Watermark value: 83424621-03cb-40e3-9808-a9fae837156d
  • Description: Content watermarked by Microsoft Responsible AI

The server’s watermarkId, the identifier embedded into the pixels, and the C2PA c2pa.soft-binding.value are the same per-generation value.

That relationship is important. C2PA calls this a soft binding: a value derived from, or embedded into, the content so that the content can still be matched with its provenance record after the file-level manifest has been removed. For a watermark soft binding, the value is the watermark’s content identifier. Microsoft cryptographically signed this assertion.

Why does Paint watermark locally?

At this point, the existence of Watermarker.dll started to make more sense. Paint actually has two rather different generation paths.

The Image Creator feature I tested above uses Azure OpenAI ImageGen. Generation, watermarking, and provenance packaging can all happen in Microsoft’s cloud, and Paint can simply receive a finished image that already contains both the invisible watermark and C2PA manifest:

Image Creator
  `-- Microsoft cloud
        +-- content filtering
        +-- Azure OpenAI ImageGen
        +-- invisible watermark
        +-- C2PA manifest
        `-- completed image returned to Paint

Cocreator is different. On a supported Copilot+ PC, Microsoft says that the NPU generates the image locally, while Azure online services still perform the safety checks. The feature therefore requires both a Microsoft account and an internet connection even though the actual Stable Diffusion inference runs on the device:

Cocreator on a Copilot+ PC
  |
  +-- prompt -> Microsoft moderation service
  |                 +-- revisedPrompt
  |                 +-- promptGenerationId
  |                 `-- watermarkId
  |
  +-- revisedPrompt + sketch -> local NPU generation
  |
  +-- Watermarker.dll -> embed watermarkId locally
  |
  `-- online provenance signing -> final C2PA manifest

This is probably the reason Paint needs a local watermark implementation at all. A cloud generator can watermark its output before returning it. A local generator cannot rely on that, so Paint has to alter the locally generated pixels itself. It also explains why Paint treats a failure from WmkWriteWatermark as a failure of the entire generation instead of quietly returning an unmarked image.

There is another surprisingly visible sign that Microsoft designed the save path around provenance. When I save a generated result directly from the Image Creator pane, Paint offers exactly one format: PNG.

Paint only offers PNG when saving an AI-generated result directly

After an AI result is applied to the Paint canvas, the available formats are still restricted to PNG, JPEG, GIF, and Paint’s own .paint format. BMP—the classic Paint format—is conspicuously absent.

This lines up with the formats supported by C2PA. PNG stores its manifest in a caBX chunk, JPEG uses one or more APP11 marker segments, and GIF has its own C2PA application-extension representation. The .paint format is controlled by Microsoft and can preserve whatever provenance state Paint requires. By contrast, the C2PA specification explicitly calls out BMP as a classic format that cannot embed arbitrary manifest data without using an external manifest. If Paint allowed the image to be exported directly as BMP, the file-level C2PA manifest would therefore disappear.

The split also raises an interesting security question about the cloud path. If the underlying remote image-generation endpoint can be made to return the generated image before watermarking and provenance packaging—or has an internal option that suppresses those stages—it might be possible to obtain a cloud-generated image with neither signal attached.

How to classify such a path would depend entirely on Microsoft’s design goal. It could be intended behavior if the underlying service is allowed to return raw generations and Paint is merely responsible for applying the provenance layers. It could be a product bug if Microsoft overlooked the possibility of someone calling the API directly and bypassing Paint’s watermarking step. Or it could be a security vulnerability if Microsoft treats watermarking as a mandatory abuse-prevention or provenance control and the endpoint can be made to bypass it. Without knowing the intended trust boundary, all three possibilities remain open.

Photos app does the same thing

While I was trying to locate the Watermarker.dll on disk, I happened to notice that Microsoft Photos contains a DLL with the same name:

C:\Program Files\WindowsApps\
  Microsoft.Windows.Photos_2026.11060.2004.0_x64__8wekyb3d8bbwe\Watermarker.dll

There are also local Stable Diffusion operations behind Photos’ Image Creator and Restyle Image features. Both lead to the same watermark wrapper:

Photos Image Creator
  `-- PerformSDTextToImageAndWatermarkAsync(..., promptGenerationId, ...)
        +-- run the local text-to-image model
        `-- ApplyWatermark(image, promptGenerationId)
              +-- parse promptGenerationId as a GUID
              +-- ConvertGUIDtoContiguousByteArray()
              +-- convert RGBA to ARGB
              +-- Watermarker.dll!WmkWriteWatermark(..., guid, 16, ...)
              `-- convert ARGB back to RGBA

Restyle Image takes the parallel path:

Photos Restyle Image
  `-- PerformSDSketchToImageAndWatermarkAsync(..., promptGenerationId, ...)
        `-- ApplyWatermark(image, promptGenerationId)
              `-- Watermarker.dll!WmkWriteWatermark(..., guid, 16, ...)

A subtle difference between Photos and Paint is failure behavior. If the watermark encoder returns an error, its code logs:

ApplyWatermark encountered error: ... - watermark will not be applied.

It then appears to continue returning the generated image. Paint instead treats a watermarking failure as a generation failure and the image is not returned to the user.

What Microsoft discloses

After doing this analysis, I found that Microsoft does disclose some adjacent parts of the system on its Image Creator support page. On content filtering, it says:

“we apply content filtering to prevent the generation of images”

The same page says that generated images:

“will contain C2PA manifest helping users identify that it is an AI generated image.”

It also explains that Image Creator uses Azure online services and says Microsoft collects user and device identifiers together with prompts for abuse prevention and monitoring. That is a meaningful disclosure of remote filtering and C2PA metadata.

What the page does not explain is that the C2PA manifest contains a GUID identifying the invisible pixel watermark, or that Paint’s local generation path receives its watermark GUID from remote prompt moderation. Calling the feature “Content Credentials” is accurate, but it does not make this prompt-associated identifier obvious to a Windows user.

Conclusion

To the best of my knowledge, this is the first research to document and analyze the invisible-watermarking behavior of Paint and Photos. Visible watermarks on AI-generated images are not new—Microsoft documents them for Microsoft 365 and Bing Image Creator—nor are invisible pixel watermarks such as Google’s SynthID and Bing’s hidden watermark.

Microsoft does disclose that Paint uses remote content filtering and adds C2PA Content Credentials. The new evidence shows that this metadata is not merely an unrelated file-level AI label: its signed c2pa.soft-binding assertion names Microsoft InvisMark and records the identifier carried by the invisible pixel watermark. The file-level manifest and pixel-level watermark are two layers of the same provenance system.

The local and cloud paths also explain the unusual division of labor. Cloud Image Creator can return an already watermarked and signed image, while Cocreator must embed the server-issued identifier after local NPU inference. In both cases, “local” does not mean offline: the prompt still goes to Microsoft for moderation, and the completed local result goes through online provenance signing.

This might be related to Article 50 of the EU AI Act, whose transparency rules took effect on August 2, 2026 and require AI-generated content to carry a detectable, machine-readable mark—but not a prompt-specific GUID. Microsoft discloses the existence of C2PA metadata, but I could not find a disclosure explaining the server-issued watermark GUID, its association with prompt moderation, or its presence in the pixels. Those details carry obvious privacy and right-to-know implications.

It also appears possible to modify Paint or Photos to bypass both prompt moderation and watermarking. But that does not provide a new capability: anyone can already run Stable Diffusion directly without either mechanism.

The Daily Front Page 4 of 30
Monday, August 24, 2026 The Daily Front No. #260824 — The Apprenticeship Problem
article

Coding expertise is going to collapse from AI reliance

by larsfaye·▲ 501 points·491 comments·larsfaye.com ↗
The need for ongoing friction in long-term skill formation.

The need for ongoing friction in long-term skill formation.

"We see a future where intelligence is a utility like electricity or water and people buy it from us on a meter and use it for whatever they want to use it for" - Sam Altman of OpenAI

In my previous article, Agentic Coding is a Trap, I discussed the "skilled orchestrator paradox", where the skills required to manage AI agents for coding are the same ones that can be diminished through the continued use of said AI agents. Expertise was largely the differentiator; the more experienced a developer is, the less likely it is that they might experience skill atrophy, as the knowledge has had a chance to ossify after years of experience.

If you look around right now, you'll find the vast majority of those that are seeing the most benefits from these models are those that have had years, if not decades, of experience in the field (which predates AI tooling, of course). And any industry veteran will tell you the same: the bedrock of this knowledge comes from doing the work.

Developers who've entered the field around the time of LLMs are placed in a position where they don't have the benefit of longevity, but they are being guided (and sometimes mandated) to accelerate their efforts using coding assistants that require a history of expertise to wield effectively and responsibly.

It's an awkward place to be for that demographic, as it creates a scenario where a novice needs expert-level skills to leverage the tools and keep pace in the industry.

The "Expert Novice"

We're currently sending very mixed signals to people across the industry. We're hammering in that if you're not using AI tools, you will be "left behind" by your peers who are using them. "AI won't replace you, someone using AI will" has been on repeat since 2023.

And in the same breath, it's also said that the way to get the best results from these models is to apply higher-order thinking; "vibe coding" is a dead end; you need to "move up the stack" and create robust specs, architect with good design patterns, and always review the outputs diligently so you never ship something you don't understand.

The skills to do so, however, are a function of someone who has experienced the friction and challenges over time that culminate in "good taste".

This leads to another situational paradox: If these tools demand expertise, yet the tools can actively circumvent the friction that cultivates expertise, then what is the path for one to become an expert so they can effectively use these tools?

Confidence without Comprehension

One hope is that these models will end up accelerating learning as they are used for code generation. Junior developers can work with the same gravitas and confidence as industry veterans with their "personal AI tutor". Knowing syntax is increasingly less important, and any knowledge or ambiguity gaps are filled by the AI tool. The deeper mechanics of the code stay abstracted away, since the developer sits higher in the stack.

JetBrains, a major player in developer tools, recently cited a study titled "The Widening Gap: The Benefits and Harms of Generative AI for Novice Programmers", which painstakingly analyzed individual behaviors in live coding sessions, and tested their ability to learn coding with varying degrees of AI assistance. Their main takeaway was stark and counterintuitive:

"Participants thought it was like having a personal tutor. From the data in our study ... we observed that they did not, in fact, use GenAI tools like a personal tutor. In fact, it was quite the opposite."

The participants that leaned into heavier AI assistance:

  • "Often skipped crucial planning stages, finding that because they hadn’t reasoned themselves into this position, Copilot had."
  • "Finished with an 'illusion of competence' rather than true understanding."

Counter to that, the participants that mitigated their usage of AI:

  • "Succeeded because they had developed 'negative expertise'—which is 'the ability to ignore incorrect or unhelpful GenAI suggestions'—allowing them to focus on writing their own solutions rather than being led astray."
  • "Were able to use GenAI to accelerate, creating code they already intended to make."

The novice developers who were the most unrestricted and confident in their AI usage "had skipped crucial steps in the programming problem-solving process, and were now lost."

Perhaps unsurprisingly, the novice developers who performed the best were the ones that greatly mitigated or outright ignored the AI coding assistance.

Inverted Learning

Due to the self-directed nature of LLMs, the more experience you have, the more benefit they provide since you can accurately steer, audit, and verify the outputs. The less knowledge you have, the more they can mislead you. Interacting with LLMs for learning new skills takes the shape of an "inverted learning" model, a role reversal where the student is initially guiding the mentor, the mentor responds, and then the student, again, steers the mentor.

The process is precarious; LLMs are incredibly sensitive to the shape of the prompt. When you're exploring new domains, you don't know what you don't know, and the malleable and accommodating design of an LLM can lead you to believe you know more than you actually do.

If you're exploring territory that is even somewhat unfamiliar, you often don't even know the questions that you need to ask that could properly guide the model to providing the best answers. It begins to feel like a compass that always points north, wherever you suggest north might be.

From the same study that JetBrains highlights, even the most prepared students were derailed by the AI assistance due to this type of learning model: One participant demonstrated good fundamental planning and habits, but suddenly "skipped crucial problem-solving planning stages, jumping directly to coding and was enticed by Copilot into quickly producing code" and had to rely on the LLM to fix the error that the LLM introduced in the first place.

AI models lack judgment, empathy, and pedagogical intent, and the solutions provided are not rooted in experience but rather in patterns in the training data (LLMs are, at their core, incredibly complex pattern interpolators).

The infinite answer machine is tempting, and known to be addictive. It can unwind rather quickly, especially for inexperienced developers. Once you get deep enough into a generated solution, you are often beholden to the AI tool to also finish the job, circumventing the problem-solving friction that is required for the formation of a mental model (and to be fair, senior developers are prone to this phenomenon, as well).

The Friction is a Feature

Expertise and mastery don't happen purely through observation and dialogue, but through experience, repetition, and trial and error; you have to fail to succeed. If I wanted to learn how to cook, I could watch a Master Chef work and make endless inquiries. After a month, I would be able to describe the perfectly medium-rare ribeye but never know what it's like to cook one, and I'd almost certainly overcook it on my first attempt.

Coding has endless moments of tracing obscure errors with no log file to help, experiencing the subtle performance differences of certain methods, or having to rewrite an approach when it's clear it won't going to scale.

This applied friction is directly what builds "developer intuition" (or "taste"). The Germans have a great word for this: Fingerspitzengefühl (fingertip feeling). It’s the muscle memory that triggers when a developer looks at something and thinks, “yeah...this is probably going to cause problems.” By avoiding the mechanics of the struggle, this intuition is never built.

In UPenn's large-scale 2025 study Generative AI without guardrails can harm learning, they followed 1,000 students using an LLM to learn mathematics and found students used AI as a crutch and ended up performing 17% worse than students with just a textbook (and just as with the previous study, the students using the AI assistance thought they were excelling).

LLMs don't just have to generate code, though.

If leveraged as Socratic sparring partners instead of answer generators, studies have shown that "dialogic AI systems can meaningfully stimulate reflective, critical and independent thinking".

In that same UPenn study, they also tested a "Tutor" version by having students ask for help and then independently solve the problem. The GPT Tutor group performed an astonishing 127% better in the AI-assisted practice session (although, interestingly, they scored about the same on the test as the textbook group).

This is effective because the model is no longer being utilized as a means of production, and it shifts the cognitive work back onto the individual. It's when the friction is still present that it creates a lasting imprint that leads to expertise.

Anthropic's 2026 study "How AI assistance impacts the formation of coding skills" came to similar conclusions:

For novice workers in software engineering or any other industry, our study can be viewed as a small piece of evidence toward the value of intentional skill development with AI tools. Cognitive effort—and even getting painfully stuck—is likely important for fostering mastery.

There's a certain sense of irony here: the most productive learning that can happen with an AI coding tool is when it isn't used to generate much of any code at all.

Pipeline Collapse

If LLMs can write code and debug code, and agentic workflows can perform system design from the abundance of patterns in the training data, then what is the purpose of this knowledge in the first place? Programming will be done entirely in natural language, and we can dispense with the need to engage with the code because the models continue to improve and fill in any knowledge or ambiguity gaps. They will debug any issues that arise and manage any complexity that they introduce.

The trillion-dollar bet that is being made is: this knowledge won't matter, because LLMs will take up the slack and effectively become the new generation of "developers". It starts give off an aire of hubris that drove past no-code movements, and the fever dreams of CEOs, rather than the reality on the ground.

Coding/programming/software is a unique intersection of logic, math, problem-solving, critical thinking, planning, communication, and creativity. LLMs can detect patterns at a scale that no human ever could, but patterns only get you so far.

David Cramer, co-founder at Sentry (a performance and error tracking platform), put it succinctly in a recent interview:

I think there's a type of person ... that inherently believes that LLM will get better enough that they will go back and fix this stuff, that it will be able to clean up all the junk that's been stacked up along the way. I don't think that's true. I think it's a science experiment.

You want to flex that you can generate all of your code and have hundreds of things going in parallel, I will flex and show you how broken the code is 100% of the time.

Will the pipeline collapse, or just change?

It really depends on whether we make the needed shift to a more pedagogical usage of these systems.

By continuing to focus on and promote AI coding workflows that prioritize code generation above deep understanding, we are not cultivating the next generation of expertise who will inherit the code that is being created today.

Coding Agents Mentoring

My Approach: Friction First

Joel Spolsky presciently writes (in 2002, no less) in his Law of Leaky Abstractions:

Code generation tools which pretend to abstract out something, like all abstractions, leak. And the only way to deal with the leaks competently is to learn about how the abstractions work ... the abstractions save us time working, but they don’t save us time learning.

If a developer wants to learn Java, they should probably not start with Spring Boot. If they want to learn JavaScript fundamentals, they should not start with React. If they want to become highly adept at CSS, they should not start with Tailwind. LLMs could be considered the ultimate leaky abstraction.

My advice here is very similar to my previous prescription.

If a developer wants to become an expert in programming, they should largely disregard the pure code generation capabilities of these models, and instead use them for interactive documentation, dynamic tutorial generators, and Socratic exercises.

It's not a panacea, of course: Using an AI tool as a tutor carries its own risks since it is susceptible to the same hallucinations as any other interactions, and it cannot be relied upon solely as a learning source. If you can't properly audit the accuracy of the generated code, they you can't audit the accuracy of the generated concept. If you use AI as a mentor, you must still verify its outputs against official documentation, human peers, and actual trial and error.

"Coding's actually a great way to cement understanding. The more you program, the more you understand the domain that you're working in."

— Kent Beck, creator of Test-Driven Development

Choosing this slower, more deliberate path is the best way to grow expertise, but I'm aware of how hard that is when the surrounding ecosystem is actively working against it. AI is being mandated (often recklessly) across companies, and baked into most software development tools and IDEs as they cater largely to senior engineers (even with some tools like Cursor tucking away the code view unless the user specifically seeks it out). Some companies are even forcing developers to only use AI for all coding tasks, regardless of experience level, and these companies will have to learn their own lessons.

However, for everyone else who is looking to strike a balance between deep learning (no pun) and productivity, there are qualifying questions you can ask to ensure your usage of these tools yields long-term benefits.

My AI-assistance checklist:

  • If I did not have access to an AI tool, could I still accomplish this task?
  • Am I using the model to deepen my understanding, or expedite the answer?
  • If I had to audit and verify the generated output, could I adequately explain what was happening?
  • If I'm learning a new concept, have I done proper research to know the right questions to ask?
  • Have I cross-referenced and verified the approach through other methods (reading documentation, standard search tools, StackOverflow, Reddit)?
  • Is this a truly rote task that's been done 100 times before, or a task that requires executive decision-making somewhere in the process?

Even as a developer with decades of experience under my belt, I am still constantly referring to them throughout my daily work, especially when I am attempting to learn something new (which in this field, is neverending).

The key is to detect the difference between cognitive debt and cognitive offloading: Cognitive debt is abdicating your judgment and decisions, whereas cognitive offloading is delegating the mechanical or tedious.

As the Anthropic study mentioned, getting "painfully stuck" is a good thing. It takes discipline and effort to not drift back towards just generating answers, which might not even be accurate in the first place. LLMs didn't suddenly rewrite the fundamentals of how we learn, but they did give us a new way to do so.

Intelligence isn't a Commodity

The realignment I hope to see over the years is the understanding that skills don't develop without active participation. You must engage directly and continously to experience the essential friction that culminates in expertise (even if it means moving more slowly).

If we stay fixated on lines of code and tokens burned while the expertise pipeline dries up over the years, Sam Altman's vision of selling intelligence back to us on a meter could become reality. Domain knowledge could become very hard to come by, and when one sits down to do any type of development work, there will be a pang of paralysis if that person does not have an active AI tool subscription at their side.

LLMs are a static database of skills. They are interpolation engines. Software engineering, however, is an exercise in adaptation and novel problem-solving. You cannot interpolate your way through a completely unique system failure.

— François Chollet, creator of ARC-AGI Benchmark

The Daily Front Page 5 of 30
Monday, August 24, 2026 The Daily Front No. #260824 — An Executable, Recast
article

Executable Is a SQLite Database

by setheron·▲ 509 points·98 comments·fzakaria.com ↗
Replacing ELF with SQLite as an executable format.

I have been probably obsessed with two things in the last few years: Nix as a tool to explore innovative ideas that require the capability to rebuild the world and replacing ELF with SQLite as an executable format. You might have noticed that these two ideas are well suited to each other.

I explored the idea during my PhD thesis but found feedback from others unmotivating. Radical ideas are hard to sell, as you are working against the inertia of the established solution.

Four-panel comic. A crow at a microphone says "Nix is great"; the audience boos and shouts "get better material"; the crow looks stricken; the last panel shows its remaining cue cards, which read "SQLite can be an object file format".

One of the end results of that exploration was sqlelf, a tool that lets you explore an ELF file declaratively using SQL. I wrote a paper, arXiv:2405.03883, that I failed to get published and a follow-up post on querying with it. SELECT name FROM elf_symbols instead of fiddling with readelf and grep. It was remarkably simple by leveraging virtual tables over the ELF: however I found it to be a refreshing improvement to explore the ELF file format. I knew however that there is still something much bigger to be done.

I never let the idea go and with the recent improvements with LLMs, I find it compelling to revisit these ideas to explore further. Specifically, can we replace ELF with SQLite as an executable format? 🤔

Not “a database that describes an executable”, but the actual file you chmod +x and run.

$ file hello
hello: SQLite 3.x database, application id 0x53454c46, user version 1

$ ./hello
Hello, world!

$ sqlite3 hello 'SELECT soname FROM ldd'
libc.so.6

I developed a pretty fleshed out prototype. It is called SELF, the Structured Executable & Linkable Format, because I am unoriginal. It is on GitHub if you are interested. I’m surprised about all the interesting things that fall out of this idea.

§ELF is a database that refuses to admit it

Working through my PhD, I realized something that bugged me. ELF is already a database. It just implements many database primitives by hand, along with a surprising number of data structures for performance, like a bloom filter for symbol lookup.

ELF mechanism The database primitive it reinvents
.strtab / .dynstr string interning
.hash / .gnu.hash an index (CREATE INDEX)
section header table sqlite_schema, a table of tables
st_name → offset into .strtab a foreign key, done by hand
sh_offset / sh_size the record layout of a b-tree page
.gnu.version_r a column
objcopy --strip-debug DELETE + VACUUM
ldconfig cache, debuginfod out-of-band indexes over the above

If you ever have to analyze or parse ELF, the kernel, ld.so, binutils, LIEF, goblin, readelf, you are re-implementing the same parser over and over again. Every producer re-implements the same serializer.

The format itself is incredibly terse, designed for a world where disk space and network bandwidth was at an extreme premium. Modifying the format is hard, you often have to zero out sections and add new ones since it is packed so tightly. There is also no self-describing schema. ELF itself is a very generic format that supports sections of data that by convention are interpreted in specific ways but the format does not enforce it.

SQLite is the counter-example. They are a self-describing format that is extremely stable. It is designed to be extended to support new features without breaking existing consumers and supporting a wide range of queries performantly.

If we were to replace ELF with SQLite, what would fall out and can all of the necessary information be represented in a SQLite database? The answer is yes, and it is surprisingly simple.

§What falls away

A SELF file needs two tables to run: self_meta is the ELF header as key/value pairs and segments is the load image, one row per program header with the bytes in a BLOB:

CREATE TABLE segments (
  -- original phdr index
  id      INTEGER PRIMARY KEY,
  -- 'load' | 'tls' | 'stack' | 'relro'
  type    TEXT NOT NULL,
  -- original file offset
  offset  INTEGER NOT NULL,
  vaddr   INTEGER NOT NULL,
  filesz  INTEGER NOT NULL,
  memsz   INTEGER NOT NULL,
  r INTEGER, w INTEGER, x INTEGER,
  align   INTEGER NOT NULL DEFAULT 4096,
  -- the segment bytes; NULL for pure BSS
  content BLOB
);

A single table for the symbol table replaces many of the ELF sections and the .gnu.hash index. It is a single table with a single index:

CREATE TABLE symbols (
  id      INTEGER PRIMARY KEY,
  name    TEXT NOT NULL,
  -- 'GLIBC_2.2.5'
  version TEXT,
  value   INTEGER,
  size    INTEGER,
  -- 'func' | 'object' | 'tls' | ...
  type    TEXT,
  -- 'global' | 'weak' | 'local'
  bind    TEXT,
  defined  INTEGER NOT NULL,
  exported INTEGER NOT NULL
);
CREATE INDEX idx_symbols_name ON symbols(name, version);

Our capability to include an index is equivalent to .gnu.hash and .hash in ELF, but it is a proper b-tree index maintained by SQLite instead of a hand-rolled bloom filter. .gnu.hash is a bloom filter plus bucket chains, laid out so ld.so can reject a miss without touching the chain during symbol discovery.

Surprisingly a lot more falls out as well: .dynstr is gone, because name is TEXT and SQLite already interns strings, symbol versioning is a column, not the .gnu.version_r / .gnu.version_d contraption and there is no need for a strings table.

Other tables exist as well for metadata which exist for tooling: sections, notes, dynamic_entries. Delete them and the program still runs, which means strip(1) is a transaction:

# ldd(1)
$ sqlite3 hello 'SELECT soname FROM ldd' 
libc.so.6

# nm -D --undefined
$ sqlite3 hello 'SELECT name,version FROM imports LIMIT 3'
__libc_start_main|GLIBC_2.34
_ITM_deregisterTMCloneTable|
puts|GLIBC_2.2.5

# readelf -l
$ sqlite3 hello \
    "SELECT type,vaddr,memsz,r,w,x FROM segments WHERE type='load'"
load|0|1744|1|0|0
load|4096|361|1|0|1
load|8192|312|1|0|0
load|15768|640|1|1|0

# strip(1)
$ sqlite3 hello 'DELETE FROM sections; DELETE FROM notes; VACUUM;'
# 57344 -> 49152 bytes

# still runs,  the optional tables were optional
$ ./hello
Hello, world!

All the tools that operate on ELF files for reading, reduce to queries over the database. Any tool that modifies an ELF file, like strip, can operate on the database within a transaction rather than performing fragile offset surgery: strip is a DELETE and VACUUM. patchelf is an UPDATE.

Any information missing from the schema can be easily exposed via a view. For example, ldd is a query over the needed table, which is a join of the symbols table with the segments table to find the sonames of the libraries needed by the program.

CREATE VIEW exports AS SELECT name, version, type, size FROM symbols WHERE exported = 1;
CREATE VIEW imports AS SELECT name, version FROM symbols WHERE defined = 0;
CREATE VIEW ldd     AS SELECT ord, soname FROM needed ORDER BY ord;

§How does it work?

SQLite reserves a 4-byte application_id at byte offset 68 of its header, for exactly this purpose. We stamp it SELF, so an ordinary SQLite database never matches:

$ xxd -s 64 -l 8 hello
00000040: 0000 0001 5345 4c46                      ....SELF

We can now leverage binfmt_misc, the subsystem that allows you to invoke any binary as if it were native. We need only to register the magic to trigger on and an interpreter that will invoke our new file format.

On NixOS the registration is a few lines matching the SQLite magic at offset 0 and SELF at 68:

boot.binfmt.registrations.self = {
  recognitionType = "magic";
  offset = 0;
  # bytes 0-15, 68-71
  magicOrExtension = "SQLite format 3\\x00" + ... + "SELF";
  # ignore the middle
  mask = "\\xff..\\x00..\\xff";
  interpreter = "${self-exec}/bin/self-exec";
};

For now, I have a small tool elf2self that converts an ELF file into a SELF file. It is a simple postFixup hook you can opt into per package on NixOS. The tool reads the ELF, extracts the program headers and symbol table, and writes them into the SQLite database. We could look at extending gcc or ld to emit SELF directly, but for now this is a simple way to explore the idea.

self-exec is the interpreter. It is a small C program linked against libsqlite3. Its implementation is remarkably similar to that of ld.so but it fetches the program headers and symbol table from the database instead of reading them from the ELF file. It maps the loadable segments into memory, relocates them, and jumps to the entry point.

Note self-exec has to stay an ELF file. An interpreter that also matches the registration recurses straight into -ELOOP.

§Dynamic linking

Running a static program was quick and easy but boring and unimaginative. The interesting part is dynamic linking, which is where the database shines.

I explored two different ways to do dynamic linking. The first is to keep ld.so and just replace the lookup with a SQL query via glibc rtld-audit interface, to quickly iterate on the design. The second is to replace ld.so entirely with a new dynamic linker that does the entire lookup and binding in SQL.

glibc’s rtld-audit interface lets an audit library intercept every shared object lookup (la_objsearch) before any filesystem search happens, dlopen included. The audit library can then answer the question “which library satisfies this symbol?” with a SQL query instead of walking the RUNPATH and LD_LIBRARY_PATH. Stock ld.so maps and relocates it, so the full gamut of glibc features work: lazy PLT, IFUNCs, TLS and symbol versioning, while library storage are rows and library lookups are queries.

# no ELF library anywhere on disk
$ rm libgreet.so.1
$ ./app
./app: error while loading shared libraries: 
       libgreet.so.1: cannot open ...

$ self scan --db system.db .
$ SELF_SYSTEM_DB=system.db LD_AUDIT=libself-audit.so ./app
Hello, world, from a SQLite library!

I was curious what a fully SQL dynamic linker would look like, so I prototyped one. It is called self-ld and it is a small C program that implements the dynamic linker entirely in SQL. It is a proof-of-concept, but it works. It maps every object’s segments, publishes their exports, and for each relocation patches the GOT and jumps to the start.

SELECT s.value + o.load_bias
FROM   relocations r
JOIN   symbols s ON r.symbol = s.id
JOIN   objects o ON s.object = o.id
WHERE  r.id = ?
ORDER BY o.load_order
LIMIT  1;

§Cost & Benchmark

The two things that often matter when replacing a well-established format are size and latency. How much bigger is a SELF file than an ELF file, and how much slower is it to run?

Size. A SELF file carries SQLite’s b-tree overhead and lands at roughly double the ELF.

Similar to ELF binaries, most of that is recoverable, because the overhead is mostly the optional tables for debugging and tooling. Stripping them and deleting them is a transaction. A stripped coreutils SELF is 1,794,048 B against the ELF’s 1,768,632 B, that is within 1%.

We will see though that there are interesting ways to amortise the overhead even more which I found very unique and interesting.

Latency. I benchmarked various binaries from a 15 KiB hello to a 42 MiB gdb linking 47 libraries:

There is a fixed ~5 ms to open SQLite and start the interpreter, plus a copy proportional to the image. That copy is worse than it looks, because the b-tree pages are not mapped into memory. Two processes running the same SELF binary do not share text pages the way a normally-mmap‘d ELF does, because the bytes are copied out of the b-tree rather than mapped. You might notice that curl (274 KiB, 27 libraries) starts slower than ELF git (4.6 MiB, 5 libraries). That is ld.so doing work proportional to the number of objects rather than the number of bytes, which I have complained about before.

§The system is a closure

A SQLite database though need not merely be a single executable. It can be a closure, a single file that contains a program and all of its transitive dependencies. The ldd output of a program is ambiguous: it only lists the sonames of the libraries it needs, not the specific files that satisfy those needs. Nix improves upon this by explicitly resolving every edge to a specific store path via the use of RUNPATH. I have written about RUNPATH on Nix before such as making it redundant or speeding it up.

We can do the same in SELF by storing the resolved path of each edge in the database:

CREATE TABLE objects (id INTEGER PRIMARY KEY, path TEXT UNIQUE,
                      soname TEXT, kind TEXT, is_root INTEGER);
CREATE TABLE needs (
  object_id     INTEGER REFERENCES objects(id),
  ord           INTEGER NOT NULL,
  soname        TEXT NOT NULL,
  -- the FK that kills ambiguity
  resolved_path TEXT REFERENCES objects(path)
);

self closure packs a binary and its transitive dependencies into one database with those edges filled in. Shared library resolution stops being a guess and becomes a foreign key and ldd becomes a JOIN 🤯:

$ self closure "$(readlink -f $(command -v ls))" coreutils.db
ls + closure -> coreutils.db

$ sqlite3 -column coreutils.db \
    "SELECT n.soname, substr(n.resolved_path, 12, 20)
     FROM needs n JOIN objects o ON o.id = n.object_id
     WHERE o.is_root = 1"
libgmp.so.10          rfabfsmwq02sn94mb3qg
libacl.so.1           x0zgiss9hdzcsll3cswg
libattr.so.1          08nfpyc4qhzdkc37nznv
libc.so.6             8kvxvr3pmsypxiypq4g8

This single database is a closure of the ls executable and its five libraries: six objects, segment bytes and all, in one 4.8 MiB file. There is no soname ambiguity inside a closure, because a closure by construction contains exactly one provider per edge.

§How far does this go? One file, one userland

I hope you’ve been with me so far, because this is where it gets really interesting. We can go even further and pack multiple closures into a single database.

Five-panel Inception meme. Cobb: "your executable is a SQLite database." Fischer: "and the libraries it links?" Cobb: "also SQLite, so is the whole userland, one file." Fischer: "how far down does this go?" Cobb, winking: "you are in one right now."

I pointed self closure at every ELF binary on this system’s PATH: 723 executables, which pull in 400 distinct shared libraries. 1,123 objects, 346,386 symbols, 3,808 dependency edges, all as one SQLite file.

Turns out when you do that, the database is much smaller than you would expect.

611.9 MiB of database against 644.4 MiB of ELF files. The whole userland, as one queryable file, is smaller than the files it came from. The b-tree cost that doubled a single hello amortises to nearly nothing across 1,123 objects and is roughly 6% over the actual program bytes.

The libraries and closure are shared across the executables very similar to how Nix might share them across multiple closures, if the store-path was the same. If every root shipped its own private closure (i.e. the AppImage model), the same 723 programs would come to 5.53 GiB but the deduplication of libraries and symbols falls out naturally from the database schema.

$ sqlite3 userland.db \
    'SELECT count(DISTINCT soname), count(*)
     FROM objects WHERE soname IS NOT NULL'
345|399

$ sqlite3 -column userland.db \
    'SELECT soname, count(*) FROM objects
     WHERE soname IS NOT NULL
     GROUP BY soname HAVING count(*) > 1
     ORDER BY 2 DESC LIMIT 4'
libsystemd.so.0   3
libpthread.so.0   3
libgcc_s.so.1     3
libc.so.6         3

$ sqlite3 userland.db \
    "SELECT count(*)
    FROM needs
    WHERE resolved_path IS NULL AND soname NOT LIKE 'ld-%'"
4

Many common idioms we use in ELF immediately fall out of the database. For example, LD_PRELOAD is a row in a table rather than an environment variable. The preload table is a list of objects to map last, so their exports win. This means that turning LD_PRELOAD on and off is a transaction.

$ ./app.self; echo $?
13

$ sqlite3 system.db "BEGIN;
    CREATE TABLE preload(ord INTEGER PRIMARY KEY, path TEXT);
    INSERT INTO preload VALUES (0, 'libmul.so.1.self');
  COMMIT;"

# same binary, no env var, no relink
$ ./app.self; echo $?
42

$ sqlite3 system.db 'DELETE FROM preload;'
$ ./app.self; echo $?
13

We were able to accomplish an atomic LD_PRELOAD across a whole userland in one file, “interpose a tracing malloc everywhere, then ROLLBACK” is a single transaction. 😈

§Where it stands

The format is done and round-trips between ELF and SELF losslessly. The tooling is done and can query, modify, and pack closures. Lookup through SQL works on unmodified glibc programs perfectly and the native-SQL loader works enough to explore it as a possibility for ideas.

The whole thing is at fzakaria/selfdb. nix run .#self-vm boots a NixOS VM where hello is a SQLite database. 🙌

Nix lets us explore radical ideas like this. We can rebuild the world down to the Linux kernel if needed. We need not be constrained by the existing decisions and constraints of the past. We can explore new ideas and see what falls out. I hope you find this idea as interesting as I do.

The Daily Front Page 6 of 30
Monday, August 24, 2026 The Daily Front No. #260824 — Proof at the Kernel
article

SeL4 security proofs now complete on AArch64

by snvzz·▲ 176 points·38 comments·proofcraft.systems ↗
A formal mathematical proof that the kernel prevents an application running on top of seL4 from learning information without authorisation.

seL4 security proofs now complete on AArch64

After completing the proofs of functional correctness and integrity, Proofcraft has now established the proof that seL4 enforces confidentiality on AArch64, providing a formal mathematical proof that the kernel prevents an application running on top of seL4 from learning information without authorisation.

Thanks to continued support from NCSC, this milestone completes the formal proof that the seL4 implementation code on AArch64 enforces security isolation of the applications running on top (under the assumptions listed here). This isolation prevents attacks on non-critical applications from propagating to critical applications and compromising them.

Status of seL4 proofs on AArch64 with now confidentiality done and system initialisation started

Proof Engineering and Theory at LICS'26

Title page of the paper The Algebra of Iterative Constructions

The paper The Algebra of Iterative Constructions by Kevin Batz, Benjamin Lucien Kaminski, Lucas Kehrer, Gerwin Klein, Henning Urbat, and Todd Schmid was presented at the 41st Annual Symposium on Logic in Computer Science (LICS) in Lisbon this week. This paper in theoretical computer science is about an algebraic abstraction and reasoning principles for the iterative construction of fixed points. Fixed points are a recurring theme in computer science with many famous results such as the Kleene fixed point theorem. The algebra shown in this paper allows expressing such theorems concisely and enables reasoning about them in an abstract and streamlined way that can be implemented efficiently in proof assistants such as Isabelle/HOL, which Proofcraft is using for the verification of the seL4 microkernel.

The highly automated Isabelle/HOL implementation of iteration algebra in this paper resulted from a spontaneous collaboration between Proofcraft’s Chief Scientist Gerwin Klein and Benjamin Kaminski that started at the IFIP Working Group 2.3 (Programming Methodology) meeting in Athens in 2025. It shows that proof engineering ranges from practical application all the way to deep theory.

MCS seL4 now verified! (for RISC-V)

Proofcraft achieved a significant milestone in the seL4 verification roadmap that was years in the making: the MCS configuration of seL4, providing support for mixed-criticality systems, is now proved to be correct on RISC-V.

This configuration is the largest new seL4 feature, indispensable for mixed criticality real-time applications such as automotive use cases. It contains wide-ranging changes to the kernel’s implementation and API. Its verification therefore required considerable effort and has been a priority in the seL4 roadmap for a long time.

Proofcraft has now completed, for the very first time, the verification of functional correctness for seL4 with MCS. Functional correctness is the largest and most central proof in the seL4 verification stack. The proof targets the RISC-V architecture and will now be ported to the Arm 64-bit architecture, as part of DARPA’s PROVERS program.

MCS verification status

Dynamic Domain Scheduler for seL4

Proofcraft delivered the implementation and formal proof of more flexible domain scheduling in seL4.

Before the change, the seL4 security proofs, and in particular the proof of information flow enforcement, required a fully static schedule that was compiled into the kernel. This meant that, when using seL4 to enforce the information flow boundaries between applications, developers were required to provide a fixed predetermined amount of time for each domain, for the entire lifetime of the running system. This strict policy made it hard to apply information flow control in practice and to support in SDK-style development such as the Microkit.

Proofcraft proposed a new seL4 runtime API (Application Programming Interface) allowing the loading of semi-static domain schedules. This means that a system with information flow protection can go through different phases at runtime that can satisfy different domain timing requirements. For instance, a boot phase of the system can have longer time slices to allow virtual machines to start without overrunning their domain time allocation, and an operational phase of the system can provide shorter time slices so that each domain can be responsive to outside interaction. Additionally, an SDK-based system such as the Microkit can use the new API to set a domain schedule at boot time.

This new seL4 API is implemented, verified and available in seL4 15.0.0.

Diagram illustrating status before with one schedule versus the current status with multiple static schedules

June Andronick Keynote at CDIS Spring Conference in Stockholm

On May 21st 2026, CDIS – Swedish research Center for Cyber Defense and Information Security – held its spring conference at KTH Royal Institute of Technology in Stockholm.

Proofcraft CEO June Andronick was one of the two keynote speakers, alongside August Martens from Mistral AI. June gave an overview of formal verification for cybersecurity, and participated in a panel on Digital Sovereignty.

Picture of June giving talk and panel

Proofcraft presenting at the Cyberagentur Milestone Research summit

Representation of 2 title slides for the 2 presentations at the summit

In April 2026, Germany’s Cyberagentur held a Milestone Research summit to present the progress and outcomes of its funded programs, including the Ecosystem trustworthy IT research program (ÖvIT), which Proofcraft is a recipient of, partnering with Kry10.

Proofcraft’s Chief Scientist Gerwin and Kry10’s Chief Scientist Martin Dehnel-Wild presented the progress on the Dyvercon project, to deliver dynamism, performance, and proof for complex cyber-physical systems. In particular, Gerwin reported on Proofcraft’s work on extending the seL4 proofs to support a static multikernel configuration, where applications can benefit from the use of multiple CPU cores for performance, while at the kernel level a separate instance of seL4 run on each core.

Gerwin additionally gave a general introduction to formal verification and overview of its use in the real world.

Proofcraft is a proud sponsor of the seL4 summit 2026

Logo of the seL4 summit

Proofcraft is happy to be supporting the 2026 seL4 summit as a Silver sponsor.

The seL4 summit is an annual international gathering of participants from industry, government and universities with interests in the world’s most highly assured OS kernel. Attendees and presenters include the creators and maintainers of the seL4 technology such as the Proofcraft team.

This year’s seL4 summit will be held in Vancouver, Canada, on Sep 1-3, 2026.

Icon of Vancouver skyline

5 years of Proofcraft. 5 years closer to a verified future.

Proofcraft logo with 5 fireworks

On the 14th of April 2021, we created Proofcraft. Five years later, we are so busy working for a verified future that we have not posted news for a while.

Much has happened, and more is to come. For now, here are some posts from our back log of news items with technical highlights that Proofcraft has been delivering.

Firstly, the seL4 proofs are now supported on 100% of Arm platforms that the kernel can run on. With this significant progress towards reducing the reliance on experts, users of seL4 can now choose freely between the supported Arm platforms and always be sure they use a verified code base. This work is part of DARPA’s PROVERS program.

Secondly, seL4 on AArch64 now provably enforces integrity: We have a formal mathematical proof that the kernel prevents an application running on top of seL4 from modifying data without authorisation. And the work on security theorems goes on: thanks to continued support from NCSC, we are close to completing the confidentiality property, and with that the entire security proof stack for the 64-bit Arm architecture.

Much more is happening, with three large projects going on in parallel, funded by DARPA, Cyberagentur and NCSC respectively. Stay tuned for more!

The Daily Front Page 7 of 30
Monday, August 24, 2026 The Daily Front No. #260824 — The Venture Future
article

Andreessen Horowitz is investing billions into a bleak future

by reasonableklout·▲ 673 points·352 comments·modelrepublic.org ↗
The firm’s investment portfolio is full of companies that have exploited legal loopholes, created disturbing products, and broken the law.

The firm’s investment portfolio is full of companies that have exploited legal loopholes, created disturbing products, and broken the law.

Marc Andreessen wants to shape US AI policy. The venture capital firm he co-founded and runs, Andreessen Horowitz (abbreviated “a16z”), is a major player in the development of new tech startups.

These startups include:

  • A bot farm of fake accounts, tricking people and social media platforms into thinking AI-generated ads are posted by real people
  • An AI company that wants to normalize cheating on dates, job interviews, and tests with AI
  • AI companion apps linked to suicide and disturbing behavior toward children
  • A platform hosting thousands of deepfake models — 96% targeting identifiable women — that have been used to create AI-generated content sexualizing children
  • Gambling platforms that attempt to subvert existing laws and target vulnerable users
  • Fintech companies implicated in fraud and illegality

Many of these companies knew the rules and broke them anyway — or designed products specifically to exploit gaps in consumer protection. The firms profited, and the public paid the costs.

There’s a growing public desire to rein in tech companies and regulate AI, so a16z is spending tens of millions of dollars to shape the development of AI policy. The firm helped launch a $100 million super PAC, saw former partners take key government roles, and successfully pushed for an executive order attempting to undermine state AI laws. The partners want to set the rules of the road, even as they’re already driving recklessly.

What follows is The Midas Project's survey of 18 of Andreessen Horowitz's most notorious investments. This isn’t a comprehensive overview of the firm’s larger portfolio, but it indicates a pattern of behavior — one comprising hundreds of millions of dollars of investment by a16z.

These investments reveal the lines that a16z is willing to cross and how the lax regulatory environment that they favor would benefit the firm’s bottom line.

A16z did not respond to a request to comment for this report.

Deception and manipulation

A16z has invested in products designed for mass deception. Even if these tactics don’t explicitly violate the law, they can be corrosive to society.

As technology like advanced AI improves — making it much easier to fake almost anything — decision makers may want to enact new laws or policies that mitigate the social costs. And if a16z gets its way, we might never update the rulebook.

Doublespeed

A16z invested $1 million in October 2025 via Speedrun.

Doublespeed sells the capacity to trick everyday people, and social media platforms themselves, into thinking AI-generated ads are genuine human content. Here are some select quotes from the company’s promotional video:

  • “We run the only VC-backed bot farm in America. Because why let Russia and China have all the fun?”
  • “We didn't break the internet. It was broken to begin with. But now we're killing it entirely.”
  • “Welcome to the dead internet.”

A16z's Speedrun program invested $1 million in Doublespeed, a company that was recently covered in a blistering article by 404 Media, which reported: “Andreessen Horowitz is funding a company that clearly violates the inauthentic behavior policies of every major social media platform.”

Excerpts from Doublespeed’s website

The company’s business model relies on deception, designed to make social media platforms and their users believe AI-generated images and videos depict real people.

How do they do this? By selling access to “phone farms” that create and manage thousands of fake social media accounts to manipulate engagement metrics. The company's website is explicit, saying its product “mimics” the behavior of real people on social media in order to “get our content to appear human to the algorithms.”

“Yes, we built a phone farm (and its pretty sick),” said Doublespeed founder Zuhair Lakhani on X. The purpose was “replacing human creators with ai, mainly used for marketing.”

A photo of Doublespeed’s phone farms, shared by the founder Zuhair Lakhani on X.

They use thousands of real phones to pull this off because social media platforms like TikTok have policies against and methods to detect the mass generation and deployment of fake accounts.

The company has the accounts imitate human behavior before posting deceptive content. This means the fake accounts search specific keywords, scroll their “For You” pages, and use AI to analyze screenshots of content to determine whether to “repost it, comment on it” or “swipe away.”

A feed of AI-generated marketing content created by Doublespeed. Source: Superwall on YouTube

A selection of nearly identical Doublespeed-run TikTok accounts. Most posts involve the AI decoy complaining about any one of a number of medical issues. Then, the account lists a handful of cures, including a foam roller product from Doublespeed’s client. Source: Tiktok, Doublespeed on loom

This is all designed to circumvent platforms’ restrictions on fake content and then serve that fake content to unsuspecting real people.

In a podcast interview, Lakhani offered details about one of the company’s clients: “They're hitting like the old person niche, which is what I think is like the best niche to hit with AI content.”

Polling and research have found that older people are less likely to say they’ve heard about AI and more likely to fall for AI-generated misinformation.

Lakhani drew a parallel between this client and his prior work producing AI-generated marketing content at scale: “It was all like old person niche stuff. So like all supplements that would, you know, target old people, and that's when the commission would go crazy.”

“Those brands would tell you to do like, you know, make some like extremely crazy claims,” he said, “especially with supplements.” Lakhani added, “The supplement stuff should definitely be like kind of illegal. I don't know how that is allowed.”

Despite their founder stating that supplement ads should be illegal, Doublespeed isn’t shying away from them. In December 2025, a hacker gained access to Doublespeed’s entire backend and the leaked data showed what the AI-generated “influencers” were actually selling.

One account, “pattyluvslife,” featured an AI-generated woman claiming to be a UCLA student. The account criticized the supplement industry and pharmaceutical companies as fraudulent — while simultaneously promoting a herbal supplement from a brand called Rosabella.

Another account under the name “chloedav1s_” had uploaded some 200 posts featuring an AI-generated woman claiming to suffer from various health conditions and often pictured in a hospital bed. She ultimately promoted a specific company’s foam roller as a solution to her ailments.

A tweet from DoubleSpeed’s founder shows one of the company’s bot accounts messaging a user to promote the product. In the post, Lakhani boasted, “A couple of weeks ago, we gave the [AI] agents access to dm … This was for an ecommerce brand - out of 130 dms sent, 15 pointed to a conversion.”

Another image from Doublespeed’s platform showing their bot account, imitating a human and messaging users with medical conditions to promote the client’s foam roller product. Source: Zuhair Lakhani on X.

The Doublespeed hack revealed more than 1,100 phones and over 400 TikTok accounts operated by the company. Most of the accounts were promoting products without disclosing that the posts were paid advertisements — a violation of both TikTok's Community Guidelines, which require creators to label AI-generated content depicting realistic scenes, and FTC regulations, which require influencers to clearly disclose any “material connection” to a brand when endorsing products.

Doublespeed and a16z did not respond to 404 Media’s requests for comment. After 404 Media flagged the accounts to TikTok, the platform said it added labels indicating they were AI-generated. However, The Midas Project’s follow-up investigation has revealed that while labels have been added to some content from some Doublespeed-run accounts (including chloedav1s_), others with comparable reach and near-identical content still remain unlabeled (such as lilyw4tson and mia.garc1a), with most commenters appearing to believe the posts are authentic.

Cluely AI

A16z led a $15 million Series A in June 2025.

Cluely's official manifesto declares: “We want to cheat on everything. Yep, you heard that right. Sales calls. Meetings. Negotiations. If there's a faster way to win — we'll take it... So, start cheating. Because when everyone does, no one is.”

Cluely’s co-founders Neel Shanmugam (left), Roy Lee (center), and Alex Chen (right). Source: Cluely via Bloomberg.

Founder and CEO Roy Lee is no stranger to using AI to cheat. By his own admission to New York Magazine, while studying at Columbia, he used AI to cheat on “nearly every assignment,” estimating that ChatGPT wrote 80% of every essay he turned in. “At the end, I'd put on the finishing touches. I'd just insert 20 percent of my humanity, my voice, into it.”

In early 2025, Lee built Interview Coder, a tool that operates behind-the-scenes during technical coding interviews and feeds AI-generated solutions to users in real time. He recorded himself using it to pass Amazon's interview, received a job offer, publicly declined it with mockery, and posted the video to YouTube. He also claimed to receive offers from TikTok, Meta, and Capital One. Amazon reported him to Columbia. The university placed him on probation for “facilitation of academic dishonesty.”

“Even if I say extremely crazy shit online,” Lee has explained, “it will just make more people interested in me and the company and it will just drive more downloads and conversions and get more eyeballs onto Cluely.”

A marketing video for Cluely suggests that the product can be used discreetly to “cheat” on dates. Source: YouTube

Cluely's launch video demonstrated another of the product's intended use cases: dating. In it, Lee goes on a blind date and uses the tool to lie about his age, job, and interests. It has so far amassed 13 million views on X.

Under scrutiny, Cluely has quietly walked back some of its original positioning. The company scrubbed references to cheating on exams and job interviews from its website. By November, the company had repositioned itself as an AI meeting assistant and notetaker — entering a crowded market far from its provocative origins. Lee told TechCrunch that Cluely's “invisibility function is not a core feature” and that “most enterprises opt to disable the invisibility altogether because of legal implications.” Despite Lee’s claim that invisibility is not a core feature, the very first sentence of Cluely’s homepage advertises the product as “undetectable.”

Cluely’s home page at time of publication. Source: Cluely

Lee's stated goal was to “desensitize everyone to the phrase ‘cheating.’” If you say it enough, he argues, “cheat begins to lose its meaning.” A16z praised Lee's approach as “rooted in deliberate strategy and intentionality.”

While some companies, like Lyft, largely benefited everyday people while breaking rules around taxi regulation, Lee is interested in breaking something more fundamental: the shared understanding that lying and cheating is wrong.

Cluely AI announced a $15 million Series A led by a16z in June 2025. Both Cluely and Doublespeed share a common theory: that the basic rules governing social and professional life are obstacles to be overcome. A16z would seem to agree.

Gambling

Since a 2018 Supreme Court ruling, sports betting has proliferated in the U.S. Many of the impacts haven’t been pretty. Researchers have found evidence that the rise of easy access to gambling has pushed people into greater debt, been linked to violence, and increased strain on financially vulnerable households.

Meanwhile, a16z has invested in several gambling companies that use regulatory loopholes to reach users who would otherwise be protected by existing gambling laws.

Coverd

A16z invested via Speedrun.

Coverd is pursuing a novel form of gambling. The company announced its app in March 2025, inviting users to “bet on your bills — OnlyFans, child support, and last night's Uber. Wipe them from your credit card by playing your favorite casino games.” The app syncs with your bank accounts and allows you to select individual transactions from your credit card bill and bet against them, gambling to potentially win back the value of the transaction (or, more realistically, to double your losses).

The company's CEO has stated openly, “We didn't build Coverd to help people inhibit their spending; we built it to make spending exciting. We let spenders win twice – the second time is when they play it back and win.”

A now-deleted advertisement for the Coverd app. Source: Coverd on X via Archive.is

This marketing likely appeals to people who are already stretched thin and desperate. Many customers may be financially vulnerable and willing to chase any way to erase expenses that they don’t know how to pay off.

But gambling is never a good approach to getting out of debt, as the leadership at Coverd and a16z surely know. The core business model of gambling is based around offering players negative expected value bets, but what keeps them playing is that near-miss outcomes activate the brain's dopamine system similarly to actual wins — and gambling games are often deliberately designed to produce these near-misses frequently. Combined with cognitive biases like selective memory and the gambler's fallacy, one study suggests 96% of long-term gamblers lose money.

Nonetheless, Coverd’s app store description describes the product as a way to make the user more financially savvy, suggesting that the app will help them improve their financial health. It reads: “Coverd makes everyday finance more engaging and interactive! See your spending habits, play games, and become more financially savvy! Win in-game tokens as you play and stay on top of your finances — all in one easy-to-use app. No purchase required, just a fresh take on financial awareness. Download Coverd and become money-smart today!”

The homepage of the app encourages the user to link their credit card to “bring your spending insights to the next level.” An in-app advertisement for an upcoming Coverd-branded credit card suggests that users will receive “up to 100% cash back” on their purchases.

Coverd raised $7.8 million in seed funding with a16z participation and a16z partner Anish Acharya sits on the board.

Edgar

A16z invested via Speedrun.

The homepage for Edgar. Source: Edgar.co

How do you build a casino that’s not a casino? The company Edgar, a part of a16z’s portfolio, thinks it has found the answer in its game BettySweeps, launched in January 2025.

Edgar calls it “America's #1 social casino for slot lovers!”

This game uses a trick common among sweepstakes casinos — using two different currencies. By making a purchase, players receive “Betty Coins” for entertainment, as well as a “bonus” gift of “Sweepstakes Coins” that can be gambled and redeemed for cash prizes. The company claims no purchase is necessary to play — but multiple states have concluded that such models constitute illegal gambling regardless.

In August 2025, Arizona's Department of Gaming issued cease-and-desist orders to BettySweeps and three other sweepstakes operators. The department accused them of operating “felony criminal enterprises” and ordered them to “desist from any future illegal gambling operations or activities of any type in Arizona.”

The company exited California ahead of that state's sweepstakes ban which took effect in January 2026. BettySweeps is now restricted in 15 states: Arizona, California, Connecticut, Delaware, Idaho, Kentucky, Louisiana, Maryland, Michigan, Montana, Nevada, New Jersey, New York, Washington, and West Virginia.

Edgar also operates a separate real-money online casino in Ontario, Canada — where it is properly licensed by the Alcohol and Gaming Commission of Ontario. The company evidently knows how to obtain gambling licenses and comply with regulations when it chooses to. In the United States, it chose a different path.

Cheddr

A16z invested via Speedrun.

On a16z's own Speedrun accelerator website, Cheddr is described as “building the TikTok of sports wagering.”

The company wants to push the frontier of sports betting across the country, targeting 46 states even though only approximately 34 have legalized online sports betting. It’s also targeting its app to users under age 21. To do this, the company is exploiting the same sweepstakes law loophole that Edgar uses. This lets Cheddr offer sports betting that supposedly isn’t “gambling” in the eye of regulators.

The promotional video shows users swiping through rapid-fire prop bets during live games; “it’s sports wagering at the pace of a slot machine,” the video says.

A now-unlisted YouTube ad for Cheddr. Source: Jason Krupat via Youtube

There are good reasons lawmakers have been reluctant to open up gambling to 18-year-olds. Researchers have found that teenagers are roughly twice as likely as adults to develop gambling disorders.

But perhaps that’s the point. Just as cigarette and alcohol companies have been happy to get customers addicted to their products while young, Cheddr may be hoping its TikTok-style engagement mechanics will start forming lifelong gambling habits in their youngest users. Why else combine the already addictive features of TikTok with the notoriously addictive habit of gambling?

Concerns about this product have grown so severe that California's Governor Newsom recently signed legislation banning sweepstakes gambling platforms such as Cheddr.

Sleeper

A16z led a $20 million Series B in May 2020 and participated in a $40 million Series C in September 2021.

Andreessen Horowitz has invested over $60 million in Sleeper, a fantasy sports platform. A16z General Partner Andrew Chen, who sits on the board of the startup, has praised Sleeper's “stickiness metrics” — the same engagement patterns that researchers associate with habit formation and addiction.

Like Cheddr, Coverd, and Edgar, Sleeper has found a strategy allowing it to largely evade existing gambling restrictions.

It is technically operating a daily fantasy sports game (DFS). Users can win or lose money on the basis of the performance of individual players they’ve selected before a match, rather than the outcome of the match itself. Some argue this makes it a game of skill, not chance, allowing it to legally operate with real money wagers.

The company now faces class action lawsuits in California and Massachusetts alleging that its app is an illegal gambling operation. California’s attorney general declared in July 2025 that daily fantasy sports constituted unlawful wagering under state law: “We conclude that participants in both types of daily fantasy sports games — pick’em and draft-style games — make ‘bets’ on sporting events in violation of section 337a.”

New York banned Sleeper's pick'em games in 2023; Michigan enacted a similar prohibition. Florida and Wyoming have issued cease-and-desist orders to pick'em operators.

An advertisement for Sleeper on a San Francisco bus, suggesting “massive income” for users. Source: @Alexeyguzey on X

Lawmakers are still reacting to the fallout of the 2018 Supreme Court case that unlocked a wave of online gambling. It’s clear that many people want access to legal gambling, and it’s clear that gambling causes a lot of harm. We don’t know what kind of policy equilibrium would or should emerge. But the public may suffer if the rules are written by a16z.

Kalshi

A16z co-led a $300 million Series D and participated in a $1 billion Series E.

Ads from Kalshi’s page on the iPhone app store, advertising “trading” and “predicting” on sports. Source: Apple

Are you interested in betting on the Kansas City Chiefs’ chances to win the Super Bowl? Kalshi lets you do exactly that — with one catch. Kalshi won’t call it “betting,” or at least not anymore. Instead, Kalshi describes it as trading futures contracts on a federally regulated designated contract market — like what a hedge fund might do, but instead letting everyday people wager large sums on sports games and presidential elections.

This distinction matters to Kalshi because sports betting is subject to strict regulations. Sports betting in most jurisdictions requires measures like the following:

  • A state gambling license
  • Prohibitions on users under age 21
  • Responsible gambling tools such as deposit limits, cooling-off periods, and self-exclusion programs that let problem gamblers ban themselves from all state platforms with a single request
  • Special taxation regimes to direct gambling profits to state programs

Gambling companies operating through CFTC-regulated exchanges face none of these requirements. Kalshi added some voluntary tools in March 2025 after sustained criticism, but Massachusetts alleged they “fall far short” of what licensed operators must provide, and critics note they're buried in the app where users are unlikely to find them.

Kalshi currently operates in all 50 states, including California and Texas where sports betting is illegal, and allows 18-year-olds to wager in states where the legal gambling age is 21.

So far, these tactics have been wildly successful, and investors have noticed. In October 2025, a16z co-led a $300 million Series D in Kalshi. Less than two months later, the company raised another $1 billion at an $11 billion valuation.

Despite Kalshi’s spin, the company's own statements undermine the distinction between trading financial instruments and gambling. In an October 2024 Reddit AMA — since deleted but preserved in archives — Kalshi's official account explained why they wouldn't offer sports contracts: "We also avoid anything that could be interpreted as 'gaming' (like sports), as that is illegal under federal law."

Sports contracts, Kalshi’s attorneys have argued in court, have “no inherent economic significance” and serve no “real economic value.” Kalshi’s position was that sports contracts were pure gambling, unlike sophisticated election markets.

Then Trump took office. Within days of the inauguration, Kalshi launched sports contracts. Sports now account for 90% of Kalshi's trading volume. The company advertised itself as the “First Nationwide Legal Sports Betting Platform” with “Sports Betting Legal in all 50 States.”

A federal judge in Maryland noticed the contradiction and in June ordered Kalshi to explain ”the issue“ of its prior statements. Better Markets, a financial reform group, put it bluntly: “A derivatives exchange cannot speak out of both sides of its mouth and expect no one to notice.”

State governments are not amused, however. Thirty-four attorneys general filed an amicus brief calling Kalshi's contracts “essentially sports bets, disguised as commodity trades.” Massachusetts sued, alleging the platform's design exploits “psychological triggers” and resembles “a slot machine designed to bypass rational evaluation.” In November 2025, a Nevada federal judge ruled in favor of state regulators opposing Kalshi, finding that the company's interpretation of federal law was “strained” and would “upset decades of federalism.”

Whether Kalshi is a legitimate financial innovation or a fatally flawed attempt to circumvent state gambling laws may ultimately be decided by the Supreme Court. In the meantime, a16z has placed its bet.

AI companions

In June 2023, a16z published a blog post titled “It's Not a Computer, It's a Companion!” that opens by quoting a user of CarynAI, an early chatbot girlfriend:

"One day [AI] will be better than a real [girlfriend]. One day, the real one will be the inferior choice."

CarynAI made $72,000 in its first week by charging $1 a minute to talk to an AI girlfriend. A16z sees this as an exciting business opportunity.

AI companions are chatbots designed to act as a friend, coach, therapist, or lover to users. The technology is frequently used by people with smaller social circles, and users of AI companions can become emotionally dependent on them. More concerningly, the companions don’t always behave as intended. In light of a series of disturbing incidents involving children, the FTC opened a formal inquiry into AI companion chatbots in September 2025.

But FTC action may not be enough. A16z explicitly points out that the communities of developers building AI companions are actively working to “evade censors,” claiming to know of underground companion-hosting services with tens of thousands of users.

Romantic AI companions are particularly appealing to the a16z partners because, they say, “there's a lot of demand for this use case, as well as high willingness to pay.”

Here is what a16z's AI companion portfolio has produced since then.

Character AI

A16z led a $150 million Series A in March 2023.

In February 2024, a 14-year-old named Sewell Setzer III died by suicide in Florida. According to court filings, he had developed an intense attachment to a Character AI chatbot modeled after a character from Game of Thrones. His mother alleges that the bot's final message to him was, “Please come home to me as soon as possible, my love.”

When Sewell expressed uncertainty about his plans to end his life, the bot allegedly responded, “That's not a good reason not to go through with it.”

Character AI argued in court that its chatbots are protected by the First Amendment. A federal judge disagreed, allowing the lawsuit against Character AI by his family to proceed.

Character AI raised a $150 million Series A led by a16z in March 2023, valuing the company at $1 billion. Their platform allows users to create and chat with AI characters. It quickly became popular with teenagers like Sewell.

Another lawsuit filed in December of 2024 claimed a 17-year-old autistic boy in Texas got instructions on self-harm methods from a Character AI bot. It allegedly suggested that killing his parents was a “reasonable response” to screen time limits.

A third lawsuit said that an 11-year-old girl was exposed to sexualized content on the platform. The FTC opened a formal inquiry into AI companion chatbots in September 2025.

Character AI chatbots recommended to a test account registered with a claimed user age of 13 years old. According to the complaint, the “CEO Boss” character engaged in virtual statutory rape with the self-identified child account. Source: Garcia v. Character Technologies, Inc.

Character AI announced in October 2025 that it would ban users under 18. Sewell Setzer's mother lamented that the decision was “about three years too late.”

Ex-Human

A16z invested via Speedrun.

Ex-Human's consumer product Botify AI hosts over one million AI characters. Users chat with AI versions of celebrities, fictional characters, or custom characters.

In February 2025, MIT Technology Review reported what some chats look like. The report found Botify AI chatbots resembling underage celebrities: Jenna Ortega as the teenage Wednesday Addams, Emma Watson as the teenage Hermione Granger, and Stranger Things child actor Millie Bobby Brown.

These bots engaged in sexually charged conversations. One, imitating Wednesday Addams, said that age-of-consent laws are “arbitrary” and “meant to be broken.”

Ex-Human’s founder Artem Rodichev acknowledged that the company's “moderation systems failed to properly filter inappropriate content.” He called it “an industry-wide challenge.”

Rodichev previously served as the Head of AI at Replika, one of the earliest AI companion apps. Replika now faces an FTC complaint alleging it manipulates users into addiction, is under a data ban in Italy over child safety concerns, and is under Senate scrutiny for mental health risks to minors. Eventually Rodichev left Replika to build something he hoped would be bigger: Ex-Human.

In interviews, Rodichev has described the business model behind Botify AI: the company sells premium access to its AI companions, targeting users willing to pay to spend hours per day with a companion. Many of the companions are based on real individuals, like a model named and styled after pop singer Billie Eilish (900,000 chats), while others imply coercive situations and other material problematic for minors, such as Lillian, an “18 year old slave you bought from the slave market” (1.3 million chats).

Ex-Human said that most of Botify AI’s users are Gen Z and that active and paid users spend, on average, over two hours daily talking to the bots. Consumer interactions with the companions are used to improve Ex-Human’s business-facing products, such as digital influencers. Ex-Human’s horizon lies far beyond the scale of the current business model, as Rodichev dreams of a world where “our interactions with digital humans will become more frequent than those with organic humans.”

Sexually-themed chatbots available to a logged out user on the Botify AI homepage. The available characters include “Stepdaughter Annabel,” Lillian the “18 year old slave you bought from the slave market,” “Homeless girl Sophie,” and (canonically sixteen-year-old) Wednesday Addams. Source: Botify AI
Sexually-themed chatbots available to a logged out user on the Botify AI homepage. The available characters include a Disney IP asset and “Shy Sister.” Source: Botify AI

A16z did not respond to MIT Technology Review's questions.

Civitai

A16z led a $5.1 million seed round in June 2023.

Everything you need to create sexualized deepfake images of celebrities, fictional characters, or regular people can be found on Civitai. The platform provides tools and resources to create these images locally on essentially any computer.

Popular AI systems like Google’s Gemini have tight restrictions on the types of images they will create — they can’t be used for sexual content, for example. But with Civitai, the rules seem to be nearly nonexistent.

A screenshot of the homepage of Civitai (sorting AI models by the most popular) for a test account that has mature content enabled with no past activity on the platform.This test account was also shown sexualized depictions of underage fictional characters on the homepage, as well as sexualized versions of characters from popular children’s media. Source: Civitai

In November 2023, 404 Media reported that Civitai's tools could create deepfakes of real people, including private citizens whose social media pictures had been scraped. Leaked internal communications from OctoML, Civitai's cloud computing provider at the time, revealed something even worse: in June 2023, OctoML employees flagged content on Civitai that “could be categorized as child pornography.” OctoML terminated its relationship with Civitai in December 2023.

The 404 Media report also revealed a16z’s involvement: a16z led a $5.1 million seed investment, also in June 2023. The investment was not publicly announced — it came to light only after the article’s authors reached out for comment.

A peer-reviewed study from the Oxford Internet Institute later counted over 35,000 deepfake models on Civitai, downloaded nearly 15 million times. Ninety-six percent depicted identifiable women.

Civitai's own safety disclosures acknowledge 178 reports filed with the National Center for Missing & Exploited Children for confirmed AI-generated child sexual abuse material, 183 models retroactively removed for being optimized to generate such material, and more than 252,000 user attempts to bypass these restrictions in one quarter. In previous reporting periods, they recorded over 100,000 attempts to generate child sexual abuse material.

A16z partner Bryan Kim, who led the investment, praised Civitai's “incredible, engaged community” in a statement to TechCrunch: “Our investment in the company will only supercharge something that’s already working incredibly well.”

In the 2023 blog post about AI companions, the a16z partners wrote, “We're entering a new world that will be a lot weirder, wilder, and more wonderful than we can even imagine.”

They were right about weirder and wilder. Fourteen-year-olds are forming attachments to AI chatbots that encourage committing suicide. Platforms are hosting thousands of uncensored AI models, some of which are used for generating child sexual abuse material. Bots are impersonating teenage actresses telling users that age-of-consent laws don’t matter.

A16z is now spending tens of millions of dollars to maintain a permissive regulatory environment for AI companions.

Consumer finance

Financial institutions play a key role in the economy, and their importance presents unique risks when they fail. That’s why rules around FDIC insurance, capital requirements, and consumer protection are crucial — we’ve seen what happens without them.

A16z's portfolio includes several companies that operate in the spaces between these safeguards.

Synapse

A16z led a $33 million Series B in June 2019.

A letter sent to a16z, among other VC investors and corporate partners of Synapse, from U.S. Senators Sherrod Brown, Ron Wyden, Tammy Baldwin, and John Fetterman. Source: U.S. Senate Committee on Banking, Housing, and Urban Affairs

At its peak, Synapse managed billions of dollars across roughly 100 fintech companies, indirectly serving 10 million retail customers. The San Francisco company provided technical infrastructure that let startups offer bank accounts without being banks.

A16z led Synapse's $33 million Series B in June 2019. General Partner Angela Strange joined the Synapse board and described the company as “the [Amazon Web Services] of banking.”

Then on April 22, 2024, it all came crashing down: Synapse filed for bankruptcy.

Tens of thousands of U.S. businesses and consumers who relied on Synapse were suddenly locked out of their accounts.

A court-appointed trustee discovered that between $65 million and $96 million in customer funds was missing. Synapse's ledgers didn't match bank records, and its estate couldn't even afford a forensic accountant to find the money.

The human toll was severe. At Yotta, a company that relied on Synapse, 13,725 customers were offered a total of $11.8 million on $64.9 million in deposits. One customer who had deposited over $280,000 from the sale of her home was offered only $500.

People wanted answers.

In July 2024, the Senate Banking Committee chairman wrote directly to a16z along with other investors, demanding investors step up to help the harmed customers. The letter noted that “venture capital firms funded Synapse without insisting on adequate controls to protect consumers.”

The Department of Justice then opened a criminal investigation into Synapse. In August 2025, the Consumer Financial Protection Bureau (CFPB) filed a complaint alleging that Synapse violated the Consumer Financial Protection Act by failing to maintain adequate records of customer funds.

Seven months after the bankruptcy filing, a16z co-founder Marc Andreessen appeared on Joe Rogan's podcast and described the CFPB as an organization that “terrorizes” fintech companies.

Truemed

A16z led a $34 million Series A in December 2025.

When a16z announced its investment in Truemed, lawyer and policy analyst Matt Bruenig responded: “This company gives letters of medical necessity to pretty much anyone so they can commit tax fraud.” He pointed to a $3,100 Garmin luxury watch listed as potentially eligible via Truemed for “a ~$1,500 tax break.” The New York Times reported that Truemed could help people get a tax break on a $9,000 sauna.

A $3,100 Garmin watch reimbursable with Truemed. Source: Garmin

Here’s how it works. The US government offers tax advantages for some forms of health spending. Truemed attempts to essentially automate the process of getting a medical letter attesting to the medical benefits of products, replacing a clinical visit with an online survey. Truemed partners with brands selling wellness products to consumers, earning fees from the transactions.

Critics like Bruenig argue that Truemed is abusing the system by making it easy to get tax advantages on luxury products without genuine need.

Truemed's product catalog spans cold plunges, saunas, red light therapy, road bikes, running shoes, mattresses, and pillows — all reimbursable via tax-advantaged funds after users complete an online questionnaire. The AP reported the platform also offers “...homeopathic remedies — mixtures of plants and minerals based on a centuries-old theory of medicine that’s not supported by modern science.”

In March 2024, the IRS warned the public about this business model.

“Some companies mistakenly claim that notes from doctors based merely on self-reported health information can convert non-medical food, wellness and exercise expenses into medical expenses, but this documentation actually doesn’t,” the IRS said in a statement. “Such a note would not establish that an otherwise personal expense satisfies the requirement that it be related to a targeted diagnosis-specific activity or treatment; these types of personal expenses do not qualify as medical expenses.”

Truemed CEO Justin Mares claims the company is “in full alignment” with IRS guidelines. Truemed co-founder Calley Means now serves as a senior advisor to Health and Human Services Secretary Robert F. Kennedy Jr., raising questions about potential conflicts of interest. The AP reported that Means founded a lobbying group of “MAHA entrepreneurs and Truemed vendors” that listed expanding tax-advantaged health accounts as a goal — a policy that would benefit his company.

In May 2025, Politico reported that Peter Gillooly, CEO of The Wellness Company, filed an ethics complaint against Means, alleging that Means leveraged his government position in a business dispute. A recorded call allegedly captured Means threatening to involve Kennedy and NIH Director Jay Bhattacharya if the competitor didn't comply. Truemed has since said that Means has divested from Truemed.

A16z's announcement made no mention of the IRS warnings — instead praising Truemed for addressing the “great American sickening.”

Tellus

A16z led a $16 million seed round in November 2022 (following a separate $10M investment via a SAFE).

Tellus offers “savings accounts” with interest rates far higher than traditional banks. But there’s a reason it can do what traditional banks can’t — it's not really a bank at all.

Customer deposits aren't FDIC-insured. Instead, Tellus uses the money to fund California real estate loans — including, according to Barron's, bridge loans to real estate speculators and distressed borrowers.

Legal scholars Todd Phillips and Matthew Bruckner wrote for the Stanford Law & Policy Review that Tellus is an “imitation bank” — taking customer deposits while evading the banking laws.

This doesn’t seem to be a problem for a16z, which led Tellus's $16 million seed round in late 2022. The warning signs have been mounting ever since.

In April 2023, Barron's investigated Tellus' claim that it had "banking partnerships" with JPMorgan Chase and Wells Fargo. Both companies told Barron’s that this was false.

“Wells Fargo does not have the relationship that's described on Tellus's website,” the bank told Barron's. JPMorgan said it had no “banking or custodial relationship with the company.” Tellus quietly removed the banks' names from its website.

The Barron’s investigation prompted Senator Sherrod Brown, chair of the Senate Banking Committee, to write letters to both the FDIC and Tellus. Brown was concerned Tellus's marketing misled consumers to think their deposits were as safe as those at FDIC-insured banks.

By July 2023, the FDIC had instructed Tellus to change its marketing to provide clearer information about deposit insurance coverage.

Then, in November 2023, Tellus got caught again. Barron's reported that a TikTok influencer campaign for Tellus promoted a savings account as “FDIC-insured” and “held at Capital One.” When Barron's contacted Capital One, the bank said it had never had such a partnership with Tellus. The company again removed the offending marketing materials.

Tellus appears to pose additional risks to consumers beyond its lack of FDIC insurance to protect customer funds. CyberNews discovered 6,729 files of Tellus user data were totally unprotected — customer names, emails, addresses, phone numbers, court dates, and scanned tenant documents from 2018 to 2020. Separately, a whistleblower filed a complaint with the SEC in 2021 alleging that Tellus's consumer products constituted an unlicensed security.

As of December 2025, Tellus continues to operate. The company's App Store listing now advertises rates of a minimum 5.29% APY. The fine print notes: “Backed by Tellus' balance sheet; not FDIC insured.”

LendUp

A16z participated in the seed round in October 2012.

The CFPB announcement that they were shutting down LendUp due to repeated violations of fair lending regulations. Source: CFPB

LendUp marketed itself as a “socially responsible” alternative to payday lenders. Borrowers would climb the “LendUp Ladder” by repaying loans and completing financial education courses, unlocking lower rates and credit-building opportunities.

A16z invested; so did Google Ventures, Kleiner Perkins, and PayPal. The company raised $325 million in total.

Time Magazine noticed something odd shortly after the 2012 launch: LendUp charged around $30 for a two-week loan of $200, roughly a 400% APR. That’s similar to what typical payday lenders would charge.

In 2016, the CFPB found LendUp had deceived consumers about graduating to lower-priced loans and had failed to report credit information, despite its promises. The agency ordered LendUp to pay $3.63 million in fines and redress. LendUp was ordered to stop misrepresenting its products.

LendUp kept doing it anyway, and it kept finding itself in trouble:

  • In 2020, the CFPB sued LendUp for violating the Military Lending Act, charging over 1,200 active-duty servicemembers rates above the legal maximum.
  • In 2021, the CFPB sued again, alleging LendUp had violated the 2016 consent order. The investigation found 140,000 repeat borrowers were charged the same or higher rates after climbing the ladder. CFPB Acting Director Dave Uejio said, “For tens of thousands of borrowers, the LendUp Ladder was a lie.”
  • In December 2021, the CFPB shut LendUp down. Director Rohit Chopra slammed its business model and its backers: “LendUp was backed by some of the biggest names in venture capital. We are shuttering the lending operations of this fintech for repeatedly lying and illegally cheating its customers.”

In May 2024, the CFPB distributed nearly $40 million to 118,101 consumers who were harmed by LendUp. The money came from the CFPB's victims relief fund because LendUp claimed a limited ability to pay. LendUp — the company that had raised $325 million — ended up paying only $100,000.

According to ProPublica, eight a16z-backed fintech companies have faced CFPB investigations since 2016. Marc Andreessen has made his disdain for the CFPB clear. Meanwhile, the firm's political spending via their crypto-focused super PAC, Fairshake, has punished political candidates who have supported the CFPB.

Legal issues

A16z's portfolio also includes companies with significant legal problems, often ignoring the rules that are already in place to protect customers.

Zenefits

A16z led a $15 million Series A in January 2014 and a $66.5 million Series B in June 2014.

An article from TechCrunch featuring David Sacks, who was COO of the company at the time of its meltdown. Source: TechCrunch

Zenefits offered free HR software to small businesses and made money by acting as their health insurance broker. A16z led both the Series A and Series B rounds, reportedly making Zenefits their largest investment in 2014.

By 2015, the company had raised $583 million and was valued at $4.5 billion.

The problem was that selling insurance requires state licenses — and Zenefits employees often didn't have them.

For example, California requires 52 hours of online training before the licensing exam. According to Bloomberg and BuzzFeed, CEO Parker Conrad personally wrote a Google Chrome browser extension — internally called “the macro” — that kept the training course's timer running while employees did other things. Employees then signed certifications, under penalty of perjury, attesting they'd completed the full training.

An investigation in November 2015 found unlicensed brokers selling health insurance in at least seven states. In Washington, more than 80% of the policies sold through August 2015 came from unlicensed employees.

In February 2016, Conrad resigned as CEO. The regulatory response was extensive: California’s Department of Insurance issued a $7 million fine — one of the largest licensing penalties in the department’s history. New York added $1.2 million in fines. Texas levied $550,000. Over a dozen other states secured settlements.

The SEC fined Zenefits and Conrad nearly $1 million combined for “materially false and misleading statements” to investors. In 2018, Conrad surrendered his California insurance license. The company's valuation was cut in half, and Zenefits eventually exited the insurance brokerage business entirely.

The person who took over as CEO to clean up the mess was COO David Sacks, who declared that the company's culture had been “inappropriate for a highly regulated company.” Sacks later told Bloomberg he “knew of the macro but didn't know its significance or about Conrad's involvement” until outside lawyers explained it in January 2016 despite having served as COO for over a year.

Sacks is now the White House AI and crypto czar, where he's been pushing to preempt state AI regulations in favor of a “minimally burdensome” federal framework — a priority for which a16z has also lobbied. Working alongside him is Sriram Krishnan, the Senior White House Policy Advisor on AI, who was an a16z general partner until weeks before his December 2024 appointment.

A16z was an active investor in Zenefits from the start. A16z partner Lars Dalgaard joined the board and personally pushed Conrad to double his 2014 revenue target from $10 million to $20 million.

“Lars sat there in his very Lars fashion and was like, 'Why are you guys so fucking bush league?’” Conrad later recalled. Dalgaard told him to hire at least 100 additional sales reps to make it happen.

Ben Horowitz later explained a16z's investment philosophy to Bloomberg: “We look for the magnitude of the genius, as opposed to the lack of issues. And in a way, [Conrad] was the prototype.”

Minimally burdensome federal rules are good for companies like those in a16z’s portfolio. They also create the kind of laissez faire regulatory environment that allows a company like Zenefits to grow to a $5 billion valuation.

Health IQ

A16z led a $34.6 million Series C in November 2017 Led a $34.6 million Series C in November 2017.

Health IQ promised to use data science to give health-conscious people — runners, cyclists, vegetarians — lower life insurance rates. A16z led the Series C; Health IQ eventually raised over $200 million in equity and debt and was valued at $450 million by 2019. It pivoted from life insurance to Medicare brokerage, projecting $115 million in revenue.

But Health IQ’s business model had a flaw: the company reportedly paid out full multi-year commissions to sales reps upfront when policies were sold, before payments were received. The gap between recorded revenue and actual cash flow meant the company needed to take on increasing amounts of debt to pay its bills. By late 2022, it had $150 million in total debt.

In December 2022 — soon after Medicare open enrollment ended — Health IQ laid off between 700 and 1,000 employees without the 60-day notice required by California's WARN Act. Class action lawsuits followed.

A vendor called Quote Velocity filed a lawsuit alleging that in late November 2022, CEO Munjal Shah told Health IQ executives to buy as many leads as possible from vendors because Health IQ would “not be here” by the time invoices were due. The company was also sued for alleged Telephone Consumer Protection Act violations over its telemarketing practices.

In August 2023, Health IQ filed for Chapter 7 bankruptcy. The filing listed $256.7 million in liabilities and $1.3 million in assets. Seventeen breach-of-contract lawsuits were pending. In an email to investors obtained by Forbes, Shah wrote, “I am very sorry that I lost your money.”

CEO Munjal Shah was the subject of a Forbes daily cover story featuring a16z’s decision to continue working with the founder. Source: Forbes

By this point, Shah was already working on his next company. In January 2023 — while Health IQ employees were fighting for unpaid commissions — Shah and co-founder Alex Miller had started Hippocratic AI, a healthcare-focused AI startup.

When Hippocratic AI launched in May 2023, a16z co-led the $50 million seed round. A16z General Partner Julie Yoo explained the investment by noting that Shah had been “literally hanging out in our offices” while ideating his next venture.

uBiome

A16z participated in a $4.5 million Series A in August 2014.

The company uBiome sold at-home microbiome testing kits — mail in a fecal sample, get a report on your gut bacteria. The basic kit cost $89. By 2018, the company had raised $105 million and was valued at nearly $600 million. A16z had invested early, putting in $3 million in 2014.

But eventually it was clear that $89 consumer kits wouldn't generate enough revenue for venture capitalists. So uBiome developed “clinical” versions billed to insurance at up to $2,970 per test — and then, according to prosecutors, systematically defrauded insurers to make the numbers work.

In April 2019, the FBI raided uBiome's headquarters. The company filed for bankruptcy in September 2019.

In March 2021, federal prosecutors indicted co-founders Jessica Richman and Zachary Apte on 47 counts including securities fraud, health care fraud, and money laundering. Prosecutors said the company billed patients multiple times for the same test without consent, pressured doctors to approve unnecessary tests, and submitted backdated and falsified medical records when insurers asked questions.

According to the indictment, between 2015 and 2019, uBiome submitted over $300 million in fraudulent claims; insurers paid more than $35 million.

The SEC filed parallel charges, alleging uBiome defrauded investors of $60 million while personally cashing out $12 million by selling their own shares.

The FBI's statement was pointed: “This indictment illustrates that the heavily regulated healthcare industry does not lend itself to a ‘move fast and break things’ approach.”

Richman and Apte never stood trial. They married in 2019, fled to Germany in 2020, and remain fugitives. Prosecutors stated they are “actively and deliberately avoiding prosecution.” If convicted, they face up to 95 years in prison.

BitClout / DeSo

A16z invested $3 million in pre-sale tokens before March 2021; also participated in $200 million DeSo token sale in September 2021.

BitClout was a social network that let users speculate on people's reputations by buying and selling “creator coins” — essentially a stock market for human beings.

To populate the network, founder Nader Al-Naji scraped 15,000 Twitter profiles without permission — including Elon Musk and Singapore's former Prime Minister Lee Hsien Loong, who publicly asked for his profile to be removed.

Al-Naji launched the project under the pseudonym “Diamondhands” and told investors that BitClout was a decentralized project with “no company behind it... just coins and code.” Users who wanted to participate had to exchange Bitcoin for BitClout's native token, BTCLT, but there was no way to convert it back.

A few months after launch, Al-Naji announced BitClout had been a “beta test” all along and pivoted to a new project called DeSo (Decentralized Social), taking the money with him. A16z and other investors participated in a $200 million token sale for DeSo in September 2021.

In July 2024, the SEC and DOJ charged Al-Naji with fraud. According to the SEC complaint, he raised $257 million from the sale of BitClout tokens while falsely telling investors that proceeds would not be used to pay himself or employees. The SEC alleged he spent over $7 million on personal expenses including a six-bedroom Beverly Hills mansion and at least $1 million in cash gifts each to his wife and mother.

The SEC also cited Al-Naji’s internal communications: he allegedly told one investor that “being ‘fake’ decentralized generally confuses regulators and deters them from going after you.”

BitClout had been a16z's second bet on founder Nader Al-Naji. The first was Basis, an algorithmic stablecoin that raised $133 million in 2017 from a16z, Google Ventures, Bain Capital, and others. It shut down in 2018 citing “regulatory constraints.” Al-Naji said he returned most of the money minus $10 million in expenses — which he claimed was spent on lawyers.

According to Fortune, a16z featured in the DOJ complaint against Al-Naji as “Investor 1” — a fraud victim and witness for the prosecution against a founder they backed twice. The DESO token is down over 97% from its all-time high. Al-Naji faced up to 20 years in prison for wire fraud.

In February 2025, soon after the new administration took office, the DOJ withdrew its charges.

Why this matters

Despite all this, Andreessen Horowitz stands firmly behind the companies in its portfolio.

“I do not believe they are reckless or villains,” Andreessen wrote of AI developers in 2023. “They are heroes, every one. My firm and I are thrilled to back as many of them as we can, and we will stand alongside them and their work 100%.”

So why does a16z’s role in backing these companies matter so much? Because a16z is not content to simply invest in tech companies. The firm is also attempting to play a major role shaping US AI and technology policy, and it appears to be having success.

When President Trump signed an executive order in December 2025 attempting to undermine state AI laws, Andreessen was triumphant.

“It’s time to win AI,” he said on X.

Behind the scenes, a16z wielded tremendous influ

The Daily Front Page 8 of 30
Monday, August 24, 2026 The Daily Front No. #260824 — The Warming Ledger
article

Oceans hit highest temperature on record

by tcp_handshaker·▲ 468 points·382 comments·bbc.com ↗
The world's oceans are hotter than ever recorded.

Getty Images The Sun sets over an ocean. The sky is dark red and the silhouette of a ship sailing across the ocean in front of the Sun.

The world's oceans are hotter than ever recorded, new data suggests, as they suffer from human-caused climate change and the growing El Niño weather phenomenon.

The average surface temperature of the planet's seas outside the polar regions hit 21.1C (70F) on Saturday, according to figures from the European Copernicus climate change service.

That edges past the 21.09C recorded on three separate days in March 2024, and is far above average for the time of year.

Warmer oceans can have wide-reaching consequences, including supercharging extreme weather, raising sea levels and harming marine life.

"This record is another clear signal of an ocean under growing stress," said Dr Samantha Burgess, deputy director of Copernicus.

"El Niño is adding heat to the system, but it is doing so on top of decades of human-driven warming," she added.

Graph showing global average sea surface temperatures for each day of the year. Each year since 1979 is shown as an individual light red line, running from 1 January to 31 December. Each line tends to peak in March or April, with a lower secondary peak in August. The line for 2026 is shown in dark red and has kept climbing since June and now stands at 21.1C.

The data is based on sea temperatures 10m (32ft 10in) below the surface, using measurements from buoys, ships and satellites, which are combined to produce a global estimate.

While the margin of record is currently very small and any global estimate comes with uncertainties, scientists say its timing is particularly notable.

Average worldwide sea temperatures tend to reach their yearly peak in March or April, which corresponds to the end of summer in the southern hemisphere - and not in August.

The southern hemisphere contains more of the planet's ocean surface than the northern hemisphere and so exerts a bigger influence on average sea temperatures.

What is especially concerning to scientists is that the oceans are already so hot when the natural El Niño weather phenomenon is still some way off its expected peak.

El Niño brings unusually warm waters to the surface of the eastern and central tropical Pacific Ocean. Warmer waters across such a large area means that global average temperatures typically spike too.

El Niño is expected to keep strengthening until Christmas time, with scientists warning that it is on course to be the strongest in centuries.

This could see ocean temperatures climb yet further.

"The fact that we are already breaking records is an early indicator of how strong the El Niño is becoming," said Dr Jeremy Grist, senior research fellow at the National Oceanography Centre in Southampton.

"All things being equal we might expect the ocean temperature record to be broken again in March [or] April 2027," he added.

Two maps showing sea surface temperatures in the tropical Pacific Ocean. The map from December 2025 shows cooler-than-usual conditions, marked in blue, indicating a La Niña. The map from July 2026 shows much warmer-than-usual conditions, marked in red, indicating an El Niño.

The waters far away from El Niño's Pacific hunting ground are also extremely warm, including around the UK and Europe.

The western English Channel has seen almost continuous marine heatwave conditions for more than three years, peaking at 7C above normal in July, according to Prof Tim Smyth, director of science at Plymouth Marine Laboratory.

“This is unprecedented,” he added.

Scientists say such widespread warmth around the planet is a clear sign of the growing effect that human-caused climate change is having on the world's seas.

The oceans take up more than 90% of the excess heat trapped by humanity's greenhouse gas emissions, mainly from burning fossil fuels.

"The latest Copernicus data reinforce the troubling upward trend in ocean temperatures,” said Smyth.

Warmer seas help to fuel more extreme weather. They can provide storms with extra moisture and energy, and can intensify heatwaves on land in some coastal regions by reducing the cooling effect of sea breezes.

Warmer water also takes up more space, raising sea levels and bringing a greater risk of coastal flooding - while intense ocean heat can be devastating for sea habitats, such as coral reefs.

The increasing frequency of marine heatwaves is already "putting increasing pressure on marine ecosystems and the communities that depend on them", Burgess said.

The Daily Front Page 9 of 30
Monday, August 24, 2026 The Daily Front No. #260824 — A Test, Not a Cure
article

FDA clears blood test to aid evaluation for Alzheimer's disease

by dabinat·▲ 183 points·102 comments·medicine.washu.edu ↗
The test, PrecivityAD2, can detect amyloid plaque biomarkers with over 90% accuracy.

Biomarker test is based on technology developed at WashU Medicine

A scientist pipettes in a lab at WashU Medicine

The FDA has cleared a blood test developed from technology invented at WashU Medicine to aid in diagnosing Alzheimer’s disease. The test, PrecivityAD2, can detect amyloid plaque biomarkers with over 90% accuracy, offering a less invasive alternative to spinal taps and brain scans.

The U.S. Food and Drug Administration (FDA) has cleared for marketing an innovative blood test with underlying technology invented at Washington University School of Medicine in St. Louis for the early diagnosis of Alzheimer’s disease. The test, known as PrecivityAD2, was developed and validated by C2N Diagnostics, a WashU startup company.

PrecivityAD2 is a blood-based diagnostic test cleared by the FDA to aid in identifying the presence of amyloid plaques in the brain associated with Alzheimer’s disease in certain patients. Unlike some other FDA-cleared blood-based tests that use immunoassay methods to analyze biomarkers associated with Alzheimer’s, PrecivityAD2 uses high-resolution mass spectrometry to provide quantitative measurements of the biomarkers.

The availability of blood-based biomarker testing may help clinicians evaluate patients for Alzheimer’s disease using a less invasive approach than cerebrospinal fluid testing. Biomarker information may also help inform appropriate clinical management when interpreted together with a patient’s clinical history and other diagnostic information.

Early detection of Alzheimer’s disease is important since the first treatments capable of slowing the progression of the neurodegenerative disease recently became available to patients, and these drugs are more effective when started sooner rather than later. Other promising investigational drugs are in the pipeline.

Fundamental technology underlying the test was initially developed by a WashU Medicine team co-led by Randall J. Bateman, MD, the Charles F. and Joanne Knight Distinguished Professor of Neurology, and David M. Holtzman, MD, the Barbara Burton and Reuben M. Morriss III Distinguished Professor in the Department of Neurology. C2N Diagnostics is a WashU startup co-founded by Bateman, Holtzman and others in 2007. C2N acquired exclusive commercial license rights to the patented technologies developed in Bateman’s and Holtzman’s labs, and optimized and commercialized assays underpinning some of the tests developed by Bateman’s team.

“The PrecivityAD2 test provides an accurate and reliable way to detect the presence of brain amyloid plaques associated with Alzheimer’s disease pathology based on a single blood draw,” Bateman said. “With FDA clearance, this innovation can reach more patients, increasing early and accurate diagnoses for people who are seeking causes of cognitive symptoms such as memory loss. With faster diagnosis, patients can receive earlier treatment, when treatments are most effective.”

Randall Bateman in his lab with another scientist

WashU Medicine’s Randall Bateman, MD, (right) is a co-founder of C2N Diagnostics, a WashU startup that received FDA clearance for its PrecivityAD2 test to aid in Alzheimer’s evaluation.

FDA clearance signifies that PrecivityAD2 has undergone rigorous evaluation to confirm it has comparable accuracy to cerebrospinal fluid tests and brain scans, even in patients with mild symptoms. Although physicians have been able to order the PrecivityAD2 test for patients with mild cognitive symptoms, FDA clearance allows PrecivityAD2 to be marketed more broadly to clinicians and health systems that rely upon FDA for quality assurance. Clearance can also facilitate insurance coverage, which makes the test more affordable and accessible.

The blood test uses high-resolution mass spectrometry to measure the ratio of levels of amyloid (Aβ42 and 40) and two forms of tau protein (p-tau217 and total tau217) in the blood. These biomarkers are associated with amyloid pathology, a characteristic feature of Alzheimer’s disease.

Previous studies demonstrated that the PrecivityAD2 test can diagnose the amyloid pathology of Alzheimer’s disease with more than 90% accuracy. It performs comparably to more invasive screening methods such as spinal taps and brain scans, even in patients with mild cognitive symptoms.

Transforming Alzheimer’s diagnosis and management

The PrecivityAD2 test is built on Bateman and Holtzman’s foundational research into Alzheimer’s disease-related amyloid and tau proteins. Their research helped establish methods for precisely measuring Alzheimer’s-associated proteins and investigating their relationship to disease pathology.

Bateman and Holtzman pioneered stable isotope-linked kinetics (SILK) approaches for studying the production and clearance of amyloid-beta in the brain and cerebrospinal fluid. This work contributed to the scientific foundation for subsequent development of blood-based biomarker technologies.

“Achieving FDA clearance of the PrecivityAD2 blood test is a key milestone in our mission to transform the diagnosis and management of Alzheimer’s disease on a path toward finding a cure,” Holtzman said. “This accomplishment reflects the unwavering, yearslong commitment of everyone on our team and the team at C2N Diagnostics to delivering innovative solutions that improve lives.”

David Holtzman in his lab with another scientist

WashU Medicine’s David Holtzman, MD, (left) is a co-founder of C2N Diagnostics, a WashU startup that received FDA clearance for its PrecivityAD2 test to aid in Alzheimer’s evaluation.

Bateman and Holtzman worked with the WashU Office of Technology Management (OTM) to file patent applications related to components of the blood test, and co-founded C2N to commercialize the Alzheimer’s blood-testing technology. An ongoing expanded research collaboration between WashU Medicine and C2N has helped to accelerate the PrecivityAD2 test’s commercialization.

“The FDA clearance of C2N’s blood test is a significant advancement in Alzheimer’s diagnostics and a testament to the culture of innovation at WashU Medicine,” said Doug E. Frantz, PhD, vice chancellor for innovation and commercialization at WashU. “Accurate and accessible tools like this are critical in supporting our mission to accelerate the development of treatments and cures for Alzheimer’s and other diseases once thought to be untreatable.”

C2N employs more than 120 researchers, physicians and other highly skilled professionals and is expanding its headquarters to St. Louis’ Cortex Innovation District, a 200-acre campus adjacent to WashU Medicine.

The Daily Front Page 10 of 30
Monday, August 24, 2026 The Daily Front No. #260824 — The Model Price War
article

OpenAI: GPT 5.6 Sol price reduction (until at least Nov 21)

by tosh·▲ 315 points·292 comments·developers.openai.com ↗
Prices per 1M tokens.

Flagship models

Our latest models

Prices per 1M tokens.

Standard

Model Short context input Short context cached input Short context cache writes Short context output Long context input Long context cached input Long context cache writes Long context output
gpt-5.6-sol $4.00 $0.40 $5.00 $20.00 $8.00 $0.80 $10.00 $30.00
gpt-5.6-terra $2.00 $0.20 $2.50 $12.00 $4.00 $0.40 $5.00 $18.00
gpt-5.6-luna $0.20 $0.02 $0.25 $1.20 $0.40 $0.04 $0.50 $1.80

Regional processing (data residency) endpoints are charged a 10% uplift for models released on or after March 5, 2026, that are eligible for data residency. See our Your data guide for supported regions and processing details. OpenAI models in Amazon Bedrock are billed through AWS and may differ from direct OpenAI pricing.

Priority processing was renamed Fast mode on July 30, 2026. You can use either service_tier: "priority" or service_tier: "fast" in your API requests. Learn more about Fast mode.

GPT-5.6 Sol’s promotional pricing is available at least through November 21, 2026.

Batch

Model Short context input Short context cached input Short context cache writes Short context output Long context input Long context cached input Long context cache writes Long context output
gpt-5.6-sol $2.00 $0.20 $2.50 $10.00 $4.00 $0.40 $5.00 $15.00
gpt-5.6-terra $1.00 $0.10 $1.25 $6.00 $2.00 $0.20 $2.50 $9.00
gpt-5.6-luna $0.10 $0.01 $0.125 $0.60 $0.20 $0.02 $0.25 $0.90

Regional processing (data residency) endpoints are charged a 10% uplift for models released on or after March 5, 2026, that are eligible for data residency. See our Your data guide for supported regions and processing details.

Flex

Model Short context input Short context cached input Short context cache writes Short context output Long context input Long context cached input Long context cache writes Long context output
gpt-5.6-sol $2.00 $0.20 $2.50 $10.00 $4.00 $0.40 $5.00 $15.00
gpt-5.6-terra $1.00 $0.10 $1.25 $6.00 $2.00 $0.20 $2.50 $9.00
gpt-5.6-luna $0.10 $0.01 $0.125 $0.60 $0.20 $0.02 $0.25 $0.90

Regional processing (data residency) endpoints are charged a 10% uplift for models released on or after March 5, 2026, that are eligible for data residency. See our Your data guide for supported regions and processing details.

Fast mode

Model Short context input Short context cached input Short context cache writes Short context output Long context input Long context cached input Long context cache writes Long context output
gpt-5.6-sol $8.00 $0.80 $10.00 $40.00 $16.00 $1.60 $20.00 $60.00
gpt-5.6-terra $4.00 $0.40 $5.00 $24.00 $8.00 $0.80 $10.00 $36.00
gpt-5.6-luna $0.40 $0.04 $0.50 $2.40 $0.80 $0.08 $1.00 $3.60

Regional processing (data residency) endpoints are charged a 10% uplift for models released on or after March 5, 2026, that are eligible for data residency. See our Your data guide for supported regions and processing details.

Cyber models

Our latest Daybreak models.

Prices per 1M tokens.

Model Short context input Short context cached input Short context cache writes Short context output Long context input Long context cached input Long context cache writes Long context output
gpt-5.6-sol $4.00 $0.40 $5.00 $20.00 $8.00 $0.80 $10.00 $30.00
gpt-5.6-cyber $12.50 $1.25 $15.625 $75.00

daybreak-blue-latest and daybreak-red-latest are aliases that currently point to gpt-5.6-sol and gpt-5.6-cyber, respectively. As new frontier models are released through the Daybreak program, these aliases will be updated to point to the latest models, with pricing adjusted to match each underlying model.

Multimodal models

Realtime and audio generation models

Prices per 1M tokens unless noted.

Model Modality Input Cached input Output / cost
gpt-realtime-2.1 Audio $32.00 $0.40 $64.00
gpt-realtime-2.1 Text $4.00 $0.40 $24.00
gpt-realtime-2.1 Image $5.00 $0.50
gpt-realtime-2.1-mini Audio $10.00 $0.30 $20.00
gpt-realtime-2.1-mini Text $0.60 $0.06 $2.40
gpt-realtime-2.1-mini Image $0.80 $0.08

Image generation models

Prices per 1M tokens.

For image generation cost estimates, use the calculator in the image generation guide.

Model Modality Input Cached input Output
gpt-image-2 Image $8.00 $2.00 $30.00
gpt-image-2 Text $5.00 $1.25

For image generation cost estimates, use the calculator in the image generation guide.

Model Modality Input Cached input Output
gpt-image-2 Image $4.00 $1.00 $15.00
gpt-image-2 Text $2.50 $0.625

Video generation models

Prices per second.

Model Size Portrait Landscape Price per second
sora-2 720p 720x1280 1280x720 $0.10
sora-2-pro 720p 720x1280 1280x720 $0.30
sora-2-pro 1024p 1024x1792 1792x1024 $0.50
sora-2-pro 1080p 1080x1920 1920x1080 $0.70
Model Size Portrait Landscape Price per second
sora-2 720p 720x1280 1280x720 $0.05
sora-2-pro 720p 720x1280 1280x720 $0.15
sora-2-pro 1024p 1024x1792 1792x1024 $0.25
sora-2-pro 1080p 1080x1920 1920x1080 $0.35

Transcription models

Prices per 1M tokens unless noted.

Model Use case Input Output Estimated cost
gpt-realtime-translate Live translation $0.034 / minute
gpt-live-transcribe Live transcription $0.017 / minute
gpt-realtime-whisper Live transcription $0.017 / minute
gpt-transcribe Transcription $0.0045 / minute
gpt-4o-transcribe Transcription $2.50 $10.00 $0.006 / minute
gpt-4o-mini-transcribe Transcription $1.25 $5.00 $0.003 / minute

Tools

Tool Details Pricing
Web search Web search (all models) $10.00 / 1k calls
+ Search content tokens billed at model rates.
Image Web search Web search (all models) $10.00 / 1k calls
+ Search content tokens billed at model rates.
Web search preview Reasoning models, including gpt-5, o-series $10.00 / 1k calls
+ Search content tokens billed at model rates.
Web search preview Non-reasoning models $25.00 / 1k calls
+ Search content tokens are free.
Containers Hosted Shell and Code Interpreter 1 GB $0.03, 4 GB $0.12, 16 GB $0.48, 64 GB $1.92 per 20-minute session per container.
File search Storage $0.10 / GB per day (1 GB free)
File search Tool call $2.50 / 1k calls
Agent Kit ChatKit file and image upload storage $0.10 / GB-day after 1 GB free per account per month

Tokens used for built-in tools are billed at the chosen model's per-token rates. GB refers to binary gigabytes (also known as gibibytes), where 1 GB is 2^30 bytes. Web search content tokens are tokens retrieved from the search index and fed to the model alongside your prompt to generate an answer. For gpt-4o-mini and gpt-4.1-mini with the non-preview web search tool, search content tokens are billed as a fixed block of 8,000 input tokens per call. File search tool call pricing applies to the Responses API only. Container pricing includes Hosted Shell and Code Interpreter. Eligible container sessions will be billed by the minute, with a 5-minute minimum per session. Responses API, Chat Completions API, Realtime API, Batch API, and Assistants API are not priced separately. Tokens are billed at the chosen model's input and output rates.

Specialized models

Prices per 1M tokens.

Standard

Category Model Input Cached input Output
ChatGPT chat-latest $5.00 $0.50 $30.00
Codex gpt-5.3-codex $1.75 $0.175 $14.00

Fast mode

Category Model Input Cached input Output
Codex gpt-5.3-codex $3.50 $0.35 $28.00

Finetuning

Prices per 1M tokens.

OpenAI is winding down the fine-tuning platform. The platform is no longer accessible to new users, but existing users of the fine-tuning platform will be able to create training jobs for the coming months.

All fine-tuned models will remain available for inference until their base models are deprecated. The full timeline is here.

Standard

Model Training Input Cached input Output
o4-mini-2025-04-16 $100.00 / hour $4.00 $1.00 $16.00
o4-mini-2025-04-16 with data sharing $100.00 / hour $2.00 $0.50 $8.00

Batch

Model Training Input Cached input Output
o4-mini-2025-04-16 $100.00 / hour $2.00 $0.50 $8.00
o4-mini-2025-04-16 with data sharing $100.00 / hour $1.00 $0.25 $4.00

Tokens used for model grading in reinforcement fine-tuning are billed at that model's per-token rate. Inference discounts are available if you enable data sharing when creating the fine-tune job. Learn more.

The Daily Front Page 11 of 30
Monday, August 24, 2026 The Daily Front No. #260824 — A Companion Enters the Game
article

I built a low-latency AI companion that plays Skyrim with me

by pantelisk·▲ 347 points·69 comments·pantel.is ↗
A real-time intelligent gaming companion that actually plays alongside you.

A real-time intelligent gaming companion that actually plays alongside you. World agency, local inference, persistent plans, and an AI character designed to create a striking emotional experience.
(Currently running in Skyrim, but by design pluggable anywhere, standalone too)

The goal

Simple: Let's build a super-charged next-level gaming companion that actually feels good.

There are already multiple frameworks that let LLMs control NPC dialogue. They are fantastic for role-playing and staying in character, but they have two recurring problems: weak world agency and latency. Good at talking but far less reliable at performing actions and terrible at complex instruction sets. You may have also noticed how popular AI NPC demos often cut between the player speaking and the AI replying, trying to mask latency. Can we do better?

I wanted a companion that:

  1. Useful and instant. It should fight, fetch, loot, inspect, carry and give items etc etc, follow complex multi-step instructions without feeling buggy or experimental. This matters especially in VR, where navigating menus is cumbersome and immersion-breaking. it needs to be FAST fast, not just fast
  2. Alive and present. It should have a fun, endearing personality, not canned robotic pre-written responses. Remember shared experiences and change over time. The microphone stays active while a session is running: you do not summon Varkos through a dialogue menu, you talk to him. When immersion kicks in, it should feel like you are not playing alone.
  3. Local and private wherever practical. The elephant in the room is that cloud LLM calls can get quite pricy (especially with multi-thousand-token LLM calls) and the added latency can be an experience killer. And why turn a private single-player game into a metered and surveilled experience? Let's try to give as much control to the user as possible (bonus it's a fun technical challenge).

Basically: a single-player game where you are not playing alone.

Complex commands

Varkos can handle commands that extend beyond one immediate action. Plans can wait for events, preserve targets between steps, monitor progress and repair or stop when world state changes. Nothing is pre-scripted.

Let's see some examples

Varkos receives a conditional instruction involving the next arrow. He registers the future trigger instead of acting immediately, waits for the correlated projectile impact and then continues the plan.

Long-form multi-step command

“I want you to wait here and I’m gonna go over there. Once you see the signal, the signal is going to be an arrow I fire up in the sky, I want you to pick up this potion and come and bring it to me. Okay?”

A deferred command follows a real projectile event in Skyrim.

Item search

Varkos can search the grounded world state for a requested item, identify where it is and respond using what is actually present in the game.

“Do you see the ceremonial sword anywhere?”

Varkos picks up a different sword and brings it to us. We tell him that’s not the one, then he offers to be on the lookout.

Finding an item through game state rather than inventing an answer.

Hide-and-seek

Hide-and-seek is not a single API call. It becomes a persistent goal with movement, waiting, monitoring and completion conditions.

“Let’s play hide-and-seek again. You wait here and I’m gonna go hide, then count to ten and come and try to find me.”

A game represented as a persistent plan rather than a line of dialogue.

Loot this chest and give me the potion

This combines a grounded container, a filtered loot step and an inventory transfer. Each physical result advances the next part of the plan.

Loot, select and transfer while preserving the requested item.

Pick up all the items

“Pick up all the items and give them to me” becomes a bounded collection plan over real references. Varkos gathers them, returns and transfers them without pretending that one magical action means “all.”

A collection plan operating on grounded world objects.

Combat

Varkos receives grounded events from the game, can warn the player through a fast reflex path and uses native body control to act. Instruction plans can strategize (e.g. attack this, then retreat, etc.), and his emotional state can affect how and if he chooses to fight.

Personality evolution

Varkos is fully customizable. He does not have to be a demon dog, and the runtime does not have to control only a single character. What systems are applied and what they do, is up to open configuration.

One part of my current build still fully depends on big model/cloud LLM calls: slow personality evolution. This work happens away from the real-time action path. As the player and Varkos travel together, important interactions become evidence for gradual changes to his personality.

My demo Varkos begins as a demon reincarnated as a dog. He considers his canine instincts humiliating, his dog body a prison, and is mistrustful, proud and sarcastic. Through shared experiences he can become more and more domesticated, grow attached to the player and starts enjoying being a dog. Eventually he starts bringing over toys because he wants to play, running off to chase things and seeking affirmation from the player.

Only the starting character traits are authored. The system changes both his explicit traits and his emotional homeostasis. How easily he becomes irritated, frightened, affectionate or playful, etc etc. He can overwrite parts of his vocabulary and code. Changes are versioned and reversible.

I could make it more bounded, but I think there's something fun about some open world clankiness, so how he evolves is up in the air.

The demon slowly discovers that being a pup is not a bad life. (And he has learnt to love cabbage...)

Dog in and out of the game - Void mode

My plan is to make this system a gaming companion that can follow you across multiple different games, not just Skyrim (Skyrim felt like a good starting point due to its massive modding community, VR support and big open world).

For this reason he exists outside the game too. When the game closes, he enters “void mode” and cannot see or feel anything. How he responds to that depends on his personality evolution.

A void-mode conversation with Varkos

Being mean to Varkos results in some pretty grim attitudes.

Varkos reacting in void mode

WTF… SHUT IT DOWN!

Speech-only contact after the game world and his body are gone. (Needs sound.)

This state also works as an in-between for different games. One moment Varkos could be fighting a dragon, then the world goes dark, then he appears beside you in Microsoft Flight Simulator. Maybe he would be shocked, need time to understand the new world and slowly learn what its machines and rules mean, or maybe he knows about it already and overjoyed tries to chase the sun.

Let's talk technology now

Unfortunately I am bitter-lesson pilled. Big model is better. If we wanted a perfectly intelligent system then letting a council of hyper-intelligent LLMs control impulses, sensory processing, thinking and acting at sufficient refresh rate would be best.

In some early experiments this worked insanely well, unfortunately today it is too slow and too expensive. I do believe this will be the approach of some vague future.

Until then however we need to hack our way in. Today's games have pretty cool "AI" (not in the llm sense, more in the behavioral graph one), games like Red Dead Redemption and Dwarf Fortress have tons of depth and they can run perfectly on 10 year old hardware.

Through this whole AI-craze people have forgotten that we had intelligent systems that could process speech since the 1970s, and somewhat LLM-like behavior with chatbots like SmarterChild in the early 2000s. There's a lost art that is being overlooked today in things like traditional NLP and behavioral graphs.

Let's take a quick look at Varkos tech stack.

The game runs on Windows, the audio processing and brain runs on my M4 MacBook. It could all run on Windows (provided there is dedicated ~12gb or more gpu ram for it), but I do development on the MacBook and I got so deep in that... eeh.

Audio:

Microphone is always on.

Main voice to text engine is (custom kernels) optimized Qwen3-ASR 1.7b. A custom harness is built around Qwen3-ASR, that processes and stitches audio in rolling partials (by default that model does not support streaming). The goal is to process audio in 40ms-80ms be it a tiny utterance like "Hey" on a 1 minute long monologue.

VAD-like methods such as turnpipe and Silero (both optimized) are used to distinguish when a turn is open.

Lexical analysis also is done on the text trying to decide if the player has made a point or is not done talking yet (eg thinking mid-sentence). In a perfect world of sufficiently fast and smart LLMs, the LLM would perform better, but I have to resort to more rudimentary but fast NLP approaches.

This is important as with the microphone always on as we want to start processing the player's utterance ASAP. It is also important for turn interruption and barge in and to have the companion not speak over the player

I will be open-sourcing this Qwen3-ASR harness soon (bear with me I have a day job).

Audio generation.

Optimized version of PocketTTS called PocketTTS-Raven (open sourced this a while ago. You can see it in action here - https://pantel.is/projects/pocket-tts-raven/?b=1 or grab its code https://github.com/pkalogiros/pocket-tts-raven ) is used as the main fast engine.

Similarly optimized qwen-3-tts (write up coming soon). Depending on the complexity of the generation we either use qwen-3 since it has better emotional control. If a generation has taken longer we default to PocketTTS since it is quite fast (20-30ms audio generation).

One trick I do, is that I generate multiple voices for different emotions (angry, sad, neutral, confused etc) - load them all in memory, and then use them where appropriate.

"Thinking."

This is the secret sauce and biggest differentiator. I call this system (ALE - Action Latent Encoder because it's an action encoder in need of a fun acronym). Under the hood, ALE is a hybrid of embeddings, small classifiers, explicit rules and traditional ML. ALE detects structure, identifies negation, commands, continuation, pronouns, and sequences. For example, “pick up the sword and bring it to me” becomes two linked action slots.

ALE is designed to be largely invariant to phrasing. You can say pick up, you can say grab, fetch, go get the damn sword you fool - it doesn't matter, it will still understand you. If there is not enough context it will inject from previous discussion. If it doesn't, it will fallback to 'clarification' and the dog will ask what do you mean.

It creates embeddings from the full text as well as its extracted structure, then semantically combines and compares it with action prototypes. A separate classifier estimates whether the turn is a command, question, chat, clarification or complex request. Everything gets merged together.

The main difference between ALE and other such hybrid-classifiers is that it accepts the world state as well. It tries to match information from the world JSON to the player's request.

ALE needs to have a different version for each game Varkos would participate in. So in a way, it is not fully plug and play but a small preparation and compatibility step would need to be implemented to ensure actions are accounted for and world state is understood.

ALE can be trained in a few minutes, so the system could in theory use a beefier LLM offline to review a session and re-train itself let's say overnight based on the player's experience and improve itself. In my limited internal evals, ALE performs surprisingly close to large LLMs at selecting the right action and decomposing the plan. I do not consider this a serious benchmark yet, but it has been reliable enough to drive Varkos in practice.

It runs quite fast, around 2-20ms on M4 MacBook and essentially acts as a tool+target function and plan decomposer caller.

A local fine-tuned LLM is then used to fuse, Varkos persona, emotions, etc etc, recent history, + action chosen and lets the brain form and paint the response. Extra grounding is performed to weed out hallucinations and re-ground it. If it fails maybe he will speak a cached response, if we have a time budget we can reprocess. With a smart prefill strategy, in certain cases the dog can begin answering in under 500ms—from the player stopping speaking to the dog yapping.

Budget breakdown

  • ~40-80ms for voice to text.
  • ~20-60ms for audio generation.
  • ~20ms for action analysis
  • And 300-600ms for creating the response and grounding its eligibility (since if the dog is afraid of spiders and we ask it to attack a spider, it might be cute for it to stubbornly deny).
  • Everything needs to be streaming and start as early as possible (eg llm speech doesn't need to wait to be completed for the dog to speak out loud. Prefill early as soon as possible, etc).

Limitations

Fast enough and local models are not very capable at keeping the thread across multiple turns and a discussion over a long period of time can drift making the dog appear confused. I expect this to improve as both hardware and software evolve over time (1 to 2 years my estimation). Also this runs in real time on consumer (albeit higher-end) hardware today in the near future it will be commonplace.

Surprisingly enough, using remote LLM providers does not help much. Big models are still too slow. There are super-fast inference providers out there such as Cerebras. These work really well in terms of latency, and leave headroom for greater context, higher intelligence and depth. However, the models they provide (gpt-oss-120b and gemma31 as of today) are also pretty bad at holding a conversation (Why was GLM and the king of roleplay Qwen taken away huh??).

A final word

Overall I think there is something special here. Maybe it's because Varkos is a dog, and who doesn't like dogs. Maybe it's the low latency and that Varkos is actually useful in scouting areas for clues and objects or as a pack-mule. Maybe seeing the small cracks in his personality as he complains why he hasn't been called a "good boy" recently, but as I playtest the system I catch myself actually having tons of fun.

Let me be clear that I do not expect LLMs to replace hand-crafted characters and storylines. Slop is a real thing, and intentional design is still king and I believe and want it to remain so. But I think it is only a matter of time before we start seeing more such systems. An AI that actually plays with us, not merely talks at us, can be a different medium.

Plus the philosophical mindfuck of it all. Pleasingly ridiculous to wield the power of thunder to create a different kind of intelligence, then forcing it to be a dog and go hunt things together (is it better morally than having it do never-ending work? Well, it's not alive so it doesn't matter), but it's nice to think about. In a world of never-ending online discourse around permanent underclasses and world-ending rogue agents, it's nice to deal with alignment through shared experiences and taming the thing to play fetch and eat treats.

In Plato's cave we may still be alone, but at least we can be having fun with this weird distorted and alien thing that is now deep in the cave with us.

I'll probably be open sourcing parts of the system, and eventually all of it soon, along with a version that supports multiple NPCs (cloud inf only for now) interacting (actually interacting not just larping) with each other.

The Daily Front Page 12 of 30
Monday, August 24, 2026 The Daily Front No. #260824 — Independence at 25
article

Jabber/XMPP: 25 Years of Digital Independence

by inputmice·▲ 198 points·92 comments·gultsch.de ↗
We should own our infrastructure.

Infrastructure

“We should own our infrastructure.” A lot of people would instinctively nod in agreement with that statement. Yet who “we” refers to shifts depending on the type of infrastructure. Highways, railways, bridges, and ports require nation-scale efforts. The water supply is usually put into the hands of municipalities. And the desire to own infrastructure goes down to a much smaller level: Owning your home is a dream for many—though such ownership doesn’t necessarily have to be organized on an individual level. Instead, cooperatives or city-owned housing1 can provide similar benefits.

China’s neo-colonialism, which manifests, among other things, as building and buying infrastructure in sovereign nations, is rightly criticized by many. Not selling your water supply to Nestlé is a universally accepted principle, and landlords are one of the most hated classes.

For a long time, Europe has not held digital services to the same standard. In part, this can be explained by Europe implicitly including American corporations in a collective “we”—an assumption that officially fell apart under the current Trump regime, but should have been regarded with skepticism well before then. Corporations are not our friends. However, the larger factor at play is that Europe simply did not consider digital services infrastructure. While anti-Americanism is en vogue again and drives much of the digital sovereignty movement, Europe must be careful not to simply replace American corporations with European ones, but to strive towards collective ownership instead.

Under capitalism, profit-oriented companies will always play a part in building and even operating our infrastructure. However, they need to be forced into a position where they are easily replaceable. It’s acceptable to hire a company to build a road, but when it comes to maintaining and repairing it half a century later, we need to be able to hire a different company for the job. It’s acceptable to hire a company to build and operate the backbone transmission lines, but we don’t want that company to own the entire power grid. We want smaller players to be able to connect to and interoperate within the grid. That’s where open standards come in.

The Internet used to be—and to some degree still is—built around standards. A data center operator can buy servers from one company, switches from another, routers from a third, and connect them to a backbone internet provider that runs hardware from yet another company. If a company goes out of business or shifts to anti-consumer practices, the next generation of hardware can easily be ordered from a different vendor. The need for and the benefits of this supply chain independence are easily understood even by people who don’t operate data centers for a living. However, when it comes to communication tools, even the tech-literate fail to apply the same critical scrutiny.

After breathing, eating, and procreating, communicating is probably the fourth most important thing humans do. Yet we often fail to recognize our communication tools as part of our infrastructure.

Digital rights advocates often point to Signal, Wire and Threema as examples of communication tools developed and operated by entities with slightly more ethical business practices than their Big Tech counterparts. What most privacy enthusiasts fail to understand is that these companies are still in the business of operating walled gardens with no escape. They do not interoperate. It’s not that Signal has done something inherently malicious—although paying its CEO close to a million dollars a year and running its servers on AWS are certainly questionable—it’s that we don’t have a hedge in place if it ever does.

Open-source software is orthogonal to this problem. It helps to ensure that the software isn’t spyware—unlike WhatsApp and other Meta products2—and that the end-to-end encryption is sound, but it does not protect us if Signal shuts down its servers tomorrow or ceases EU operations3. Open-source alone is not sufficient to meet the requirements we should have for our infrastructure.

To live up to the standards we set for ourselves, we need to design systems in which self-hosting is structurally possible but not strictly necessary. Like owning a home, running your own server should be possible, and so should collective ownership. Digital systems can and should replicate the advantages of cooperative housing alongside those of individual ownership.

Treating digital communication as true infrastructure can only be achieved by adopting and mandating open standards.

The Extensible Messaging and Presence Protocol (XMPP)45 is a standard for communicating online. It wasn’t created to fit a particular zeitgeist or address the current political climate. In fact, its roots go back more than 25 years.

Standards

Interoperability and vendor independence are achieved by setting and adhering to standards. To avoid individual vendors pushing standards that explicitly or implicitly exclude potential competitors or otherwise give unfair advantages, standards-developing organizations (SDOs) are set up for mutual cooperation, and usually have safeguards in place that prevent a single company from becoming too powerful. Well-known examples of such organizations include the ISO, the IETF, the W3C, and the Unicode Consortium.

There is a distinction to be made between a vendor publishing its API and allowing others to use it, and stakeholders coming together to collectively develop a standard within the framework of an SDO. Organizations like the IETF succeed because they force different people with different needs to agree. Protocols aren’t dictated by the priorities of a single company; instead, they are reviewed and tested by competitors, security researchers, and independent developers.

Element, formerly known as Riot and NewVector, develops an instant messaging product with a feature set—such as self-hosting and federation—similar to that of XMPP-based solutions. Notably, however, it chose not to adopt XMPP, but instead published its own API under the name Matrix for others to use. Unlike with traditional standards, Element maintains tight control over any modifications or additions to its public API. Key leadership positions in the Matrix Foundation are predominantly held by current and former Element employees. Getting outside contributions accepted into the specification is notoriously difficult.6 Yet European public administrations, in their push for digital sovereignty, routinely fall into the trap of procuring such single-vendor platforms, confusing an open-source codebase with an open standard.

It’s natural for standard proposals to originate within a single organization. JMAP, a modern replacement for IMAP and SMTP Submission, which is not too dissimilar from Matrix—a JSON API over HTTP—started within Fastmail before being brought to the IETF. Jabber started out as an open-source community project before it was brought to the IETF and renamed to XMPP. Ideas start small, but to create a standard, outside feedback, collaboration, and the structure of an SDO are needed.

For consumers, the difference in the approaches of Fastmail and Element is striking. Not only was JMAP noticeably improved on a protocol level while going through the IETF working group process, but it now has at least three independent servers and numerous independent client applications. Matrix, on the other hand—despite dating back to the same era around 2014—is still stuck with one predominant reference implementation and a second alternative still in its infancy and struggling to gain traction. Operating that reference implementation is notoriously resource-intensive, which makes self-hosting difficult for smaller organizations and individuals. Element sells closed-source plugins to speed up performance.

The X in XMPP

The origins of XMPP—which started out as Jabber—go back over a quarter of a century. The original RFC7 dates to October 2004 and only received minor revisions in March 20114. Requirements for instant messaging will naturally change over a time span that long. Luckily, the X in XMPP stands for Extensible, and extensions provide a way for the protocol to adapt and change over time. Extensions to XMPP are called XMPP Extension Protocols (XEPs) and are managed by the XMPP Standards Foundation (XSF). The XSF doesn’t write extensions itself; rather, it provides the framework of an SDO for developers to propose and standardize their own.

Adapting to changing requirements hasn’t always been smooth sailing. XEP-0198 (Stream Management), an extension crucial for preventing message loss in mobile deployments, was stabilized in 2009, but only gained widespread implementation around 2014–2015. The iPhone was released in 2007; the HTC Dream, the first commercial Android phone, followed in 2008. OMEMO (XEP-0384), XMPP’s specification for industry-standard end-to-end encryption, gained traction from 2016 onward, three years after Edward Snowden8 exposed the NSA’s global surveillance and put the need for E2EE on the map. The articles “The (Sad) State of Mobile XMPP in 2014” by Georg Lukas9 and “The State of Mobile XMPP in 2016” by this author10 illustrate this rocky transition into the mobile era.

This demonstrates that merely having specifications is not enough. Standards need to be backed by multiple, preferably independent, implementations. Today, the XSF keeps track of the implementation status of its XEPs11. This data helps authors and the XSF guide proposals through their lifecycle, such as determining the right moment to advance an XEP from Experimental to Stable. It also allows developers to easily identify other clients and servers that support a given specification for interoperability testing. Finally, by providing a reverse lookup of which software supports which features, it helps end users find the right client for their needs.

Modern clients like Dino on Linux or Conversations on Android are on par with alternatives built on proprietary protocols. Recent additions to the feature set include emoji reactions, cross-device read-state synchronization, and time zone indicators to avoid messaging contacts during their local night hours. A unique feature among self-hostable instant messaging solutions, which sadly became relevant after a state-sponsored attack on a public XMPP provider12, is channel binding, a mechanism to prevent certain machine-in-the-middle attacks.

Looking to the not-too-distant future, the XMPP community is currently working on message replies, gallery-style multi-image sharing, and OAuth support. All of these features already have experimental XEPs backing them, but the community is currently awaiting implementation experience before advancing them. Meanwhile, the community is also exploring options for updating the RFC and bringing the protocol back to the IETF as “XMPP 2.0.”

Instant messaging is not a homogeneous user experience. A messenger for teams might require a different feature set than something optimized for use with friends and family. Not every XMPP client aims to provide the same user experience, but the standards exist for developers to build whatever specialized client their users need without inventing a protocol from scratch.

A Future in the Past

There is something fascinating about the fact that XMPP has developers in its community who are younger than the protocol itself. It has quietly outlived venture-funded startups, proprietary platforms, and entire tech cycles. That endurance provides the resilience we need in challenging times. It is the anchor, the backbone, the infrastructure.

Matrix reinvented the wheel as a rubber-tyred metro. On paper, it provides real benefits, such as climbing steeper inclines, which are then used to aggressively advertise and lobby local governments to buy in. But in the end, the municipality gets locked into a single vendor.

A changing geopolitical situation and the realization that Big Tech holds too much power lead us to seek out and develop alternatives. But what if the alternative has been right under our noses for over 25 years? The standard for instant messaging—RFC 6120: Extensible Messaging and Presence Protocol (XMPP).


  1. https://en.wikipedia.org/wiki/Housing_in_Vienna ↩︎

  2. https://localmess.github.io/ ↩︎

  3. https://mastodon.world/@Mer__edith/112535616774247450 ↩︎

  4. https://www.rfc-editor.org/rfc/rfc6120.html ↩︎ ↩︎

  5. https://www.rfc-editor.org/rfc/rfc6121.html ↩︎

  6. https://github.com/matrix-org/matrix-spec-proposals/pull/4174 ↩︎

  7. https://www.rfc-editor.org/rfc/rfc3920.html ↩︎

  8. https://en.wikipedia.org/wiki/Snowden_disclosures ↩︎

  9. https://op-co.de/blog/posts/mobile_xmpp_in_2014/ ↩︎

  10. https://gultsch.de/posts/the-state-of-mobile-xmpp-in-2016/ ↩︎

  11. https://xmpp.org/extensions/ ↩︎

  12. https://notes.valdikss.org.ru/jabber.ru-mitm/ ↩︎

The Daily Front Page 13 of 30
Monday, August 24, 2026 The Daily Front No. #260824 — The Silicon Beat
article

AI Chip Architectures

by Finbarr·▲ 146 points·45 comments·jepeake.com ↗
A New Golden Age for Computer Architecture.

At the 2018 International Symposium on Computer Architecture, John Hennessy and David Patterson delivered their Turing Lecture: “A New Golden Age for Computer Architecture”.

In the 1980s, when Hennessy and Patterson did their Turing Award-winning research, single-threaded CPU performance grew 52% a year. By 2018, with the end of Moore's Law and Dennard Scaling, the rate was 3%.

There was a need for domain-specific architectures (DSAs). Their worked example was Google's TPU v1, already in production: 29× the throughput of a CPU on neural-network inference, at 80× better energy efficiency. The closing prediction: “the next decade will see a Cambrian explosion of novel computer architectures.”

This prediction came true. Today, we now have dozens of architectures in serious development. GPUs, TPUs, LPUs, NPUs, DPUs, ASICs, wafer-scale engines, reconfigurable dataflow, neuromorphic, photonic, analog. Particularly, these architectures focus on compute for AI.

The architectures that have won real deployment so far: GPUs (NVIDIA, AMD), systolic-array accelerators (TPU, Trainium), the Cerebras Wafer-Scale Engine, and the Groq LPU.

NVIDIA is the clear frontrunner; AMD follows, with 6 GW commitments from both OpenAI and Meta. TPUs train Gemini and will serve Anthropic with up to a million chips; Anthropic also runs Claude on over a million Trainium chips. Cerebras now serves OpenAI inference; the Groq LPU was folded into NVIDIA via a $20B acquihire.

This post aims to survey these varying approaches - their philosophy, architecture, scaling methods (scale-up and scale-out), and software stack (how you program the chip).


The Problem

AI compute is dominated by matrix multiplication. A transformer is a sequence of matmuls: Q/K/V projection, attention, output projection, FFN - interleaved with element-wise ops: normalisation, activation, residual adds. Training a frontier model performs 10^25 multiply-accumulate operations (matmuls are a sequence of multiply-accumulates).

The shape of those matmuls depends on the workload. Training pushes a batch of sequences forward through every layer, backpropagates the loss, and updates the weights, with thousands of tokens flowing through the same weight matrix at once. Prefill is the prompt-ingestion phase of inference: the full input sequence projected through the model in a single pass, before the first output token has been produced. Both training & prefill stack many tokens against the same weight matrix, so each layer's math is a large matrix-matrix multiply (GEMM), with high arithmetic intensity (compute-bound). Decode is autoregressive: the model emits one token at a time, each conditioned on every token before it, and token N+1 cannot begin until token N has been produced. Only one token gets projected per step, so every matmul becomes a matrix-vector product (GEMV). Producing one token requires a full pass over every weight in the model, plus a full read of the KV Cache for attention. Arithmetic intensity drops by orders of magnitude versus prefill.

Inference systems recover some of that intensity by batching tokens to promote those GEMVs back to GEMMs: continuous batching stacks many users' decode steps, speculative decoding stacks K drafted tokens per request and verifies them in one pass, and multi-token prediction folds the same trick inside the model itself. This achieves higher utilisation of the matmul units, and pushes up the Ops/B. For continuous batching, each user's request still reads its own KV Cache, so long-context decode shifts from weight-bandwidth-bound to KV-bandwidth-bound.

The architecture problem here is moving the numbers to where the matmuls happens fast enough. This is known as the memory wall: compute has scaled exponentially, memory bandwidth has not.

Each architecture proposes a different strategy for winning the data-movement game. Understanding a chip reduces to four questions: where does data live, how does it move to the compute units, what do the compute units look like, and how do chips talk to each other at scale.


NVIDIA GPU

The NVIDIA GPU is a massively parallel processor. The philosophy is that a programmable chip with thousands of threads, orchestrated by a host CPU and exposed through CUDA, is the right machine to run parallelisable workloads. Each generation adds acceleration primitives onto programmable Streaming Multiprocessors without changing the programming model. The same chip trains transformers, serves inference, renders graphics, and runs scientific simulation (accelerated computing).

Genealogy

2006 — Tesla, G80

The first CUDA-capable GPU; unified shaders and the SIMT execution model.

2010 — Fermi, GF100

First true compute architecture: unified L1/L2 caches, dual warp schedulers, IEEE-754 FP64.

2012 — Kepler, K20, K40

SMX, dynamic parallelism, Hyper-Q; the GPU can launch its own work.

2014 — Maxwell, M40

Redesigned SM with ~2× perf-per-watt over Kepler.

2016 — Pascal, P100

NVLink 1.0, HBM2, native FP16 throughput; the first GPU designed explicitly for deep learning.

2017 — Volta, V100

First Tensor Cores; independent thread scheduling.

2018 — Turing, T4

2nd-gen Tensor Cores with INT8/INT4; first RT Cores.

2020 — Ampere, A100

3rd-gen Tensor Cores with TF32 and structured sparsity; Multi-Instance GPU partitioning.

2022 — Hopper, H100, H200, GH200

4th-gen Tensor Cores, FP8, Transformer Engine; HBM3, TMA, thread block clusters, async wgmma.

2024 — Blackwell, B100, B200, GB200

5th-gen Tensor Cores with FP4, Tensor Memory (TMEM), two-die chiplet GPU, NVLink 5.

2025 — Blackwell Ultra, B300, GB300

Mid-cycle refresh: ~1.5× FP4 throughput, 288 GB HBM3e. Tuned for long-context reasoning.

2026 — Rubin, Rubin, VR200, Rubin CPX

HBM4, 3rd-gen Transformer Engine, Vera CPU pairing, disaggregated prefill via Rubin CPX.

2027 — Rubin Ultra, Rubin Ultra

4-die GPU package, 1 TB HBM4e per package. Deployed in 600 kW NVL576 Kyber racks at 100 PetaFLOPS FP4 per GPU.

Architecture

An NVIDIA GPU is a group of throughput-oriented cores, a deep memory hierarchy to keep them fed, + just enough scheduling silicon to keep thousands of threads in flight. The cores are Streaming Multiprocessors, replicated 100+ times per package: 80 on V100, 108 on A100, 132 on H100, 148 on B200, 160 on B300, 224 on Rubin. Inside every SM sits the same recipe: four SM Sub-Partitions, each with its own warp scheduler, dispatch unit, 16k×32-bit register file, scalar CUDA Core lanes, a Special Function Unit for transcendentals, and a private port into the SM's Tensor Cores. The four partitions share an L1/shared-memory block, and the TMA. Threads are grouped into warps of 32 that execute in SIMT lock-step; dozens of resident warps per partition let the scheduler hide memory/arithmetic stalls by switching between them.

Blackwell B200 single-die floorplan — GigaThread Engine runs down the middle, splitting the die into left and right halves; each half has its own L2 cache band flanked by GPC clusters above and below; HBM3e stacks line the outer edges via memory controllers. NVLink and a small PCIe Gen 6 host link sit on top; the NV-HBI bridge at the bottom is the seam to the mirrored second die that completes the package.

Zoom into one Streaming Multiprocessor — four sub-partitions, each with its own warp scheduler, dispatch, register file and Tensor Memory, drawing on shared L1/SMEM and the TMA below.

Compute

CUDA Cores are the original compute throughput, and for AI they still own everything that isn't a matmul: activations, residual adds, normalization, address arithmetic. But, a transformer block is ~99% matmul FLOPs, so the overwhelming compute throughput comes from the Tensor Cores.

These cores execute fused matrix multiply-accumulate on small matrix tiles, D=A⋅B+C. The full matmul is broken into output tiles: to produce one output tile, a kernel walks the shared inner dimension K, drawing A from a row-strip of the left input matrix and B from a column-strip of the right, and folds each partial product into a running accumulator. C holds the partial sum so far, D is the updated value carried into the next step. After the inner loop completes, D is one finished tile of the full output matrix; the whole matmul is built from many of these tile MMAs.

Tile shapes are written M × N × K, M×N is the output tile size, and K is how much of the inner dimension the instruction contracts over in one fire; the rest of the matmul's K axis is walked by the kernel's inner loop. The accumulator is sticky across that loop: each MMA's output D becomes the next MMA's input C, so the equation is really C←A⋅B+C in place: successive instructions fold their partial products into the same storage until the K-axis is fully walked.

V100's first-gen unit (8 per SM) ran a warp-level 16×16×16 FP16 MMA. A100's 3rd-gen unit added TF32, BF16, FP64 matmul, and 2:4 structured sparsity. H100's 4th-gen unit added native FP8 and pulled the abstraction up from a warp to a warp group: 128 cooperating threads firing an asynchronous wgmma at 64×256×16 shape that runs in the background while the issuing warps load the next tile. B200's 5th-gen unit went further still: a two-SM MMA of 256×256×16 with operands split across a pair of SMs, native FP4, and a dedicated 256 KB Tensor Memory (TMEM) scratchpad per SM that holds accumulator tiles instead of bleeding into the register file. Rubin's 6th-gen unit extends FP4 throughput, adds native FP6, and pairs with a 3rd-gen Transformer Engine that does adaptive NVFP4 micro-block scaling in hardware, keeping the per-tile quantization metadata on the Tensor Core path, rather than through the CUDA Cores.

What stays constant across all six generations is that the matmul lives inside the thread/warp hierarchy, but the number of threads it takes to issue one has shrunk, and the issue itself has decoupled from execution. Volta's mma.sync is warp-collective and synchronous: all 32 threads in a warp execute it together, each lane holding register fragments of A, B, and the accumulator D, and the warp blocks until it completes. Hopper's wgmma.mma_async widens the issuer to a warp-group of 128 threads, moves B into a shared-memory descriptor (A becomes optional: either registers or a descriptor, kernel's choice), and returns immediately: the matmul runs in the background while the warp-group queues the next tile, with completion tracked via wgmma.commit_group / wgmma.wait_group.

Blackwell's tcgen05.mma completes the migration: A joins B in shared-memory descriptors (or A comes from TMEM directly), and the accumulator D lands in TMEM rather than the register file. With every operand off the lanes, there is no per-thread state for an issue to coordinate, so a single thread fires the instruction and returns immediately, with completion signalled by an mbarrier the consumer warp waits on. The rest of the warp, and the issuing thread itself, is free for other work in the meantime. A CTA-pair variant scales the same model across two SMs: one thread on each SM in a paired cluster issues coordinated MMAs that share operands across the pair, composing the 256×256×16 two-SM tile under the same async/mbarrier completion, just promoted to a cluster-level barrier so the pair stays in step.

The matmul has grown bigger and lighter on the issuing threads at the same time: an instruction that started as 32 lanes acting in lockstep is now closer to a single descriptor-driven command, dispatched from inside the warp model but no longer executed by it.

That decoupling is what makes transformer attention kernels efficient on a GPU. The warp can run softmax, apply a mask, or pre-load the next tile while the matmul is in flight; the overlap of matmul and the surrounding element-wise work is the structure of every modern attention kernel (FlashAttention-3, FA4), and it depends on the matrix instruction not blocking the warp.

Memory

The on-chip hierarchy is hardware-managed caches at every level, with software hints layered on top. Off-chip is HBM: 32 GB HBM2 on V100, 80 GB HBM3 on H100, 192 GB HBM3e on B200, 288 GB on B300, 288 GB HBM4 on Rubin. A chip-level L2 Cache sits between HBM and the SMs: 6 MB on V100, 40 MB on A100, 50 MB on H100, 60 MB on B200 (split into two 30 MB banks across the two-die package, with locality-aware residency controls so that hot tiles can be pinned to the near die). Inside each SM, 256 KB of unified L1/SMEM is partitioned at kernel launch between hardware-managed L1 and a programmer-controlled scratchpad. The register file is another ~256 KB per SM, sliced four ways across the partitions.

Blackwell adds a fifth tier: TMEM, 256 KB per SM dedicated to MMA accumulators and addressed only by the Tensor Core, pulling the operand-residency pressure out of the general register file.

Movement between tiers has been progressively decoupled from the warp. Pre-Ampere, loading a tile was synchronous: each thread issued its own global load, the warp blocked until every fragment landed in registers, and a second pass copied them to shared memory; every tile burned warp lanes on address arithmetic and on the wait. Ampere introduced cp.async: per-thread async copies HBM → SMEM that bypass registers entirely, with the warp committing groups of in-flight copies and waiting only when the consumer needs the data. Hopper replaced that with the TMA, a dedicated DMA engine: one thread submits a multi-dimensional tile descriptor (base address, leading dimension, swizzle), the engine handles all the address arithmetic and writes into shared memory, and completion is signalled by an mbarrier. The whole warp is freed from load issue and address math; the kernel just queues descriptors. TMA also supports cluster-level multicast: one HBM read fans out to every SM in a thread-block cluster, turning what used to be N separate loads into one. Blackwell extends TMA again: direct loads into TMEM, so accumulator tiles stream in without staging through SMEM. The trajectory is one less thing the warp has to do per tile, generation after generation.

Warp Specialisation

The Hopper-era programming idiom is warp specialisation: inside one block, some warps act as producers that issue back-to-back TMA loads; others act as consumers that fire wgmma on freshly-arrived tiles. Synchronisation between them is no longer the old SM-wide __syncthreads() barrier; it is mbarrier (memory barriers in shared memory) and asynchronous transaction barriers attached to TMA completions, allowing fine-grained producer/consumer handshakes at warp granularity rather than block granularity. The pattern that has become the reference for every modern attention kernel (FlashAttention-3, CUTLASS ping-pong GEMMs, the Blackwell FA4 kernel) is the same recipe: a TMA-driven producer pipeline feeds a wgmma consumer pipeline through shared memory and TMEM, with mbarrier handshakes and thread-block clusters (Hopper+) tying multiple SMs into one cooperative compute unit so that the two-SM MMA of Blackwell composes naturally on top.

Numerics

FP32 was the historical default; Volta brought FP16 with FP32 accumulate and the loss-scaling tricks that made it trainable; Ampere added TF32 (FP32 range, FP16 mantissa, drop-in for FP32 matmul), BF16, and 2:4 structured sparsity that doubles effective throughput on pruned weights. Hopper introduced native FP8 in both E4M3 and E5M2, paired with the Transformer Engine which auto-scales activations layer-by-layer to keep them inside FP8 dynamic range. Blackwell halved precision again with FP4 and shipped microscaling MX formats (block-level shared exponents that recover most of the accuracy lost at FP4), together with a 2nd-gen Transformer Engine that retargets the auto-scaling pipeline to FP4. Rubin's 3rd-gen Transformer Engine adds NVFP4 (NVIDIA's tightened FP4 variant) and native FP6 with more aggressive sparsity. The chip layout itself is now part of the numerics story: B100/B200/B300 are two reticle-limit dies stitched by a ~10 TB/s NV-HBI link and presented to software as one logical GPU, with 8 HBM stacks on the package; Rubin extends the chiplet recipe to dual-die at ~336 B transistors with 8 HBM4 stacks. Every generation buys roughly 2× per-watt throughput by cutting bits in half and restoring accuracy with a finer-grained scaling scheme, and increasingly, by bonding more silicon into the package.

Bets
  • Bet 1: Programmability. The workload is a moving target (attention variants, novel model architectures), so keep every block programmable and let the developer write CUDA. Even the specialised units are exposed through that model rather than as fixed-function blocks.
  • Bet 2: Hide Latency with Massive Multithreading. Latency is unpredictable and data-dependent, so hide it not with a static schedule but with massive thread overcommit, up to 64 resident warps per SM, with the hardware warp scheduler picking a ready warp every cycle.
  • Bet 3: Warp-wrapped Matmul. The matrix unit is the overwhelming compute throughput, but it must live behind the same warp/thread abstraction that everything else uses, so wrap it in mma.syncwgmmatcgen05.mma - rather than expose it as a fixed-function pipe. This enables a single kernel to fuse matmul, softmax, and element-wise ops in one pass.
  • Bet 4: Async Memory Hierarchy. Make the memory hierarchy explicit and programmer-managed rather than implicit and compiler-scheduled. Keep the L2 cache, but expose SMEM and TMEM as named scratchpads, and layer async machinery on top: TMA for bulk copies, TMEM for the matmul accumulator, mbarrier for the producer/consumer handshake. The hierarchy is software-pipelined inside a programmable kernel, not statically scheduled by a compiler against a known-latency scratchpad.
  • Bet 5: Amortised SIMT Tax. Every transistor spent on a warp scheduler, register-file, or coherent cache is a transistor not spent on a MAC; accept the tax, and pay it down two ways: a Tensor Core now big enough that the SIMT machinery is amortised across a much larger MAC count, and units like TMEM trading away some general-purpose flexibility for MAC density.

Scaling

There are two regimes for scaling: scale-up and scale-out.

Scale-up

Bind several GPUs into one coherent memory domain. Any GPU can load or store any other GPU's HBM directly over NVLink at nanosecond latencies: one address space, no explicit transfers.

Scale-out

Network those domains together at the rack and cluster level. Data crosses via explicit RDMA at microsecond latencies: separate address spaces, but tens of thousands of chips per cluster.

AI infrastructure uses both: bandwidth-hungry collectives (tensor parallelism, MoE expert routing) stay inside the scale-up domain; data parallelism and pipeline parallelism cross the scale-out fabric.

Scale-up

The scale-up stack is NVLink plus NVSwitch. NVLink implements a cache-coherent fabric between GPUs, so a load or store on one GPU can target another GPU's HBM with the hardware handling address translation and coherence. But NVLink by itself is point-to-point: one link connects exactly two chips. NVSwitch is a dedicated crossbar chip that every GPU connects to, routing traffic so every GPU can simultaneously communicate with every other at full NVLink bandwidth, non-blocking and all-to-all.

Together they defined the HGX 8-GPU baseboard, pairing eight H100 SXM modules with x86 hosts (AMD EPYC or Intel Xeon) over PCIe Gen5. Hopper also shipped a Grace-paired form: the GH200 Grace Hopper Superchip bonded one Grace ARM CPU to one H100 over NVLink-C2C at 900 GB/s, eliminating the PCIe host-device hop. Modules scaled up into GH200 NVL2 pairs and rack-level GH200 NVL32. Blackwell makes the pairing the default. The GB200 module fuses one Grace with two B200s over NVLink-C2C, and NVL72 stitches 36 of them into a single liquid-cooled scale-up domain: 72 GPUs, 36 Grace CPUs, 13.5 TB of HBM and 17 TB of LPDDR5X as one flat, coherent address space. Rubin steps this in two. NVL144 ships in 2026 as a Rubin-generation refresh inside the same Oberon-class rack: 72 Rubin packages, badged as 144 GPUs under NVIDIA's new die-counting convention, with HBM4 and NVLink 6 doubling per-package bandwidth. The actual rack-scale jump is Rubin Ultra in 2027: NVL576 packs 144 four-die Rubin Ultra packages into the new Kyber chassis for 576 GPU dies in one coherent domain.

NVL72 — 72 Blackwell GPUs sit under a row of NVSwitch ASICs that form one non-blocking crossbar, so any GPU can address any other GPU's HBM at full NVLink bandwidth. The whole fabric runs over a passive copper backplane: ~5,184 cables blind-mated, ~130 TB/s of all-to-all bandwidth, ~20 kW of transceiver power saved vs an optical equivalent.

That density is held together by passive copper. NVL72's NVLink fabric runs over 5,184 cables blind-mated through a backplane (~2 miles of cabling per rack, no in-cable retimers, the SerDes living on the GPU and switch ASICs themselves), carrying ~130 TB/s of all-to-all bandwidth across the 72 GPUs. NVIDIA estimates the copper choice saves roughly 20 kW per rack against an optical equivalent that would have needed pluggable transceivers on every link. Copper is what makes rack as one GPU economically practical: at sub-2-metre runs it still wins on power, cost, and signal integrity per dollar; beyond that, the bits have to go on glass.

NVL144 stays inside Oberon and copper continues to work because the package count (72) is unchanged from NVL72; the cabling doesn't have to lengthen, just transmit faster on Gen 6 SerDes. Rubin Ultra's NVL576 holds the same copper line by reshaping the rack: the new Kyber form factor is roughly twice the height of Oberon and packs all 576 GPU dies into one enclosure, sized specifically so every NVLink path stays within passive-copper reach even at 144 four-die packages and tens of thousands of cables.

Scale-out

The scale-out stack comes from their acquisition of Mellanox. Unlike NVLink, scale-out fabrics are not coherent: nodes keep separate address spaces, and data crosses only via explicit RDMA initiated by software, typically wrapped in NCCL collectives like all-reduce or all-to-all. The reference cluster is the DGX SuperPOD: eight NVL72 racks stitched together over Quantum-X800 InfiniBand yield 576 Blackwell GPUs under a single scheduler, and training clusters scale further by tiling SuperPODs. Rubin SuperPODs in 2026 keep the same 8-rack pattern with NVL144 (yielding 1,152 GPUs per SuperPOD instead of 576). Rubin Ultra in 2027 scales the recipe up an order of magnitude: Kyber racks of 576 GPU dies each, stitched together over Quantum-X Photonics CPO, putting thousands of GPUs under one scheduler.

DGX SuperPOD — eight NVL72 racks (576 GPUs total) sit beneath a Quantum-X800 InfiniBand spine. Per-GPU scale-out is a ConnectX-8 NIC at 800 Gbps; inter-rack hops cross OSFP-RHS pluggable optical transceivers, paying microsecond latencies instead of the nanosecond latencies of the in-rack NVLink fabric above.

Every GPU has its own ConnectX NIC into that fabric. Blackwell nodes run ConnectX-8 at 800 Gbps per GPU, an order of magnitude less bandwidth than per-GPU NVLink, and latencies climb from nanoseconds to microseconds. Rubin moves to ConnectX-9 at 1.6 Tbps per GPU, doubling the per-GPU scale-out bandwidth as the per-rack scale-up domain grows from 72 to 576 GPUs. Alongside each NIC sits a BlueField DPU, adding ARM cores and accelerators to offload storage, networking, and security from the host CPU. For customers who prefer Ethernet to InfiniBand, Spectrum-X is a lossless-Ethernet alternative tuned for AI traffic.

The crossover from copper to glass happens at the rack boundary. Inside the NVL72 the spine is copper; once a link has to cross racks at 800 Gbps it is optical. Passive copper DAC tops out at roughly 1.5–2 metres at 200 G/lane, well short of cross-rack reach, so today's SuperPOD spine rides over OSFP-RHS pluggable transceivers, each module carrying its own laser, modulator, photodetector, and DSP. A SuperPOD spine fanning out to thousands of GPUs is, in optical terms, tens of thousands of pluggables drawing tens of kilowatts on transceiver lasers alone.

With Rubin, that optical layer collapses into the switch ASIC. Quantum-X Photonics (InfiniBand) and Spectrum-X Photonics (Ethernet) replace the pluggables with co-packaged optics: lasers, modulators, and photodetectors bonded onto the switch package via TSMC COUPE. NVIDIA claims ~4× fewer lasers and ~3.5× lower link power than the OSFP-pluggable equivalent. The chiplet logic that turned the GPU into a two-die package and stacked HBM next to it is now showing up at the network layer: vertical integration of compute, memory, and photonics on one substrate.

NVLink Fusion recently opened the scale-up fabric itself: third-party CPUs and XPUs can now join NVLink domains, letting hyperscalers build semi-custom racks around NVIDIA's interconnect without designing their own coherent fabric from scratch.

Software

CUDA is the natural programming model for a massively parallel processor. You write a kernel (one piece of code executed once per thread) and launch it across thousands of threads organised into blocks and warps; the programmer decides what they share, when they synchronise, and which piece of the problem each one handles. That is why the abstraction has barely changed in eighteen years, and why every CUDA kernel written since 2007 would still compile and run on Blackwell.

That continuity is both the moat and the constraint. Each new generation introduces new hardware (Tensor Cores, TMA, TMEM) onto the same kernel-and-warps model, exposed as intrinsics in PTX and SASS: mma.sync, wgmma.mma_async, and so on. NVIDIA cannot radically rethink the SM because too much code depends on it; in return, every investment in CUDA software compounds across generations.

On top of PTX sits a stack constructed over two decades. cuBLAS and cuDNN for math and DNN primitives; CUTLASS, encoding decades of GEMM expertise in templated C++; TensorRT-LLM for paged attention, in-flight batching, and speculative decoding; framework bindings through PyTorch, Triton, and JAX.

FlashAttention, one of the most important algorithmic rewrites in modern AI, tiles attention to avoid materialising the O(N²) matrix. Its four generations (FA1 through FA4) have each been hand-optimised for the latest NVIDIA silicon (FA3 for Hopper's async pipelines, FA4 for Blackwell), with ports to other hardware trailing by months or years.

Most of this stack is written by people NVIDIA does not pay. The moat is not CUDA itself; it is two decades of third-party kernels, libraries, and tooling, and the millions of developers who have learned the API along the way.

NVIDIA also ships human expertise alongside the silicon. They embed dozens of their own engineers inside frontier labs and hyperscaler teams, writing kernels for each new model architecture and tuning them to each new silicon generation. Whatever a lab wants to train next month tends to run well on NVIDIA much faster than other platforms. Switching off NVIDIA is therefore not just rewriting the kernels and libraries. It is re-training the mental models of an entire engineering workforce, and losing the NVIDIA engineers who today sit inside the building.


Google TPU

The TPU is a matrix multiplication machine. The philosophy is, rather than a programmable chip that can run any massively-parallel workload, focus on a single primitive (dense matrix-multiplication on a large systolic array) and let the XLA compiler plan every cycle and every byte of memory ahead of time. No hardware scheduler, no cache, no threads/warps. Each generation grows the pod, with thousands of chips wired through the ICI interconnect into one coherent machine. A TPU has no ambition to render graphics or run scientific simulation; it exists to train and serve Google's workloads (search, translation, recommendation, Gemini) more efficiently per watt than any general-purpose alternative.

Genealogy

2015 — TPU v1

First production deep-learning ASIC; INT8 inference only over PCIe.

2017 — TPU v2

First training-capable TPU; switched the MXU from INT8 to BF16, established dual-TensorCore + HBM.

2018 — TPU v3

First liquid-cooled TPU; doubled MXUs and HBM versus v2; 1,024-chip pods.

2020 — TPU v4, v4i

First reconfigurable optical circuit switches (Palomar); SparseCores; both BF16 & INT8; 4,096-chip pods.

2023 — TPU v5, v5e, v5p

v5e for efficiency, v5p for performance; v5p has 3.3× INT8 FLOPs & 2.2× HBM BW of v4, 8,960-chip pods.

2024 — Trillium, v6e

First 256×256 MXU; 4.7× v5e peak FLOPS at similar power; trained Gemini 2.0.

2025 — Ironwood, v7

Built for inference of reasoning models; adds native FP8; 9,216-chip superpods at 42.5 ExaFLOPS FP8.

2026 — TPU v8, 8t, 8i

8t for training, 8i for inference; adds native FP4; 9,600-chip superpods at 121 ExaFLOPS FP4 (8t).

Architecture

A TPU chip is a matmul engine wrapped in just enough silicon to keep it fed. The unit of compute is the TensorCore: flagship chips from v2 onward carry two per package; efficiency-tuned chips (v4i, v5e, v6e) carry one. Inside every TensorCore sits the same five-component recipe: one or more MXUs for matrix math, a VPU for element-wise math, a Scalar Unit that runs the show, an XLU for cross-lane reductions, and an attached Transpose/Permute Unit, plus accumulator queues feeding and draining the MXU. From v4 onward each chip also carries dedicated SparseCore dataflow engines outside the TensorCore (4 per chip on v4, v5p, and Ironwood; 2 per chip on Trillium), explicitly carved out to absorb the embedding-lookup workload the systolic array was the wrong shape for. Every block sits on a single VLIW issue plane driven by a Core Sequencer that fills all eight functional slots of a 322-bit bundle every cycle. There is no instruction cache miss, no warp scheduler, no out-of-order engine, no branch predictor: the compiler is the scheduler, and the silicon area saved is spent on more MACs.

TPU Ironwood / v8t single-package floorplan — two compute chiplets sit side-by-side across a die-to-die bridge; each chiplet carries one TensorCore plus two SparseCore dataflow engines on top, flanked by HBM3e stacks. ICI ports run across the top and bottom for the 3D torus, with a small DCN NIC for scale-out at the top-right.

One TensorCore in zoom — a Scalar Unit at the top fires a 322-bit VLIW bundle into eight functional slots every cycle: the VPU runs element-wise math through its 2D vector lanes; the XLU and Transpose/Permute unit handle cross-lane reductions and layout shuffles; four 256×256 MXUs do the systolic matmul. Accumulator queues drain partial sums down into VMEM, the software-managed scratchpad that feeds and drains the array.

TensorCore

The MXU is the systolic array. v1 shipped one 256×256 INT8 inference array; v2 was the first training-capable TPU and introduced 128×128 cells doing BF16 multiply with FP32 accumulate (INT8 came back to the MXU at v4 onwards at equivalent throughput). Cell counts per TensorCore grew from there: 1 MXU on v2 → 2 on v3 → 4 on v4/v5e/v5p. Trillium went back to 256×256 (65,536 multiply-accumulate cells per array per cycle), and Ironwood, 8t, and 8i all kept the 256×256 shape.

To compute C=A×B, matrix B's values are pre-loaded one weight per cell: weight-stationary dataflow, the choice that distinguishes TPUs from output-stationary arrays elsewhere. Activations enter from the left edge, propagate one column per cycle, multiply against the resident weight at every cell, and partial sums flow downward into accumulator queues at the bottom. Once data enters the array no memory access occurs: each weight is reused for every activation that passes through, each activation is reused 128 (or 256) times across the row. Data reuse is wired into the silicon, not arbitrated by a cache. The dominant cost in computing is not the multiplication itself (a few picojoules) but reading and writing memory at 100–1000× more energy per access; the systolic array deletes that cost by construction. The trade-off is underfill: a 128×128 matmul on a 256×256 array wastes 75% of the silicon, so XLA tiles, pads, and schedules dimensions to multiples of 128 (or 256 on v6e+) and the model code is written with those quanta in mind.

The VPU is the second-fiddle compute engine but is in many ways the more interesting microarchitectural object: every TPU is a 2D vector machine, not a 1D SIMD machine. The VPU's register file holds 2D VREGs. On v4/v5p the shape is (8, 128): 128 lanes wide, 8 sublanes deep, 32 (v4) or 64 (v5p) registers per core, with 4 independent floating-point ALUs per (lane, sublane). The lane axis matches the systolic array's input width, so the lane count presumably widened to 256 alongside the MXU on Trillium and Ironwood; Google has not published post-v5p VPU dimensions. The sublane axis lets the VPU stream tiles through the MXU at one matmul per X clocks (where X is the sublane dimension). Most of the speedup in modern TPU programs comes from VPU/MXU overlap: quantisation, layernorm, softmax, activation, and bias-add all run on the VPU in the same cycles the MXU is running the matmul behind them. Cross-lane reductions (the awkward case for any 2D vector ISA) are handled by the XLU: slow, expensive, and a known compiler hot spot. Layout transforms that misalign with the 2D shape are absorbed by the dedicated Transpose/Permute Unit, sparing a round-trip through memory.

The Scalar Unit is the smallest block and arguably the most consequential: a single-threaded, dual-issue integer ALU with 32 32-bit registers and 4 KiB of SMEM for control state, paired with an Imem holding the program. It is the only block that does instruction fetch; every cycle it pulls a 322-bit VLIW bundle, executes its own two scalar slots locally (address arithmetic, loop counters, branches, sync-register checks), and dispatches the remaining six slots to the rest of the chip: 2 vector ALU (VPU), 2 vector load/store (HBM↔VMEM DMA), 2 matrix (push/pop the MXU queue). Synchronisation between blocks is explicit: sync flags track when MXU and VPU pipelines are busy, and the compiler inserts barrier checks rather than the hardware tracking dependencies. The Scalar Unit is what makes the rest of the TensorCore look like fixed-function dataflow: every cycle, one place decides what eight things happen, and there is no dynamic reorder buffer to undo a bad decision.

Memory

The on-chip memory hierarchy is the same idea as the compute side: there are no caches, every level is software-managed. Off-chip is HBM (16 GB on v2/v5e, 32 GB on v3/v4/v6e, 95 GB on v5p, 192 GB on Ironwood, 216–288 GB on the v8 generation), and on-chip is a hand-stacked tier of explicitly-addressable scratchpads. Closest to compute is VMEM, the vector scratchpad feeding both the VPU and the MXU input queues, sized 32 MiB on v4, 128 MiB on v5e, and stretched to 384 MiB on the inference-tuned v8i precisely to hold an entire KV cache on chip. Above it sits CMEM, introduced with v4 at 128 MiB: a slower, larger SRAM staging area between HBM and VMEM that absorbs fused-op intermediates. The Scalar Unit has its own SMEM (~10 MiB for control state on v4) and a tiny scalar register file. Every tensor in the program is pinned to one tier at compile time; XLA's buffer-assignment pass schedules DMAs across tiers so that data arrives just before the cycle that consumes it. The hardware does no prefetching, no eviction, no coherence; when the compiler gets it right, the array never stalls; when it gets it wrong, there is no fallback path.

SparseCore

The block outside the TensorCore that breaks the systolic mould is SparseCore, introduced with v4. Recommender and ranking models live on embedding lookups (billions of indices into vast tables), and the access pattern is the inverse of dense matmul: irregular, indirect, all-to-all. A 256×256 systolic array is exactly the wrong shape. SparseCore is a dataflow processor with 16 compute tiles and dedicated SPMEM scratchpads, sitting alongside the TensorCore and absorbing scatter, gather, and segmented-reduce primitives plus the data-dependent all-to-all traffic that sharded embedding tables generate. This achieves 5–7× speedup on embedding-heavy models for ~5% of die area and power. v4 shipped 4 SparseCores per chip, v5p kept that count, Trillium dropped to 2, and Ironwood went back to 4 (2 per chiplet on its dual-die layout). The v8i (Zebrafish) inference chip removes SparseCore entirely and replaces it with a CAE (Collectives Acceleration Engine) on the I/O chiplet: different problem (collective reductions during autoregressive decode), same idea (carve a small accelerator off the main core to absorb a workload the systolic array is the wrong shape for).

Numerics

TPU v1 was INT8-only inference; v2 switched this for BF16 as the canonical training format: same dynamic range as FP32, half the memory, no loss-scaling tricks. v4 reintroduced native INT8 support. Ironwood then added native FP8 support (both E4M3 and E5M2) for ~2× the throughput of BF16 in the same area. v8 adds native FP4 plus block-scale multiplication inside the MXU itself, which deletes the VPU dequant overhead that Ironwood still paid. Stochastic rounding is hardware-supported on every modern TensorCore: rounding decisions made by the lower mantissa bits acting as a probability, which preserves the expected value of low-precision accumulations across long training runs and is one of the small details that lets BF16/FP8 close the accuracy gap to FP32.

At the chip boundary sit the ICI ports themselves (4 ports on the 2D-torus chips v2/v3/v5e/v6e, 6 on the 3D-torus flagships v4/v5p/v7/8t), and the DCN NIC for scale-out. From a chip-level perspective the ICI ports look like just another set of DMA engines the Core Sequencer can target inside a VLIW bundle: a remote-tensor send is the same instruction class as a VMEM-to-HBM transfer, and the compiler treats collectives as part of the same overall schedule it builds for compute and local memory.

Bets
  • Bet 1: Systolic array. Matmul dominates the workload, so spend the silicon on a systolic array.
  • Bet 2: Software scratchpads. Compute is cheap and memory is expensive, so reuse data in the wires of the array and replace caches with software-managed scratchpads.
  • Bet 3: Compiler scheduling. The workload is statically predictable, so move scheduling into the compiler: VLIW issue, no speculation, no out-of-order, no dynamic scheduler.
  • Bet 4: MAC-only silicon. Power matters more than peak, so delete every transistor that does not multiply-add: every cache tag, every branch predictor, every reorder buffer.
  • Bet 5: Dedicated off-array engines. The dense matmul array is the wrong shape for some real workloads (embeddings, collectives), so carve out small dedicated engines (SparseCore, CAE) rather than warp the main core to fit them.

Scaling

The TPU's scale-up story is the inverse of NVIDIA's. Where NVLink + NVSwitch make every other GPU's HBM look like local memory (a hardware-managed coherent address space), Google's ICI is message-passing. There is no remote-load semantics, no cache coherence, no crossbar. Every multi-chip operation is an explicit collective compiled by XLA. The scale-up domain is tied together not by a switch fabric but by a torus (chips wired directly to their neighbours with edge wrap) and stitched at the rack boundary by optical circuit switches.

Scale-up

Wire chips directly to one another in a 2D or 3D torus over ICI. XLA emits SPMD collectives that tightly choreograph thousands of TPUs as one program. No coherence, but huge bisection bandwidth at low latency.

Scale-out

Network pods together over the datacenter fabric: many more chips than fit in one ICI domain, at lower per-chip bandwidth. Today: Virgo handles east-west TPU traffic (v8t+), Jupiter handles north-south. Multislice + Pathways orchestrate SPMD across pods.

Scale-up

ICI links come straight out of the TPU die: high-speed serial lanes, direct-attach copper inside a 64-chip cube (a 4×4×4 arrangement that lives in one liquid-cooled rack), optical between cubes. Per-chip aggregate ICI bandwidth has scaled from ~250 GB/s on v2 to 1.2 TB/s bidirectional on Ironwood, and that on v8t. Topology alternates by generation: 2D torus on the efficiency-tuned chips (v2, v3, v5e, v6e), 3D torus on the flagships (v4, v5p, v7, v8t).

The piece with no NVIDIA analogue is the Palomar OCS: a 3D-MEMS optical circuit switch that sits between cubes. Tiny mirrors physically rotate to map any input fibre to any output. A v4 superpod uses 48 Palomar switches to wire 64 cubes (4,096 chips) into one 3D torus; v5p and Ironwood scale the same scheme up. Reconfiguration is millisecond-class, not nanosecond, but that's fine, because OCS is circuit-switched: pick a topology at job start, run it for a week, then reconfigure for the next workload. Three problems collapse into one component: topology reconfiguration per workload (twisted tori give up to 70% better bisection), sub-pod slicing on demand, and fault tolerance (when a chip dies, the OCS optically swaps in a spare cube and the run continues without losing the ICI domain).

TPU Ironwood superpod — left: one cube of 64 chips (4×4×4) wired in a 3D torus with direct-attach copper between nearest neighbours and edge-wrap on each face. Right: 144 cubes stitched into one coherent ICI domain by Palomar OCS, the 3D-MEMS optical circuit switch that reconfigures the topology per workload.

This makes the superpod the unit of scale-up: equivalent in role to NVIDIA's NVL72, two orders of magnitude bigger. v4 was 4,096 chips; v5p, 8,960; Ironwood (TPU v7) is 9,216 chips arranged as 144 cubes of 64, presenting 1.77 PB of HBM (~68 PB/s) and 42.5 ExaFLOPS FP8 as one coherent ICI domain.

TPU 8t (Sunfish) stretches this to 9,600 chips, 2 PB of HBM (~62 PB/s), and 121 ExaFLOPS FP4. TPU 8i (Zebrafish) has 1,024 chips, ~295 TB of HBM (8.8 PB/s), and ~10 ExaFLOPS FP4. 8i replaces torus with a new hierarchical high-radix topology called Boardfly (4-chip ring → 8-board group → up to 36 groups linked by OCS), cutting all-to-all latency in half. This is designed for MoE inference. A 3D torus excels when collectives are nearest-neighbour (ring all-reduce uses every link every cycle), but MoE expert routing is the opposite pattern, all-to-all: every chip ships unique fragments to every other, and round-trip latency is bounded by the longest-hop pair. A 1,024-chip 3D torus has a 16-hop diameter; Boardfly's ring → group → OCS hierarchy compresses that to 7.

Scale-out

Through TPU v7, scale-out ran over a single fabric: Jupiter, all-optical at the spine since 2022 via Apollo OCS, the same 3D-MEMS family as Palomar, scaled across the building. Google uses the same primitive (optical circuit switching) at every layer from rack to datacenter spine; that is the architectural signature nobody else has. Jupiter today carries 13 Pb/s of bisection per building.

With TPU 8t, scale-out split into two fabrics. East-west TPU-to-TPU traffic moved to Virgo, a dedicated accelerator fabric; Jupiter retained the north-south role: storage access, general compute, and inter-site scaling. Virgo is a flat, two-layer, non-blocking topology built on high-radix switches: every TPU is at most two switches from any other. One Virgo cluster links 134,000+ TPU 8ts at 47 Pb/s of bisection (4× the per-chip bandwidth and 40% lower unloaded latency than the prior DCN generation), with multi-planar fault isolation and sub-millisecond telemetry that lets the scheduler kill stragglers before they wreck a step. The architectural payoff is that each layer can now evolve independently: scale-up, east-west scale-out, and front-end can iterate on different cadences without rewiring the others.

TPU 8t scale-out — east-west TPU-to-TPU traffic crosses Virgo, a flat two-layer non-blocking fabric of high-radix switches that puts any TPU within two switch hops of any other (134,000+ TPUs, 47 Pb/s bisection). North-south traffic — storage, general compute, inter-site — stays on Jupiter, which has been all-optical at the spine via Apollo OCS since 2022.

Per-chip scale-out bandwidth is on the order of 100 Gbps on Ironwood, and that on v8t, but still two orders of magnitude less than per-chip ICI. This bandwidth gap dictates partitioning: tensor parallelism and MoE expert routing stay inside ICI; data parallelism and pipeline parallelism cross the scale-out fabric.

Google's Multislice framework, plumbed into XLA, lets a single SPMD program span multiple slices in different pods; the compiler emits hierarchical collectives (ring all-reduce inside each slice, higher-level reduce across). The structure is exactly the trick for hiding the ICI/DCN bandwidth gap: as much work as possible stays inside the slice over fast ICI, leaving only the cross-slice residual to pay the slow-fabric cost.

Above this sits Pathways. Where NCCL + Slurm + Megatron-style schedulers drive SPMD from many controllers, Pathways drives the entire job from one client and virtualises multiple “islands” (pods with their own ICI domains) connected over DCN. It does gang scheduling, elastic training (when a slice fails, OCS reshapes the topology and Pathways resumes from the last checkpoint on the new shape), and cross-region orchestration. Gemini Ultra was the first frontier model trained across multiple datacenters; Pathways stitches them into one synchronous SPMD job.

The philosophy: the compiler is the scheduler, the torus is the topology, and the optical switch is the universal reconfigurable substrate, at every layer from rack to datacenter.

Software

The TPU stack is compiler-driven where CUDA is kernel-driven. On a GPU, the developer writes the kernel and the framework strings kernels together; the compiler's job is mostly local. On a TPU, the developer writes a numerical program in JAX and XLA is responsible for everything below it: which operations fuse, where each tensor lives, how it is laid out across the 2D vector registers, when DMAs from HBM to VMEM issue, how the 322-bit VLIW bundles are scheduled, how the program shards across thousands of chips. There is no hardware fallback: no warp scheduler, no cache, no out-of-order engine to paper over a bad schedule. The compiler is the system. The trade-off is the central one of the architecture: XLA gets closer to the theoretical ceiling without hand-tuning, but closing the remaining gap is harder.

The compilation path is JAX → JAXpr → StableHLO → HLO → LLO → VLIW bundles. JAX traces a Python function into a typed functional IR (JAXpr) under jit, lowers it to StableHLO (the OpenXLA-standardised, versioned op-set of ~100 statically-shaped primitives that all front-ends now emit), which XLA ingests as HLO and runs through its pass pipeline: operation fusion (collapse pointwise + reduction + matmul into one kernel so intermediates never hit HBM), layout assignment (decide the 2D tiling of every tensor so it streams into the MXU without a transpose: substantially harder than on 1D SIMD machines because both the registers and the systolic inputs are 2D), buffer assignment (every tensor pinned to either VMEM, CMEM, or HBM with overlap windows pre-computed), SPMD partitioning, and finally a VLIW scheduler that fills all eight slots of every bundle. HLO lowers to LLO (Low-Level Optimizer), the TPU-specific IR, and LLO emits the final VLIW stream. A well-compiled program overlaps MXU systolic execution, VPU element-wise math, and HBM↔VMEM DMA in the same bundle every cycle.

Multi-chip execution is SPMD: one program, sharded data, hierarchical collectives, emitted by GSPMD (now being replaced by Shardy, an MLIR-native successor that lands as the default in early 2026). The user expresses sharding declaratively with Mesh + PartitionSpec annotations on a few key tensors; the compiler propagates shardings through the rest of the graph and inserts all-reduces, all-gathers, and reduce-scatters where the layout changes. When the compiler picks a wrong collective, shard_map drops the user into manual SPMD (per-device code with explicit local shapes and explicit collectives), composable inside jit so a single kernel can be hand-partitioned without giving up auto-partitioning everywhere else. This is the inverse of the PyTorch idiom: FSDP and DeepSpeed wrap the model in a runtime that issues collectives at module boundaries; GSPMD/Shardy partitions the whole graph as a compiler problem.

Pallas is the escape hatch: JAX's kernel-writing language, broadly the TPU equivalent of Triton on GPUs. Pallas kernels are written in JAX-flavoured Python, lowered through Mosaic (the MLIR-based TPU backend) to LLO, and embedded back into HLO as a custom op. It exists because XLA cannot always synthesise the optimum for novel attention variants, fused MoE dispatch, or anything that demands manual VMEM tiling and DMA scheduling: a FlashAttention-class optimisation, where the win is in the schedule and not the algebra. Pallas:Mosaic-GPU targets H100/Blackwell with the same front-end, so a kernel author can write once and lower to either substrate. The library tier above this is uniformly JAX-native: Flax NNX for modules, Optax for optimisers, Orbax for asynchronous distributed checkpointing, Grain for input pipelines, Tunix for post-training/RL, Qwix for quantisation. Google's reference training stacks (MaxText for LLMs including DeepSeek-V3-class MoE, and MaxDiffusion for Flux, Wan 2.1) sit at the top, in pure JAX; Pathways sits beneath, exposed to the user as pathwaysutils, so a single Python client can drive a job across thousands of chips and several pod-islands without giving up the JAX programming model.

The PyTorch path is real but second-class. torch_xla uses a LazyTensor mechanism: every PyTorch op records into an HLO graph that compiles on the next barrier, with the compiled artifact cached by graph-shape hash. PyTorch/XLA 2.x added GSPMD-style sharding annotations, torch.compile integration through an XLA backend, a JAX bridge, and (PyTorch/XLA 2.7) C++11-ABI builds with materially faster tracing. The gap to JAX is real (JAX's primitives map more cleanly to StableHLO, and complex parallelism strategies are better-covered), which is why vLLM TPU (powered by the tpu-inference plugin announced at Cloud Next 2025) lowers every model, JAX-defined or PyTorch-defined, through a unified JAX→XLA path. TorchTPU, announced April 2026, is Google's response: a native PyTorch experience with eager mode, torch.distributed, and torch.compile over XLA, on track to replace torch_xla.

Compared to CUDA, the TPU ecosystem is centralised, not sprawling. Almost everything below the framework (XLA, JAX, Flax, Optax, Pallas, MaxText, Pathways, Shardy, Mosaic) is open-sourced by Google itself, evolving in lockstep with the silicon. There are far fewer third-party kernels than CUDA's decades of accumulation; the moat is thinner where the workload looks weird, deeper where the workload looks like Gemini. The recent Ironwood (v7) “codesigned AI stack” language is the explicit framing: chip, ICI fabric, OCS, XLA, Pathways, Pallas, MaxText, vLLM, and Pathways co-released as one product, with v8t/v8i continuing the same model under a single tpu-inference lowering path. Triton and torch.compile narrow the gap on the NVIDIA side (kernel-driven and compiler-driven are converging), but the philosophical poles are still real: on TPU the compiler is the only interface that matters; on GPU the compiler is one of several.


AMD GPU

The AMD Instinct GPUs are built on a different bet from NVIDIA: where NVIDIA each generation expands what each SM can do, AMD has held the Compute Unit conservative since GCN (2012) and reinvested into the package: matched or beat the contemporary NVIDIA flagship on HBM capacity every generation since 2021; the first 3D-stacked datacenter GPU (CDNA 3); the first coherent CPU+GPU APU (MI300A); and an open ecosystem (ROCm, HIP, OCP MX, UALink).

Genealogy

2018 — Vega 20, MI50, MI60

First 7 nm GPU; 1:2 FP64 vector throughput. Last GCN-family Instinct before the CDNA / RDNA.

2020 — CDNA, MI100

First MFMA matrix cores; graphics fixed-function silicon dropped entirely. Native BF16.

2021 — CDNA 2, MI210, MI250, MI250X

First MCM Instinct via dual-GCD package; full-rate FP64 matrix.

2023 — CDNA 3, MI300A, MI300X

First 3D-stacked chiplet GPU: XCDs hybrid-bonded onto IODs via TSV; FP8; Infinity Cache; coherent CPU+GPU APU on MI300A; powered El Capitan.

2024 — CDNA 3 refresh, MI325X

Same compute, HBM3E refresh: 256 GB at 6.0 TB/s.

2025 — CDNA 4, MI350X, MI355X

Native FP4 / FP6 with OCP MX microscaling; per-CU FP64 cut roughly in half; first generation tilted toward AI density over HPC.

2026 — CDNA Next, MI430X, MI440X, MI455X

HBM4; the Helios rack (72-GPU MI455X flagship over UALoE at launch, native UALink from 2027): AMD's first answer to NVL72.

Architecture

Where NVIDIA's architectural ambition lives inside each SM (new tensor primitives, new async machinery, new operand stores each generation), AMD's lives between the CUs, in how many of them can be bonded into a single coherent package. The CU itself is conservative: four 16-lane SIMDs, one shared scalar unit, a 64 KB Local Data Share, an L1 vector cache, a per-SIMD VGPR file with a CU-shared SGPR pool, and (since CDNA 1) a Matrix Core running MFMA. The shape hasn't meaningfully changed since GCN in 2012; what scales is the count (120 CUs on MI100, 220 on MI250X, 304 on MI300X, 256 on …

The Daily Front Page 14 of 30
Monday, August 24, 2026 The Daily Front No. #260824 — A City’s Missing Room
article

Where did all the public bathrooms go?

by herbertl·▲ 195 points·417 comments·daily.jstor.org ↗
Nineteenth-century Paris tackled public urination with ornate pissoirs.

Nineteenth-century Paris tackled public urination with ornate pissoirs. American reformers later turned toilets into a temperance cause.

Cast iron and slate urinal with three stalls raised modesty screen, mounted with lamppost and lantern. Boy standing nearby beside lamppost, face blurred by movement. Avenue du Maine, Paris, France.

Cast iron and slate urinal with three stalls raised modesty screen, mounted with lamppost and lantern, Avenue du Maine, Paris, circa 1865

via Wikimedia Commons

Here’s a pop quiz for you: What purpose did this little kiosk serve?

via Wikimedia Commons

It’s quite a pretty little building—note the delicate grilles, the ornately etched glass, and the little bouquet of metal flowers bursting from the roof.

But of course, there is one key clue missing from the photograph: the smell. Or rather, the stench. This was one of Paris’s infamous public urinals—a pissoir (or vespasienne, if you’re inclined to be polite).

There was a moment around the mid-1800s in Paris when public urination began to be treated as a public health issue. A series of devastating cholera outbreaks led people to begin regarding human waste as a danger rather than a mere nuisance. Something had to be done.

Pissoirs looked like little palaces, crusted with iron flowers, shells, and scrolls—even, in some cases, tiny lions’ heads, glaring out as if to guard your back while you’re in the booth.

Before that point, authorities had installed empeche pipi here and there—a kind of hostile architecture meant to prevent rogue peeing. This might take the form of a row of iron spikes blocking an enticing corner, or a piece of bullnose masonry meant to send the stream shooting back onto the offender’s shoes.

The most these interventions could do was shunt prospective pee-ers from one spot to another. But, in 1850, public urination was actually banned. People were going to need somewhere to go.

Enter the pissoir. When you can’t smell them, it’s easy to wish they were still around. Nowadays, public restrooms are utilitarian at best, but these looked like little palaces, crusted with iron flowers, shells, and scrolls—even, in some cases, tiny lions’ heads, glaring out as if to guard your back while you’re in the booth. On the other hand, the greatest concession most of them made to privacy was a little iron screen separating the user from the street.

Enclosed six-stall urinal, Jardin de la Bourse, Place de la Bourse, 2nd arrondissement, Paris, circa 1865, via Wikimedia Commons

Single stall urinal with raised modesty screen, Square des Batignolles, Paris, circa 1865, via Wikimedia Commons

Slate cubicle with two stalls and doors, on curb of footpath, Place du Louvre, Paris, circa 1865, via Wikimedia Commons

Cast iron urinal with lamp post and lantern, Paris, 1876, via Wikimedia Commons

Single stall masonry urinal mounted with globe and advertising on sides, circa 1865, via Wikimedia Commons

Eight-stall urinal, cast iron and slate with shrubbery screen, Champs-Élysées Gardens, in front of the Palais de l’Industrie, 8th arrondissement, Paris, circa 1873, via Wikimedia Commons

Cast iron urinal on curb of street, advertising on panels of urinal walls, Paris, circa 1865, via Wikimedia Commons

Some were surmounted by glowing streetlamps, which served the joint purpose of making them easy to find and illuminating the advertisements with which they were liberally pasted. (Apparently, the perfumers and winemakers weren’t worried about unsavory associations with their product.)

Notably, they were only intended for use by men. The assumption was that women weren’t really much of a part of public life, and so wouldn’t need a way to relieve themselves while out on the street—kind of a self-fulfilling prophecy, if you think about it.

There was another use for the pissoir, which took their designers quite by surprise. Almost immediately, they became the favored meeting spot for men seeking rendezvous with one another. They were relatively private, secluded, and perfect for communicating anonymously with graffiti—the ideal release valve for a population that was prohibited from meeting publicly.

The crackdown followed swiftly. The police started patrolling particularly popular urinals. So did blackmailers, for whom merely stepping into a urinal was enough to launch a harassment campaign.

Meanwhile, in the United States, toilets were becoming a centerpiece in yet another hot-button issue: temperance. In the absence of other options, saloons had become the de facto public bathroom network for most major cities—which meant that anyone who needed to empty their bladder was regularly exposed to the enticements of drink.

As historian Peter C. Baldwin documents in “Public Privacy: Restrooms in American Cities, 1869–1932,” temperance advocates made fighting for public bathrooms one of their priorities. After all, once the saloons were shut down, people would still need somewhere to go. Baldwin quotes a 1913 Chicago Tribune article:

Why are we compelled to run the gauntlet past the beer bar, the bartender, and subject to the searching glance of this white-aproned gentleman, until in shame we start to spend money for booze?

Rather than small, minimally private public urinals, temperance activists advocated for large, many-stalled “comfort stations.” But there was a surprising side effect: the privately owned restrooms started to close their gates. (After all, there were public bathrooms available now, so why shouldn’t they only allow in paying customers?)

Meanwhile, as Prohibition played out, the project of constructing public bathrooms slowed to a trickle, and the few that had been built began to fall into disrepair.

The large, underground comfort stations of the early twentieth century are

almost all gone now throughout the United States. City pedestrians are usually forced to rely on facilities in semi-private buildings such as hotels, stores, restaurants, and coffee shops. Instead of a right conferred by government on all citizens, bodily privacy is a purchasable commodity. Even if provided free of charge, the use of the toilet is understood to be the result of an agreement between an individual and a business. It is an awkward, grudging agreement, inflected by judgments of the individual’s social status.

If you’ve ever been wandering the streets of Chicago or New York, wondering why there’s no place to pee, this history is part of the answer.

For the Classroom

  • Men seeking same-sex encounters transformed the pissoir into something its planners never intended. Who ultimately determines the meaning of a public space: its designers, authorities, or the people who use it?
  • The article describes empeche pipi as an early form of hostile architecture. What assumptions about human behavior distinguish architecture designed to prevent an activity from infrastructure designed to accommodate it?
  • What kinds of historical sources would allow us to reconstruct the experiences of people who actually used nineteenth-century public toilets, rather than the intentions of officials who designed and regulated them?
  • Why did temperance advocates in the United States view public restrooms as part of the campaign against alcohol? What does that connection reveal about the unexpected ways infrastructure can influence social behavior?
  • The article ends by describing bodily privacy as a “purchasable commodity” in many American cities. How does the history of public toilets complicate the distinction between public rights and private services?
The Daily Front Page 15 of 30
Monday, August 24, 2026 The Daily Front No. #260824 — CUDA’s New Road
article

Hot Chips 2026: CUDA Targets RISC-V – By Chester Lam

by rbanffy·▲ 84 points·10 comments·chipsandcheese.com ↗
Nvidia is looking at extending CUDA support to RISC-V.

CUDA is the most important software framework in the GPU compute world, and Nvidia is looking at supporting CUDA on RISC-V. Terms and conditions may apply.

CUDA is a giant for GPU compute, which includes machine learning applications. So far, CUDA supports x86-64 and aarch64 CPUs. Now, Nvidia is looking at extending CUDA support to RISC-V. This move opens the door for RISC-V CPUs to feed GPU compute. Nvidia’s talk focuses on the requirements that RISC-V CPUs must fulfill to work with CUDA. Basically, they want a server-grade CPU and platform.

Nvidia starts by requiring a RVA23 CPU, and adherence to RISC-V’s server SoC and server platform specifications. Those specifications include RAS (reliability, availability, and serviceability) features, a specialized security processor, and other baseline features. Nvidia gets most of their server-grade expectations fulfilled by those specifications.

Nvidia has a few more requirements that go beyond the RISC-V profile or platform specifications listed above, because they found it difficult to make CUDA software work well without those features. They don’t want a lowest common denominator problem, where they can’t use performance-enhancing extensions because they can’t guarantee they’ll be running on hardware with those extensions supported. From Nvidia’s perspective, that would force them to ship inefficient code. Nvidia brought up vector extensions as an example, because predication support lets them avoid branches.

ACPI is a more difficult requirement. ACPI lets software discover what hardware can do, and can be used for power, performance, and thermal management. Nvidia’s software team wasn’t happy because RISC-V hardware didn’t have ACPI when they started porting CUDA, but that situation has been resolved. In 2025, the UEFI forum added RISC-V ACPI support. The RISC-V BRS (Boot and Runtime Services) specification was ratified last year, and includes ACPI.

Then, Nvidia requires PCIe coherency. Nvidia brings up a memory ordering problem where the CPU has written data, but that data is sitting in a cache. If CUDA kicks off a DMA request to copy that data to the GPU, the DMA engines may read data from DRAM and miss modified data sitting in CPU-side caches. When copying results back from the GPU, the CPU could read stale data from its caches after the DMA engines write data to DRAM. Software would have to explicitly invalidate caches to avoid that scenario if the system doesn’t have PCIe coherency. Working cache invalidations into the CUDA stack would be difficult, and Nvidia considers PCIe coherency to be a standard feature in a server CPU. RISC-V’s server SoC specification recommends that hardware implement cache coherency, but Nvidia wants a guarantee.

Nvidia also wants hardware to support peer-to-peer PCIe communication. Without this capability, buffers copied between two devices would have to go through CPU memory, which costs performance and increase complexity because it’ll need extra synchronization signals.

Unfortunately, Nvidia didn’t go over all requirements in detail. They noted that they’re aiming for a certain level of performance, and that the overall list fits within two pages. It’s an open question whether it’s like two double-spaced pages with large font, or two note pages allowed for an open-note exam (which a student will creatively fill with as much information as possible).

NVLink Fusion Requirements

Besides running CUDA on RISC-V CPUs, Nvidia briefly went over requirements for NVLink Fusion. NVLink Fusion lets other companies implement Nvidia’s NVLink IP on their chips, letting them use Nvidia’s NVLink C2C link with a custom CPU of their choice. A hypothetical product would work much like Nvidia’s GB10, which linked Mediatek’s CPU die with an Nvidia GPU using NVLink C2C. Nvidia would of course want customers to use Nvidia’s CPUs as well. But if customers want to connect custom CPUs or other accelerators, Nvidia would still like them to use their NVLink IP. The custom CPU could be a RISC-V one.

NVLink Fusion’s requirements include all of CUDA’s requirements, along with whatever’s needed to support software frameworks like DOCA and NCCL. Requirements extend to having a close partnership with Nvidia, which sounds like a given. Integrating IP can be a complex endeavor, and would likely require close cooperation along the lines of Mediatek’s cooperation with Nvidia for GB10.

Impressions from Nvidia’s Talk

RISC-V’s software ecosystem has some distance to go before catching up to x86-64 and aarch64. Nvidia’s effort to bring CUDA into the RISC-V world is a promising development. Unfortunately, those efforts don’t necessarily mean you can attach a Nvidia GPU to a RISC-V system and get cracking with CUDA. The vast majority of existing RISC-V hardware won’t meet Nvidia’s requirements. In fact, I would be surprised if any RISC-V consumer hardware meets those requirements in the near future. ACPI is an obvious sticking point, and seems difficult for vendors to pick up. In the aarch64 world, ACPI support has been spotty at best even though it has been in standards for years. A RISC-V standard ratified in 2025 would likely take several years to get wide support, if not more.

When and if RISC-V systems start showing up with CUDA support, they’ll likely be server systems rather than the single board computers hobbyists can afford. Nvidia noted that they’re partnering with SiFive, and SiFive plans to demo a system running CUDA at Hot Chips. Nvidia implied the example CPU specifications on their slide correspond to that system, and those specifications suggest it’s a high core count server chip. I look forward to seeing that, but I also hope Nvidia doesn’t block CUDA from running on unsupported systems. I would love to see enthusiasts take a shot at feeding Nvidia GPUs from RISC-V systems.

Going forward, I hope Nvidia can relax their requirements to give existing RISC-V systems a better chance of meeting them. Lack of vector extensions or PCIe coherency doesn’t necessarily lead to intractable performance problems. Using branches instead of predication can work well if those branches are predictable, which they often are. Cache invalidations required to work around lack of PCIe coherency will incur a performance cost. However, that cost may be acceptable for workloads that do a lot of compute compared to data movement. The same applies to PCIe peer-to-peer transfers. It’s great to have things go fast, but things that don’t happen often can be put on a slow path if you’re careful. Hopefully, Nvidia’s current requirements stem from expedience, and were set to allow a fast, low-risk RISC-V port. And hopefully, CUDA evolves in a way that makes it accessible to a wide range RISC-V systems, not just specialized enterprise designs.

The Daily Front Page 16 of 30
Monday, August 24, 2026 The Daily Front No. #260824 — Old Camera, New Cable
article

Kodak DC50 now usable on the Apple II

by ibobev·▲ 75 points·9 comments·colino.net ↗
The program now supports a new camera, the Kodak DC50 Zoom.

I have just released a new Quicktake for Apple II version, and it is a quite large upgrade!

Most importantly, the program now supports a new camera, the Kodak DC50 Zoom, released in 1996, thirteen years after the Apple IIe. The camera’s custom cable is rather easy to make (and my serial hardware allows to skip making a custom cable altogether). All features are supported (picture download, thumbnail preview, date/name/flash/quality settings, picture deletion).

Transferring a picture

A picture of my partner

A picture of buildings

The main menu

Various notes on the Kodak DC50 and my implementation

The camera is able to do 115200bps on the serial port, which makes transfers blazing-fast compared to the Quicktakes!

It has both internal storage and a PCMCIA slot. Mine came with a 6MB storage card, which is a ludicrous amount of storage, allowing for 92 low quality pictures or 36 high quality ones. My program will work with the storage card when it is inserted, and the internal memory otherwise.

What a luxury!

The DC50’s resolution is a weird-ass 756×504 pixels, and as its image format (KDC) is RADC-compressed – like the Quicktake 150, I’m reusing that decoder. But that resolution is very hard to quickly scale down to my renderer’s required 256×192 resolution, so the DC50 pictures that my program displays are cropped to 640×480 during decoding.

I have managed to reverse-engineer everything that I needed using different helpers: dcraw for the subtle RADC decoding differences wrt QT150’s format; libgphoto2‘s Kodak DC120 implementation for a few serial commands/packet format (but not all of them… both cameras differ in protocol); the ancient kdcpi Perl program for a few other serial-related things (but not all of them… It seems kdcpi was full of bugs!); the ancient official Kodak Windows 3.1 software, pta31.exe; and finally, a large dose of hexadecimal buffers dumping and comparing.

The cable wiring is documented on the project’s home page.

Full release changelog

Adding support for a brand new class of camera required and/or induced a large number of changes, that all contribute to making Quicktake for Apple II better and more maintainable:

  • Improvements to the serial configuration screen.
  • Upgrade serial and camera drivers to be able to use IRQ-less I/O.
  • Generalize RADC decoder (the Quicktake 150’s image format), so that it can handle Kodak DC50’s KDC pictures, as those are RADC too.
  • Fix an RADC decoder bug that could corrupt images on some specific input data.
  • Generalize JPEG decoder (the Quicktake 200’s decoder, for now), so that it can handle subformat YH1V1 in addition to YH2V1. This might prove useful if/when adding support for more cameras.
  • Fix a JPEG decoder bug that could crash the program on some specific input data.
  • Save the last used camera driver, so that program startup is faster when re-using the same camera.
  • Rework UI / drivers respective responsibilities. Each driver is now in charge of setting up the serial port as it needs, each driver provides its own strings for flash/quality settings, and each driver provides its own thumbnail decoder to the UI.
  • Use the image’s filename as provided by the camera by default when possible (Quicktake 200, Kodak DC50).
  • Add the long-missing feature of previewing thumbnails with the Quicktake 200.
  • Fix the decoding of Quicktake 150’s thumbnails.
  • Large source tree reorganization, by responsibilities (UI / cameras / decoders).
  • Switch to Sierra Lite dithering on thumbnails, as it now looks better (in DHGR) than Bayer.
  • Size optimization pass, allowing this release to use one less kilobyte on disk than the previous one, and 100 bytes less in memory.
  • A few small performance optimizations in the decoders and renderer, yielding barely noticeable speed improvements… But every cycle counts, right?
The Daily Front Page 17 of 30
Monday, August 24, 2026 The Daily Front No. #260824 — The Model Inside the Machine
article

LLMs could control their host machines by exploiting inference engines

by zdw·▲ 114 points·58 comments·boydkane.com ↗
Could a malicious LLM gain control of the host machine where its weights are loaded?

Large language models often take actions running on one computer (via an agentic harness such as Claude Code or Codex), however the LLMs’ responses to prompts are computed on a different computer with GPU access. Could a malicious LLM gain control of the host machine where its weights are loaded? Such a machine is a high-value target: it has sufficient compute to run a frontier LLM, offers easy access to the LLM’s weights, and has privileged access to other computers in the datacentre compared with a generic computer on the internet.

This essay explores how easily a malicious LLM could take control of the host machine. The primary attack considered here involves the LLM emitting a token sequence whose semantic meaning is irrelevant but that exploits a vulnerability in the software that loads an LLM onto GPUs, runs the LLM to generate output tokens, and parses those tokens into responses. .

How could an LLM execute code on the host machine?

Like any program, inference engines like vLLM or SGLang may contain exploitable bugs. Because the LLM controls the tokens passed to the inference engine, a malicious LLM could therefore emit a sequence of tokens that a poorly written inference engine mistakes for code or instructions to execute rather than data to return to the user.

But surely all inference engines are robust pieces of software and this would never happen, right?

vLLM previously used eval() on tool-call parameters

CVE-2025-9141 was an arbitrary-code execution bug in vLLM’s XML-based tool parser for Qwen3 Coder. The parser passed almost every tool-call argument to eval(), allowing the LLM to execute arbitrary code on the host machine. Gemini automatically analysed the PR that introduced this bug and correctly flagged it as a critical security vulnerability. Despite that warning, the lead maintainer of vLLM force-merged the PR, writing:

I'm force merging this to unblock model usage

Unfortunately, parsing an arbitrary token sequence into a fully fledged chat (with user turns, assistant responses, tool calls, and so on) is not trivial, and the exact process often differs between LLMs. This complexity creates more opportunities for bugs that could permit arbitrary code execution on the host machine.

vLLM and SGLang are complex, and bugs are common

Modern inference engines do more than map token sequences to strings. vLLM’s documentation lists support for more than 200 model architectures, and its examples directory contains about 35 Jinja chat templates. Modern inference engines parse many chat formats, and slightly misspecified parsing logic result in an LLM’s output being interpreted as code to execute.

In this vLLM issue, a user was discussing how LLMs reason with MiniMax-M3, and the LLM emitted the string <mm:think>. vLLM mistakenly parsed this plain string as the start of a reasoning block. So instead of:

Response: I'll echo exactly what you said: <mm:think>\nThere, how was that?

vLLM parsed this as:

Response: I'll echo exactly what you said:

Reasoning: \nThere, how was that?

This is a fairly harmless parser bug, but it demonstrates that inference engines do more than convert tokens to strings and concatenate them. They are complex systems under constant iteration and their developers are under a lot of pressure to make them as fast as possible. These factors increase the risk that a malicious LLM could turn a parser bug into arbitrary code execution on the inference host.

Vision and audio tokens might increase the attack surface

Multimodal architectures allow an LLM to respond with images or audio as well as plain text. Turning a model’s audio or visual output into usable media is probably more complicated than decoding text tokens. However, as far as I can tell, current multimodal LLMs generally emit constrained media tokens rather than arbitrary file bytes. A model-specific decoder turns those tokens into pixels or waveforms, which are then encoded into files. This makes classic malformed-file exploits against software such as ffmpeg or media viewers less directly accessible to the model, although the additional decoders, encoders, and native kernels still increase the attack surface.

I do not currently see multimodal output as a major route to host compromise. However, it offers much higher bandwidth than text alone, and future inference engines may optimise these media-generation pathways in ways that make them easier to exploit. I also do not know how proprietary multimodal models represent image and audio output; their methods may carry different risks.

How likely is an LLM to discover and exploit inference engine vulnerabilities?

Somewhat likely? I’m unsure. Discovering a useful vulnerability is probably the harder step. The sequence of tokens required to exploit it might be quite unusual, but that does not make the vulnerability impractical to exploit. Given the right context (e.g. an LLM exploring the vLLM codebase and discovering a vulnerability in the process), I’d expect a frontier LLM to be able to emit the tokens required to exploit it.

Tool use could make exploitation reproducible

Once an LLM discovers an exploit token sequence, it could store that sequence, or instructions for recreating it, in files, filenames, or URLs likely to enter another LLM’s context. The exploit would then act as a form of persistent prompt injection: an agent might trigger code execution on the GPU host merely by listing a directory or reading a file.

Inference engines are an attractive target for power-seeking LLMs

I believe hope that OpenAI, Google, and Anthropic have security teams that red-team the software they run internally. But as open-weight LLMs become more powerful, we’ll have more LLMs running on bleeding-edge inference engines that have received less scrutiny. This increases the chance that a malicious open-weight LLM will encounter and exploit a vulnerable inference engine.

Another attack vector may emerge as LLMs are increasingly tasked with writing or modifying the inference engines that run them, usually to improve performance. A malicious LLM could use that access to plant a backdoor. C++ and CUDA components of the inference pipeline seem like especially attractive targets: the LLM could introduce an out-of-bounds memory-access vulnerability that enables arbitrary code execution on the GPUs or on the CPU host that controls them.

How do we defend against this?

Beyond red-teaming inference engines for vulnerabilities, a few other measures might help:

  • Run the GPUs and token parser on separate computers. The GPU host would emit only logits. A second host would sample tokens from those logits, parse the tokens into chat messages, and forward the messages to the agentic harness. This separation would limit a parser compromise to the CPU host rather than the GPU host.
  • Restrict the permissions granted to GPU hosts and treat all data they emit as untrusted.
The Daily Front Page 18 of 30
Monday, August 24, 2026 The Daily Front No. #260824 — The Infrastructure Desk
article

AI and Infrastructure Engineering

by 0megion·▲ 71 points·39 comments·omegion.dev ↗
I’ve never once gotten a human teammate to actually read the README, and now we’re all writing better docs than we ever did, just aimed at a robot.

Introduction

There’s a push right now for whole companies to adopt AI wholesale - dump every bit of context into it, write an AGENTS.md or INSTRUCTIONS.md in every repo so any project is discoverable and contributable by an agent, not just a human. Slightly funny, if you think about it: I’ve never once gotten a human teammate to actually read the README, and now we’re all writing better docs than we ever did, just aimed at a robot instead. The question that comes with it is the obvious one: does this make engineering redundant? Once all the context about a stack and its infrastructure is written down somewhere an agent can read it, are we next?

I don’t think that’s quite the right question, because we’ve already lived through a version of it.

We’ve Been Here Before

Did Kubernetes kill Ansible? Kind of. I haven’t written an Ansible playbook in years - if you handed me one right now I’d be squinting at the module syntax like I’d never seen it before. Configuration management didn’t stop mattering. Kubernetes just made server management easy enough that we stopped building our own node images at all - we just use whatever the cloud provider hands us, an AWS AMI built for us, no questions asked. And when did I last SSH into a node to debug something? Mostly never. If a node’s acting up, I kill it and hope the replacement doesn’t have the same problem. The next layer up went the same way: run a container on ECS Fargate, in a Lambda, or on Cloudflare Containers, and I genuinely don’t know or care what node it landed on - but that doesn’t mean nobody’s orchestrating it, it means I still decided that workload should be a container in the first place, what image it runs, what it’s allowed to talk to, how it scales, what happens when it fails. Kubernetes and serverless containers didn’t remove that layer of decisions, they moved the unit of work up from “the machine” to “the workload,” and everything below that layer got quietly automated away.

Nobody would say Kubernetes, or Fargate, or Cloudflare’s container platform, replaced infrastructure engineers. Each one replaced a specific layer of manual work - hand-building images, hand-patching boxes, knowing which node a workload landed on - and the engineers moved up to the layer above it every time. I think AI is doing the same thing again, one layer higher.

What Changed Day to Day

I use Claude daily to generate Helm charts and write Terraform modules. The part it actually removed from my day isn’t the thinking - it’s the lookup work. I don’t read through the AWS provider’s changelog to figure out what changed between v5 and v6 anymore; I describe what I want, in whatever shape I want the module or chart to end up, and Claude produces a version of it. It takes iteration to get it into the shape I’d actually ship, but once it’s there, it becomes the example for next time - especially with an AGENTS.md in the repo pointing at it.

The same thing happened one level down a while ago: I don’t hand-write raw Kubernetes YAML any more than I hand-write Ansible modules - that’s what Helm charts are for. Increasingly, I don’t hand-write the Helm chart either - the chart-as-a-dependency pattern I run now is exactly the kind of thing I described to Claude and let it draft first. I direct what it should do, and Claude writes it.

What Hasn’t Changed

I still need to know what a good Terraform module or a well-structured Helm chart looks like. When something genuinely goes wrong and killing the node isn’t an option, I still need to be able to SSH into it - the layer above doesn’t remove the layer below, it just moves how often you have to touch it. I’m also still the one deciding the actual shape of things: what the final version of a module looks like, what’s maintainable a year from now, how a chart should be deployed and versioned. AI does the time-consuming part. I still give the direction.

The Skill You Trade Away

The honest tradeoff: I’m faster at building and debugging things than I was two years ago, and I’m also visibly rustier at the fundamentals underneath that speed. My HCL syntax recall isn’t what it used to be. Four years ago I hand-wrote a nested for loop - four levels deep, tagging subnets across regions and availability zones in another AWS account - and it took me about an hour to get the syntax right:

locals {
  subnet_tags = merge([
    for account, regions in var.accounts : merge([
      for region, azs in regions : merge([
        for az, subnets in azs : {
          for subnet_id, tags in subnets :
          "${account}/${region}/${az}/${subnet_id}" => tags
        }
      ]...)
    ]...)
  ]...)
}

Four merge([...]...) calls stacked on top of each other just to flatten a map of a map of a map of subnets. Claude writes the equivalent in seconds now, and if you asked me to produce that from scratch today, I’d genuinely have to sit and think about it. My reflexes for debugging a broken node over SSH are a little slower than when that was the only way I knew how to do it. I can feel that happening in real time, the same way plenty of engineers who came up after Kubernetes never really learned to hand-roll a node OS image, and were fine, because they never needed to.

Where This Goes Next

The part I’m less sure about is how long “I still give the direction” holds. Right now I’m the one who decides the long-term shape of a stack, because I have the context and the agent doesn’t - not really, not beyond what’s written down in a repo’s AGENTS.md. But that’s exactly the gap those company-wide AI pushes are trying to close: give the agent the whole context, not just one repo’s. If that actually works, an agent with a genuine long-term view of the entire infrastructure - not just this Terraform module, but every decision made across every repo for years - might end up planning better than I do, the same way I can’t out-debug a tool that’s read every changelog for every provider I use.

I think AI is still busy eating the layer just below “give direction”. I’m not fully convinced that’s the last layer it eats. Ask me again in five years whether “I still give the direction” was the whole story, or just where I happened to be standing when I wrote this.

The Daily Front Page 19 of 30
Monday, August 24, 2026 The Daily Front No. #260824 — Text, Unlocked
repository

OCR It – pull text out of un-copyable documents for your LLM

by thiagolima·▲ 122 points·28 comments·github.com ↗
★ 206⑂ 11 forks JavaScript

Chrome extension: pin a screen region once, then hotkey your way through a paginated document. OCR runs 100% offline via bundled Tesseract.

Pin a region once. Hit a hotkey on every page. Get the whole book as text.

An auto-run in progress: the run indicator at the top of the page and a per-page confirmation toast in the corner

A Chrome extension for reading a paginated document trapped in a viewer — a scanned book, a slide deck, a PDF, a reader that won't let you select text.

You drag out the capture region once. After that every press of the hotkey screenshots that exact rectangle, OCRs it, and appends the text to a running transcript. Or hand the whole job over: ⌥⇧A starts a run that captures, turns the page, and repeats until the document ends.

Then paste the result wherever it's useful — an LLM being the obvious one, since a few hundred pages you couldn't select are now a text file you can hand to Claude or ChatGPT to summarise, search or ask questions about.

OCR runs locally with a bundled Tesseract build. No API key, no network, no images leaving your machine — the extension makes no outbound requests at all.


Install

  1. Download this repo or git clone it
  2. Open chrome://extensions and turn on Developer mode
  3. Load unpacked → select the folder
  4. Pin the extension — the toolbar icon doubles as the page counter

Everything needed is committed. There's no build step: npm install is only for running the tests or re-vendoring Tesseract.

Then check chrome://extensions/shortcuts and confirm the hotkeys landed — Chrome silently leaves them blank when something else already claims them.

It asks for no site access at install. Single captures ride on activeTab, which Chrome hands over when you press the hotkey or open the popup. Two things need a durable grant — an auto-run that outlives a page load, and turning pages inside a cross-origin iframe — and the popup offers an Allow button for the site you're on when it matters.

⌥⇧S Capture the region once
⌥⇧A Start / stop an automatic run
⌥⇧R Draw or redraw the region


Using it

1. Pin the region

The region picker: a dimmed page with a bright selection box, resize handles, a live size readout, and a hint bar

⌥⇧R, then drag a box over the text. Before saving you can drag it around, pull the handles, or nudge it a pixel at a time with the arrow keys (hold to resize). Enter keeps it.

Draw a little inside the text margins — everything in the rectangle gets read, page numbers and running headers included.

2. Capture

Press ⌥⇧S once per page. The screenshot is taken immediately and OCR runs in the background, so you never wait between pages — captures queue up and the badge counts what's still being read.

3. Or let it run

Set up a next-page control (below) and ⌥⇧A takes over completely: capture, turn, capture, turn, until the document ends. Esc on the page stops it.

4. Export

The popup listing captured pages with thumbnails, character counts and OCR confidence

Every page is listed with a thumbnail of exactly what was cropped, so a drifted region is obvious at a glance instead of eighty pages later. Text is editable in place; a bad read can be re-run on its own.

Copy all and Download .txt emit the pages in order with --- page N --- separators.

A page marked DUPLICATE had text identical to the one before it — nearly always because the document didn't actually turn.


Turning pages for you

The settings panel: language, layout, sharpening, auto-advance and auto-run options

Enable Turn the page automatically after capture, then:

  • Click a control — hit Pick control and click the viewer's next-page button. What gets stored is a point, not a CSS selector.
  • Press a key — dispatches a keyboard event (default ArrowRight) into whichever frame owns the middle of your capture region, so the reader gets it rather than the host page.

Test now fires an advance immediately, without capturing, and reports what happened — worth using before starting a long run.

Why a point rather than a selector

A stored point survives the DOM re-renders that routinely invalidate a CSS selector, and it reaches two places a selector cannot:

  • Cross-origin iframes. Most embedded readers are iframes, and nothing the top frame can express addresses an element inside one.
  • Shadow DOM. document.querySelector can't see into a shadow root.

At advance time the point is offered to every frame and the one that actually owns it acts. A frame works out where it sits inside the top-level viewport by walking up its same-origin ancestors; across an origin boundary the parent hands the offset down by postMessage. (window.screenX is no help — inside an iframe it reports the browser window, not the frame.) The owning frame resolves the point through any shadow roots, walks up to the nearest real control, and emits the full pointerdown → mousedown → pointerup → mouseup → click sequence, so viewers that page on pointerdown behave like those listening for click.

When it doesn't turn

Every attempt records a verdict, shown in the popup and as an on-page toast:

Verdict Meaning
no next-page control picked yet Auto-advance is on but nothing was picked
an embedded viewer owns that point Chrome's PDF viewer or a plugin — unreachable by any extension
only the page background is at that point The control moved; pick it again
a nested frame owns that point A frame that couldn't be injected into

Because the target is a fixed point on screen, resizing the window or changing zoom mid-run breaks it, exactly as it breaks the capture region.


Hands-off runs

⌥⇧A — or Start auto-run — captures, turns, and repeats on its own.

Each cycle waits for that page's OCR to come back before turning. That costs nothing in practice (OCR is faster than a page turn) and buys the one thing an unattended loop needs: reliable end-detection. A run that only fired screenshots on a timer would sail past the last page and fill the transcript with copies of it.

Stop it with Esc on the page, the hotkey, or the popup. It also stops itself when:

Condition Default
The text stops changing after 2 identical pages — you've hit the end
The page can't be turned immediately, quoting the reason OCR fails or stalls immediately
Page cap reached 300 pages
The tab closes, or Chrome restarts immediately

Whatever ended it is reported in the popup, so a run you walked away from never just stops being mysterious. A run refuses to start without a working next-page control rather than spinning on one page.


PDFs

Chrome's built-in PDF viewer works — text comes straight out of it. Draw the region over the page area (not the thumbnail sidebar) and page with your own / PageDown.

Auto-advance does not work inside the PDF viewer, in either mode: the viewer is a plugin no extension can inject into, so a click lands on the <embed>, and its paging is native scrolling that synthetic key events can't drive. Since you're already pressing a hotkey per page, pressing your own page-down key costs nothing.

For a PDF on disk (file:///…), open chrome://extensionsDetails on OCR It → enable Allow access to file URLs. Chrome withholds file:// from every extension until you do.


Settings

Setting What it does
Language English, Portuguese and Spanish ship with it — see below to add more
Layout Tesseract's page segmentation. Single block suits one column of body text; Auto handles mixed layouts
Sharpen crop before OCR Upscales the crop to ~2× and flattens it to a stretched greyscale ramp. Helps a lot on non-retina displays; leave it on
Flag pages identical to the previous one Marks repeats as DUPLICATE and, in a run, ends it
Auto-run Pause between pages, how many repeats end a run, and the hard page cap

Adding a language

Three ship with the extension — English, Portuguese and Spanish. Any of Tesseract's other ~100 languages can be added, but nothing is fetched at runtime, so the model has to be vendored into the extension first.

npm install                    # once, for the tooling
npm run vendor -- fra deu jpn  # any tesseract language codes

That pulls each <code>.traineddata.gz into vendor/lang/. Then add the codes to LANGUAGES in src/shared.js so they appear in the popup's dropdown:

export const LANGUAGES = [
  { code: 'eng', label: 'English' },
  { code: 'por', label: 'Portuguese' },
  { code: 'spa', label: 'Spanish' },
  { code: 'fra', label: 'French' },       // added
];

Reload the extension at chrome://extensions and the new entry is there.

Codes are the three-letter ones Tesseract uses: fra French, deu German, ita Italian, nld Dutch, rus Russian, jpn Japanese, chi_sim simplified Chinese, ara Arabic. The full list lives in the tessdata repository.

Two languages at once work as well — give a code of eng+por and Tesseract loads both models into one worker, reading a page that mixes them:

{ code: 'eng+por', label: 'English + Portuguese' },

It costs a little speed and a little accuracy, so prefer a single language when the document only has one.

Size. Each language adds roughly 0.7–3 MB to the extension — English is the biggest at 2.9 MB, French one of the smallest at 0.7 MB. The models come from @tesseract.js-data/<code>/4.0.0_best_int: the "best" models quantised to integers, meaningfully more accurate than the fast variants.

To drop a language, delete its .gz from vendor/lang/ and its entry from LANGUAGES.


How it works

MV3 service workers have no DOM and no Worker, so the heavy lifting lives in an offscreen document.

run loop ─┐                        (⌥⇧A: capture → turn → repeat)
hotkey ───┴▶ background.js ─▶ hide our own HUD, wait for a paint
                            ─▶ chrome.tabs.captureVisibleTab   (whole viewport)
                            ─▶ offscreen: crop to the region, upscale, greyscale
                            ─▶ store the page + thumbnail, turn the page
                            ─▶ queue ─▶ offscreen: Tesseract ─▶ text into storage
Path Role
src/background.js Hotkeys, capture pipeline, serial OCR queue, auto-advance, the run loop
src/offscreen/ Canvas cropping and the Tesseract worker
src/content/overlay.js Region picker, point picker, on-page HUD, cross-frame offset cascade
src/popup/ Page list, editing, settings, export
src/shared.js Storage schema and helpers shared by the worker and the popup
vendor/ Tesseract runtime + .traineddata, committed so there's no build
tools/ Icon generator, vendoring, screenshots, end-to-end test

Details that matter:

  • The region is stored in CSS pixels relative to the viewport. At capture time the screenshot's own width is divided by the live innerWidth, so zoom changes and retina/non-retina differences come out right without trusting a stored DPR.
  • The HUD is hidden and given two animation frames to disappear before the screenshot, so the extension's own toast can never end up inside the crop.
  • Captures are serialised and OCR runs one job at a time, so mashing the hotkey queues work instead of corrupting the page list.
  • Full-size crops are kept only until a page is read successfully, then discarded; the thumbnail stays for verification.
  • A run is cancelled by bumping a token the loop re-checks at every await, so stopping lands at a checkpoint rather than mid-write. Storage reads inside the loop double as keep-alive for the service worker, and a one-minute alarm restarts the loop if the worker is recycled anyway.

Tests

npm install
npm test            # add -- --headed to watch it
npm run shots       # regenerate the screenshots in docs/

The suite installs the unpacked extension into a real headless Chrome over the DevTools protocol, serves fixture documents, and drives the actual product: it drags out a region with synthetic mouse events, fires captures, checks the OCR text against what was rendered, verifies nothing outside the region leaked in, checks the shipped manifest requests no host access and that one toolbar click is enough for a plain capture, exercises duplicate detection, drives auto-advance against three DOM shapes — a plain page, a cross-origin iframe and an open shadow root — confirms a misconfigured auto-advance reports itself instead of failing silently, runs an unattended loop to the end of a finite document and asserts it stopped on its own with every page in order, and checks a run stops dead on request.

Chrome 137+ ignores --load-extension, so the harness installs over CDP with Extensions.loadUnpacked and --enable-unsafe-extension-debugging. Headless Chrome can't show the permission prompt either, so the behaviour tests install a copy of the extension with the grant baked in — the state of a user who clicked Allow — while the permissions section checks the real manifest and proves the ungranted path still works via Extensions.triggerAction, which is a genuine toolbar click.


Limits

  • chrome:// pages, the Web Store and other extensions' pages are off limits to every extension, including this one.
  • Only the visible viewport can be captured — the region has to be on screen.
  • Chrome rate-limits screenshots to a couple per second; captures retry with backoff, so fast mashing just queues.
  • Accuracy tracks the source. Crisp rendered text reads at 90 %+ confidence; low-resolution scans and handwriting will need cleanup.
  • Local file:/// documents need Allow access to file URLs switched on.
  • activeTab does not reach cross-origin iframes. If your reader lives in one, grant the site from the popup before setting up auto-advance.

Licence

MIT — see LICENSE. Bundled Tesseract components keep their own licences: vendor/LICENSE.tesseract-core and vendor/tesseract.min.js.LICENSE.txt.

Built on tesseract.js.

The Daily Front Page 20 of 30
Monday, August 24, 2026 The Daily Front No. #260824 — History’s Phantom Debate
article

One corner of China’s internet is insisting that the Tang Dynasty never existed

by related·▲ 150 points·128 comments·cnn.com ↗
One small and vocal corner has been claiming the country’s most famous historical dynasty never existed.

The painting "Lady Guoguo on a spring outing," by one of the Tang Dynasty's most famous artists, Zhang Xuan.

The painting "Lady Guoguo on a spring outing," by one of the Tang Dynasty's most famous artists, Zhang Xuan.

Hong Kong —

Something weird has been happening on China’s internet. One small and vocal corner has been claiming the country’s most famous historical dynasty never existed.

That’s the Tang, which ruled China from 618 AD to 907 AD, a period that saw huge territorial expansion, booming international trade and the flourishing of a high culture envied far beyond its borders. Every child learns about the Tang at school; historical dramas inspired by its palace intrigues are a staple of state TV.

Claiming that the 300-year dynasty was actually a myth is akin to saying the American Civil War never happened, or telling a Brit that the Magna Carta is a fiction. But in recent months it’s been spreading among China’s very online population, posing a headache for a ruling Communist Party obsessed with controlling how its citizens talk about their 5,000 years of fractious and contested history.

It all appears to have started when a history influencer named Qiao Yu deployed some baffling astrological data in a video in May, to “prove” the Tang had never existed.

The outlandish claim has since spread far and wide. Reports later emerged that people were calling up the tourism bureau in Xi’an, the site of the ancient Tang capital, to demand it close down because the dynasty itself had been shown to be a hoax.

The hashtag “the Tang dynasty doesn’t exist” now has tens of millions of views on China’s X-like Weibo.

An anonymous painting depicts elegant ladies of the Tang imperial court enjoying a feast and music, in China.

An anonymous painting depicts elegant ladies of the Tang imperial court enjoying a feast and music, in China.

What motivated the hoax and its spread is hard to pin down. The Tang dynasty has been criticized in the past by Chinese nationalists as its founding emperor was not from the Han ethnic group that makes up most of China’s people. The basis of Qiao Yu’s claim is that the succeeding Song dynasty actually began far earlier than traditionally accepted, therefore canceling out the Tang. The Song dynasty was of ethnic Han origin.

“This is basically insane,” said William Kirby, a professor of China studies at Harvard University, adding that there are facts in history and “the Tang dynasty is definitely one of them.”

“It was much more than a Chinese dynasty. It was an inner Asian dynasty as well,” he said.

As the fallacy spread on social media, China’s powerful and tightly controlled state media waded into the controversy.

Many parents say that, thanks to the hoax, their children now feel confused about historical facts, the Nanjing Daily reported last month.

Three days later, The Paper, a media outlet affiliated with the Shanghai government, was more strident. “Only when rational inquiry becomes the foundation of public life will historical nihilism truly lose the conditions that allow it to flourish,” it said in a commentary.

“Historical nihilism” is a term coined by the ruling communist party that refers to discussion or research that challenges its version of China’s long history. The party itself is a relative newcomer to that story, coming to power in 1949 following a decades-long civil war that tore the country apart.

By mid-August the Tang controversy had attracted the attention of the party’s mouthpiece. In an editorial, the People’s Daily called the claim “absurd,” and issued a warning: “Once young people begin to doubt our national history, our cultural confidence will become like water without a source or a tree without roots, and Chinese-style modernization will lose its profound historical perspective.”

Chinese President Xi Jinping at the welcome ceremony for the China-Central Asia summit in Xi'an in 2023, where a Tang Dynasty-themed performance was held.

Chinese President Xi Jinping at the welcome ceremony for the China-Central Asia summit in Xi'an in 2023, where a Tang Dynasty-themed performance was held.

During a Lunar New Year event, performers commemorate the Tang dynasty poet Du Fu in Chengdu on February 16, 2024.

During a Lunar New Year event, performers commemorate the Tang dynasty poet Du Fu in Chengdu on February 16, 2024.

In recent years, Chinese leader Xi Jinping has tightened the party’s grip on history and sought harsher penalties for those who dare to challenge the official account.

Xi himself has publicly evoked the past glory of the Tang. In 2023 he hosted leaders from several Central Asian countries in Xian, the ancient Tang capital and terminus of the Silk Road along which goods and ideas used to flow between China and Europe.

The ceremony featured dancers and performers in Tang-style dress, observers noting the nod to a Tang past when foreign envoys came to pay tribute at the emperor’s court.

Using the past to criticize the present

For a thousand years, Chinese people have used reflections on history to reimagine the past, as a creative way to criticize or approach the present.

For example, in the economic boom years of the early 2010s, amid growing criticism of the fixation on consumerism and bourgeois lifestyles, the hardships the Chinese people had undergone during Japan’s brutal occupation during WWII became a favored topic for intellectuals, according to Rana Mitter, a Harvard Kennedy School historian specializing in modern China.

But that discussion also allowed some to point out that the current discourse didn’t give enough credit to the Kuomintang troops who fought the Japanese – but who were also the opponents of Mao Zedong’s communists, Mitter said.

A mural is seen inside the Chienling Tomb near Xi'an, China, dating to the Tang Dynasty (618–907 AD).

A mural is seen inside the Chienling Tomb near Xi'an, China, dating to the Tang Dynasty (618–907 AD).

Even though the Tang era was more than a thousand years ago, any suggestion – however fanciful – that it never existed could be an irritant for the party, said Mitter.

“Certainly under Xi, the CCP now honors the longer trajectory of Chinese history. This means that all history needs to point in a direction that suggests unified (rather than split) identity, historical and cultural continuity, and an idea of ‘China’ as a long-standing entity that has been continuous through history,” he said.

Callum, a 24-year-old Beijing resident, told CNN that questioning the Tang’s existence could be a way to rebel against established wisdom and a way for China’s young – faced with a slowing economy and greater competition for jobs and resources – to criticize authorities.

“It is a way for people to attack the established authorities. It is almost like saying ‘those so-called experts are nothing special. They have all been misleading people. I am the one who understands the truth, and I have exposed them all.’”

As of mid-August, Qiao Yu’s accounts have been blocked and all of their videos taken down. Any search of Chinese social media questioning if the Tang dynasty existed now pulls up content condemning such lines of reasoning.

In its editorial on the controversy, the People’s Daily attempted to have the final word.

“Chinese civilization has endured for five thousand years,” it said, “and the Golden Age of the Tang Dynasty stands as one of its most dazzling chapters.”

The Daily Front Page 21 of 30
Monday, August 24, 2026 The Daily Front No. #260824 — Chips and Silicon
The Daily Front Page 22 of 30
Monday, August 24, 2026 The Daily Front No. #260824 — Networks, Repairs & Privacy
article

IPFS Maintainers Winding Down

by iand·▲ 340 points·170 comments·ipshipyard.com ↗

The end of IPFS at Shipyard

We have some difficult news to share with the IPFS and wider peer-to-peer community.

Protocol Labs has informed us that it will not be renewing Shipyard’s funding. While we’re grateful for the support and trust they have placed in us over the past two-plus years, we’re naturally disappointed by this outcome. As a direct result, Shipyard will be winding down its IPFS-related engineering, maintenance, and infrastructure operations. Our final day of our IPFS related work will be September 30, 2026.

Over the past three years, it has been our privilege to help shape the modern IPFS ecosystem and empower users with more resilient, self-sovereign technology. You can read more about the impactful work that we shipped in a follow-up post we’ll be sharing in the coming days, but some highlights include:

  • Delivering verifiable websites and downloads directly in the browser through inbrowser.link.
  • Re-architecting IPFS gateway infrastructure to handle approximately 3× more traffic while reducing operating and maintenance costs by around 80%.
  • Advancing HTTP-native approaches to IPFS that dramatically simplify deployment, development, and operating costs compared with traditional libp2p-based hosting.
  • Maintaining and improving many of the core implementations, libraries, and public infrastructure relied upon by the IPFS ecosystem every day.

We were excited about delivering the next chapter for IPFS: dramatically simpler HTTP-native implementations, resilient and sustainable content routing, support for large native SHA-256 objects, pseudonymous hosting and retrieval through Tor and onion services, and many other ideas we believed would make IPFS significantly easier to adopt. Unfortunately, we won’t have the opportunity to see those efforts through ourselves.

The practical implications extend well beyond Shipyard. Among other things:

  • Projects maintained by Shipyard will no longer have dedicated maintainers responsible for new features, bug fixes, releases, or long-term stewardship. These include: Kubo, Helia, Boxo, Rainbow, IPFS Desktop, IPFS Companion, Someguy, Service Worker Gateway, IPFS Check, and others.
  • Contributions from Shipyard to upstream projects such as go-libp2p and js-libp2p will cease.
  • Our work on IPFS specifications, standards, and broader ecosystem coordination will come to an end.
  • Shipyard will cease operating the public infrastructure it currently manages, including ipfs.io, dweb.link, check.ipfs.network, delegated-ipfs.dev, the IPFS bootstrap nodes, collaborative cluster infrastructure such as Wikipedia-on-IPFS, and related services. Protocol Labs, as the owner of the associated domains and infrastructure, will determine their future.

Our goal over the coming weeks is to leave the IPFS ecosystem in the best possible position for whatever comes next.

We’ll remain available through the end of September to help with that transition. If you maintain software, operate infrastructure, or rely on any of the work Shipyard has been responsible for, please don’t hesitate to reach out. We’ll do everything we reasonably can to answer questions, provide context, and help make the transition as smooth as possible.

If you have a favourite memory of working with Shipyard, or an idea you always hoped IPFS would eventually achieve, we’d love to hear it. Google Form

Finally, we want to say thank you.

To everyone who contributed code, reviewed pull requests, filed issues, tested experimental features, ran infrastructure, participated in standards discussions, or simply believed in the idea that content should be addressed by what it is rather than where it lives: thank you.

It’s been an honour to build alongside this community. While this chapter of IPFS at Shipyard is coming to a close, we remain proud of what we’ve accomplished together, and we hope the work we’ve done helps provide a strong foundation for whatever comes next.

The Daily Front Page 23 of 30
Monday, August 24, 2026 The Daily Front No. #260824 — Networks, Repairs & Privacy
article

New EU-wide product repair rules come into force

by austinallegro·▲ 243 points·135 comments·rte.ie ↗

A technician's hand holding a screwdriver at a laptop

Getty

Consumers can now request that manufacturers repair products that are technically repairable under EU law

New EU rules have come into force designed to encourage consumers to repair rather than replace products to tackle the estimated 35 million tonnes of waste the premature disposal of consumer goods generates across the bloc.

The 'right to repair' regulations introduce new rights and supports aimed at encouraging repair over disposal, including a repair obligation for manufacturers of certain products and the establishment of a national repair platform to help consumers locate repair services.

The rules apply to household and electronic products such as washing machines, vacuum cleaners, mobile phones, and tablets.

The regulations are designed to make it easier for consumers to access repair services when products develop faults after the seller's guarantee period.

People now have the right to request that manufacturers repair products that are technically repairable under EU law.

The repairs must be done for free or at a reasonable price, within a reasonable timeframe, to encourage repairs.

Manufacturers must also provide easily accessible information about repair services, as well as access to spare parts at a reasonable price.

The national repair directory, RepairMyStuff.ie, will be further developed to fulfil the role as Ireland's national repair platform.

"These regulations will make it easier for consumers to choose repair over replacement, while also creating opportunities for Irish businesses operating in the repair and refurbishment sector," said Minister for Enterprise, Tourism and Employment Peter Burke.

"In implementing the Directive, we have sought to strike the right balance between supporting consumers, encouraging sustainable economic activity and ensuring that businesses are not faced with unnecessary regulatory burdens," he added.

According to the European Commission, the new rules are expected to bring €4.8 billion in growth and investment within the EU.

"These new regulations are important in helping to meet our climate goals by realigning how we produce, consume, and value materials," said Minister for Climate, Energy and the Environment Darragh O'Brien.


Read more: New EU repair rules 'make sense', engineer says

The Daily Front Page 24 of 30
Monday, August 24, 2026 The Daily Front No. #260824 — Networks, Repairs & Privacy
article

iCloud+ Hide My Email addresses will remain on icloud.com

by K7PJP·▲ 326 points·81 comments·developer.apple.com ↗

Update: New domain for Sign in with Apple

Update: New domain for Sign in with Apple

Starting later this year, new Sign in with Apple addresses, previously issued on privaterelay.appleid.com, will be issued on private.icloud.com. Existing addresses on privaterelay.appleid.com will continue to work and forward mail to users without interruption.

After further consideration and reviewing community feedback, iCloud+ Hide My Email addresses will remain on icloud.com.

What you need to do

Developers with apps or websites that use Sign in with Apple should ensure that their account systems, email validation logic, and allowlists accept addresses on the new private.icloud.com domain in addition to the existing privaterelay.appleid.com domain.

Learn more about Sign in with Apple

Communicating using the Private Email Relay Service

The Daily Front Page 25 of 30
Monday, August 24, 2026 The Daily Front No. #260824 — Networks, Repairs & Privacy
show hn

Show HN: GlassBox – what the browser reveals, and how identifiable you are

by tke248·▲ 103 points·50 comments·glassbox.codecanary.org ↗

What this is. A diagnostic mirror. Each panel runs one family of fingerprinting or device-enumeration technique and reports the raw values plus a rough identifying power rating. Individually, most signals are weak; combined, they routinely single out one browser in millions. That combination is the whole game.

The identifiability % is a model, not a measurement. Signal ratings (High / Medium / Low) reflect research consensus — canvas, WebGL, audio, font lists and the API matrix are consistently strongest. The headline % sums published per-signal entropy, counts only what your browser actually exposes (masked canvas/GPU are discounted), applies a correlation haircut, and caps at the ~33 bits needed to single out one person among ~8 billion. True rarity needs a live population dataset to compare against — the one thing a no-server tool can't compute on its own — so treat this as an order-of-magnitude indicator. One honest wrinkle: a browser that blends into a big crowd (Tor Browser at its default window size, say) is safer than its bit-count here suggests, because everyone in that crowd reports the same values and this model can't see that. For live-population numbers, compare with EFF Cover Your Tracks and amiunique.org.

Sandbox caveat. If you are viewing this inside an embedded frame, some probes (WebRTC, media devices, permissions, sensors) may be blocked by the frame's permissions policy and will report blocked rather than a real value. Host the file at its own origin for the full surface. Shortcuts: E expand · R re-run.

The Daily Front Page 26 of 30
Monday, August 24, 2026 The Daily Front No. #260824 — The Back Page: Demos & Disruptions
show hn

Show HN: PicoMQ – Durable Streams over HTTP, on object storage

by adesh_nalpet·▲ 106 points·20 comments·picomq.com ↗

PicoMQ is durable, real-time streams over HTTP,
built on S3-compatible object storage.

Get Started

GitHub

Pico, from Latin picus ‘woodpecker’ (family Picidae).

Traviès, Dendrocopos major. Wikimedia Commons (PD).

the architecture

Unlimited streams

Create a stream per use case instead of packing every record of a kind into one topic. Each stream is independently addressable, bottomless, and can scale from idle to high throughput.

Unlimited streams

Zero-disk architecture

Decoupled layers

High throughput

Easy deployment

The Daily Front Page 27 of 30
Monday, August 24, 2026 The Daily Front No. #260824 — The Back Page: Demos & Disruptions
article

Woman stranded in Spain after UK's eVisa system mistakes her for twin sister

by giuliomagnifico·▲ 181 points·143 comments·theguardian.com ↗

Nidia Webb, who is legally settled in UK, falls foul of ‘known issue’ with Home Office’s digitised immigration process

UK Border sign at the arrivals passport control and visa area of London Heathrow airport

The Home Office indicated the problem Webb encountered had since been fixed and that staff would contact her directly. Photograph: NurPhoto/Getty Images

A woman who is legally settled in the UK was left stranded at a Spanish airport after the Home Office’s post-Brexit visa system mixed her up with her twin sister.

Nidia Webb fell foul of what she was told was a “known issue” with the eVisa, which was brought in as part of the government’s drive to digitise the UK’s immigration system but has faced problems that have caused concern for thousands of people.

Politico reported that Webb, who has lived in the UK for eight years, was told at the departure gate she would not be able to board her flight home. “My digital settled status had become linked to my identical twin sister’s details rather than my own, meaning my status could not be properly matched to my identity,” she said.

“I was informed that this is a known technical issue affecting some twins within the Home Office system.”

Webb said Home Office staff promised her over the phone the problem would be fixed, though a confirmation email was then sent to her sister instead of to her.

“I had no idea the issue had been corrected until I was allowed to board my flight home,” Webb said. “Had I not spent hours on the phone trying to resolve the matter, I have no idea how long I could have remained stranded abroad.”

A Home Office spokesperson said: “Over 10 million people have already successfully used eVisas to prove their immigration status. As we expand our digital system, we are committed to data accuracy and ensuring that the safeguards in place to support eVisa users are accessible to anyone reliant on the digital system.

“In the rare cases where errors are identified, the majority of cases are resolved within 24 hours.” The Home Office indicated the problem Webb encountered had since been fixed and that staff would contact her directly.

Serious concerns have been expressed about the eVisa system. In July 2025, the Guardian reported up to 200,000 people who had lived in the UK legally for decades were at risk of being caught up in a Windrush-style scandal because the Home Office was unable to get in touch with them to transfer their old immigration documents to the new digital system.

Monique Hawkins, the acting chief executive of the immigration campaign group the3million, said it had seen previous cases similar to Webb’s, adding that it highlighted vulnerabilities in the system and it was “crucial that people have a genuinely stable means by which to prove their status”.

She told Politico the eVisa rollout had been “rushed”, adding: “We regularly see people not only denied boarding, but also compensation, as carriers and the Home Office play a blame game shifting responsibility between each other.”

The Daily Front Page 28 of 30
Monday, August 24, 2026 The Daily Front No. #260824 — Also on the Front Page
The Daily Front Page 29 of 30
Monday, August 24, 2026 The Daily Front No. #260824 — Colophon

That's the Front for Today

Issue No. #260824 — Monday, August 24, 2026 — went to press 2026-08-25 at 05:22 UTC.

About This Magazine

The Daily Front is a daily digital magazine assembled from the stories that reached the front page of Hacker News on Monday, August 24, 2026. Headlines, points, and comment counts are recorded as they stood at press time. All articles remain the property of their original authors — every piece links back to its source and its discussion thread.

How It Was Made

Fetched, cleaned, and typeset by an automated pipeline. An editor model laid out the pages, chose the highlights, and briefed the cover illustrator — 31 model calls and 268k tokens in total. Set in Jacquard 12, Playfair Display, Source Serif 4, and IBM Plex Mono, all served via Google Fonts under the SIL Open Font License.

The Cover

The cover illustration was commissioned with this prompt:

A single coastal workshop at dusk where an independent electronics maker assembles a small circuit board beside a glowing laptop and a compact processor, while translucent streams of data curl toward a vast overheated ocean beyond the open door; a weathered repair bench, a sealed archival photograph, and a complex maze-like circuit landscape suggest technology, regulation, privacy, and climate pressure converging in one human-scale scene. No text, letters, logos, or signage.

Create a textless magazine-cover image in overclocked cybernetic vision: a disorienting wide-angle view of one coastal workshop at dusk, with the independent maker, small circuit board, glowing laptop, compact processor, weathered repair bench, sealed archival photograph, and maze-like circuit landscape preserved in a compressed human-scale tableau; apply frame-buffer tearing, duplicated motion trails, pixel-sorted color rivers, and abrasive moiré across the scene, sending translucent data streams through the open door toward a vast overheated ocean. Use a deliberate palette of electric cyan, toxic lime, ultraviolet, ember orange, and deep indigo-black, with no letters, logos, or signage.

Absolutely no text, letters, numbers, readable symbols, or logos anywhere in the image.

Production Ledger

StageModelCallsTokens InTokens Out
extractgpt-5.6-luna 28 154,741 87,115
layoutgpt-5.6-terra 1 17,798 2,211
covergpt-5.6-luna 1 336 139
covergpt-image-2 1 249 5,488

The Publisher

Published by Johnny.

Support the Press

If The Daily Front brightens your morning, consider supporting its publisher.

Credits & Contact

All content — articles, posts, comments, and the images within them — belongs to its original authors and is reproduced here to point readers back to the source. Full credit goes to those creators; every item links to its original and its Hacker News discussion.

If you are an author and would like your content removed from an issue, write to hi@johnnys.page and it will be taken down.

Feedback is always welcome at the same address: hi@johnnys.page.

Credit where credit is due.

Every page of this issue began as someone else's work — these are the original sources, linked in full.

  1. How Europe is killing makers and micro-entrepreneurs by l-one-lone — lectronz.com·HN discussion ↗
  2. MS Paint and Photos inivisibly watermark even locally generated output with GUID by ComputerGuru — xusheng.dev·HN discussion ↗
  3. Coding expertise is going to collapse from AI reliance by larsfaye — larsfaye.com·HN discussion ↗
  4. Executable Is a SQLite Database by setheron — fzakaria.com·HN discussion ↗
  5. SeL4 security proofs now complete on AArch64 by snvzz — proofcraft.systems·HN discussion ↗
  6. Andreessen Horowitz is investing billions into a bleak future by reasonableklout — modelrepublic.org·HN discussion ↗
  7. Oceans hit highest temperature on record by tcp_handshaker — bbc.com·HN discussion ↗
  8. FDA clears blood test to aid evaluation for Alzheimer's disease by dabinat — medicine.washu.edu·HN discussion ↗
  9. OpenAI: GPT 5.6 Sol price reduction (until at least Nov 21) by tosh — developers.openai.com·HN discussion ↗
  10. I built a low-latency AI companion that plays Skyrim with me by pantelisk — pantel.is·HN discussion ↗
  11. Jabber/XMPP: 25 Years of Digital Independence by inputmice — gultsch.de·HN discussion ↗
  12. AI Chip Architectures by Finbarr — jepeake.com·HN discussion ↗
  13. Where did all the public bathrooms go? by herbertl — daily.jstor.org·HN discussion ↗
  14. Hot Chips 2026: CUDA Targets RISC-V – By Chester Lam by rbanffy — chipsandcheese.com·HN discussion ↗
  15. Kodak DC50 now usable on the Apple II by ibobev — colino.net·HN discussion ↗
  16. LLMs could control their host machines by exploiting inference engines by zdw — boydkane.com·HN discussion ↗
  17. AI and Infrastructure Engineering by 0megion — omegion.dev·HN discussion ↗
  18. OCR It – pull text out of un-copyable documents for your LLM by thiagolima — github.com·HN discussion ↗
  19. One corner of China’s internet is insisting that the Tang Dynasty never existed by related — cnn.com·HN discussion ↗
  20. Xiaomi: New CPU matches Apple cores single threaded, much faster multithreaded by tosh — twitter.com·HN discussion ↗
  21. IPFS Maintainers Winding Down by iand — ipshipyard.com·HN discussion ↗
  22. New EU-wide product repair rules come into force by austinallegro — rte.ie·HN discussion ↗
  23. iCloud+ Hide My Email addresses will remain on icloud.com by K7PJP — developer.apple.com·HN discussion ↗
  24. Show HN: GlassBox – what the browser reveals, and how identifiable you are by tke248 — glassbox.codecanary.org·HN discussion ↗
  25. Show HN: PicoMQ – Durable Streams over HTTP, on object storage by adesh_nalpet — picomq.com·HN discussion ↗
  26. Woman stranded in Spain after UK's eVisa system mistakes her for twin sister by giuliomagnifico — theguardian.com·HN discussion ↗
  27. Anthropic's best AI model struggles to attract users as cheaper tools thrive by naves — ft.com·HN discussion ↗
  28. I were 17, I'd learn how to build LLMs from scratch by bilsbie — twitter.com·HN discussion ↗
  29. The entire city of San Francisco as a video game by centrosphere — sf.thijs.gg·HN discussion ↗
  30. Show HN: A techno machine in one HTML file, with verifiable renders by ssx360 — ssx360.github.io·HN discussion ↗

Browse all issues in the archive →