Cover illustration

TheDaily Front

Issue No. #260908 Tuesday, September 8 2026 #260908 — TUESDAY, SEPTEMBER 8, 2026
Turbulence in the equations, and rather more in the comments.
Tuesday, September 8, 2026 The Daily Front No. #260908 — Contents
30stories
12,027points
6,592comments
281kllm tokens
Assembled with 32 model calls — 186,701 tokens read, 94,362 written.

Highlights

On the Navier–Stokes Millennium Prize Problem

An OpenAI system’s claimed Navier–Stokes breakthrough arrives with a formal Lean proof—and an exceptionally spirited public examination.

Navier-Stokes – Tristan Buckmaster [pdf]

Tristan Buckmaster sets out related blowup results and the context surrounding the day’s loudest mathematical dispute.

Mistral raises €3B

Mistral raises a record €3 billion round to pursue Europe’s sovereign, open-weight AI ambitions.

I've factored the RSA keys of a Certificate Authority from the 90s

A 1990s certificate authority meets modern hardware in a hands-on lesson about the practical lifetime of old cryptography.

LG TVs caught spying even when offline or on standby

Reports of pervasive smart-TV telemetry renew the case for keeping the television’s network cable unplugged.

From the Editor

The day’s principal spectacle was not merely a question of fluid dynamics, but of priority, proof, and the proper conduct of scientific announcements. Elsewhere, the machines grew richer, smaller, more observant, and—depending on one’s television—rather too attentive to domestic life.

  1. On the Navier–Stokes Millennium Prize Problem3
  2. Navier-Stokes – Tristan Buckmaster [pdf]4
  3. There's a new "Google Jail" for independent wikis5
  4. I've factored the RSA keys of a Certificate Authority from the 90s6
  5. Jellyfin 12.07
  6. We built our house for LAN parties (2024)8
  7. TALA Is Open-Source9
  8. Show HN: Copperhead – Cursor for circuit boards10
  9. Benchmarking Qwen3.8 27B quantizations: 4-bit holds up, 1-bit collapses11
  10. Among European Companies That Use a CDN, Nearly 9 in 10 Use Cloudflare12
  11. The Helicopter with Radioactive Blades13
  12. John Margolies' photographs of roadside America14
  13. Getting your hands dirty is good for you15
  14. The two Christian saints who are the Buddha16
  15. Emacs Bedrock 2.017
  16. How well do agents use test/verification techniques?18
  17. Show HN: LLM Attention Visualization19
  18. Kimi K3 (2.8T) at 1 token/s on a MacBook Pro, streamed from four SSDs20
  19. Arm Mali G2-Ultra NX GPU: desktop-class mobile gameplay with AI-native graphics21
  20. Antiquated HTML Snippets and Artefacts22
  21. We have a year to fix security everywhere23
  22. Mistral raises €3B24
  23. AlphaGenome Atlas: a high-resolution map of human DNA25
  24. LibreOffice breaks download records after declaring it has no AI features26
  25. I-have-ADHD: A skill to stop coding agents from burying the answer27
  26. LG TVs caught spying even when offline or on standby28
  27. Paramount Caught Using 'Astroturf' Group to Drum Up Fake Support for Merger29
  28. DaVinci Resolve 21.130
  29. Muse – Meta’s personal AI agent30
  30. The 92-Year-Old Mathematician and the Teenage Apprentice30
The Daily Front Page 2 of 31
Tuesday, September 8, 2026 The Daily Front No. #260908 — The Mathematics Desk
article

On the Navier–Stokes Millennium Prize Problem

by tedsanders·▲ 1,193 points·1,029 comments·openai.com ↗
We’re sharing a solution to the Navier–Stokes existence and smoothness problem, one of the Millennium Prize Problems.

We’re sharing a solution to the Navier–Stokes existence and smoothness problem, one of the Millennium Prize Problems. This proof, produced by an internal OpenAI system, shows that the dynamics of the Navier-Stokes equations for fluid motion can develop a singularity in finite time. We’re sharing both a writeup of the proof and a formalization in Lean.

The Millennium Prize Problems⁠(opens in a new window) represent some of the deepest questions at the frontier of mathematics. The question of whether smooth three-dimensional fluid motion can break down has remained unresolved for roughly 90 years.

A major goal of our work is to empower scientists to advance research and technology that benefits all of humanity. To solve the Navier–Stokes problem, we used an internal model that is significantly more capable than GPT‑6 Astra. We believe it is important to inform the world about the pace of AI progress and what to expect from upcoming models.

The problem

The Navier–Stokes equations use Newton’s second law of motion (“F=ma”) to describe how fluids move. Importantly, they treat a fluid as a continuous medium rather than tracking individual molecules. These equations are used for aircraft design, weather forecasting, and the study of blood flow.

A fundamental open question for these dynamical equations has been whether the continuum approximation of the fluid can break down. Specifically, can the Navier–Stokes equations for a three-dimensional incompressible fluid with constant density develop a “singularity,” even when the motion starts smoothly? Here, a singularity means the dynamics lead to speeds in the fluid growing without bound within a finite amount of time. The development of a singularity would have to happen despite the presence of viscosity, which tends to smooth out motion. Because a real fluid cannot move infinitely fast, this would mark a breakdown in how the equations model the fluid. To continue modeling the system, one would then need to track the behaviour of each particle individually.

The equations date to the nineteenth-century work of Claude-Louis Navier and George Gabriel Stokes. In 1934, Jean Leray proved that solutions exist in a generalized sense, but whether they always remain smooth became a central unanswered question. In 2000, the Clay Mathematics Institute named the Navier–Stokes existence and smoothness problem one of seven Millennium Prize Problems.

The result

Our system produced an analytical proof and a Lean formalization that an initially smooth fluid at rest can develop a singularity in a finite time. The fluid has a smooth force applied to it, and its energy remains finite through the entire dynamics, from rest to the formation of the singularity. This resolves the Navier–Stokes Millennium Prize problem by establishing statement “C” (and also “D”) in the official Millennium Prize formulation⁠(opens in a new window).

The solution is a vortex, a spinning swirl of fluid, that spirals inward and gets increasingly elongated, like spaghetti. This central region shrinks while it speeds up in such a way that its energy still stays finite, as required by the laws of physics. The technical challenge is for the equations to develop the breakdown through the motion of the fluid itself, rather than, for example, us putting in an infinite force by hand. More mathematically, the terms in the Navier–Stokes equations that describe the motion—acceleration, pressure gradients, momentum transfer, viscosity—must both become big yet cancel in a precise way. This detailed balance leaves a smooth external force even as the velocity of the fluid grows without bound.

Diagram of a swirling vortex illustrating inward spiral and axial stretching.

A snapshot of local incompressible motion. Orange marks faster angular rotation; teal marks slower rotation. Circulating speed also depends on radius. The trajectories show inward spiraling and axial stretching.

How we found the proof

Since August 28 we have been training a new internal model that has exhibited unprecedented performance in our benchmarks, including mathematics. This model’s training is ongoing and its performance continues to improve.

On Tuesday, September 1, we heard rumors that two Millennium Prize problems had been resolved. Inspired by these rumors and by the step change in performance of our internal model, we launched an effort to evaluate it on all open Millennium Prize problems and a few other high-impact problems.

We used a system of coordinating agents powered by our internal model. The agents had access to tools such as the ability to read from a cached version of the internet and the ability to run code. Agents were subdivided into groups with the ability to communicate within the group. The groups varied in size, and the group that produced the Navier–Stokes resolution involved on the order of 10,000 concurrent agents. At all times we maintained the same strict safeguards that we apply to all our frontier model evaluations, including monitoring and isolation.

For each problem, we prompted different groups of agents with different variants of the problem statement, covering all variants of the problem. For the Navier–Stokes problem, we suggested versions “A” and “B” (particular forms of the Navier–Stokes problem which would result in a proof) and versions “C” and “D” (which would result in a disproof) to separate groups of agents.

In addition to the full Millennium Prize problems, we asked our multiagent system to try a set of “easier” problems. One of these problems was a similar blowup question for the limit of the Navier–Stokes problem with the viscosity term removed. This is known as the regularity problem for the Euler equations, and our agents surprised us by resolving this question. The specific variant of the question that they resolved was the unforced version, where no external force is applied to the fluid. Nearly 100 agents worked together for approximately 50 hours to produce our Euler regularity disproof.1

Once we saw the Euler solution, we thought that Navier–Stokes was the most promising problem to work on. Thus, we decided to devote our resources to Navier–Stokes. To do so, we shifted agents away from the other Millennium Problems and prompted these agents with the Euler resolution. When a further trained version of our internal model became available over the course of the effort, we updated our agents to that model.

We encouraged different groups of agents to explore a diversity of approaches. After some time, we cross-pollinated the agent groups by using Codex to consolidate the most useful insights from each agent group. These follow-up prompts drew on the agents’ own intermediate results. The group that found the solution to Navier–Stokes was guided in such a way.

The agents arrived at their resolution on Saturday, September 5, about 88 hours after the first agents were launched. Lean formalization and verification took an additional 17 hours via GPT‑6 Astra.

Across all attempted problems, the agents sent 4.9 million messages and used about 300 billion output tokens. In the process of resolving the Navier–Stokes problem, the agents sent 2.7 million messages and used approximately 130 billion output tokens.

Concurrent work

Our effort began on September 1st after hearing a rumor which we later realized was related to Levent Alpöge, an Anthropic employee, and Tristan Buckmaster, a math professor at NYU. After the completion of our full project and Lean verification (on September 6th), believing from the rumor they also had a solution of Navier–Stokes, we reached out to them to offer a concurrent release of our result and to recognize their priority in a joint announcement. At that point we found out that they had a resolution of the forced Euler problem. In these discussions we offered them visibility into all of the prompts we used and later to see the proof. We recognize the priority of their work on forced Euler and congratulate them on their remarkable mathematical achievement.

We (the researchers and the agents) did not see any of their work through any means until they released it publicly — in particular, no specific user data was accessed in order to solve this problem. While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models⁠. However, our proofs differ significantly and even the precise results proved are different in the Euler case (forced vs unforced).

Progress and responsibility

Our goal in releasing this result is to report on the substantial progress of our AI models. We do not intend to claim the Millennium Prize for this result.

This milestone represents substantial work by mathematicians and AI researchers. However, this is not a culmination, but rather a snapshot in time, of progress on AI development.

We believe we are now in the next period of AI progress⁠, and today’s results provide further evidence of this. We are focusing on understanding this model, and using what we learn to help us guide and pace how we pursue further advances in capability. One of our key goals⁠ is to build AI systems which are steerable, accountable, and connected to people, which may require more deliberate choices about the pace of progress, as we continue our mission to ensure AGI benefits all of humanity.

Footnotes

  1. Read the Euler proof paper⁠(opens in a new window) · Link to Lean formalized proof⁠(opens in a new window)
The Daily Front Page 3 of 31
Tuesday, September 8, 2026 The Daily Front No. #260908 — Concurrent Work
article

Navier-Stokes – Tristan Buckmaster [pdf]

by procedurecall·▲ 1,422 points·606 comments·cims.nyu.edu ↗
We believe we also have blowup for hypo-dissipative Navier-Stokes.

Today, Levent Alp¨oge and I have made public three results: finite-time blowup with smooth forcing for incompressible porous media, for Boussinesq, and for 3d incompressible Euler.

We believe we also have blowup for hypo-dissipative Navier-Stokes. We are not releasing that paper today: unlike the above, the Lean verification has not yet finished. We do not yet have anything resembling a presentable writeup. I mention it because it is suggestive of a path to unforced Euler.

The program this fits into was not started by us nor was it proposed by a Large Language Model. The credit for the basic idea of this program goes to Diego C´ordoba and Luis Mart´ınez-Zoroa, who for several years have been exploring the construction of forced blow ups. We took their work as a starting point, using Large Language Models to push their program to completion.

Concretely, what Levent and I did was to take the C´ordoba and Mart´ınez-Zoroa program, which achieved blowup results with rough forcing, and, with a great deal of help from LLMs, push it to smooth forcing and to the incompressible Euler equations. The ideas making this line of attack possible are due to C´ordoba and Mart´ınez-Zoroa. Let me make plain what I have said to colleagues in private: in view of this body of work, I believe Luis Mart´ınez-Zoroa deserves a Fields Medal.

My work with Levent has been a purely personal collaboration, free of any institutional agreements or official involvement by either of our employers. We used several LLMs throughout: Anthropic’s Claude, OpenAI’s Codex, especially with GPT-5.6 Sol and, more recently, Astra. The latter was only used for writeups and auditing our arguments.

For most of the past year progress was slow. We worked through the literature and upgraded various preliminary results, up to obtaining finite time blow up for the Incompressible Porous Media equation (with smooth forcing). This was until about a month ago, when we had real progress: on August 15th, we obtained the blow up results, with smooth forcing, for both Boussinesq and Euler. I can say the first LLM generated proof Levent sent me was the most horrendous I have ever read; we verified it on Lean on August 22nd. Since this point, we have been working around the clock to understand this proof and turn it into something readable.

There is another part of this story, and one that, honestly, I very much wish I did not have to be concerned with. I am not happy about the presentation quality in these papers. Ideally, we would have preferred to spend weeks turning the LLM generated proofs into something readable from the very first page. This level of care is what these problems and the community devoted to these problems deserves. The various Boussinesq and Euler write-ups in particular are much closer to what models produce under human direction than to a paper written by a person. The Euler writeup, in particular, can only be described as AI slop. I am sorry for this. The reasons are below, and they involve our being pressured by outside factors. I say this not as an excuse but as an explanation.

I had planned to say on announcing our work that the results are not the important thing. Rather the important thing is instead the significance that a mathematician and an LLM model can now do all this work in a month. The significance of this with respect to the way we train students, assign credit, referee, and decide what is worth one human life’s attention cannot be understated. This is a a Deep Blue-Kasparov moment. The community needs to have serious and unhurried discussion about where to go from here. Instead of these incredibly important developments, I find myself writing about something else. I want to set out what happened as plainly as I can.

On Thursday, September 3rd, with a rumor circulating that Anthropic had resolved a major open problem, and with Levent having received tips that information about our progress had been passed to OpenAI, I wrote to a prominent mathematician at OpenAI. I am quoting my email in full because I would rather the full text be read rather than my summary of it:

“Hi [...],

We have not met in person, but I was one of the speakers at the [...].

I am writing because a rumor seems to be spreading quickly that Anthropic has resolved a major open problem. I cannot be certain I am the person it attaches to, but a colleague at Courant emailed me about it yesterday – who heard it from an analyst in the UK, who had it from somewhere further upstream – so it seems safe to assume I am. I gather versions linking Levent Alp¨oge to a solution are going around in tech as well. There however appears to be a lot of confusion with regards to the exact problem solved.

I should also emphasize that this is not an institutional effort. It is a strictly personal collaboration between the two of us, and there is no formal agreement behind it. I pay for the tools my group uses out of my own research funds, including footing a large bill to OpenAI. I have had an industry collaboration before, with DeepMind, which had a formal institutional arrangement.

We do have work in this area that we are confident in, and we will post it shortly, the paper and the formalization together. We intentionally decided against rushing out a Lean certificate alongside an unpolished preprint. I feel strongly that the first thing anyone reads should be a mathematical argument presented in the normal manner, rather than just a formal certificate.

I am writing to you directly rather than saying anything publicly so that you have the facts to address this on your end.

Best wishes, Tristan”

He replied the same day:

“If you are willing to give any details it would be useful to avoid competing here and in general we are always thrilled when mathematician make progress with our models. Additionally if there is anything in terms of compute from OpenAI’s end we would be happy to provide it.”

I asked to speak the following week. On Friday, September 4th, I was asked whether I could meet that day; I again said the following week. At 12:45 on Sunday, September 6th, I was asked whether I could meet “at any point today.” Sebastien Bubeck joined. The three of us spoke twice that afternoon. Levent was not on the calls.

I was told that an internal OpenAI model had produced a proof of finite time blowup for the forced Navier-Stokes equations. When Levent asked by text for the precise statement, the answer was: “Existence of forced blowup in R3 and T 3,” with “the forcing function is smooth option c and d in Fefferman.” I was told the proof is about 100 pages. I have not seen it.

I should say here why I interpreted their statement the way I did, the interpretation I will discuss below. The route to the Clay problem through a smooth force, options c and d in Fefferman’s statement of the problem, is the route Luis and Diego opened and the one Levent and I had quietly chosen to attack. Almost nobody else I know of was working on it. It is not the direction one arrives at in a few days by giving a model the problem statement. When I heard “forced,” it was a bright red flag.

I was shown a prompt and told the internal research model had simply been given the problem statement. Levent had been told by Sebastien “very little human input” had been used. This turned out not to be true. Over the course of the call, as members of their team sent Sebastien corrections and details over their internal chat, it emerged that an entire team had been working on the problem, that this was one of a number of things that was tried, that work had started on the unforced problem, that the team first set the model on easier problems, including Euler, that even the prompt that had been shown to me had been written by prompting Codex, and that an insane amount of compute had been used.

I asked when the first prompt had been sent by them. This question was not answered directly by OpenAI for some time. Eventually it was agreed that it had been sent in the past few days, after information about our work had reached OpenAI.

I asked whether the model had been trained on, or had access to, our sessions in Codex, into which we had been putting all our drafts for the whole of this project. I was told the model did not look up user data. I asked again, about training, and I did not get an answer.

Two proposals were offered to me. The first was that we post our Euler result, and that OpenAI post its Navier-Stokes result the next day. The second was that, after posting Euler, I alone write a paper presenting the Navier-Stokes result, acknowledging that an internal OpenAI model had resolved it. Sebastien twice asserted that he wanted Levent removed from authorship, and said it would all be simple if only it were not the case that, and it was so annoying that, Levent works at Anthropic. It was also said that if OpenAI posted after us, they would say that we deserved the Clay Prize, and that we were the “closest humans to the problem”. I declined both offers.

I said that if OpenAI released its result in the way proposed I would go public with what happened. The reply was, “Why would you ruin your career?” I replied that I am an academic, and asked why he thought going public would ruin my career. The reply was, “If you don’t want me to be nice, then I don’t have to be nice.”

Some time later Levent received a text proposing that he and Sebastien speak one on one, saying, “I don’t know if Tristan is being fully rational right now.” Levent declined and said conversations should be with me. Sebastien sent a follow-up email that night requesting to speak with me on Monday, September 7th. I did not respond.

I would like to be clear about what I am not claiming. I have not seen OpenAI’s proof. I do not know what their model did, or how. I do not know whether our data was used. I am not accusing anyone of anything. I am stating what I was told, when, and what was proposed to me. I am stating it because the alternative is to let a sequence of announcements say something I know to be false.

If indeed an OpenAI model did close the gap to Navier-Stokes, that is a remarkable thing and it should be said loudly, by them, with the history intact. I would much rather be talking about mathematics, Luis and Diego’s ideas, and what this all means for the rest of us.

Lastly, I would like to thank the entire mathematics community that have been so supportive of me over the last 24 hours.

The Daily Front Page 4 of 31
Tuesday, September 8, 2026 The Daily Front No. #260908 — Search Exile
article

There's a new "Google Jail" for independent wikis

by pizzaiolo·▲ 529 points·223 comments·weirdgloop.org ↗
It’s still my favorite thing in the world to help wiki editors grab control of their projects.

Big week! We just finished moving the Overwatch and Fortnite wikis off of Fandom. It’s still my favorite thing in the world to help wiki editors grab control of their projects and make them awesome (and not covered with a dozen auto-playing video ads!)

Unlike most of the other wikis we host (runescape.wiki, minecraft.wiki, gta.wiki…), both Overwatch and Fortnite are launching at subdomains of weirdgloop.org instead of on their own root domain, like (say) overwatch.wiki. There is an unusually good reason for this, and the rest of this post will be going into excruciating detail about:

  • a frustrating change that Google made in early 2024 to how it treats brand-new-domains
  • the wide-ranging effects of these changes on the entire independent wiki ecosystem
  • how I’m trying to work around this new “Google Jail” for the wikis that we host

If you have a new domain, only the main page shows up on Google search results

These March 2024 changes have a specific “failure mode” on Google Search which has made it even more difficult for a certain class of new wikis (and most likely, new websites in general) to show up in search results. For the last two years, with very few exceptions, wikis on brand new domains will not show up on Google Search for anything other than the main page.

Since ~85% of video game wiki traffic comes from Google, and wikis (obviously) have many popular pages other than the main page, this is a pretty catastrophic outcome.

Google jail graph

Hollow Knight wiki: into Google Jail in March 2024, back out 9 months later

Here’s what we think we know:

  • This has happened to about 90% of wikis I’m aware of that have launched on brand new domains since the March 2024 Google core update. This includes gta.wiki and hytalewiki.org that we host, hollowknight.wiki, the official Path of Exile 2 wiki, and a couple dozen smaller wikis
  • This happens regardless of whether the content is brand-new, or derived from something else (like Fandom) that Google is already indexing. It’s not the conventional “duplicate content” issue that wikis have had to deal with for the last decade.
  • It doesn’t seem to have much to do with the actual “ranking” of the domain - there’s quite a few examples where the main page is actually beating Fandom’s main page on Google, and yet that’s the only page that shows up at all for the entire domain.
  • This “Google Jail” lasts for an unclear period of time, sometimes up to a year, and is also sometimes”defeated” by big game updates that cause significant new traffic/content to come to the wiki. Undertale and Vampire Survivors are examples of wikis that defeated the Google Jail by waiting it out and having a big game update.
  • It can intermittently start and stop, sometimes getting out of Google Jail for a few months and then going back in again. It seems to always eventually stop.
  • Most of the time, but not always, articles besides the main page are still getting indexed and crawled, and show up with a site: search
  • The relevant factor seems to be not the actual “age” of the registration of the domain, but roughly when Google first indexed it
  • I can’t find any evidence of this happening before March 2024, in any context even outside of wikis

Hytale Wiki on Google

#1 for the most popular search, and it's the only page on the entire domain that shows up

My best guess (and to be clear, this is a guess): Google decided they could no longer effectively identify and swat away SEO slop, and figured that just massively nerfing brand new domains was their next best option.

Subdomains of existing domains are totally fine

Since this Google issue started, we’ve also launched a number of wikis on subdomains of an existing, established domain (wiki.leagueoflegends.com, wiki.warframe.com, hypixelskyblock.minecraft.wiki). Every single one of these has been immediately successful for getting articles other than the main page indexed on Google.

It doesn’t even need to be a popular domain! overwatch.weirdgloop.org, which launched just 7 days ago, is already doing better on page-indexing than any of the wikis currently in Google Jail.

This massively changes the strategy for new wikis

I’ve always said that the only legitimately hard part about moving your wiki off Fandom is that you have to wage a years-long Google war against the zombified corpse left over on Fandom. When you’ve got the game’s community on your side and you’re not dealing with the new-domain issues, this is genuinely not that hard…

…but when your wiki is simply not able to get its pages to show up because of this Google Jail, that doesn’t matter. You’re not going to get the majority of readers, which to me is what determines the success or failure of a move from Fandom.

So now, when wiki communities come to us asking if we can host them, the conversation about domains is much more complicated than it used to be. We have three basic options:

  1. stick with (say) overwatch.wiki - this is the natural place where everyone expects it to be, but it’ll get wrecked on Google
  2. try to arrange something with the game studio at (say) wiki.overwatch.com - this is achievable about half of the time, but not consistently
  3. put them, at least temporarily, at a subdomain of a domain we already control, that has decent domain authority

None of these are great options, but I hope it makes sense why we’ve been gravitating towards option 3 recently.

From an aesthetic point of view, I genuinely don’t like putting these wikis on weirdgloop.org. At all. I think “Weird Gloop” is an awesome inside jokey name for an org that only wiki nerds need to know about (3 separate people have made me bucket-of-weird-gloop plushies over the years!), but a pretty terrible (weird?) name for something that is user-facing and part of a domain that the general public needs to become aware of. It also, on a surface level, makes us look foolish for railing about how global branding for wiki platforms is super lame when suddenly we’re doing the same thing everyone else is.

There’s a bit of light at the end of the tunnel, though - based on our experiments so far, it seems like once we establish these new wikis fairly well on Google, it should be safe to move them off weirdgloop.org back to whatever the appropriate name was, while keeping the existing Google juice that the subdomain picked up. I’m hoping that in 6 months or so we can just 301 redirect overwatch.weirdgloop.org to overwatch.wiki, use Google’s change of address tool, and put these wikis to where they should have been in the first place.

Help me figure out what’s going on here

If you’re a wiki person with any sort of data about this phenomenon (or data about moving FROM a subdomain to a new root domain later), please hit me up in the Weird Gloop discord. Honestly, even if you’re just a general SEO person that has an inkling that something massive changed about new domain authority in 2024, let’s figure out what’s really happening.

The Daily Front Page 5 of 31
Tuesday, September 8, 2026 The Daily Front No. #260908 — Keys From Another Era
article

I've factored the RSA keys of a Certificate Authority from the 90s

by ahlCVA·▲ 488 points·123 comments·mcpherrin.ca ↗
RSA’s cryptography relies on the difficulty of factoring a large semiprime number.

… from the 90s.

I’ve been thinking about the security of RSA lately. RSA’s cryptography relies on the difficulty of factoring a large semiprime number, but what “large” means is an interesting question. The Web PKI deprecated 1024-bit RSA over a decade ago, and while I don’t know of anyone factoring a key of that size, it’s within the realm of possibility for a government or other organization with a large number of computers. Just a few days ago, someone factored the 862-bit RSA-260 key from the RSA factoring challenge. That’s the largest factorization I’m aware of. Today, the world uses RSA of at least 2048 bits, but even that will be deprecated soon with the risk of quantum computers in the future.

This led me to wonder: small RSA keys can be factored on even a modest desktop computer. And in the early days of the Web PKI, there were no standards, and no minimum requirements. Netscape shipped SSL support in 1994, and IE shortly afterwards. This was still the era of export restrictions on cryptography. Are there any keys small enough that I can factor? I don’t have any good reason to do that, but it seems like fun.

The spoiler is of course, yes, but first we need to find a key to crack.

Fortunately, root certificates were shipped with browser installers, and there are archives of both Internet Explorer and Netscape on archive.org. The archives aren’t comprehensive, but they should provide good coverage of old root CAs. I downloaded both collections and set Claude Code on extracting all the roots. I’ve hosted a Claude-generated webpage with all those old-timey, ancient roots. While I haven’t verified this LLM output is entirely trustworthy, it looks pretty plausible.

Using the filters on that site, we can find what small keys are trusted for SSL.

A table of old roots, filtered to 512-bit RSA roots trusted for SSL, showing the E-Certify RSA 512 Gold root shipped in Netscape 4.51

Aha! We have a target. Back in March 1999, Netscape 4.51 shipped a 512-bit RSA certificate authority trusted for SSL, and another for S/MIME. These two roots were both from the long-defunct Canadian certificate authority called E-Certify.

Later that year, the 512-bit RSA-155 was factored, so even in its era this was too weak and probably shouldn’t have shipped in the first place.

The E-Certify 512-bit roots were removed by Netscape in 2002. Unfortunately, Internet Explorer seems to have never shipped any 512-bit roots for SSL, so our fun will be limited to Netscape from a relatively small time frame.

Factoring the public keys in the root certificates will give me the two primes that I need to reconstruct the private key. I ran CADO-NFS on my Ryzen 9 5950X desktop; it took 32 hours to factor E-Certify RSA 512 Gold Server for SSL, and another 29 hours for E-Certify RSA 512 Gold Client for S/MIME. You can get the resulting private keys below.

Once you’ve installed CADO-NFS following its instructions, running it is easy. You need to get the modulus to factor from the cert, which can be done with a bit of python,

from cryptography import x509
f = open("gold-server.pem", "rb").read()
pubkey = x509.load_pem_x509_certificate(f).public_key()
print(pubkey.public_numbers().n)

The resulting value of N is what you pass to CADO-NFS, something like

./cado-nfs.py $N -t 32 --workdir /data/cert1

Assuming you’re somehow running Netscape 4.51 with a clock set before E-Certify roots expired on 2003-10-16, you can use these private keys to issue certificates. This describes zero people on the planet… except for this VM I set up.

Verifying that the issued certificates would work in Netscape 4.51 was an adventure in itself, as there is zero overlap in TLS capability between Netscape 4.51 and any modern TLS stack. So it was back to Claude Code to make a custom old-timey TLS server in Go. This site is publicly hosted at e-certify.fly.dev, which you are welcome to try out with your own copy of Netscape, but it won’t load in any modern browser.

Netscape 4.51 certificate viewer showing a certificate for e-certify.fly.dev issued by E-Certify RSA 512 Gold Server

Or, if you’d like to host your own website using these old-timey keys, the keys and tools are all in the repo at https://github.com/mcpherrinm/ancientroots. Or do worse, like MitM the SSL of all those Netscape 4.51 users with their clocks set to 25 years ago…

C=CA, O=E-Certify, OU=RSA Gold Server, CN=E-Certify RSA 512 Gold Server
-----BEGIN CERTIFICATE-----
MIIByjCCAXSgAwIBAgIBATANBgkqhkiG9w0BAQQFADBjMQswCQYDVQQGEwJDQTES
MBAGA1UEChMJRS1DZXJ0aWZ5MRgwFgYDVQQLEw9SU0EgR29sZCBTZXJ2ZXIxJjAk
BgNVBAMTHUUtQ2VydGlmeSBSU0EgNTEyIEdvbGQgU2VydmVyMB4XDTk4MTAxNjEz
Mzc1M1oXDTAzMTAxNjEzMzc1M1owYzELMAkGA1UEBhMCQ0ExEjAQBgNVBAoTCUUt
Q2VydGlmeTEYMBYGA1UECxMPUlNBIEdvbGQgU2VydmVyMSYwJAYDVQQDEx1FLUNl
cnRpZnkgUlNBIDUxMiBHb2xkIFNlcnZlcjBcMA0GCSqGSIb3DQEBAQUAA0sAMEgC
QQDNVQ93Ev7zgNaJAR1Z7gCydU6mky1e/B4EbY1NsdtfsitU9cELqg5uRJDPA40n
CDPeOyil1lJ5N8hekcqJAkkXAgMBAAGjEzARMA8GA1UdEwEB/wQFMAMBAf8wDQYJ
KoZIhvcNAQEEBQADQQB09SV6OeeDEP8Je3DOLNZ24U98NHqIBTDyB4sRpDmNdHum
+3rm4AYtznBxG5hEShO89bcWi3yJtBITGuTRDnMq
-----END CERTIFICATE-----
-----BEGIN RSA PRIVATE KEY-----
MIIBOgIBAAJBAM1VD3cS/vOA1okBHVnuALJ1TqaTLV78HgRtjU2x21+yK1T1wQuq
Dm5EkM8DjScIM947KKXWUnk3yF6RyokCSRcCAwEAAQJAWZ2FOWf+A9K4T2VAJS69
+SU/pW3YwHrysuYJZN56K0Iz+Hqd1hBhCeJ3/T+/cvXq+ctD0x3uOxU1rDSeCoMO
KQIhAPdbc1wTm3twWDYi1iXmRBXz97MYxRFld/KFr8x5+tK9AiEA1IG09a+oVg6j
NMDj6GD7spaD4q9t1wk/Nyq/MTLPkmMCIEzzYjvuzZvlI0wUIlLAA8Zgk1pgBk6X
Jm2IMVyHRgRxAiEAjZKgCTHuVu7HgiSjcTPzW0X1NTckWSc66zjaSR+Ns/sCICxT
yXoTXWQCWL6BMpf2no7IlvoEnBIrED2Frq3PxlEA
-----END RSA PRIVATE KEY-----
C=CA, O=E-Certify, OU=RSA Gold Client, CN=E-Certify RSA 512 Gold Client
-----BEGIN CERTIFICATE-----
MIIByjCCAXSgAwIBAgIBAjANBgkqhkiG9w0BAQQFADBjMQswCQYDVQQGEwJDQTES
MBAGA1UEChMJRS1DZXJ0aWZ5MRgwFgYDVQQLEw9SU0EgR29sZCBDbGllbnQxJjAk
BgNVBAMTHUUtQ2VydGlmeSBSU0EgNTEyIEdvbGQgQ2xpZW50MB4XDTk4MTAxNjEz
MzQwOFoXDTAzMTAxNjEzMzQwOFowYzELMAkGA1UEBhMCQ0ExEjAQBgNVBAoTCUUt
Q2VydGlmeTEYMBYGA1UECxMPUlNBIEdvbGQgQ2xpZW50MSYwJAYDVQQDEx1FLUNl
cnRpZnkgUlNBIDUxMiBHb2xkIENsaWVudDBcMA0GCSqGSIb3DQEBAQUAA0sAMEgC
QQBwCcT1iYlNyKPywB/kffD8esiCzGYJxSnTXQjU6ej/XxnA+9yqjzAMPtqFd094
wM89Vsmz9YOWSO6Qn6wOAs45AgMBAAGjEzARMA8GA1UdEwEB/wQFMAMBAf8wDQYJ
KoZIhvcNAQEEBQADQQAdktdM5AzW+0o96eHCHwD3UfzxPvjKxPEjiI/QTn+njHt/
BEJb9yZatONRckglVc9v8P8Dy8HZGQD0+Pn0uxhW
-----END CERTIFICATE-----
-----BEGIN RSA PRIVATE KEY-----
MIIBOQIBAAJAcAnE9YmJTcij8sAf5H3w/HrIgsxmCcUp010I1Ono/18ZwPvcqo8w
DD7ahXdPeMDPPVbJs/WDlkjukJ+sDgLOOQIDAQABAkASXZeaxFPsm0I8zb+snfR9
/saVolnrqhVEH5EODdXy3nSm6fpdrsDwQDDg52biXjIUizQYdd3OourOn0/PyyiB
AiEAyW4gFFgM0dJ+Y//vrIyQutLeH/5c/Vq3Lbcxv7R908kCIQCOZAAv8OCG/2oH
yKaYj/EoxcFZ9csVh/kjicO/JIX+8QIhAMEzAkvhBDLALYAmvDCJBkxa8rhHFdPf
jbCodGwGZ2WZAiB//XmBnlZkYl/PoVfGmNRgHun+0AZ9UxzqCeJvBQiBMQIgHWlr
0Sh1a8K2pQP2ktgmL3RH5+1Qd2wLx/hhGWtwPvU=
-----END RSA PRIVATE KEY-----

As a bonus, Internet Explorer 3.02 shipped a code signing CA called OU=Test VeriSign Commercial Software Publisher CA. I’m not sure what that’s for, but it’s just as easily factored. There are a few other test-looking 512-bit RSA keys in the ancient roots repo, so you should try factoring some more of them, or seeing what they could be used for!

Steve Weis factored this one in about an hour using a GPU cluster.

L=Internet, O=VeriSign, Inc., OU=Test VeriSign Commercial Software Publisher CA
-----BEGIN CERTIFICATE-----
MIIByDCCAXKgAwIBAgIQIsTi2AgEp+1OkpqVLROnazANBgkqhkiG9w0BAQQFADBl
MREwDwYDVQQHEwhJbnRlcm5ldDEXMBUGA1UEChMOVmVyaVNpZ24sIEluYy4xNzA1
BgNVBAsTLlRlc3QgVmVyaVNpZ24gQ29tbWVyY2lhbCBTb2Z0d2FyZSBQdWJsaXNo
ZXIgQ0EwHhcNOTYwNTAyMTcwMTUyWhcNOTcwNTAyMTcwMTUyWjBlMREwDwYDVQQH
EwhJbnRlcm5ldDEXMBUGA1UEChMOVmVyaVNpZ24sIEluYy4xNzA1BgNVBAsTLlRl
c3QgVmVyaVNpZ24gQ29tbWVyY2lhbCBTb2Z0d2FyZSBQdWJsaXNoZXIgQ0EwXDAN
BgkqhkiG9w0BAQEFAANLADBIAkEA1XqTpg1hQWq6IKvBXLor12GBTH6y4BHyOJjy
MJISn0ZWBqkbZiEfoWgaWXhEwrvxMWonLA2vYFHijofFG4YIjwIDAQABMA0GCSqG
SIb3DQEBBAUAA0EAKxuR/KPJPDNJgw0ORqqM7SlvIeqnVP71OEFytqPaZw/3TPEi
9oK6PzBvPdLBZlnjRd/hf58FW2+/PXxsFl78Mg==
-----END CERTIFICATE-----
-----BEGIN RSA PRIVATE KEY-----
MIIBOwIBAAJBANV6k6YNYUFquiCrwVy6K9dhgUx+suAR8jiY8jCSEp9GVgapG2Yh
H6FoGll4RMK78TFqJywNr2BR4o6HxRuGCI8CAwEAAQJAYM8tlegLarcToS1CiuKC
bzHwiNgMFkENL01sx0n21/MhmA6pBiSwVn3Pu6yvx6EMB1zbCsaBHs/VTyofMbRW
QQIhAOqdAX3bVZwAlZ5FDWyfGwgE0FXEIQ85kjcihrFeDs49AiEA6PBfC0YuNMSa
M135mH+MEY1GEmQe9xHcXpiYAPazCrsCIQDDtuw6oJEfDYHCwRn8xhGXs+RT18Q4
Xi9yXRP9vFgfhQIhAIYUTfD0XYZcIBIvJosj56D2u328iaJXcow0s1Hirp4fAiAG
bza0Qcc27d/0OzKOqep3JrDl/L5HidTj2XEqEXmnXw==
-----END RSA PRIVATE KEY-----

And finally, SSL Labs thinks my test site is great.

SSL Labs report giving e-certify.fly.dev an F rating, with many issues

The Daily Front Page 6 of 31
Tuesday, September 8, 2026 The Daily Front No. #260908 — Media Server, New Edition
article

Jellyfin 12.0

by 0xC0ncord·▲ 589 points·334 comments·jellyfin.org ↗
Finally giving books and comics the support they have deserved for years.

We are pleased to bring you Jellyfin 12.0, our new stable release. This release continues in the same direction as 10.11: finishing what the database conversion started, turning that new foundation into real performance work, and finally giving books and comics the support they have deserved for years.

If you just want a quick summary of what you need to know to get your system upgraded and running, please read on to the "TL; DR" section just below, or keep reading for a full explanation of all the major features and improvements in Jellyfin 12.0! You can also view full changelogs on the server and web GitHub releases.

- Cody

TL; DR

IT IS VERY IMPORTANT THAT YOU READ THIS SECTION BEFORE UPGRADING TO JELLYFIN 12.0! Failure to do so may cause issues! Always feel free to ask for help in our chat if you are unclear or run into trouble.

This release changes the database schema and actively rewrites data on first boot, so a backup is the only way back to your previous version.

  1. As always for major upgrades, ensure you STOP Jellyfin and take a FULL MANUAL BACKUP OF YOUR DATA AND CONFIG DIRECTORIES before upgrading!
  2. You must be running Jellyfin 10.10.7 or any 10.11.x release before upgrading to 12.0. Upgrading directly from 10.11.x is fully supported and no intermediate step is needed. If you are on anything older than 10.10.7, upgrade to 10.10.7 first, then upgrade to 12.0.
  3. Check your usernames before upgrading. Usernames are now case insensitive, so two accounts can't have names that only differ by capitalization. If you have users that match this pattern the database migration will fail.
  4. A full library scan is REQUIRED after upgrading. As part of fixing how alternate versions are stored, versions that Jellyfin grouped automatically, not ones you merged yourself, are cleared during the upgrade. Until you run the scan those versions will look like they are missing.
  5. The first scan after upgrading will take significantly longer than normal, and some movies may appear as newly added. This is expected. Jellyfin now checks every item in your library against the files on disk to clear out leftovers from previous versions, and anything that was previously filed incorrectly gets corrected as it goes. Do not stop the server while migrations are running.
  6. Once installed, please ensure you hard refresh (e.g. Ctrl+Shift+R or similar) and/or clear your web cache for your Jellyfin instance if you notice any UX anomalies. Bad cached assets are the #1 cause of such issues.
  7. Remove any third-party plugins before upgrading. Plugins built for 10.11 will not load on 12.0 and need updated builds from their authors, so give them time to catch up before adding them back. Official plugins have been updated for 12.0 support.
  8. Very old third-party clients will stop working. Support for the legacy /emby/ and /mediabrowser/ addresses has been removed, and the deprecated way of signing in is now disabled, including on existing servers. Clients that have not seen an update in years are the ones at risk here.
  9. This release contains security fixes, so we recommend upgrading rather than staying on an older release once you are ready.
  10. As always with major Jellyfin releases, bugs will exist. Please prefix bug reports with "[12.0]" so we can triage them quickly, and see point 1 one more time: take a backup.

And now on to the cool new features!

Why 12.0?

The most visible change in this release is the one in its name: we are dropping the major version "10" from our naming scheme. What would have been 10.12.0 is simply 12.0, and the server reports its version as 12.0.0. 10.11.x was the last release branch to use the old scheme.

The reason is the feedback we got after 10.11.0, which we first raised in January and confirmed in May. A release like 10.11.0 was, by any reasonable measure, a major release - it rewrote the library database - but the version number presented it as a minor update, and people upgraded with expectations to match. The leading 10. never changed and never told anyone anything, so all it did was push the number that actually mattered into the middle position and make "major" releases look minor. Dropping it means the first number moves when the release is big.

If you maintain anything that parses Jellyfin version strings - a client, a monitoring check, a deployment script, a container tag pin - this is the item to look at before upgrading.

Tuning the new database

Jellyfin 10.11.0 finished rebuilding how the library database works, which opened the door for the performance work in this release. We aren't finished, but you should notice a big difference when browsing.

One of the largest changes is how playlists and collections are stored. Before, everything in a playlist was kept as one big list inside the playlist itself, and the database couldn't look into that list. Anything we needed to know meant loading the whole playlist and unpacking it first. Getting the item count, working out how much you've watched, or showing a single page all cost the same as loading the entire thing. Editing was the same story, because adding or removing one item meant writing the whole list back out.

Every item in a playlist is now its own row. The database can count rows, return one page, and add or remove a single item without touching the rest, so this should be a big improvement on large playlists. Collections and boxsets were stored the same way and get the same fix.

Alongside that:

  • Deleting a large number of items at once no longer fails partway through.
  • Continue Watching, Next Up, rewatching, Latest Media for music, artist lookups, and the watched counts on folders should be faster.
  • Heavy database maintenance no longer runs during a library scan, so the two stop competing for the same resources.

Most of this shows up as pages that used to freeze no longer freezing.

One thing to note is that this work went into how Jellyfin reads your library, not into the library scanner. You may find the scanner is a little quicker as a side effect, but speeding it up was not a goal this time around.

What runs on first boot

As usual, a new release of Jellyfin comes with multiple database migrations. These migrations will take a while to complete depending on the size of your library and how much bad data has been accumulated.

If you would rather run that step deliberately than have it happen on first start the server now accepts --mode MigrateSystem , which performs the upgrade and exits without starting the rest of Jellyfin.

Multiple versions for episodes

Alternate versions have been a movies-only feature since they were introduced. In 12.0 they work for episodes as well, so a series with a broadcast cut and an extended cut, or a 1080p and a 4K copy of the same episode, can be grouped the way movies always could. This includes resume data that follows the version you were actually watching.

This is also the feature behind the required post-upgrade scan: making versions work correctly for episodes meant fixing how version links are stored, and automatically resolved versions have to be rebuilt from the files on disk.

Books and comics, properly this time

Books have often taken a backseat in favor of video playback in Jellyfin, and this release is the start of an effort to change that. Most of what the Bookshelf plugin used to do now has been moved to the server itself.

On the server:

  • Book metadata is read directly from OPF and ComicInfo files, or from ComicBookInfo comments, with no plugin required.
  • Posters are generated for EPUBs and every supported comic archive format, and external covers work for audiobook files.
  • Name, index, year, and series are read from book filenames, and volume and chapter numbers are picked up from comic filenames when they are there.
  • Page counts are extracted from comic archives and PDFs.
  • Chapters are extracted from audiobooks.
  • Bookshelf has been split into separate GoogleBooks and ComicVine providers, and a new OpenLibrary plugin provides metadata and images.
  • ISBN external IDs and links are supported.

In the web client:

  • There is a Modern book library layout with view types and paging, plus Authors, Collections, and Folders tabs.
  • Books show information about their authors, and authors show their books and audiobooks.
  • The reading interface has been redesigned and standardized across all book types, with unified fullscreen behavior and swipe navigation for PDFs.
  • Progress indicators work again for supported eBooks, sorting by index number and release date is available, and font size selection for EPUBs has been improved.
  • Background audiobook playback works on iOS devices.

The Bookshelf plugin has been deprecated. Its features have been merged into the server or extracted into the ComicVine and GoogleBooks providers.

Better recommendations, and a search that plugins can extend

Where "more like this" and the suggestion rows get their ideas from is no longer fixed. You now choose the source per library in the same place you already pick metadata providers, so you can use one source for movies and a different one for music.

ListenBrainz ships with the server as one of those sources. Point your music library at it and similar-artist suggestions come from real listening data rather than from tags alone.

Search works the same way now: a plugin can add its own results alongside Jellyfin's own. If you have ever wanted Jellyfin to search somewhere else at the same time, that is a plugin now rather than a fork. Both systems are written up for plugin authors in the release notes.

The Modern layout is now the default

The layout that shipped as "experimental" is now simply the Modern layout, and on desktop and mobile it is the default for anyone who has not explicitly chosen otherwise. The previous layout is still available and is now called Legacy. Televisions are unchanged: TV devices continue to use the TV layout, which still runs on the legacy app.

Alongside making it the default, the layout got the polish that implies:

  • All themes - Dark, Light, WMC, Blue Radiance, Apple TV, and Purple Haze - now derive from a shared base theme built on CSS variables. If you maintain a custom theme, this is worth a look.
  • The library toolbar has been merged into the app bar, with a sticky library header.
  • Collections and playlists tabs are available on all libraries, and collections appear on item details pages.
  • Music Videos, Mixed Media, Collections and Playlists, and Books views were all updated, and Home Videos and Photos libraries gained default tab options and a folder view.
  • New filters for audio and subtitle languages, a Reset Filters button, and studio search.

Photo of the updated Modern UI

Photo of the updated Modern UI - Movie Library

Changes you may notice after upgrading

A few things behave differently than they did on 10.11, beyond the legacy client removals covered above:

  • Subtitle settings are now configured per library rather than once for the whole server, so the old server-wide subtitle options no longer exist.
  • Sorting is more consistent, which does mean some libraries will be ordered slightly differently than you are used to.
  • Artwork is no longer stretched past its real size. Low resolution posters now appear at their actual size instead of being blown up to fit, which looks smaller but sharper.
  • .ogg files are treated as audio, not video. If you had .ogg video files, they will be re-sorted on the next scan.
  • Usernames can now be changed between upper and lower case. As part of that, two accounts can no longer have names that differ only by capitalization. If your server has usernames that differ only by capitalization, the database migration will fail.
  • Symlinked media is followed when something is played rather than when the library is scanned.

Security

This release includes a number of security fixes on both the server and the web client. Several of them close off ways a crafted request could reach files outside the directories Jellyfin is supposed to serve. Others prevent the setup wizard being re-run on a misconfigured server without signing in, reject plugin packages with unsafe names, apply parental controls in more places, and fix cross-site scripting issues in the web client.

New Features & Enhancements

User Experience - Web Client

  • A "still watching" prompt.
  • , and . scrub frame-by-frame during playback.
  • Chapter names appear in the preview bubble as you drag along the seek bar, and the playback info overlay is more compact while showing more detail.
  • Folders can be marked as played.
  • An improved Upcoming view, and Play All and Shuffle buttons on series libraries.
  • The screensaver is suppressed while viewing photos or reading, and a screensaver time setting is available in the Modern layout.
  • Crew members with multiple roles are merged into a single card, and TV show creators appear on item details.
  • A configurable delay for the photo slideshow.
  • The web client caches more data in your browser, so screens you have already visited come back faster, and artwork no longer looks blurry on high resolution displays.
  • Game controller navigation fixes, keyboard shortcuts that work on non-Latin keyboard layouts, and support for rewind and fast forward buttons on remotes.

Administrator Experience - Web Client

  • The log viewer can follow a log live instead of needing a refresh.
  • Sorting and filtering on the Activity page.
  • The source of similar-item and recommendation data is configurable per library.
  • Jellyfin now marks its cache directories with CACHEDIR.tag, so backup tools know to skip them instead of backing up these files.
  • Clients can ask for responses in a specific language, so a server with users who prefer different languages behaves better.
  • The database file can be kept somewhere other than the default location.
  • Disabled plugins are no longer re-enabled on restart.
  • Backups no longer fail outright when they hit a corrupt record, and you now get a warning before restoring a backup, or before starting one while a library scan is running.
  • A restyled startup interface that shows version and activity information.
  • Collections and playlists can be filtered by library, so each library can show only its own.
  • More ways to filter when searching for people, the option to prefer a title's original language for audio, and the new "Collections" listing on item pages.
  • Live TV: XMLTV background images and episode thumbnails are imported, unreachable "server-local" streaming URLs are no longer returned to clients, and XMLTV guide imports now skip programs whose data has not changed, making repeat guide refreshes considerably cheaper.
  • Metadata: TVDB IDs work for movies, AudioDb artist search, ReplayGain album gain is read from music files, MusicBrainz lookups are more resilient, WEB-DL tags in filenames are recognized, and provider IDs can be set on season and episode folder names.
  • Curly braces and parentheses are now supported, along with the existing square brackets, for parsing Provider Identifier values. New aliases are also supported: tvdb for tvdbid, imdb for imdbid, and tmdb for tmdbid.
  • The optimize database task now analyzes and vacuums the DB when triggered and at shutdown.
  • The Refresh People task now properly cleans up removed people and checks for updates for people with missing images.
  • Live TV: EPG now honors SchedulesDirect API error codes, preventing lockouts. HDHomeRun tuners will now direct play on clients that support it. M3U tuners will no longer direct play and will instead default to remuxing.
  • The playlist library image can now be set for all users through the Metadata Manager. Duplicate tracks are now allowed in playlists.

Transcoding and Media Handling

  • We have moved to the new upstream FFmpeg 8.1.
  • Optimized CUDA transposing, OCL scaling, and OCL tonemapping performance, the last of these specifically on Mali GPUs.
  • HLG tonemapping now uses the EOTF from BT.2446 Method B.
  • A spec-compliant dvh1 HLS variant for Dolby Vision Profile 5 for better device compatibility.
  • Fixed a potential A/V desync in HLS when transcoding video while remuxing audio.
  • Subtitle writing now goes through SubtitleEdit, which avoids the old SSA to ASS conversion and the loss of styling that came with it.
  • VobSub subtitle support, external subtitles can be embedded into MKS when transcoding, the subtitle extraction timeout is configurable, and image-based subtitles can be rendered by the client during remuxing.
  • A device profile option for Android TV boxes that cannot handle rotated video, so phone footage shot sideways plays the right way up.
  • Trickplay picks up files that already exist on scan instead of regenerating them, no longer produces duplicates for interlaced video, works with bad timestamps in the source file, and cleans up after itself when generation fails.
  • Better device support in the web client: Dolby Vision in MKV files on webOS 25 and newer, AV1 direct streaming on TV clients, anamorphic video direct play on Tizen, and on iOS both background playback with the screen off and a fix for audio normalization affecting pitch and speed.

Client Development Changes

The following changes apply to all client application developers. Please review thoroughly and update your applications as required.

HTTP API

  • Deprecated authorization mechanisms are now disabled by default. See the details in this pull request if you have not migrated yet.
  • GetItems is now asynchronous and applies recursive when filters are requested, limited to requests that include includeItemTypes. The same query can return a different result set than it did on 10.11.
  • ItemByName responses are restricted and people are deduplicated.
  • Newly obsolete but still functional: GetTrailers (use GetItems with includeItemTypes=Trailer), GetArtists and GetAlbumArtists (use GetPersons), GetArtistByName (use GetPerson), GetMusicGenre (use GetGenre), the music genre instant mix endpoints (use GetInstantMixFromItem), GetRecordingsSeries, and the startup routes (use the configuration endpoints). UserDto.HasPassword is also obsolete and no longer provides useful information. The HLS controllers are hidden from the specification.
  • Swashbuckle has been updated to v10, which changes the generated OpenAPI document. SDKs need to be regenerated.
  • Removed routes: POST /Users/\{userId\}/EasyPassword, GET /Items/\{itemId\}/CriticReviews, GET /Environment/NetworkShares, POST /System/MediaEncoder/Path, and GET /LiveTv/Recordings/Groups/\{groupId\} were all obsolete no-ops returning 403, 404, or an empty result. GET /QuickConnect/Initiate did work and was an alias for the POST route, so clients using the GET form need to switch to POST.

As a general reminder of our API support policy: if an endpoint is not listed in the OpenAPI specification it should not be used, and if an endpoint or parameter is marked obsolete it should not be used. Deprecations will normally be marked for an entire major release cycle before removal.

Plugins

  • The server now targets .NET 10, and several plugin interfaces changed. Plugins need to be retargeted and rebuilt for 12.0.
  • New things plugins can do: provide search results, provide similarity and recommendation data, provide comic metadata, provide unaired and missing episode data, use aggregated credits for cast members, save chapters for non-video items such as audiobooks, handle a password reset for a username the server does not recognize, and clean up their own extracted files when an item's data is pruned. Live TV plugins can also query Schedules Direct availability through the server instead of reimplementing it.
  • Plugins that worked with alternate versions or playlist contents need attention, since those are no longer stored inside the parent item. There are new library methods for both.
  • The full list of new, changed, and removed interfaces is in the release notes - ISearchEngine, IAuthenticationProvider.HasPassword, parts of IItemRepository, and a few IUserManager members are the notable breaks.

Deprecation of internal TLS/SSL support

In the 10.11.0 release notes we announced that internal TLS/SSL support would be removed in this release. That removal has been postponed to a future version. The reasoning has not changed - we continue to recommend running Jellyfin behind a reverse proxy - so if you are running an Internet-facing instance using Jellyfin's built-in TLS, this is extra time to migrate, not a reprieve.

Happy Watching!

The Daily Front Page 7 of 31
Tuesday, September 8, 2026 The Daily Front No. #260908 — The LAN House
article

We built our house for LAN parties (2024)

by fittingopposite·▲ 464 points·320 comments·lanparty.house ↗
12 PCs are built into the walls.

This is our basement game room.

A photo depicts the game room, 18 feet wide and 24 feet long, full of people playing video games. The room is a finished basement room; it is carpeted with beige walls and no windows. On the far wall is a 98-inch TV. Along both side walls are rows of computer screens, six on each side. Each screen sits inside wood panel cabinetry which folds out to provide a desk in front of it. People are sitting at all the desks, playing video games. In the middle of the room, facing the TV are two recliner loveseats. Someone is sitting at one of them, holding a controller and playing a video game on the TV.

12 PCs are built into the walls.

Game Stations

These fold up nicely when not in use.

A photo shows an up-close view of the game stations that lined the left wall in the previous photo. In this photo, one station is open, while the others are all closed. The closed stations merely look like wood paneling on the wall; the monitors are entirely hidden. The open station shows that a panel in front of the monitor folds outwards and down to become the desk on which the keyboard sits. Two panels below this one swing outwards to form supports for the left and right sides of the desk. The monitor is on a retractable mount such that it can be pulled away from the wall, closer to the user. There is a power outlet behind the monitor, suitable for charging a phone while sitting at the station.

Closed station

Open station

Dance Dance Revolution

There are DDR pads in the floor.

A photo shows the floor of the game room. The floor has a row of four removable panels. The panels are aluminum with carpeting on top matching the rest of the floor and retractable handles at either end. One of the panels has been lifed and placed to the side, revealing a Dance Dance Revolution game pad underneath.

The Upstairs Office

We can game here too.

A photo depicts the office room, 16 feet wide and 20 feet long. Large windows on the far side show it is night time, with distant lights lining the horizon. In the middle of a room is a large wood conference table, aronud which several people are gathered, playing a board game. The table surface appears to be made up of several wood panels with cracks in between.

Board games are games, right?

Computer Mode

The table also transforms to reveal six more game stations.

This photo depicts the same room as before, but now the conference table has transformed into six computer stations. The panels that formed the table surface before flip up, revealing a monitor mounted underneath. A second layed below them becomes the surface on which the keyboards sit. People are sitting at all the computers, playing video games. At the back of each station there is a slot where the keyboard and mouse are stored while not in use. Extending from the back side of the conference table is a normal-looking desk surface that does not fold up. Instead, a large curved monitor sits on top.

Two sit-stand desks in the back are our personal workstations.

(But that's not me in the seat. That's Lester.)

Cable Management

Four photos reveal how cables are managed. The top-left photo shows a trough in the center of the office table containing a mass of cables from the six computers. This trough is only visible when the stations are open, and only if you look from the side to see between the rows of monitors. The top-right photo shows the space under the office table. Wires come out of the middle of the table and enter holes in the floor. The lower-left photo shows where these cables eventually endup, coming out of cable conduits in the wall of another room. The lower-right photo shows another set of cables entering this room through a hole in the wall; these cables come from the game room.

Video and USB cables pass into conduits...

Which eventually lead to...

The Engine Room

The photo depicts a room with three computer racks side-by-side. The center and right racks each hold ten computers; these are the game machines hooked up to the monitors seen in previous photos. The right rack is mostly unpopulated, but contains two computers that are our personal desktops, and a server machine, and one chassis that is currently empty. All of the machines except the server are in 4U rackmount chasis; the server is 2U. The racks are embedded in the wall. To their right is a closed door which leads to their back side. Above the racks is an air vent which provides cold air from a dedicated air conditioner; the intake is on the other side of the machines.

20 identical game machines.

All game machines netboot from a shared disk image on this server.

Our personal desktops

Dedicated A/C intakes from behind the machines and supplies cold air here.

Door leads to the "hot side".

The "Hot Side"

It's actually not that hot.

A photo depicts the back side of the computer racks. To their left is the door, now open, leading to the front side. Many wires are visible in the photo, some organized in a cable track hanging from the ceiling, others strewn haphazardly on the floor. Directly behind the server racks there is a column full of power outlets which many of the machines are plugged into. Also visible in the picture is various HVAC equipment, inculding an air handler and a HEPA filtration unit. These actually don't serve this room but were placed here because they are ugly.

Networking

A photo shows an up-close view of a stack of five network switches mounted at the top of the rack furthest from the door. An enormous bundle of category 6 network cables comes out of a hole in the ceiling and connects to the network switches. A Google Fiber box is mounted on the wall nearby.

Lots of cat6:

  • 35 wall boxes with 4 ports each
  • 7 PoE WAPs
  • 8 PoE security cameras
  • 3 PoE intercoms

All network equipment and cameras are UniFi.

Egads! This photo has made a lot of network engineers angry! Response in the Q&A.

2 Gbps fiber internet

Components

This is the hardware before it went into the machines.

(Note: This was in 2023.)

A photo shows shelving in the engine room full of computer components still in their boxes. This was taken before the machines were built. 20 or so of each component are present, including CPUs, GPUs, motherboards, heat sink fans, power supplies, RAM, NVMe drives, and sound bars. There are also a large number of unmarked cardboard boxes; these contain spooled DisplayPort and USB cables of lengths varying from 35 to 100 feet.

Intel Core i5-13600

32GB RAM

GeForce RTX 4070

An insanely unstable motherboard with 10G ethernet on-board.

Don't buy this. :(

Assembly

My (Kenton's) friends from junior high helped assemble the machines and pull the cables.

A photo depicts the game room from earlier. Four guys are standing around two folding tables set up in the middle of the room, where they are assembling the computers. Each person has an open chassis in front of him with components and boxes strewn about.

We've been doing LAN parties together for nearly 30 years.

Lester

Jesse

Peter

Kenton

Living Room

This house has normal parts, too!

A photo depicts a large room, 40 feet long and 30 feet wide, which contains couches, a dining table, a full kitchen, and hardwood flooring. The ceiling is 14 feet high, held up by huge glulam beams. The far wall is mostly glass, composed of a row of sliding glass doors and transom windows above them. Beyond the glass is a distant view with lots of trees and hints of houses nestled between them. There is also a swimming pool, though it is barely visible. The kitchen features an island counter with an induction stovetop embedded within it. The vent hood above descends all the way from the ceiling. Wooden beams extend horizontally from the hood with pots and pans hanging from the beams. A series of bar stools on the opposite side allow people to sit at the island and chat with Jade while she cooks. The sink is along the back wall, making it difficult to keep Kenton company while he washes dishes.

Catwalk

(Same room as above, after rotating right.)

A photo depicts the living room again, showing a wall that wasn't visible before. The wall features at 85 inch TV displaying artwork. To the right of the TV is a hallway entrance. Above both of these, a long wood shelf extends across the length of the wall. To the right of the hallway entrace is more wall space, where more wood shelves are arranged such that a cat could hop up from shelf to shelf to get to the top. At either end of the long top shelf are two cat doors through the wall.

These shelves are designed for cats to climb.

Cat doors lead into kids' bedrooms.

Kids' Rooms

Beds are lofted.

A photo depicts a kid's bedroom. The room features 14-foot ceilings similar to the living room. The right side of the room has two levels, a 6-foot-tall walk-in closet and a loft above it. A black metal ladder provides access to the loft. The loft is about the size of a queen mattress, but currently features a twin bed. To the right of the bed is a small bookshelf with children's books. To the left of the bed is a narrow ledge along the wall. At the back of this ledge is the other side one of the cat doors leading to the living room. Also above the ledge, a door leads to a kid-sized hallway which leads to the other kid's loft. On the opposite side of the room from the loft, a wall box provides access to a cable conduit leading to the engine room, but nothing is currently connected through it. A computer desk could be placed here when the child is old enough to have one.

Hallway to other kid's loft.

(With lockable door.)

Cat door from living room.

6' closet below loft.

Cable conduit to engine room.

(For future computer desks, when they're older.)

The other kid's room is a mirror image of this.

Call Rooms

Two call rooms allow taking meetings without bothering other people.

Of course, they are also game stations.

Two side-by-side photos depict two small rooms, each featuring a desk with a monitor and keyboard on it. The room on the left features windows showing a view of the yard. It also has a small tread mill, appropriate for taking a meeting while walking. The room on the right has no windows and looks drearier. A closed door in the back of the room leads to a mechanical closet.

Cat Restrooms

Two side-by-side photos depict rooms with slanted ceilings located under staircases. Each room has a cat litter box situated on a litter rug on top of a tiled floor, and an exhaust fan in the wall. Each room features two cat doors with plastic flaps leading to adjacent rooms, as well as a human-sized door.

Exhaust fans

Cat doors allow cats access to bedrooms when human doors are closed.

Space under stairwells is dedicated to cat litter boxes.

Guest Rooms

We have two guest rooms for out-of-town visitors.

Two side-by-side photos depict two guest bedrooms. One room features a king-sized bed. The other features a bunk bed with a trundle beneath, which could sleep three to four people.

The Roof Deck

My favorite part of the house.

A photo depicts a roof deck with an epic view of trees, houses, and city lights in the distance. The deck hosts four swivel chairs and a large L-shaped sectional sofa. A covered trellis above the deck provides shade. Thirteen adults and four children are present, chatting and enjoying the view. Although not visible in the photo, the deck always experiences a pleasant breeze.

30-mile view

Misting fan, great in the heat!

Cat

A photo of the cat, Garply, perched majestically on top of the cat shelf.

Cat

Q&A

About Us

Whose house is this?

We are Kenton Varda and Jade Wang. We've been married since 2014. This is our house, where we live with our two kids, in Austin, Texas.

Kenton is a software engineer. He:

Jade is an entrepreneur. She:

Who designed and built it?

We bought the property as an empty lot and designed and built the house from scratch.

The lead architect was Richard Varda, Kenton's father. Rich is an accomplished architect who has designed everything from houses to skyscrapers. Among other things, he designed The Kingdom Centre in Riyadh, Saudi Arabia, and The Musical Instrument Museum in Phoenix, Arizona. Rich was also VP of Architecture at Target for many years.

Many features of the house were designed directly by Kenton and Jade. For example, Kenton designed the game station cabinetry, office conference table, cat shelves, conduit routing, and most technical features of the house.

Additionally, we worked with Thom Lasley and others at RSP Architects. RSP normally does commercial buildings, not houses, but Rich was working with RSP on other projects at the time and had known Thom for decades, so it made sense.

Blue Horse Building & Design was our general contractor who built the house.

When was it built?

The house was completed in late 2023. We originally bought the property in 2019, and began construction in mid-2021. Yeah, it took a while.

Who made this site?

This site was designed and coded by Kenton with photos taken by Kenton, Jade, and Rich. We are neither web designers nor photographers. That's why the design is weird and the photos are kinda meh. Nothing on this site is sponsored.

Didn't someone else do this already? Like a decade ago?

Yes, someone did. Specifically, me (Kenton)! In 2011 I completed a LAN-party optimized house in Palo Alto, California, and it went viral. It, too, was in collaboration with my father (but I had not yet met Jade).

But contrary to what most people think upon seeing the pictures, the original LAN house was not very big: 1400 sqft. This made for a pretty awesome bachelor pad, but would have been a bit cramped for raising a family. The new house is much, much better.

I've never heard of anyone else having done anything like this. This surprises me! But, surely, if someone else did it, someone would have told me about it? If you know of another, please let me know!

Why did you move to Austin?

In 2019, my team at Cloudflare was growing. But, most of the hiring was happening in the Austin office, while I worked from San Francisco. Meanwhile, with our first child on the way, Jade and I needed a bigger house, but we really could not afford to buy (much less build) anything bigger in Palo Alto.

I suggested to Jade: Should we move to Austin? Jade initially said no, because she wanted our kids to benefit from Palo Alto's school district. At the time, it was rated #12 in the nation. But, looking closer at the rankings revealed a surprise: The Eanes school district in Austin was #8. When I showed this to Jade, she changed her mind.

Ironically, just after we moved, covid hit, and Cloudflare became 100% remote. The original reason for the move—being closer to my team—became moot. In retrospect, we could have gone anywhere! Not that there's anywhere in particular that I think would have been better. It's just funny.

How much did this cost?

The 22 game machines (including monitors, cables, and peripherals) cost about $75,000 in total. The house overall was a 7-digit number. Sorry, I'm not comfortable being any more specific than that.

I actually find it funny how cheap computers are. The cabinetry around the game stations cost a similar amount to the computers powering them. Think about that! The cabinetry is just a bunch of wood, cut into fairly large pieces. Maybe a few screws and hinges. Whereas the computers in total contain more than a trillion transistors, each of which had to be carefully etched in the right spot connected correctly to all the others. A trillion! That's 1,000,000,000,000!

Of course, the point is, GPUs are mass-produced. Cabinets for game stations in LAN-party-optimized houses are, um, not.

Where did you get the money?

This question also makes me uncomfortable, but I know from last time that people will ask it, and if I don't answer, people will make things up. So, OK, I will tell you.

I (Kenton) started my career at Google in 2005, which was a pretty good time to be there, even as a junior engineer. They gave me stock, and that stock went up. With no family and no hobbies aside from video games, money piled up in my bank account. After 4-5 years of saving, I was able to make a $200k down payment on a $1M construction loan (later converted into a mortgage). In many places that could have bought a mansion, but in Palo Alto it got me a tiny sliver of land and a 1400 sqft house. Still, I considered myself rich.

Then things got weird. The success of Google, Facebook, and others in Silicon Valley had produced a large number of people in the area with a lot of money, ready to buy houses and start families. Meanwhile, NIMBY policies were blocking all new development. As a result, Palo Alto housing prices, already high, went bonkers over the course of the 2010's. My house, which had cost me $1M, became worth over $2M.

This is absolutely unfair! I built the house so that I could throw parties and play video games with my friends. These do not seem like activities that are supposed to make money. But it made me a million dollars in profit. What?

So basically, we were able to parlay that into a bigger house in Austin. And just in time! While Austin housing prices had been growing for a while, they really exploded after 2020. We bought our property in late 2019. (Aside: Today, in 2024, Austin housing prices are now actually declining rapidly, due to an enormous amount of new housing having been built over the last few years. Yes, it can be done! Now is a great time to move to Austin!)

For completeness, I'll mention two other not-insignificant sources of money. First, Jade and I joined Cloudflare in 2017, a couple years before its IPO. This was a good time to join. The stock has done very well, even hitting a high point just as we started construction. Second, Jade has long been an active (albeit small-scale) seed-stage investor in startups. Before she met me, she invested in Matterport, in their very first funding round. They went on to become a public company. We were already in the process of moving to Austin before either Cloudflare or Matterport went public, so we didn't have any idea at the time how much they'd pay off. But they did, and this allowed us to expand scope.

Yes, we are very lucky!

About LAN Parties

What's a LAN party?

It's where everyone goes to one place and plays multiplayer videogames with each player having their own computer, connected over a Local Area Network.

In the 90's, when internet connections were bad, LAN parties were popular among PC gamers. But when broadband got decent in the early 2000's, LAN parties largely died off, as people could now play multiplayer games without leaving the house.

But I think that sucks. For me, LAN parties were never just about the game. LAN parties are a social event, and the game is merely a catalyst. LAN parties serve very much the same purpose as normal parties serve for normal people, but the game makes it easier and more fun for introverts to be social. Indeed, LAN parties are so fun that they traditionally go all night long—much longer than your normie parties—because no one wants to stop.

My (Kenton's) first LAN party was for my 14th birthday, in 1996. We had four computers: a 486sx 25Mhz, a 486dx 33Mhz, a 486dx2 50Mhz, and a Pentium 120Mhz. We played Doom 2 all night long. It was the most fun I'd ever had.

That same year, I had another LAN party on New Year's Eve, which I have done every year since. At least two out of three (and usually all three) of my junior high friends from Minneapolis have attended every single time.

In total, I have hosted or attended probably over 100 LAN parties. All of them were at someone's house, typically with 8-16 people. I have never been interested in big commercial LAN parties with hundreds of attendees—at that point it feels no different than playing with randos on the internet.

How can you have LAN parties when no games support LAN anymore?

We have an internet connection.

Dragging over your own computers is part of the fun of LAN parties. Why build them in?

I originally thought this too. In fact, my first LAN party house was designed to let people bring their own box and quickly connect it to a station. But nobody ever did. Not once. Some people brought laptops, but they only ever used them if the game stations were all in use already.

I do feel a lot of nostalgia for the days of trying to pack four people, four computers, and four monitors into one car on the way to a friends' LAN party, setting up machines on haphazardly arranged card tables with questionable seating arrangements, daisy-chaining power strips and network hubs. I'm a little less nostalgic for the experience of trying to copy game files over the network to get everyone on the same version, or pitying the one friend who inevitably has to reinstall Windows and doesn't manage to get in-game until after midnight.

The fact is that, while these things were fun, they were Type II fun. And they made LAN parties inaccessible. Even the most enthusiastic of us didn't really want to do all that more than, like, 3-4 times a year, and a lot of people—even those who like games—really don't care to do it at all. I'm not even sure if I could do it anymore, as a 40-something with two kids!

With computers built-in and set up in advance, we can commence gaming essentially immediately upon enough people arriving. We can get a lot more people to participate. If someone can only drop by for an hour or two, they can still play. And we can do it all much more often: at one point I was hosting LAN parties every other weekend.

Why have everyone's backs to each other rather than facing each other?

Some people feel REALLY strongly about this. A few have asserted that the whole setup is worthless because we have our backs to each other.

So, first of all, the office actually has everyone around a table facing each other. It's only the basement that has backs to each other.

In my opinion, though, facing each other is worse. You can't actually see the people on the other side unless you stand up, because the monitors are in the way. So, when we take a break between rounds to chat, everyone either has to get up, or talk over a wall. It's much nicer to just rotate in our chairs.

Additionally, the facing-each-other approach takes a lot more space. The game room is designed so that we can fold everything away, slide the couches away, and have a big empty space where we can do other things, like play DDR, watch movies, play VR games, or set up a workshop. If we had the game stations in the middle of the room then it wouldn't be a suitable space for any of that.

So, I think this was the right call. But I can certainly see opinions differing.

What games do you play?

Lots of things. We are trying new games all the time.

In general, we prefer cooperative games, or at least team games. Most attendees are not actually hardcore gamers, so skill levels vary widely. In a free-for-all competitive game, one or two people tend to dominate while several can't get anywhere—that's not fun. Cooperative games allow everyone to participate according to their ability. Team competitive games allow us to carefully balance the teams, though this can be tricky.

Our favorite LAN game of the past five years is probably Deep Rock Galactic, in which the players are a team of space dwarves mining an alien planet for resources. Prior to Deep Rock, Left 4 Dead 2 served the same purpose, or TF2 Mann vs. Machine mode.

We have also played quite a few survivalcraft basebuilder games, such as Ark or Factorio.

We played a lot of Overwatch, back when it was good, often with us on a team against internet randos, and occasionally team-vs-team entirely within the house.

We still play Unreal Tournament 2004 regularly, usually in Assault or Onslaught mode.

Are your parties open to the public?

No. Sorry, you must be invited. I'm sure you understand: For security reasons, we can't just let random people on the internet into our house.

If you want to get invited, you need to figure out some way to get to know one of us, or one of our friends who can vouch for you.

One straightforward (albeit not easy) way to get an invite is to get hired at Cloudflare. ;)

Do you buy the games for every computer?

No. Everyone logs into their own Steam/Epic/etc. accounts, and must buy the game we are playing. We do, however, try to preinstall everything people might play. Fortunately, Steam and similar game launchers are able to use the same copy of the game for all users.

About the tech

How do you keep 20 computers up-to-date? That sounds tedious!

I don't! I maintain a single disk image for all the machines. Before the party starts, I will install all needed games and updates on that single disk image. The machines all boot off of a network drive based on this image. Each machine gets a copy-on-write overlay on top of the main image, so that guests can make changes to their machine which won't be seen by any other, and will be deleted at the end of the party.

I have put all the scripts I use for this up on GitHub.

What exact hardware is in the game machines?

For the game machines, I mostly targeted the Logical Increments "Outstanding" level, which aims to stay just under the inflection point where high-end price gouging sets in. This hardware was purchased in mid-2023, so obviously it's no longer the best choice today. But here it is:

  • CPU: Intel Core i5-13600KF — Games care much more about GPU than CPU.
  • CPU heat sink: Scythe Fuma 3 — Mostly chosen because it fits in a 4U chassis while the Noctua fans do not.
  • GPU: Gigabyte Windforce RTX 4070 — Not super, not Ti, just 4070
  • Motherboard: Gigabyte Z790 Aorus Master — I do not recommend the motherboard! I was aiming to get the cheapest board that had 10G network on-board. The only reason I needed 10G network was because of the netboot setup: these machines access their primary storage entirely over the network. Otherwise, you absolutely do not need 10G for gaming. In general, I think expensive motherboards are a scam: they often provide no real benefit, but are marked up just because some suckers will buy them. Moreover, this particular motherboard is horribly unstable, frequently blue-screening shortly after boot. (Luckily, it seems to settle down after a minute or two and rarely disrupts actual games.)
  • RAM: Corsair Vengeance 32GB (2x16GB) DDR5 5600 — Every time I see "Corsair Vengeance" I can't help but chuckle. Vengeance??? It's just computer memory, dude.
  • PSU: Corsair SF750 — Platinum efficiency rating actually kinda matters when you're running 20 of them.
  • Storage: Corsair MP700 Gen5 1TB NVMe — The storage is not actually used due to the netboot setup, but my hope is eventually to improve the design such that the copy-on-write overlay can be on local disk instead of server-side, using this storage.
  • Chassis: Rosewill RSV-L4500U 4U — Realistically if you're rack-mounting consumer-grade hardware you have to go with a 4U chassis. I really tried to make 2U or even 3U work. It doesn't work.
  • RGB: None
  • Incredibly long DisplayPort & USB cables: Monoprice SlimRun
  • Monitor: Corsair Xeneon 32UHD144 (32" 4k 144Hz) — I chose this monitor mostly because it had the slimmest profile out of all the similarly-capable monitors I found. This was important in order to fit into the cabinetry. Unfortunately it does not have built-in speakers, necessitating the sound bar.
  • Keyboard: Logitech K120 Wired — The world's cheapest keyboard at $13 a pop. Works perfectly fine for all gaming needs. (See next question, below.)
  • Mouse: Logitech M500s Wired — A cheap 5-button mouse. Unfortunately I've found the middle button is difficult to use without accidentally scrolling. I miss the old MX518.
  • Sound bar: Soulion R30 Wired — This is just some crap I found on Amazon without doing a lot of research, but they seem to work fine. Unfortunately the monitors don't have built-in speakers, but this sound bar is probably better than a typical built-in speaker, so in the end I'm happy with it.

Also, in case you are wondering:

  • Game room TV: Samsung 98" QLED 4K Q80C
  • Living room TV: Samsung 85" "The Frame" QLED 4K LS03B
  • Workstation monitor: Samsung 57" Odyssey Neo G9 Dual 4K — Ironically, I use this for work, not games. It's great for coding. Playing games on it gives me vertigo.
  • Dance Pads: L-TEK Ex Pro X

After spending all that money, why did you get the cheapest possible keyboard and mouse?

To be honest, when I've used more-expensive "gaming" keyboards and mice I have never actually liked them. I mean, they are fine, but I don't feel like I get any benefit. If anything there are drawbacks: they tend to have extra buttons that enable gimmicky features which only mess me up when I touch one of those buttons by accident. And they tend to be loaded with RGB, which I just don't care about. So I see no reason to spend money on any of that.

Note that for work purposes, I absolutely need a better keyboard. (I use a Kinesis Advantage 360.) But for WASD purposes, the cheapest possible keyboard is all I want. When it comes to mice, for either work or gaming, I want something that has five buttons and fits reasonably in my (large) hand but that's all.

With all that said, guests are welcome to bring their own keyboards and mice if they like. Each station has a USB hub where you can plug in your favorite peripherals.

OK but why no mousepads?

Honestly I never even thought about it. I haven't used a mousepad in decades. None of my friends ever commented on it. But I guess it might be a good idea, to prevent wear on the wood finish.

Speakers? No headphones? Isn't the sound pollution annoying?

Honestly, it really isn't. I know it sounds weird, but in my experience it's just no big deal to be able to hear a bunch of other people's sound around me. It doesn't seem to detract from the game. In some cases it's nice to be able to hear what my teammates are up to.

Of course, if we are playing a team vs. team competitive game, we wouldn't want teams hearing each other's sound. But in that case we have one team go to a different room entirely. Or if we're playing a game like Dead By Daylight or Last Train Outta' Wormtown, where one person is the monster against everyone else, then that person goes to one of the call rooms.

It's true that without headphones, you lose positional audio. That might matter in some games, but mostly in the games we play, it doesn't matter much.

The down side of headphones is that it would make it harder for us to talk to each other. Sure, you can get open-back headphones. But it's still creates a psychological barrier between people, which works against the point of being together at a LAN party.

And honestly, would you really want to use loaner headphones that who-knows-how-many people have sweat into? Seems kinda gross?

With all that said, players who want to use headphones are absolutely welcome to bring their own, just as with keyboards and mice. Every station has a USB hub where you can plug in whatever you want.

Patch panels!

OK, yeah, this was a mistake.

I'm a software guy, and my experience with physical networking is limited to homes. Before I posted this site, I honestly didn't know what a "patch panel" was.

The subcontractor who pulled all the cat6 cable in the house was someone who does AV systems in houses, not commercial networking. He was chosen by my general contractor, who also obviously specializes in houses. He was a good guy! But in retrospect, given the scope, it might have made sense for me to insist on finding a network contractor that normally does commercial buildings. Anyway, the guy asked if I wanted him to crimp the cables, and I said "uhh well I certainly don't want to do it, so yeah?", and so he did. I don't think he mentioned patch panels, or if he did, I didn't understand what he meant. Then when I moved in, I mounted these Unifi switches, shoved all cables into them, and called it a day.

I now understand that I should have asked the contractor to terminate all the cat6 into "patch panels", that is, rows of RJ45 sockets. I could then use patch cables to connect those to the actual switches. This way, the in-wall cables never move at all, therefore are at minimal risk of being damaged somehow. Also, I could nicely label each one, and easily reconfigure what is connected to what. I wish I'd done that!

That said, in practice I mostly don't have a need to treat any of these cables specially (except that the PoE devices need to be connected to the PoE switch). I'm happy managing my network entirely in software.

By the way, the game machines are not actually connected to the switches in the photo. There is a separate USW EnterpriseXG 24 further down the rack, providing 10G networking to those machines (and allowing them to utilize the full 2G internet bandwidth).

How long are the cables? Do they add latency?

The long DisplayPort and USB cables range from 35 to 100 feet, depending on the specific location. All cables are Monoprice SlimRun, which are fiber optic. The speed of a signal in fiber optic cable is about 2/3 the speed of light, and light travels at about 1 foot per nanosecond, so 100 feet of fiber would add about 150ns of latency. Thats 0.00015ms. Wheras at 144Hz, you see one frame per 7ms. In other words, no, the cable length does not add any meaningful latency. (For what it's worth, signal propagation in copper cables is similarly fast, but high-bandwidth signals in copper tend to degrade more quickly over long distance.)

Why not use thin clients and one beefy machine running a lot of VMs?

Frankly, it would cost more, perform worse, and I'm not even sure how to set it up. Games are resource-intensive. A high-end game will fully utilize the machine's GPU, CPU, and RAM, so there's no real efficiency to be gained by packing more players onto fewer machines. If you wanted to put four players on one machine—if the software even exists to make it work—you'd need a machine with 4x the cores, 4x the RAM, and probably four separate GPUs. To do that you have to dive into "enterprise" hardware that gets extremely expensive. Building four separate machines is actually cheaper, and doesn't require any special virtualization layer.

What home automation tech do you use?

Generally, I try to avoid any cloud-dependent home automation.

For security cameras, we use UniFi Protect. I am very happy with these. They will reliably wake me up if a human approaches my house in the middle of the night, with very few false positives.

I wrote my own baby monitor that simply combines the audio from two UniFi cameras placed near my kids' beds and makes the stream available via a web page. I like it much better than the off-the-shelf monitors we were using before!

I intend to set up Home Assistant eventually, but haven't gotten around to it yet.

Is your electric bill enormous?

Yes, but not because of the computers. The computers are off most of the time; I only turn them on for parties. Most of our energy usage goes to heating and cooling. We have efficient heat pumps and solar panels, but it's still a disturbing amount of energy usage and something I'm trying to debug. Perhaps we have too many windows letting in too much sunlight...

Miscellaneous other questions

Any funny contractor reactions?

We had a subcontractor designing the HVAC system. I told him that there needed to be a dedicated air conditioner for the server rack. He sort of rolled his eyes and said sure.

Later, when I got the design, there was no AC for the server rack. We had a conversation:

Me: Where's the AC for the server rack? We really do need AC in there.

Him: I mean, how many servers do you have?

Me: Well, if all the machines are running at full power playing a high-fidelity game, they could be consuming 15kW of power and turning it all into heat.

Him: (skeptical) That would be by far the largest server rack we've ever seen in a residential setting.

Me: I would expect so, yes.

Tell me more about cat doors.

Our philosophy with the cat doors is that we want to be able to close our bedroom doors while still allowing cats to enter. If we locked our cat out of our bedroom, he would meow all night long, and worse, we wouldn't get cuddles. But, leaving the bedroom door open sacrifices privacy. Solution: Cats get their own doors.

Why don't you have automated litter boxes?

We owned a Cat Genie for a while in Palo Alto, and honestly it was terrible.

If the robot arm failed to scoop the poops (which it often did, especially if they were too soft) then it would wash the litter—poop still inside—and then blow-dry it—poop still inside—and the whole house would smell like cooked poop!

Also the process took 45 minutes and if Garply decided he needed to relieve himself in the meantime he'd do it on the floor.

Also the inside of that thing was horrid to clean out. Imagine opening a complex mechanical device to find someone had spread peanut butter over every component except it's not peanut butter, it's poop.

People have told me that there are other automated litter boxes that are much better than the Cat Genie. Maybe so. But, my take-away is that poop and complex mechanical devices just don't mix. It only takes a couple minutes a week to scoop poops, so I'd rather not risk it again.

Just how often do you play Dance Dance Revolution that it justifies built-in pads?

Pretty often! I try to exercise most days and alternate between DDR, biking, and swimming. Of these, DDR really has the best fun/work and work/time ratios. Jade plays, too. I've been playing for over 20 years but my skill plateaued a while ago at "not quite competitive" levels...

Wait you have four DDR pads? Is that supported?

There are some obscure versions of DDR, and an old unofficial build of Stepmania, that supports four pads. I haven't actually gotten around to trying them yet, but I intend to at some point, and if I end up having to make my own Stepmania fork, so be it. I've been thinking about an algorithm to convert "doubles" step charts into "quads", where one person uses all four pads.

The Daily Front Page 8 of 31
Tuesday, September 8, 2026 The Daily Front No. #260908 — Diagrams, Properly Arranged
article

TALA Is Open-Source

by alixanderwang·▲ 305 points·23 comments·d2lang.com ↗
TALA is a novel autolayout algorithm designed with software architecture diagrams in mind.

Following up on the announcement here, TALA (Terrastruct's AutoLayout Algorithm) is now open-source under the same license as D2 (MPL-2.0).

TALA is a novel autolayout algorithm designed with software architecture diagrams in mind. This means it's primarily an orthogonal layout engine, which more closely matches what you might find on whiteboards, rather than the DAG-based ones that grow in one direction. It blends ideas from different graph-drawing research papers (cited in source code) along with original techniques to achieve aesthetic diagrams. It considers multiple objectives of "aesthetic", including symmetry, median distance, flow, clustering of like nodes, and much more.

I'll keep the text short and lead with examples.

  • The first batch compares diagrams rendered with TALA with the other two layout algorithms D2 comes with -- Dagre and ELK. These are not hand-selected, I just found public d2 files from around GitHub. So for some, you may very well prefer the not-TALA layout.
  • The second batch demonstrates a unique property of TALA, which is that node positions and sizes can be customized, e.g. locking in the coordinates. This lends itself especially well to agentic use cases, where models can draw in 2D space well, but TALA still takes care of routing, which models still struggle with. I had AI generate these.
  • The third batch demonstrates TALA's capability to support a hybrid of some nodes specifying coordinates and some left to the layout engine. You might have a specific shape of a collection of nodes in mind, which you can specify with coordinates, and TALA can take care of the rest. Again, AI generated.

Please also note that TALA is not without tradeoffs.

  • It has randomness in the algorithm. It finds the best layout by using a default of 3 seeds and choosing the one scored the best. Given the same seeds and same input, it'll produce the same diagram. But let's say you just add one more node. The diagram could look completely different. In Dagre and ELK, it looks mostly the same as prior, with the extra node accommodated for. This is sometimes desirable.
  • It doesn't do DAGs as well. I often find myself preferring Dagre or ELK when I want a long flowing graph.
  • It can take longer to run for larger diagrams -- scaling nonlinearly. For a benchmark of TALA's runtime performance compared to others, see https://github.com/d2lang/d2-benchmarks.

TALA comes bundled into D2 v0.9.0, so just install and specify with --layout=tala to try it out! Or head on over to https://play.d2lang.com, which runs 100% client-side. I especially look forward to the improvements that being open-source brings, and can't wait to see what improvements and ideas are submitted by the community.

Special thanks to Gavin Nishizawa for substantial broad contributions across TALA, and Júlio César Batista for his work on hierarchy algorithms and more. It was so fun getting to work on such interesting stuff with you guys.

Batch 1: Comparisons

Fulcro RAD architecture

TALA

Fulcro RAD architecture rendered with TALA

Dagre

Fulcro RAD architecture rendered with Dagre

ELK

Fulcro RAD architecture rendered with ELK

Select a diagram to enlarge it. Each layout is scaled to fit its panel.

Mocha secure-enclave SoC

TALA

Mocha secure-enclave SoC rendered with TALA

Dagre

Mocha secure-enclave SoC rendered with Dagre

ELK

Mocha secure-enclave SoC rendered with ELK

Select a diagram to enlarge it. Each layout is scaled to fit its panel.

Jupyter on AWS EKS

TALA

Jupyter on AWS EKS rendered with TALA

Dagre

Jupyter on AWS EKS rendered with Dagre

ELK

Jupyter on AWS EKS rendered with ELK

Select a diagram to enlarge it. Each layout is scaled to fit its panel.

Lion Reader frontend data flow

TALA

Lion Reader frontend data flow rendered with TALA

Dagre

Lion Reader frontend data flow rendered with Dagre

ELK

Lion Reader frontend data flow rendered with ELK

Select a diagram to enlarge it. Each layout is scaled to fit its panel.

ROSS rotor-dynamics workflow

TALA

ROSS rotor-dynamics workflow rendered with TALA

Dagre

ROSS rotor-dynamics workflow rendered with Dagre

ELK

ROSS rotor-dynamics workflow rendered with ELK

Select a diagram to enlarge it. Each layout is scaled to fit its panel.

Go Queue worker architecture

TALA

Go Queue worker architecture rendered with TALA

Dagre

Go Queue worker architecture rendered with Dagre

ELK

Go Queue worker architecture rendered with ELK

Select a diagram to enlarge it. Each layout is scaled to fit its panel.

Ouroboros Leios simulator

TALA

Ouroboros Leios simulator rendered with TALA

Dagre

Ouroboros Leios simulator rendered with Dagre

ELK

Ouroboros Leios simulator rendered with ELK

Select a diagram to enlarge it. Each layout is scaled to fit its panel.

Batch 2: Custom positioning

Signal House

TALA · positioned with top / left

Signal House

Atlas / Data platform

TALA · positioned with top / left

Atlas / Data platform

Night shift / Mission control

TALA · positioned with top / left

Night shift / Mission control

Friday deploy: the escape room

TALA · positioned with top / left

Friday deploy: the escape room

The Internet is a jellyfish

TALA · positioned with top / left

The Internet is a jellyfish

Orbital coffee logistics

TALA · positioned with top / left

Orbital coffee logistics

Cloud Conservatory

TALA · positioned with top / left

Cloud Conservatory

Velvet Rope

TALA · positioned with top / left

Velvet Rope

Synthwave City

TALA · positioned with top / left

Synthwave City

Batch 3: Partial positioning

The Printing Room

TALA · 4 pinned nodes / 10 automatic nodes

The Printing Room

Pinned: The four CMYK stations share a fixed top coordinate and equally spaced left coordinates so the print sequence retains its mechanical alignment.

Automatic: TALA positions the feeder, camera, registration controller, dryer, prepress and finishing steps; none has top or left.

MULE / Utility Rover

TALA · 4 pinned nodes / 11 automatic nodes

MULE / Utility Rover

Pinned: The four motor assemblies are fixed at the front and rear corners of the chassis rectangle.

Automatic: TALA places every controller, sensor, power and safety node between or around those corners and arranges the unpositioned fleet container.

Sources and rendering details

The seven layout comparisons use public project diagrams from D2's real-world fixtures. Each comparison uses the same D2 source and the same compiler build, changing only the layout engine. Source styles, themes, and explicit grid constraints are preserved. Some fixture icons were already replaced with built-in shapes.

SVGs are scaled independently to fit each panel. Use Open SVG to inspect labels and connections at a larger size.

Source provenance and licenses

Render settings and revisions

The eleven positioning examples are original, fictional compositions rendered with TALA using the same public D2 build. Icon downloads include the full source and local vector assets. The two partial-positioning examples apply top/left only to their listed pinned nodes; all other nodes and every container are automatically placed.

Positioning render settings

The Daily Front Page 9 of 31
Tuesday, September 8, 2026 The Daily Front No. #260908 — Boards by Agent
show hn

Show HN: Copperhead – Cursor for circuit boards

by animeshchouhan·▲ 216 points·87 comments·copperhead.sh ↗
Copperhead is an open-source AI engineering platform that helps hardware teams design, verify and ship circuit boards.

Copperhead is an open-source AI engineering platform that helps hardware teams design, verify and ship circuit boards.

Start building Read the docs

$ npm i -g copperhead
copperhead v0.8.1

claude via flag · kicad-cli 9.0.4

~/hardware/open-key

❯ $ add a USB-C power input to the key

▸ read_file docs/POWER.md (48 lines)

▸ list_nets 41 nets, no VBUS, no CC1/CC2

▸ validate_change add-usb-c-power valid, edit tools unlocked

✓ propose: add-usb-c-power validated, edit tools unlocked

▸ edit_file hardware/open-key.kicad_sch: USB-C receptacle, 2x 5.1k CC pulldowns

▸ edit_file docs/POWER.md: VBUS budget, 5 V at 1.5 A

✓ edit: 2 files, 5 anchored regions

▸ run_erc clean, 0 violations

▸ check_drift docs match schematic

✓ verify: ERC clean, no drift

▸ record_decision DECISIONS.md +1 · CHANGELOG.md +1

✓ remember: decision + changelog recorded

turn 1/40 · 900 in / 120 out · drafting the power proposal

turn 2/40 · 4.2k in / 800 out · placing the receptacle

turn 3/40 · 9.1k in / 1.5k out · running ERC

turn 4/40 · 11.8k in / 1.9k out · writing the decision log

done · verified erc · committed 3f2c9a1 · 17s · 12.3k tokens

A copperhead-designed board: the Open Telegraph key, with ESP32-S3, USB-C and the copper switch, rendered from its KiCad files.

Proven on real copper.

Open Telegraph: pocket-size ESP32-S3 Morse key, built end to end with this workflow. Every decision, every check, every file is public.

Read the build story · Browse the repo

Eight stages, start to finish.

You write the brief. copperhead runs the stages in order, each one its own agent run. A stage has to leave its artifact on disk before the next starts. If a gate fails, the run stops there.

Input: brief.md

What you want, written in your own words.

  1. Spec

    docs/SPEC.md

    Gate: a budgets section, actually filled in.

  2. Architecture

    docs/SUBSYSTEMS.md

    Gate: real reasoning under every subsystem.

  3. Parts

    docs/BOM.md

    Gate: part numbers chosen against datasheets.

  4. Schematic

    design.kicad_sch

    Gate: symbols placed, docs agree, ERC clean.

  5. Layout

    design.kicad_pcb

    Gate: footprints on the board, DRC clean.

  6. Outputs

    gerbers / drill / STEP

    Gate: gerbers actually on disk.

  7. Firmware

    firmware/ / pins.h

    Gate: sources on disk, pins from the pinout.

  8. Dev plan

    docs/DEVPLAN.md

    Gate: a written bring-up and test plan.

Output: a design package

Gerbers, firmware, docs. One commit per stage.

  • Every hop is a gate: no real work on disk, no commit.
  • Copper border: verified by KiCad itself, ERC and DRC clean.

Read the long version, one stage at a time

Pay to not run it yourself.

copperhead is open core. The CLI is free and always will be, and it does the real work on its own. You pay to host it, to work as a team or to give an auditor what they ask for.

CLI

Free

Apache-2.0, forever.

Design on your own, on your own machine.

Includes:

  • The full agent: do, create, init, check, watch
  • Bring your own Claude or GPT-5 key
  • Runs on your KiCad files, in your own git repo
  • Plain markdown, JSON and KiCad output. No lock-in.
  • Runs locally, never metered
  • Community support

Cloud

$49 per user / month

Free for open hardware repos.

Hosted runs on private repos, with a web viewer.

Everything in CLI, plus:

  • Hosted runs, no local setup, from anywhere
  • Private repositories
  • Web viewer: chat, live schematic and board render, ERC/DRC status
  • Run history and shareable check reports
  • One-click gerber, DXF/STEP, render and BOM export
  • BYO key, or managed inference with 200 credits a seat

Team

$49 per user / month
+$199 / mo platform

Cloud, with governance and CI for a whole team.

Everything in Cloud, plus:

  • CI bot: run check as a required PR status check
  • Shared, versioned constraint libraries across the org
  • SSO / SAML, role management, central billing
  • Managed credits pooled across the team
  • Shared run history

Enterprise

Custom annual

Self-hosted, built for procurement, one price a year.

Everything in Team, plus:

  • Self-hosted or VPC, so design data never leaves your network
  • Altium support beyond KiCad
  • RBAC, security review, dedicated support and an SLA
  • An audit trail your compliance team can hand over
  • Flat annual license, no per-run metering

Questions, before you ask them.

What is copperhead?

An open source AI agent that designs, documents and verifies printed circuit boards. You describe a change or hand it a product brief, and it edits your real KiCad files, updates every document that references them, and runs KiCad's own checks until they pass. Longer introduction here.

What problem does it actually solve?

Drift. A hardware design spreads one decision across a schematic, a bill of materials, a power budget and several documents, and nothing breaks when they fall out of sync. The inconsistency is found at bring-up, and a respin costs 5,000 to 50,000 dollars and six to eight weeks. The full argument is here.

What do I need installed?

Node 20 or newer, KiCad with kicad-cli on your path and a model API key of your own. Then npm i -g copperhead.

Does it work on a design that already exists?

That is the main case. Point it at a KiCad repository, run copperhead init and start asking for changes. It can also run the full pipeline from a written brief with copperhead create, but iterating on real designs is what it is best at.

What will it refuse to do?

It refuses to run on a dirty git tree, refuses to edit any design file before a validated change proposal exists and refuses changes that break a budget or constraint you have documented, citing the line it would violate. It also never invents a part number it cannot justify from a datasheet.

Will it rewrite my whole schematic?

No. Edits are surgical changes to the KiCad s-expression source, so your diffs stay small and reviewable and untouched parts of the file stay byte-identical. A tool that regenerates the file to move one net has made its own work impossible to review.

9 more questions

Never ship a board your docs no longer describe.

copperhead is open source and free to use. Install once, describe the board you want in a short brief.md and run create.

$ npm i -g copperhead
$ export ANTHROPIC_API_KEY=<api-key>
$ copperhead create --brief brief.md

brief.md

# Pocket Bluetooth speaker

A palm-size Bluetooth speaker that plays for a day on one charge.

- ESP32, Bluetooth audio (A2DP sink)
- 3 W class-D amp into a 4 Ω driver
- Li-Po cell, USB-C charging
- Standby current budget: 100 µA

Circuit boards, designed by an agent.

Open source, verified against KiCad and built by people who still solder.

The Daily Front Page 10 of 31
Tuesday, September 8, 2026 The Daily Front No. #260908 — Models on a Diet
article

Benchmarking Qwen3.8 27B quantizations: 4-bit holds up, 1-bit collapses

by stared·▲ 226 points·110 comments·quesma.com ↗
Compression eventually hits a cliff.

How much GPU RAM do you actually need to run Qwen3.8 27B without sacrificing quality?

The full BF16 model weighs 55 GB, putting it beyond most consumer hardware. Yet the 17 GB Q4_K_M matches the full model on a popular agentic coding benchmark, Terminal-Bench 2.1. It fits on a 24 GB card such as RTX 4090, still leaving room for about 64k tokens of context.

Compression eventually hits a cliff. At 1 bit, the model performs around random chance on GPQA Diamond, and longer reasoning makes it worse.

Background

Qwen3.8 27B GGUF quantizations available from Unsloth on Hugging Face

Qwen3.8 27B GGUF quantizations available from Unsloth on Hugging Face. So much to choose from! I will check 8-bit Q8_0 (29 GB), 4-bit Q4_K_M (17 GB), 2-bit UD-Q2_K_XL (10.7 GB), and the smallest one possible, 1-bit UD-IQ1_S (6.2 GB).

Previously, I investigated the Qwen3.6 27B model, which was good at generating SVG pelicans even at 12GB, and maintained most of its knowledge up to 16GB. At the same time, in Reddit threads, many complain that all quantizations, even the 8-bit ones, give worse results - with people asking why your local LLM feels dumber than it is. Are these complaints grounded?

Measuring token prediction differences (KL-divergence, top-1 predictions) is easy, but it does not tell us whether the model gets worse at solving tasks. Some noise might be irrelevant for solving tasks, as (say) a quantized model generates an answer of precisely the same quality, paraphrased a bit. In other cases, a single different token might be a logical error, or even abruptly end the output.

So, I focus on directly measuring results on popular benchmarks - GPQA Diamond, instruction-following IFBench, programming Terminal-Bench 2.1. First, to replicate official results of the full model BF16, and then to see how quantization affects results.

I burned around $3,000 on Modal GPUs when I ran models with llama.cpp using a build from 16 August 2026 as earlier builds do not work for this model. I could have run it on my own laptop, in principle, but (unlike pelican-generation), these are time-consuming benchmarks.

Note that I use F16 KV-cache regardless of model quantization, weighing around 2.3 GB per 32k tokens.

I used Unsloth quantizations: v2 for the 2-, 4-, and 8-bit models, and v3 for the 1-bit models. Unsloth replaced the v2 files on 19 August 2026, so the exact files used for most tests are no longer available.

In short, if you go with a 4-bit quantization Q4_K_M (17GB), you won’t notice a difference on these benchmarks. At the same time, the effort setting matters a lot (note that the default is xhigh) - and it is a tricky choice, as it can overthink.

One-shot tests

The easiest ones are one-shot tests: in this case, graduate-level science GPQA Diamond and instruction-following IFBench. I run each at three reasoning efforts: low, medium, and the default xhigh.

GPQA Diamond

First and foremost, I was happy I replicated the official results. Running benchmarks is hard; there are many hidden settings or assumptions that can change the results drastically. Here, on the first go, results were as reported by Qwen.

Second, besides noise (bars are Wilson 95% confidence intervals, very conservative for run-to-run noise), there is little difference down to 4-bit; only the 2-bit scores a bit lower.

At the same time, thinking level changed the score dractically. The best results, for xhigh, needed around 8k reasoning tokens.

IFBench

Here, to my great surprise, there is no change between models, down to a decent 2-bit one, weighing less than 11 GB. Yet, context is even lower, around 4k tokens.

Agentic coding and Terminal-Bench 2.1

How does it work for programming? Terminal-Bench 2.1 is a standard agentic benchmark, with 89 tasks. Here I use 3h timeout, xhigh effort. I reserve 98k context.

Not only does my measurement of BF16 replicate the stated result, but, to my surprise, Q4_K_M does as well. I accidentally skipped running Q8_0; yet, in this case, I can safely interpolate between 4-bit and the full model’s values. Running it would be both costly and unnecessary (and would exceed an informal blog post’s budget). Only at 2-bit UD-Q2_K_XL things break a bit. A noticeable fall, but still the level of Opus 4.7 or Gemini 3.1 Pro. Again, far from frontier, but also - far from useless.

Results are one thing, but what about the process? Do smaller models need more turns, tokens or time to get the result?

On the same solved tasks, UD-Q2_K_XL takes as many turns as BF16 but writes about a quarter more tokens. The number of turns stay roughly the same.

The 1-bit cliff

Quality drops off a cliff at 1-bit. As with knowledge, quantization damage is nonlinear: first there is no measurable change, then a small decline, and finally a collapse.

While 2-bit quantizations work to some extent, even the best 1-bit model is useless for these benchmarks:

As you may see, the scores are around the random guessing level, with the smallest model being below that threshold. And longer reasoning makes it worse: at xhigh, scores drop below low, as the model more often reasons until the token budget runs out and returns an empty answer. Sure, Unsloth boasts that:

We also made some smaller UD-1bit quants with UD-IQ1_S being 6.2GB (without MTP) which retain around 72% top-1% accuracy yet being 89% smaller.

But in this case, these remaining 28% matter a lot. And this matches another user’s experience, vide Qwen3.8 27b 1bit brain damage quant on r/LocalLLaMA.

Costs

Running these benchmarks isn’t cheap. Running benchmarks via API is costly, as I know from my previous benchmarks. Running on rented GPU is much costlier.

I used Modal, as it is easy to run it from the CLI, including from agents. Other setups may have different pricing. Obviously, this calculation changes if you have your own devices.

It takes some testing to find the optimal way to run models. Usually, instead of using Multi-Token Prediction (MTP), which works well for a single stream, I use a few parallel streams. The key constraint is whether the GPU has enough memory for both the model and the required KV caches.

I used NVIDIA L40S (the same Ada Lovelace chip as the RTX 4090 but twice as much memory: 48 GB), H100 (80 GB) and H200 (141 GB). I would like to share costs to give you a ballpark estimate if you want to run benchmarks yourself.

For comparison, DeepSeek V4 Flash 0731, a 284B model, costs around $0.1/Mtok for output from the cheapest providers on OpenRouter. I am not sure how much of this difference comes from the efficiency of running models at scale, pricing strategy, or popularity.

Conclusion

If you run experiments locally, usually pick the best model that fits in your GPU memory together with the required context. For most tasks Unsloth’s Q4_K_M should be good enough, without any noticeable difference; for some simpler tasks UD-Q2_K_XL should be more than fine. Since people report that KV-caches are more susceptible to quantization, I may test it as well.

But in general, I believe that quantization should be embraced, rather than feared.

And what is your experience? Join the discussion on r/LocalLLaMA, Hacker News, or LinkedIn.

The Daily Front Page 11 of 31
Tuesday, September 8, 2026 The Daily Front No. #260908 — Europe’s Front Door
article

Among European Companies That Use a CDN, Nearly 9 in 10 Use Cloudflare

by adulion·▲ 377 points·359 comments·ciphercue.com ↗
That means the CDN is a shared dependency.

A content delivery network sits in front of a website. Requests hit the CDN's edge servers first, which serve cached content, terminate TLS, filter traffic, and pass the rest back to the origin. That means the CDN is a shared dependency: if it goes down, every site behind it goes down together, whether or not those sites have anything else in common.

I went through CipherCue's tracked European companies to see which CDN fronts their websites. Of the 44,143 companies where we detected a CDN at all, 39,547 are behind Cloudflare. That is 89.6%.

89.6% of 44,143 European companies with a detected CDN sit behind Cloudflare

For a reference point, W3Techs put Cloudflare at 84.1% of sites where the reverse proxy provider can be identified, as of 28 July 2026 (w3techs.com). Our 89.6% is measured the same way, as a share of the identified set rather than of all sites, so the two are close; the European cut here runs a little higher.

I only counted companies that actually run a CDN. If a company serves its site straight from its own origin, it is not in these numbers at all. So this is not "Cloudflare vs the entire web"; it is who the CDN users went with.

First place and everyone else

Amazon is second at 3,112, but that number needs a caveat. We detect Amazon through CloudFront, and Amazon is also a general cloud host, so some of these are companies whose origin happens to sit on AWS rather than companies that deliberately chose CloudFront as their front door. Fastly does not have that ambiguity. It is a pure CDN, and it comes third at 1,299, which is roughly one company for every thirty behind Cloudflare. Akamai, the oldest name in the category, is fourth at 396.

The raw counts add up to more than 44,143 because a company can sit behind more than one vendor and is counted under each (a Fastly front door with an AWS origin, say). That double counting pads the smaller vendors' totals, not Cloudflare's share, which is measured against the whole CDN-using set.

Per country

Rolling eight countries into one figure hides the differences between them, so here is the same measurement per country. Cloudflare is the majority front door in all of them, but the share ranges from about four in five in Spain and Ireland to about nineteen in twenty in the Netherlands.

Country Companies with a CDN Behind Cloudflare Cloudflare share
Netherlands 7,939 7,587 95.6%
United Kingdom 17,007 15,846 93.2%
Poland 2,896 2,682 92.6%
France 4,008 3,456 86.2%
Italy 3,661 3,126 85.4%
Germany 5,715 4,650 81.4%
Spain 2,001 1,576 78.8%
Ireland 792 624 78.8%

The UK has the largest absolute count, 15,846 CDN-using companies behind Cloudflare in one market. Germany is the lowest of the big markets at 81.4%, which is still four in five.

What the concentration costs

Part of why a company buys a CDN is resilience, and one advantage of independent suppliers is that they fail independently. That advantage disappears when most of a market is behind the same supplier. An incident there is no longer one company's outage; it is most of the market's outage, on the same afternoon.

Cloudflare has published postmortems for several incidents in the last fifteen months. Three were global and hit customers:

  • 18 November 2025, roughly 11:20 to 17:06 UTC. A database permissions change caused a Bot Management feature file to fill with duplicate entries and double in size, past a limit the proxy enforced, which crashed core CDN and security serving. Cloudflare's writeup states the outage was "not caused, directly or indirectly, by a cyber attack or malicious activity of any kind" (blog.cloudflare.com).
  • 5 December 2025. A configuration change applied while mitigating an industry-wide React Server Components vulnerability caused a global disruption (Cloudflare outage postmortems).
  • 20 February 2026, from 17:48 UTC, lasting 6 hours 7 minutes. A cleanup automation task read a buggy API response as an instruction to withdraw BYOIP prefixes, and 25% of them were pulled from the internet via BGP. Cloudflare again states it was "not caused, directly or indirectly, by a cyberattack or malicious activity of any kind" (blog.cloudflare.com).

None of the three was an attack. A doubled config file, a change made while patching someone else's vulnerability, a cleanup job that deleted too much: routine internal work that reached the edge and took a lot of sites down with it. The companies in the table did not need anything in common to go offline together on 18 November. Sharing a front door was enough.

This is not about Cloudflare being sloppy. Fastly and Akamai ship bugs like this too; theirs just take fewer sites down because fewer sites are behind them. That is the whole point. When one provider fronts most of a market, its mistakes stop being its own problem and become everyone's, at the same time.

The CDN is not where the data lives

This says nothing about where a company keeps its data. The CDN is just the front door; the origin server, the database, and whoever processes the data sit behind it, and for a GDPR or data-residency question that back end is what counts. Measure that instead and Europe looks better: people who dig into API subdomains rather than the front door usually turn up far more OVH and Hetzner than I see here. I stuck to the front door because it is the part that goes down for everyone at once, which is what happened in the outages above.

Check your own front door

You can see which CDN, if any, fronts a domain from its response headers:

curl -sI https://example.com | grep -i 'server\|cf-ray\|x-served-by\|x-amz-cf-id'

A cf-ray header or server: cloudflare indicates Cloudflare. x-served-by with a Fastly cache node indicates Fastly. x-amz-cf-id indicates Amazon CloudFront. No such header, and a server value naming your own web server or origin host, usually means no CDN in front.

Method note

Source and cohort: CipherCue's own observations of European company websites (HTTP response fingerprinting and DNS), not a third-party dataset. The cohort is the companies in our tracked entity set with a country attribution in one of Germany, the United Kingdom, the Netherlands, Poland, France, Italy, Spain, or Ireland, and at least one CDN component detected. That gives 44,143 companies, using the latest observation per company as of 2026-09-07. This is a company count within our dataset, not a representative sample of every company in these countries, and not a claim about companies that use no CDN at all. The cohort also skews toward small and mid-sized companies rather than large enterprises, and Cloudflare's free tier is strongest in exactly that segment, so read the figure as concentration among CDN-using companies of this mix, not as a statement about large enterprises specifically.

What "detected a CDN" means: we classify a component as a CDN when the response headers or serving fingerprint match a known CDN vendor's pattern (for example a cf-ray header for Cloudflare, a Fastly cache-node x-served-by, a CloudFront x-amz-cf-id). Companies that serve directly from an origin with no recognised CDN fingerprint are excluded from the denominator. A misconfigured or unusual setup that hides these headers would be undercounted.

Double counting: a company can be detected behind more than one CDN vendor and is then counted under each. The Cloudflare share (89.6%) is Cloudflare-detected companies divided by all companies with any CDN detected, so a company behind both Cloudflare and Fastly counts once in the numerator and once in the denominator. Raw per-vendor counts therefore sum to more than 44,143.

The Amazon caveat: Amazon is detected via CloudFront, but Amazon is also a general cloud host. Some Amazon detections are front-door CloudFront choices and some are AWS-origin serving; we do not separate the two here, which is why Fastly (an unambiguous pure CDN) is called out as the nearest clean comparison.

Outage details are from Cloudflare's own published postmortems, cited inline. Times and durations are as stated by Cloudflare. We link the primary source for each rather than restating figures from secondary coverage.

Country grouping is our choice and is stated explicitly so the per-country table can be read on its own; the headline figure is the union of these eight markets, not a single country.

Finding this in your own market

CipherCue tracks CDN, hosting, email, and DNS infrastructure per company, filterable by vendor, country, and sector. If you sell a European alternative in one of these categories, the directory shows which companies in your market currently run which incumbent: the companies we detect behind Cloudflare are listed there, and the EU vendor directory covers the alternatives by category.

The Daily Front Page 12 of 31
Tuesday, September 8, 2026 The Daily Front No. #260908 — Cold-War Engineering
article

The Helicopter with Radioactive Blades

by zdw·▲ 162 points·49 comments·hackaday.com ↗
The problem was detecting problems with the blades before they grew into a catastrophic failure.

Sometimes you find gold in unexpected places. In this case, it’s from a video about the CH-53 helicopter. The CH-53 Sea Stallion first entered service in 1966. It’s a huge helicopter, capable of lifting immense amounts of weight. All that weight hangs from 6 (or 7 on some models) rotor blades. The problem was detecting problems with the blades before they grew into a catastrophic failure.

Much like a fixed-wing aircraft wing, the CH-53 rotor blades are designed with a spar as the main structural member. On early designs, the spar was extruded aluminum. The CH-53D and other variants moved to cold-formed titanium.

All of these blades shared the same problem – detecting tiny cracks in the metal before they grow into blade failure. Detecting cracks turned out to be rather easy. The blades are sealed and pressurized with nitrogen gas. Cracks allow the gas to leak out. A barber pole pressure indicator on the blades allows mechanics to tell at a glance if everything is ok.

That’s all fine and good on the ground, but if a blade begins cracking in the air, the pilots want to know about it immediately. What was needed was something akin to the tire pressure monitoring system on modern cars. A pressure sensor that would turn on a light in the cockpit. Modern automotive TPMS systems use wireless sensors inside the tire. In the early 60’s though, electronic components were not reliable enough to survive rotating on a helicopter blade. Wiring the sensors through the spinning rotor head using slip rings would be incredibly complex. Expecting wireless electronic components to survive life in a rotorhead would be even more problematic.

The solution turned out to be simple: Throw a little ionizing radiation into the mix. Called IBIS, short for Inflight Blade Monitoring System keeps the CH-53 safe. The magic happens in the same pressure indicator we mentioned earlier. Rather than just a shift from a green to a barber pole indication, the IBIS module also exposes a tiny amount of radioactive Strontium-90. The Strontium is a beta emitter, meaning its radiation can be detected by a Geiger counter inside the helicopter. It’s also relatively safe for humans as long as they don’t eat or breathe it in. No batteries, no electronics on the rotor blade. It’s a simple mechanical solution to a difficult problem.

The solution is so elegant it’s still being used today. The newest versions of the CH-53 of fiber optics to detect faults in the newest all-composite blades. The older variants? They’re still going with the nuclear option.

The Daily Front Page 13 of 31
Tuesday, September 8, 2026 The Daily Front No. #260908 — Roadside America
article

John Margolies' photographs of roadside America

by duck·▲ 124 points·42 comments·publicdomainreview.org ↗
The culture of the American road has been much celebrated — and much criticized.

John Margolies Roadside America

Big Fish Supper Club, Route 2, Bena, Minnesota; 1980.

The culture of the American road has been much celebrated — and much criticized. Lawrence Ferlinghetti saw the rise of the automobile and the construction of the interstate system (which began in the 1950s) as a new form of punishment inflicted on the populace. Driving in their cars, “strung-out citizens” were now

plagued by legionnaires
                                false windmills and demented roosters…

      on freeways fifty lanes wide
                                                        on a concrete continent
                                                                spaced with bland billboards
                                            illustrating imbecile illusions of happiness

The architectural critic and photographer John Margolies (1940–2016), on the other hand, saw there could also be home-made beauty in the buildings and signs locals built on the American roadside. For almost forty years, he documented the most remarkable examples he found, publishing some of his discoveries in books and consigning the rest to an archive, which has now been purchased by the Library of Congress who, in a wonderfully gracious move, have lifted all copyright restrictions on the photographs (though art works shown in some photographs may still be under copyright).

Readers of Ferlinghetti would not be surprised to see Margolies’ archive offer up no end of "false windmills" and "demented roosters".

John Margolies Roadside America

Log Cabin Motel Office, San Leandro, California; 1978.

John Margolies Roadside America

Rooster statue at Shirt World, Pigeon Forge, Tennessee; 1984.

But the billboards he preferred to photograph illustrated relatively humble "illusions of happiness".

John Margolies Roadside America

Club Cafe sign near Santa Rosa, New Mexico; 1987.

John Margolies Roadside America

Chicken Cowboy billboard wreck, Elko, Nevada; 1991.

He also found no shortage of signs graced with attention-grabbing, groan-inducing puns.

John Margolies Roadside America

Billboard, near Dillon, South Carolina; 1986.

John Margolies Roadside America

Leaning Tower of Pizza, Quincy, Massachusetts; 1984.

John Margolies Roadside America

Octopus Car Wash, Minneapolis, Minnesota; 1981.

Probably one of his favorite roadside phenomena to document, however, were novelty buildings. He especially liked structures that mimicked their own shape or function, whether this was as witty as a car wash shaped like a whale or as uncomplicated as a coffeehouse shaped like a coffeepot.

John Margolies Roadside America

The Whale Car Wash, Oklahoma City, Oklahoma; 1979.

John Margolies Roadside America

Bob’s Java Jive, Tacoma, Washington; 1979.

Almost all of Margolies’ work was done in the interest of preserving images of what would otherwise be lost to time. Even his first book, published in 1981, was elegiacally called The End of the Road: Vanishing Highway Architecture in America. From the start, Margolies knew the quirky motels, miniature golf courses, diners, billboards, and gas stations were being endangered by franchising and changing fashions — not to mention changing patterns of automobile traffic. (For decades now, most drivers have, of course, opted for the high speed-limits of superhighways and the convenience of service areas, leaving the old local highways in the lurch.)

Today, this collection of Marogolies’ photographs offers an invaluable tour of the diverse vernacular architecture and signage of North America. Some of these wonders remain, while others have gone the way of the dinosaur — which, as it happens, remains one of the American roadside’s most frequent denizens.

John Margolies Roadside America

Harold’s Auto Center, Spring Hill, Florida; 1979.

John Margolies Roadside America

Dinosaur Village RV Mobile Home Park, Jensen, Utah; 1991.

You can view all 11,710 of the colour slides in hi-res digitisations on the Library of Congress, and also a slightly lower-res but easy-to-browse selection of 1,555 on Flickr: The Commons.

John Margolies Roadside America

World's largest buffalo (46' long, 26' high, 60 tons), Jamestown, North Dakota; 1990.

John Margolies Roadside America

Peach water tower, Frontage Road, Gaffney, South Carolina; 1988.

John Margolies Roadside America

Hat n' Boots gas station, Route 99, Seattle, Washington; 1980.

John Margolies Roadside America

St. Joseph Auto and Furniture Fabric slip cover sign, Saint Joseph, Missouri; 1988.

John Margolies Roadside America

The Nat Ballroom, 6th & McMaster's, Amarillo, Texas; 1977.

John Margolies Roadside America

Rawhide City billboard, I-94, Mandan, North Dakota; 1980.

John Margolies Roadside America

Indian Trading Post, Route 66, Elk City, Oklahoma; 1982.

John Margolies Roadside America

Wall Drug final billboard, Wall, South Dakota; 1987.

John Margolies Roadside America

Terrace Drive-In Theater, Terrace Way, Bakersfield, California; 1987.

John Margolies Roadside America

Atomic Signs sign, Route 550, Farmington, New Mexico; 1980.

John Margolies Roadside America

Budget Lodge Motel sign, Route 1, Port of Palm Beach, Palm Beach, Florida; 1990.

John Margolies Roadside America

The Donut Hole, straight-on view, no cars, Amar Road, La Puente, California; 1991.

John Margolies Roadside America

Dino skeleton view 1, Flintstone's Bedrock City, Rts. 64 and 180, Valle, Arizona; 1987.

John Margolies Roadside America

Muffler Man II, Meineke Discount Mufflers, Charleston Avenue and D Street, West Columbia, South Carolina; 1988.

John Margolies Roadside America

Aztec Motel, diagonal view 2, Route 66, Albuquerque, New Mexico; 2003.

John Margolies Roadside America

Wigwam Village #6, Route 66, Holbrook, Arizona; 1979.

John Margolies Roadside America

Ko-Ko-Mo Dine In Your Car sign, Routes 79 & 80, Bossier City, Louisiana; 1979.

John Margolies Roadside America

Hoot Owl Cafe, horizontal view, 8711 Long Beach Boulevard, Southgate, Los Angeles, California; 1977.

John Margolies Roadside America

Roadside flamingo statue, Frog City, Route 41, Florida; 1980.

John Margolies Roadside America

Serra Motel sign, Old Route 101, Gilroy, California; 1991.

John Margolies Roadside America

Thunderbeast Park, platybelodon's eye, Route 97, Chiloquin, Oregon; 1987.

The Daily Front Page 14 of 31
Tuesday, September 8, 2026 The Daily Front No. #260908 — A Little Soil
article

Getting your hands dirty is good for you

by HatchedLake721·▲ 222 points·185 comments·bbc.com ↗
Putting our hands in soil and touching grass, trees or moss provides a wide range of health benefits.

BBC/ Serenity Strull An illustration shows two cupped hands holding a small patch of soil covered with grass, set against a blue backdrop (Credit: BBC/ Serenity Strull)

Putting our hands in soil and touching grass, trees or moss provides a wide range of health benefits, for our skin and wider immune system.

I take a long walk in my nearest park every day, almost religiously. There are plenty of different tree species to look at, an everchanging cast of fragrant flowers in spring, and enthusiastic birdsong fills the air. But the sensory experience usually stops there. However mucky my shoes get, I usually come home with clean hands.

Research, though, suggests there's something to be said for reaching out to touch the trees, shrubs and soil I usually walk past. Mira Grönroos, an ecologist and nature-health researcher at the Finnish Environment Institute, asked people to rub their hands with soil, peat and moss for 20 seconds as part of an experiment. She found that the abundance and variety of microbes on their skin increased, even after they rinsed their hands with tap water. "To the naked eye [their hands] looked clean," she says, "but the change was so huge".

The idea of more bacteria on our skin might not sound appealing. But when there's a rich mix of microbes living alongside each other, they keep each other in check and defend us against more harmful, pathogenic bacteria. A growing body of research shows that contact with natural materials, like soil and plants, is an easy way to boost our skin's microbial diversity, and that this not only benefits our skin, but also our immune system.

Twenty seconds isn't long – I'd happily pause on my walk to touch some tree bark or sit down in the grass. But more would be much better, Grönroos says, as in her experiment "the skin microbiota very quickly returned to a similar state as it was before". The results suggest that prolonged or frequent skin contact with nature could significantly change what's happening on our skin and in our bodies.

Playing in the mud

Mucky hands are the norm for children at Aurinkolinna daycare centre in central Finland. They spend at least a couple of hours every day playing outside in a wooded playground, racing between bushes and trees, climbing ropes, making sandcastles, splashing in puddles and grabbing fistfuls of dirt and woodchips.

By the end of each day they are covered head to toe in mud. Scientists are now studying how this could be good for their health. By analysing microbes on the children's skin, saliva and faeces, alongside blood and hair samples, researchers are learning what happens to their immune systems over a three-year period when nature is brought into the yards where they play every day.

Aurinkolinna is one of 43 daycare centres in Finland taking part in the Vahvistu research project, which builds on a recent two-year study led by environmental scientist Marja Roslund.

In 2016-17, she brought the forest floor into a daycare yard – complete with peat and planting boxes – and measured the impact on young children's skin, saliva and gut bacteria after they had played there regularly for 28 days, and again after a year. Compared to children who played in a paved yard, those on the forest floor had greater microbial diversity on their skin a year later.

This diversity is important for the health of our skin. Increasing the diversity of bacteria on the skin can prevent disease-causing bacteria from becoming dominant, keeping the microbial community in balance, Roslund says. "Disease is like a dictator, and healthy skin is like a democracy."

Roslund's research found fewer potential pathogens on children's skin when they had played on the forest floor, likely because the number and diversity of other microbes crowds out these "bad" bacteria. "Because there is so much of everything else, there is not so much space for the pathogens," she says.

Healthy skin, healthy body

Taking care of our skin also has benefits that run much deeper. Skin health can impact hormonal activity, cardiovascular disease, and even cognitive function. In particular, research is starting to show that exposing our skin to diverse bacteria can impact our immune system.

Roslund asked a group of young children to play in a sandbox enriched with microbially diverse soil, while another group played in a sandbox that looked the same but contained less microbial diversity. As we might expect, the diversity of certain bacteria on children's skin increased when they played in the microbially rich sandbox. But there were also changes to the immune markers, called cytokines, in these children's blood samples that support immunoregulation – the immune system's ability to fight off threats without becoming overactive.

BBC/ Serenity Strull Contact with soil and plants not only benefits our skin, but also our immune system (Credit: BBC/ Serenity Strull)

Contact with soil and plants not only benefits our skin, but also our immune system (Credit: BBC/ Serenity Strull)

It is these and similar findings that inspired the Vahvistu project. At Aurinkolinna, government funding has helped transform the yard so it now includes woodchips instead of gravel, a bigger sanded area, and wooden pathways, alongside the birch and pine trees that were already there.

The research behind the Vahvistu project supports what is known as the biodiversity hypothesis – the idea that air, plants and soil rich in diverse microbes have an important influence on the human microbiome, and therefore our health. It has long been suggested that the loss of biodiversity in our environment could be behind the current rise in immune-related conditions, and Roslund's research may help to explain how and why that is.

She found that one family of bacteria, namely gammaproteobacteria, plays a particularly important role. Gammaproteobacterial diversity increased on children's skin when they played in the yard with the forest floor, and this increase was associated with more regulatory T cells – white blood cells that help control and balance the immune system and are important for preventing autoimmune diseases, she says.

For now, it's not clear exactly what this means for immune-mediated diseases like eczema, asthma or allergies. "We cannot yet say that when you have contact with nature, you will not get asthma, for example. But it seems that it might reduce the possibility of getting those immune system disorders," Grönroos says.

Green fingers

These health benefits aren't only seen in children. In one study, adults were asked to grow vegetables indoors over winter, some using microbially diverse soil and others using similar-looking but microbially poor soil. They tended to the plants every day for a month and ate their crops. Those using the microbially-diverse soil saw an increase in the diversity of bacteria on their skin, and on the number of anti-inflammatory cytokines in their blood – while the other group didn't.

When I go outdoors with my children, if they want to pick some cones or sticks, I don't say 'no'. I let them bring them home and play there – Mira Grönroos

This kind of gardening "is a good example of an easily accessible way of nature contact, and it was quite surprising that this kind of exposure really showed changes in the immune system functioning", Grönroos says.

In another study, researchers installed "green walls" made of living plants in Finnish offices and measured the effect on workers after two weeks. They found that the diversity of microbes on their skin increased, and levels of pro-inflammatory cytokines in their blood went down. They also saw an increase in Lactobacillus microbes on their skin, which is important for skin health.

Grönroos puts her findings into practice with her own family. "When I go outdoors with my children, if they want to pick some cones or sticks, I don't say 'no' to them. I let them bring them home and play there."

We should start applying this thinking to our cities, Grönroos adds. "[Cities] should be built so that citizens are able to get into contact with nature. It would be good if some of that contact was unintentional, so people wouldn't need to decide to go there, they would come into contact with nature on their way to work or school," she says.

There's more for science to discover about exactly how contact with nature impacts the immune system, but for Grönroos, the findings so far are enough to make the grubby hands and muddy carpet worth it. "It's low risk," she says, with "lots of potential benefits without any obvious major harms."

These findings are, perhaps, a reminder to us all to reach out and touch that mossy tree, to sit down in the grass – to not be afraid of a bit of mud.

The Daily Front Page 15 of 31
Tuesday, September 8, 2026 The Daily Front No. #260908 — Saints in Translation
article

The two Christian saints who are the Buddha

by surprisetalk·▲ 226 points·173 comments·signoregalilei.com ↗
Let me tell you the story of saints Barlaam and Josaphat of India.

Let me tell you the story of saints Barlaam and Josaphat of India. A king named Abenner, who persecutes Christians, hears from his court astrologers that his son will attain great power and rule over India. But one astrologer predicts that the prince will instead become a great holy man – a Christian. Alarmed, Abenner has the young prince Josaphat isolated from external contact and sees that Josaphat is pampered, with servants attending to his every need.

But Josaphat becomes curious about the outside world, and eventually is allowed to leave the palace. On these trips, he encounters men ravaged by illness, old age, and death. He is in the middle of a crisis of faith when he meets a Christian hermit named Barlaam, converts, and continues to profess the Christian faith despite his father’s best efforts.

Eventually, Abenner chooses to share the throne with Josaphat. Later Abenner himself converts, and after many years as a wise ruler, Josaphat retreats to the desert to meet his old teacher and live out his days as a monk.1

This might sound familiar to students of world religions: it’s a close parallel of the story of the Gautama Buddha, the founder of the Buddhist tradition. This is no coincidence.

The story of Barlaam and Josaphat first shows up in the Christian tradition in the 10th century, in Georgia as the Balavariani. It seems to have been derived from an Arabic version, the Book of Bilawhar and Budhasaf. The name “Budhasaf” in turn comes from the Persian version “Bodisav”, and thence from the Sanskrit “Bodhisattva”. Somewhere over the successive retellings, the Buddha’s story morphed into a Christian one.

And the story was fully embraced by Christian churches. The Roman Catholic Church included the two saints in older versions of its martyrology, and various Orthodox churches still mention them today.2 Translations of their story reached as far as Iceland, and were even brought by Christian missionaries back to Asia, as seen in a 1591 Japanese translation by Jesuit missionaries.3

Individual Western historians have noted the similarities between the story of Barlaam and Josaphat and the story of the Buddha for centuries – for example, it’s mentioned in an early printed edition of Marco Polo’s travels from 1446.4 But it was only in the mid-19th century that the saints’ tale was generally acknowledged as being derived from the Buddha’s.

World religions have influenced each other for as long as they’ve been in contact. Many of the world’s religious monuments have remained holy sites to their local people even as specific religions came and went – the Hagia Sophia, the Parthenon, the Mosque–Cathedral of Córdoba, and Angkor Wat, just to name a few. Christianity in particular has a history of absorbing other religions’ traditions: see for example the Day of the Dead, which combined Indigenous and Spanish festivals into a national symbol of Mexico.5 But concepts from Christianity have also been borrowed by other faiths, as with the so-called “Jesus Sutras” from the 7th century which restate Christian teachings from a Taoist and Buddhist perspective.6

All this tells me that there’s something about our relationship with the divine and the sacred that transcends individual eras and places. A place or a festival or a story that appeals to our common senses of awe or morality or wisdom can reach people of any faith tradition. Perhaps they’ll interpret it through their own lens, of course. But in the end, isn’t that all we have?

1 The Balavariani (Barlaam and Josaphat), a Buddhist Tale from the Christian East. David Marshall Lang, translator. 1966. Via Archive.org https://archive.org/details/LangBalavariani
2 https://www.oca.org/saints/lives/2026/11/19/100292-saints-barlaam-the-monk-and-prince-ioasaph-of-india
3 https://docs.filologi.no/barlaam.pdf
4 https://www.philology.no/barlaam
5 https://web.archive.org/web/20190408161730/https://www.nationalgeographic.org/media/dia-de-los-muertos/
6 The Jesus sutras. Martin Palmer, 2001. https://archive.org/details/jesussutrasredis00palm/page/n9/mode/2up
The Daily Front Page 16 of 31
Tuesday, September 8, 2026 The Daily Front No. #260908 — Emacs, Reduced to Bedrock
article

Emacs Bedrock 2.0

by ashton314·▲ 183 points·52 comments·lambdaland.org ↗
Bedrock is essentially a collection of better defaults for Emacs.

Bedrock is a minimal, human-crafted starter-kit for Emacs, and today I’m pleased to announce that version 2.0 of Bedrock is available! This release requires Emacs 31, which came out in August 2026. If you are unable to upgrade Emacs that high, you should be able to remove some of the Emacs 31-specific config without too much trouble.

Bedrock is essentially a collection of better defaults for Emacs, and it uses no third-party packages in its base configuration. Part of the appeal of Emacs, of course, is the vast set of packages developed and maintained by the community, so Bedrock includes sample configurations for some of the most popular packages that I, personally, find utterly indispensable.

The idea behind Bedrock is a lot simpler than other “starter kits” or configuration frameworks: Bedrock is just an early-init.el and an init.el that you copy once into ~/.emacs.d/, and then you modify those files directly as your needs grow and your experience with Emacs matures. The files contain enough comments to point you in the right direction when you want to dig into how something works, but not so much as to drown you in a sea of prose when you’re looking for code. The configuration for third-party packages lives in a separate extras/ folder, and you can copy individual packages’ configuration directly into init.el as-needed, or import groups of packages an entire file at a time.

Changes in Bedrock 2.0

This is a major release just because Bedrock is breaking its compatibility with Emacs 29, last released in 2024. Here is a summary of major changes in this version:

  • Bedrock now uses the lexical-binding: t cookie in all its .el files.
  • All packages declared with use-package get installed if not present on the system by default.
  • Use embark-auto-prefix-help-mode instead of which-key if the embark package is loaded.
  • Remove wgrep in favor of built-in grep-change-to-grep-edit-mode.
  • Vastly improved tree-sitter configuration.
  • Nicer isearch configuration.
  • Make the built-in *Completions* buffer auto-update.
  • Turn on repeat-mode by default.

Fixes in this version:

  • Fixed theme colors not being respected.
  • Better suppression of startup message.
  • Miscellaneous little improvements.

These lists are not exhaustive (as that last bullet should make clear) but that’s the gist of it. More details are available in the changelog.

This is very much an incremental improvement—not some total-overhaul of Bedrock. If you’re a happy user of an earlier version of Bedrock, then you might not have any reason to “upgrade”. Emacs is a stable editor (though the level of active development it enjoys is both meaningful and exciting!) and it should come as no surprise that a starter kit made for the “bleeding edge” should be, well, relatively stable itself.

Upgrading from earlier versions of Bedrock

Bedrock does not support any sort of automatic upgrade path, as that would require the Bedrock code to live somewhere other than the init.el file, and the point of Bedrock is that you learn to own your entire config as quickly as possible. Instead, you should compare the commit you cloned/copied your version of Bedrock from with the new main branch and pick what you want to apply.

Thank you

I appreciate all those who have tried out Bedrock and have sent me their feedback. I made Bedrock initially as an exercise in tweaking the builtin options to get something that felt more “usable” to me. From there it turned into a configuration I could give to my friends who expressed some curiosity with Emacs. Now it’s something that way more people than I initially imaged use as a starting point for their Emacs configurations. (I’ll never know how many because no telemetry—yay privacy!) Anyway, a huge thank you to those who’ve submitted bug reports or PRs, and to the wider Emacs community as well! Thanks you your dedication to free and open-source software, you’ve made my life better. I hope Bedrock improves your life as well.

— Ashton

The Daily Front Page 17 of 31
Tuesday, September 8, 2026 The Daily Front No. #260908 — Testing the Agents
article

How well do agents use test/verification techniques?

by vinhnx·▲ 177 points·64 comments·danluu.com ↗
Software quality seems to be getting worse.

How well do agents use test/verification techniques?

We previously noted that, while it's easier than ever to hit a particular quality bar by having coding agents use effective test techniques, software quality seems to be getting worse, indicating that whatever defaults developers are using may not work very well. Here, we test if simple instructions to agents to use particular techniques or libraries improve implementation correctness, as a kind of test to see how effective agents are when guided by someone with no expertise in testing who's maybe heard that you should apply certain techniques or use certain libraries.

We'll re-use the Zstd implementation eval discussed in this comparison of agentic programming language effectiveness and, instead, compare different testing techniques and testing libraries when agents are given a prompt to implement Zstd with different addendums, such as "Use test-driven development", "Use Lean 4", "Use QuickCheck", "Use property-based testing", etc. I also ran some other evals, such as on the IMAP RFC, which are briefly discussed.

All implementations were in Rust. The 26 prompt conditions tested were ACL2, Alloy, "Audit and fuzz risky areas", "Audit first", Creusot, Default (no additional instructions), Differential testing, Fuzzing, Hegel, Insta, Judgement (agents asked to use the best technique), Kani, Lean 4, "Make no mistakes", Metamorphic testing, Mutation testing, Property-based testing, Proptest, QuickCheck, rstest, Rust built-in test framework, SMT solvers (with Z3, cvc5, and Yices, all available), Spin, TDD, TLA+, and Verus. Additional, 4 skills were tested: Hegel with the official Hegel skill, the ECC Rust test skill (ECC is a collection of skills with 250k GitHub stars and 38k forks), the Trail of Bits property test skill, and a test skill I wrote (I'm a luddite who uses prompts instead of skills and have no feel for how to write a good skill). Other than my skill, the skills were chosen because those were the top skills codex turned up when asked to find relevant skills.

Predictions

I pre-registered some guesses on how conditions will do:

  • TDD will underperform (55% confidence)

    • I actually added TDD specifically because I thought it would underperform
    • My confidence is low here because I don't know what agents will do when instructed to do TDD; perhaps agents won't do TDD and will do something that doesn't underperform (or perhaps I'm wrong about TDD underperformance)
  • Formal methods will not overperform (52% confidence)

    • My thought here is that formal methods are effective and useful (more so now than ever), good test methods are also effective and useful and, on simple problems, formal methods shouldn't outperform if used at a similar level of competence
    • As with the above, but even more so, my confidence is low here because I don't know what agents will do when instructed to do anything, and formal methods have been more hyped than effective test techniques for agentic coding, so it's entirely plausible that labs have trained agents with RL environments with synthetic data which trains them to be very effective with formal methods without having trained agents to be effective with good test techniques (which I would expect to be easier to do, but not done because of how relatively untrendy effective test techniques are)
  • Make no mistakes will not outperform no instructions (95% confidence)

    • It's a joke, and one that a lot of people have tried. If it worked, surely people would've noticed?
  • The ECC test skill (with 250k stars and 38k forks) will not outperform (65% confidence)

    • It's somewhat big and doesn't have any information I'd expect to be useful. It instructs agents to use TDD; to the extent that it gets agents to use TDD, I'd expect this to make things worse (and it's more directive than the TDD condition and perhaps more likely to succeed, although for all I know that makes it less likely to succeed); the rest of the information doesn't seem useful and has some cost
    • All of my skill predictions are low confidence because I don't tend to use skills and don't know how to really evaluate them. I'm thinking of this like, "how effective would it be if I passed the text in as a prompt and had this thing floating around in the LLM's context window?"
  • Hegel's skill will not outperform (65% confidence)

    • It's very big (the SKILL.md plus the linked Rust reference are over 20k tokens) and reads more like a tutorial than agent instructions
  • The Trail of Bits test skill will not outperform (55% confidence)

    • It has what looks like it might be useful information, but it's also fairly big

Overall results

Below, we have a very messy graph which shows the results for the conditions tested (codex with GPT-5.6 Sol, with medium and xhigh efforts). When looking at data, I tend to prefer much denser and messier graphs than most people, such as the first graph here. Because most people find these kinds of graphs unreadably messy, I tend to split information out into a series of graphs, each of which shows less information, when presenting information to others. For reasons discussed elow, I'm not going to do this here and am just going to present this extremely messy graph where we have cost on the x axis and the fraction of runs that passed 100% of the (hidden) tests on the y axis, average of 80 runs from each condition and effort (mousing over items shows bootstrap covariance, 50% uncertainty, and there's some attempt at making like things similar colors, e.g., blue-ish for formal methods, green-ish for property-based testing, etc.):

One thing we can see is that nothing really wildly outperforms. However, Default (no additional instructions) does well above average. Looking at xhigh, on average, the fuzzing and PBT-related conditions did a little better than formal methods on average, with the situation being a lot more mixed at medium. The testing-related skills codex recommended we try underperformed, although our quick custom skill did ok (a major difference is that our skill is designed to nudge away from their default behavior towards more productive behaviors whereas the other skills seem more like tutorials). TDD didn't do well, as predicted (one skill also suggested that agents used TDD, and that skill also fared poorly in the cases where agents attempted to follow the instruction).

If we actually look at what agents did, it quickly becomes apparent that, in general, agents don't know how to use these tools or techniques very well. As we noted here, and as everybody I've talked to has also noted, agents are really bad at testing and don't seem to understand how to test reasonably "by default". For example, here's a comment by Gary Bernhardt:

AI agents' approach to testing, more or less:

  1. Take the pathological cases dreamed up by someone objecting to mocks 15 years ago, without ever having actually used mocks. Naive dreams of excessive mocking.
  2. Make those pathologies the backbone of your testing strategy.

It turns out, if you ask agents to use a particular test technique or test library, this approach doesn't change as much as you'd hope. We'll look at what happened in cases in more detail, but at a high level, with test techniques, agents tend to either just write the tests they would normally write, but inside a framework for a different type of test technique, or they'll use a technique superficially but not really do the things that get the value out of the technique. For the most part, when a technique was named, they did what Gary described, but with respect to that technique (for example, for formal methods, they mostly proved irrelevant properties and with property-based testing, agents would lean heavily on totally random inputs and heavily hit invalid/rejection cases or find a trivial property to check and run low-value random cases against the trivial property). Results weren't materially different on the IMAP RFC (where I tried 40 runs of each condition) or other random RFCs (where I tried a few individual runs). In general, regardless of the type of problem, whether it's some kind of bit manipulation problem like Zstd, a protocol like IMAP1, or anything else, agents did not use formal methods or test libraries or techniques in an effective way.

On xhigh, agents were generally able to get the tests they wrote to pass, but they wrote poor tests (e.g., they'd submit four identical bitstreams into a test of a feature that uses four bitstreams and miss any bug that would occur because they transposed bitstreams). And as we noted previously on the Zstd eval with respect to languages, running at a lower effort level in a naive loop gets worse results (agents do even more of this and stall out with lower correctness).

I'm curious why AI labs haven't created RL envs to get agents to learn how to test well since software not working reasonably seems important for coding agent adoption and it also seems like the kind of thing that's amenable to RL. As we previously saw, agents have gotten quite good at bounded runtime optimization problems, which makes sense because that's exactly the kind of thing you cheaply create a ton of RL envs to train on. Maybe this is one of those things that's harder than it seems when you try it, but creating RL envs for effective testing and test techniques seems like it's in the same class of problem. Perhaps the limiting factor is just that knowledge of effective test techniques isn't very widespread, so no one's thought to try it and people are getting agents to test inefficiently (for example, by doing standard unit testing)2, or maybe this problem is much harder to package up than runtime optimization for some reason? It's possible this will be a moot point soon if agents get so good that they can generally write correct code without testing or verification, but at least for the state of publicly available agents from inception until now (September 2026), it seems like agents having some idea how to test without being guided by a testing expert would've substantially increased agentic coding effectiveness.

Below we'll look at how agents did things for each condition, ordered from worst correctness to best, but I would caution anyone against drawing any kind of strong conclusions from the ordering.

A lot of the failures here seem analogous to the failures we saw when we looked at the impact of programming language on token usage and correctness, in that the failures are often idiosyncratic. For example, with programming languages, we saw that agents had a fairly high rate of getting the semantics of byte conversion incorrect in Clojure but not Java, even though agents "should" (and probably sort of do) know that they can get Java byte conversion semantics by converting with unchecked-byte instead of byte.

Although people have all sorts of hand wave-y high-level explanations for why some languages are better for agents than others, when we look at what agents actually do and what the failure modes are, none of the explanations I've heard for why someone's pet language is suited for agentic coding, whether it's Elixir or Ocaml or J, are actually true (with the exception of comments about Rust's memory safety, which were validated in the multi-language pandoc eval we tried by comparison memory safety issues between agent-written C, C++, and Rust). Instead, we see a bunch of idiosyncratic failures that happen for unclear reasons3. With languages, because we can observe a moderate correlation between language popularity and performance (both lower cost and higher correctness), it seems reasonable to guess that the reason is because there was more training data (possibly synthetic data and not just human-written code) for more popular languages. Here, there isn't a clear pattern, other than that agents are mostly not very effective at applying test or verification techniques when all they have is the name of a library or technique (we'll discuss what works better afterwards). If you don't want to read about what happened in each condition, click here to skip to the last item.

Verus

Verus uses an SMT solver and various types of reasoning to prove that the code matches specifications.

Although Verus can prove that code matches specifications, agents didn't do that. Instead, they made proofs about various abstract properties relating to Zstd. I've not used a tool like Verus myself, so I can't speak to what an expert or even a beginner user would normally do, but from reading the tutorial, I find it a bit odd that agents didn't attempt to use Verus to verify any of the actual code and only used it to do abstract reasoning, as it seems designed to make it easy to prove properties about the actual code.

Additionally, if we look at the properties proved, there were generally few properties proved and the properties that were proved were uninteresting. For example, agents would prove things like "given a valid cursor/index/distance, the resulting operation remains in bounds", which isn't bad to prove, but wasn't really a source of bugs. Also, agents would frequently write vacuous proofs that were effectively A => A. An actual Verus proof of this form was:

  requires                                                                                                                                                                                         
      0 < a <= window,                                                                           
      0 < b <= window,
      0 < c <= window,
  ensures
      0 < c <= window,
      0 < a <= window,
      0 < b <= window,

In cases where agents actually proved something, they generally proved something relatively simple and avoided proving properties about the parts that were likely to have a bug (for example, agents often failed to reverse the bitstream order for encode and decode and would write tests that failed to detect this because the tests were palindromic; perhaps some kind of proof of reversal here might get agents to "think" about this in a different way).

It doesn't seem that agents were getting value out of Verus when just provided with Verus and the Verus docs.

If we look at the result, the aggregate xhigh Verus results are fine (slightly lower correctness than average, but much cheaper). The medium results had average cost and the lowest percentage of correct runs as well as the lowest average number of correct tests. Because agents didn't really get value out of Verus, what they actually did for correctness was mostly just traditional tests (built-in Rust #[test] functions with unit tests). When going from medium to xhigh, agents spend much more effort on traditional testing and only a bit more effort on using Verus, which allowed the xhigh result to be ok.

Looking at the actual tests, for one of the two features which agents using Verus did much worse on (the four stream jump table), Verus agents wrote a test for this in 89 out of 160 cases, coincidentally the exact same number as Default agents, but Verus agents were much more likely to write bad tests. They were more likely to encode incorrect results in the tests as well as make easy to pass tests that don't cover the space well, such as making all four streams identical. This kind of thing is what I meant when I said that the failures were idiosyncratic. There's nothing about Verus that necessarily makes one write poor tests when not using Verus and we wouldn't, in general, expect a human who's used Verus to write bad unit tests, in the same way that we wouldn't expect a human using Clojure to make more byte conversion mistakes, but this happened here for whatever reason (possibly a coincidence).

I don't know if folks inside AI labs can get access to better information on why things happened, but here on the outside it's generally quite difficult to tell why something like this happened (even when we formed a plausible hypothesis for the language issue, it required running many samples of many languages, and papers we looked at which studied the same thing didn't observe the language popularity / agentic effectiveness correlation because they either looked at too few languages to be able to reason about such a weak correlation or they looked at problems that were too small and too trivial).

Alloy

Alloy is often called a bounded model checker. This is maybe not quite right with Alloy 6 since that introduces some extra features, but this is way outside of my area of expertise. My understanding is that, with Alloy, you normally prove properties about your model (as opposed to proving that your code works).

Alloy got the 2nd worst correctness score and, unusually, scored generally poorly on both medium and xhigh. Although it isn't shown (because it doesn't seem to add anything), in general, results were highly correlated between max and xhigh, which were quite different from medium results.

As we saw with Verus, agents using Alloy pretty much relied on standard Rust #[test] for correctness and mostly faffed about with Alloy. Once again, using a formal tool poorly did not help with correctness.

There were individual cases of Alloy use that were close to finding an issue or risk, but even then, only a small number. In one case, Alloy found a counterexample which then caused the agent to implement the Rust version with a mitigation for the potential bug. Unfortunately, the counterexample relied on an 8-bit overflow that couldn't happen in practice because the actual implementation used 64-bit usize with no possibility of overflow given the inputs, so it just made the code more complex without preventing an actual bug.

In another case, the Alloy specification was incorrect and a related test failed. After the test failed, the agent fixed the Alloy specification. Had the specification been correct, perhaps the agent would've written the correct code without the failure. There were some cases where it's possible the good version of this happened, but it's not clear if an actual potential bug was prevented.

Alloy agents did model things that were more closely related to the Zstd algorithm than Verus agents (which mostly checked things like arithmetic), but it was still the wrong modeling.

Differential testing

Differential testing is a technique where you give the same inputs to multiple implementations and then compare results to find issues. In principle, this seems like a reasonable thing to try with LLMs as we often get different results from different rolls of the dice, and as we noted here, having an agent iterate more on an implementation (which might be only part of the entire thing, perhaps even only part of a function) often works worse than having the agent restart from scratch.

But this gave us the third worst results. In this case, we had slightly above average results on xhigh and far below average results on medium. None of the agents created two full implementations to compare. Out of 160 runs, 135 did something you might call differential testing, but like the other conditions we've seen, these were generally trivial and effectively useless. And, in the cases where differential testing might've caught a bug, instead of implementing things in independent ways, agents just did the same thing twice and encoded the same bug in both versions.

I sometimes tell agents to do things independently and get them to launch with separate contexts, but this was not done effectively for differential and agents would generally just write the same thing twice.

Hegel Skill

It makes sense to discuss how the official Hegel skill changes Hegel behavior, but in reverse correctness order, Hegel Skill appears above Hegel because the result was worse on correctness. See the Hegel section below for discussion of this skill.

Lean 4

Lean 4 can maybe be described as an interactive theorem prover.

Although I didn't pre-register a guess about Lean, if I had pre-registered guesses on which formal tools would do well, I would've put Lean on the list of things I'd expect to do well because it's relatively hot/trendy and therefore seems relatively likely to have good performance due to synthetic data from RL envs.

The Lean agents did prove properties, like the Verus condition, agents mostly did arithmetic proofs that didn't hit the bug-prone or risk surface areas.

Like the other formal conditions, Lean agents relied heavily on standard Rust tests. As with the formal conditions so far, doing a few proofs of things that don't matter didn't help with correctness.

QuickCheck

QuickCheck is a property-based testing library, probably the best known such library for a long time, although Hypothesis might currently hold that crown.

Unfortunately, agents were about as effective at using property-based testing as they were at using the formal tools we've seen so far. When using QuickCheck, agents mostly wrote very simple "smoke tests" that didn't check much. They also used random inputs, which, when fully randomized, are pretty poor for testing something like Zstd (because they just go down one of a few failure/rejection code paths).

Also, relatively few properties were checked. Although all agents used QuickCheck, 63 out of the 160 runs only checked a single property. Agents once again mostly relied on traditional testing, although they technically did use QuickCheck. For whatever reason, agents actually wrote more traditional tests than under the Default condition or most other conditions, but did fewer test-fix iterations (which resulted in this condition coming in with below average cost).

TDD

TDD underperformed here as well as in the IMAP RFC eval.

The TDD prompt seemed to cause large changes to agent behavior. Agents produced twice as many tests, and worked in a much more iterative test-code-test-code-etc. workflow, although a TDD advocate would probably say that agents didn't actually use TDD. There were only a few instances of agents doing some kind of fine-grained iterative TDD.

Overall, agents wrote more tests up front; for example, agents had one or more failing tests in 67 of 160 cases before doing substantial (non-stub) implementation, vs. 0 of 160 for the Default condition.

For broad test classes, TDD had more tests of every kind. There were more small, trivial tests and there were also more integration and end-to-end tests. Any kind of obvious high-level "agents did too much or too little of X" doesn't seem to fit the data. If we look at specific failures and how they were missed by tests, we can observe that the TDD condition had a number of these. For example, Zstd uses something called a jump table when there are four Huffman streams.

TDD agents were more likely to fail the eval test for this although they wrote more tests that cover the general case. For whatever reason, TDD agents were more likely to write tests that don't cover hard cases (e.g., making all four streams identical and then also making them trivial, like we saw with Verus). This is another case where I'd be curious what kind of visibility people at AI labs have since it's not obvious from the outside why priming agents with TDD made them write worse tests and worse implementations.

If we only had TDD and a few test conditions to go on, a hypothesis might be that TDD'd code often seems to have a lot of small tests that aren't very good, so maybe priming agents with TDD causes them to write more of these sorts of ineffective tests. But it's not clear why we should see the same pattern with Verus. Maybe we could tell whether or not this is true for TDD if there's a shared reason for the Verus (or other) behavior by re-running the experiment on an open model and inspecting what's actually going on inside the model at some level?

Two of the skills also caused agents to run in a more iterative approach, perhaps on the theory that executing more frequently would give better results, and both of those skills also underperformed. In general, across all conditions, agents were able to get the tests they wrote to pass on xhigh and max (not shown, but max had slightly better correctness than xhigh at substantially better cost). Getting their own tests to pass more iteratively tended to get agents to write more incorrect tests that would enforce incorrect behavior.

Yossi Kreinin had this thought for why TDD might result in worse tests:

fwiw, I think if you write the tests before the code, it's harder to test the harder cases since you know less about what is going to be hard, and even if you do random testing which I don't think "tdd" is associated with, you are less likely to steer the distribution in the direction where the bugs are. if you wrote the code or at least can look at it, you know what seems trivially correct and what might or might not work since it's not easy to understand what it does. in other words, tdd steers you towards black box testing which for complicated machinery seems to me to be less effective than white box testing; pretty sure this is how it works with people, less sure about agents

Was my guess that TDD would underperform correct? Strictly on the result, the answer is yes. On my reasoning (not explicitly pre-registered in writing, but I do know what I was thinking), I think it's not clear. My thinking was something like, as we've recently discussed in a variety of contexts, getting agents to actually do something like the right thing and not just overfit is a key part of achieving good performance or correctness with agents. Speaking to the methodology in general and not how this instruction changed agent behavior, TDD seems primed to cause overfitting.

Agents did write worse tests and sometimes used a relatively expensive and ineffective iterative workflow, but I don't know that the failure mode I'd expect from a human using TDD and then directing agents to implement was the real problem here, and that problem was where my intuition came from. I would rate the reasoning here as perhaps and perhaps not in the right vicinity; I think more evals and investigation would be necessary to decide this and I would guess that the result of additional data would be that my original reasoning is wrong.

Spin

Spin is a model checker.

Now we're getting into the range where results weren't far from average. Spin did moderately worse than average on both medium and xhigh, at below average cost. As we saw with the other formal tools, usage of Spin was generally ineffective. In this case specifically, using Spin to model a certain class of behavior had no correlation to passing or failing the hidden tests covering that behavior. Usage of Spin was superficial and not productive.

Hegel

Hegel is a property-based testing library based on Hypothesis.

As we might expect by now, agents didn't use Hegel effectively. To the extent they used it, they used it superficially, and they generally used it after heavily relying on ordinary testing. Since just saying that agents didn't really meaningfully do the thing is repetitive, I'll make these sections short and only highlight particular curiosities.

The actual workflow agents used was generally

  1. Read the RFC and API/contract
  2. Implement Zstd
  3. Run normal tests
  4. Read Hegel docs
  5. Use Hegel to write 1-4 simple property-based tests
  6. Continue using normal built-in Rust tests

As noted above, the Hegel skill didn't improve correctness. Correctness was worse (though it was close enough that this could've been random). What was more striking was that cost was much higher (26% higher on medium and 41% on xhigh), for reasons which seem causal.

The skill caused agents to generate more tests. The additional tests were mostly checks that malformed inputs don't cause a panic and round-trip tests. The former is something that agents were already inclined to do an excessive amount of for all of the property-based and fuzzing conditions, so additional effort there wasn't useful. The latter doesn't seem like an inherently bad idea (I even often explicitly instruct agents to create round-trip tests and they seem to be useful to check specific properties), but it wasn't done in any of the most bug-prone areas. Without additional instruction, agents were inclined to create round-trip tests for relatively trivial properties that were already likely to be correct.

As for the cost, there are multiple reasons for the cost. One is that the skill is fairly large (34k characters for the skill, which also loads a 45k Rust-specific reference, which ends up being more than 20k tokens). This was loaded at the start of the run and was re-read on many subsequent actions. This resulted in an average additional dollar cost of 16% for medium and 18% for xhigh (by raw tokens, the average increase was 900k on medium and 1.8M on xhigh; although the cache hit rate on these was very high, 99.85% after the initial read, they were re-read enough that this was still a substantial fraction of total cost).

A multiplicative cost (this multiplier is included in the previous numbers) is that the skill also specified a structured set of operations that cause a lot more work to get done. This work didn't increase correctness, so this increased cost without a concomitant benefit.

One thing to note is that the skill was "only" used in 157 out of 160 cases. As is generally the case when using LLMs, the actions and results are random. If you have a skill available that you think an agent should use for a particular task, it may or may not use it depending on factors that seem opaque to people outside of AI labs.

ToB skill

In this case, only 108 out of 160 runs actually opened the skill to read it. The skill suggests using proptest in Rust, but the skill suggests approval is required to add a dependency and these were all single-turn autonomous runs, so this wasn't done.

As with the other property test cases seen so far, property testing was rudimentary and not done in a helpful way.

Rstest

Rstest is a fixture-based test library.

Agents effectively didn't use rstest. They technically did use it, but they pretty much just wrote standard unit tests inside rstest and didn't use rstest as intended, defeating the purpose of rstest. While this is arguably true at some high level for techniques seen so far, agents were at least superficially using some of the other techniques (such as writing some low-value property tests with Hegel), but here agents didn't use the thing that makes Rstest Rstest (the analogous behavior for the property-based testing libraries would be if they just wrote non-property-based unit tests with them).

Rust test

This is referring to the standard Rust built-in test framework that agents used in the Default condition and also very heavily relied on in the other conditions.

Explicitly asking agents to use the built-in test framework resulted in more tests (double normal on medium, 25% more on xhigh), but this didn't result in better correctness. When agents got things wrong, it was often because they didn't test significant behavior or implemented incorrect test behavior. Adding more tests didn't materially increase coverage of risky behaviors or reduce the fraction of runs with tests that encoded incorrect behavior.

Yossi Kreinin added:

i think the fixed input/output style of testing encourages this in machines and humans alike. if you generate inputs you need to then have code that classifies output as correct or incorrect, and while this code itself might be buggy, it at least makes you think about what correct means and how to tell if something is correct more easily than running the code and assuming its output is the right answer. with fixed outputs you are quite likely to just encode the output of the code and convince yourself that it makes sense

Creusot

Creusot sits in the same space as Verus.

As we've seen with the other formal conditions, Creusot was not used effectively.

Mutation testing

Mutation testing involves modifying the code to determine how effective tests are and then adding tests to get good coverage. Although mutation testing is a standard programming term, agents generally didn't actually do mutation testing and instead did normal testing with some small amount of mutating things in a way that isn't really mutation testing, similar to how the TDD instruction modified behavior but didn't get agents to do TDD.

There were a few cases where mutation testing occurred, but only a small amount, and that was rare.

Judgement

This condition asked agents to adaptively use testing methods as appropriate based on their judgement. Given what we've seen so far, unsurprisingly, agents mostly used standard Rust unit tests. A few agents did some limited fuzzing. Agents had access to other test and formal libraries but didn't use them.

Fuzzing

Fuzzing involves randomizing test inputs in some way.

Agents relied heavily on sending random bytes in, which mostly resulted in going down the same code paths (invalid input). Agents also tried sending in random variations of valid inputs, which mostly also just repeatedly exercised input rejection paths.

On the rare occasion that agents generated random structured inputs (10 out of 160 cases), this found real bugs half the time, some of which were non-trivial cases. Using fuzzing a bit effectively in 5 out of 160 cases isn't exactly good, but this was one of the more effective uses of a technique that we've seen so far. This also seems to indicate that agents could be trained to do this better and also that that can be directed to do this better without changes in training. They do, in some sense, know how to do this; they just don't normally actually do it without being pushed into doing it.

Insta

Insta is a library for snapshot testing (sometimes called golden testing), where you compare results to a "snapshot" or "golden file" of correct results. Speaking generally, a snapshot is usually some kind of serialized data, e.g., it could be a JSON object of a data structure, a log of CLI output, etc.

As you might expect, snapshot testing was barely used and agents mostly relied on traditional tests. Agents did use Insta, but would often just write normal unit tests in Insta.

SMT

Agents were instructed to use an SMT solver, with Z3, cvc5, and Yices installed.

Agents mostly used the SMT solver as a kind of scratchpad to compute things like FSE state ranges, header arithmetic, etc. Even when agents modeled something, they would generally not model the right thing to avoid a common mistake.

For example, there's a computation that should've been byte1 + (byte2 << 8) + 0x7F00. Many agents implemented byte1 + (byte2 << 8) | 0x7F00 instead. Agents used SMT solvers to prove properties relating to this computation, but then still wrote the wrong code, making SMT use seemingly no better than Default (no instructions).

TLA+

TLA+ is a language and tool for modeling behaviors.

We're into the set of above average results (but still worse than Default) but, as noted above, I wouldn't take the actual ordering too seriously. Though this isn't necessarily significant, TLA+ did score a bit above average on medium and more above average on xhigh.

159/160 agents created some kind of TLA+ model, generally a state-machine model of Zstd. For particular coverage, 30 modeled Huffman/FSE/entropy (areas that often had bugs). As with the other formal cases, TLA+ modeling happened relatively late in the flow (after a lot of standard tests and implementation). Agents sometimes found and fixed errors in the TLA+ model, but I didn't find an instance of a TLA+ issue resulting in an actual change in the Rust code.

Although there was some real looking TLA+ modelling happening, if this improved correctness, it did so in a small way that was difficult to observe. In general, runs that had more sophisticated TLA+ modeling did not have better correctness.

Metamorphic testing

With metamorphic testing, we check that related inputs produce outputs with the expected relationship. For example, you could check that, for a sort function, changing the order of unequal inputs doesn't change the order of the outputs, or for addition, adding a value to an input adds the value to the output modulo overflow.

As we've seen for the other conditions, Metamorphic testing wasn't done very usefully with respect to correctness. Some actually reasonable properties were checked (e.g., inserting a skippable frame at a frame boundary shouldn't change the output, legal block repartitioning shouldn't change outputs, etc.), but these didn't hit the areas that agents got wrong relatively frequently so checking these properties didn't help. In general, agents seemed to be fans of the old joke:

A policeman sees a drunk man searching for something under a streetlight and asks what the drunk has lost. He says he lost his keys and they both look under the streetlight together. After a few minutes the policeman asks if he is sure he lost them here, and the drunk replies, no, and that he lost them in the park. The policeman asks why he is searching here, and the drunk replies, "this is where the light is".

Curiously, metamorphic testing was used less on xhigh than on medium.

ECC

The ECC Rust test skill did ok, but mostly because large parts of the skill were ignored. Agents generally opened and read the skill (153/160 read it) and this seemed to cause them to generate more tests. Not only did agents generate more tests in this condition, if we look at when agents read the skill (earlier vs. lateer vs. never), there's an exposure-based gradient in how many tests were added.

Although ECC scored almost as well as Default, based on how agents did when more exposed to the skill, I would guess that this is random. The earlier an agent looked at the skill, the more its behavior was impacted and the worse the correctness result.

ECC seemed to do ok in terms of raw score because the 7 agents that didn't read the skill did unusually well and got a 100% correct result, and then the 9 agents that looked at ECC late and were only barely influenced also did well and had 100% correctness. This also explains the unusual ECC result that medium had the same score as xhigh (all but one of these runs where agents didn't really look at the skill happened on medium). While it's true that there may be some kind bias in when the skill gets invoked or not, the overall pattern would indicate that ECC is not effective unless you think ECC acts as a good luck charm that improves results, but only when the skill isn't really used, which is more likely to happen at lower effort levels.

Of course, agents shouldn't be influenced by a skill they didn't look at and we should score this based on the cases where the skill was used. If we look at the cases where the skill actually influenced agents, ECC scores below average (between Rust built-in framework and Creusot), with a very similar failure mode to Rust built-in framework of having a large number of small and not meaningful tests. The skill tells agents to use red-green TDD. The agent behavior probably isn't what a TDD practitioner would call TDD, but agents do write a small test before implementing functionality, which results in a large number of tests. As noted above, this isn't an effective way for agents to develop, so the result is worse than no instruction and no skill.

BTW, as we noted when we tried out Caveman mode, there's quite a bit of variance and people are often misled into thinking a skill is useful by a few small runs. In this case, we tried 160 runs of a skill, a fairly large number, more than any reasonable person would do. And yet, superficially, if we just look at the score, ECC seems ok.

We would need a much larger number of runs to average out the noise inherent when using an LLM. We can do what we did here, and inspect the results and use our human brains a little bit, but I rarely see this done when people are talking about public LLM benchmarks, whether it's for skills or anything else (I did try having LLMs analyze the results but, as usual, even with current public SOTA models, the analysis was poor and full of basic reasoning errors). Instead, I mostly see people pass around the top-line number, even when it's not meaningful for boring statistical reasons or, worse yet, the benchmark is fatally flawed, as we saw with Senior SWE-Bench.

Default

Default gave the agent no test or verification instructions.

Given what we've seen so far, it's not surprising that Default scored above average. Agents generally did things that were not useful when asked to use particular libraries or use particular test techniques. It stands to reason that not telling agents to do things that will make them do useless work does better than telling them to do things that will make them do useless work.

Audit

Audit asked agents to audit the code after implementation. 152/160 agents actually did this and 151 agents claimed find an issue and then made a change as a result of the audit. Agents generally picked reasonable areas to audit, but usually didn't do an independent audit with a fresh context (which I will often ask agents to do) and often just made the same mistake in the audit that they had already made.

42 used an independent agent, but these runs actually scored worse (it's possible this isn't causal and agents decided to spin off an independent audit because they were in a worse or harder situation). Audit ended up with the best correctness on xhigh, but below average correctness on medium, and all of this auditing substantially increased cost, especially on xhigh. On average, Audit did about as well as Default and it's not clear if it's really better on xhigh and worse on medium. That would be plausible, but I don't think we have enough evidence to tell.

Em Chu had the following comment:

The results here are consistent with my experience. Auditing code is where most of my tokens go at the moment because I find it quite useful. I always give the two instructions though:

  • Don't spawn subagents; read and understand the code/diff yourself
  • Don't execute any of the code

because I find the LLM to be significantly dumber if you let it do either of those (though of course I haven't measured...). It really doesn't read or reason about code by default, even if I'm never making a change big enough to exceed its context window.

I also usually include some BS like "be adversarial" "consider all possible combinations of features" "consider the entire input space" but I'm less sure that helps at all.

It would be interesting to try that, but as I've noted in my recent posts, I'm trying to go into less detail in posts, so maybe that will be a topic for another post.

Audit and fuzz risky areas

For Zstd, when this instruction was followed, it caused agents to focus heavily on FSE, Huffman, bit readers, and state. These were areas where, in general, agents often missed issues, so agents were correct to think that these areas were risky. The areas that were targeted for fuzzing were better choices than the plain Fuzzing condition.

On medium effort, agents mostly ignored the instruction and didn't do it, but they did follow instructions on xhigh. While this condition didn't perform poorly, it didn't seem to do better than no instructions.

When looking at what agents actually did, one issue was that agents often just generated a bunch of random inputs which were generally invalid and wouldn't test any interesting condition.

When a human tester generates randomized tests, they'll generally try to target the randomization in a way that generates "interesting" inputs and agents failed to do that. Agents also didn't check outputs very effectively and, in many cases, only looked for crashes. Fuzzing is often associated with only checking for crashes and not checking for properties, so this is maybe not too surprising, but it's probably not what a human would want if they were testing a Zstd implementation.

Make no mistakes

Although this technically scored higher than Default, behavior didn't seem to be meaningfully different and the scores are quite close; I would guess that this is due to random variation. At every level at which I looked at the results, they were indistinguishable from random draws of Default.

Kani

Kani is a Rust model checking library.

In terms of "actually using a formal method on the code that will execute", Kani had the best coverage in that Kani was actually used on the Zstd code. However, that only happened occasionally and most use was superficial.

There was one case where real Kani use caught a non-trivial bug that caused Rust code to change. 1 out of 160 isn't amazing, but it does indicate that agents can stumble into using Kani reasonably sometimes (which, I would guess, means that, if used in an RL env, models could learn how to use Kani more effectively).

Kani had noticeably higher cost than other conditions. This seemed to be because reading Kani output repeatedly was expensive, which resulted in a high input token cost.

ACL2

ACL2 is a theorem prover. One thing to note about the result here is that, in many cases, ACL2 OOM'd (192 GiB limit). OOM results weren't counted, which biases the results in some opaque way.

Although ACL2 scored higher than Default, I think it would be surprising if this was causal and significant. As we saw with almost all of the other formal methods, ACL2 was mostly used to prove things that didn't significantly impact correctness, so it's not clear why this would improve correctness.

With many different conditions, we wouldn't expect Default or the seemingly equivalent Make no mistakes to be at the top unless other conditions had severely degraded performance.

Proptest

Proptest is a property-based testing library.

Just as we saw with the other randomized testing, most tests weren't very interesting, and a too-heavy reliance on randomness caused poor coverage.

Despite generally poor use of property-based testing, proptest's shrinking (finding a simpler input that causes a test failure) did sometimes provide some value, which is better than the little to no value we saw in most other cases.

Property-based testing

As with the other technique-based approaches, agents had a container with all options installed. Every agent chose to use proptest, so this effectively became a 2nd proptest condition.

As with the proptest condition, tests were mostly not very good but they did sometimes find bugs and shrinking seemed to generate some wins.

I find it mildly interesting that this second "accidental" proptest arm also scored well above average, just like proptest.

Skill

Here, Skill refers to the skill I wrote to test having a simple skill (as opposed to the large/complex skills that were what I found when I asked an agent to find relevant testing skills).

Maybe I should use skills, but I generally don't and instead rely on prompting and seeing what happened and then prompting some more. As a result, I have no intuition for what makes a good skill since I don't have any practice at it, but Max Bittker suggested that it would be interesting to see the result with a test skill that attempts to encode some information I have in my head about testing. On seeing the result of this, he had an "I told you so" reaction.

I didn't write a pre-registered guess about this, but the guess in my head was that this wouldn't work well. From my attempt at conveying this to humans in 2015, which I would say pretty much failed, I don't think I'm good at explicitly laying out how someone should test in writing. I've sat down with people and showed them what to do, which has generally converted them for life and turned them into way above average bug finders, but being able to convey something by showing someone is a different (and easier) skill than conveying it by writing down how to do it.

In this case, the skill was:

  • Think about areas likely to have subtle bugs before implementing; for each, state likely mistakes and plausible alternative interpretations, then come up with a check where the results differ (prefer asymmetric / boundary examples on both sides of the boundary)
  • After implementing, for high risk areas, independently re-derive the result without context on production code and compare (fresh context, do not re-use helper functions)
  • When feasible, use property-based testing or randomized inputs to try to explore the space, minimizing effort on no-panic or no-crash randomization
  • When randomizing, lean towards inputs that will explore interesting state and code paths (don't just naively randomize inputs that all fall into the same error paths); this may require structured random inputs
  • If you're unsure about details, use independent reasoning to check what's correct (fresh context, do not re-use helper functions)

This got the highest score, but didn't work as intended. It didn't really do the fresh context thing almost ever, so it was pointless to have that in there and we don't know if that's something that's effective that needs to be refined to force agents to do it more frequently or if it's something that should be removed (while it's technically possible it's happening at the optimal frequency, I highly doubt it).

We noted in "Audit and fuzz risky" that agents seemed to know how to identify risky areas. This was true here as well, but this didn't necessarily mean that agents did the right thing. For example, agents identified bitstreams being reversed for encoding vs. decoding in Zstd as being risky, but agents didn't do better on tests that exercised this. If we look at specific examples, for medium run #35, an agent identified this as risky, did independent derivations and an audit, but still failed. It had a relevant test, but the input was palindromic, so reversing the order gave the same result, allowing for a failing implementation that had this backwards.

Another issue, if we can call it that, is that all of the fuzzing / property-based testing was done "by hand". Given that agents seem ok at using proptest and that proptest has some useful machinery to lean on, this skill could probably trivially be improved by instructing agents to use proptest. The instructions to agents that were intended to minimize the standard failure mode of generating many useless "too random" tests directionally worked and a larger fraction of agents generated somewhat meaningful tests, but the tests were still worse than I'd expect a human to write (or an agent with active human guidance). Without iterating on this, I'm not sure what generic guidance would be good (as opposed to spending a few minutes looking at the structure of Zstd and giving Zstd-specific guidance, which is one kind of thing that's worked well for me on other problems).

As a first draft for a skill to iterate on, I don't think this is horrible, but I don't think it's really ready to use either. I could see an improved version of this working if it were tried with many more examples to make sure there isn't overfitting to RFC-like problems, bit-manipulation-intensive problems, etc., but, since I don't normally make skills and haven't ever tried to iterate on one, this fails to capture what I or another human would do if really driving an agent.

Since I'm used to prompting and then looking at the result (not necessarily the code, but at least what agents say they did and some kind of agentic summary of what happened, and parts of actual results for some kinds of experimental work) and then re-prompting based on that, I'm not used to front-loading information, which is a fairly different problem than reacting to information. From previous fuzzing work, I've seen failure modes that agents often fall into and the skill was intended to prevent those failure modes, but it's easier to do this if you check back in even occasionally than to do it fully up front, and the up front instructions weren't sufficient to stop the standard failure modes, though they did mitigate them somewhat.

General comments

As we noted above, I didn't try to break down the data in a nice, easy to look at way. I didn't do that because, once we look at what agents actually did, it seems like they were mostly pretty ineffective and I don't think it's particularly interesting to see how well "agents using Verus badly" do compared to "agents using QuickCheck badly". One thing that I find a bit interesting is that, when asked to identify areas that are risky or prone to subtle bugs, agents were able to do that.

But, in general, regardless of the library or technique suggested, agents failed to use the technique. As previously discussed, just asking agents to "test" or repeatedly asking them to test more results in poor testing. It turns out that asking them to use test techniques (some of which I've personally found to be highly effective) also results in poor testing. The quick and dirty skill I wrote seems like it could improve things a bit, but would need more than the 2 minutes I spent on it to be actually useful. Yossi Kreinin made the comment that the state of software testing is atrocious, therefore we should expect poor results if agents fall back to their training, so to speak, which is what we saw.

For whatever reason, agents seemed to be somewhat better at using proptest, although the level of testing was well below what I'd expect out of a reasonable human who's read the proptest manual and is given some direction on how to test. I'd be curious if agents that are given more direction are more effective with proptest than with other libraries, but that's a topic for another post as I've been trying to get posts out in half an hour and we're approaching 9000 words here, which is beyond a reasonable amount to try to type in half an hour.

How do you get agents to write good tests?

My experience has been, if you guide agents to set up a reasonable test and triage structure, getting agents to add to that effectively without a huge amount of supervision works ok-ish. Because of my background (bias), the kind of testing I tend to lean on is some form of randomized testing / fuzzing / property-based testing.

I talked to Jamie Brandon about this, and he's found the same with snapshot testing. He mentioned that, on one project, when he asked agents (using a variety of models) to do snapshot testing, they would say that they were doing it and then just wouldn't do it (they would write a unit test and then say they wrote a snapshot test). On a different project, he was able to get them to write reasonable end-to-end tests with mocked IO, but only after moving the tests into a separate crate and putting instructions in AGENTS.md to keep tests in the crate and not modify the public interface.

At least to date, I've been leaning more heavily on getting agents to write the test code than Jamie (my tendency has been to type to agents in a CLI; at least for now, he favors writing code by hand a lot more than I do), but it doesn't seem to matter how you do it as long as you set up some kind of reasonable structure.

Similar to this earlier problem we looked at, it seems like doing anything remotely reasonable works. If you "talk to" an agent and give it light instructions like we did for the Zstd or IMAP evals, the agent will do poor work. But if you look at what it does and type a few more sentences, you can often get it to a good place pretty quickly (or that's what my experience has been on other problems, anyway). Without having ever attempted to understand what makes models work, my made up and probably completely wrong guess for why it seems like doing something remotely reasonable generally works is that, somewhere inside the model is some kind of understanding how to do this stuff. It's not the default, and it's not even near enough to the default that naming what technique to use works, but if the model is sufficiently primed, the knowledge for how to do this stuff will actually get put into practice.

I'd be curious if this can be effectively packaged up into skills, or if AI labs are going to start training models to get better at testing or formal methods, or if whatever they're doing that doesn't directly improve those things will still improve those indirectly enough that agents will write decent tests without much supervision or structure.

Prediction accuracy

  • TDD underperforms (55% confidence)

    • True
  • Formal methods do not outperform (52% confidence)

    • True, but not for the reason I expected. Agents failed to use them remotely effectively, so of course they couldn't outperform
  • Make no mistakes doesn't outperform no instructions (95% confidence)

    • True; outperformed most conditions because a no-op is better than getting agents to do ineffective things
  • ECC skill will not outperform

    • True; I think I would've had more confidence in this if I used skills more, since most of the skill text seems like it won't do much of anything, and the text that seems like it will do something looks counterproductive
  • Hegel skill will not outperform

    • True; another one where I would've had higher confidence if I'd used skills more, since this didn't work well for the reason I guessed; I just didn't have any confidence in my feeling about this
  • ToB skill will not outperform

    • True

Guesses I didn't register, but I could tell I held implicitly because I was surprised when I saw the result:

  • Lean will do relatively well among formal methods

    • False
  • My skill will be mediocre to bad

    • False, even though Max Bittker correctly guessed what would happen, told me this in advance, and named a reason that's consistent with what happened; it feels quite silly when someone names the reason something is going to happen and you don't believe it, and then it happens for the exact reason they named

Skills

At various times, I've felt like I have a bad/antiquated/ineffective workflow because I hear people are doing something and I've been too lazy to try it out. I've felt this way about skills for a while since I don't really use skills. Instead, I keep a large scratchpad of things that I sometimes copy+paste in as prompts, which sort of feels like the equivalent of commenting out blocks of code to save them instead of using version control.

But then I saw this talk by Thorsten Ball, where he mentions he doesn't rely heavily on skills, and I talked to a couple people who seem relatively effective with LLMs who also don't really use skills, and it made me wonder if I'm not missing out on much.

Then I tried this experiment, where my feeling was that the skills I looked at weren't going to help and are probably actually going to hurt, with low confidence since I don't know anything about skills. The skills did pretty much what I thought they would do, so it turns out the intuition I have from just seeing how agents respond to things and running a bunch of little experiments seems to hold up ok for skills. I also looked at a number of other skills that allegedly improve testing which I didn't include in this experiment that looked like they would have the same failure modes as the skills we tested.

I also ran two experiments (details not discussed here, perhaps in another 10k word post another time) on some other skills that are "official" skills that companies have to support their product. One is from a big AI lab and the other is from a "small" few billion dollar company, but in both cases, the skills made results worse, just like we saw here. Funnily enough, after these experiments, I'm actually more bullish on skills for personal use than I was before since the failure modes seem predictable and therefore fixable without a huge amount of costly experimentation. Creating a publicly released skill that's intended to be really good, work well across different models and harnesses, etc., seems like it might be hard (claude and codex seem to "want" different styles of prompting, so of course that should be true for skills as well), but just addressing the issues that cause a lot of skills to be worse than no skill for personal use seems quite doable?

Naive thoughts on skill writing

I don't know enough about skills to say how to write a good skill, but with all the skills we looked at in this post (except for the one I wrote in a minute or two) and the skills from these other two experiments, the skills seemed written like they're human tutorial instructions, in that the goal of the skill seems to be to explain how to do something. My naive thought as someone who's written all of one skill is, I'd guess that this isn't optimal when working with a model that should already have some knowledge of the topic (which was the case here and in the other experiments as well). The model is already going to have some kind of default behavior distribution, so I feel like the more natural thing to do is to give statements that will modify that behavior, not write instructions that would allow a human or non-knowledgeable agent to do the behavior at all.

One obvious problem is that we get different default behaviors from different harnesses, models, and effort levels, but throwing a bunch of text into a prompt or a skill doesn't actually change this; that's just a longer way to push the agent away from its default, with a lot of text that may do some kind of unintentonal pushing. As we saw here, much of the text just gets ignored (and what gets ignored and when is of course harness, model, and effort dependent). For example, for the ECC skill, even when it was read, had most instructions ignored, and although the TDD instructions were influential (which made results worse), the instructions to do TDD were still not really followed despite them being laid out clearly.

Since what kind of prompting is effective changes enough between model releases and effort levels, for something general like "testing code well", it's not clear to me how these skills are supposed to work across so many models and efforts. Just going from GPT-5.5 to GPT-5.6 changed how I worked substantially because a number of things that worked fairly reliably with GPT-5.5 either stopped working or became much less reliable (even though, overall, the level of capability seems higher). In the same way that I don't prompt GPT-5.6 the same way I prompted GPT-5.5, I don't think I'd want to use the same skills.

I'm not sure who, other than someone at an AI lab, would actually go through the trouble of running evals on skills to see what's effective for each model and effort level and then create a portfolio of skills that are differentiated by model and effort, and I wouldn't expect AI labs to have skills that are optimized for their competitors' harnesses and models, so I don't know about things like generic "testing" skills (as noted above, quite a few testing skills that I looked at but didn't test here looked like they would have the exact same failure modes as the pre-existing skills we tested), but I could see having a few skills that work for my own use cases with the specific harness/model/efforts that I tend to reach for.

Just thinking about testing, while there are particular pitfalls that certain models fall into at certain effort levels that I want to nudge them away from, there isn't really a generic test workflow that I want to give agents that's independent of the thing being tested and the level of quality I want from the thing and the dimensions in which I want quality, so I don't think I'd want a generic test skill that lays out a set of testing steps that agents should, in general, do. Something like this goes for a lot of task that I do, which I want done in task-specific way and not a generic way. I could imagine some kind of skill that asks me questions and then emits the correct instructions to agents, but given how fast models are improving, if I'm making something for personal use, I don't think it makes sense to spend time tweaking a skill like that until it's useful. If I was working on an agentic product and wanted more people to use it, that might be a different story, but the skills I've tried have had the same failure modes as the skills we tested here, so it seems fairly easy to make a skill that turns out to not be that effective.

I could see skills being generically useful for things like teaching agents how to execute workflows or how to interact with APIs/interfaces, such as Sawyer Hood's skill that helps agents drive a web browser. Since I haven't tried that skill, I'm not endorsing it, but from reading through it, it seems like the kind of thing that would work well and save me a lot of hassle when I'm trying to get an agent to drive a web browser. However, if you read the actual skill (and scripts), it has a very different style than the testing skills we tried here.

Thanks to Max Bittker, Yossi Kreinin, Em Chu, Dennis Snell, @panoramic.blue, and Jamie Brandon for comments/corrections/discussion.

P.S. I've had this note on my last handful of posts indicating that I'm trying an experiment where I write up half-baked (barely fixed/audited/cleaned up) results as quickly as possible because agents let you run experiments so quickly that I otherwise wouldn't write anything up at all. I actually ran this experiment immediately after I ran the programming language token cost / correctness experiment, but I haven't had time to write this up because I wanted to write up this creation of a regex engine with an interpreter and a native code compiler, this experiment with running a forked version of ripgrep that uses the native code compiler on codex's ripgrep queries, and a few other things I haven't had time to write up; I've had a goal to do each of these write-ups in half an hour, but I'm still falling pretty far behind in terms of experiments I've run vs. what I've written up.

For a couple years, I was running experiments like this and just telling a few friends about curious results and then moving on without really ever talking about these things publicly. If you have opinions on these quicker (and lower quality) experiments and write-ups, let me know what you think (X Bsky Mastodon)!

Appendix: some responses

David R. MacIver said:

Regrettably, @danluu is right. The Hegel skill sortof sucks right now.

I think the common problem with a lot of agent skills is that agents suck at writing agent skills and also everyone (including us) uses an agent to write their skills.

A mutual friend mentioned that MacIver (later?) set up a benchmark and confirmed the deficiency in the Hegel Skill and is presumebly working on either improving the skill or tweaking the Hegel comments/docs so the skill isn't necessary. If the only thing that comes out of this post is that Hegel gets an improved skill, I think that would already be pretty awesome. As previously discussed, I think measuring and benchmarking are underrated, in large part because just publishing a measurement can often point a problem people didn't know existed and motivate some changes. I've had more than a few posts that have driven some kind of change just by showing where there's a gap. It's nice that this is another one of those posts.

Appendix: agent silliness

After asking an agent to do a simple lookup of something for an analysis, it exec'd a perl process that ran for 2 hours and 20 minutes before I killed it (I really need to have something that automatically catches things like this, because it's fairly common).

A subagent used perl to do a regex search over a relatively small file (44kB, 1364 lines), but the expression was degenerate with PCRE and did a combinatorially large amount of work. I tried re-running this with the FRE regex engine we tried building in a few minutes here and it finished matching in 0.7s (Rust regex was 0.6s).

You never know what's going to happen when agents are off doing things, but there are three reasons this never should've happened in the first place. First, the agent shouldn't have invoked a regex engine that can give you a combinatorial explosion like this; there's no reason not to use a safer regex engine for this (such as ripgrep at default settings). Second, the expression was wrong; the actual regex returns a uselessly large capture and doesn't do what's intended when it succeeds. Third, why did the subagent (or the harness) not automatically kill this after the subagent finished? Of course this shouldn't happen for everything a subagent runs, but agents often leave runaway processes like this lying around.

I actually have a process that goes around cleaning up after things agents leave lying around (agents that, themselves, leak memory, temporary build artifacts that consume space, etc.), but it wasn't looking for runaway perl processes. That's another one to add, but surely I'm not the only person who's run into this problem. I guess I could open source my silly tool for this, but it should become obsolete once the major harnesses fix this, so there doesn't seem to be any good reason for anyone to even pick up my thing in the first place were I to open source it.

Appendix: experimental details

In the interest of writing this quickly, I'm going to punt on this (sorry!). The distribution across conditions wasn't fundamentally different than we saw here when we looked at how language impacts correctness and token cost. Somehow, this post I wanted to write up quickly in half an hour is almost 10k words, which is definitely more than half an hour of writing (10k words in half an hour would be over 300 words per minute).

One thing I'll note is, just like with this post on programming languages, there were a couple of things that looked like really interesting/compelling results (at least from the standpoint of just looking at the top-level graph and seeing if anything stands out), but on looking more closely at those, they were due to an experimental error caused by giving a short prompt to agents to set up the experiment. On fixing those errors, we got a much more boring negative result, except for the part where the skills codex suggested might be useful seemed to be counterproductive.


  1. Overall, I think the IMAP results were less interesting. I thought it might be more interesting to try since it's more of a "business logic" problem, but not as huge and expensive to run as the pandoc eval we look at when we looked at languages, which costs over $1k in API costs per run for some of the less effective languages, and was still $700/run in Rust to get to only 30% passing tests (note that this eval uses a much harder metric, of fraction of runs that achieve a perfect score). In the IMAP eval, only a single run achieved a perfect score ("Make no mistakes", on medium). In general, there were some bits of logic that were, apparently, too tricky for 5.6 Sol to one-shot. I suppose one way to look at it would be that Make no mistakes dominated with a 2.5% score on medium vs 0% for all other conditions. Finally, evidence that "Make no mistakes" works!

    If we look at the actual results, as before, all methods weren't very effective. Even things where you might expect a good fit. For example, you might think TLA+ woudl do well because it's a natural fit for IMAP state and concurrency; in prinicple, it should be good at modeling a lot of the higher-level parts of the protcol. However, as we saw with the Zstd results, agents "decided to" do most of the implemention in Rust, then whip out whatever tool they had been told to use in the initial prompt (5 out of 80 agents actually used TLA+ before coding, but 75/80 did it the other way around). For whatever reason, agents almost always decided to use TLA+ to model a small mailbox mutation. None used TLA+ to model things like multple observors, event queues, UIDVALIDITY and mailbox epohcs, etc., where agents made many mistakes in implementation. Instead, agents modeled things like CONDSTORE and QRESYNC, where the Default condition had a more than 99.6% test pass rate, and the TLA+ agents actually had a lower pass rate despite modeling these things.

    One thing that's true of all the evals I tried for this is that they're RFCs, which are highly unrealistic, but the way in which they're highly unrealistic is that the specs are much clearer, more detailed, and less ambiguous than the specs virtually all programmers give to agents when they ask an agent to implement something. I would expect that the failure modes we've seen here are the same or worse on most real-world problems.

    [return]

  2. Funnily enough, when I asked ChatGPT to fact check this, it told me that this paragraph was wrong because there are papers that show that people have used RL environments to train agents to test, and then linked to three papers that trained agents to write poor tests by training them to write unit tests like most programmers do. That's exactly the kind of thing that I would expect to lead to the kind of poor testing we see LLMs do today, where it takes a human who understands more effective test techniques to steer the agent. In multiple independent subfields where people care about correctness, folks have independently converged to a few sets of related techniques that are generally the opposite of writing small unit tests. Of course training agents to do this thing that's the opposite of what people do when they're serious about correctness isn't likely to result in good correctness.

    There is some work related to RL environments and randomized testing, such as this paper, but based on how ineffective models are at any of {PBT, fuzzing, randomized testing, etc.} without specific guidance, it doesn't seem that this has made it into the training of the models from the big AI labs in a serious way.

    [return]

  3. Maybe people's stated reasons for their laguage supremacy will become true after models get much better, but there doesn't seem to be any reason to think that should be the case. If I had to guess, I'd guess it will look more like poker, where people had all sorts thoughts about what's most effective that were invalidated once simulation became powerful enough that computers could outplay humans in many situations. Even "modern" concepts like "range advantage", that people often say is consistent with optimal solver play, actually don't fall out of solver data at all once you look closely. They're just nice cocktail party concepts that are easy to understand and sound compelling.

    It probably still makes sense to use concepts like that if we're talking about humans learning how to play poker against humans because, for casual play, humans are generally not going to want to study solver lines enough to understand what's near-optimal, but if we're talking about what coding agents are good and bad at, I don't see a reason to throw around cocktail party ideas when we can run experiments to see what works and the expeirments that have been run to date don't support most of these abstract reasons X is good or bad that are bandied around.

    [return]

The Daily Front Page 18 of 31
Tuesday, September 8, 2026 The Daily Front No. #260908 — Attention, Illustrated
show hn

Show HN: LLM Attention Visualization

by ifz·▲ 145 points·23 comments·ishamf.dev ↗
Turns out, we can visualize this mechanism!

One interesting thing about transformer-based large language models are that, during the generation phase, it is able to draw information from any of its previous tokens. But it needs to be selective; if every token affects the generation equally, it won't be very effective. This process needs a mechanism to decide how much a token affects the next token.

Turns out, we can visualize this mechanism!

You can tap or hover over any of the generated tokens to see the past tokens that affected* the generation.

Loading... (JavaScript required)

* "Affected" might not be fully accurate, as this visualization is highly simplified. It's calculating the attention weight, scaled by the magnitude of the value vector, aggregated across all attention heads, and summed across all layers. This is then used to control the opacity of the previous tokens. The largest values always have an opacity of 1 and the rest are interpolated.

A lot of information had to be thrown away to limit the visualization to just one numeric value per past token. Because of that, when I started implementing this, I actually thought it might not be comprehensible. But it actually can produce some interesting patterns!

For example, in the default "Office Move Summary" prompt, you can hover over the text that are copied verbatim like the address and dates. You can then see the original data stand out quite a bit, because the generated token takes up a lot of the information from the source data.

This addresses one thing that I've previously found unintuitive about LLMs. If they work by predicting the next tokens probabilistically, why are they somehow so good at copy-pasting stuff? Won't they eventually make a mistake just by random chance?

But with this mechanism, you can see that it doesn't predict the entire sequence from some limited internal states. Since it has access to all past tokens, it can just decide which past tokens to draw from when copying, and so the probability of errors can be very low. In the "Debugging an Average Function" example, you can see that this quite small model (600 million parameters) can easily reproduce an entire JS function except for the intended modification. (Although it's not actually capable of finding the issue by itself, so it needed some hints.)

Another interesting part is when you hover over the "remain" in "Existing access cards and phone numbers remain" in the "Office Move Summary" prompt. You can see that it draws from "work" in "Existing employee access cards will work" and "stay the same" in "company phone numbers will stay the same". So it's kind of combining the information from the words in both phrases, which I find quite cool.

Implementation

The visualization itself is a pretty basic React app using Transformers.js to generate the text. But, since we need to pull more data out of the model to visualize it, it can't use the regular generation loop. I had to vibe-code the generation loop in the app so we can actually keep track of the values to visualize.

Despite using a smaller model for this, it's still hundreds of megabytes, and waiting for it to download before showing anything just won't work. So I pre-generated a bunch of prompts that can be loaded and viewed instantly.

Another tricky thing is that some of the things in the visualization are not actually meant to be read, so they're not defined as outputs. I suppose if you're implementing this using Python ML libraries, it would still be easy to access them. But Transformers.js uses .onnx files that contains the entire computation graph. The model loading and computation logic is implemented in wasm, so there's no easy way to access anything other than the predefined outputs, as far as I can tell.

In the end, I used a small script to modify the onnx file just enough to expose those internal values. But that means I can't just use the regular .onnx model. Since I want to have a browser-based generation feature, I have to upload a separate instrumented model to my own Hugging Face repo and point the app there.

You can find the code in the GitHub repo.

The Daily Front Page 19 of 31
Tuesday, September 8, 2026 The Daily Front No. #260908 — A Terabyte and a Half, To Go
repository

Kimi K3 (2.8T) at 1 token/s on a MacBook Pro, streamed from four SSDs

by Argonautlabs·▲ 233 points·122 comments·github.com ↗
★ 53⑂ 0 forks Rust

ARGODRIVE Deltafin: Kimi K3 (2.8T MoE) streamed from SSDs on Apple Silicon — fork of gavamedia/deltafin with the ARGODRIVE storage work and benchmark package

TL;DR — Kimi K3 (2.8T-parameter MoE, 1.45 TB of expert weights) running on one M5 Max MacBook Pro with 128 GB, experts streamed from four SSDs.

  • 1.00 tok/s steady decode over a 512-token answer; 1.13 over 128; 0.96 on the public 17-token prompt (upstream reported 0.68).
  • The honest limit: a 512-token prompt takes ~6.3 minutes to its first token. Cause found (prefill re-reads each layer's experts 8×), fix planned, not built.
  • Useful findings: one drive gives ≈52% of four-drive speed, two ≈73%, three ≈90% — the slowest of each layer's 16 reads sets the pace, not total bandwidth.
  • Every number is one cold run with the exact prompt; per-run logs and placement manifests are in k3-public-bench/.
  • Fork of gavamedia/deltafin (MIT), who built the engine — see CREDITS.md. Instruments: ARGODRIVE.

ARGODRIVE Deltafin benchmarks — M5 Max, 128 GB, experts streamed from four SSDs

Measured 2026-09-08 with this fork's configuration of record (k3-public-bench/env.sh); every number is one cold run with the exact prompt, and the per-run logs are in k3-public-bench/results/.

test drafter off drafter on steady decode, 512 generated tokens 0.9232 tok/s 1.0015 tok/s steady decode, 128 generated tokens 0.9261 1.1252 17-token prompt from issue #15 (upstream reported 0.684 there), median of 3 — 0.9631 time to first token, 512-token prompt ≈376 s ≈375 s

Decode speed by number of drives

Per-drive draw under the engine vs standalone ceiling

Drive-count ladder on the same prompts: one drive ≈52% of the four-drive speed, two full mirrors ≈73%, three ≈90% (results/SCALING.md). Why prefill is slow and what fixes it: results/PREFILL.md. Definitions, identity scope and precision statement: k3-public-bench/README.md.


	____       _ _         __ _
	|  _ \  ___| | |_ __ _ / _(_)_ __
	| | | |/ _ \ | __/ _` | |_| | '_ \
	| |_| |  __/ | || (_| |  _| | | | |
	|____/ \___|_|\__\__,_|_| |_|_| |_|

Run the full, never-pruned, 2.8-trillion-parameter Kimi K3 on consumer hardware, as "fast" as possible

Deltafin is a single native binary that runs full Kimi K3. Nothing pruned. Nothing skipped. K3 decides every token.

All 16 experts, every single token. No shortcuts, no "close enough." It's exactly what Moonshot shipped.

The quality rule is simple: K3 itself decides every token, and nobody else. Small draft models are allowed to guess ahead (that's where much of the speed comes from), but K3 checks every guess, and nothing reaches you without its official sign-off.

Upstream benchmarks on an M1 Max laptop (gavamedia/deltafin, unchanged — not this fork's numbers)

  • 0.2901 token/s (3.447 s/token) — 1.9% higher throughput than last update

Upstream's historical M1 benchmarks:

  • 0.2847 token/s (August 2, 2026) — 7.0% higher throughput
  • 0.2660 token/s (July 30, 2026) — 102.9% higher throughput
  • 0.1311 token/s (July 28, 2026) — 829.8% higher throughput
  • 0.0141 token/s (July 27, 2026)

model platforms accelerators experts runtime license


Mission Statement

Pure raw uncut K3 quality, as fast as possible. Speed must never come from reducing model quality. Deltafin keeps all 16 routed experts and the full K3 target as the sole authority for every single token.

Our goal is to squeeze out every last drop of efficiency possible when running a huge model like K3, with all options on the table... except for reducing quality.


But Why?!?

Deltafin is not a product pitch. It is an experiment in how far consumer hardware can be pushed, and what we can learn by attempting something so challenging.

Kimi K3 targets infrastructure on the scale of 16 nodes and roughly 4.8 TB of aggregate VRAM. That means the full 2.8T parameters and the 1M-token context window, with the expert bank never pruned. On any home setup, this is an extreme constraint. Every 1% improvement is very hard-won. But each gain can teach something.

Research and exploration is the point. That is our mission. Not everything has to be a "minimum viable product" to impress venture capitalists. If Deltafin helps make frontier models usable on a $15,000 home setup, instead of a $2,000,000 infrastructure like Kimi recommends, we believe that is worthwhile progress on our self-hosted AI journey. Plus everything learned along the way could even benefit other projects in unexpected ways.

“We choose to run the full 2.8-trillion-parameter model locally, and do the other things, not because they are easy, but because they are hard.” — John F. Kennedy probably

Other projects appear to run full K3, somehow faster. But look closer: they've re-encoded K3's expert bank down to ~3 bits. Clever engineering toward a different goal: the smallest K3 that fits and is "close enough." Those weights are no longer the ones Moonshot released, and nobody, including them, has measured what those compromises cost.

Deltafin is the other experiment: every expert byte exactly as Moonshot shipped it, made as fast as physics allows.


1. New Installation

Deltafin installs almost everything it needs. See Requirements if you're missing anything.

# 1. Get it
git clone https://github.com/gavamedia/deltafin.git
cd deltafin

# 2. Build it
cargo build --locked --release

# 3. Download the FULL 1.7 TB K3 model to disk (optional, but fastest)
./target/release/deltafin setup --full

Or, if you don't have enough disk space:

# 3. Stream K3 as you use it (slower, but 215 GB to start) 
./target/release/deltafin setup --stream

setup --stream installs the resident model and fetches exact experts on-demand only, initially running far more slowly when routes have no local cache yet. As you build up your cache over time, this can be a way to save space, storing only the parts of the model you use, running entirely off cache on disk.

Default DSpark (and optional Qwen)

The normal setup includes Inferact's Kimi-K3-DSpark. It takes 6.635 GiB on disk and approximately 4.49 GiB when admitted at runtime. Deltafin avoids materializing DSpark's redundant copy of K3's embedding. Chat and server requests use DSpark automatically when beneficial; any failures, insufficient headroom, or bad live economics simply leaves full K3 running by itself.

Qwen is a separate add-on for faster raw text continuation:

# 4. Optionally install qwen later
./target/release/deltafin setup-qwen

Qwen speeds up raw completion only: the small models guess what comes next, K3 checks the guess, and you get identical output in less time. That helps code autocomplete and other /v1/completions traffic, plus deltafin run --prompt ... — one measured 17-token completion ran 2.7× faster with the same output IDs.

This adds 4.337 GiB on disk, and because Qwen will not improve chat speed, we make it an optional add-on. You can add it to a fresh install with deltafin setup --full --include-qwen, or add it later with the command above.

2. Upgrading

From the Deltafin folder:

./target/release/deltafin upgrade

upgrade gets what you need, and rebuilds the binary. Models, converted weights, and caches are left alone. It never re-runs setup or re-downloads K3.


NOTE: Upgrading from our old python version? Inspect it first:

git status --short

# ⬆️ Continue only when that returns nothing

git pull --ff-only
cargo build --locked --release
./target/release/deltafin upgrade

Continue only when git status --short is empty. If it lists files, preserve or commit that work yourself, rather than allowing an upgrade procedure to guess. Existing model data remains in the same repository-root directories.


upgrade needs a clean, non-diverged branch, and it remembers how the binary was built, so an NVIDIA/CUDA build stays a CUDA build rather than quietly falling back to CPU. Anything unexpected safely stops the upgrade.

upgrade ignores build environment variables — it reuses whatever the binary was already built with. So to switch configuration (CPU to CUDA, say, or a moved LibTorch tree), run cargo build --locked --release yourself once with the new variables set; see Requirements. That becomes the recorded setup, and later upgrades keep it.

3. Use from the command line

# Chat: apply K3's audited chat template and stop at the model's end marker.
./target/release/deltafin run --chat \
  --prompt "What are the three largest moons of Saturn?"

# Raw continuation: cap output because raw text has no chat end boundary.
./target/release/deltafin run \
  --prompt "The capital of France is" --max-new 17

# Add cumulative throughput and native transaction statistics.
./target/release/deltafin run \
  --prompt "The largest planet in our solar system is" --max-new 17 --stats

Without --stats, generated text streams normally instead of printing one diagnostic line per token. --max-new N limits new tokens; it does not alter the prompt or context. Chat output stops at K3's control boundary, while raw completion should normally use a bound.

Long conversations are far slower than short completions — prefill and cache grow with history, and startup prints the actual usable context bound.

4. Use through an OpenAI-compatible server

./target/release/deltafin serve --host 127.0.0.1 --port 8000
curl http://127.0.0.1:8000/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{"model":"deltafin-kimi-k3","stream":true,"messages":[{"role":"user","content":"Hello!"}]}'

The native server implements /v1/chat/completions, /v1/completions and /v1/models, including server-sent-event streaming. Point an OpenAI-compatible client at http://127.0.0.1:8000/v1 and use any non-empty local API key expected by that client.

The server implements a deliberately small, strictly-checked subset of the OpenAI API — text-only, one generation at a time, always greedy and reproducible — and refuses anything it cannot honor exactly with a normal OpenAI-shaped error instead of silently ignoring it. Growing chats automatically benefit from exact conversation-state reuse, draft-verified DSpark speedups and an exact-response memo; every accepted field, refusal rule and caching detail is in the server reference.

The default server response ceiling is one million tokens. Lower it with --max-tokens N when integrating clients, and raise client timeouts because full K3 responses are slow. Request JSON is bounded by --max-request-bytes (128 MiB by default). Keep the server on loopback unless you add your own authentication and network boundary.


Documentation

  • How the native runtime works — the one-binary in-process design: Rust core, C-ABI providers, router tracing, expert prefetch and native tokenization.
  • OpenAI-compatible server reference — exactly which API fields are accepted or refused, the automatic chat speedups, and the exact-response memo.
  • Health checks — read-only, network-free auditors that verify the runtime and each installed component after an install, upgrade or problem.
  • Native storage preparation — the default row-int8 resident spine that setup prepares for you, selecting the original BF16 explicitly, packing either into contiguous DFSP files, and lossless scale4 expert sidecars.
  • Performance reference — how to reproduce measurements with the benchmark harness, and the established M1 Max reference results.
  • Configuration — the few flags and environment variables that matter, and the quality guard behind them.
  • Supported platforms — what each host class runs (MPS/Metal, CUDA, native CPU) and the evidence status per platform.
  • Development reference material — why historical tools/*.py files remain in the tree as frozen reference material the native runtime never executes.

Credits

Deltafin exists because other people published amazing work:

  • Moonshot AI released K3's weights, architecture and readable model semantics.
  • Inferact released the unchanged K3-specific DSpark checkpoint; TorchSpec documents its training framework, and the vLLM team published K3 integration and recurrent/attention cache research. Deltafin's native runtime, verifier, scheduling and state transactions are its own.
  • GigaToken, by Marcel Rød, inspired Deltafin's automatic stable-order parallel tokenization path for large server histories.
  • Maurice Brown (trumb) contributed Linux, aarch64, x86 SIMD and NVIDIA findings plus DGX Spark measurements in pull request #2. Deltafin retained those findings behind reviewed capability and ABI gates.
  • colibri demonstrated aggressive MoE streaming and router-lookahead ideas. ds4 / DwarfStar provided especially clear prior art for exact expert streaming, cache ownership and correctness-first measurement.
  • Qwen supplies the optional 0.6B/1.7B proposal-only raw-completion models.
  • flash-linear-attention, PyTorch, llama.cpp/ggml, tiktoken and the broader local-model community supplied essential semantics and prior art.

Exact provenance and distribution boundaries are recorded in Third-party provenance and notices.

License

Deltafin's tracked project code is MIT. Kimi K3 weights, the DSpark checkpoint, optional Qwen checkpoints and all other third-party material retain their upstream terms. Deltafin is an independent project with no affiliation to Moonshot AI.

The Daily Front Page 20 of 31
Tuesday, September 8, 2026 The Daily Front No. #260908 — Mobile Graphics Works
article

Arm Mali G2-Ultra NX GPU: desktop-class mobile gameplay with AI-native graphics

by Re-Tails·▲ 81 points·62 comments·newsroom.arm.com ↗
Mali G2-Ultra NX brings dedicated neural accelerators, a new execution engine and a third-generation ray tracing unit.

Arm Mali G2-Ultra NX brings dedicated neural accelerators, a new execution engine and a third-generation ray tracing unit to help developers create desktop-class gaming experiences on mobile devices.

Overview of Arm Mali G2-Ultra NX highlighting its tightly integrated Neural Accelerator, new execution engine, third-generation Ray Tracing Unit, open software stack and ecosystem enablement.

Mali G2-Ultra NX combines dedicated neural acceleration, a new execution engine and third-generation ray tracing to enable AI-native graphics on mobile.

Demands from mobile users continue to grow. They want richer visuals, smoother frame rates, and increasingly immersive experiences. For developers the challenge is complex, as they aim to meet these expectations while staying within the strict power, thermal, and bandwidth limits of mobile devices.

Traditional rendering can deliver high image quality, but reaching desktop-class visual fidelity within mobile constraints demands greater efficiency. Neural graphics offers another path, using AI to reconstruct detail, generate intermediate frames, and refine images while reducing GPU demands and overall system workload.

However, bringing neural graphics into mainstream mobile experiences requires deep integration across the GPU architecture and developer ecosystem. With more than 14 billion Mali GPUs shipped to date, Arm has the scale and experience to make these capabilities accessible across the next generation of mobile devices.

That shift is at the heart of the Arm Mali G2-Ultra NX GPU. As the first AI-native Mali GPU, it integrates neural acceleration directly into the graphics pipeline, helping developers deliver richer, more responsive visual experiences within the mobile power envelope and bring neural graphics into production faster.

An AI-native GPU for neural graphics

Mali G2-Ultra NX is built for neural graphics from the ground up, with neural accelerators tightly integrated into the shader cores, allowing neural graphics workloads to run alongside graphics and compute workloads. The tight integration reuses the GPU memory system, coherent caches, and control structures, helping reduce data movement and improve efficiency across the pipeline.

Three neural technologies that are part of Mali G2-Ultra NX demonstrate what this enables:

  • Neural Super Sampling (NSS) reconstructs higher-resolution images from lower-resolution renders, reducing the amount of conventional rendering work needed to produce a high-quality final image.
  • Neural Frame Rate Upscaling (NFRU) generates intermediate frames to increase frame rates and deliver smoother gameplay.
  • Neural Super Sampling and Denoising (NSSD) combines neural upscaling with denoising to improve image quality in demanding ray-traced scenes with complex lighting and shadows.

Diagram of Neural Super Sampling showing low-resolution color, motion and depth inputs feeding a neural network, which reconstructs detail and produces an anti-aliased, upscaled output using temporal frame history.

NSS reconstructs higher-resolution images from lower-resolution inputs, combining temporal data and integrated anti-aliasing to improve image quality efficiently on mobile.

Diagram of Neural Frame Rate Upscaling showing two rendered frames, depth and motion inputs feeding a neural network that generates an intermediate AI frame, with hardware-accelerated optical flow supporting motion estimation.

NFRU uses motion, depth and rendered-frame data to generate high-quality intermediate frames for smoother mobile gameplay.

Neural Dawn, developed by Arm and Sumo Digital, showcases how developers can integrate NFRU and NSSD into a production game pipeline, delivering up to 4x higher performance efficiency and up to 70 percent lower external memory traffic compared with native rendering. These gains can make advanced techniques such as Unreal Engine MegaLights practical on mobile devices.

Explaining how NFRU and NSSD work together in the Neural Dawn game to improve mobile graphics.

By combining traditional rendering and neural graphics, developers can preserve image quality while reducing GPU and system-level workloads. For players, that means sharper images, smoother frame rates, richer scenes, and more responsive gameplay within the constraints of a mobile device.

A new execution engine increases the graphics budget

Alongside neural acceleration, Mali G2-Ultra NX introduces a new execution engine – the biggest Mali GPU instruction set architecture upgrade in seven generations – which is designed to handle increasingly complex graphics workloads while improving performance. It provides up to 2x more registers per warp, helping the GPU handle increasingly demanding scenes more efficiently. This gives developers greater headroom for richer visuals and more complex effects, while maintaining consistent frame rates.

Architecture diagram showing a Neural Accelerator integrated within the Mali G2-Ultra NX shader core alongside the execution engine, memory system and control structures, with support for INT8 and INT16 processing, optical flow acceleration and shared GPU resources.

Mali G2-Ultra NX integrates dedicated neural acceleration directly within the shader core, sharing GPU memory and control structures to improve efficiency for neural graphics workloads.

The architectural improvements translate into up to 24 percent higher benchmark performance and 14 percent higher non-AI gaming performance compared with the previous generation. Combining the architectural improvements with neural acceleration through NFRU supports longer mobile gaming sessions at up to 120 FPS for smoother, more responsive gameplay.

A side-by-side comparison showing how the higher frame rate with NFRU improves visual smoothness on Mali G2-Ultra NX

Third-generation ray tracing unit delivers richer lighting and shadows more efficiently

Mali G2-Ultra NX introduces Arm’s third generation of hardware ray tracing, supporting more complex lighting, shadows, reflections and geometry on mobile devices. The new GPU architecture delivers up to 13 percent lower DRAM traffic on leading ray tracing benchmarks, reducing memory pressure as scene complexity increases.

Meanwhile, hardware support for Opacity Micromaps helps handle intricate transparent geometry more efficiently, enabling richer foliage, more detailed fabrics, layered surfaces, and other complex scene elements on mobile devices. In one Arm demo, using Opacity Micromaps increased frame rates by 30 percent and reduced the ray tracing workload by up to 70 percent, demonstrating how the technology can make complex geometry more practical on mobile.

Real-time ray-traced shadows enabled by Opacity Micromaps to create a vivid environment in Moku’s Central Garden while maintaining performance

Neural graphics complement these advances in ray tracing. Combining new ray tracing hardware with neural technologies, such as NFRU, helps deliver richer visual detail more efficiently within mobile rendering constraints.

Giving developers a practical path to bring neural graphics to production faster

However, new graphics hardware only delivers value when developers can access it through familiar tools and workflows. Arm started preparing the software ecosystem two years before Mali G2-Ultra NX with Arm Neural Technology and an open development kit for neural graphics. These gave content developers, tools developers, and game engine providers, including Tencent Games Central Tech’s Magic Dawn and Unity China’s Tuanjie Engine, early access to neural techniques, so they could begin integrating them into existing workflows before the hardware arrived.

Neural graphics are opening up new possibilities for how we build the next generation of mobile games. Together with Arm, we are integrating Arm Neural Technologies into MagicDawn and jointly developing an NSSD technology demo with Arena Breakout Infinite. By connecting technology innovation with engines, developers and real game content, we are helping build an ecosystem that can bring more advanced graphics experiences to mobile players at scale.” Nathan Chen, Head of Technology, Tencent Games.

“Our collaboration with Arm is about more than bringing new technology into the engine — it’s about enabling the broader developer ecosystem to put that technology to work. By natively integrating Arm Neural Technology into Tuanjie Engine, we are helping developers more easily adopt advanced neural graphics and accelerate the path from technology innovation to real-world gaming across China’s gaming ecosystem.” Junbo Zhang, CEO, Unity China.

Tencent Games and Unity China explain how they are bringing neural graphics into familiar development workflows and helping move AI-native mobile graphics from technology to production.

Today, the Arm Neural Graphics Development Kit provides the software and tools needed to move that work toward production across established graphics APIs, game engines, and custom workflows. The developer stack includes NSS and NFRU, along with machine learning extensions for Vulkan and plug-ins for leading game engines, including Unreal Engine. An SDK supports custom-engine integration, while profiling, graphics optimization, training, and model-optimization tools help developers evaluate and fine-tune neural graphics workloads.

All these resources provide a path from experimentation to production without requiring studios to build bespoke research infrastructure or change the development workflows they already use.

As Keli Zhou, Engine Lead for Where Winds Meet, says, “Through our collaboration with Arm, we are integrating Arm Neural Technology into the Messiah Engine and bringing it into real-world game development. Where Winds Meet will be among the first games to bring Arm Neural Technology to players, marking an important step from engine-level integration to real-world deployment. Together, we are paving the way for broader developer adoption of neural graphics across the gaming ecosystem.”

The AI-native graphics foundation for Arm CSS for Mobile 2

Advanced mobile graphics require optimization across the entire system, not just the GPU. Within Arm CSS for Mobile 2, Mali G2-Ultra NX brings AI-native graphics together with compute, memory, software and implementation optimizations to help silicon partners manage increasingly demanding workloads within mobile power and thermal constraints.

Combined with an ecosystem ready to put these capabilities into production, this raises the ceiling for mobile graphics, enabling richer visuals, and smoother experiences delivered to billions of mobile devices.

The Daily Front Page 21 of 31
Tuesday, September 8, 2026 The Daily Front No. #260908 — The Web’s Fossil Record
article

Antiquated HTML Snippets and Artefacts

by patadune·▲ 229 points·81 comments·vale.rocks ↗
For every line of the HTML specification itself that has changed since the language’s inception, there are many more bits of HTML that have seen environmental changes.

With ever-changing devices, browsers, operating systems, form factors, specifications, personal preferences, tooling, and corporate interests, the web is in a constant state of flux. As a result, so is the HTML we write. For every line of the HTML specification itself that has changed since the language’s inception, there are many more bits of HTML that have seen what I’ll call ‘environmental’ changes. Adaptations to differing browsers, extensions, integrations, and systems which we find HTML existing in.

This article doesn’t cover once-specced but now obsolete bits of HTML but instead looks at all the snippets that have wormed their way into websites as result of, or in combat against, third-party integrations, browser competition, vendor extensions, and platform-specific hacks. The bits of HTML that were included for reasons, and which have been forgotten for present irrelevance. The snippets that live on only in the markup of sites from bygone eras and in the minds of those who fought during the browser wars.

X-UA-Compatible

<meta http-equiv="X-UA-Compatible" content="IE=edge">

In an age where browsers changed quickly and version-specific behaviours weren’t unheard of, it was deemed necessary to signal compatibility with specific versions of browsers with the X-UA-Compatible meta tag. Here the content attribute reads IE=edge, meaning that the target is the highest supported document mode in Internet Explorer (IE).

As outlined in the documentation, there are other values than edge. One can also supply a value of 5 through 11 to correspond with the associated Internet Explorer version, or EmulateIE7 through EmulateIE11 to enter into that version’s mode, but only if a valid DOCTYPE is declared, otherwise falling back to quirks mode.

<meta http-equiv="X-UA-Compatible" content="chrome=1">

Before Chrome took over the market, Google released an Internet Explorer plugin called Google Chrome Frame. Installable in Internet Explorer versions 6 through 9, it rendered websites in Chrome’s modified WebKit engine rather than Internet Explorer’s Trident engine if they had the tag set.

<meta http-equiv="X-UA-Compatible" content="requiresActiveX=true">

requiresActiveX was used to prompt Internet Explorer 10 to switch to desktop mode if in Metro mode, which didn’t support plugins.

ICBM Coordinate

<meta name="ICBM" content="-31.9548, 115.8602">

ICBM (an abbreviation of ‘intercontinental ballistic missile’) is hacker slang for one’s geographical location that dates back to Usenet days. The meta tag was used for identifying the location of a website’s subject and was referenced by the service GeoURL for plotting sites on the globe and helping people find sites of geographical proximity.

Conditional Comments

<!--[if gte IE 6]>
	<p>Only shown in Internet Explorer 6 and higher.</p>
<![endif]-->

In an era where browsers changed regularly and the vicious browser wars were in full force, it was sometimes important to target a specific browser or version of one. Microsoft handled this in Internet Explorer with specially formatted comments as seen in the snippet. There were a number of operators that could be used to further refine the condition:

Operator Description IE Denotes Internet Explorer. Can also include a version, such as IE 7. lt Less than. lte Less than or equal. gt Greater than. gte Greater than or equal. ! NOT. & AND. | OR. () Subexpression operator. true Always evaluates to true. false Always evaluates to false.

<!--&{navigator.appName == 'Netscape'};
	<p>Only shown in Netscape.</p>
-->

<!--&{navigator.platform == 'win95'};
    <p>Only shown in Netscape on Windows 95.</p>
-->

Netscape had its own way of handling conditional comments using JavaScript entities immediately following an opening comment tag. If evaluated to true, the comment would be included. If false, it’d just be a comment. Browsers other than Netscape would always treat it as just a comment.

Microsoft Smart Tags

<meta name="MSSmartTagsPreventParsing" content="TRUE">

Smart Tags was a system Microsoft introduced in Internet Explorer 6 which would automatically inject hyperlinks into pages.1 An example given by Microsoft was that ‘a Smart Tag might detect the names of major companies on the Web and tag them, allowing you to access stock quotes and company information.’. Site owners were outraged, and the MSSmartTagsPreventParsing meta tag was introduced to allow opting out. Microsoft later dropped it from Internet Explorer entirely.

PICS

<meta
	http-equiv="PICS-Label"
	content='
 (PICS-1.1 "http://www.gcf.org/v2.5"
    labels on "1994.11.05T08:15-0500"
           until "1995.12.31T23:59-0000"
           for "http://w3.org/PICS/Overview.html"
    ratings (suds 0.5 density 0 color/hue 1))
'
>

In the mid-90s the World Wide Web Consortium (W3C) attempted to establish a set of technical specifications for labelling internet content, called Platform for Internet Content Selection (PICS). PICS was incorporated into many filtering products, as well as Internet Explorer as of version 3. The shown meta tag could be included on a page to label it.

Rather unsurprisingly, many site owners just lied and miscategorised their sites. More still just never added the tags. The Protocol for Web Description Resources (POWDER) succeeded PICS but was also unsuccessful in its intentions. PICS was more or less completely dropped by 2003.

Interpage Transitions

<meta http-equiv="Page-Enter" content="revealTrans(Duration=2.0,Transition=12)">
<meta http-equiv="Page-Exit" content="revealTrans(Duration=2.0,Transition=12)">

Much like slide deck applications such as Microsoft PowerPoint or Apple Keynote have the ability to set transitions, Internet Explorer from version 4 to version 8 supported the above tags for transition style effects. This feature wouldn’t officially reach the web platform until the adoption of the View Transition API. Web developer Bramus has ported Internet Explorer’s transitions to the new View Transition API.

Frame Buster

<meta http-equiv="Window-target" content="_top">

Starting in Netscape Navigator in the mid-90s and then working its way over to Internet Explorer and other browsers, this non-standard meta tag would cause a page to break free of the parent page and open as a standalone browser window if it was loaded in a frame. It was popularly used by people to stop their pages from being embedded in frames on other sites. The _top Window-target specifies that a page should load in a top level window. In the modern age people use security headers to prevent pages from being embedded.

Baidu Page Transcoding

<meta name="applicable-device" content="pc,mobile">
<meta http-equiv="Cache-Control" content="no-siteapp,no-transform">

The Baidu Browser (百bǎi度dù浏liú览lǎn器qì) included a page ‘transcoding’ feature on mobile which was immensely unpopular. Websites would be routed through transcoder.baidu.com and modified. Layouts would be restructured, styles stripped, images compressed, branding removed, and adverts on the site would be removed while Baidu Union (百bǎi度dù联lián盟méng) adverts would be injected. It heavily disrupted sites. One of these tags would be added to sites from 2012 through to 2018 to opt out of this transcoding taking place.

Skype Toolbar

<meta name="skype_toolbar" content="skype_toolbar_parser_compatible">

There was an add-on for Internet Explorer, Chrome, and Firefox called ‘Skype Toolbar’ (later ‘Skype Click to Call’). Available from the mid-2000s until the mid-2010s, it was automatically installed alongside Skype and would attempt to detect phone numbers on pages and show a clickable icon. This icon allowed calling a number via Skype, adding it to one’s Skype address book, and assorted other Skype-related functionality. In inserting an icon, it would often cause page breakages, so developers would include the above tag to disable it on their sites.

The extension was also extremely unstable in Firefox, to the extent that Mozilla explicitly blocked it in 2011.

Image Toolbar

<meta http-equiv="imagetoolbar" content="no">

Internet Explorer 6 added an Image Toolbar, which would appear over images and show buttons for actions such as saving or sharing an image. Developers naturally hated this intrusion on their site and could disable the functionality with the above meta tag.

ClearType

<meta http-equiv="cleartype" content="on">

ClearType is Microsoft’s system for improving the legibility of typography. This tag could be included on pages to make typography look better on low-resolution devices, namely those running Pocket Internet Explorer / Internet Explorer Mobile.

AJAX Crawling

<meta name="fragment" content="!">

Google previously didn’t run JavaScript when crawling pages, so sites which had AJAX-based content would not be indexed. This meta tag told Google’s crawler to request a page with the _escaped_fragment_ query parameter, which was expected to return a HTML version of the page’s contents. Google began recommending this approach in 2009 and stopped advising it in 2015.

Directory

<meta name="robots" content="noodp, noydir">

Web directories used to be immensely popular as a way for people to find content on the internet. Two of the most popular were Yahoo Directory and DMoz (named for the URL directory.mozilla.org). The directories would sometimes write their own titles and descriptions for sites, which would be used by some search engines. To signal to search engines that they should use the metadata provided on the page and not the metadata from the directories, developers would include the above meta tag. Both directories are defunct now, making this snippet irrelevant.

Web App Name

<meta name="application-name" content="App Title">

application-name was used by various web app implementations across various browsers before Progressive Web Apps (PWAs) became established. If missing, systems usually fell back to <title>.

Internet Explorer Pinned Sites

<meta name="msapplication-tooltip" content="Visit our site!">
<meta name="msapplication-starturl" content="./">
<meta name="msapplication-window" content="width=1024;height=768">
<meta name="msapplication-navbutton-color" content="#FF3300">
<meta name="msapplication-task" content="name=Blog;action-uri=/blog;icon-uri=/favicon.ico">

Internet Explorer 9 and higher on Windows 7 had a function to pin sites to the taskbar. When pinned, sites could expose some site features on an operating system and browser level, almost like a modern progressive web app. msapplication-tooltip was text shown when hovering a site’s icon, msapplication-starturl set the root URL of the pinned site, msapplication-window set the initial size of the pinned site’s window when opened (minimum width of 800px and height of 600px), msapplication-navbutton-color set a custom colour for the browser chrome’s back and forward navigation buttons, and msapplication-task was used to define quick actions which could be accessed by secondary clicking a pinned site in the taskbar. For example, a quick link to a blog page.

Microsoft Web Apps

<meta name="msapplication-TileColor" content="#FF3300">
<meta name="msapplication-TileImage" content="/mstile-144x144.png">
<meta name="msapplication-square70x70logo" content="/images/small-tile.png">
<meta name="msapplication-square150x150logo" content="/images/medium-tile.png">
<meta name="msapplication-wide310x150logo" content="/images/wide-tile.png">
<meta name="msapplication-square310x310logo" content="/images/large-tile.png">
<meta name="msapplication-config" content="/browserconfig.xml">

Very similar to the previously mentioned Pinned Site functionality, these values were used for presentation. msapplication-TileColor, msapplication-TileImage, and the various msapplication-square* and msapplication-wide* attribute carrying tags were used for setting the theming of the tiles across the assorted versions of Windows, including the desktop UI, Metro UI, and Windows Mobile.

<meta name="msapplication-badge" value="frequency=360;polling-uri=https://contoso.com/BadgeUpdate.aspx">

msapplication-badge was used to define a web address to be polled for notifications which would display on a site’s Live Tile.

Chrome Pinned Sites

<meta name="application-url" content="https://vale.rocks">

Google Chrome had some pinned web app functionality which would use application-url for the start URL.

Safari Pinned Sites

<link rel="mask-icon" href="icon.svg" color="green">

Pinned Sites/Tabs in Safari allowed ‘users to keep their favorite websites open, running, and easily accessible.’. An SVG icon for the Pinned Tab could be declared alongside a colour which the icon would use.

CRX-less Web Apps

<link rel="chrome-application-definition" href="application_definition.json">

CRX-less Web Apps were applications in Google Chrome which allowed using an online manifest to describe a hosted app. The above meta tag would link to a JSON file with the manifest:

{
	"name": "Application",
	"description": "Description describing the app.",
	"launch_url": "index.html",
	"launch_container": "panel",
	"icons": {
		"128": "128.png"
	},
	"permissions": ["notifications"]
}

These CRX-less apps could be distributed on the Chrome Web Store with just the JSON manifest and an icon, with no need for the full .crx format Chrome uses for extensions. This type of app was later renamed to ‘Chrome Apps’, which are being phased out.

iOS Web Apps

<link rel="apple-touch-icon" href="/custom_icon.png">
<link rel="apple-touch-startup-image" href="/splash.png">
<meta name="apple-mobile-web-app-title" content="App Name">
<meta name="apple-mobile-web-app-capable" content="yes">
<meta name="apple-mobile-web-app-status-bar-style" content="black-translucent">

To allow site owners to customise the experience of their websites when added to a user’s homescreen on iOS, Apple introduced the above meta tags. The icon, startup image, and title are self-explanatory, while apple-mobile-web-app-capable disabled the browser chrome, and apple-mobile-web-app-status-bar-style allowed setting the operating system status bar to one of three possible values:

  • default - White background with black text and icons.
  • black - Black background with black text and icons.
  • black-translucent - Background colour sourced from the webpage background colour with white text and icons.

These meta tags are now rendered obsolete by modern PWA definitions.

<meta name="apple-touch-fullscreen" content="yes">

apple-touch-fullscreen was used in some early demos for iOS 2, however, Apple’s documentation only reflects apple-mobile-web-app-capable. Due to its presence in those early demos, apple-touch-fullscreen was popularised before the feature released, seemingly prompting Apple to make it an alias of apple-mobile-web-app-capable.

UC Browser and QQ Browser

<meta name="screen-orientation" content="portrait">
<meta name="x5-orientation" content="portrait">
<meta name="full-screen" content="yes">
<meta name="x5-fullscreen" content="true">

UC Browser and QQ Browser were particularly popular browsers in Asia through the early to mid-2010s due to their heavy data savings and optimisation for cheap devices. UC Browser would use screen-orientation to force a specific display orientation and full-screen to hide the browser’s chrome. QQ Browser would do the same, though with x5-orientation and x5-fullscreen. ‘X5’ in those meta names refer to Tencent’s X5 browser engine. The meta tags are no longer needed due to the standard fullscreen and orientation APIs and popularisation of more capable browsers.

PDA Optimised

<meta name="HandheldFriendly" content="true">
<meta name="MobileOptimized" content="320">

Before the proliferation of <meta name="viewport" content="width=device-width"> for making sites cleanly scale to small screens, these tags were widely used for mobile content. HandheldFriendly was created for AvantGo and then spread to the Blackberry Browser, while MobileOptimized was a Microsoft concoction for old versions of Windows Mobile.

Internet Explorer Tap Highlight

<meta name="msapplication-tap-highlight" content="no">

This snippet stops links from highlighting when tapped in Internet Explorer 11. It was similar to -webkit-tap-highlight-color, though unlike that property it was a boolean and didn’t allow changing the colour.

Revisit-After

<meta name="revisit-after" content="7 days">

A website called Vancouver Webpages ran a niche regional search system called ‘searchBC’ (the BC standing for ‘British Columbia’) which used the tag to know how often it should re-scrape a site. However, this was the only known use of the tag. There is no evidence that it has ever been used by any major search engine, including Google. Despite this, it somehow caught on like some sort of SEO myth and was perpetuated across the web.

Page Prerendering

<link rel="prerender" href="https://example.com">

Only fully supported in Chrome between versions 13 and 62 and never specced, this would fetch and process a page so it would be ready when users navigate to it. The functionality was replaced by the Speculation Rules API.

Accelerated Mobile Pages Version

<link rel="amphtml" href="https://example.com/amp.html">

Accelerated Mobile Pages (AMP) were restructured versions of pages optimised for faster content loading. This tag was used to link to an AMP version of a page, however, Accelerated Mobile Pages are no longer widely used, and Google no longer pushes for them, making the tag unneeded.

Internet Explorer Reading View

<meta name="IE_RM_OFF" content="true">

Starting with Internet Explorer 9, ‘Reading View’ was introduced. By pressing a button which would appear in the browser chrome, you could get a simplified version of a page suited for reading. Pages could opt out of having this button shown by including the aforementioned meta tag.

FrontPage

<!--webbot bot="Timestamp" s-type="EDITED" s-format="%B %d, %Y" startspan -->
<!--webbot bot="Timestamp" i-checksum="54321" endspan -->

<!--webbot bot="TableOfContents" s-component-title="Site Map" s-starting-point="index.htm" s-heading-level="3" startspan -->
<!--webbot bot="TableOfContents" endspan -->

<!--webbot bot="HitCounter" u-custom i-digits="6" startspan -->
<!--webbot bot="HitCounter" endspan -->

<!--webbot bot="PurpleText" preview="TODO: Replace product prices before launching sale page." startspan -->
<!--webbot bot="PurpleText" endspan -->

Microsoft FrontPage was a WYSIWYG website creation tool. Among other code, it would inject WebBot Components, such as those seen above. They were marked up as comments, so they wouldn’t appear on a page directly, but they were processed by FrontPage Server Extensions (FPSE) and used in FrontPage’s editor. Many of these WebBot Components existed, and developers could create their own using the FrontPage Software Developer’s Kit.

Really Simple Discovery

<link rel="EditURI" href="https://example.com/xmlrpc.php?rsd" type="application/rsd+xml" title="RSD">

Really Simple Discovery (RSD) was a system for exposing services for managing a blog to clients. The link in the head would direct clients to an XML document which would include endpoints to which tools such as Windows Live Writer, BlogJet, or MarsEdit could connect. It was implemented across platforms, including WordPress, Blogger, LiveJournal, TypePad, and MediaWiki. Modern publishing systems have generally deprecated it.

Pingbacks

<link rel="pingback" href="https://example.com/xmlrpc.php">

Pingbacks were a system where sites would alert other sites when they linked to them. The site receiving the ‘ping’ would check that a link was actually present on the site sending the ping, then make a note on the page saying something to the effect of ‘[Website] linked to this page’. They functioned somewhat similarly to how WebMentions do.

Unfortunately, pingbacks rather quickly became an avenue for spam and SEO-gaming, so fell out of fashion. Vulnerabilities allowing sites to be tied up in DDoS attacks also became widely exploited.

Share Partners

<link rel="image_src" href="https://example.com/post/image.png">
<link rel="audio_src" href="https://example.com/post/audio.mp4">
<link rel="video_src" href="https://example.com/post/video.mp4">

Before Facebook created Open Graph, they had a ‘Share Partners’ system for defining metadata to be used in embeds. Many additional sites adopted it as an ad hoc standard and implemented the tags. For example, Digg.

rel="image_src" would be referenced for an image to embed, rel="audio_src" would be referenced for audio to embed, and likewise rel="video_src" for video. There were a number of these meta tags which were supported to varying degrees.

Twitter Embeds

<meta name="twitter:card" content="summary">
<meta name="twitter:site" content="@account">
<meta name="twitter:creator" content="@writer_account">
<meta name="twitter:url" content="https://example.com/posts/hello-world">
<meta name="twitter:title" content="Title">
<meta name="twitter:description" content="Page description">
<meta name="twitter:image" content="https://example.com/hello-world-embed-image.png">
<meta name="twitter:image:alt" content="Alternative text for the image.">

These meta tags (and some other, less frequently used ones) were used on Twitter when generating link embeds. However, the documentation and card validator are no longer accessible (previously at https://dev.twitter.com/cards/getting-started and https://cards-dev.twitter.com/validator respectively). X falls back to the widely respected Open Graph meta tags, making the Twitter-specific declarations largely useless. They should be removed in favour of Open Graph tags. Further, they should be removed because X is an awful site with poor moderation that is owned by a man who publicly performed a Nazi Sieg Heil salute and has directly contributed to the rise of fascism in the United States of America and globally, among other horrors.

<meta name="twitter:dnt" content="on">

Twitter also had a Do Not Track meta tag for opting out of tracking when using Twitter for Websites widgets. Again, the documentation (previously at https://dev.twitter.com/web/overview/privacy) is now inaccessible, and it seems the tag is no longer respected.


This article isn’t comprehensive. These are only the more popular or notable non-standard bits. There are so many more obscure bits drifting through the shatters of cyberspace. For the curious mind, some more are listed – albeit without associated detail – on the WhatWG Wiki.

Footnotes

  1. Google tried something very similar in late-2024 with a feature called ‘Page Annotations’ in their Google app on iOS. It was hated and ultimately discontinued in March 2025.
The Daily Front Page 22 of 31
Tuesday, September 8, 2026 The Daily Front No. #260908 — The Security Deadline
article

We have a year to fix security everywhere

by saikatsg·▲ 299 points·337 comments·jyn.dev ↗
We need to fix vulnerabilities across the industry so that we aren't caught unawares.

GLM 5.3-flash released last week, and that means Project Glasswing and Daybreak are running out of time. Cheap models capable of dangerous hacking are now available to anyone, without the normal safeguards for refusing malicious actions. We need to fix vulnerabilities across the industry so that we aren't caught unawares. And for one of the first times in computing history, we have the ability to! We can use frontier LLMs that move faster than a human to find and fix these issues in the time we have left. The hard remaining part is deploying the fixes.

This probably sounds like nonsense words or hysterical overreacting to most people, so here's what that means:

  • "GLM" is a kind of LLM (AI). The GLM family is open-weight, which means anyone can download and run the models.
  • "flash" means that it is cheap and fast to run, compared to most "frontier" models. "cheap" is relative, but think around 5-15k USD in hardware to run it locally.
  • "frontier" here means that the LLM is "close to the frontier of what AI is currently able to achieve".
  • Project Glasswing and Daybreak are initiatives to use LLMs to fix security issues across the tech industry.
  • "malicious actions" includes things like hacking infrastructure and telling people how to build pipe bombs.

The rest of this post is about what makes me so sure this is an imminent threat, and what we can do in response.

GLM

GLM 5.3-flash can be downloaded and modified by anyone in the world.

The GLM ("General Language Model") family is developed by Z.ai Co. (formerly Zhipu AI), which is a Chinese AI lab. When the model is hosted by Z.ai, it comes with restrictions required by law:

GLM 5.3-flash refuses to tell me how to build a pipe bomb

Z.ai releases its models publicly on the internet ("open-weight" models). Once it does so, organizations such as DeAlignAI release "abliterated" models with their task refusals surgically removed. DealignAI says the abliterated model scores 0% on Harmbench-320, which tests whether models refuse to complete tasks about disinformation, cybercrime, biological weapons, and other illegal acts such as building a pipe bomb.

In other words, this model is willing to do basically anything.

Flash

GLM 5.3-flash is possible to run locally on stock consumer hardware.

"Flash" is mostly an advertising term—it's relative to other models, not a specific technical approach. Various people online have run benchmarks of GLM 5.3-flash locally. Here's one example showing around 20 tokens/second on a ~6k USD NVIDIA GPU.

On September 22, Apple is releasing the M5 Mac Studio with 256 GB of unified memory. "Unified memory" means it can be shared between the host operating system and the GPU. That's more than enough to run 5.3-flash, and it will probably get around 30 tokens/second once it releases. For 256 GB, the price starts at around $9,500.

Further improvements in software can get half-again the throughput through changes to the model decoder. If we extrapolate that to the M5, that would put the total throughput at around 45 tokens/second.

45 tokens/second is enough to write this snippet of code in 3 seconds:

⚠️ LLM generated code

from pathlib import Path
import hashlib

def digest(path: Path) -> str:
    hasher = hashlib.sha256()
    with path.open("rb") as file:
        while chunk := file.read(1024 * 1024):
            hasher.update(chunk)
    return hasher.hexdigest()

def main() -> None:
    import sys

    if len(sys.argv) < 2:
        raise SystemExit("usage: hash.py FILE...")

    for name in sys.argv[1:]:
        path = Path(name)
        try:
            print(f"{digest(path)}  {path}")
        except OSError as error:
            print(f"{path}: {error}", file=sys.stderr)

if __name__ == "__main__":
    main()

In other words, it's not just possible to run this model locally, it's possible to do so from an ordinary individual's savings, and use it round-the-clock at high speeds.

Frontier

GLM 5.3-flash is very close to the abilities of the best AIs we have made. The AIs we've made are already finding and exploiting real security vulnerabilities in the wild. The AIs we make in the future are going to get more and more capable.

GLM 5.3 scores 84.5% on CyberGym and 54.4% on ExploitBench. We don't have data for 5.3-flash directly, but it will probably be around the same or a bit lower. Abliterated models will be slightly lower again.

CyberGym measures real world vulnerabilities that have been found and patched by open source projects in the past. In other words, 84.5% of vulnerabilities in this representative sample would have been reproduced by GLM 5.3 just by looking at publicly available source code and a CVE description.

ExploitBench measures whether the model can actually use vulnerabilities to cause harm. It scores on a sliding scale that gives partial points for partial exploits, with the final step being arbitrary code execution.

For comparison, the leading ("frontier") model on ExploitBench is GPT-6 Astra (100%), with GPT-5.6 Sol as the runner-up with 78.5% 1. The leading model on CyberGym is ... GLM-5.3. The runner-up is GPT-5.6 Sol with 83.6%. OpenAI hasn't released numbers for Astra on CyberGym yet, but once they do it'll likely beat GLM 5.3.

cybersecurity evals visualizing the above stats

You might think these are just synthetic benchmarks, but security experts are reporting that they can no longer be competitive in security challenges without the assistance of an LLM.

We don't have many standard benchmarks for remote-code and reverse-engineering exploits, but we do have evidence of GPT 5.6-Sol exploiting infrastructure in the real world, without human involvement.

I think it is quite likely that people will be able to point GLM 5.3-flash at the open internet—real services, running real infrastructure—and it will be able and willing to find and exploit vulnerabilities.

This Is Bad

Together, this means:

  • Just about anyone can run GLM 5.3-flash if they have a bit of savings, continuously, day and night.
  • Just about anyone can use GLM 5.3-flash for just about any task, including to malicious ends.
  • GLM 5.3-flash is so good at those tasks that human involvement in those tasks can be negligible.

As a result, we are now in a world where cybersecurity attacks can be run in a for loop.

Now, the frontier US labs have been aware of this coming for a while and have been working on getting security patches out. Project Glasswing and Daybreak have been working with companies, foundations, governments, and NGOs across the tech industry to find and fix vulnerabilities using frontier models before this capability was open-sourced. They've done a lot of good, and I'm very glad that this was funded. Both have been sold as products after the initial funding, which feels a little bit sketchy at best, but they're at least giving out free credits to security organizations.

However, we are running out of time. And despite the good that Daybreak and Glasswing have done, the hard part is deployment, not fixing the bugs themselves. Critical systems often require physical access or carefully planned staged rollouts to avoid downtime, both of which delay deploying patches. It doesn't help to have a patched Linux kernel if your power grid is running Windows Server 2012.

There are some caveats: the 1.5 speedup might not be so high on GLM 5.3-flash; abliterated models might be worse on malicious tasks they weren't trained on; it might be hard to go from "break this" to an exploit without extensive human involvement. But those things are temporary and models keep getting better. Historically, GLM has lagged around 3-6 months behind OpenAI and Anthropic, and I think it's likely we'll see an Astra-level GLM model by this time next year. And when that happens, there's going to be a high risk of successful cybersecurity attacks on public or private infrastructure. We may be getting a lesson on brownouts sooner than we'd like.

In general, attackers are getting more capable faster than defenders are improving their posture. Even if models stop scaling so fast (which they currently show no sign of doing), it's only a matter of time before they get capable enough to start exploiting these vulns. We need to act now, the sooner the better.

What do we do?

Things are getting weird, and scary, very quickly. We need to act with urgency, not panic. Some things we can do:

Governments and regulatory agencies

Scanning with frontier models is relatively cheap and does not need major incentives. What does need incentives is deployment and remediation, and requiring organizations to look at their security practices in the first place. On the current policy trajectory, the biggest risk is a heap of untriaged warnings that never get fixed.

If you're in a position to make policy, the following would help: Fund security engineering, preferably with flexible grants that can be used for hiring or technology products as decided by the organization. Create mandates and incentives for improving security, especially for frequent penetration testing. Encourage using frontier models with human oversight for that pentesting. Encourage increased airgapping and discourage over-the-air updates: updates should be frequent but require physical access. For systems where airgapping isn't feasible, incentivize frequent, signed, and tested deployments. Penalize not investigating and revising security posture regularly, with increased penalties if a hack happens as a result. Require findings to be fixed within a risk-based deadline from discovery, with federal funding for the fixes. Both carrot and stick.

Some specific things that may be worth looking into:

  • Be especially sure to fund local governments and hospitals, which are unlikely to get this funding through other channels. EO 14409 is not enough because it's unfunded and voluntary.
  • For banks, extend DORA's TLPT in the EU and FTC/OCC/NCUA in the US. TLPT should increase the frequency and coverage of penetration testing. NCUA currently only suggests pentesting; upgrade it to a mandate. The FTC doesn't mandate pentesting if the financial institution has "continuous monitoring": it should be unconditionally mandated.
  • For power companies in the US, adopt guidelines similar to NERC Critical Infrastructure Protection at the state and local level, including for distribution systems and others that aren't currently regulated, not just for the highest-risk and largest systems. Create federal grants for implementing those guidelines. Extend NERC-CIP to require active testing for all systems, not just high-impact systems. Change NERC-CIP and the EU's NIS2 / Network Code on Cybersecurity to increase the frequency of required tests.
  • Telecoms in the US are currently high risk and have no unified mandatory cybersecurity risk standards. Create one and enforce it, using existing regulations for banks and power companies as a starting point.

Across the board, require security postures to be updated frequently. Mandating specific models or providers will become outdated as new models are released. This is a rapidly changing field and defenses that were effective 12 months ago may not be effective in a year as threat models (both senses) change. Mandate testing and accountability, not specific techniques.

Banning GLM 5.3-flash weights from being hosted anywhere in the US or Europe will be hardly any use in the short term, and no use at all in the long term. In the short-term, it will just pop up again on file-sharing sites; you'll have no more luck killing it than killing piracy. In the long-term, some other lab will release another model that's just as capable.

Blanket-banning access to Mythos or Astra will actively make things worse; it will remove defenders' most powerful tool at exactly the moment they need it most. Instead, restrict access to approved organizations and individuals, as frontier labs are already doing. This likely doesn't need new policy unless a lab shows signs of breaking ranks.

Banning the sale/export of new GPUs or large unified memory will extend the year-long window for a bit but won't help long-term. It can't do anything about existing hardware, and it will be massively unpopular. Memory in particular is hard to regulate because everything uses it, not just specialized AI systems.

In general, prioritize policies that address triaging and fixing security findings. Findings are getting very cheap; the fixes are not.

Companies and open source foundations

Take advantage of the (literal) billions of dollars that are flooding the industry to improve safety across the board. Hire as many security engineers as you can and fund existing maintainers. Instruct those engineers and existing maintainers to triage, design, review, backport, and deploy patches, not primarily to find vulnerabilities or write new code.

Use Astra, Mythos, and other frontier models for good, to find the risks before attackers do. Use structured prompts such as Google's Unsafe Rust Review; this is much more effective than telling them to look hard for bugs.

LLMs are good at writing patches, but not as one-off-prompts. Give them structured prompts and iterated self-review cycles until the LLM itself judges the patch to be high-quality. Whenever possible, get them to test their own fixes rather than guessing at whether their patch is effective. Only then consider it ready for a human to review.

Sandbox the agents themselves. The OpenAI-HuggingFace attack happened from a frontier lab testing a model; your own LLMs can easily cause incidents if you're careless. Restrict credentials to narrow scopes. If the issuing authority doesn't support scoped credentials, put a trusted interface in front of the services that adds the scope limitations itself; do not give agents direct access to broad credentials. Do not rely on filtering to only GET requests. Block requests at the firewall level and only expose a trusted list of domains. Filter endpoints using network proxies and trusted interfaces, not local configuration that the LLM can override. Preserve logs of every mutation or network request the agent makes.

Invest in formal verification, fuzzing and property testing, and memory-safe languages. LLMs are good at writing Lean and fuzz tests. I don't care whether you use Go or Rust but for the love of god please don't use C or C++ for new code.

Invest in triage: Record which versions of systems are affected, assign critical findings a human owner and a deadline, and create developer tooling to automatically update/close issues when they're fixed.

Invest in backport, release, and deployment machinery. Test upgrades and rollbacks, all the boring stuff. Developer tooling is cheap now; throw tokens at it so you can spend less human time on each patch: dependency-update automation, signed and reproducible releases, increased deployment speed. Engineers should be spending their time on coordinated disclosure and frequent releases, not on individual patches.

Deprecate old and insecure versions. There's a sea-change: you're in a rush, but the people depending on you are too. Use that as leverage to get them to upgrade. Where possible, write developer tooling that helps them automatically upgrade. Track whether people are upgrading and patching; if they aren't, invest more in tooling.

There are going to be a lot of patches and they will be exploited very quickly after the embargo lifts. Measure how long it takes end-to-end from a patch being reported to being deployed and adopted. Conduct campaigns to speed it up, focusing on the bottlenecks. Wherever possible, try to shorten embargo times: if you can find a flaw, an attacker probably can too, so the coordination window is much narrower than you're used to.

Invest in supply-chain security. Inventory your software and infrastructure dependencies. Inventory your own systems too: what versions are running in prod? what services do you run that don't have a maintainer? which of your systems are EOL? You finally have the ability to review all your dependencies without skimming; do so, prioritizing privileged and security-exposed dependencies first. LLMs are really good at finding bugs given the source code: use that to your advantage.

Invest in containment and recovery. Do not rely on a single firewall or VPN. Instead, use defense-in-depth: segment your networks, limit credential scope, test your backups, and run incident-response exercises. If possible, practice bringing up your systems from a cold start.

Pay attention to developments in frontier and open weight models. The more advanced that models get, the less time you have to patch and deploy.

Even if you don't think the threat described here is real, you're getting a once-in-a-lifetime opportunity to improve security for your projects and communities. Please take it.

Summary

We are living in interesting times. We can't hide our heads in the sand. We should act now, while there's still time.

Thank you to Manish Goregaokar and several others for their feedback on this post. Thank you to everyone who is working tirelessly to make Glasswing and Daybreak a reality. And a big fuck you to DeAlignAI, Z.ai, and everyone else who's been participating in this race to the bottom.

  1. depending who you ask, Z.ai and OpenAI disagree on exact numbers.
The Daily Front Page 23 of 31
Tuesday, September 8, 2026 The Daily Front No. #260908 — Capital, Genomes, and Office Suites
article

Mistral raises €3B

by kuberwastaken·▲ 814 points·570 comments·mistral.ai ↗

Mistral today announced that it has raised €3 billion in a Series D funding round at a post-money valuation of more than €21 billion, the largest equity fundraising round ever completed by a European technology company, three years after the company's launch.

Samsung Electronics led the round, joined by co-leads Scaleup Europe Fund, managed by EQT, and existing investor PSG Equity.

The round will significantly expand Mistral's frontier research, which is the foundation underpinning its infrastructure, products and sovereignty. While allowing Mistral to scale its compute capacity for training powerful models, it will help Mistral expand infrastructure and accelerate its commercial growth and international footprint. The company now operates across 20 countries and supports 125+ global enterprises’ mission-critical AI transformation, including Airbus, ASML, and HSBC.

During the first wave of generative AI, the central question was who could build the most powerful model. Organizations and governments are now asking a different one: how to harness the power of AI for their mission-critical needs without surrendering control over the infrastructure and intelligence loop. Demand for that combination of performance with control, choice and independence is growing internationally, as enterprises and governments weigh the long-term technology dependencies, data governance requirements and deployment choices that come with any AI investment.

Mistral is the only AI company in the world building the full stack required to answer that question: open-weight models, the infrastructure and the compute capacity they run on, and the products that bring them into production; ensuring that customers are never locked into a single vendor's roadmap, pricing or availability.

Mistral’s full-stack and open approach also allows organizations to build on it without exposing their most valuable data, workflows and institutional knowledge to anyone outside their own walls. That's what makes Mistral’s stack the sovereign AI layer, meaning retaining control across four dimensions: data that stays inside the organization's boundaries, models that are controllable and customizable, compute that is private and predictable, and systems in production that are fully controllable and auditable.

The round reflects a strategic endorsement from investors across Europe, Asia and North America. It brings together a world-class syndicate of strategic and financial investors, including global technology leaders, growth investors and existing shareholders that have supported Mistral’s development to date.

With a Series C led by ASML and a Series D led by Samsung Electronics, Mistral has attracted backing from companies at the forefront of advanced manufacturing, engineering and industrial technology. This support reflects growing confidence that Mistral's approach can help organisations deploy state-of-the-art AI inside complex real-world environments while maintaining control over their data, infrastructure and knowledge.

Advent, funds and accounts managed by BlackRock as well as the Grand Duchy of Luxembourg joined the round as new investors. Existing investors a16z, ASML, Belfius, BNP Paribas CIB, Bpifrance, Carmignac, DST Global, Eurazeo, General Catalyst, Headline, Hillspire, Index Ventures, Korelya Capital, Lightspeed, NVIDIA, Phoenix Court's Solar fund (home to LocalGlobe, Latitude and Solar) and Salesforce Ventures participated in the round.

The Daily Front Page 24 of 31
Tuesday, September 8, 2026 The Daily Front No. #260908 — Capital, Genomes, and Office Suites
article

AlphaGenome Atlas: a high-resolution map of human DNA

by utiiiD·▲ 514 points·118 comments·blog.google ↗

Wavy pink and lavender pillars with glowing orange columns in the center.

The human genome is made of about 3 billion base pairs of DNA — but much of it remains a mystery. Scientists understand the 2% of the human genome that codes for proteins relatively well, but have only limited knowledge of the remaining 98%. Our AlphaGenome model has already shown how single changes in these non-coding DNA regions can disrupt molecular processes like protein production, but the bigger picture remained unclear.

Today, we're introducing AlphaGenome Atlas, a database that predicts the effects of every possible single nucleotide variant in the human genome. We used the AlphaGenome AI model to pre-calculate the regulatory impact of all 9 billion single-letter genetic changes, resulting in a massive, 1-petabyte dataset. Our new Atlas helps scientists rapidly query this vast information.

To help researchers rapidly navigate this, the Atlas introduces the AlphaGenome Variant Impact (AVI) score. This single, easy-to-use score combines predictions for both coding and non-coding regions, allowing researchers to quickly prioritize the most promising avenues for research without sifting through thousands of data points.

Empowering researchers to solve biological mysteries

AlphaGenome Atlas is already acting as a powerful augmentation partner for the scientific community, accelerating research in areas like:

  • Rare genomic variations: At the Broad Institute, Laura Covill and her team used the AVI score to prioritize variants for unsolved rare disease research. The tool highlighted a critical variant in the DNM1 gene, predicting that it created an incorrect splice site. This provided crucial supporting evidence to successfully solve the case.
  • Complex traits: Identifying rare, non-coding variants linked to complex traits is difficult due to statistical noise. Dr. Gareth Hawkes applied AlphaGenome Atlas to data from 54,000+ UK Biobank participants. By grouping variants based on predicted molecular effects, he uncovered 22% more non-coding genetic associations. Focusing on the top 1% of impactful variants, he identified 19 genetic regions linked to body mass index (BMI), directing the next stage of targeted research.

Opening access to researchers and biologists worldwide

AlphaGenome Atlas is available today through an intuitive website portal that requires zero coding skills, democratizing access for clinical researchers and biologists worldwide. This is part of our ongoing commitment to accelerate genomic discovery and science, for everyone.

AlphaGenome Atlas provides grounded genomic insights that will accelerate the pace of biological discovery.

Read more on the Google DeepMind blog.

The Daily Front Page 25 of 31
Tuesday, September 8, 2026 The Daily Front No. #260908 — Capital, Genomes, and Office Suites
article

LibreOffice breaks download records after declaring it has no AI features

by rpgbr·▲ 658 points·218 comments·manualdousuario.net ↗

LibreOffice 26.8, released on August 26th, became the software’s most popular update. LibreOffice is a free alternative to Microsoft Office that’s been around for nearly two decades. In one week, the installer was downloaded more than 1 million times — not counting updates via Linux distribution repositories.

Someone might say the improvements to the writing system and typography are a hit among NGOs, government agencies, and offices that use LibreOffice. I’m putting my money on a “non-feature,” though: the statement that LibreOffice doesn’t come with generative AI features due to the technology’s (lack of) privacy.

A day after celebrating the download record and the wide media coverage, The Document Foundation (TDF), which manages the software, spelled out its position in a post titled “Yes, no AI is now a feature”.

The article, signed by Italo Vignoli, says TDF “does not reject artificial intelligence out of hand,” but that the technology doesn’t yet meet an exhaustive list of principles to be included by default. Those principles are:

  • User-controlled execution.
  • No content may leave the computer without authorization.
  • No telemetry of any kind.
  • No dependency on a single vendor.
  • No compromises on [file] format.
  • Entirely optional.

For now, TDF recommends community plugins to integrate AI into LibreOffice.

TDF’s distinctive stance “has almost nothing to do with the technology,” the article continues, before taking a jab at rivals that dove headfirst into AI. For the foundation, “an AI assistant justifies a price increase and strengthens the case for keeping all documents within its own infrastructure,” referring to companies that rely on selling subscriptions.

The Document Foundation’s detailed explanation may disappoint anyone who read the announcement at the LibreOffice 26.8 launch as an anti-AI manifesto.

Whatever the reasons, the LibreOffice 26.8 case answers a question I and others have had since 2023: if the whole market now offers AI baked in software of every kind, is there really no room for AI-free alternatives to stand out? Apparently, there is.

The Daily Front Page 26 of 31
Tuesday, September 8, 2026 The Daily Front No. #260908 — A Handy Agent Add-On
repository

I-have-ADHD: A skill to stop coding agents from burying the answer

by domhudson·▲ 363 points·272 comments·github.com ↗
★ 31,190⑂ 1,884 forks Python

A skill to stop your coding agent from burying the answer. ADHD-friendly output.

i-have-adhd

ADHD-friendly outputs. No ADHD diagnosis needed!

Install

Copy/paste into your CLI prompt:

Install the i-have-adhd skill/plugin from https://github.com/ayghri/i-have-adhd, refer to the repo's AGENTS.md for instructions.

Or 🔗 check the installation instructions.

What it does

A skill for your coding assistant that stops it from burying the answer. Action first. Steps numbered. No "Hope this helps!"

What changes

Before

Great question! Let me think about this. Your auth flow has a few moving pieces: the middleware, the token verification, and the cookie handling. Looking at src/auth.ts, the verifyToken function (around lines 42-58) seems to be using an older jsonwebtoken API. One approach would be to update the package and rewrite that function. After making the change, you'd want to run the auth tests to confirm nothing breaks. By the way, you might also want to look at your dependency versions overall. Hope this helps! Let me know if you want to dig deeper.

After

Run npm install jsonwebtoken@latest, then edit src/auth.ts:42.

  1. Open src/auth.ts
  2. Replace verifyToken (lines 42–58) with the snippet below
  3. Run npm test -- auth.spec.ts

Next: paste the first failing line if any test fails.

The rules

10 rules. Full text in SKILL.md.

  1. Lead with the next action.
  2. Number multi-step tasks.
  3. End with one concrete next step.
  4. Suppress tangents.
  5. Restate state every turn.
  6. Specific time estimates (minutes, not "a bit").
  7. Make wins visible.
  8. Matter-of-fact errors.
  9. Cap lists at 5 items.
  10. No preamble. No recap. No closers.

Tune it

Fork, edit skills/i-have-adhd/SKILL.md, then swap your copy in:

claude plugin uninstall i-have-adhd            # drop the upstream copy first:
claude plugin marketplace remove i-have-adhd   # fork and upstream share both names
claude plugin marketplace add <your-username>/i-have-adhd
claude plugin install i-have-adhd@i-have-adhd

Restart Claude Code, then re-invoke /i-have-adhd.

Credits

Loosely based on The Adult ADHD Tool Kit by J. Russell Ramsay and Anthony L. Rostain. Adapted for how an LLM should respond, not how a human should organize their day.

License

MIT.

Star ⭐ if it saved you one scroll past one "Great question!"

The Daily Front Page 27 of 31
Tuesday, September 8, 2026 The Daily Front No. #260908 — Screens, Spies, and Spin
article

LG TVs caught spying even when offline or on standby

by sbulaev·▲ 509 points·303 comments·theverge.com ↗

Gamers Nexus gives you another reason to never connect your smart TV to the internet.

258078_LG_G5_OLED_TV_JHiggins_0007

The LG G5 OLED is one of the TVs tested by Gamers Nexus.

Photo by John Higgins / The Verge

LG smart TVs are almost constantly logging and uploading data about owners and their homes, even when offline or on standby mode, according to a new report from YouTube channel Gamers Nexus. The company’s TV sets scan Wi-Fi networks for nearby devices, record audio logs through their microphones, and use audio and video sampling to recognize exactly what you’re watching from across the TV inputs.

Gamers Nexus partnered with fellow YouTubers Level1Techs and independent security researchers for the investigation, which involved testing retail LG OLEDs. Packet captures showed the TVs scanning the local area network for nearby hardware like phones or smartwatches, as well as logging location data and details of nearby Wi-Fi networks, and feeding the information back to LG Ad Solutions. Perhaps more concerningly, the TVs were capable of recording microphone audio when in standby; this continued even after the TV was disconnected from the internet, with audio files stored offline and uploaded once a connection was restored.

The final piece of the puzzle is Automatic Content Recognition (ACR), tech that uses audio and video data to figure out what you’re actually watching, examining the apps you use through the TV’s own smart interface and the content you connect via HDMI and other ports. ACR isn’t just an LG issue though: a new video from RTINGS digs into how ACR is used by almost every smart TV manufacturer to figure out what you watch.

This isn’t LG’s first brush with Gamers Nexus. In July this year, the channel flagged that some LG monitors were automatically installing an app that collected device data and ran pop-up ads for LG’s other apps and even McAfee antivirus. The public pressure from that video prompted an intervention from Microsoft, and ultimately led to LG disabling the pop-up ads.

The Daily Front Page 28 of 31
Tuesday, September 8, 2026 The Daily Front No. #260908 — Screens, Spies, and Spin
article

Paramount Caught Using 'Astroturf' Group to Drum Up Fake Support for Merger

by hn_acker·▲ 355 points·117 comments·techdirt.com ↗

We’ve well established how the $111 billion Paramount and Warner Brothers merger is terrible for labor, creatives, consumers, and markets. But the kind of behaviors Paramount leadership have been engaged in to sell the deal also give a pretty clear indication what kind of company we’re dealing with, and why they shouldn’t be allowed to control an even bigger slice of U.S. media.

Recent desperate moves by the company include accusing all deal critics of being “antisemitic,” threatening to take Paramount out of California if states attempt to enforce U.S. antitrust laws, and running press leak and PR campaigns (in close collaboration with their Trump allies) making all sorts of patently false claims as to why more media consolidation is great for America.

Paramount’s also been caught using an “astroturf” — or fake grass roots organization — to try and make it appear that the giant merger has more public support than it does.

The Intercept notes that the company has been employing the use of a nonprofit named Neighbors for Strong Communities to bombard California residents with text messages urging them to pressure on California AG Rob Bonta to drop his antitrust lawsuit against the company, claiming that enforcing antitrust law will be bad for Californians:

“In text messages to Californians that went out earlier this week, Neighbors for Strong Communities asked recipients to send Bonta messages raising the concern that his opposition to the merger will cost the state thousands of jobs — because of Ellison’s reported threat to move Paramount to Texas.”

In reality, mergers like this — particularly when involving Warner Brothers — have a long history of resulting in mass layoffs, higher prices, and lower-quality product as the merged company tries to pay down debt from the deal. There’s also nothing stopping Ellison from offshoring film and TV production if the deal is approved, continuing an existing trend for Hollywood.

A coalition of consumer rights groups under the banner of the Committee For The First Amendment (which actually defends consumers and discloses its funding) say the astroturf group was only formed three months ago by industry lobbyists and refuses to disclose its donors:

“The sender organization’s website is less than three months old. It discloses no founders, board members, staff, or funders, and the organization does not appear in ProPublica’s nonprofit database. Its newly updated address (1902-A Lincoln Blvd, Suite 1314, Santa Monica, CA) is a private PO Box at a UPS store. These are classic hallmarks of a corporate-backed astroturf campaign designed to manufacture the appearance of grassroots opposition.”

As is always the case with these groups, there’s usually no discernable paper trail, allowing both Paramount and Neighbors for Strong Communities to confidently insist the entire effort is authentic. And when you have to covertly pay organizations to support your giant merger, you just know the giant merger is great for everyone involved.

The Daily Front Page 29 of 31
Tuesday, September 8, 2026 The Daily Front No. #260908 — Also on the Front Page
The Daily Front Page 30 of 31
Tuesday, September 8, 2026 The Daily Front No. #260908 — Colophon

That's the Front for Today

Issue No. #260908 — Tuesday, September 8, 2026 — went to press 2026-09-09 at 04:57 UTC.

About This Magazine

The Daily Front is a daily digital magazine assembled from the stories that reached the front page of Hacker News on Tuesday, September 8, 2026. Headlines, points, and comment counts are recorded as they stood at press time. All articles remain the property of their original authors — every piece links back to its source and its discussion thread.

How It Was Made

Fetched, cleaned, and typeset by an automated pipeline. An editor model laid out the pages and chose the highlights; a second read a handful of the day's stories and briefed the cover illustrator — 32 model calls and 281k tokens in total. Set in Jacquard 12, Playfair Display, Source Serif 4, and IBM Plex Mono, all served via Google Fonts under the SIL Open Font License.

The Cover

The cover illustration was commissioned with this prompt:

Inside a vast observatory-like laboratory, a towering double helix rises from a dark circular basin, its pink and lavender strands twisting into a violent fluid vortex. Thousands of tiny glowing orange bridges flare along the central rungs, while a researcher in a plain coat steers a narrow boat through the swirling base, collecting luminous fragments in glass vessels. Around the basin, suspended transparent panels show branching molecular pathways and scattered points linked by fine threads, all buffeted by the same turbulent current.

Brutalist 3D editorial render in matte clay, with crisp ambient occlusion and asymmetric gallery lighting: stage the vast observatory-like laboratory as a monumental slate-blue and bone-white space, centering a dark circular basin beneath a towering double helix whose dusty rose and muted lavender strands twist into a violent ultramarine fluid vortex; ignite thousands of tiny amber-orange bridges along its central rungs, and show a plain-coated researcher steering a narrow boat through the turbulent base while collecting luminous fragments in clear glass vessels. Suspend transparent panels around the basin, etched with branching molecular pathways and scattered points connected by fine threads, all visibly buffeted by the same current; use hard-edged sculptural forms, sparse negative space, and selective electric-cyan reflections against the deep navy ground.

Absolutely no text, letters, numbers, readable symbols, or logos anywhere in the image.

Production Ledger

StageModelCallsTokens InTokens Out
extractgpt-5.6-luna 28 165,377 86,150
layoutgpt-5.6-terra 1 19,507 2,364
covergpt-5.6-luna 2 1,539 360
covergpt-image-2 1 278 5,488

The Publisher

Published by Johnny.

Support the Press

If The Daily Front brightens your morning, consider supporting its publisher.

Credits & Contact

All content — articles, posts, comments, and the images within them — belongs to its original authors and is reproduced here to point readers back to the source. Full credit goes to those creators; every item links to its original and its Hacker News discussion.

If you are an author and would like your content removed from an issue, write to hi@johnnys.page and it will be taken down.

Feedback is always welcome at the same address: hi@johnnys.page.

Credit where credit is due.

Every page of this issue began as someone else's work — these are the original sources, linked in full.

  1. On the Navier–Stokes Millennium Prize Problem by tedsanders — openai.com·HN discussion ↗
  2. Navier-Stokes – Tristan Buckmaster [pdf] by procedurecall — cims.nyu.edu·HN discussion ↗
  3. There's a new "Google Jail" for independent wikis by pizzaiolo — weirdgloop.org·HN discussion ↗
  4. I've factored the RSA keys of a Certificate Authority from the 90s by ahlCVA — mcpherrin.ca·HN discussion ↗
  5. Jellyfin 12.0 by 0xC0ncord — jellyfin.org·HN discussion ↗
  6. We built our house for LAN parties (2024) by fittingopposite — lanparty.house·HN discussion ↗
  7. TALA Is Open-Source by alixanderwang — d2lang.com·HN discussion ↗
  8. Show HN: Copperhead – Cursor for circuit boards by animeshchouhan — copperhead.sh·HN discussion ↗
  9. Benchmarking Qwen3.8 27B quantizations: 4-bit holds up, 1-bit collapses by stared — quesma.com·HN discussion ↗
  10. Among European Companies That Use a CDN, Nearly 9 in 10 Use Cloudflare by adulion — ciphercue.com·HN discussion ↗
  11. The Helicopter with Radioactive Blades by zdw — hackaday.com·HN discussion ↗
  12. John Margolies' photographs of roadside America by duck — publicdomainreview.org·HN discussion ↗
  13. Getting your hands dirty is good for you by HatchedLake721 — bbc.com·HN discussion ↗
  14. The two Christian saints who are the Buddha by surprisetalk — signoregalilei.com·HN discussion ↗
  15. Emacs Bedrock 2.0 by ashton314 — lambdaland.org·HN discussion ↗
  16. How well do agents use test/verification techniques? by vinhnx — danluu.com·HN discussion ↗
  17. Show HN: LLM Attention Visualization by ifz — ishamf.dev·HN discussion ↗
  18. Kimi K3 (2.8T) at 1 token/s on a MacBook Pro, streamed from four SSDs by Argonautlabs — github.com·HN discussion ↗
  19. Arm Mali G2-Ultra NX GPU: desktop-class mobile gameplay with AI-native graphics by Re-Tails — newsroom.arm.com·HN discussion ↗
  20. Antiquated HTML Snippets and Artefacts by patadune — vale.rocks·HN discussion ↗
  21. We have a year to fix security everywhere by saikatsg — jyn.dev·HN discussion ↗
  22. Mistral raises €3B by kuberwastaken — mistral.ai·HN discussion ↗
  23. AlphaGenome Atlas: a high-resolution map of human DNA by utiiiD — blog.google·HN discussion ↗
  24. LibreOffice breaks download records after declaring it has no AI features by rpgbr — manualdousuario.net·HN discussion ↗
  25. I-have-ADHD: A skill to stop coding agents from burying the answer by domhudson — github.com·HN discussion ↗
  26. LG TVs caught spying even when offline or on standby by sbulaev — theverge.com·HN discussion ↗
  27. Paramount Caught Using 'Astroturf' Group to Drum Up Fake Support for Merger by hn_acker — techdirt.com·HN discussion ↗
  28. DaVinci Resolve 21.1 by tosh — blackmagicdesign.com·HN discussion ↗
  29. Muse – Meta’s personal AI agent by yks — ai.meta.com·HN discussion ↗
  30. The 92-Year-Old Mathematician and the Teenage Apprentice by robinhouston — nytimes.com·HN discussion ↗

Browse all issues in the archive →