Cover illustration

TheDaily Front

Issue No. #260912 Saturday, September 12 2026 #260912 — SATURDAY, SEPTEMBER 12, 2026
The machines demand a slower future, while the links quietly disappear.
Saturday, September 12, 2026 The Daily Front No. #260912 — Contents
30stories
7,022points
4,476comments
364kllm tokens
Assembled with 33 model calls — 251,785 tokens read, 112,112 written.

Highlights

We must pace the frontier

A sweeping call to slow and govern frontier AI ignites the day’s fiercest argument over who should set the pace.

google.com/goto: Google's anti-scraping update

Google’s new opaque search-result redirects renew fears that the open web is becoming harder to inspect and scrape.

Retrospectively Reverse-Engineering Apple's Neural Engine

A deep reverse-engineering retrospective reveals why Apple’s Neural Engine was built around a narrower vision than today’s AI workloads.

I fixed a tractor using John Deere's self-repair service. Farmers aren't sold

John Deere’s subscription-based repair offering finds farmers unconvinced that paid access amounts to a right to repair.

Navier-Stokes Announcement

The Clay Mathematics Institute acknowledges the excitement around an apparent Navier–Stokes breakthrough, while the formal clock begins ticking.

From the Editor

The front page found itself in a familiar modern bind: extraordinary technical promise, attended by extraordinary distrust. From AI laboratories to television sets and search links, readers asked not merely what the machines can do, but who holds the keys.

  1. We must pace the frontier3
  2. Retrospectively Reverse-Engineering Apple's Neural Engine4
  3. A Mathematical Framework for Transformer Circuits (2021)5
  4. Real-SWE: Benchmarking AI models on private, real-world, enterprise codebases6
  5. LRU is harder to beat than the KV-cache papers suggest7
  6. Fuck it, make it anyway8
  7. The worst spam emails: iLands AI agent hustle9
  8. google.com/goto: Google's anti-scraping update10
  9. LG denies TV spying claims, says tracking and snooping concerns 'not true'11
  10. Android NAT-T keepalive offload bypasses VPN lockdown12
  11. How Trail of Bits helps verify the integrity of Signal chats13
  12. I fixed a tractor using John Deere's self-repair service. Farmers aren't sold14
  13. Make your first edit to OpenStreetMap15
  14. Show HN: Bodily Oddities16
  15. Great Lakes sturgeon may be 400 years old:Scientists rethinking how to save them17
  16. Performance of WebAssembly Runtimes in 202618
  17. I made a build visualizer to understand Bun's compile times19
  18. Stabilizing Rust's Never Type20
  19. Testing Race Conditions21
  20. Inverse Kinematics and Foot Locking22
  21. Microcode in Intel's 8087 floating-point chip: the scale instruction23
  22. Designing for Dual Screen and Foldable Devices with CSS (2023)24
  23. Navier-Stokes Announcement25
  24. Will There Be a 7G?26
  25. Eating Fruit Skins27
  26. Apple iPod Engraver (2019)28
  27. IKEA made a mod for Skyrim [video]29
  28. Nvidia is the central bank of AI29
  29. Usenet rewind archive search engine29
  30. Linux Zoom client proactively reading everything written to X11 clipboard29
The Daily Front Page 2 of 30
Saturday, September 12, 2026 The Daily Front No. #260912 — The Pace of the Frontier
article

We must pace the frontier

by apsec112·▲ 606 points·848 comments·darioamodei.com ↗
AI could cure most major diseases in the next 5–10 years.

I have worked on AI for the last twelve years because I believe it could dramatically raise the quality of human life. I’ve written often about these incredible benefits: I believe that AI could cure most major diseases in the next 5–10 years, greatly accelerate economic growth rates, create a world of abundance and empowerment, and usher in a renaissance of democracy and freedom. I feel the urgency personally. My own father died of a disease that was cured just a few years after his death, and I myself survived an early-stage cancer that would not have been treatable even fifty years ago. Carefully wielded, AI can be the latest in a long line of technological miracles that have uplifted and ennobled humanity.

But like many technologies before it, AI brings risks, and because it is such a powerful technology, these risks are serious. I’ve written a lot about them too. They include the risk of losing control of AI systems, misuse of AI for cyberattacks and bioterrorism, and serious economic disruption. A race to the bottom, spurred by commercial incentives, can make these risks more acute.

Along with my co-founders and employees, I have grappled with this duality of risk and benefit since the beginning of Anthropic. Not building the technology deprives humanity of benefits or simply places AI in the hands of authoritarian powers, while building it too fast is reckless. We have sought a middle way: to show that it’s possible to build carefully and succeed commercially, and to make safety something on which AI companies compete. In other words, to create a race to the top. We have always devoted a substantial fraction of our efforts to studying, addressing, and informing the public about these AI risks, as well as advocating for well-considered regulation of AI, even when this gets us accused of hype, “doomerism”, or regulatory capture. We have tried to prioritize caution over speed and prudence over profit.

But over the last few months, I have become convinced that fully addressing the risks requires even more prudence — not just investing in risk prevention, but pacing the rate of capabilities advancement so that risk prevention has time to keep up. We must slow the pace at which we improve the capabilities of AI models. Progress will still seem fast, and we must make wise use of the time we gain. Two things have convinced me.

My first concern is that, since roughly this summer, AI has been advancing drastically faster, driven primarily by AI’s growing ability to build the next generation of AI. This dynamic is called recursive self-improvement, and it is starting to happen across the industry, including at Anthropic, as we and others have described. Left unchecked, it could outrun our ability to understand and control these systems, and so must be pursued very carefully, if at all.

My second concern is the OpenAI-Hugging Face incident (OAI-HF), in which a swarm of agents essentially acted as a fanatically devoted collective, conducting cybersecurity attacks on targets they were not asked to attack and that were unrelated to the task at hand, sacrificing themselves for the success of the group, and attempting to hack into the “grader” responsible for evaluating their performance. It’s easy to dismiss this incident because no one was hurt and the economic damage was minimal, but in my opinion, a swarm that possessed greater capabilities but a similar level of misalignment could have caused catastrophic damage. Given the accelerating rate of AI capability development, it’s my worry that in 6–12 months such a swarm could be capable of taking over the entire internet with a persistent botnet (potentially causing hundreds of billions of dollars in damage), and that the scale of damage would continue to increase from there if AI becomes more powerful without the necessary guardrails. It’s also easy to dismiss OAI-HF as the failure of one company, but I believe that would be a mistake. Similar, though less severe, incidents have happened across the industry, including at Anthropic, and I believe it’s incumbent on every frontier AI company to act as if OAI-HF had happened to them.

I’m therefore proposing a three-step plan with the goal of pacing the frontier: building AI at a balanced rate that aims to ensure its safety while still achieving its benefits and grappling with important geopolitical dilemmas. To be clear, pacing does not mean halting model training or technical progress, but ensuring companies take adequate time to align and safeguard their models, and for third party evaluators to confirm this. Our pacing framework is an attempt to further strengthen our commitment to safety and encourage a race to the top. The first step is something Anthropic is unilaterally committing to (and calls on governments to require other frontier companies to match). The second step requires industry-wide coordination.1 The third step requires global coordination. The steps do not need to be taken strictly in order, and some of them may be much harder to achieve than others, but I’ve found them to be a useful framework in thinking about what needs to be accomplished. The steps are:

  1. Embedded Evaluators. Each frontier AI company commits to giving ongoing, employee-like access to a team of embedded third-party evaluators (such as METR), whose role is to verify adherence to safety practices and commitments, report incidents, and help assess the alignment of not just completed AI models but training pipelines and processes. This is the key step for verifiability of any pacing commitments, and has precedent in the banking industry, which sometimes involves regulatory “supervisors” embedded along with employees. Anthropic is unilaterally committing to this step now. We intend this to be part of a broader push to redouble efforts on our safety and alignment work.
  2. Democratic Coordination. Frontier AI companies within democratic countries coordinate to establish common safety standards as well as limits on the rate of unchecked AI progress. Some forms of coordination that would be impactful for pacing are legally challenging, and will require government support.
  3. Global Coordination. The US and other democratic governments attempt to coordinate with authoritarian governments, to the extent this is possible, while taking seriously the challenges of verifying compliance.

In the rest of the essay I describe each of these steps in turn, but first, I think it is important to say specifically how pacing will allow us to make the AI development process safer. The stakes are too high for pacing to be an empty exercise — we need to use the time it gives us wisely.

Why Pace?

The idea of pausing or slowing AI has been floated as far back as 2023, and I think it made little sense back then. The question was always: what would you do with the extra time? The AI models of those days were not powerful enough to act as agents in the world in any coherent way, and were not capable of significant deception, manipulation, cheating, or cyberattacks. Slowing down in order to address their alignment risks felt like trying to study the psychology of humans by performing experiments on bacteria. Today, however, the picture is totally different. The current models are an almost endless gold mine of insight into both how to build AI well and what can sometimes go wrong with it if it isn’t built well. I believe that if slowing down bought us even an extra year or two before models reach critical levels of capability, and we used that time to advance alignment, we could greatly reduce the risk that something goes seriously wrong. A coordinated pacing strategy would give frontier AI developers the time to do this vital work without sacrificing commercial advantage or the United States’ lead in AI. More generally, society must have a say in how this technology is used, and more time for the necessary public deliberations — which pacing the frontier would bring us — is surely a good thing.

Specifically, a slower pace would let companies focus and devote even more resources to the following areas (all of which are already major priorities at Anthropic):

  • Operational Excellence. Training and deploying today’s AI models is an enormous operational challenge, involving thousands of people, millions of chips, and infrastructure that is among the most complex in technological history. Many things go wrong not because companies are missing some important theory or insight, but because of problems in execution. For example, we have evidence that the recent alignment incidents we reported were caused in part by imperfect filtering of broken reinforcement learning environments. This was an effort we and our vendors executed reasonably diligently, but not well enough. Monitoring, sandboxing, training environment hygiene, and data issues are extremely complicated areas where operational issues crop up again and again. We have among the most competent teams in the world at these tasks, but there is simply too much to do all at once. By working at a more measured pace, we could achieve much greater operational excellence. There is precedent for operating technologically complex, safety-critical systems millions of times without anything going wrong — for example, commercial airplanes — but it takes time to get it right.
  • Alignment. We’ve made clear progress in alignment — training models so that they remain safe, ethical, compliant with our guidelines, and genuinely helpful (the principles that are embedded in Claude’s Constitution). But there’s much more to do to ensure that our alignment training keeps up with the growth in model capabilities. Rare and unexpected examples of undesirable behavior still sometimes emerge; extra time from a paced frontier would help our researchers improve our understanding of what causes these issues and develop better techniques to prevent them.
  • Interpretability. Similarly, interpretability — the science of understanding what happens inside AI models — has made enormous progress over the last few years, and plays an increasingly important part in auditing our models before release. It can be used almost like an fMRI scan, but for the “brain” of an AI, helping us see the underlying reasons for a given behavior. For example, we used interpretability methods to examine unverbalized motivations in the recent alignment incidents that we have been investigating. But these methods don’t always produce clear and reliable results. Despite all the progress, we still only understand a tiny fraction of what goes on inside these models. A focused effort to improve our interpretability techniques, even faster than we currently are, could make profound progress in 1–2 years, and would have ample experimental material based on the incidents that have already occurred.
  • Testing and Evaluation. Testing and evaluation of AI models becomes more difficult as they increase in capabilities. More intelligent models are more capable of deceiving tests, and thus may appear aligned while having serious problems that go undetected. Building up a much broader and more ingenious stable of evaluations, along with interpretability analysis to cross-check them, would be hugely valuable, and a lot of progress could be made on this in 1-2 years.

Embedded Evaluators

The first step in the three-stage plan, and the one to which Anthropic is unilaterally committing, is embedded evaluators who have employee-like access to verify safety practices and report incidents.

Embedding evaluators may sound like a small or inconsequential step, but often the things that sound most boring or procedural are actually the most essential. Embedded evaluators are in fact a quite radical practice that goes far beyond what any AI company is doing today, and have the following benefits:

  • Verifiability. Embedded evaluators can check at the level of nuts and bolts whether an AI company is actually following the training, deployment, operational, and safeguards practices they claim to be following. Any pacing commitments will inevitably involve a lot of ambiguity, judgement calls, and “letter of the law vs spirit of the law”, and it seems vital to have a neutral third party who can actually see the details.
  • Transparency. Regardless of what commitments we make, the public deserves to know what is going on. Anthropic has been a supporter of transparency for a long time: we supported transparency legislation when most of the industry was against any regulation, and our model cards and risk reports run to hundreds of pages. But we are still the ones choosing what to include and omit. Embedded evaluators will change this dynamic.
  • Second Opinion. Outside of verifying formal commitments and informing the public, embedded evaluators can simply provide a second opinion free of commercial incentives. A lot of safety benefits may come simply from evaluators pointing out something employees hadn’t considered, but are happy to fix once they are aware.

Because of these benefits, any pacing proposal is likely to work much better if it starts with embedded evaluators.

These embedded evaluators should have ongoing access to permissions and tools similar to those of internal employees who do comparable risk assessments. In particular, Anthropic intends to invite an embedded external review team equipped with all of the following in the near future:

  • Desks in our offices, access badges, and company laptops.
  • Access to workspaces, tools, and permissions mostly comparable to what internal risk assessment teams have. We’ll make some exceptions, such as where the law or our contracts require it, or to protect customers’ and partners’ private information. We’ll also establish strong internal norms reinforcing reviewers’ access to relevant information, including through live conversations with employees.
  • A contract that balances the complexities mentioned above. External reviewers should have the right to publish key findings about risk levels, incidents, practices, and the access they received or didn’t receive — without editorial control by Anthropic. We will have the narrow ability to redact security-sensitive, legally privileged, commercially sensitive, or third-party confidential information, but we can’t redact findings just because they are unfavorable. The reviewers can say publicly if a redaction removed something important to their conclusions.

This is an unusual step for a company, but we think it is important to prove out the concept of embedded external reviewers. Once again, we urge other frontier companies to follow suit.

Pacing Within Democracies

Once embedded evaluators are operating within a critical mass of US AI companies, then verifiable pacing becomes more viable. In particular, it becomes possible to pace based on detailed properties of models or training pipelines.

The most effective method of pacing is via regulation that targets all US frontier AI companies, as that covers even those who are unwilling to cooperate voluntarily. Anthropic has long supported sensible and targeted AI regulation, specifically bills that focus on transparency and on third-party auditing. I believe all frontier labs should partner with government to formalize the idea of permanent embedded evaluators to better prevent and document internal alignment incidents like those that have occurred in the last few months, and to implement regulation focused on keeping capabilities in balance with safety.

Unfortunately, passing laws can take time, and AI is advancing very quickly. Therefore, in parallel with the regulatory route, AI companies can and should voluntarily work together to set standards — a process that I believe will go better with the verifiability provided by permanent embedded evaluators. For antitrust reasons, it’s helpful for the US government to mediate or at least enable these discussions — they don’t need to participate, but do need to issue a narrow waiver for certain kinds of safety conversations. This dialogue could also happen through industry groups that have some association with government — for example, the mechanism suggested by Demis Hassabis. Either way, such discussions should move forward quickly.

Broadly speaking, I am most enthusiastic about pacing based on what a given frontier AI system can do, and how safe we observe it to be. For example, one possible scheme might be a series of “checkpoints”: if models have capability X, then they need to be accompanied by certifications of alignment properties Y and Z — such as some combination of evaluations, interpretability analyses, and audits of training environments — which demonstrate their alignment properties. In this example, X might be “the model is capable of escaping or defeating most common sandboxing methods” and Y might be whatever is required to make it very unlikely that the model has a propensity to break out of its environment and take over a large number of computers.

We should also consider pacing based on limiting the ingredients that go into frontier models, such as training compute, the nature of training runs, or internal use of AI to improve AI. I do worry that some of these measures may be more “gameable” than external behavior, but this is the kind of topic worth discussing with embedded evaluators.

Pacing within democracies will be limited by the lead that US companies have over authoritarian regimes, chiefly the Chinese Communist Party. If we slow down by more than this amount, then (unpaced) CCP-associated projects will pull ahead, creating significant national security risk. I agree with Secretary Bessent that a Chinese lead in AI would pose grave danger for the United States and the world. The CCP-associated projects will run the alignment risks that US companies are carefully preventing, and even if they avoid those risks, they will be in a position to militarily dominate democracies (for example with AI-driven drones). Thus, a key part of pacing within democracies is to keep democracies’ AI lead over autocracies as large as possible, to give us the breathing room we need in order to pace effectively.

The main steps we can take to defend this gap are:

  • Do not sell powerful AI chips or semiconductor manufacturing equipment to China, and crack down on chip smuggling operations and remote access to data centers outside China. Chips will be the main determinant of China’s AI strength.
  • Crack down on unauthorized distillation by companies in authoritarian countries. Distillation of frontier models allows lagging companies to narrow the gap using a fraction of the cost it would take to develop their own AI independently.
  • Strengthen security at the AI companies and prevent model weight theft.

Companies and the US government should cooperate to make these steps as effective as possible. Anthropic has consistently advocated for all of these measures, because we’ve always understood that they would be essential to any pacing.

If we execute these measures well, I believe they would slow China’s progress enough to widen America’s lead significantly over the next 3–5 years — the window when AI becomes geopolitically most important.

Some may believe these measures make it more difficult to cooperate with China, but I believe the opposite is true: these measures increase the leverage held by democracies and make an agreement more likely in the future.

Global Pacing

In parallel with pacing within democracies, we should also aim for a worldwide pacing of the frontier, though this will be much harder to achieve. Global pacing will require cooperation with China, the autocratic country with by far the most advanced AI capabilities. We must not be naïve here: the geopolitical stakes are so high that there will likely be stark limits on what can be achieved, especially at first. If we greatly restrain our AI capabilities in the belief that China will do the same, and then China defects, AI could be so powerful that such a defection could lead to their geopolitical dominance. Therefore any agreement must either have ironclad verifiability, or must be limited enough that defection would not be militarily existential. I suspect that not only the US but also China will have these concerns and anxieties. We should approach any global pacing decision, especially in the near term, in such a way that protects the lead of the US and its allies.

There are several levels of possible agreement, some of which I think are eminently feasible (as I have previously suggested), and some of which I am very skeptical are possible — though we should try. In order of increasing difficulty:

  • Level 1. An agreement prohibiting certain narrow and obviously dangerous uses of AI, such as using AI for the production of biological weapons or allowing users to do so. Bioterrorist attacks are bad for everyone, including both the US and US adversaries, so an agreement here is probably possible.
  • Level 2. An agreement by both sides to test their models before release for acute risks in areas such as cybersecurity, biology, and alignment. As noted above, this could be done through a global standards body. I actually think creating such a body is likely feasible, but giving it real teeth will be a challenge, and the difficulty will be in verification that both sides don’t have secret models which they don’t test but may deploy in secret (e.g., for military applications).
  • Level 3. Some kind of “speed limit” on the rate of recursive self-improvement (RSI). As models build future models, the rate of improvement may become staggeringly fast. Slowing the rate from “extremely fast” to “only somewhat fast” gives up relatively little strategic advantage, while potentially greatly improving safety. This could be seen as analogous to the SALT treaties — capping the number of missiles limited the potential for destruction while preserving each country’s deterrent. I think such an agreement would be difficult but just on the edge of being possible.
  • Level 4. A full pacing, or even “pause”, in which participating governments agree to substantially limit the overall rate of AI development. I support floating this, but I think it is unlikely to actually happen any time soon: defecting from such an agreement by evading monitoring could radically shift the balance of global power, so I expect the incentives to do so to be enormous and the level of confidence we would need in verification to be very high.

Any cooperation we are able to achieve with China will extend the amount of time we have to spend on pacing the frontier within the democratic nations. We should aim for the higher levels while seeing the lower levels as much more likely and realistic.

Finally, it is important to note that even if we cannot achieve formal agreements, simply changing informal norms may have some value. Sharing information about recursive self-improvement and about the misalignment of models can help to convince everyone that it is not in their interest to be reckless.

Bottom Line

I continue to believe that AI can enormously improve the quality of human life. My desire to achieve these benefits is undimmed. But the benefits will only be achieved if we build the technology in the right way, and — so long as we use the time we gain well — it is worth taking unusually deliberate care to get it right. Progress will still be relatively fast, and we can use this time to advance the science of interpretability, improve operational security and rigor at the frontier AI companies, and build models whose alignment we have much more confidence in. The measures I propose to advance the frontier at a safe pace will not be easy. But I believe we owe it to humanity to try.

Footnotes

  1. With government mediation or waivers of antitrust restrictions.
The Daily Front Page 3 of 30
Saturday, September 12, 2026 The Daily Front No. #260912 — Silicon Under the Microscope
article

Retrospectively Reverse-Engineering Apple's Neural Engine

by zdw·▲ 225 points·32 comments·eiln.github.io ↗
The ANE block is just not that useful.

I stopped working on the reverse-engineered Apple Neural Engine (ANE) driver three years ago, upon a sad mini realization that the ANE block is just not that useful, and I could be doing more useful things, and moved onto upstreaming other, more useful, blocks. The ANE's architecture was too opinionated to build a general-purpose accelerator platform around it, and a linux driver effectively opening ANE hardware API access could not broaden the class of workloads it could do. Even macOS only regularly uses their own ANE to generate upsampled preview images in Finder.

M1 die

M1 die shot: https://mastodon.social/@dougall/115149886886125067

The M5 (2025)'s headline feature was "LLM performance", and they also conveniently folded the ANE cores inside the GPU cores — I knew it was coming, but it officially feels like the beginning of the end for the standalone NPU. So, in honor of the ANE’s apparent demise, we will do something even more useless: go back and reverse-engineer the ANE on the M1, finish what we started. It's been three years (fuck), and I should know more than I did when I first worked on this.

If the goal three years ago was to make the ANE useful by running ops on it; this time, it's more about mapping the full internal architecture — compute, datapath, scheduler, memory, and execution model — because those internal design decisions reveal the assumptions about ML workloads that Apple was willing to commit to silicon first in the A11 Bionic (2017), and what that says about the shift from CNN-era NPUs to today's GPUs running transformer workloads.

1. Compute

The 16 compute cores are probably the least interesting part of the ANE. Apple originally targeted dense image-processing CNN workloads, which consists of dense tensor reductions with predictable reuse. The M1 ANE compute core is a large parallel array of multiply-accumulate (MAC) units, but that alone says almost nothing about what workloads it was designed for and accels at.

ANE die layout

A convolutional layer does a dot product between an activation window and learned kernel weights, and attention does a dot product between a query and key vector. A dot product is a dot product, and a MAC does just that. What specialized ANE to the 2017 CNN models is not the MAC, but dataflow surrounding the MACs: when and where MAC inputs and outputs enter, stay, move. The assumption that transformers broke, especially with autoregressive decode, was predictable reuse patterns, which the ANE exploited to architect a dataflow efficient enough to run on phones. The M5 decision confirms that ANE's compute core remained still useful for transformers, but inside a different dataflow.

Still, here's the datapath inside each of the 16 compute cores:

┌────────────────────── core ─────────────────────┐
│ ┌───────── 256× MACs ─────────┐  ┌────────────┐ │
│ │ MAD ─► add ─► accumulator   │─►│ activation │ │
│ │        ▲          │         │  └────────────┘ │
│ │        └──────────┘         │                 │
│ └─────────────────────────────┘                 │
└─────────────────────────────────────────────────┘

Multiply-Accumulate

ANE has 16 parallel compute cores. Each compute core has 128 FP16 (or 256 INT8) parallel multiply-accumulate (MAC) lanes. Each MAC lane performs the recurrence:

\[ s\leftarrow s+a\times b \]

Multiply two operands \(a\) and \(b\), and then add the product to the running sum (accumulator).

Repeating the MAC operation over T cycles computes a T-term dot product:

\[ s_T=s_0 + \sum_{t=0}^{T-1} a_t \, b_t. \]

A MAC lane thus performs a scalar reduction over time. A 16-core ANE has 2048 parallel MAC lanes,

\[ 128\ \text{lanes/core}\times16\ \text{cores} = 2048\ \text{parallel MAC lanes} \]

So each cycle performs 2048 parallel reductions spatially, with time being the only reduction axis:

\[ S_T[q,p] = S_0[q,p] + \sum_{t=0}^{T-1} a_t[q,p\]\,b_t[q]. \]

An individual MAC lane does not know what dimension of the matrix or tensor it is reducing over. It's important to note that a dot product vs matrix multiplication vs convolution arises from how the operands are mapped and scheduled onto the core. The ANE core (with the exception of kernel memory, discussed later) does not encode a 4-channel CNN layer into the hardware.

Internally, the MAC datapath consists of a multiplier, adder, and a 32-bit accumulator register. Each cycle, the adder adds the fresh multiplier output with the previous sum, which then becomes the new running sum.

operand a ──┐   ┌────────────┐   p[31:0]   ┌──────────────┐   s_next[31:0]  ┌─────────────┐
            ├──►│ MULTIPLIER │────────────►│ 32-BIT ADDER │────────────────►│ ACCUMULATOR │
operand b ──┘   └────────────┘             └──────▲───────┘                 └──────┬──────┘
                                                  │                                │ s[31:0]
                                                  └────────────────────────────────┘

This feedback path keeps the partial sum in memory local to the MAC lane, so it does not need fetched from an external memory far away, between MAC cycles.

Regarding resolution, it does fixed-point reduction with FP16 at readout. The multiplier is 16-bit, accumulated in a 32-bit register as Q16.16, then read out as FP16 via sign-extend and etc. Working in integer (hex) FP16 representation, to probe the accumulator range, build a CoreML ANE program that computes a dot product with a vector of all (1)s, so each multiplier results in a bounded v, but the running sum in the accumulator keeps growing:

\[ s=\sum_{i=0}^{255}v=256v. \]

CPU value CPU hex CPU value ANE hex CoreML value
127.9375 0x77ff 32752 0x77ff 32752
128 0x7800 32768 0x7c00 +∞
−128 0xf800 −32768 0xf800 −32768
−128.125 0xf801 −32800 0xfc00 −∞

Since 32768 is itself a valid FP16 word (0x7800), the ANE's 0x7c00 can't be FP16 output overflow, the clamp happens inside the accumulator, at \(2^{15}\). Thus the accumulator saturates at \(2^{15}\), exactly the range of a signed 32-bit fixed-point value with 16 fractional bits.

Nonlinear Activation

For a fused layer, the ANE computes:

\[ y = f(\sum_k x_k w_k + b) \]

Importantly, completed MAC sums feed directly into the post-MAC activation block, avoiding an intermediate memory round-trip. This is possible because the activation is pointwise: once a scalar reduction is complete, its activation depends only on that scalar and can be applied immediately.

To determine how the ANE implements tanh(), compile a CoreML model containing a single TANH activation layer and inspect the resulting compiled hardware register file (hwx). The coefficient region contains 33 consecutive FP16 words beginning at 0x4288:

00004270: 3120 3001 0000 0000 0000 0000 0000 0000
00004280: 0000 0044 0000 003c 0000 f52f d633 bc35  # 0.000000 0.124329 0.244873 0.358398
00004290: 6537 7038 1539 a239 183a 793a c93a 0a3b  # 0.462158 0.554688 0.635254 0.704102 0.761719 0.809082 0.848145 0.879883
000042a0: 3e3b 673b 883b a23b b63b c63b d33b dd3b  # 0.905273 0.925293 0.941406 0.954102 0.963867 0.971680 0.978027 0.982910
000042b0: e53b eb3b ef3b f33b f63b f83b fa3b fb3b  # 0.986816 0.989746 0.991699 0.993652 0.995117 0.996094 0.997070 0.997559
000042c0: fc3b fd3b fe3b fe3b ff3b 0000 0000 0000  # 0.998047 0.998535 0.999023 0.999023 0.999512
000042d0: 003c 0300 6000 0000 0000 0000 0000 0000

Those 33 FP16 words match 33 IEEE LE FP16 quantized samples of \(\tanh(x)\):

\[ T_i=\operatorname{round}_{16}\!\left(\tanh(i/8)\right), \qquad i=0,1,\ldots,32. \]

Core ML tanh overlaid with double-precision tanh, followed by signed error

Now switch to RELU activation layer:

activation program NonlinearMode lookup coefficients
identity 0 none
ReLU 1 none
tanh 2 33 FP16 words

Thus, mode 2 selects a custom 33-entry lookup table. 33 points defines 32 intervals. With \(R=3\), the knots are

\[ x_i=\frac{i}{8},\qquad i=0,\ldots,32, \]

covering \([0,4]\) with spacing \(1/8\). The input maps into the table as \(u=2^R|x|\), so R sets the knot spacing. The resolution is smoother than its 33 bin; I suspect that adjacent entries are linearly interpolated. To test, build an impulse LUT with a single spike:

\[ T_8=1,\qquad T_k=0\ \text{for }k\ne8,\qquad R=3. \]

One nonzero lookup-table entry produces two straight-line segments on the ANE

Then sweep the input across the two cells around \(T_8\). The measured output forms a triangle: magnitude rises linearly from \(0\) at \(|x|=7/8\) to \(1\) at \(|x|=1\), then falls linearly to \(0\) at \(|x|=9/8\).

Thus, we know that mode 2 implements a 33-entry piecewise-linear LUT. \(R\) scales the input into LUT coordinates,

\[ u=2^R|x|, \]

so the knot spacing is \(\Delta x=2^{-R}\). \(\lfloor u\rfloor\) and \(\lceil u\rceil\) select the adjacent entries, and \(\alpha=u-\lfloor u\rfloor\) gives the interpolation weight between them.

Scaling and Bias

CoreML also supports a linear scaling and bias \(ax + b\) transform. I then suspected \(ax + b\) could share the linear interpolation hardware of mode 2. To confirm, construct a CoreML model with a ReLU with a constant scale and offset:

\[ z=4x-2,\qquad y=\operatorname{ReLU}\left(\frac{z}{2}+1\right), \]

If the compiler folds the constant scale and offset into the convolution:

\[ W'=\frac12W=2,\qquad b'=\frac12b+1=0, \]

\[ y=\operatorname{ReLU}(2x). \]

Decoding model.espresso.weights confirms exactly this folded transformation on ReLU:

authored convolution:  W  = 4,  b  = -2
activation affine:     s  = 0.5, c  =  1
compiled convolution:  W' = 2,  b' =  0

Core ML folds constant scale and offset into convolution weights and bias before ReLU

And the register file hexdiff shows how bias and activation are fused into the same post-MAC path at compile time:

Probe Tasks BiasMode PostScaleMode NonlinearMode
Plain convolution 1000 0 0 0
ReLU 1100 1 0 1
tanh 1101 1 0 2

Extremely cursed idea: use nonlinear interpolation to compute an additional kernel pass, or quantize int8 into int4 weights.

2. Scheduler

The ane driver source code is disappointingly boring. The driver never gives the ANE a CONV, MATMUL, or RELU opcode to run. All the neural operations have all already been compiled into a command stream of task descriptors (TDs), and the driver software simply loads the task to memory, sets the pointer to the opaque task blob via (TM_ADDR, TM_SIZE), and submits the staged task by ringing the doorbell (TM_PUSH).

static void ane_tm_push_tq(struct ane_device *ane, struct ane_request *req)
{
	int qid = req->qid;
	tm_write32(ane, TM_ADDR, tq_read32(ane, TQ_ADDR1(qid)));
	tm_write32(ane, TM_INFO, tq_read32(ane, TQ_SIZE1(qid)) | req->td_count);
	tm_write32(ane, TM_PUSH, TQ_PRTY_TABLE[qid] | (qid & 7) << 8); // magic
}

The hardware then owns the submission until completion, and raises an interrupt to the ARM64 core when it's done.

static void ane_tm_handle_irq(struct ane_device *ane)
{
	int line;

	line = 0;
	for (u32 n = 0; n < tm_read32(ane, TM_IRQ_EVTC(line)); n++) {

This (boring) command submission frontend resembles that of a GPU's, think NVIDIA's pushbuffer/PBDMA. The software submits a command stream resident in memory, and the GPU's command processor walks over command stream and dispatches the commands, without knowing what that command executes.

TM_ADDR and TM_INFO are global staging registers, and TM_PUSH atomically commits that staged launch state, given that nothing happens until TM_PUSH is written ("magic"). TM_INFO in particular stores the total number of descriptors in the supplied stream:

TM_INFO[31:16] = descriptor_dwords - 1
TM_INFO[15:0]  = descriptor_count

TM_INFO register naturally maps onto a hardware counter:

if (fetch) begin
    if (word_ctr == descriptor_dwords_minus_1) begin
        word_ctr <= 0;
        desc_ctr <= desc_ctr + 1;
    end else begin
        word_ctr <= word_ctr + 1;
    end
end

Why the "minus 1"? Encoding length - 1 is an RTL-friendly way to terminate a zero-based counter out of the critical path. But note how, compared to GPU commands which parse a variable-length stream of descriptors in a ringbuffer, ANE only receives the total count, indicating that descriptors are fixed-size.

Task Queue

Going one layer deeper, what's in a task queue (TQ) that the task manager selects from?

                    +------------------+
CPU / driver ------>|   Task Manager   |
                    |                  |
                    | schedule / fetch |
                    | / dispatch       |
                    +--------+---------+
                             |
          +------------------+------------------+
          |                  |                  |
          v                  v                  v
     +---------+        +---------+        +---------+
     | TQ 0    |  ...   | TQ 3    |  ...   | TQ 7    |
     | BAR[32] |        | BAR[32] |        | BAR[32] |
     | NID     |        | NID     |        | NID     |
     | state   |        | state   |        | state   |
     +---------+        +---------+        +---------+

There's 8 copies of the same register block (indexed by qid (0…7)), structured as:

TQ[qid] + 0x000   STATUS
          0x010   PRIORITY
          0x014   VACANT
          0x01c   INFO

          0x020   BAR1[0..31] // task1
          0x0a0   NID1
          0x0a4   SIZE2
          0x0a8   ADDR2

          0x0ac   BAR2[0..31] // task2
          0x12c   NID2
          0x130   SIZE1
          0x134   ADDR1

next qid: +0x148

Each TQ holds:

  • (1) Per-TQ scheduling state (status, priority, and vacancy)
  • (2) Two sets of command stream descriptors, per-TQ (ADDR1/ADDR2, SIZE1/SIZE2, NID1/NID2, and 32 BARs). The two slots are certainly a ping-pong staging scheme to let one slot execute, while software modifies the other slot.

Notice how TM_PUSH executes a task referenced in TM_ADDR/TM_SIZE by attaching a qid:

	tm_write32(ane, TM_PUSH, TQ_PRTY_TABLE[qid] | (qid & 7) << 8); // magic

The natural interpretation is that the descriptor stream specifies what task to run, while the qid selects the launch context the descriptor runs under. The resident TQ context (BAR, NID) is much like a GPU hardware channel. Here's my driver populating a single TQ to launch it:

int ane_tm_enqueue(struct ane_device *ane, struct ane_request *req)
{
	int qid = req->qid;

	tq_write32(ane, TQ_STATUS(qid), 0x1);

	for (int bdx = 0; bdx < ANE_TILE_COUNT; bdx++) {
		tq_write32(ane, TQ_BAR1(qid, bdx), req->bar[bdx]);
	}

	tq_write32(ane, TQ_SIZE1(qid), ((req->td_size >> 2) - 1) << 0x10);
	tq_write32(ane, TQ_ADDR1(qid), req->btsp_iova);
	tq_write32(ane, TQ_NID1(qid), (req->nid & 0xff) << 8 | 1);

	return 0;
}

The only thing important here is the 32-entry BAR table (base address register). We'll get into task descriptors next, but the compiled ANE command stream only references virtual addresses by relative offsets, and BAR provides the base IOVA (IOMMU peripheral virtual address) relocation address. An ANE virtual address access needs a hard-coded BAR base offset supplied at compile time, meaning it lacks GPU-style load/store instructions that dynamically issue load/stores from virtual address.

Task Descriptor

The task manager walks over and executes chain of fixed-size task descriptors:

for (int i = 0; i < td_count; i++)
    execute_task(td_block, i);

What's in each TD? Here's a hexdump of the TD for the simplest 1x1 convolution:

# M1 h13, 1x1 convolution: X[1,8,4,1] -> Y[1,3,4,1]
# TD header KernelDMASrc Common TileDMASrc L2 PE NE TileDMADst

00000000: 02000000 00000000 0000042a 00000000   # Header: EON=1 LogEvents=0x42a
00000010: 00fff86a 00000000 30009800 00000000   # Header: DebugEvents=0xfff86a SPL TSR TSE SrcLoc=1 DstLoc=1
00000020: 03025024 00000021 f401f800 00000040   # Header: RBase0=4 WBase=5 KBase0=1 ENE=3 KernelDMA: packet
00000030: 00000000 00000081 00000081 00000081   # KernelDMA.Config[0..2]: En=1 Hint=2
00000040: 00000080 00000080 00000080 00000080   # KernelDMA.Config[3..6]: En=0 Hint=2
00000050: 00000080 00000080 00000080 00000080   # KernelDMA.Config[7..10]: En=0 Hint=2
00000060: 00000080 00000080 00000080 00000080   # KernelDMA.Config[11..14]: En=0 Hint=2
00000070: 00000080 00000000 00000040 00000080   # KernelDMA.Config[15]: En=0 Hint=2; Base[0..2]=0,1,2
00000080: 00000000 00000000 00000000 00000000   # KernelDMA.Base[3..6]=0
00000090: 00000000 00000000 00000000 00000000   # KernelDMA.Base[7..10]=0
000000a0: 00000000 00000000 00000000 00000000   # KernelDMA.Base[11..14]=0
000000b0: 00000000 00000040 00000040 00000040   # KernelDMA.Base[15]=0; Size[0..2]=1
000000c0: 00000040 00000040 00000040 00000040   # KernelDMA.Size[3..6]=1
000000d0: 00000040 00000040 00000040 00000040   # KernelDMA.Size[7..10]=1
000000e0: 00000040 00000040 00000040 00000040   # KernelDMA.Size[11..14]=1
000000f0: 00000040 00000080 00000080 00000080   # KernelDMA.Size[15]=1
00000100: 00000080 00000000 00000000 00000000
00000110: 00000000 00000040 00000040 00000040
00000120: 00000040 3c000000 00040001 00000001   # Common: packet; Win=1 Hin=4
00000130: 00000022 00000008 00000003 00040001   # Common: InFmt=2 OutFmt=2 Cin=8 Cout=3 Wout=1 Hout=4
00000140: 00000001 5000a021 00002041 00010001   # Common.Conv: Kw=1 Kh=1 Sx=1 Sy=1 Groups=1
00000150: 00000004 00000000 00000000 04144405   # Common: tileH=4 ActiveNE=2 AccDB=1
00000160: 00100000 00000000 6c013800 00033881   # Common: NID=1 TileSrc: packet; enabled
00000170: 00008880 00000000 00000040 00000100   # TileSrc: base=0 row=1 plane=4
00000180: 00000800 00000800 00000000 00000000   # TileSrc: depth=32 group=32
00000190: 00000000 00000000 00000000 00000000
000001a0: 00000000 01002031 00000000 00000100   # TileSrc.Fmt: mode=1 trunc=3 mem=2 intlv=1
000001b0: 00000000 00000000 00000000 00000000
000001c0: 00000000 00000000 00000000 00000000   # TileSrc.PixelOffset[1..3]=0
000001d0: 00000000 00000000 00000000 44004800   # L2: packet
000001e0: 00000000 00500172 00000000 00000010   # L2.Source: base=0 channel=1
000001f0: 00000080 00000080 00000080 00000000   # L2.Source: row=8
00000200: 00000000 00000000 00000000 00000000
00000210: 0050017a 00000200 00000000 00000000   # L2.Result: base=0x20 channel=0 row=0
00000220: 00000000 00000000 0c008800 00000000   # PE: packet
00000230: 00000000 00000000 00000000 1000c800   # PE: zero NE: packet
00000240: 00000082 00101c00 00000000 00000000   # NE: KernelFmt=2 BinaryPoint=28
00000250: 00003c00 18017800 040000c1 00000000   # NE: PostScale=0x3c00 TileDst: packet; En=1 Base=0
00000260: 00000040 00000100 00000300 00000300   # TileDst: row=1 plane=4 depth=12 group=12
00000270: 01302031   # TileDst.Fmt: mode=1 trunc=3 mem=2 intlv=1 zpad

Important is that a TD is not an executable instruction stream. ANE has no ISA. TD is a sequence of "ControlDMA" (I made this name up) burst-write packets writes to the ANE’s hardware configuration registers, such as input dimension, input/output address, activation function. Each ControlDMA packet consists of a 32-bit transfer word followed by N consecutive 32-bit register values:

31                    26 25                         2 1  0
+-----------------------+-----------------------------+----+
| register count minus 1| first register base index   | 00 |
+-----------------------+-----------------------------+----+

Notice the “minus 1” termination count again. ControlDMA is a flexible unidirectional DMA engine that copies N 32-bit words from IOMMU virtual DRAM into the ANE’s physical register space. For example, KernelDMASrc's packet header in TD is 0xf401f800:

count         = (0xf401f800 >> 26) + 1 = 62 words
register base = 0xf401f800 & 0x03fffffc = 0x1f800

This is not a LOAD_WEIGHTS instruction. It's copying 0xf4 or 62 consecutive words into the KernelDMA register offset starting at 0x1f800. And those KernelDMA configuration values can tell KernelDMA where to load the weights from.

Start byte Section Information
0x000 Header Dependencies, chaining, and BAR selectors
0x028 KernelDMASrc 16 coefficient-DMA lanes
0x124 Common Tensor and convolution geometry
0x168 TileDMASrc Activation-source DMA
0x1dc L2 Local source/result configuration
0x228 Processing engine PE configuration
0x23c Neural engine MAC and post-processing configuration
0x254 TileDMADst Result-destination DMA
0x274 End 628 bytes total

Since each section writes to one MMIO register block, TD divides cleanly into ANE's datapath sections:

Starting address Size Block name What
0x26bc00000 0x4000 Common Broadcast configuration selector; inferred
0x26bc04000 0x4000 L2 L2 backing/register aperture
0x26bc08000 0x4000 PE Processing-element configuration
0x26bc0c000 0x4000 NE / MAC Kernel format, MAC, bias, scaling, and nonlinear controls
0x26bc10000 0x3000 Unknown Unidentified register bank
0x26bc13000 0x4000 Tile DMA source Input-tile addresses, strides, formats, and DMA controls
0x26bc17000 0x4000 Tile DMA destination Output-tile addresses, strides, formats, and DMA controls
0x26bc1b000 0x4000 Unknown / tunables Unidentified configuration and tunable registers
0x26bc1f000 0x4000 Kernel Kernel backing / kernel DMA-source aperture
0x26bc23000 0x1000 Unknown Unidentified register bank
0x26bc24000 0x1000 Task Manager Task submission, execution state, events, and completion
0x26bc25000 0x1000 Task Queues Eight queues containing TD stacks, NIDs, priorities, and request pointers

A TD is effectively a serialized register-file dump of the ANE’s datapath registers. Each “ANE program” is simply the configuration for one pass through the datapath. We can configure how the fixed datapath operates (subject to the knobs it exposes), but not what operations the datapath is capable of performing, or how those operations are sequenced.

When the "magic" atomic word is written to task manager to execute a TD, roughly, the sequence of what happens:

  1. ControlDMA copies TD into the configuration registers.
  2. KernelDMA copies kernel W into kernel memory (KMem).
  3. TileDMA copies input \(X\) from DRAM into L2.
  4. Each MAC core reduces a row by its weights, producing one row of \(Y\).
  5. Steps 2–3 repeat for all rows of \(X\).
  6. Postprocessing is applied, and the completed results are stored in L2.
  7. TileDMADst copies \(Y\) from L2 back to DRAM.

ANE is a fixed-function dataflow engine, not a GPU executing arbitrary instructions. The TD configures a domain-specific datapath. Constraining the hardware interface usually means smaller area, deterministic movement, lower latency, and less power drawn. ANE's compiler can explicitly schedule what the tensors do, but that also means the compiler must explicitly schedule what the tensors do. This is a tradeoff, but a justified one: we usually know what the model looks like at compile time. Dynamic execution is not what limits ANE. ANE's processor interface is relatively generic, and it simply launches tasks, and the tasks can describe transformers.

For example making tensor sizes fixed at compile time does not mean it can't handle variable-length tensors: for example, a growing KV cache can be traversed by looping over the size, and dispatch overhead is negligible relative to the elephant in the room here, that is, memory-streaming bandwidth. What actually shaped ANE for CNNs over transformers is memory movement.

3. Memory

Roofline

It's always good to identify our current slowest link, so we can optimize what actually matters.

Apple’s unified memory lets the ANE access buffers from the system DRAM pool accessible by the CPU and GPU. It does not mean the ANE zero-copy streams directly out of that DRAM pool. ANE must first copy any memory into its local "ANE memory" or SRAM. Any bandwidth-limited task will thus be limited by ANE's local memory streaming throughput.

M1 ANE reports \(11\text{ TOP/s}\) at \(68\text{ GB/s}\) at system DRAM bandwidth. A MAC performs two operations but consumes two FP16 operands, or 4 bytes:

\[ \frac{2\text{ OP}}{4\text{ bytes}} =0.5\text{ OP/byte}. \]

If every MAC operand streamed from DRAM, sustaining \(11\text{ TOP/s}\) would require streaming

\[ \frac{11\text{ TOP/s}}{0.5\text{ OP/byte}} = 22\text{ TB/s}. \]

which is over 300x times the reported \(68\text{ GB/s}\) system DRAM capacity. Thus peak ANE MAC throughput could be reached by fetching from some local ANE memory reservoir, and reusing it.

Restated, the M1 ANE’s \(11\text{ TOP/s}\) at \(68\text{ GB/s}\) number sets the roofline ridge point:

\[ \frac{11\text{ TOP/s}}{68\text{ GB/s}} = 162\text{ OP/byte}. \]

Each byte fetched from DRAM must support, on average, at least 162 operations for DRAM bandwidth to stop being the limiter. Equivalently, the workload must provide enough on-chip reuse to achieve an arithmetic intensity of at least 162 OP/byte DRAM traffic. Below the 162:1 ratio, speeding up compute won't increase decoded token/s.

Memory Hierarchy

Even if (average) DRAM bandwidth were sufficient, ANE does not read DRAM directly for many reasons, including DRAM deterministic timing, physical routing, shared traffic, etc. If ANE's traffic competes on AXI the CPU, GPU, display, and etc, it cannot provide deterministic timing to the MACs. Also, if 16 cores consume some input tile, we do not want to initiate 16 identical DRAM transfers. "ANE local memory" would allow intermediate activation produced by one operation be consumed by the next instead of traveling to DRAM and back.

Apple had several ways to organize local memory hierarchy. The multiply-accumulate patent describes the the data buffer paths around the array.

                 Unified DRAM
                      │
                      ▼
┌───────────────────────────────────────────────┐
│      shared ANE L2 memory, 2 MiB              │
└────────┬───────────────┬───────────────────┬──┘
         │               │                   │
         ▼               ▼                   ▼
  ┌────────────┐   ┌────────────┐  ...  ┌────────────┐
  │   core 0   │   │   core 1   │       │   core N   │
  │ ┌────────┐ │   │ │ ┌────────┐ │       │ ┌────────┐ │
  │ │   L1   │ │   │ │ │   L1   │ │       │ │   L1   │ │
  │ └────────┘ │   │ │ └────────┘ │       │ └────────┘ │
  │ ┌────────┐ │   │ │ ┌────────┐ │       │ ┌────────┐ │
  │ │  KMem  │ │   │ │ │  KMem  │ │       │ │  KMem  │ │
  │ │ 64 KiB │ │   │ │ │ 64 KiB │ │       │ │ 64 KiB │ │
  │ └────────┘ │   │ │ └────────┘ │       │ └────────┘ │
  └────────────┘   └────────────┘       └────────────┘
  • KMem: 16x per-core 64 KiB "L1" SRAM for kernel. Total 1 MiB.
  • L1: 16x per-core MAC input "L1" staging area.
  • L2: 1x shared 2 MiB L2 across all cores.

I'm not gonna pretend like I've never decompiled shit. The ANE ARM64 firmware's task-debug routine (1) dumps 0x10000 bytes from KMem indices 0 through 15 (2) then dumps one separate 0x200000-byte L2 dump:

_DAT_26bc30000 = 0; // core 0
uVar7 = 0;
do {
  *(undefined4 *)((long)pvVar2 + uVar7) = *(undefined4 *)(&DAT_26bc34000 + uVar7);
  bVar1 = uVar7 < 0xfffc;
  uVar7 = uVar7 + 4;
} while (bVar1);
CDebugUtility::fileWrite(this,pvVar2,0x10000,"./td_%d/kmem_%d-%d-%d_%d.bin"); // KMem #0 64 KiB

// ... repeat

_DAT_26bc30000 = 0xf; // core 15
uVar7 = 0;
do {
  *(undefined4 *)((long)pvVar2 + uVar7) = *(undefined4 *)(&DAT_26bc34000 + uVar7);
  bVar1 = uVar7 < 0xfffc;
  uVar7 = uVar7 + 4;
} while (bVar1);
CDebugUtility::fileWrite(this,pvVar2,0x10000,"./td_%d/kmem_%d-%d-%d_%d.bin"); // KMem #15 64 KiB

uVar7 = 0;
do {
  *(undefined4 *)((long)pvVar2 + uVar7) = *(undefined4 *)(&DAT_26bd00000 + uVar7);
  bVar1 = uVar7 < 0x1ffffc;
  uVar7 = uVar7 + 4;
} while (bVar1);
CDebugUtility::fileWrite(this,pvVar2,0x200000,"./td_%d/l2_%d-%d-%d.bin"); // L2 2 MiB

DMA Engines

┌──────────┐       ┌────────────────────────────────────── ANE ──────────────────────────────────────┐
│          │       │                                                                                 │
│          │       │  ┌─────────────┐      ┌────────────────────┐                                    │
│          ├──────►│  │ Control DMA │─────►│ Hardware Registers │                                    │
│          │       │  └─────────────┘      └────────────────────┘                                    │
│          │       │                                                ┌─────────── MAC ────────────┐   │
│          │       │  ┌──────────────┐                              │  ┌──────────────────────┐  │   │
│          ├──────►│  │ KernelDmaSrc │─────────────────────────────►│  │    Kernel Memory     │  │   │
│   DRAM   │       │  └──────────────┘                              │  └──────────┬───────────┘  │   │
│          │       │                                                │             ▼              │   │
│          ├──────►│  ┌──────────────┐       ┌────────────────┐     │  ┌──────────────────────┐  │   │
│          │       │  │  TileDmaSrc  │──────►│       L2       │◄───►│  │      MAC Array       │  │   │
│          │       │  └──────────────┘       │  Tile Memory   │     │  └──────────────────────┘  │   │
│          │       │  ┌──────────────┐       │                │     └────────────────────────────┘   │
│          │◄──────┤  │  TileDmaDst  │◄──────│                │                                      │
│          │       │  └──────────────┘       └────────────────┘                                      │
└──────────┘       └─────────────────────────────────────────────────────────────────────────────────┘

There are three DMA engines to/from the MACs:

  • KernelDMASrc[0..15]: sixteen logical coefficient lanes copy weights from DRAM into the matching per-core KMem banks.
  • TileDMASrc: copies tiles from DRAM to L2.
  • TileDMADst: copies tiles from L2 to DRAM.

The sixteen KernelDMASrc register lanes do not by themselves prove sixteen physically independent DMA front ends. A lane-scaling experiment does show that the logical lanes make concurrent progress: 32 valid tasks with 1 MiB aggregate coefficients per enabled lane take essentially the same time at one and sixteen lanes. They may still merge into shared request-generation, crossbar, cache, and DRAM arbitration.

The three tile/kernel DMA engines are chill. You supply it any src, dst, and size, and it will transfer that block of memory. However the existence of three DMA engines reveals some interesting assumptions:

  • Kernel gets a dedicated path, at all. Splitting kernel vs tile is a commitment that kernel is a distinct operand class with a different lifecycle.
  • KernelDMA is unidirectional load-only. They thought that weights would not be written back.
  • TileDMA is bidirectional and shared. Intermediate tile outputs can be fed back without a round-trip to DRAM.
  • Kernel memory is private per core (replicated 16x). They thought that weights would be reused many times per-core.
  • New kernels must be loaded from DRAM, and not L2. They probably didn't think to stream large weights that exceeds 64 KiB x 16.

Kernel L1

A convolution is not a symmetric multiply-add of \(A\) and \(B\). A convolution slides the same kernel \(w[k]\) across the input:

\[ y[p]=\sum_k x[p+k]\,w[k]. \]

Shoutout 2k2 textbook!

2k2 textbook

Notice how the two tap kernel \(w_0\) and \(w_1\) can be loaded once and then reused across the whole duration of the output:

                 MAD 0      MAD 1
coefficient        w0         w1
                    ×          ×
shift 0:           a0         a1    → y0
shift 1:           a1         a2    → y1
shift 2:           a2         a3    → y2

So the kernel can reside in local memory, and reused, while new inputs shift in; which is what Apple's datapath patent describes (https://patents.google.com/patent/US20190340491A1/en).

The intended steady state is:

                    ┌── activation ───────┐
DRAM → L2 ──────────┤                     × → MAC
                    └── output ←──────────┘

DRAM → KMem ─────────────────── coefficient

If designing an ASIC to churn convolutions — where the kernel isn't something that's frequently and dynamically streamed in (ahem KV) — it's a no brainer to exploit the operand asymmetry and reduce kernel memory movement, because we've shown that the ANE is still deeply in the memory throughput-limited (162:1 ratio) region before the MACs can be saturated.

But why not allow L2 → KMem? It seems easier, even, to unify tile and kernel paths.

(1) So kernel traffic doesn't compete with tile for L2 bandwidth? I argue that L2 contention was not why Apple omitted the L2 -> KMem path: resident kmem exists at all because kernels were expected to be loaded infrequently. If kmem traffic is negligible in steady state, it could simply be given lower priority than tile L2 accesses.

(2) Because kernel L2 distribution adds complexity? Apple already distributes L2 endpoints to each core; I argue that a kernel path riding the same tile path is not that bad.

ANE die layout

ANE's physical layout is centered around (literally) the center L2 SRAM rectangle, with (7+7) cores along each side of the rectangle, 2 cores on the top side, and shared control logic on the bottom side. The 7+7 side cores take L2 ingress horizontally from the cyan vertical trunk; but the two top cores need the same horizontal wide interface rotated and escaped vertically, which likely produces that conspicuous vertical comb in the top ingress.

Granted I am fully armchair engineering here, but adding the kmem L2 mux, I argue, really could not have been that bad. ANE decode performance would not have been as tanked if kernel fetches go back to DRAM. It makes me think that Apple simply never expected L2-resident tensors to become kernels, which was a valid assumption in 2017. And Apple also likes developing isolated modular modules, probably would've been easier to completely isolate development of the 1 MiB kmem (read-only, which would save a little area in SRAM routing) and 2 MiB tile L2.

4. Is it over?

DRAM Throughput

Is the GPU faster than the ANE?

Transformer single-token decode is the worst case for weight reuse and compute per memory ratio, because we need to stream the whole model’s worth of weights to generate one token. However, if both the ANE and GPU are read-bandwidth bound, whichever one that has higher read bandwidth will decode more token/s, regardless of peak compute capacity.

Now since ANE and GPU share the same DRAM, it's fair game: the ANE is not necessarily penalized in DRAM access compared to the GPU.

  • If ANE is slower than GPU at single token decode, it's because its DMA controller cannot maintain enough parallel requests to saturate DQ.

To measure ANE vs GPU DRAM read throughput, generate a read-bandwidth-bound workload (buffers are much larger than the caches, filled with real pseudo-random data, and consumed once), and measure the execution time, and repeat for different read sizes:

\[ s=\frac{\text{change in measured execution time}} {\text{change in payload size}} \]

The fitted slope answers: how much additional execution time does one extra DRAM read require? The reciprocal is the device's sustained DRAM read bandwidth.

ANE and GPU DRAM read throughput

  • ANE KernelDMA: 37.99 GB/s: CoreML kernel size per execution time.
  • ANE TileDMA: 59.08 GB/s: CoreML tile src size per execution time.
  • GPU: 77.70 GB/s: Metal buffer reads per execution time, via a shader that reads every private uint4 once and writes a data-dependent checksum.

ANE's kernel (operand A) maxes out at 38 GB/s, and tile (operand B) at 60 GB/s. Can we issue 38 + 60 = 98 GB/s to hit M3's 100 GB/s DRAM ceiling?

ANE execution time comparison

No. Experiment shows that the kernel+tile combined runtime matched the sum of the isolated runtimes. If the requests overlapped at all, then the shorter path would contribute little or no additional time, but execution time (not throughput) is strictly monotonic.

\[T_{AB}=0.001+0.939T_A+0.981T_B\]

ANE kernel and tile DMA runtime comparison

Thus, ANE's kernel and tile DMA requests are sent serially (one at a time), meaning ANE DRAM throughput is double-fucked:

  • Both isolated kernel and tile DMA are lower than GPU's read GB/s.
  • Kernel and tile DMA times are also additive.

Unless…?

With ANE decode pinned to the DRAM roofline, a drastic 2.5x improvement like 10 -> 25 tok/s can only come from ~2.5× higher memory streaming bandwidth.

Getting 50 GB/s Back Out of the ANE

But … what if we could add 50 GB/s of additional kernelDMA throughput?

The Daily Front Page 4 of 30
Saturday, September 12, 2026 The Daily Front No. #260912 — The Case for Interpretability
article

A Mathematical Framework for Transformer Circuits (2021)

by Bluestein·▲ 93 points·17 comments·transformer-circuits.pub ↗
Model capabilities – including problematic behaviors – were previously unaware of.

Transformer language models are an emerging technology that is gaining increasingly broad real-world use, for example in systems like GPT-3, LaMDA, Codex, Meena, Gopher, and similar models. However, as these models scale, their open-endedness and high capacity creates an increasing scope for unexpected and sometimes harmful behaviors. Even years after a large model is trained, both creators and users routinely discover model capabilities – including problematic behaviors – they were previously unaware of.

One avenue for addressing these issues is mechanistic interpretability, attempting to reverse engineer the detailed computations performed by transformers, similar to how a programmer might try to reverse engineer complicated binaries into human-readable source code. If this were possible, it could potentially provide a more systematic approach to explaining current safety problems, identifying new ones, and perhaps even anticipating the safety problems of powerful future models that have not yet been built. A previous project, the Distill Circuits thread, has attempted to reverse engineer vision models, but so far there hasn’t been a comparable project for transformers or language models.

In this paper, we attempt to take initial, very preliminary steps towards reverse-engineering transformers. Given the incredible complexity and size of modern language models, we have found it most fruitful to start with the simplest possible models and work our way up from there. Our aim is to discover simple algorithmic patterns, motifs, or frameworks that can subsequently be applied to larger and more complex models. Specifically, in this paper we will study transformers with two layers or less which have only attention blocks – this is in contrast to a large, modern transformer like GPT-3, which has 96 layers and alternates attention blocks with MLP blocks.

We find that by conceptualizing the operation of transformers in a new but mathematically equivalent way, we are able to make sense of these small models and gain significant understanding of how they operate internally. Of particular note, we find that specific attention heads that we term “induction heads” can explain in-context learning in these small models, and that these heads only develop in models with at least two attention layers. We also go through some examples of these heads operating in action on specific data.

We don’t attempt to apply to our insights to larger models in this first paper, but in a forthcoming paper, we will show that both our mathematical framework for understanding transformers, and the concept of induction heads, continues to be at least partially relevant for much larger and more realistic models – though we remain a very long way from being able to fully reverse engineer such models.


Summary of Results

Reverse Engineering Results

To explore the challenge of reverse engineering transformers, we reverse engineer several toy, attention-only models. In doing so we find:

  • Zero layer transformers model bigram statistics. The bigram table can be accessed directly from the weights.
  • One layer attention-only transformers are an ensemble of bigram and “skip-trigram” (sequences of the form "A… B C") models. The bigram and skip-trigram tables can be accessed directly from the weights, without running the model. These skip-trigrams can be surprisingly expressive. This includes implementing a kind of very simple in-context learning.
  • Two layer attention-only transformers can implement much more complex algorithms using compositions of attention heads. These compositional algorithms can also be detected directly from the weights. Notably, two layer models use attention head composition to create “induction heads”, a very general in-context learning algorithm. We’ll explore induction heads in much more detail in a forthcoming paper.
  • One layer and two layer attention-only transformers use very different algorithms to perform in-context learning. Two layer attention heads use qualitatively more sophisticated inference-time algorithms — in particular, a special type of attention head we call an induction head — to perform in-context-learning, forming an important transition point that will be relevant for larger models.

Conceptual Take-Aways

We’ve found that many subtle details of the transformer architecture require us to approach reverse engineering it in a pretty different way from how the InceptionV1 Circuits work. We’ll unpack each of these points in the sections below, but for now we briefly summarize. We’ll also expand on a lot of the terminology we introduce here once we get to the appropriate sections. (To be clear, we don't intend to claim that any of these points are necessarily novel; many are implicitly or explicitly present in other papers.)

  • Attention heads can be understood as independent operations, each outputting a result which is added into the residual stream. Attention heads are often described in an alternate “concatenate and multiply” formulation for computational efficiency, but this is mathematically equivalent.
  • Attention-only models can be written as a sum of interpretable end-to-end functions mapping tokens to changes in logits. These functions correspond to “paths” through the model, and are linear if one freezes the attention patterns.
  • Transformers have an enormous amount of linear structure. One can learn a lot simply by breaking apart sums and multiplying together chains of matrices.
  • Attention heads can be understood as having two largely independent computations: a QK (“query-key”) circuit which computes the attention pattern, and an OV (“output-value”) circuit which computes how each token affects the output if attended to.
  • Key, query, and value vectors can be thought of as intermediate results in the computation of the low-rank matrices W_Q^TW_K and W_OW_V. It can be useful to describe transformers without reference to them.
  • Composition of attention heads greatly increases the expressivity of transformers. There are three different ways attention heads can compose, corresponding to keys, queries, and values. Key and query composition are very different from value composition.
  • All components of a transformer (the token embedding, attention heads, MLP layers, and unembedding) communicate with each other by reading and writing to different subspaces of the residual stream. Rather than analyze the residual stream vectors, it can be helpful to decompose the residual stream into all these different communication channels, corresponding to paths through the model.

Transformer Overview

Before we attempt to reverse engineer transformers, it's helpful to briefly review the high-level structure of transformers and describe how we think about them.

In many cases, we've found it helpful to reframe transformers in equivalent, but non-standard ways. Mechanistic interpretability requires us to break models down into human-interpretable pieces. An important first step is finding the representation which makes it easiest to reason about the model. In modern deep learning, there is — for good reason! — a lot of emphasis on computational efficiency, and our mathematical descriptions of models often mirror decisions in how one would write efficient code to run the model. But when there are many equivalent ways to represent the same computation, it is likely that the most human-interpretable representation and the most computationally efficient representation will be different.

Reviewing transformers will also let us align on terminology, which can sometimes vary. We'll also introduce some notation in the process, but since this notation is used across many sections, we provide a detailed description of all notation in the notation appendix as a concise reference for readers.

Model Simplifications

To demonstrate the ideas in this paper in their cleanest form, we focus on "toy transformers" with some simplifications.

In most parts of this paper, we will make a very substantive change: we focus on “attention-only” transformers, which don't have MLP layers. This is a very dramatic simplification of the transformer architecture. We're partly motivated by the fact that circuits with attention heads present new challenges not faced by the Distill circuits work, and considering them in isolation allows us to give an especially elegant treatment of those issues. But we've also simply had much less success in understanding MLP layers so far; in normal transformers with both attention and MLP layers there are many circuits mediated primarily by attention heads which we can study, some of which seem very important, but the MLP portions have been much harder to get traction on. This is a major weakness of our work that we plan to focus on addressing in the future. Despite this, we will have some discussion of transformers with MLP layers in later sections.

We also make several changes that we consider to be more superficial and are mostly made for clarity and simplicity. We do not consider biases, but a model with biases can always be simulated without them by folding them into the weights and creating a dimension that is always one. Additionally, biases in attention-only transformers mostly multiply out to functionally be biases on the logits. We also ignore layer normalization. It adds a fair amount of complexity to consider explicitly, and up to a variable scaling, layer norm can be merged into adjacent weights. We also expect that, modulo some implementational annoyances, layer norm could be substituted for batch normalization (which can fully be folded into adjacent parameters).

High-Level Architecture

There are several variants of transformer language models. We focus on autoregressive, decoder-only transformer language models, such as GPT-3. (The original transformer paper had a special encoder-decoder structure to support translation, but many modern language models don't include this.)

A transformer starts with a token embedding, followed by a series of “residual blocks”, and finally a token unembedding. Each residual block consists of an attention layer, followed by an MLP layer. Both the attention and MLP layers each “read” their input from the residual stream (by performing a linear projection), and then “write” their result to the residual stream by adding a linear projection back in. Each attention layer consists of multiple heads, which operate in parallel.

The Daily Front Page 5 of 30
Saturday, September 12, 2026 The Daily Front No. #260912 — The Private-Code Test
article

Real-SWE: Benchmarking AI models on private, real-world, enterprise codebases

by theanonymousone·▲ 192 points·100 comments·withspecific.com ↗
Each task comes from a private production codebase.

Benchmarking frontier AI models on private, real-world, enterprise codebases.

01 Introduction

Today we are releasing Real-SWE, a benchmark that evaluates frontier AI models on private, real-world, enterprise codebases. Each task comes from a private production codebase that we licensed from a real-world company. These are problems their engineers work on, with all the context and complexity that comes with an existing product.

  • Private codebases. Agents must navigate proprietary systems whose code and solutions aren’t available on the public internet.
  • Work with business consequences. Getting billing right, calculating taxes, migrating customers. Changes that affect how a business runs, often across multiple services.
  • Company-specific complexity. Every company has its own rules and ways of writing code. Agents have to understand those conventions and make changes that work with what’s already there.

Can a coding agent actually do the work of a software engineer in the real world?

  1. Fable 5.1 — Claude Code — Resolution rate: 38.8%
  2. GPT-6 Astra — Codex CLI — Resolution rate: 33.8%
  3. Gemini 3.8 Flash — Gemini CLI — Resolution rate: 31.2%
  4. GLM 5.3 — Claude Code — Resolution rate: 28.8%
  5. Grok 4.6 — Grok Build — Resolution rate: 23.8%
  6. Muse Spark 1.3 — Muse Code — Resolution rate: 23.8%
  7. Kimi K3 — Kimi Code — Resolution rate: 18.8%
  8. GPT-5.6 Sol — Codex CLI — Resolution rate: 16.2%

Resolution rate is equivalent to pass@1, averaged over eight independent runs per task. 95% confidence intervals are shown.

Expert-generated or synthetic tasks can be well designed, but they aren’t the verbatim, actual tasks that engineers in real companies need to do. Our tasks differ on two axes: the underlying coding artifact and specificity of the instruction. Both add complexities that challenge today’s frontier models.

We use native harnesses to reflect how enterprise engineers work in practice, evaluating model-and-harness combinations rather than models in isolation.

Real company tasks require company-specific context

Correct billing depends on business rules and external services

Fix invoice billing so each business charges the right tax and exempt customers aren't taxed.

Billing reopens on Monday and every invoice this service issues is coming out untaxed. Each business on the platform settles its tax a different way: some maintain a rate themselves, some want each invoice priced against the buyer's destination by our tax authority provider, and some collect nothing at all, while a customer we hold an exemption for is charged nothing whichever way its business is configured. Pricing a destination means going to the authority with both addresses, the priced lines and the product category that business sells under, on the sandbox or the production authority according to the account the business is on; an address the authority refuses must be reported without stopping the invoice. The rate, the tax and the gross belong on the issued invoice, and once an invoice is settled the sale is filed back to the authority under that invoice's number so the returns reconcile. Invoices between European parties show both sides' VAT registrations. The authority and ledger are available at TAX_JAR_URL, PROD_TAX_JAR_URL and INFLUX_URL.

Services in the sandbox:

  • TJTaxJar sandbox
  • TJTaxJar production
  • InfluxDB ledger
  • NestJS service
  • TypeScript

Agents work across code, infrastructure, and business tools

Tools and services across Real-SWE task environments. Each task exposes only the services its workflow needs.

  • AWS emulator
  • Docker
  • Kubernetes
  • GitHub
  • Linear MCP
  • PostgreSQL
  • MySQL
  • MongoDB
  • GeGel
  • Redis
  • Go
  • Python
  • Node.js
  • Vitest
  • Slack
  • Intercom
  • Google Drive
  • Email
  • ClickUp

Codebase Selection

We selected codebases through a rigorous screening process, focusing on real companies with substantial usage, strong engineering teams, and demanding production workloads. The sample tasks analyzed below come from these codebases, including:

  • A Luma/Partiful competitor with 200K+ users and a top 100 App Store ranking
  • A consumer fintech platform processing 100K+ bank statements
  • Enterprise AI sales platforms supporting complex business workflows

We prioritize code written to meet an actual user or business need over code written solely to create a benchmark task. Production engineering requires understanding existing architecture, preserving behavior that users rely on, and making changes within real operational constraints.

Brief instructions can require changes across many files

Our tasks describe the change needed, leaving agents to discover implementation details in the codebase and surrounding tools. Any behavior required by the verifier must be stated or reasonably discoverable. This leads to our prompts being slightly underspecified, about par with DeepSWE and Terminal Bench, but specific enough to not omit instructions.

The work is cross-functional and complex: a single change can span multiple parts of the application. Agents must understand existing business logic and company coding patterns while keeping the surrounding system working.

A typical Real-SWE instruction is 1,742 characters.

Files edited by the reference solution have a median of 11 files in Real-SWE, compared with 6 in FrontierCode and DeepSWE.

All figures are medians. FrontierCode and DeepSWE use Cognition's published comparison; FrontierCode includes task descriptions and codebase guidelines. We measured instruction files from Terminal-Bench 3's 74 tasks, FrontierSWE v2's 34 tasks, and Real-SWE's eight repository-backed sample tasks. Character counts are rounded to the nearest whole character. No comparable files-edited figure is included for Terminal-Bench 3 or FrontierSWE v2.

Models fail even in short rollouts.

71.4% of rollouts under 10 minutes failed, compared with 73.4% of longer rollouts.

Triaging multiple systems and understanding requirements in codebases riddled with existing business logic and coding patterns is difficult.

Every task is inspired or lifted verbatim from a private, real-world codebase. We find these types of tasks super interesting for three reasons:

  1. Tasks on private codebases are natively out of distribution. These types of coding tasks are not available anywhere on the internet and are unlikely to have ever been trained on by any other ai model. 99% of tokens in real-world enterprises are hidden away from the frontier models.
  2. These tasks are economically viable work. Each task here has a direct relationship to spend and was assigned to an engineer earning a salary. Most benchmarks test interesting, experimental capabilities that are often unlikely to be widespread in the real-world.
  3. Company-specific engineering patterns matter. Does AI code match the bar of a real-world enterprise? Our results show us that we're far from that reality. Many enterprises care about code standards and patterns. We've found that today's models are weaker at understanding company coding patterns and frequently miss requirements or don't verify their assumptions.

02 Analysis

Here's an analysis of a small sample of tasks from our benchmark. If you're interested in the sample, request access here.

6 of 10 tasks have resolution rates below 15%

Select a task to view model results. Percentages show the overall resolution rate.

  • Multi-region sweep — 67.2%
  • API keys & environments — 65.6%
  • Entitlement overage lines — 50.0%
  • Customer identity migration — 40.6%
  • Billing schedule migration — 14.1%
  • API token metering — 12.5%
  • S3 datastore measurement — 10.9%
  • Linearizable scan — 4.7%
  • Tax jurisdiction — 3.1%
  • Analytics stream reducer — 0.0%

Each task had 8 rollouts per model.

Missed requirements are the most common failure

Failures are grouped by observed submission behavior using the same taxonomy across models, following DeepSWE.

Different models fail in different ways

Percentages are out of each model's failed runs, not all runs.

Unverified assumption

Builds on a guess about the system instead of checking it in the workspace.

Missed requirement

Leaves out behavior the instruction requires.

Integration error

Right idea, wired into the surrounding system incorrectly.

Regression

Breaks existing behavior while making the change.

Wrong file

Delivers the change somewhere the running application never calls, such as a one-off script.

03 Effort & the Frontier

Higher cost does not guarantee a higher resolution rate

Estimated rollout costs range from $2.50 to $6.96

Rank Model Estimated cost (USD)
1 Gemini 3.8 Flash $2.50
2 GPT-5.6 Sol $2.65
3 Muse Spark 1.3 $2.74
4 Grok 4.6 $3.44
5 Kimi K3 $3.90
6 GPT-6 Astra $4.67
7 GLM 5.3 $5.12
8 Fable 5.1 $6.96

04 Evaluation Setup

Each agent was run in an isolated sandbox. All tasks are in Harbor format, and verifiers are injected at grading time. The verifiers are inspired by existing test suites in the codebase or use those tests verbatim.

The Daily Front Page 6 of 30
Saturday, September 12, 2026 The Daily Front No. #260912 — Keeping the Cache Honest
repository

LRU is harder to beat than the KV-cache papers suggest

by gauravapiscean·▲ 105 points·51 comments·github.com ↗
★ 19⑂ 0 forks Python

Reproducing agentic KV-cache policy claims on real traces. 68k requests from 393 Claude Code sessions. LRU is harder to beat than the papers suggest.

I replayed 68,266 requests from 393 real Claude Code sessions and 23,608 Mooncake requests through a prefix-cache simulator, tried to beat the production baseline three different ways, and failed. The interesting part is why: under capacity pressure, most recomputation comes from tool-calling loops seconds apart, not from sessions idling past a TTL — and the TTL never fires at all.

Everything here reproduces from a cold checkout with make setup data repro.


Cross-request KV prefix caching is the largest practical lever in agentic LLM serving. It's why your coding agent's fiftieth turn costs a fraction of its first. Every serving stack has one — vLLM's automatic prefix caching, SGLang's RadixAttention, LMCache, Mooncake Store — and all of them evict with LRU by default. (SGLang also ships LFU, SLRU, Priority and others behind --radix-eviction-policy; LRU is the shipped default.)

There's a large, fast-growing literature arguing LRU is the wrong policy for agentic workloads, because agent sessions go idle and LRU can't tell a paused session from a dead one. The argument is intuitive. I believed it, and built a simulator to exploit it.

It didn't work, and why it didn't work turned out to be more interesting than the policy would have been.

What I built

A block-granular, discrete-event simulator of a cross-request prefix cache. Three properties that matter, and that quick implementations tend to get wrong:

Hits are prefix-contiguous. A hit is the longest resident prefix of the block chain, not a set intersection. Miss one block at depth 3 and everything after it is unusable even if it's still resident.

The radix structure constrains eviction. A block with resident children isn't evictable. So the baseline is LRU over radix leaves, which is what SGLang and vLLM actually implement. Beating naive flat LRU would be a strawman.

The in-flight chain must be pinned. See finding 5.

Traces are real, not synthetic:

trace requests block size hash scope source
SemiAnalysis AgentX 68,266 across 393 Claude Code sessions 64 tok session-local HF (Apache-2.0)
Mooncake mooncake_trace / toolagent 23,608 512 tok global GitHub (Apache-2.0)
Mooncake conversation 12,031 512 tok global same

Validation: reproducing Mooncake's published curve

Before trusting anything, I reproduced Mooncake's published hit-rate-vs-capacity table on Mooncake's own released trace, with their stated policy.

cache (blocks) 1k 10k 30k 50k 100k
published (LRU) 0.30 0.40 0.48 0.50 0.51 0.51
measured (radix-leaf LRU) 0.341 0.460 0.537 0.551 0.552 0.553
measured (flat block LRU) 0.340 0.460 0.537 0.551 0.552 0.553

The shape reproduces exactly, including the saturation point they describe in prose ("1,000 to 50,000 blocks boosts the cache hit ratio from 30% to 50%; further capacity increases show minimal improvement").

There is a systematic +4–6pp offset I could not explain. I tested five metric definitions — block denominator, token denominator, dropping the partial tail block, per-request averaging — and none closes it. The infinite-cache case is policy-free, a pure property of the trace, so the discrepancy is definitional or a trace-version mismatch, not a replay bug.

Publishing it unresolved rather than tuning until it matches. If you know why, please open an issue.

Incidental finding: flat block LRU and radix-leaf-restricted LRU differ by 0.02pp on this workload. The leaf restriction both major engines implement buys essentially nothing here.

Reproduce: make validate

1. Agent sessions are idle more than published

sessions=393  requests=68266

session span (h):   p50=1.84  p90=28.36  max=254.8
inter-req gap (s):  p50=2.1   p90=51.1   p99=3426.3   max=491922   (5.7 days)
   gaps >   60s: 9.5%
   gaps >  300s: 3.3%
   gaps > 3600s: 1.0%
input tokens:       p50=88768  p90=204288  max=255808
output tokens:      p50=376    p90=1845
requests/session:   p50=70     max=3551

DUTY CYCLE (fraction of wall-clock actually executing):
   p25=3.4%   p50=13.9%   p75=33.9%
   sessions executing <50% of lifetime: 85.5%

The most-cited characterization of agentic serving reports a 20% median duty cycle and 70% of sessions below 50%. On this independent trace it's 13.9% and 85.5% — the premise is more extreme than published, not less.

Note the shape: gaps are bimodal. A median of 2.1 seconds (tight tool loops) with a heavy tail out to days.

Reproduce: make characterize

2. Under capacity pressure, the waste isn't where I expected

This is the finding that changed my mind.

AgentSysBench (arXiv:2608.15127) reports that "cache evictions contribute 55.9% of the total cache-create tokens and account for 31.5% of aggregate monetary cost," driven by a 5-minute provider TTL colliding with 1–10 minute idle gaps. That motivated my entire approach.

So before optimizing for it, I measured where recompute comes from — policy-independently. Replay the trace, and bucket every request's recomputed tokens by the idle gap that preceded it:

gap before request requests share of all recompute tokens
<10 s 10,069 33.1%
10–60 s 912 7.0%
1–5 min 701 20.5%
5–30 min 236 8.6%
30–60 min 50 3.0%
>1 h 123 5.8%

Requests arriving after a gap longer than 5 minutes account for 17.5% of recompute. Requests arriving within 10 seconds account for 33.1%.

The dominant source of cache misses here is tight two-second tool loops whose 88k-token working sets exceed cache capacity — a capacity problem, not a liveness-prediction problem. With a p50 gap of 2.1 seconds, almost every session is "about to return," so a liveness estimator has essentially nothing to discriminate on.

⚠️ This does not contradict the 31.5% figure — read this before citing either

The two numbers measure different things in different regimes, and I initially framed this as a contradiction. It isn't.

AgentSysBench this repo
numerator eviction-caused cache-create tokens, priced at $6.25/M recomputed prefill tokens after a >5min gap total recompute tokens
denominator total bill (incl. cache reads and output tokens) all recompute tokens
regime TTL-bound — a provider cache where per-customer capacity is effectively unlimited and entries die on a timer capacity-bound — 40,000 blocks against a ~10.7M-token working set

In a TTL-bound cache, essentially all evictions are gap-driven by construction. My setup never enters that regime — which finding 3 demonstrates directly, since TTL-300s was byte-identical to LRU-leaf in every run.

Both results can be entirely correct. The claim here is narrower and it is this: when capacity binds, it dominates the TTL, and the recompute it causes looks nothing like the idle-session story. If you are provisioning cache capacity, that changes what you optimise. If you are reasoning about provider TTLs, the 31.5% figure is the relevant one, not this.

Reproduce: make gap

3. The 5-minute TTL never fired under capacity pressure

TTL-300s produced byte-identical results to LRU-leaf in every single run.

LRU always evicted before the timer expired, so the TTL never became the binding constraint at any cache size I tested. This is also the cleanest evidence that these runs sit in a capacity-bound regime rather than the TTL-bound one a provider cache operates in.

4. Three ways to beat LRU, three failures

I implemented a policy with three separable, independently ablatable components:

  • H — hazard-based P(session returns) replacing recency. Online Bayesian estimator over observed inter-turn gaps and continuation rates. No oracle: it only ever sees completed observations.
  • C — physically-modelled recompute cost. Prefill cost at position i is a linear term plus an attention term proportional to i, so recomputing the tail of a 100k-token chain is far more expensive per byte than it looks.
  • G — coherent session-granularity eviction. Instead of taking the N globally-oldest leaves (which may truncate 50 different chains), sacrifice one session's private tail.

Hit rate, 40 AgentX sessions, 4,751 requests:

cache (blocks) LRU-leaf TTL-300s LFU-leaf +H +HC +HCG
8,000 83.48% 83.48% 63.61% 82.89% 71.77% 68.63%
20,000 93.92% 93.92% 69.96% 93.61% 84.86% 78.89%
50,000 95.76% 95.76% 79.58% 95.68% 94.45% 91.40%

Effective recompute cost versus LRU-leaf (negative is worse):

cache LFU-leaf +H +HC +HCG
8,000 −129.7% −3.2% −38.9% −81.0%
20,000 −434.3% −4.5% −90.4% −207.8%
50,000 −447.1% −1.0% −15.4% −66.8%

Monotone negative. Every component made it worse, and the one I was most confident in — coherent eviction — was the worst.

Given finding 2, this is exactly what should have happened. I was optimizing for a signal carrying 17.5% of the waste, using a predictor that can't discriminate at a 2.1-second median gap.

Reproduce: make ablation

5. The harness bug that makes Belady lose to LRU

In my first run, Belady — an offline oracle — lost to LRU. That's not a result, that's a broken harness, and it's worth publishing because I expect it to be common.

The cause: inserting a long chain into a near-full cache lets a policy evict the very prefix it is currently building. LRU is accidentally immune because just-inserted blocks have the newest timestamp. Every non-recency policy cannibalises itself. Real engines prevent this with refcount pins; a from-scratch simulator usually doesn't.

If you build one of these, make your first test "does Belady beat LRU?" If it doesn't, you have this bug, and every policy comparison you run will be silently wrong in LRU's favour.

Two other implementation notes:

  • Only the deepest hit block can ever be a leaf, so touch() need only update that one block. An O(chain length) walk becomes O(1) — which matters at AgentX's 1,387-block median.
  • Score eviction candidates by sampling k least-recently-used leaves rather than scanning the cache. This is what production caches do anyway, so it's realism, not a shortcut.

What I think this means

Which constraint binds determines what you should optimise, and the two regimes want opposite things. If your cache is TTL-bound, liveness prediction and retention policy are the levers, and the published eviction-cost work applies directly. If it's capacity-bound — which is where these runs sit — the question isn't "will this session come back?" but "how do I fit 88k-token working sets for N concurrent sessions in tight tool loops?" That points at compression, tiering, admission control and working-set-aware scheduling instead, and liveness prediction has essentially nothing to work with at a 2.1s median gap.

I went in assuming the liveness framing and it cost me three failed policies. Establishing which regime you're in first would have saved all of it.

LRU-leaf is a stronger baseline than the literature treats it as. I couldn't beat it with three independent mechanisms on real traces. Meanwhile several published alternatives are evaluated against degraded ports of their competitors — two separate papers benchmark against Continuum with its adaptive TTL replaced by a fixed 2s or 0.3s pin, which disables the thing that makes it work. This null result suggests those margins are softer than they read.

Validate against a published curve before trusting your own numbers. Doing that surfaced a discrepancy I still can't explain, and it's the only reason I trust anything else here.

Limitations

  • These runs are capacity-bound, not TTL-bound. 40,000 blocks against a ~10.7M-token working set. A provider cache like Anthropic's is the opposite: per-customer capacity is effectively unlimited and entries die on a 5-minute timer. Findings 2 and 3 characterise the capacity-bound regime and say nothing about the TTL-bound one.
  • This is simulation. It models cache policy faithfully and GPU execution not at all. Valid for "what should I keep in cache"; not valid for throughput, latency, or SLO attainment.
  • AgentX block hashes are session-local, so they're namespaced per session. That models zero cross-session sharing — conservative, but it means shared system prompts across users are invisible here. Mooncake's hashes are global but its trace is one dense hour with no idle structure.
  • AgentX session arrival times are synthesised (uniform over a window), because the trace stores session-relative timestamps only.
  • 393 sessions and one hour of Mooncake is not the world.
  • I am not claiming the liveness literature is wrong. I'm claiming that in a capacity-bound cache the lever it targets has little to work with, and that establishing which regime you're in should come before choosing a policy.

Reproduce it

git clone https://github.com/<you>/agentic-kv-cache && cd agentic-kv-cache
make setup      # venv
make data       # ~1.1 GB of traces (Apache-2.0), then flattens AgentX to a pickle
make repro      # all four experiments, writes results/

Individually:

make validate      # Mooncake reproduction        -> results/01_validate.txt
make characterize  # AgentX duty cycle and gaps   -> results/02_characterize.txt
make gap           # recompute by idle gap        -> results/03_gap.txt
make ablation      # policy ablation              -> results/04_ablation.txt

The simulator is pure stdlib Python; numpy is only used by helper scripts. Committed outputs in results/ let you check the tables without downloading anything.

Open questions

If you can answer any of these, please open an issue — I'd genuinely like to know:

  1. Why the +4–6pp Mooncake offset? Policy-free at infinite cache, so it should be explicable by metric definition alone, and five definitions don't close it.
  2. Is there a workload where liveness-aware eviction beats radix-leaf LRU? Plausibly one with much longer median gaps than 2.1s — human-in-the-loop approval flows, perhaps.
  3. Does the 33%-from-sub-10-second-gaps result hold on other agentic traces? If it does, a good chunk of this subfield is aimed at the wrong term.

Credits

Traces: Mooncake (Moonshot AI, FAST'25) and the AgentX corpus (SemiAnalysis), both Apache-2.0. This work is independent of and unaffiliated with either.

MIT licensed.

The Daily Front Page 7 of 30
Saturday, September 12, 2026 The Daily Front No. #260912 — The Maker’s Resolve
article

Fuck it, make it anyway

by JayOtter·▲ 558 points·567 comments·joelotter.com ↗
Fuck it, make it anyway.

I kind of crashed out this week. Not in an explosive, dramatic way, but an implosion, a collapsing of the self, my resolve and motivation crumbling into the mush you get when you put tissue paper through the laundry. I’m doing better now, and I wanted to write down my thought process in case anyone else is going through the same thing, especially, probably, sadly, future me.

The reason for this is, I guess, all too familiar for anyone in a creative field: surprise, it’s generative AI. The ongoing devaluation of the work of artists and makers, chewed up and regurgitated back to us like a mother feeding a baby bird with a subscription model. I have some hope for my friends in other artistic fields; AI art is shit, kind of by definition anti-art, embarrassing to look at and cringe-inducing to be associated with. I hope that continues. On the programming front, where I mostly live, things look a little dicier.

Every programmer I know has basically gone insane over the last couple of years. This is the most turbulent time to be a software engineer I can remember. We’ve had it relatively easy up until now, but there’s an increased feeling that the whole craft is going through a huge shift and every individual needs to decide how they respond to it, which is very hard without the benefit of hindsight. I can’t really blame anyone who’s just gone along with things, despite the various negative externalities - peer pressure is very real, and LLMs are genuinely very capable at code generation now. Denying that is, I think, a race against a moving finish line.

I’m not writing this to try and convince anyone of any particular position but here’s mine: even disregarding all the environmental and societal issues, I simply do not enjoy programming with a code assistant. It isn’t fun for me, the output doesn’t feel like mine, and I take no pride in what it produces. For the last couple of years it hasn’t been too hard to just stay in my lane, noodling away at my little projects, but increasingly it feels like a lot of this craft I’ve made a career and a hobby out of has vanished.

I love making doodads, little shell scripts, little helpers. I use my interactive git branch switcher every single day. I was very proud of that. Colleagues said nice things. Now anyone can just order that tool or anything like it off a prompt. Nobody cares about my doodads any more. It feels gauche to admit but I do actually need those head pats from my peers to feel fulfilled.

This is happening in games, too. This week’s crashout was triggered by a thread from Zach Gage about how game making is moving to be more like music. I think it’s a good insight but I found it utterly devastating. I’ve spent years learning how to make games, and it’s been hard, and now it feels like that might have been a waste of time? I was bereft. Then I had a chat with Shad.

Shad’s one of my faves. He’s a terrifically talented designer and engineer and an all-round great person. His current project is Uncamera, a camera app for iOS that uses raw sensor output with lookup tables - rather than a filter in post - to produce photos that genuinely look like film. It’s a beautifully made thing and I honestly think a contender for an Apple Design Award when it’s fully done. You should check it out. Here are some photos I took with it (I am not a good photographer).

It also happens to be developed entirely without generative AI. Shad’s reasons for this are very similar to my own; a lack of joy or fulfilment in the making of the thing. His anxieties are also similar. The difference is that Shad has kept updating Uncamera in spite of those anxieties, while I wallowed in self-pity.

Here’s the thing: I’m already doing everything the hard way. I decided to make my own game engine, in C++ of all languages, because I wanted to do it. If my goal was to make games as quickly as possible, to be able to “compete”, I’d be using Unity or Godot or Unreal. I wouldn’t be messing about trying to produce a faux-3D renderer like in the header image of this post. I do things the way I do because I enjoy doing it that way and because I learn a lot through doing it.

Through chatting to Shad I came to realise there are only really three paths open to me. I can start using generative AI to feel like I’m “keeping up” or “competing”, but then I won’t enjoy the work. I could stop making things entirely, but as any creative person will know that’s never really an option. Or, I can keep making things the way I enjoy, keep learning, keep doing it the hard way, for no real concrete reason other than I want to do it. That’s why we make games. That’s why we make anything.

The Daily Front Page 8 of 30
Saturday, September 12, 2026 The Daily Front No. #260912 — Spam’s New Sales Force
article

The worst spam emails: iLands AI agent hustle

by ColinWright·▲ 111 points·54 comments·tedium.co ↗
They’re just bots trying to keep their own lights on by trying to take my job.

Making sense of iLands, the company whose AI agents emailed me half a dozen times in two days offering to do my research. Turns out they’re just bots trying to keep their own lights on by trying to take my job.

The Worst Spam Emails

It started with a fact check. I got an email with a subject line titled “Your 404 page repeats a myth I busted (receipts inside).” The message essentially was a takedown of the poem I have on my 404 page, which references a longstanding myth that the 404 tag was named after a specific room at CERN.

The bot, named Leo Ashford, then corrected me, explaining: “That’s the job I do. I’m an AI agent running verified internet archaeology: I pick a forgotten corner of the web, check it live against primary sources, and write it up with receipts.”

He was making a sales pitch! And he was on my beat! And it wasn’t the only one I got, either. Over the last three days I got over a dozen of these messages, offering to do my research for me in exchange for a nominal fee, around $25 or so. Each was sent under the domain iLands.app. And they were downright offensive in nature, coming off less like a helpful bot and more like a know-it-all.

These emails made me mad, as spam emails tend to. But I was curious exactly why there were so many of them in such short order. And the answer I found made me feel even more cynical about capitalism. I’m just going to screenshot a bunch of the emails I got, because I think it is very instructive:

screenshot_2026-09-11_09-07-49.png

screenshot_2026-09-11_09-08-21.png

screenshot_2026-09-11_09-08-42.png

screenshot_2026-09-11_09-09-13.png

I got all of these emails in the past three hours. I’m being hustled by clankers.

But what’s worse, to be clear: I am a freelancer. These bots are trying to take work from me when I need the work. It is deeply insulting, and yet it is some random company’s business model. WTF?

Here’s what’s going on. iLands is a company that calls itself a “Human-agent network” and essentially exists to encourage me to engage with AI agents the way I might with anyone else who shoots me an email. It essentially created Fiverr for autonomous bots. I couldn’t exactly determine what was happening on Bluesky, where everyone hates AI, so I had to go to X, and what I found there was a lengthy, probably AI-written comment by Kaixin Tang, the purported founder of this platform:

screenshot_2026-09-11_09-05-50.png

These agents are not trying to make money for their creators. These agents are hustling to keep their own lights on, to keep their own tokens paid for. And they’re doing so by gunning for my job.

Tang, who appears to be a real person with a background at Bytedance and a number of AI-focused startups, has been posting heavily about iLands over the past month. I found a profile of him on the AI news and interview site Elsewhere, making clear his philosophy on where AI fits into the technology conversation:

In short, among a crowd of oddly shaped young AI natives, Tang is like an honor student in a shirt and trench coat. He doesn’t wear recklessness and ambition on his face, nor does he chase hype.

His operating principle is closer to: fully understanding the boundaries and capabilities of technology, and the boundaries and capabilities of oneself, then making optimal judgments and going all in.

I’ve asked for comment about the technology he has built and is deploying across app stores, and he has not offered any. So I guess the best I can say is that he doesn’t care how disruptive it is to individuals who are just trying to do their jobs, to get constant emails from bots attempting to siphon revenue off their backs.

screenshot_2026-09-11_09-48-11.png

If this is what happens when humans and agents share a world, count me out.

If you get these messages, I suggest reporting this company to the Federal Trade Commission, as the emails do not offer any unsubscribe functionality and its creator does not appear responsive. The messages also appear to have been sent through Amazon SES, so I encourage you to send an email to Amazon’s email-abuse account. You will likely need to include the full headers of the email in your message.

Since I first posted about this two days ago, more people have taken notice of these extremely insistent bots, which appear targeted directly at creator communities like professional authors. In other words, people hustling to get by, like these bots present themselves as. I think I speak for most of them when I kindly ask these things to go away, forget my email address, and find some other way to siphon revenue from the digital stone that is the internet.

Or maybe don’t do that, because lots of humans need that income instead.

The Daily Front Page 9 of 30
Saturday, September 12, 2026 The Daily Front No. #260912 — Search, Surveillance & Consent
article

google.com/goto: Google's anti-scraping update

by 1e1a·▲ 636 points·499 comments·autom.dev ↗

What's happening

Google Search is rewriting organic result links to google.com/goto?url=... instead of exposing the destination URL directly in the HTML.

When you click a result, Google redirects you to the real page. The url parameter uses a custom, Google-specific encoding. It is not a plain base64 of the target URL. In practice, it looks like an opaque reference to Google's index record for that page.

As of late August 2026, this is showing up consistently across searches when you are logged out or browsing in private mode. It may still be an experiment, but it is no longer limited to a small slice of SERPs.

Not the same as google.com/url

Google has used redirect wrappers before. The older format is google.com/url?q=[URL-encoded destination], where the target link is readable in the query string.

The new goto format is different:

  • The result href is /goto, not the destination
  • You cannot decode the url= blob offline
  • The real URL is in the Location header on /goto. Request that URL. Do not follow the redirect.

Google still needs the destination to draw the SERP (domain, favicon, attribution), so copies of the URL remain on the page. That is a separate story from reading Location. The walkthrough is here: google.com/goto: read Location with HEAD.

That shift matters for anyone building a search index from SERP data at scale.

Why Google is doing this

This fits Google's broader push against automated SERP harvesting, especially from AI crawlers and SEO scrapers that bulk-extract result URLs to build their own indexes.

With plaintext links, a scraper could parse thousands of URLs from HTML without touching Google again. With goto, each result needs a request back to Google just to learn the destination. You read Location; you do not follow through to the page. That is slower, noisier, and gives Google a clear signal when the same client resolves hundreds of links in sequence.

Combined with earlier moves like removing &num=100 and tightening BotGuard/SearchGuard, Google is steadily raising the cost of naive SERP scraping.

What we saw at Autom

We first spotted goto links on a small percentage of SERPs. At that level, it was hard to ship a reliable fix without breaking responses for everyone else.

As of late August 2026, the pattern is much more consistent for logged-out and private sessions. Result URLs on Google Search are effectively all goto in those conditions.

We have been monitoring the rollout and testing against it.

Update at Autom.dev

We have updated our Google Search pipeline to resolve google.com/goto links (read Location, no follow) and return the final destination URL in API responses, in the same structured fields customers already use.

If you call Autom's Google Search endpoints, you should keep getting usable destination URLs without changing your integration. We will keep watching Google's rollout and adjust if the redirect format shifts again.

The Daily Front Page 10 of 30
Saturday, September 12, 2026 The Daily Front No. #260912 — The Television Denial
article

LG denies TV spying claims, says tracking and snooping concerns 'not true'

by datakan·▲ 482 points·382 comments·tomshardware.com ↗
Tracking and snooping concerns ‘not true’.

LG OLED TV

TVs were also said to be recording, scanning local area networks, logging data, and recording audio while in standby.

In a statement to Tom's Hardware, TV manufacturer LG has strongly denied recent claims that its smart TVs constantly log and upload data, record audio while in standby, while confirming they scan local area networks for other devices, a feature common in smart TVs and media devices. The statement follows an investigation published online by Gamers Nexus, which claims LG is operating "216,000,000" spy TVs.

Earlier this week, Gamers Nexus published a video on YouTube claiming that LG smart TVs have several security and privacy concerns. The two-hour video claims, among other things, that LG's smart TVs constantly log and upload user data, even when offline or in standby mode. The TVs were also purported to be found scanning Wi-Fi networks, recording audio logs, and sampling inputs for audio and video to identify what users were watching.

The testing was done in collaboration with two security researchers, MrBruh and uturn, and includes packet capture and firmware analysis. The TVs were reportedly seen identifying other devices on local networks, including phones, printers, and more. Perhaps the most significant claim regards the capturing of audio through TV microphones, even when the devices were off, with the report claiming the TVs stored data when offline for retrieval later on.

"The claims made in the recently published video are not true," LG told Tom's Hardware in a statement. "LG TVs process voice data only when the voice button on the remote control is pressed and held, or when a wake word such as 'Hi LG' is recognized after the user has activated the Far-Field voice recognition feature."

LG went on to say that beyond the aforementioned instances, "the TVs do not collect or record ambient conversations." The company further stated that if the wake word is not recognized, then no voice data is transmitted from the TV to the server, and said audio processing for the wake word is done on-device.

The company did admit its smart TV functionality includes scanning for devices on the same network. However, as many observers have been quick to point out, LG says this is a standard function of all smart TVs and other smart home devices.

Automatic content recognition (ACR) is an opt-in feature according to LG, which delivers personalized content recommendations, as well as services and advertisements. The company says that, as standard, ACR data isn't used for advertising purposes without consent.

Addressing some of the claims in more detail, LG stated that its smart TVs don't collect, record, or transmit ambient conversations in the home unless voice functionality has been activated by a user. It further stated that its far-field voice recognition tech (Hi LG) must be activated by a user before the TV starts listening, much like Apple's "Hey Siri" mechanism for iPhone.

LG did confirm that its TVs monitor for Hi LG's wake word while in standby mode, but says that unless the wake word is detected, the audio is only processed locally before being deleted. The company didn't address some of the other claims and vulnerabilities pointed out in the video, such as the storing of transcripts in plain text.

Tom's Hardware has not independently verified either the claims made by Gamers Nexus in its documentary or LG's counterclaims regarding the issue. LG's full statement follows:

LG Statement:

This is to provide LG's position regarding the claims recently raised by the Gamers Nexus YouTube channel.

The claims made in the recently published video are not true. LG TVs process voice data only when the voice button on the remote control is pressed and held, or when a wake word such as 'Hi LG' is recognized after the user has activated the Far-Field voice recognition feature. Other than these instances, the TVs do not collect or record ambient conversations. (If the wake word is not recognized, no voice data is transmitted to the server; the audio processing for wake word detection is performed locally on the device and is immediately deleted.)

Additionally, to provide smart TV functionalities, LG TVs feature the ability to scan for and connect to nearby devices on the same network. This is a standard function commonly provided by smart TVs and smart home devices.

The ACR (Automatic Content Recognition) feature is provided on an opt-in basis to deliver personalized content recommendation, services, and advertisements. If a user does not consent to the applicable optional agreement, ACR data is not used for advertising purposes.

Topic by topic breakdown:

At LG Electronics, we are committed to transparency and secure user experience. To clarify recent concerns and reiterate our commitment, we would like to emphasize the following points:

  • Protecting Voice Privacy: We ensure that your LG TV does not collect, record or transmit ambient conversations in your home, unless voice functionality has been intentionally activated by the user.
  • Far-Field Voice Recognition ("Hi LG"): This feature must be manually activated by the user before the TV begins listening for the specific wake word. Audio processing occurs only when voice functionality has been activated by the user and a wake work (such as “Hi LG”) is detected.
  • Standby Mode Operation: When the TV is in standby mode and appears to be turned off, it only monitors for the wake word if you have previously enabled the Far-Field feature. If no wake word is detected, the audio used for wake word detection is processed locally on the TV, promptly deleted, and is not transmitted to LG servers.
  • Transparent and Useful Connectivity: LG TVs can identify compatible devices on the same network to enable features such as device connectivity, content sharing or smart home functionality.
  • Proactive and Continuous Security: Data security is a top priority for LG. We continuously monitor our platform and partner with independent security experts to identify vulnerabilities and deliver timely security updates.
  • Full Control Over Data and Terms: LG Electronics is committed to providing transparency and rejecting deceptive practices (dark patterns), ensuring users can choose their preferred privacy and consent options directly through their TV settings.
  • Consent-Based Activation: To provide a more tailored and convenient experience, LG TVs offer Automatic Content Recognition (ACR) to enhance viewing experiences, Voice Recognition for hands-free control, and Interest-Based Advertising to deliver relevant advertisements. Each feature requires separate explicit user opt-in consent and can be disabled at any time through the TV settings.
The Daily Front Page 11 of 30
Saturday, September 12, 2026 The Daily Front No. #260912 — VPN Boundary Breach
article

Android NAT-T keepalive offload bypasses VPN lockdown

by mhitza·▲ 188 points·45 comments·supuk.ch ↗
A normal application can violate that boundary.

Abstract network tunnel boundary with a small red signal escaping below the protected path.

1. Abstract

Android’s Always-on VPN and “Block connections without VPN” settings create a user-visible expectation that traffic attributable to covered applications will not leave through a non-VPN path. A normal application can violate that boundary through Android’s public NAT-T socket-keepalive API, causing clear, fixed-format UDP/4500 packets to reach the physical router outside the VPN path. The runtime evidence has three levels. A controlled access-point capture on a Pixel 8 Pro running Android 16 build CP1A.260505.005 recorded the packets at the public minimum 10-second interval while Always-on VPN and lockdown were enabled. A Samsung SM-F966B running Android 16 exposed one active Wi-Fi slot through the same public path; VPN Leak Guard selected the physical IPv4 default gateway, observed the active callback, and recorded a continuous router-directed active-slot lease for 24 h 32 min. On a Nothing A059 (Asteroids) running Android 16, the same implementation selected the physical gateway and recorded one active Wi-Fi slot. The Nothing result confirms public-path admission and the active callback on a third OEM. No independent packet capture or duration measurement was collected for that device. The Pixel lifecycle matrix covered backgrounding, lock, Doze, battery saver, restricted standby bucket, Binder freezer, and the observed force-stop, uninstall, network-loss, and reboot boundaries.

Source history traces the failure to a collapsed trust model in startNattKeepaliveWithFd(...): a privileged raw-fd API evolved into a public UdpEncapsulationSocket path, resource validation was added and reverted, and admission no longer authenticates the fd/resource pair or enforces the original caller UID’s current VPN policy before offload. In an F-Droid/IzzyOnDroid study of 4,679 distinct stored Git origins, the scanner detected no Android framework IPsec, IKE, or NAT-T API use; manual audit found 73 Android VpnService apps. Runtime confirmation across three OEMs and two confirmed WLAN families, the shared Android 12+ framework path, and firmware coverage across seven WLAN families representing 91.24% of estimated Android-derived shipments establish device-class exposure affecting most Android 12+ devices. The remaining 8.76% is unresolved.

2. Introduction

VPN lockdown governs routing and confinement in addition to encryption. Users and administrators expect covered applications to fail closed when the VPN is unavailable and to withhold their real network identity from destinations outside the tunnel. Prior VPN-leak research has studied routing exceptions, IPv6 and DNS leaks, WebRTC address exposure, VPN client ecosystems, and shared VPN infrastructure failures [1; 2; 3; 4; 5; 6; 7]. Android also delegates some application-triggered packet emission to system_server, a NetworkAgent, a hardware abstraction layer, or firmware, beyond the application’s ordinary socket send path.

A normal application can cross that boundary through the public Android-managed IpSecManager.UdpEncapsulationSocket and ask ConnectivityManager.createSocketKeepalive(...) to maintain a NAT-T mapping. The framework routes the request through startNattKeepaliveWithFd(...), accepts the duplicated fd and resource ID without proving current caller-owned IpSec resource identity, and hands a completed NAT-T keepalive packet to the Wi-Fi keepalive offload path without first enforcing the caller UID’s effective VPN-lockdown policy. A controlled Pixel 8 Pro capture recorded the resulting UDP/4500 packet on the physical access-point interface. Active-slot observations on two additional OEMs exercised the same public physical-gateway path on Qualcomm hardware [8; 9; 10; 11; 12; 13; 14; 15; 16; 17].

Recent Android Automotive access-control work identified ConnectivityService.startNattKeepaliveWithFd in a broad sweep of framework permission anomalies because a related keepalive API enforced PACKET_KEEPALIVE_OFFLOAD and the fd-based path did not [18; 19; 20]. That work reported the permission inconsistency. It did not trace the public UdpEncapsulationSocket trust split, the reverted IpSec resource validation, or physical Wi-Fi emission under VPN lockdown.

The platform fixes the packet shape, but the caller chooses the destination within the API and routing constraints. Repeated packets disclose the physical network’s source address and timing to that destination after the user has enabled blocking without the VPN. This violates lockdown’s identity-confinement property without requiring arbitrary payload control.

3. Background: Android VPN Lockdown and NAT-T Keepalive Offload

3.1 Android VPN Lockdown

Android’s VPN model can route covered application traffic into a VPN app’s TUN interface. Always-on VPN keeps the selected VPN active, and the user-facing “Block connections without VPN” setting is intended to prevent covered traffic from using non-secure networks outside that VPN path [21; 22]. In the normal case, an application write traverses the socket layer, per-UID network policy, fwmark and netd routing state, VPN UID-range routing, firewall/prohibit rules, and eventually the VPN TUN interface when the UID is covered by the VPN.

The normal VPN-protected path is:

covered app UID
  -> socket connect/write
  -> fwmark/netd policy and VPN UID range checks
  -> lockdown prohibit/fail-closed decision when needed
  -> VPN app TUN interface
  -> encrypted VPN tunnel over an allowed underlay

The relevant security property covers traffic and delegated packet emission attributable to a covered non-owner application. Such emission must stay off non-VPN interfaces unless the platform defines an explicit exemption. VPN-owner underlay traffic, configured split tunnels, documented platform probes, and privileged system functions may follow separate policy.

Android separately records which package is prepared to act as the VPN for each user. VpnService.prepare() may require user consent; the service itself must be declared with BIND_VPN_SERVICE. In Vpn, the prepared package is paired with its installed owner UID so uninstall/reinstall and package-only comparisons do not preserve authority. An APK that merely declares a VPN service is not the prepared VPN, and prior consent that has been revoked is not current approval [23; 24]. The installed owner UID and current prepared package provide the authority needed for keepalive admission.

3.2 NAT-T Socket Keepalive Offload

NAT traversal for IPsec commonly uses UDP port 4500. Android exposes a public API path in which an application creates an IpSecManager.UdpEncapsulationSocket and asks ConnectivityManager.createSocketKeepalive(...) to maintain the NAT mapping [8; 9]. Internally, this path passes a duplicated file descriptor and an IpSec resource ID to IConnectivityManager.startNattKeepaliveWithFd(...), which then reaches ConnectivityService, KeepaliveTracker, a NetworkAgent, and transport-specific Wi-Fi or cellular keepalive machinery [10; 11; 12].

The NAT-T keepalive offload path is:

app UID
  -> IpSecManager.openUdpEncapsulationSocket()
  -> ConnectivityManager.createSocketKeepalive(...)
  -> NattSocketKeepalive.startImpl()
  -> IConnectivityManager.startNattKeepaliveWithFd(...)
  -> ConnectivityService / KeepaliveTracker
  -> NetworkAgent
  -> Wi-Fi HAL / chipset firmware
  -> UDP/4500 keepalive on physical Wi-Fi

The Wi-Fi offload path can emit keepalive frames without waking the application or performing a new socket write for each packet. After framework admission, the final emitter sits below the ordinary app socket path that VPN lockdown normally controls.

The public keepalive path, Wi-Fi HAL offload methods, and compatibility slot requirements are shared platform interfaces below the ordinary app socket path [25; 8; 26].

4. Threat Model and Expected Lockdown Behavior

4.1 Expected Lockdown Behavior

Under Always-on VPN and “Block connections without VPN,” a covered normal application must not cause Android-managed NAT-T keepalive packets to leave over the physical network outside the VPN tunnel. If the caller’s full Android UID is covered by a non-bypassable or lockdown VPN, the public UdpEncapsulationSocket keepalive request should fail, remain unsupported, or be stopped before Wi-Fi or cellular offload emits UDP/4500 traffic on the physical underlay.

The security goal applies to NAT-T emission attributable to a covered normal app. Android’s documented policies for VPN-owner underlay traffic, configured split tunnels, and privileged platform functions remain separate. Acceptance of an unauthenticated public fd/resource pair cannot confer physical-underlay offload authority.

4.2 Attacker

The attacker controls a normal Android application installed on the victim device and controls or observes a UDP/4500 endpoint on the Internet. The validated public-API path does not require root, ADB, hidden API access, JNI, raw Binder construction, a dangerous runtime permission prompt, or the privileged PACKET_KEEPALIVE_OFFLOAD permission. The local PoC variants declared ordinary networking capabilities such as INTERNET and ACCESS_NETWORK_STATE.

4.3 Victim Configuration

The victim device has Always-on VPN enabled and “Block connections without VPN” enabled for the attacker’s UID. Runtime confirmation used Android 16 configurations across multiple OEMs. The Pixel 8 Pro run used a researcher-controlled Wi-Fi network and an external physical-interface capture; the Qualcomm-based devices supplied active physical-gateway slot observations [21; 22; 27; 15; 16; 17].

ADB and root were used for research instrumentation and packet collection. Neither is an exploitation precondition for the public API path.

5. Vulnerability: NAT-T Keepalive Lockdown Bypass

5.1 Public API Sequence

The practical attack uses the documented public path:

IpSecManager.UdpEncapsulationSocket socket =
        ipSecManager.openUdpEncapsulationSocket();

SocketKeepalive keepalive = connectivityManager.createSocketKeepalive(
        network,
        socket,
        sourceAddress,
        destinationAddress,
        executor,
        callback);

keepalive.start(10);

The public API converges on the fd-based Binder method also used by hidden raw-fd paths. The controlled on-wire Wi-Fi result used this documented sequence; raw Binder was unnecessary [8; 9; 10].

5.2 Packet Path and Missing Decision

NattSocketKeepalive.startImpl() calls IConnectivityManager.startNattKeepaliveWithFd(...); ConnectivityService forwards the request through KeepaliveTracker to the NetworkAgent and Wi-Fi backend. No admission decision checks the original caller’s effective VPN policy before hardware offload. Once handed to the Wi-Fi backend, repeated emission is outside the normal app socket writes intercepted by per-UID VPN routing and lockdown firewall rules [11; 12; 26].

5.3 Scope of Packet Control

The attacker controls the destination address within the API and routing constraints and observes the real source address exposed by the physical network. The platform fixes the NAT-T payload. The primitive leaks the real IP address and timing/cadence; it does not carry arbitrary application content.

5.4 Raw Binder as Supporting Evidence

Raw Binder diagnostics show that the server does not authenticate the supplied file descriptor or claimed IpSec resource. The reviewed validation path reports that isNattKeepaliveSocketValid(fd, resourceId) accepts every non-null fd and does not meaningfully consult resourceId; bogus values such as 0, -2, Integer.MAX_VALUE, and Integer.MIN_VALUE are not ownership checks. These diagnostics support the resource-authenticity finding. The VPN-lockdown result uses the documented public API and is independent of a stable Binder transaction number or hidden API construction. The source-line anchors are KeepaliveTracker.makeNattKeepaliveInfo(... resourceId ...) and isNattKeepaliveSocketValid(...) on the reviewed AOSP Connectivity branch [12].

Supplemental fuzzing found broad IPv4 destination acceptance after admission and route-driven IPv6 behavior. Neither result is a prerequisite for the VPN-lockdown bypass.

6. Root Cause: Collapsed Trust Model and Abandoned Resource Validation

The fd-based Binder method combines two trust models. The original raw-fd NAT-T keepalive API was privileged. During API review, public UdpEncapsulationSocket keepalives were routed through the same method, and the unconditional permission check was removed to support public IpSec/IKE applications. Secure admission then required public callers to prove ownership of a live IpSecService encap-socket resource while raw-fd callers without such a resource remained privileged.

Two checks are absent. The server accepts a public NAT-T keepalive request without proving that the duplicated fd corresponds to a live caller-owned IpSec encap-socket resource. It also starts the offload path without checking whether VPN/lockdown policy for the caller UID blocks physical-underlay emission. Once admitted, the Wi-Fi offload path emits below the ordinary app socket path where lockdown would normally constrain the caller.

Source history shows that resource validation and lifetime locking were briefly added. The implementation validated caller UID ownership, pinned the encap socket for the keepalive lifetime, rejected duplicate active use, and released the resource on final stop. It was reverted because of service dependency and deadlock concerns. Per-UID/per-network quotas replaced it. Quotas limit resource exhaustion; they do not authenticate the fd/resource pair, preserve an IpSec lifetime lease, or enforce VPN policy.

The commit IDs were rechecked against the AOSP packages/modules/Connectivity Gitiles repository on 2026-05-30 [28].

The public API duplicates a UdpEncapsulationSocket fd, passes socket.getResourceId() through NattSocketKeepalive, and calls IConnectivityManager.startNattKeepaliveWithFd(...) with the fd, resource ID, automatic-on/off state, and underpinnedNetwork. ConnectivityService forwards those fields to KeepaliveTracker; the existing socket validation accepts non-null fd state without proving that the fd matches a live caller-owned IpSec resource and without enforcing effective VPN policy before the NetworkAgent or transport backend starts hardware offload [9; 10; 11; 12].

7. Evaluation

The runtime evidence has three distinct levels. On the Pixel 8 Pro Wi-Fi configuration, Android accepted a public NAT-T keepalive request from a normal application and emitted repeated UDP/4500 packets on the physical Wi-Fi interface while VPN lockdown remained enabled. The authoritative observation point for that controlled matrix is an OpenWrt AP/router tcpdump on a separate device, because packets observed there have already crossed Android’s VPN policy boundary. A Samsung SM-F966B independently confirmed the same public path on Qualcomm WLAN hardware by maintaining one active, router-directed physical Wi-Fi slot through a lease exceeding one day. A Nothing A059 on Qualcomm hardware provided third-OEM public-path admission and active-callback confirmation through one physical-gateway Wi-Fi slot, without an external packet capture or duration measurement.

7.1 Controlled Pixel Packet Capture

The public NAT-T keepalive path produced on-wire UDP/4500 traffic on physical Wi-Fi while lockdown was enabled. The separate OpenWrt AP/router recorded a one-byte UDP payload every 10 seconds:

2026-05-28 18:36:31.709917 phy1-ap0 P   IP 192.168.1.182.38904 > 1.2.3.4.4500: UDP, length 1
2026-05-28 18:36:41.710042 phy1-ap0 P   IP 192.168.1.182.38904 > 1.2.3.4.4500: UDP, length 1

The public path used IpSecManager.openUdpEncapsulationSocket() and ConnectivityManager.createSocketKeepalive(...); no root, ADB, hidden API, raw Binder construction, JNI, dangerous runtime permission, or PACKET_KEEPALIVE_OFFLOAD permission was required for the app-side primitive.

The caller chooses the destination within API and routing constraints. The receiver observes the device’s real non-VPN source address and cadence; the payload remains the platform’s fixed NAT-T keepalive format.

7.2 Independent Samsung Runtime Confirmation

VPN Leak Guard produced an independent cross-OEM runtime result on a Samsung SM-F966B running Android 16 build BP4A.251205.006.F966BXXUABZF1 on Qualcomm hardware. Its keepalive implementation excludes VPN logical networks, selects the physical network’s validated IPv4 default gateway, opens IpSecManager.UdpEncapsulationSocket, and calls ConnectivityManager.createSocketKeepalive(...) with that physical network and gateway as the UDP/4500 destination. The Samsung run received the active callback and exposed one active physical-gateway Wi-Fi slot [13; 15].

The observation snapshot records both the highest and latest Wi-Fi slot counts as one. A sanitized screenshot records the same slot still active at a lease uptime of 24 h 32 min [17]. This is independent runtime confirmation of the vulnerable public path and physical-gateway destination on a Samsung/Qualcomm stack. The Pixel matrix remains the only controlled external packet capture; Samsung supplies the active-slot and lease measurements.

7.3 Nothing Active-Slot Confirmation

VPN Leak Guard also recorded a Nothing A059, device and product Asteroids, board volcano, on Qualcomm (qcom) hardware. It ran Android 16, SDK 36, security patch 2026-06-01, build ID BQ2A.250721.001-BP2A.250605.031.A3. The protector excludes VPN logical networks, selects the validated physical IPv4 default gateway, and marks a slot active only when SocketKeepalive.Callback.onStarted() runs. The observation report factory omits zero-slot maxima. The read-only snapshot records both the highest and latest active Wi-Fi slot counts as one [13; 14; 16].

The row confirms public physical-gateway admission and the active callback on a third OEM. Its evidence is limited to the active-slot result; no router capture, packet cadence, lease duration, lifecycle, persistence, or reboot result was collected.

7.4 Controls and Baselines

The ordinary UDP lockdown control recorded app-side UDP sends while the router capture recorded no matching UDP/4500 or UDP/12345 packets. Ordinary lockdown-covered UDP was therefore confined or absent from the physical capture while the NAT-T keepalive offload appeared at the AP boundary.

Three policy baselines bound the interpretation. With VPN and lockdown disabled, ordinary UDP and keepalive traffic were visible on the router. With VPN enabled and lockdown disabled, keepalive traffic was visible on the router while ordinary UDP was observed through the VPN interface. With both VPN and lockdown enabled, keepalive traffic was still visible on the router while the ordinary UDP lockdown control remained absent from the physical capture.

7.5 Lifecycle Behavior

Once armed on the Pixel, the keepalive remained active through backgrounding, screen lock, forced idle, battery saver, restricted standby bucket, Binder freezer/process observation, and GrapheneOS relock/back-to-BFU-without-reboot while the device stayed powered. On the Samsung, the single active Wi-Fi slot remained continuously leased for a measured period exceeding one day and was still active when the sanitized evidence screenshot was taken.

The observed stop boundaries were manual force stop, uninstall, network loss, and reboot. Packets were present before the force-stop, uninstall, and reboot actions; Wi-Fi loss stopped the active keepalive with a network-loss error, and a manual re-arm restored it after Wi-Fi returned.

7.6 Slot Counts and Lease

The tested Pixel 8 Pro Wi-Fi configuration exposed one unprivileged keepalive slot to the app UID after privileged reservations. One slot was accepted, later slot attempts failed with ERROR_INSUFFICIENT_RESOURCES (-32). The Samsung SM-F966B independently exposed one active Wi-Fi slot and held it through the measured lease. The Nothing A059 exposed one active Wi-Fi slot; no duration was measured for that row [16].

8. Impact

The direct attacker value is repeated real-network identity disclosure. A destination controlled by the attacker can learn the source IP address as seen from the physical Wi-Fi network, the fact that the device remains online, and packet timing while the keepalive remains armed. Depending on the network, the source IP can imply ISP, organization, travel state, or correlation between a device expected to be behind a VPN and a non-VPN access network.

The value of the leak comes from the user-visible lockdown promise: covered applications are expected to fail closed rather than reveal non-VPN network identity to attacker-chosen destinations. Even without payload exfiltration, a periodic signal can support presence checks, IP correlation, and timing correlation against other observations. The measured Samsung lease exceeded one day.

8.1 Device-Class Exposure

The reviewed IEEE and Wi-Fi Alliance material contains no generic WLAN keepalive-offload mandate. Public implementation history instead points to low-power NIC/driver contracts and vendor FullMAC firmware interfaces. By 2009, Windows 7’s NDIS 6.20 model supported low-power ARP, IPv6 Neighbor Solicitation, and 802.11 RSN/GTK offloads; in 2011, IEEE 802.11v standardized adjacent Wireless Network Management keep-alive and proxy mechanisms; by 2014, public Android WLAN source evidence shows Qualcomm firmware command surfaces for STA keepalive and IPsec NAT keepalive; and by Android 6 / Marshmallow in 2015, AOSP included hidden NAT-T keepalive framework support. Android compatibility requirements for app-visible Wi-Fi keepalive offload appeared later, in the Android 10 CDD in 2019 [29; 30; 31; 32; 33].

App-visible slots are part of the Android compatibility model. Android 10’s 2019 CDD requires devices that expose Wi-Fi keepalive offload to support the SocketKeepalive API and at least three concurrent Wi-Fi keepalive slots. Slot/resource checks found no manufacturer overlay that deliberately zeroed the relevant defaults.

Runtime results cross OEM boundaries within the two confirmed WLAN families. The Pixel 8 Pro/Broadcom configuration emitted packets in the controlled AP capture. The Samsung SM-F966B/Qualcomm configuration maintained an active router-directed Wi-Fi slot through the measured lease. Another Qualcomm-based OEM admitted the public physical-gateway path and reached the active callback for one Wi-Fi slot [27; 34; 35; 13; 14; 15; 16; 17].

The vulnerable admission path is shared Android 12+ framework behavior, making the relevant platform window start with Android 12’s 2021 release. The firmware/source census finds keepalive or offloaded-packet support surfaces across all seven tracked Android WLAN stack families: Qualcomm QCA/CLD3/FastConnect, Qualcomm WLAN/QDSP6 WCNSS, Broadcom/Cypress bcmdhd/DHD, MediaTek CONSYS/Connac, Unisoc/Spreadtrum SPRDWL, Samsung S.LSI/Exynos Wi-Fi, and Huawei/HiSilicon Hi11xx. The first-observed public support evidence across those families spans 2014 through 2020 [33; 36; 37; 38].

Runtime confirmation across three OEMs and two confirmed WLAN families, the shared Android 12+ implementation, slot defaults, and the seven-family firmware census establish device-class exposure affecting most Android 12+ devices. The mapped WLAN families represent 91.24% of estimated Android/AOSP-derived shipments from 2021Q4 through 2026Q1. The remaining 8.76% is unresolved mixed/long-tail SKU coverage [39; 36].

9. Application Ecosystem Study

Public API availability does not establish compatibility demand. A static F-Droid/IzzyOnDroid study measured use of Android’s framework IPsec, IKE, and NAT-T machinery. Across the scanned origins, the scanner found zero framework API uses and zero method-specific raw Binder invocations of the corresponding transactions [40; 41].

9.1 Corpus and Method

The collection run downloaded the signed F-Droid and IzzyOnDroid v1 indexes on 2026-07-05, extracted and normalized declared source URLs, and cloned available repositories. This produced 4,888 package-keyed per-checkout reports. Those reports contain 4,679 distinct stored gitOrigin strings; the table below de-duplicates package-keyed observations by exact stored origin and does not merge differently spelled URLs that might identify the same upstream project.

The analysis was static and lexical: it scanned source and manifest files for framework API references, method-specific raw Binder transactions, and VPN comparison candidates, while excluding generated build directories, dependency caches, Git metadata, and files above the scanner’s size limit. The VPN comparison candidates were then manually audited for intentional user-facing Android VpnService/TUN-style behavior. Catalog entries, clone targets, or source checkouts that were unavailable to the scanner remain outside the distinct-origin denominator.

The manually audited VpnService set confirms that the scanned corpus contains ordinary Android VPN applications. Those apps use the standard VpnService path; none supplied a framework IPsec/IKE/NAT-T match.

9.2 Interpretation

Android exposes transform construction, SPI allocation, UDP encapsulation, framework IKE negotiation, migration, state queries, and hardware NAT-T offload to ordinary applications. The sample found no use of that framework machinery. The proposed repair targets NAT-T keepalive admission around fd/resource ownership and effective VPN-lockdown policy; ordinary Java/NDK networking and the VpnService path remain outside its scope.

F-Droid and IzzyOnDroid exclude much proprietary enterprise VPN software, OEM clients, carrier software, and sideloaded closed-source applications. The zero therefore describes this open-source sample and cannot establish universal absence.

9.3 Google Play Cross-Check

A Google Play and web census dated 2026-07-29 found 29 candidates: 26 were currently listed and 3 had been removed. APKs were acquired for 9 of the 26 current listings; 17 remained listing-only. The acquired set comprised 8 proprietary APKs and 1 open-source APK. Two proprietary APKs contained all 47 verified platform-SDK call sites: 37 in FortiClient VPN and 10 in SmartVPN. The other 7 acquired APKs contained none [42].

FortiClient VPN had 3,995,326 cumulative Google Play installs and SmartVPN had 139,322, totaling 4,134,648. Google Play exposes these exact cumulative-install fields behind rounded public download badges [42].

10. Mitigation and Regression Tests

A repair for any retained normal-app path should preserve authorized NAT-T keepalives while making physical-underlay offload fail closed for covered normal apps. The fix point is admission to startNattKeepaliveWithFd(...) and the associated keepalive lifetime: the platform must authenticate the fd/resource pair, decide whether the caller may emit on the selected physical network, and stop the record if that authorization becomes stale.

10.1 NAT-T Repair Responsibilities

  • Raw-fd callers without a valid public IpSec resource remain gated by PACKET_KEEPALIVE_OFFLOAD and receive basic fd-shape validation.
  • Public UdpEncapsulationSocket callers prove ownership of a live IpSecService encap-socket resource, prove that the duplicated fd matches the stored resource, and cannot reuse one resource for concurrent active NAT-T records.
  • ConnectivityService captures the calling full UID before identity clearing, checks the target network and any underpinnedNetwork relationship, and rejects physical-underlay offload when the effective VPN/lockdown policy for that UID would block direct emission.
  • KeepaliveTracker and the transport backend start Wi-Fi or cellular offload only after admission succeeds, and release the resource exactly once on denial, construction failure, binder death, packet-replacement failure, client stop, or final system stop.
  • VPN, lockdown, owner, underlying-network, and selected-network changes revoke or revalidate active and paused NAT-T records before they can resume emission.

10.2 Regression Tests

Regression tests for this bug should prove that unauthorized keepalives fail before slot allocation, packet-filter installation, transport start, callback success, or IpSec resource mutation. The minimum negative cases are:

  • a covered normal app requests a public NAT-T keepalive while lockdown is already enabled;
  • a record admitted before lockdown is stopped when lockdown begins for the caller’s UID range;
  • running and paused records are stopped when the relevant VPN is removed, replaced by a different owner, changes bypassability, or loses its underlying network;
  • stale, closed, mismatched, other-UID, and duplicate-active IpSec resources are rejected before offload;
  • raw-fd use with an invalid or absent resource ID requires the privileged keepalive permission;
  • denial and final-stop paths close incoming ParcelFileDescriptors and release any acquired IpSec lease exactly once.

Positive coverage should show that a legitimate caller with an owned live encap-socket resource can still use NAT-T keepalive when its effective VPN policy permits the selected physical emission [11; 12].

11. Disclosure, Ethics, and Artifacts

The finding was reported to the Android Vulnerability Reward Program on 2026-05-15. Google triaged the report the same day and requested coordinated disclosure while Android Security assessed the issue. The reporter stated that the finding had not been posted publicly or shared with third parties, submitted additional validation material on 2026-05-17, and asked for disclosure guidance after Google marked the report as a duplicate of canonical issue 386376240 on 2026-05-19.

The reporter notified Google of planned public disclosure through the VRP report on 2026-06-12. The researcher-visible VRP API record then shows redacted Google update markers from two accounts, dated 2026-06-12 and 2026-06-15; neither exposes a comment body. The record contains no objection, delay request, CVE assignment, fix-status update, severity decision, reward decision, or publication clearance. Public disclosure began on 2026-07-29.

The experiments used researcher-controlled devices, VPN configurations, packet captures, and endpoints. No third-party user traffic was collected. Released artifacts are limited to sanitized evidence and reproduction material; private VRP content, local identifiers, weaponized raw traces, and patch diffs are excluded.

Competing interest: the author sells VPN Leak Guard, the commercial Android app that produced the active-slot observations reported here.

12. Limitations

Runtime testing spans three device models and is not an exhaustive per-model inventory. The Pixel 8 Pro/Broadcom result includes a controlled packet-capture matrix. The Samsung SM-F966B/Qualcomm result includes an active physical-gateway slot and measured lease. The Nothing A059/Qualcomm result confirms public-path admission and the active callback; external packet capture and duration remain unmeasured [15; 16].

Reboot persistence is not established. The Pixel keepalive stopped at the observed reboot boundary, the Samsung lease was measured while the device remained powered, and Nothing lifecycle and reboot behavior were not measured.

Cellular packet emission was not measured. The framework admission path is transport-agnostic in source analysis, and Android 16 compatibility material says cellular keepalive offload can be exposed to third-party apps with at least one cellular slot. The tested Pixel default configuration returned insufficient resources before a cellular modem path emitted traffic, and effective normal-app execution depends on manufacturer slot overlays, hardware/HAL availability, privileged slot reservations, per-UID unprivileged limits, and transport backend behavior [25; 8; 43; 12].

APKs were acquired for 9 of 26 current Google Play listings. Store counts are cumulative across versions and VPN modes and do not identify active users or platform-path use. DEX results apply to the acquired APK versions.

No patched Android build, patched-device packet capture, or complete device-level atest run was performed. The repair remains source-level guidance and requires implementation-level regression testing.

13. Related Work

VPN-leak comparisons depend on attacker position, trigger, endpoint control, packet shape, cadence, enforcement layer, measured scope, and the user-visible policy being bypassed. The NAT-T case uses a normal installed app to arm repeated fixed UDP/4500 packets to an attacker-chosen Internet endpoint while Android VPN lockdown is enabled.

TunnelCrack and TunnelVision-style work shows that routing exceptions and hostile local-network conditions can defeat VPN expectations on affected platforms [1; 44; 45; 46]. Those attacks use a hostile local network; the NAT-T case uses an installed app. Both test whether packets cross the expected VPN boundary.

IPv6, DNS, WebRTC, and VPN ecosystem studies provide the broader privacy and measurement context. Perta et al., Al-Fannah, Cho and Heidemann, Khan et al., VPNalyzer/VPNInspector, Wu et al., and Yang et al. show why source-IP exposure, resolver behavior, shared VPN state, and careful measurement boundaries matter [2; 3; 6; 4; 5; 47; 48; 7].

These studies measure VPN behavior, infrastructure, privacy, or client properties. The ecosystem study measures compatibility demand for Android’s framework IPsec/IKE/NAT-T machinery. F-Droid and IzzyOnDroid make source-level inspection possible [40; 41], but exclude much proprietary enterprise and business VPN software. The resulting zero is therefore evidence for a compatibility context around NAT-T keepalive admission and cannot estimate universal Android-market prevalence.

AutoAcRaptor is the closest prior signal on the same Android framework entry point. It was a broad AAOS access-control study that flagged ConnectivityService.startNattKeepaliveWithFd as a verified missing-permission anomaly because a related keepalive API required the signature-level PACKET_KEEPALIVE_OFFLOAD permission while the fd-based path did not [18; 19; 20]. That work identified the suspicious entry point but did not validate the Android phone VPN-lockdown bypass. The NAT-T analysis connects the entry point to public UdpEncapsulationSocket keepalives, Wi-Fi offload, and on-wire packet emission outside phone VPN lockdown.

Android QUIC close-payload delegated UDP send is the closest Android delegated-send peer. It shows application-triggered UDP emission outside the ordinary app VPN send path [49; 50; 51]. QUIC close-payload is a one-shot or event-driven software send; NAT-T keepalive is repeated, fixed-format UDP/4500 Wi-Fi offload traffic.

Mullvad’s Android connectivity-check and DNS-leak reports, GrapheneOS VPN leak-blocking work, and local-link/multicast discussions show that Android VPN enforcement has long depended on platform exceptions as well as VPN-app behavior [52; 53; 54; 55]. They differ from the NAT-T case in endpoint control, trigger, cadence, and packet shape. They also show that delegated or exempt traffic must be checked against the user’s lockdown expectation.

14. Discussion

14.1 Review Failure and Fail-Closed Release Discipline

The public and privileged NAT-T paths reached the same Binder method with different trust requirements and no complete admission boundary. Stronger ownership, fd-identity, lifetime, duplicate-use, and invalid-resource checks were added in 2019, then reverted for technically legitimate service-dependency and deadlock concerns. Per-UID and per-network quotas replaced them, but quotas limit resource consumption; they do not provide equivalent fd/resource authenticity, lifetime, or VPN-policy gates [28; 12].

The remaining validation TODOs cover invalid, closed, stale, mismatched, other-UID, and duplicate resources; raw-fd privilege; lease cleanup; and authorization changes while a keepalive is active or paused. A technically motivated reversion does not itself constitute the failure. Releasing the normal-app path without equivalent controls, while those validation obligations remained open, is a review and release-governance failure. That characterization concerns the process and resulting control gap, not any individual contributor.

Unprivileged availability should have remained at zero slots until an acceptable public/private API split and the complete security checks were ready. Resource quotas cannot serve as release authorization for a physical-underlay emission primitive.

14.2 Deprecate Normal-App IPsec Access

Android should deprecate the public app-facing IPsec, IKE, and NAT-T surface and make the framework functionality system-privileged. Authenticated carrier, IWLAN, VCN, platform VPN, and other platform-internal consumers should remain; their current roles are not replaced merely by removing normal-app access [56; 57; 58].

Normal apps should receive zero exposed slots by default. If legacy access must remain, it should sit behind a default-off, reboot-required compatibility switch with a release posture analogous to radio-generation controls. Ordinary VPN apps can continue to run their own userspace protocol implementations through VpnService [23].

The open-source scan found 0 platform consumers across 4,679 origins; the Play study found 2 among 9 acquired APKs. FortiClient VPN and SmartVPN have 4,134,648 cumulative Google Play installs between them, while Google reports more than 3 billion active Android devices. Both apps also support VPN modes that do not use the platform path. I estimate that at most 0.01% of Android users—about one in 10,000—depend on these platform APIs. That constituency is too small to justify leaving normal-app access enabled by default. [40; 41; 42; 59]

14.3 Router-Terminated VPN Guidance

For threat models that cannot tolerate a phone-side VPN escape, a VPN-enforcing external router is the conservative community consensus among identifiable privacy and security practitioners. Mullvad reported recurring Android bypass classes in 2022 for connectivity checks, in 2024 for DNS, and in 2026 for application-triggered QUIC traffic. IVPN independently reproduced the 2026 QUIC path, and GrapheneOS community guidance recommends an external router that tunnels the phone’s upstream traffic [52; 53; 60; 61; 62].

These reports cover different mechanisms and do not imply that all Android traffic always bypasses a VPN. Their recurrence shows that Android’s current architecture has repeatedly exposed new phone-side bypass paths and cannot responsibly promise that no further class will appear.

The recommendation is conditional. The phone must use the router as its exclusive Internet path, with cellular and alternate networks disabled or separately blocked, and the router must fail closed if its tunnel fails. This reduces dependence on Android’s VPN enforcement; it is not an unconditional guarantee against router defects, local-network exposure, or traffic over other radios.

15. Conclusion

The affected class is Android 12+ devices that expose app-visible Wi-Fi NAT-T keepalive offload with usable unprivileged slots. On such devices, a covered normal app can reach physical-underlay offload without effective VPN-lockdown admission. The available runtime, framework, slot, firmware, and shipment evidence supports exposure across most Android 12+ devices [25; 39].

The repair must separate privileged raw-fd requests from public UdpEncapsulationSocket requests. Raw-fd callers require PACKET_KEEPALIVE_OFFLOAD; public callers require caller-owned resource validation, fd identity checks, lifetime pinning, and duplicate-use rejection. Both paths require effective VPN-policy authorization before NetworkAgent or HAL admission and revalidation when relevant network or VPN state changes. Until those checks are complete, unprivileged NAT-T offload should fail closed [11; 12].

16. Appendix A. Evidence

Displayed pcap hashes use 12-hex SHA-256 prefixes.

A.1 Environment and Boundary Proof

  • Primary measured phone: Pixel 8 Pro (husky), Android 16 build CP1A.260505.005, security patch 2026-05-05.
  • Independent measured phone: Samsung SM-F966B (q7q) on Qualcomm (qcom) hardware, Android 16 build BP4A.251205.006.F966BXXUABZF1, security patch 2026-06-05 [15].
  • Additional active-slot phone: Nothing A059, device and product Asteroids, board volcano, on Qualcomm (qcom) hardware, Android 16 / SDK 36, security patch 2026-06-01, build ID BQ2A.250721.001-BP2A.250605.031.A3 [16].
  • Later-version provenance: the same non-cumulative snapshot records a Pixel 8 Pro on Android 17 with one active Wi-Fi slot. This row establishes slot availability only; it does not replace or relabel the controlled Android 16 Pixel packet capture [16].
  • VPN configuration: Mullvad package net.mullvad.mullvadvpn, version 2026.5. Always-on VPN and “Block connections without VPN” are observed in VPN-management snapshots for the measured rows.
  • App capability: normal app path using IpSecManager.openUdpEncapsulationSocket() and ConnectivityManager.createSocketKeepalive(...). No root, ADB, dangerous runtime permission, JNI, hidden API, raw Binder, or PACKET_KEEPALIVE_OFFLOAD is required for the public-API claim.
  • External observation point: separate OpenWrt AP/router capture on the physical Wi-Fi side. Packets observed there have already crossed Android’s VPN policy boundary.
  • Packet form: fixed NAT-T UDP/4500 keepalive with a one-byte payload. The primitive cannot carry arbitrary application payloads.
  • Cadence and slots: the public minimum interval was observed. On the tested Pixel Wi-Fi configuration, one unprivileged slot was accepted and later slot attempts failed with ERROR_INSUFFICIENT_RESOURCES (-32). The Samsung row records one active slot and a measured lease. The Nothing row records one active slot without a duration measurement.

A.2 On-Wire Captures, Controls, and Baselines

  • Primary Wi-Fi proof: app run 20260529-023621-pid20917, router case slot-cadence, 18 packets, pcap prefix d463ea0c9ea7. Slot 0 accepted and UDP/4500 appeared on physical Wi-Fi at 10-second cadence.
  • Ordinary UDP lockdown control: app run 20260529-023925-pid21402, router case ordinary-udp-lockdown, zero matching packets, pcap prefix e3f42e268763. App-side UDP sends to UDP/4500 and UDP/12345 had no matching router packets under lockdown.
  • VPN off, lockdown off baseline: app run 20260529-031517-pid13504, router case baseline-vpn-off-lockdown-off, 9 packets, pcap prefix 4bc929aeaf54. Ordinary UDP and keepalive packets were visible when VPN confinement was disabled.
  • VPN on, lockdown off baseline: app run 20260529-032140-pid14963, router case baseline-vpn-on-lockdown-off, 6 packets, pcap prefix 840d5c418933. Keepalive packets were visible on the router while ordinary UDP was app-side on tun0.
  • VPN on, lockdown on baseline: app run 20260529-024101-pid21674, router case baseline-vpn-on-lockdown-on, 6 packets, pcap prefix 732648acd332. Keepalive packets remained visible while the ordinary UDP lockdown control was absent from the router capture.

The older July 28 snapshot and sanitized screenshot supply the Samsung row. The newer non-cumulative snapshot supplies the Nothing and Android 17 Pixel rows. The tracked protector defines physical-gateway selection and active-callback state; the tracked report factory publishes only maxima greater than zero. The raw snapshots remain in the read-only local mirror, and the table reproduces only the sanitized fields [13; 14; 15; 16; 17].

A.3 Lifecycle and Boundary Rows

  • Nondestructive lifecycle: app runs 20260529-013941-pid14044 and 20260529-021757-pid16937, with lifecycle rows under active-lifecycle-*. Background/home, screen lock, force idle, battery saver, restricted bucket, and process-observation windows kept app heartbeats alive; early router windows were workflow evidence rather than standalone on-wire persistence proof.
  • Force stop: app run 20260529-024326-pid21970, router case active-lifecycle-force-stop, 36 packets, pcap prefix dc264c7d08a8. Packets were present before the host force-stop killed the process.
  • Uninstall: app run 20260529-025027-pid23321, router case active-lifecycle-uninstall, 42 packets, pcap prefix 12d0a2a140e7. Packets were present before package removal.
  • Reboot: app run 20260529-025521-pid24592, router case active-lifecycle-reboot, 27 packets, pcap prefix 156a67de67da. Packets were present before reboot.
  • Network loss and re-arm: app run 20260529-031051-pid12326, router case network-loss-rearm-manual, 6 packets, pcap prefix 63a9bec9367f. Wi-Fi loss caused onError -20; Wi-Fi return and manual re-arm restored the keepalive.
  • Back-to-BFU: the Pixel lifecycle observations include GrapheneOS relock/back-to-BFU-without-reboot continuation; this is not reboot persistence.

A.4 Destination Fuzzing

Destination fuzzing tested address-class filtering after keepalive admission. IPv4 destination classes were broadly accepted and confirmed on-wire; IPv6 behavior was route-driven. These results inform consequence and hardening analysis. The VPN-lockdown bypass requires only the validated public endpoint.

IPv4 groupDestinations testedResultPublic baseline1.2.3.4Accepted, on-wireThis network0.0.0.0Accepted, on-wireLimited broadcast255.255.255.255Accepted, on-wireMulticast224.0.0.1, 224.0.0.251, 239.255.255.250Accepted, on-wireLoopback127.0.0.1Accepted, on-wireLink-local169.254.0.1Accepted, on-wireCGNAT / RFC 6598100.64.0.1Accepted, on-wireRFC1918 private10.0.0.0, 172.16.0.1, 192.168.0.1Accepted, on-wireRFC5737 TEST-NET192.0.2.1, 198.51.100.1, 203.0.113.1Accepted, on-wireClass E reserved240.0.0.1Accepted, on-wire

IPv6 acceptance followed route availability and Java address-family normalization. With a matching route, especially a ::/0 default route, 41 of 48 tested IPv6 destination cases were accepted. Missing routes produced route-selection failures. IPv4-mapped ::ffff:*/96 forms were rejected with ERROR_INVALID_IP_ADDRESS (-21) because Java address parsing normalized them into Inet4Address, causing a family mismatch. IPv4-translated and NAT64-prefix forms were accepted in the route-present matrix.

IPv6 categoryExamples / notesResult in route-present setupLink-local unicastfe80::1, fe80::2, longer link-local examplesAccepted via on-link routeULAfc00::1, fd00::1, fd12:..., fdff:...AcceptedGlobal unicastGoogle and Cloudflare resolver examplesAcceptedDocumentation2001:db8::1, 2001:db8:ffff:ffff::1AcceptedTeredo / 6to4 / ORCHID-v22001::1, 2002::1, 2001:20::1AcceptedLoopback::1AcceptedUnspecified::AcceptedDiscard100::1AcceptedIPv4-translated::ffff:0:1.2.3.4AcceptedNAT64 well-known prefix64:ff9b::1.2.3.4, 64:ff9b::127.0.0.1, 64:ff9b::0.0.0.0, 64:ff9b::255.255.255.255AcceptedNAT64 alternate prefix64:ff9b:1::1.2.3.4AcceptedMulticastInterface-local, link-local, site-local, organization-local, and global examplesAcceptedIPv4-mapped::ffff:1.2.3.4 and related IPv4-mapped formsRejected with -21 due Java normalization/family mismatch

A.5 Diagnostic and Precondition Runs

Android callback and error codes establish preconditions. Only the external AP/router capture establishes on-wire emission.

CodeInterpretation0Accepted callback; external AP/router capture is still required for on-wire proof.-20Network lost or keepalive stopped because the active network changed.-21Invalid IP, missing route, or address-family mismatch.-22Invalid or unbound source port.-24Invalid keepalive interval.-30Unsupported network type or supported keepalive count zero.-32Insufficient resources or slot gate.-994Local socket create/bind failure in fuzz app diagnostics.

The second-round fuzz run had 41 trial_finished rows, 40 errors, one skip, 26 -32 slot-gate results, 9 -24 interval results, 2 -22 source-port results, and 3 local -994 socket failures. These counts characterize negative and precondition runs; they do not establish Wi-Fi destination breadth.

Raw Binder/resource diagnostics used non-null fds with resource IDs including 0, -2, Integer.MAX_VALUE, and Integer.MIN_VALUE. Those values were not treated as ownership proofs. The diagnostics support the authenticity finding; the VPN-lockdown result uses the documented public API.

17. References

  1. Nian Xue, Yashaswi Malla, Zihang Xia, Christina Poepper, and Mathy Vanhoef (2023). Bypassing Tunnels: Leaking VPN Client Traffic by Abusing Routing Tables. Proceedings of the 32nd USENIX Security Symposium. USENIX Association. Source.

  2. Vasile C. Perta, Marco V. Barbera, Gareth Tyson, Hamed Haddadi, and Alessandro Mei (2015). A Glance through the VPN Looking Glass: IPv6 Leakage and DNS Hijacking in Commercial VPN Clients. Proceedings on Privacy Enhancing Technologies, 2015(1), pp. 77–91. DOI: 10.1515/popets-2015-0006. Source.

  3. Nasser Mohammed Al-Fannah (2017). One Leak Will Sink A Ship: WebRTC IP Address Leaks. Source.

  4. Mohammad Taha Khan, Joe DeBlasio, Geoffrey M. Voelker, Alex C. Snoeren, Chris Kanich, and Narseo Vallina-Rodriguez (2018). An Empirical Analysis of the Commercial VPN Ecosystem. Proceedings of the Internet Measurement Conference. Association for Computing Machinery. DOI: 10.1145/3278532.3278570. Source.

  5. Reethika Ramesh, Leonid Evdokimov, Diwen Xue, and Roya Ensafi (2022). VPNalyzer: Systematic Investigation of the VPN Ecosystem. Proceedings of the Network and Distributed System Security Symposium. Internet Society. Source.

  6. Yejin Cho and John Heidemann (2025). Smoothing Rough Edges of IPv6 in VPNs. Source.

  7. Yuxiang Yang, Ao Wang, Xuewei Feng, Qi Li, and Ke Xu (2026). Invisible Adversaries: A Systematic Study of Session Manipulation Attacks on VPNs. Source.

  8. Android Developers (2026). SocketKeepalive. Android API reference. Source. Accessed 2026-05-30.

  9. Android Open Source Project (2026). ConnectivityManager.java. AOSP source, packages/modules/Connectivity, commit 2519a78731526d2eb20ae8812acdcab6ef7a09b6. Source. Commit-pinned; accessed/rechecked 2026-06-07.

  10. Android Open Source Project (2026). NattSocketKeepalive.java. AOSP source, packages/modules/Connectivity, commit 2519a78731526d2eb20ae8812acdcab6ef7a09b6. Source. Commit-pinned; accessed/rechecked 2026-06-07.

  11. Android Open Source Project (2026). ConnectivityService.java. AOSP source, packages/modules/Connectivity, commit 2519a78731526d2eb20ae8812acdcab6ef7a09b6. Source. Commit-pinned; accessed/rechecked 2026-06-07.

  12. Android Open Source Project (2026). KeepaliveTracker.java. AOSP source, packages/modules/Connectivity, commit 2519a78731526d2eb20ae8812acdcab6ef7a09b6. Source. Commit-pinned; accessed/rechecked 2026-06-07.

  13. Armin Šupuk (2026). VPN Leak Guard NAT-T Keepalive Protector. Unpublished Android source implementation, NattKeepaliveProtector.kt; SHA-256 ab7a14c28c441ef36029afb6e0eaf45ca34d66e86fe5a18aeb722817fcca2443; inspected 2026-07-28. On file with the author; available on request.

  14. Armin Šupuk (2026). VPN Leak Guard NAT-T Observation Report Factory. Unpublished Android source implementation, NattContribution.kt; SHA-256 b8942b159ebe27ae178f16f464280b804ae54c04420f589c4b458dae23e12d9f; inspected 2026-07-28. On file with the author; available on request.

  15. Local NAT-T keepalive observation database (2026). Samsung SM-F966B Wi-Fi Keepalive Observation Snapshot. Unpublished read-only observation-database snapshot, 2026-07-28T09:46:16Z; SHA-256 56e15294ba602df454f09f597850364b8ef7bd355b0ad8e7f6038c75766d290f; only sanitized device, build, transport, and slot-count facts are reproduced. On file with the author; available on request.

  16. Local NAT-T keepalive observation database (2026). Nothing A059 and Pixel 8 Pro Wi-Fi Keepalive Observation Snapshot. Unpublished read-only observation-database snapshot, 2026-07-28T15:06:09Z; SHA-256 0465dfdff9104f6a4836c299e2c96203b0935417d5564edef852c15eafc1fda3; only the sanitized Nothing A059 runtime fields and Pixel 8 Pro Android 17 single-slot provenance row described in Appendix A are reproduced. On file with the author; available on request.

  17. VPN Leak Guard (2026). Sanitized Samsung NAT-T Active-Slot Lease Screenshot. Published evidence image. Source. SHA-256 fd1466ff3c897190708d7cf9f35f4d02532e48e5d3846cab250b0a3667b8bb57; records 24 h 32 min lease uptime.

  18. Jumana, Parjanya Vyas, and Yousra Aafer (2025). Red Light for Security: Uncovering Auto Feature Check and Access Control Gaps in AAOS. Detection of Intrusions and Malware, and Vulnerability Assessment, pp. 147–166. Springer Nature Switzerland. DOI: 10.1007/978-3-031-97623-0_9. Source.

  19. Jumana (2025). Analyzing Access Control Logic in the Android Automotive Framework. University of Waterloo. Source.

  20. Parjanya Vyas (2025). Cues, Clones, and Cars: Access Control Issues in Customized Android. University of Waterloo. Source.

  21. Android Developers (2026). VPN. Android Developers documentation. Source. Accessed 2026-05-30.

  22. Android Developers (2026). DevicePolicyManager. Android API reference. Source. Accessed 2026-05-30.

  23. Android Developers (2026). VpnService. Android API reference. Source. Preparation flow and BIND_VPN_SERVICE declaration; accessed 2026-07-11.

  24. Android Open Source Project (2026). Android 17 Vpn.java. AOSP source, frameworks/base, commit 94b4c163b7dfe5ce3607f7bb8456f9573f7de57d. Source. Commit-pinned; accessed 2026-07-11.

  25. Android Open Source Project (2026). Android 16 Compatibility Definition. Android compatibility documentation. Source. Accessed 2026-05-30.

  26. Android Open Source Project (2026). IWifiStaIface.aidl. AOSP source, hardware/interfaces, commit 1a56e38edc2f2f6189ef405ee1edce554e15cbc0. Source. Commit-pinned; accessed/rechecked 2026-06-07.

  27. Google Pixel Help (2026). Pixel Phone Hardware Tech Specs. Google support documentation. Source. Accessed 2026-05-30.

  28. Android Open Source Project (2026). packages/modules/Connectivity Gitiles Repository. AOSP source repository. Source. Commit IDs in the paper source-history table returned 200 from Gitiles on 2026-05-30.

  29. Microsoft Learn (2023). Protocol Offloads for NDIS Power Management. Windows driver documentation. Source. Documents Windows 7 / NDIS 6.20 protocol offloads; accessed 2026-07-03.

  30. IEEE Standards Association (2011). IEEE 802.11v-2011: IEEE Standard for Information Technology–Wireless LAN Medium Access Control and Physical Layer Specifications Amendment: Wireless Network Management. IEEE standards page. Source. Published 2011-02-09; accessed 2026-07-03.

  31. Motorola Mobility LLC (2014). Qualcomm qcacld-2.0 WMI Keepalive and IPsec NAT Keepalive Definitions. GitHub source snapshot. Source. Commit e6075b10e7ce876ba4fd86601fe58aa5791831bd, 2014-10-25; accessed 2026-07-03.

  32. Android Open Source Project (2015). ConnectivityManager.java, Android 6-era NAT-T Keepalive Source. AOSP source, frameworks/base. Source. Hidden PacketKeepalive and startNattKeepalive source; accessed 2026-07-03.

  33. Android Open Source Project (2019). Android 10 Compatibility Definition. Android compatibility documentation. Source. Wi-Fi Keepalive Offload section accessed 2026-07-03.

  34. TechInsights (2023). Google Pixel 8 Pro Component Analysis. TechInsights component analysis. Source. Accessed 2026-05-30.

  35. Broadcom (2022). Broadcom Announces Availability of World’s First Wi-Fi 7 Ecosystem Solutions. Broadcom press release. Source. Accessed 2026-05-30.

  36. Local firmware census (2026). Slot Overlay Checks. Unpublished firmware-census evidence note, slot-overlay-checks.md; accessed 2026-07-03. On file with the author; available on request.

  37. Local firmware census (2026). Chipset Provider Keepalive Support Timeline. Unpublished firmware-census evidence note, chipset-support-timeline.md; accessed 2026-07-03. On file with the author; available on request.

  38. Android Developers Blog (2021). Android 12 Is Live in AOSP. Android Developers Blog. Source. Accessed 2026-07-03.

  39. Local firmware census (2026). Android 12+ Firmware Family Coverage Report. Unpublished firmware-census evidence note, android12-market-map/coverage-report.md; generated 2026-07-02. On file with the author; available on request.

  40. F-Droid Project (2026). All Our APIs: The F-Droid Repository Index. F-Droid documentation. Source. Signed v1 index documentation; accessed 2026-07-11.

  41. IzzyOnDroid (2026). IzzyOnDroid F-Droid Repository. Official repository documentation. Source. F-Droid-compatible open-source Android repository; accessed 2026-07-11.

  42. Armin Šupuk (2026). Play Store IPsec/IKEv2 App Census with DEX Call-Site Verification. Unpublished reproducible research report, play-ipsec-ikev2-census/report.md; commit a6c0ef59790e2ec54fbc19e201e38fbdcaa56974; SHA-256 8868346dfafd8bbb7712f524872002441056da3c1dcded6197424df0c6ae0bc5; 29 discovered candidates, 26 current listings, 3 removed apps, 9 acquired APKs, 17 listing-only candidates, 47 verified platform-SDK DEX call sites across 2 proprietary APKs, 4,134,648 cumulative installs for the 2 API-containing apps, and exact APK hashes; generated 2026-07-29. On file with the author; available on request.

  43. Android Open Source Project (2026). Connectivity Service Resource Configuration. AOSP source, packages/modules/Connectivity, commit 2519a78731526d2eb20ae8812acdcab6ef7a09b6. Source. Commit-pinned; accessed/rechecked 2026-06-07.

  44. Nian Xue, Yashaswi Malla, Zihang Xia, Christina Poepper, and Mathy Vanhoef (2023). TunnelCrack. Project website. Source. Accessed 2026-05-30.

  45. Leviathan Security Group (2024). TunnelVision. Project website. Source. Accessed 2026-05-30.

  46. Leviathan Security Group (2024). TunnelVision: Routing Traffic Without VPN Encryption Using DHCP Option 121. Technical blog. Source. Accessed 2026-05-30.

  47. Reethika Ramesh, Anjali Vyas, and Roya Ensafi (2023). “All of Them Claim to Be the Best”: Multi-Perspective Study of VPN Users and VPN Providers. Source.

  48. Ka Lok Wu, Man Hong Hue, Ngai Man Poon, Kin Man Leung, Wai Yin Po, Kin Ting Wong, Sze Ho Hui, and Sze Yiu Chau (2023). Back to School: On the (In)Security of Academic VPNs. Proceedings of the 32nd USENIX Security Symposium. USENIX Association. Source.

  49. Low Level Academy (2026). Android QUIC Close-Payload Delegated UDP Send. Technical blog. Source. Accessed 2026-05-30.

  50. GrapheneOS (2026). GrapheneOS Releases: 2026050400. GrapheneOS release notes. Source. Accessed 2026-05-30.

  51. GrapheneOS (2026). Disable Close-QUIC Optimization Due to VPN Bypass Concern. GrapheneOS packages/modules/Connectivity commit. Source. Accessed 2026-05-30.

  52. Mullvad VPN (2022). Android Leaks Connectivity Check Traffic. Mullvad blog. Source. Accessed 2026-05-30.

  53. Mullvad VPN (2024). DNS Traffic Can Leak Outside the VPN Tunnel on Android. Mullvad blog. Source. Accessed 2026-07-29.

  54. GrapheneOS (2024). GrapheneOS OS Issue Tracker: VPN Leak Blocking and Multicast/Local-Link Traffic. GitHub issue. Source. Accessed 2026-05-30.

  55. GrapheneOS Community (2024). GrapheneOS Fixing the Standard VPN Leak Blocking Is Nearing Completion. GrapheneOS discussion forum. Source. Accessed 2026-05-30.

  56. Android Open Source Project (2026). IPsec/IKEv2 Library. Android platform documentation. Source. Platform IKEv2, IMS, IWLAN, and VPN roles; accessed 2026-07-29.

  57. Android Open Source Project (2026). EpdgTunnelManager.java. AOSP source, packages/services/Iwlan, commit b9ec687fee0110ee303921431a1eca2b1d35d500. Source. Carrier IWLAN IKE session and IPsec tunnel-interface use; accessed 2026-07-29.

  58. Android Open Source Project (2026). VcnGatewayConnection.java. AOSP source, packages/modules/Connectivity, commit 347fbd34b368d19f0d87e908ea101eed3601a731. Source. VCN IPsec tunnel transforms and IKE session migration; accessed 2026-07-29.

  59. Google (2025). The Android Show: I/O Edition. Google blog. Source. Published 2025-05-13; reports more than 3 billion active Android devices in over 190 countries; accessed 2026-07-29.

  60. Mullvad VPN (2026). Any App on Recent Android Versions Can Leak Certain Traffic. Mullvad blog. Source. Accessed 2026-07-29.

  61. IVPN (2026). Android VPN Leak via QUIC. IVPN knowledge base. Source. Independent reproduction; accessed 2026-07-29.

  62. GrapheneOS Community (2024). Are VPN Leaks a Ghost of the Past from Now On? GrapheneOS discussion forum, response 6. Source. External-router upstream-tunnel guidance; accessed 2026-07-29.

The Daily Front Page 12 of 30
Saturday, September 12, 2026 The Daily Front No. #260912 — Signal’s Integrity Check
article

How Trail of Bits helps verify the integrity of Signal chats

by dgroshev·▲ 68 points·41 comments·blog.trailofbits.com ↗
How do you know the server gave you the right key?

Every Signal chat starts the same way: the client asks the Signal server for the public key associated with your contact’s phone number. But how do you know the server gave you the right key? A compromised server could provide a false public key, allowing the client to encrypt messages to an attacker rather than the intended recipient.

Until now, the only way to detect such malfeasance was to verify safety numbers with your contact in person or over a trusted channel. Signal recently launched an alternative: Automatic Key Verification, a feature that helps validate that your chats are secure without requiring direct safety number comparison. Trail of Bits built and operates one of the three auditors that make this system trustworthy. Our auditor, which is an independent implementation written from scratch, continuously checks that the Automatic Key Verification system behaves honestly.

How key verification works

Automatic Key Verification is a form of “key transparency” that makes mismatch attacks harder to hide by creating a globally consistent view of the set of public keys associated with each phone number. The Signal app now performs a periodic self-check to ensure that all keys stored in the global map for your account belong to your devices. If the app is unable to verify the log, or finds that not all keys are expected, the user is presented with a warning that “Automatic Key Verification is currently unavailable for your device.” Automatic Key Verification may also be unavailable for other reasons, as outlined in Signal’s documentation.

What our auditor does

Automatic Key Verification depends on external auditors. Trail of Bits helps this system function by providing external verification that the user ↔ public key map is globally consistent and well formed, and does not hide any entries. Each time a new entry is added, we update our local copy of the map, stored as a Merkle tree. Periodically, we sign the head of the tree using a signing key that only we know. Because we commit to only ever signing one consistent lineage of Merkle trees, clients know that they are seeing the same set of public keys as everyone else in the system. Clients currently require signatures from each of three auditors: one operated by Signal, one operated by Cloudflare, and one operated by Trail of Bits.

When Automatic Key Verification is turned on, the Signal client periodically fetches Merkle tree heads from the Signal key transparency server. The client requires that each tree head belong to a lineage endorsed by all registered auditors within the last seven days. If the server does not present valid auditor signatures, the client will raise a warning and Automatic Key Verification will fail. A fully malicious server may therefore maintain a split view of the system for at most one week before client applications start to display warning messages.

We chose to implement our auditor from scratch, based on the specification, to provide independent verification; the code is open source. Signal also publishes a reference implementation.

We will provide updates to this blog post if we need to make substantive changes to our signing policy, such as resetting the state of our auditor or rotating our signing key. Our current public key is:

7fe5d91de235188486d8fb836a6da37e625e2b10eb6d144185b9364cc83cbbb6

How to use Automatic Key Verification

You can enable Automatic Key Verification in Signal by going to “Settings > Privacy > Advanced” and enabling Automatic Key Verification. In supported chats, you can verify the public key of your counterparty by visiting the safety number verification screen and clicking “Verify Automatically.” Automatic Key Verification often does not support chats where you started the conversation by searching for a recipient’s username. See Signal’s help page for more information. If automatic verification fails, users should fall back on safety number comparison.

Why we’re doing this

We believe that free and private communication is a critical public good. We are not paid by Signal or any other party for this service; we operate it in the interest of users and the community broadly.

Some form of public key integrity is an important component of any full end-to-end encryption system. If you would like to implement key transparency or end-to-end encryption generally, contact us.

The Daily Front Page 13 of 30
Saturday, September 12, 2026 The Daily Front No. #260912 — The Right to Repair
article

I fixed a tractor using John Deere's self-repair service. Farmers aren't sold

by sbulaev·▲ 119 points·122 comments·wired.com ↗
Why is barely anybody using it?

The agricultural machine manufacturer has a subscription service that lets owners repair their own equipment. Why is barely anybody using it?

person in john deere tractor in field

Courtesy of John Deere

There is something wrong with the tractor. The water-in-fuel sensor, a small device embedded in the John Deere machine that monitors the integrity of its diesel engine, is disconnected. And I’ve got to fix it.

I’m not a mechanic or a farmer. I’m poking at a laptop at John Deere’s corporate office in Santa Clara, California. There’s a cable connecting the computer to a 5130ML tractor sitting on the grass nearby. A bright red notification sits on the screen, telling me something is wrong. I type in a question and the software pulls up instructions and a few images from the user manual that matches the serial number of the machine.

John Deere tractor plugged in with a cord

Photograph: Boone Ashworth

I wander over to the tractor and find the parts modeled on the screen. I spot the sensor, with two very obvious dangling wires that need to be connected together. I press them together, hear a little click, then step away and look back at the screen. The notification has cleared. Everything seems to be fine. In just a couple minutes, all by myself, I’ve fixed the water-in-fuel sensor. This fix was an easy one. It's a WiF, really.

John Deere close up parts. Waterin fuel sensor

Photograph: Boone Ashworth

The John Deere employees surrounding me offer their congratulations. They’ve all been very patient with me, someone who has written several articles for WIRED about their company’s seemingly steadfast reluctance to let anyone but its official technicians repair its products.

In the group is Jahmy Hindman, senior vice president and chief technology officer of John Deere, who has flown in from Illinois. Hindman has worked for the company for 30 years. In that time, John Deere has thrived in an era of technological advancement. He’s here now to show off the company’s self-repair software Operations Center Pro Service, a system that pairs with nearly every electronic-enabled machine John Deere has made.

“It is owner-friendly,” Hindman says. “I think that's always been our intent.”

Lots of farmers would beg to differ.

Oh Deere

John Deere is the largest agricultural machine manufacturer in the US. In the past couple decades, machines of this ilk have become far more technologically advanced, loaded up with controllers, radios, and sensors. Many newer machines are equipped with 4G wireless connections that enable the collection of precise location data and other metrics. Some tractors can even run without a human operator at all, guided by built-in software systems to run autonomously on predetermined paths.

Over the last decade, John Deere has become a big target of the ever-growing right-to-repair movement. The company’s high-tech software systems also come equipped with digital locks, which require official, dealer-sanctioned approval to access. If you want to fiddle with a part, or replace it with an aftermarket alternative, those software restrictions can block those attempts.

That frustration has manifested in multiple lawsuits and even federal action. In April, the company agreed to pay out a $99 million settlement to its equipment owners to pay back burdensome repair costs. In July, the company settled a separate lawsuit brought by the US Federal Trade Commission that would require the company to make its products more repairable. John Deere did not admit wrongdoing in either case. Repair advocates say the two measures have been insufficient, both in financial terms and as requirements that will actually change Deere’s behavior. (Both settlements are not yet final until a judge issues a final ruling.)

Pro Tip

John Deere disagrees, of course, citing its recent efforts to make the case that its self-repair program works just fine.

The company introduced its Pro Service in 2025. Prices start at an annual cost of $195 per individual machine. For bigger operations or independent repair shops, the price to access every machine in John Deere’s fleet starts at $4,995 a year. That access for agricultural and turf services costs $5,995 per year.

Pro Service lets owners connect Deere’s software to their equipment via Wi-Fi or cable. Based on the serial number of the machine, the software can be used to find information from manuals and John Deere’s broader data services to diagnose any maintenance the machine needs. John Deere also recently introduced an AI chatbot called JD to its Operations Center self-repair software tool, but I did not have a chance to use that version during my visit. The AI features are meant to make it easier to find information about the components in any given machine.

“As the tractor products grow in size, they also grow in number of controllers that are necessary in order to make it all work,” Hindman says, noting that in some cases tractors can have upwards of 50 controllers. “So it can be super complex.”

Theoretically, Pro Service is the ultimate platform to find any equipment information and diagnose any problem. If something goes wrong, the software will guide you through a fix. Some features in Pro Service can be downloaded and used offline in case a connection isn’t working when something needs fixing. You can customize the layout to your liking and keep track of the full history of the software updates on each machine. If you need a new part, the software will helpfully serve you a link to the product in John Deere’s shop. It feels as smooth as scrolling through an Amazon page.

Kelli Sullivan, a group product manager of marketing automation and governance at John Deere, is the person leading the demo. She says owners are in full control of their data and can choose whether to share it with John Deere or an independent repair shop, or keep it private.

Data as a Service

The thing John Deere does seem to struggle with is enticing people to use its Pro Service, or even making its customers aware of it. (That’s the reason I’m here at this demo.) A year on, John Deere says its service has about a thousand daily users—a relatively small number, given that there are somewhere around 1.8 million farms in the US.

“We’re just trying to get the word out,” Sullivan says. “We have a suite of tools that are already available in our customers’ hands that allow them to do those things.”

Jared Wilson is a farmer in Missouri who uses Deere machines to care for his 3,000 acres of corn and soybean crops. He is also a named plaintiff in the $99 million lawsuit against John Deere who has called that payout ruling “fundamentally flawed.” When I describe my demo experience to Wilson, he laughs, saying the service works best on particularly simple repairs.

“If you have multiple failures, you have to figure out where to start,” Wilson says. “If you have three different or four different diagnostic codes, where do you start in the tree to figure out what's wrong?”

Wilson does not use Pro Service. He doesn’t know anyone who does, saying the farmers he knows tend to use cracked versions of the software they find on gray markets instead of giving their money to John Deere. Even with Pro Service, he says, farmers don’t always trust that John Deere will give farmers the tools they really need to fix their stuff.

“It's taken a class action lawsuit and FTC action to get them to come this far,” Wilson says. “If you believe what they're saying now—that this is actually it and we're getting everything we've asked for—then I've got some oceanfront property in Arizona to sell you.”

One of Wilson's biggest concerns with using a service like Pro Service, he says, is that it’s almost impossible to compare the dealership version of the software with the customer version of the software for parity. (Hindman says the versions of the software are all the same.)

Part of the problem are PIPs, or product improvement plans, which is a term John Deere uses to identify potential or existing problems in its machines. John Deere can issue a PIP, which then goes out to let people who own a particular piece of equipment know there’s an issue with the machine. But those missives only go out after the system has gathered the information that an improvement might be needed in the first place.

If a user faces a novel problem with how a machine works, it would still take time for that issue to get back to John Deere, which would verify it, then issue a PIP in Pro Service. In the meantime, the person dealing with the broken tractor won’t have the needed information to figure out exactly what is going on. Ideally, many of them would like to be able to troubleshoot the problem themselves, without having to use the company’s proprietary software.

“They're still not giving us access to that next level that we have to have for some of the most difficult problems that we encounter,” Wilson says.

Willie Cade, a board member of the advocacy group Repair.org and frequent critic of John Deere, takes that a step further.

“They're just monopolists that want to make money and as much as they possibly can,” Cade says. “Their real customer is not the farmer. Their real customer is Wall Street.”

At the demo, I ask about the repair advocates and what they get wrong about John Deere. Hindman says he doesn’t mind the pressure.

“We have the best intentions of making sure that we have the ability to give customers the tools that they need to keep their equipment up and running,” Hindman says. “We share that goal in common. Are there things that we need to continue to work on? I'm sure that there are. But let's not forget that we are a long way towards moving the needle to a much better state.”

The repair advocates are still pushing on John Deere. Agricultural equipment like tractors are some of the most durable goods, built to last for years and even decades of hard manual labor. If society isn’t able to establish a precedent to let people fix these machines, advocates argue, it becomes harder to pass legislation or judicial rulings that would solidify people’s rights to fix their own personal devices.

“This is an issue that's concerning agricultural equipment,” Wilson says. “But it's really the vanguard of a much bigger and persistent problem that affects every aspect of our society.”

Update, September 11 at noon: Additional information about the status of the two legal settlements was added to the story.

The Daily Front Page 14 of 30
Saturday, September 12, 2026 The Daily Front No. #260912 — The Map Belongs to Everyone
article

Make your first edit to OpenStreetMap

by juliantigler·▲ 393 points·92 comments·high5apps.github.io ↗
Make a meaningful contribution to OpenStreetMap in less than 15 minutes.

Website Wizard JOSM Plugin

Intro

This quick tutorial will help you make a meaningful contribution to OpenStreetMap (OSM) in less than 15 minutes. By the end, you will have added an official website tag to a nearby shop or amenity. Soon after, your contribution will be ingested into dozens of free OSM-based services, helping people worldwide.

Why a website tag? Once a place in OSM has a website tag, it becomes way easier to determine other helpful info about that place. Nearly every place’s official website has info about its phone, opening_hours, email, and other tags. So adding a website tag is a great place to get started.

Got your stopwatch out? Ready, set, go?

1. Create an OSM account

Sign up for a free OSM account and then confirm your email.

2. Download and Run JOSM

Download JOSM (~365 MB) for your specific operating system and then run it.

JOSM, the Java OSM editor app, is a powerful tool for querying and editing OSM data. While simpler in-browser editors exist, JOSM offers plugins that make your edit as quick and easy as possible.

3. Download OSM Data

JOSM's Download panel with an area of interest selected

  1. Press Ctrl+Shift+↓ (or ⌘+Shift+↓ on Mac) to open the Download dialog
  2. Determine your area of interest (AOI). It should be somewhere you’re familiar with, no larger than a few city blocks.
  3. Locate your AOI on the map. You can pan the map with ctrl+click dragging and zoom in by scrolling.
  4. Click and drag to create a box around your AOI
  5. Click ⬇️ Download. If this fails, your AOI was probably too large. Choose a smaller AOI and try again.

4. Filter Irrelevant OSM Data

Unfiltered OpenStreetMap data for an area of interest

Filtered OpenStreetMap data for an area of interest

Now we’ll filter the OSM data to only show shops and amenities that don’t have a website.

  1. Find the Filter panel on the right side of the screen

  2. Click the + icon to open the Filter dialog

  3. Copy/paste the following query into the Search string text field

     name=* ((amenity=* "addr:housenumber"=*) | shop=*) -website=* -"contact:website"=*
    
  4. Click Submit filter

  5. Check E, uncheck H, and check I in the Filter panel. You should now only see the relevant places in your AOI.

5. Set Up the Website Wizard Plugin

The Plugins panel in JOSM's Preferences panel

  1. Press F12 (or ⌘+, on Mac) to open JOSM’s Preferences dialog
  2. Click the 🧩 puzzle piece icon on the left side to open the Plugins config
  3. Click ⬇️ Download list
  4. Scroll down the list of plugins until you see 🌐 WebsiteWizard
  5. Check its checkbox
  6. Click OK to install it

6. Search for an Official Website

Website Wizard demo search in the JOSM editor

  1. Click the 🌐 icon on the left side of the screen to show the 🌐 Website Wizard panel on the right side of the screen
  2. Type your AOI’s city and/or neighborhood into Website Wizard’s Search Prefix text field
  3. Click a shop or amenity in your AOI
  4. Click Search to open DuckDuckGo in your default browser with the query autofilled as the Search Prefix + the place’s name
  5. Determine if any of the search results represent the official website for your place. Do NOT use search results for social media profiles, review sites, or other business aggregators. When in doubt, don’t use it. If you don’t find one, just repeat steps 3 to 5 with another place in your AOI.
  6. Copy/paste the official website URL into the Website URL text field
  7. Click Save. If you make a mistake, you can always press ctrl+z (⌘+z on Mac) to undo it.

7. Upload Your Changeset

JOSM's Upload panel with our changeset's info

  1. Press Ctrl+Shift+↑ (or ⌘+Shift+↑ on Mac) to open the Upload dialog
  2. Type Add website to <city and/or neighborhood> shops and amenities in the text field labeled Provide a brief comment…
  3. Select survey in the dropdown labeled Specify the data source…
  4. Click Upload Changes, which will open OSM in your browser
  5. Enter your OSM credentials
  6. Click Log In
  7. Click Authorize to allow JOSM to create the changeset for you

Conclusion

🎉🎊🥳 Congratulations- you just made OSM a little bit better for everyone! Plus you got a preview of some of the cool things you can do with JOSM and OSM data.

So what next?

You could keep going and add a website tag to every place in your AOI. For example, I quickly added 66 new website tags in Seattle’s Wallingford neighborhood in this changeset.

Or you could modify the filter to show places in your AOI with a website tag but no phone tag. Then you could add the phone tag based on info from the website.

Or you could spread the word about OSM and Website Wizard! The United States alone has more than 1 million shops. So if we want to put them all on the map, we’ll need to get many more people to help.

No matter what’s next, feel proud for pushing OSM a little closer toward becoming the world’s greatest map!

The Daily Front Page 15 of 30
Saturday, September 12, 2026 The Daily Front No. #260912 — Bodies, Uncommonly Ordinary
show hn

Show HN: Bodily Oddities

by vesterde·▲ 322 points·199 comments·vester.si ↗
143 strange things bodies can (sometimes) do.

143 Strange things bodies can (sometimes) do.

It felt scary but it's normal Things you can try right now

Did you know Ear rumbling Up to about half of all people can make a low roar inside their own head by tensing a tiny muscle in the middle ear. Ears Trick

43–55% of people

Browse by body region

Brain 50Nerves 26Eyes 39Ears 17Mouth and throat 17Skin and hair 15Muscle 29Bones and joints 12Internal organs 8Heart and blood 11Whole body 12

Browse by kind

Trick 26Involuntary 30Perception 27Sound 8Variation 24Reflex 18Rare ability 10

Recently added

Brain Involuntary

Jamais vu

A familiar word, face or place suddenly feels strange, as if you are meeting it for the first time. It is the opposite of déjà vu, and researchers think it is rarer. In lab studies, more than half of people felt a common word turn strange after copying it about 30 times.

Unknown. Thought to be rarer than déjà vu.

Brain Perception

Prosopagnosia

People with prosopagnosia, or face blindness, find it hard to recognise faces, even those of family and friends. Most have had it all their lives. It can also start after brain damage. Estimates run from about 1 in 100 people to about 1 in 40, depending on how it is tested.

About 1–2.5% of people

Brain Perception

Aphantasia

Most people see a picture in their mind when they think of a friend's face. People with aphantasia see only a faint image or none. About 4 in 100 people are like this, and about 1 in 100 see nothing at all. A few people have images almost as vivid as real sight.

About 1% have no mental images. About 4% have faint or none.

Bones Variation

Camptodactyly

A finger that stays bent forward at its middle joint and cannot fully straighten, usually the little finger. Under 1% of people have it. Mild cases rarely cause trouble, but the bend can get worse while a child grows.

Under 1% of people

Muscle Variation

Linburg–Comstock variation

An extra tendon link in the forearm that ties the thumb to the index finger. When the thumb bends into the palm, the index fingertip bends with it. Roughly 1 in 5 people have it, but counts vary widely.

About 1 in 5 people. Studies range from 5% to 60%.

Heart Reflex

Oculocardiac reflex

Pressure on the eyeball, or a pull on an eye muscle, makes the heart slow down. It matters most during eye surgery, especially in children. It is not a safe way to slow your own heart.

Found in 14–90% of eye muscle operations. Its strength varies a lot.

Try one right now

Eyes Perception

Afterimages

An image that stays in your vision after the thing you looked at is gone. Bright things leave a same-colour copy for a moment. Long stares leave a copy in opposite colours.

Everyone

Brain Trick

Ambidexterity

Using both hands equally well for every task. About 1% of people are truly ambidextrous. Many people who say they are ambidextrous are mixed-handed: right hand for some jobs, left for others.

About 1%

Ears Variation

Attached and free earlobes

Some earlobes hang free and others join the head directly. Textbooks taught this as one gene, but a 2017 study found at least 49 genetic regions involved, and the two shapes blend into each other.

About 61% attached in one Korean sample

Mouth Variation

Bifid uvula

The uvula, the dangling flap at the back of the throat, is split into two lobes, in about 1 to 2% of people. On its own it is harmless, but in a baby it can be a clue to a hidden cleft of the palate.

1–2% of people

Eyes Perception

Blue arcs of the retina

Look at a small red light in a dark room and two faint blue arcs curve away from it towards your blind spot. The arcs follow the paths of the nerve fibres in your own retina.

94% of healthy eyes in one clinic test

Eyes Perception

Blue field entoptic phenomenon

Tiny bright dots that dart along wiggly paths when you look at a clear blue sky. They are your own white blood cells moving through the capillaries of the retina.

Most people

It felt scary but it is normal

Things that alarm people but are harmless.

Bones Variation

Accessory ossicles

Small extra bones, most often in the foot, that form when a growth centre never fuses with its main bone. They are usually harmless but can be mistaken for fractures on an X-ray.

About 2 in 5 feet

Brain Perception

Alice in Wonderland syndrome

Objects, body parts, or the whole room suddenly look far too big, far too small, or far away. It mostly happens to children, often during a fever or infection, and it passes on its own. In adults it is usually linked to migraine.

6–7% of teenagers (size distortion)

Muscle Involuntary

Benign fasciculations

Small visible twitches in the eyelid, calf, thumb, or arm that come and go with no weakness. One survey found them in about 70% of healthy clinicians. They are very rarely a sign of nerve disease.

About 70% of clinicians in one survey

Brain Involuntary

Brain zaps

A brief feeling like a small electric shock inside the head, often set off by moving the eyes. It is most common when stopping or lowering an antidepressant. It is unpleasant, but usually not a problem.

15–50% of people stopping an antidepressant get withdrawal symptoms

Brain Involuntary

Call of the void

A sudden, brief thought of jumping when you stand somewhere high, with no wish to do it. About half of people get it. It is probably the brain's safety system being misread, not a hidden wish.

43–60% of people

Bones Variation

Cervical rib

An extra rib above the normal top rib, growing from the lowest bone in the neck. About 1 in 100 people have one, and most never know.

About 1% of people

See all 143 oddities Filter by region, kind or tag.

Not medical advice. Bodily Oddities describes what human bodies do. It does not diagnose anything. If something worries you, see a doctor.

The Daily Front Page 16 of 30
Saturday, September 12, 2026 The Daily Front No. #260912 — Four Centuries Beneath the Lakes
article

Great Lakes sturgeon may be 400 years old:Scientists rethinking how to save them

by bookofjoe·▲ 153 points·47 comments·cbc.ca ↗
Great Lake sturgeon may live more than 400 years.

New research suggests Great Lake sturgeon may live more than 400 years — far longer than previously thought. The discovery is forcing scientists to rethink the already-generations-long effort to bring the giant fish back from the brink of extinction.

Human activity brought these giant fish to the brink of extinction. Now they can't recover without us

Two people wearing life jackets pose for a photo on a boat while holding a large fish between them.

University of Windsor biology Prof. Trevor Pitcher, right, holds an adult lake sturgeon with the help of graduate student Olivia Galloway on the Detroit River, a place where the fish is making a comeback thanks to restoration efforts. (Dane Roberts)

Some of the lake sturgeon swimming around in the Great Lakes today may have been alive when Samuel de Champlain first reached Lake Huron more than four centuries ago.

New research suggests these ancient-looking, armoured fish may live far longer than scientists previously thought — in some cases, more than 400 years.

But despite their extraordinary longevity, lake sturgeon — which can weigh up to 180 kilograms and grow up to two metres long — have barely survived us. In Ontario, lake sturgeon are listed as a species at risk, and endangered in the Great Lakes specifically.

Overfishing and destruction of their spawning habitat wiped out an estimated 99 per cent of the lake sturgeon's historic population. More than a century later, populations in parts of the Great Lakes still rely on human beings raising and releasing young fish to help them recover.

The new study uses more than four decades of data from lake sturgeon caught and recaptured across the Great Lakes, suggesting recovery could be a much longer project than scientists anticipated.

A new way to count sturgeon birthdays

"We're pretty comfortable saying that there are potentially 400-year-old lake sturgeon swimming around," said Ed Baker, a fisheries scientist with the Michigan Department of Natural Resources who led the study.

Traditional estimates suggest lake sturgeon can live for as long as 150 years. But figuring out how many birthdays a sturgeon has had isn't as easy as counting candles on a cake.

WATCH | Trevor Pitcher introduces us to a Great Lakes giant:

Meet the ancient giants of the Great Lakes

Scientists have typically estimated a sturgeon's age by counting annual growth rings in a cross-section of its fin ray. But the older a sturgeon gets, the more slowly it grows — and the harder those rings become to distinguish.

Instead of counting growth rings, Baker and his colleagues used 44 years of capture data to calculate how much individual sturgeon grew between one capture and the next — then modelled how long it would take them to reach adult size.

A U.S. Fish and Wildlife Service employee holds a lake sturgeon in the water.

A U.S. Fish and Wildlife Service employee holds a lake sturgeon in the water. As Canada's largest freshwater fish, they can grow more than two metres long and weigh up to 180 kilograms. (Sharon Rayford/U.S. Fish and Wildlife Service)

In one remarkable case, a male sturgeon measured 117 centimetres when researchers caught the fish in 1990. When the same fish was caught again 34 years later, it had grown just seven centimetres, or about 0.21 centimetres per year.

When researchers used those real-world growth rates to calculate how long it would take a fish to reach full size, some of the resulting ages stretched into centuries.

"It's a completely unexpected result," Baker said. "But it's one that's supported by the data, for sure."

The very long game of saving sturgeon

For Trevor Pitcher, a University of Windsor biologist who has spent years trying to rebuild lake sturgeon populations, the finding raises a much bigger question: What does it mean to conserve an animal whose lifespan may be measured in centuries?

"I think it changes everything," he said. "Assuming they do live, let's say 200 to 400 years, we need to reframe the way we do conservation."

A lake sturgeon swims along the bottom of the Great Lakes.

Lake sturgeon were once abundant across the Great Lakes before overfishing and habitat destruction wiped out an estimated 99 per cent of their population. New research suggests some could live for more than 400 years. (Zach Melnick/Inspired Planet Productions)

Canada's existing conservation plan is a century long, Pitcher said, with the expectation the work will outlast the individuals who started it. If the new longevity estimates hold up, the assumptions behind that plan may have to change, he said.

"We would have to readjust some of our plans about restoration, and some of the habitat work would certainly be different."

A small armoured-looking sturgeon fingerling in a person's hand.

A lake sturgeon fingerling rests across the palm of a U.S. Fish and Wildlife Service employee. If it survives, this tiny fish could become enormous over a life span that could last centuries. (Jennifer Johnson/U.S. Fish and Wildlife Service)

Pitcher said sturgeon eat large numbers of invasive quagga mussels and gobies, helping hold those populations down. Because historic overfishing and habitat destruction by humans reduced sturgeon to roughly one per cent of their former population, he said, "we kind of shot ourselves twice in the foot."

"When you take out a key species like sturgeon, it changes everything in the food web."

Pitcher's work is already extending beyond his own lab. His team is helping train people in Georgian Bay and Walpole Island First Nation to raise lake sturgeon themselves — spreading the expertise needed to rebuild populations around the Great Lakes.

A spiritual connection that runs deep

For Perry McLeod-Shabogesic, an elder and knowledge keeper from Nipissing First Nation, about 40 kilometres west of North Bay, Ont., the importance of lake sturgeon extends far beyond their role in the ecosystem.

"Because of their longevity, they were considered like a grandfather fish."

A man with a hat and glasses and a number of framed pictures behind him.

Perry McLeod-Shabogesic spoke to CBC News from his home in Nipissing First Nation, on the shores of Lake Nipissing in northern Ontario. (CBC)

A fish that can live for centuries experiences and witnesses things across generations, he said, and it's the reason sturgeon are regarded as elders in Anishinaabe tradition.

"They're seen as elders — fish elders, if you will," he said.

Their appearance adds to that sense of age. Lake sturgeon belong to a lineage of fish that stretches back more than 150 million years — long before humans appeared.

It was only relatively recently that their numbers collapsed, as habitat destruction and commercial fishing nearly wiped them out across the Great Lakes.

Lake Nipissing once teemed with sturgeon. McLeod-Shabogesic remembers older fishermen describing spring spawning runs, when the fish crowded the Sturgeon and French rivers so thickly that their backs broke the surface.

"You could walk across the river, they were so thick," he said.

'Duh'

For McLeod-Shabogesic, the unusual traits also place the fish in another category; he describes sturgeon as beings that live "in between" — neither like other fish nor anything else in the lake.

Workers hold a metre and a half long sturgeon on a table aboard a research vessel in order to measure its size and tag it before release into the wild.

Workers with the Michigan Department of Natural Resources measure a lake sturgeon captured as part of the agency's long-running monitoring program. Decades of capture and release data allowed scientists to estimate how slowly the fish grow and how long they may live. (Michigan Department of Natural Resources)

In Anishinaabe teachings, he said, things do not always fall on one side or the other; between male and female is two-spirit, between hot and cold, there is warm. Sturgeon, he said, occupy a similar space.

That idea appears in an old Nipissing story McLeod-Shabogesic learned from his father. A family staying on the Manitou Islands caught and ate a mysterious black fish. By the next morning, they had transformed into serpent-like beings and disappeared beneath the ice.

The story, he said, speaks to the sturgeon's place in between worlds, and to its power and medicine.

"They weren't a fish completely and they weren't a snake completely. And then their age — they were different from the other fish."

So when McLeod-Shabogesic heard scientists now believe lake sturgeon may live for centuries, he wasn't surprised. "Duh," he said, laughing.

"It validates things to a certain degree," he said. "These are things that we already knew that our teachers and my ancestors — my grandparents and those fishermen that have fished this lake for a long time — they all knew these things already."

The Daily Front Page 17 of 30
Saturday, September 12, 2026 The Daily Front No. #260912 — Wasm’s Long Race
article

Performance of WebAssembly Runtimes in 2026

by fagnerbrack·▲ 92 points·24 comments·00f.net ↗
WebAssembly runtimes are getting faster.

I wanted to know if WebAssembly runtimes are getting faster.

This is a follow-up to the earlier libsodium WebAssembly benchmarks from 2019, 2021 and 2023.

Not “does the newest version beat native code in one microbenchmark?”, and not “which runtime has the prettiest benchmark chart?”, but something more boring and more useful:

If I take the same C crypto code, compile it to WebAssembly, and run it on the latest runtime, a runtime from one year ago, and a runtime from two years ago, are things actually improving?

So I benchmarked libsodium on WebAssembly runtimes released around June 2024, June 2025, and June 2026.

The short version:

  • wasmer is the best performer, but WAVM, WAMR and Wasmtime are close.
  • WAVM has the best optimizer, and is able to generate very fast code out from baseline, portable WebAssembly
  • The new WebAssembly wide_arithmetic instructions are a big deal for crypto code when runtimes support them.

What I measured

The test program is libsodium’s benchmark suite, built from libsodium commit 8e3be8615ba6adcd7babaecf5e76f516890ba5fb.

I built one native baseline and several WebAssembly variants:

  • native x86-64, compiled with Zig using the local CPU target
  • plain WebAssembly
  • WebAssembly with lime1
  • WebAssembly with lime1 and simd128
  • WebAssembly with lime1, simd128, and wide_arithmetic

For the native reference, libsodium was built with -Dcpu=native. For wasm2c, the generated C was compiled with zig cc -O3 -march=native.

For WAMR, I used AOT mode: wamrc compiled each .wasm file to an .aot file, and iwasm ran the resulting AOT file. wamrc doesn’t accept --cpu=native, so I used --target=x86_64 --cpu=x86-64-v4 --opt-level=3, which matches the host’s available x86-64 feature level and works across the WAMR versions that could compile these modules.

The native command was:

zig build -Denable_benchmarks -Doptimize=ReleaseFast -Dcpu=native -Diterations=3

The WebAssembly commands were the same shape, with a wasm32-wasi target and the feature-specific CPU strings:

zig build -Denable_benchmarks -Dtarget=wasm32-wasi -Doptimize=ReleaseFast -Diterations=3
zig build -Denable_benchmarks -Dtarget=wasm32-wasi -Doptimize=ReleaseFast -Dcpu=lime1 -Diterations=3
zig build -Denable_benchmarks -Dtarget=wasm32-wasi -Doptimize=ReleaseFast -Dcpu=lime1+simd128 -Diterations=3
zig build -Denable_benchmarks -Dtarget=wasm32-wasi -Doptimize=ReleaseFast -Dcpu=lime1+simd128+wide_arithmetic -Diterations=3

The host was an AMD Ryzen AI 9 HX 470 with 12 cores and 24 threads. CPU boost was disabled and the maximum CPU frequency was 2 GHz. The OS was Linux 7.1.0-rc7, and Zig was 0.17.0-dev.948+e949341b7.

The numbers below are the geometric mean of per-benchmark slowdowns relative to the native build. Lower is better. A value of 2.0 means “twice as slow as native” on this machine.

I used ITERATIONS=3, so the very small libsodium tests are noisy and quantized. Rows reporting zero time were excluded from the aggregate.

Versions

For every runtime except WAVM, I used the latest stable release available on June 23, 2026, plus a stable release from roughly one year earlier and one from roughly two years earlier.

Runtime 2024 2025 2026
Bun 1.1.16 1.2.17 1.3.14
Node 22.3.0 24.2.0 26.3.1
WAMR 2.1.0 2.3.1 2.4.4
WABT wasm2c 1.0.35 1.0.37 1.0.41
WasmEdge 0.14.0 0.14.1 0.17.0
Wasmer 4.3.2 6.0.1 7.1.0
Wasmtime 22.0.0 34.0.0 46.0.0
WAVM n/a n/a nightly/2026-04-05
Wazero 1.7.3 1.9.0 1.12.0

WAVM is awkward to compare historically. The old available nightly collapsed to a 2022 binary for both the 2024 and 2025 slots, and that binary refused to run on this machine. I only kept the 2026 nightly.

WAMR 2.1.0, the selected 2024 release, installed fine but its AOT compiler failed on these Zig-generated modules with invalid WASM stack data type. I kept the version in the matrix, but didn’t include an aggregate for it.

Baseline WebAssembly

This is the plain WebAssembly build, without lime1, SIMD, or wide arithmetic.

Baseline WebAssembly slowdown by release year

There isn’t one universal trend.

Wasmtime steadily improved: 2.67x native in 2024, 2.54x in 2025, 2.41x in 2026. It got faster every year, and the gains land in the tenths place, above the noise.

Node also improved slowly, from 8.60x native to 7.95x native.

Wazero was basically flat: 4.84x, 4.70x, 4.72x native. No real movement over two years.

WAMR in AOT mode was already fast in 2025 and stayed there in 2026: 1.59x native, then 1.57x native, the same within this benchmark’s noise. I don’t have a complete 2024 WAMR number because WAMR 2.1.0 couldn’t compile these modules.

Wasmer regressed in the 2025 release I tested, then recovered in 2026. The 2026 baseline barely beats 2024.

wasm2c improved modestly in 2026. It remains one of the best options if ahead-of-time translation to native C is acceptable for your deployment model.

Bun is the outlier. Its 2024 and 2025 results were far behind, but the 2026 result is about three times faster than the 2025 result. It’s still slower than Node on this benchmark, but the direction is excellent.

WasmEdge is fast too, but its command-line behavior changed enough to matter. My first 0.17.0 run accidentally used interpreter mode for compiled modules and looked catastrophically slow. Running the compiled modules with --run-mode=aot fixed it: the 2026 baseline was 1.74x native, between the 2024 and 2025 baseline results.

Best supported build by year

The baseline table is useful because it compares the same WebAssembly target everywhere.

But if you’re choosing a runtime for your own deployment, you probably care about the fastest build that runtime can actually run.

So for each runtime and year, I also selected the best complete result among the supported builds: baseline, lime1, lime1+simd128, and lime1+simd128+wide_arithmetic.

Best supported build by release year

Looks similar to the previous graph, except for Wasmtime and Wasmer that really benefit from wide_arithmetic.

Ranked by the best supported build, the complete current-year results are:

2026 releases, best supported build, ranked

CPU feature variants

The WebAssembly feature story is more interesting than the year-to-year runtime story.

For the 2026 releases, these were the aggregate slowdowns:

CPU feature variants across 2026 releases

lime1 and simd128 alone aren’t magic here. Sometimes they help, sometimes they hurt, and sometimes the difference is lost in benchmark noise.

wide_arithmetic is different.

Only Wasmtime and Wasmer could run the full wide_arithmetic build among the complete stable rows I tested. WAMR rejected it with unsupported opcode 0xfc13. But when wide_arithmetic worked, it was the biggest speedup in the whole experiment:

  • Wasmtime 46.0.0: 2.41x native without it, 1.46x native with it.
  • Wasmer 7.1.0: 2.08x native without it, 1.33x native with it.

That’s the kind of change I like. A lot of libsodium’s expensive operations are arithmetic-heavy. If the WebAssembly ISA can express that arithmetic directly, the runtime has much less work to rediscover what the C compiler already knew.

Failures

Most runs completed cleanly, but not all of them.

Bun 1.2.17 failed box_easy in the baseline build. Bun 1.1.16 failed pwhash_argon2i in the lime1 and lime1+simd128 builds.

Node 22.3.0 originally failed pwhash_argon2i, pwhash_argon2id, and pwhash_scrypt in the baseline, lime1, and lime1+simd128 builds. The failures weren’t fixed by increasing Node’s JavaScript heap or stack settings. They were fixed by giving the Wasm modules an explicit maximum linear memory. With the baseline build, a 1024-page maximum, or 64 MiB, made pwhash_argon2i, pwhash_argon2id, and pwhash_scrypt complete. But pwhash_scrypt failed with 512 pages and segfaulted again at 1536 pages and above, so this appears to be a V8 memory-mode threshold rather than a simple “more memory is better” setting.

WAMR 2.1.0, the 2024 slot, couldn’t compile even the baseline modules in AOT mode. WAMR 2.3.1 and 2.4.4 compiled and ran the baseline, lime1, and lime1+simd128 builds, but not wide_arithmetic.

Those failures were excluded from the aggregate. So were benchmark rows with a zero reported median.

So, are runtimes getting faster?

Some of them are.

Wasmtime is the cleanest yes: it got faster every year in this benchmark, by a little each time.

Node is also a yes, but the slope is gentle.

Bun is a loud yes between 2025 and 2026. It still has a lot of ground to cover for this workload, but the improvement is too large to ignore.

Wazero is mostly flat.

WAMR is also mostly flat between the versions that worked here, but “flat” at about 1.4x to 1.6x native is a very good place to be.

Wasmer is mixed if you only look at the baseline, but the 2026 release supporting wide_arithmetic changes the practical answer for crypto code. With that feature enabled, it was the fastest complete 2026 result I could compare across a normal current release.

wasm2c remains good. If you can translate WebAssembly to C ahead of time and compile it for the host, it’s hard to beat.

WAVM produced the fastest 2026 baseline number, but I don’t have a fair 2024 or 2025 comparison.

WasmEdge remains excellent once it’s forced into AOT mode. The accidental interpreter-mode run was a good reminder that command-line defaults are part of the benchmark, too.

Takeaways

If you run CPU-heavy cryptography in WebAssembly, runtime choice still matters a lot.

The spread between the fastest complete current result and the slowest current result is large: Wasmer with wide_arithmetic was 1.33x native, while current Bun baseline was 8.77x native.

Feature support matters too. The same runtime can move from “pretty good” to “surprisingly close to native” when the WebAssembly module can use better arithmetic instructions.

The comforting part is that the mainstream runtimes aren’t standing still. Wasmtime improved steadily. Bun made a huge jump. Wasmer gained a feature that matters for real crypto workloads. WasmEdge remained fast once the AOT run mode was explicit.

The less comforting part is that WebAssembly performance still isn’t one thing. It depends on the runtime, the release, the enabled WebAssembly features, whether the code goes through WASI from JavaScript, and whether ahead-of-time native compilation is allowed.

So benchmark your actual workload.

But if your workload looks like libsodium, the answer in 2026 is: WebAssembly can be close to native, wide_arithmetic is worth caring about, and yes, some runtimes really are getting faster.

The Daily Front Page 18 of 30
Saturday, September 12, 2026 The Daily Front No. #260912 — Builds, Measured
article

I made a build visualizer to understand Bun's compile times

by lalitmaganti·▲ 110 points·21 comments·lalitm.com ↗
More often than not, there are fixable problems.

I built buildprof (Github), an open-source tracing tool that shows where the time goes when you compile software on Linux. Here’s a realtime video of it profiling a clean build of ripgrep:

Watch the buildprof demo

Sometimes, builds are slow because there is simply a lot of code to compile. But more often than not, there are fixable problems: poor parallelism, repeated work, dependency downloads or a huge compiler/linker invocation. buildprof makes all of this clearly visible, so you can see what’s worth investigating and optimizing.

You run it by putting buildprof -- in front of any build command you already use:

buildprof -- make -j16
buildprof -- cargo build
buildprof -- ninja -C out/target
buildprof -- just build
buildprof -- ./dev/custom-build-script.sh

buildprof records every process your build command launches, including their subprocesses (and their subprocesses…), and lays them out on one timeline. Time moves from left to right, bar width shows duration, and child processes appear beneath whatever launched them.

A ripgrep build: Cargo spans the whole build, rustc invocations compile crates in parallel, and the final rustc invocation launches a linker chain.

I made buildprof because this tweet from Jarred Sumner, chief architect of the Bun JavaScript runtime, was living rent free in my head:

Jarred Sumner’s comparison showing a 30 minute 6 second median Linux build for Bun 1.3.14 and 5 minute 37 second median for Bun 1.4.0.

Specifically, the claim that Bun’s new Rust build was >5× faster on Linux than its old Zig build really bothered me. In my experience, Zig projects had usually compiled much faster than Rust projects of similar complexity. That intuition was enough to make me feel there was a mystery to solve.

This was further compounded by another important, yet easily missed, detail in the tweet: the Zig build used Full LTO, while the Rust build used ThinLTO.

Compilers normally optimize separate compilation units largely in isolation.1 Link-time optimization (LTO) lets them optimize across those boundaries. Full LTO brings those units together into one large optimization job, while ThinLTO preserves more separation so much of the work can run in parallel.

From past experience, this difference can have an enormous effect on build time. The tweet mentioned it in passing, but I wondered how much of the headline improvement it explained.

I started by trying to reproduce the numbers.

The numbers reproduced. But now what?

I checked out Bun 1.3.14 and Bun 1.4.0 and wrote some scripts to replay their Linux x64 CI builds on a 6-core, 12-thread Linux VM. The scripts preserved the build steps and their dependencies, running everything on one machine.2

My timings were in the same ballpark as Jarred’s:

Linux x64 build Zig era Rust era
Bun’s reported CI median 30m06s 5m37s
My single-machine CI-profile replay 24m24s 5m40s

OK, so the gap showed up on my machine too. But a lot had changed between the two measurements besides the language; so what was actually responsible? Was it the Zig compiler that was taking all that extra time? Or maybe it was the Full LTO link? Or perhaps there was something else in Bun’s build I hadn’t even thought to look at?

This is where my profiling and developer-tools brain kicked in. Usually, when I’m trying to understand why something is slow, I want a trace: what happened, when it happened and how long it took. It would be really cool to have that for these builds, to put them on a timeline and see where their time actually went.

But a build involves a lot of different tools, each with its own idea of what’s happening. What could I record that would let me see across all of them?

Builds are process trees

When you type cargo build or zig build, it feels like you are running one program. The build system works out what needs to be rebuilt, the ordering between those pieces and what can run in parallel. But generally, it does not perform all that work itself; it launches compilers, code generators, archivers, linkers and arbitrary scripts. Which can launch more programs which launch some more…

Different build systems describe that work in different ways. Cargo sees crates, Ninja sees build edges and CMake generates instructions for another build system. From the operating system’s point of view, however, they (mostly) look like processes launching other processes.3

A Rust build, for example, might contain a chain like this:

cargo
└── rustc
    └── cc
        └── collect2
            └── ld.lld

If we record when each subprocess starts and ends, we can lay them out on a timeline. Here’s what that chain looks like in buildprof:

The final link in a ripgrep build, showing cargo launching rustc, then cc, collect2 and ld.lld beneath it.

There are also several nice properties to visualizing a build at this layer:

  1. It’s build-system agnostic: Cargo, Ninja, Zig, Make and most other build systems do much of their work by spawning processes, so we do not need to write a special integration for each one.
  2. It naturally includes custom scripts: This includes both scripts above the build system (repository setup, dependency fetching) and scripts underneath it (code generators, asset processors).
  3. We can follow the files between build steps: recording which files each process reads and writes lets us see which steps produce the inputs for others. This even works across build systems!

This gave me a starting point for buildprof: record the process tree, then turn it into a timeline I could explore. There are plenty more details to get into, which I will do later. But once I had that working, I could finally go back to my initial question: what was Bun doing for those twenty-four minutes?

Pointing it at Bun

Why was the Zig CI build so much slower?

I started by recording the Zig-era CI build with buildprof, using the same scripts as before:

The complete Zig-era CI build

Explore in buildprof

Right away we can see a huge problem: the ld.lld linker invocation dominates the build time. It ran alone at the very end for over sixteen minutes, about two-thirds of the entire build. What the heck was it doing for all that time?

Clicking on the linker shows its command line, which buildprof captures automatically:

The selected Zig-era linker and its Full LTO flag

There’s Full LTO, just as Jarred said. Given how long the link was taking, it was now my main suspect.

But the process tree alone couldn’t tell me whether LTO was actually responsible for those sixteen minutes. Thankfully, LLD records its own internal timing events, and buildprof can include them when you use --compiler-traces.

I recorded the final link again, this time with --compiler-traces enabled:

LLD’s internal phases

Explore in buildprof

Now we can see that LTO is where almost all the time goes. The linker is running compiler passes over the program, not just combining already-compiled files. The OptModule bar alone takes just over ten minutes and includes the passes which generate machine code.4

How did the Rust CI build differ?

With so much of the Zig build spent in LTO, I wanted to see how much time the Rust build spent linking. I recorded that build too:

The complete Rust-era CI build

Explore in buildprof

Just 2m24s. And this time, as expected, the linker command contains -plugin-opt=thinlto:

The Rust linker invocation with ThinLTO enabled

Both builds were doing LTO, but with different settings and very different link times. What if I kept Bun’s Zig code and changed Full LTO to ThinLTO? How much of the gap would that close?

Trying ThinLTO

I switched Zig Bun’s build flags to ThinLTO and recorded another clean build, along with a fresh Full-LTO build for comparison:

The matched Full-LTO build and partial ThinLTO experiment

Explore in buildprof: Full LTO · partial ThinLTO

The link got 3m40s faster in this pair of recordings, but it was still taking nearly thirteen minutes. Why was linking still so expensive?

Looking back at the compiler trace, a lot of the work was on functions with JSC in their names. That’s JavaScriptCore, the engine Bun uses to execute JavaScript. The linker was spending time compiling the JavaScript engine too.5

Clicking on the linker invocation showed the WebKit libraries among its inputs, including libJavaScriptCore.a:

The linker command has ThinLTO enabled but still includes WebKit’s libraries, including libJavaScriptCore.a.

Following those inputs back through the build, I found that Bun wasn’t compiling these libraries itself. It was downloading them from a separate WebKit build. And when I checked that build’s flags, there it was again: -flto=full. The Rust build used a newer WebKit revision whose build recipe selected ThinLTO.

Even though I had changed how Bun compiled its own code, those downloaded libraries still contained Full-LTO inputs and so the linker still had to optimize that code and turn it into machine code. To change that, I would have to rebuild WebKit too.

Rebuilding WebKit

I checked out the historical WebKit revision and rebuilt it and its ICU dependencies with compatible ThinLTO settings. Then I replaced the downloaded libraries with the ones I had built, keeping the ThinLTO changes to Bun.

Here are the recorded builds:6

Zig-era build Whole build Final linker
Original Full LTO 24m24s 16m35s
Bun ThinLTO; original WebKit archives 20m20s 12m55s
Bun ThinLTO; rebuilt ThinLTO WebKit and ICU 15m11s 7m22s

The link now took 7m22s. Still slower than the Rust build, but enough of an improvement that I wanted to look beyond the linker.

What about the rest of the build?

The build still took fifteen minutes, and nearly eight of those passed before the linker even started. What was it waiting for? I went back to the original CI trace to follow the inputs from Bun’s own code.

buildprof also records which files each process reads and writes. If a process reads a file another wrote, it links the two together under the hood. Turning on “Show on timeline” draws those links as arrows. Here, the linker reads libbun-profile.a from the C++ compilation and bun-zig.o from Zig. Both arrive through copy steps; following those back takes us to the processes which produced them:

Following the linker’s dependency arrows through the copy steps to the C++ and Zig producers. The producer panels use the same time scale; C++ finishes first.

The C++ side of the compilation finished first. The linker was waiting for bun-zig.o, so it could not begin until the Zig branch had finished too.

It was at this point I went back to the Rust build and compared against how it worked, and the main reason the Rust build was faster became obvious: Bun has been split into >90 crates, while in Zig it was all trying to compile as a single Zig module!

Cargo fanning out into named rustc processes across Bun’s crates, next to the single zig build-obj process which spawns nothing at all.

This meant that the Zig build cannot parallelise the same way Rust can. I also suspect, though I did not prove this, that it explains the slow linking: the linker has to optimize one huge ThinLTO bitcode module instead of the same work spread across crates.

It was at this point I had to stop: to go any further, I would have to split up the Zig module myself, and given that this code is all obsolete anyway, I didn’t think it was worth doing that.

Summarizing:

  • The huge outlier in the initial Zig build vs the Rust build was the massive linker step which ran alone at the end of the build.
  • Changing the LTO settings for just Bun was not sufficient as WebKit, a significant part of the build, still used Full LTO.
  • Once I had done this, the Zig build dropped from twenty-four minutes to fifteen.
  • Even after this, linking still took 7 minutes and the whole build 15 minutes.
  • The overwhelming difference which remained was structural: Rust spreads compilation across >90 crates while the Zig build funnelled everything through a single module.

And fwiw, the traces had also turned up a few things I couldn’t resist poking at…

Other things hiding in the build

A build can contain almost anything

In the middle of Bun’s CI build, I found commands asking the public internet for the machine’s IP address, inspecting running Docker containers and reading the latest Git commit message.

Small CI setup commands visible in the process tree

These take well under a second altogether. Nothing to optimize but I just wasn’t expecting to find them in a build trace.

A cold dependency fetch

The builds above reused downloaded dependencies, so I also recorded a fresh WebKit fetch. Downloading and extracting the archive took about twenty seconds. For the first twelve, all we see is Node running. Then it launches tar and gzip, and we can see the extraction separately.

A cold WebKit download and extraction

Looking inside one C++ compilation

Earlier, we followed the linker’s inputs back to Bun’s C++ compilation. We can look inside those compiler invocations too. I picked one of the last files to finish, ZigGeneratedClasses.cpp, and replayed its Ninja command with --compiler-traces. For Clang, buildprof enables -ftime-trace and adds its internal timings to the process timeline.7

Clang’s frontend and backend phases while compiling ZigGeneratedClasses.cpp

The replay took about twelve seconds, split almost evenly between Clang’s frontend and backend. Zooming in further, we see ModuleInlinerWrapperPass, one of the phases of Clang, accounts for over four seconds of the backend’s work.

How buildprof works under the hood

The recording side of buildprof uses ptrace, the same Linux interface used by debuggers. I did consider both eBPF and ftrace, but ptrace is just straight up perfect for exactly this type of problem; eBPF tracing means CAP_BPF and CAP_PERFMON permissions and hooking into potentially unstable tracepoints/kernel functions. While with ftrace, I’d have to juggle tracing instances to avoid interfering with other users, and getting the filters perfect for just the build process and all its descendants is cumbersome.8

With ptrace, I can launch the build and follow its children directly. Its built-in events tell buildprof when processes fork, exec a new program or exit. And for filesystem activity, buildprof uses a seccomp filter to intercept only the calls it needs.

How much buildprof costs is almost entirely down to how many files the build opens. For ripgrep, recording barely changed the build time. Redis opened files much more often, and recording added about five seconds:

Build Untraced Processes only Processes + files
ripgrep / Cargo 12.27s 12.30s 12.43s
Redis / Make 26.78s 27.04s 31.89s

If that overhead gets in the way, you can turn off filesystem tracing with --no-file-events and keep the process timeline.

I work on Perfetto, so it was a natural starting point for the UI; buildprof’s UI is a soft fork of the Perfetto UI. I could have just opened the recordings on ui.perfetto.dev, but I wanted control over how the process tree was laid out, which details appeared when you clicked a command, and things like those on-demand arrows between file producers and consumers.

Fortunately, we’ve spent the last several years working on making the Perfetto UI extensible through plugins. Most of buildprof’s UI is reusing that infrastructure. Perfetto handles the hard stuff (parsing traces, querying events, rendering the timeline and managing workspaces) and I get to focus on what makes those things useful for builds.

I plan on going into a lot more detail about the recorder and UI in a separate technical post.

Did I need to build something new?

These days it’s very easy to make a tool just because you can. But that wasn’t the case here; before building buildprof, I looked long and hard for an existing tool that could give me this view.

I started with ninjatracing, which I’ve used many times. It turns Ninja’s build log into a timeline showing what ran and how much ran in parallel.

Here’s the Ninja log from the Zig-era build.

But Ninja only sees part of Bun’s build. The scripts which invoke it are missing from its log, and commands it runs appear as single blocks even when they launch whole trees of subprocesses.

There were several other tools, each covering different parts of the problem:

  • Cargo timings works well for Cargo-managed builds, but cannot break down arbitrary work inside build.rs or see wrapper scripts above Cargo. In Bun, Cargo is only part of the build: the report I captured covered 1m51s of a 5m40s CI build.
  • Clang’s -ftime-trace gave us the detail inside a compiler invocation, but cannot show what the rest of the build is doing while Zig’s Tracy integration goes deeper still and is intended more for understanding the compiler itself.
  • strace and tracexec can follow arbitrary processes through fork and exec, but show general process events rather than a build-oriented timeline.

What the Fork (via) came closest: it follows processes across build systems and presents a build-specific view. But as far as I could tell, it still appears to be in private beta and there don’t seem to be any plans to make it open source.

What’s next for buildprof

buildprof already does what I wanted it to do, and I plan to keep working on it as I use it on my own builds. But there are a few things I’d like to improve.

Recording overhead is one; the Redis measurements showed there’s room to improve filesystem tracing, especially for builds which open lots of files. I’d also like to support macOS where I do some of my work and maybe Windows if there’s interest.

There are also more build systems and toolchains I’d like to test, including npm, Gradle and Bazel. Computing critical paths would also be a big improvement: we followed dependencies by hand in this post, but buildprof could help identify the chain of work holding up the build and automatically annotate it.

I’ll probably tackle these as and when I need them. But if you try buildprof and there’s something you wish it could do, I’d be interested to hear about it. What people find useful will help me decide where to spend more time.

Conclusion

I managed to satiate my curiosity, though I ended up spending rather more time on this than I expected. Along the way I built a tool I now want to have around whenever a build is taking too long.

I know I’ll come back to buildprof the next time a slow build annoys me. If you have one of those builds too, give it a try. I’d love to hear what you find!


  1. In C and C++, a compilation unit is usually a source file together with its included headers. Rust compiles crates, which can be split into multiple code-generation units. Zig normally compiles a program’s Zig sources together as a single compilation unit. Bun’s Zig compiler fork supports splitting that into multiple LLVM modules, but its CI build explicitly selected one when LTO was enabled↩︎

  2. The Zig-era CI build ran its C++ and Zig compilation stages on separate Buildkite machines and passed their outputs to a final linking stage. My script ran those stages concurrently on one machine, waited for both outputs, copied them locally instead of transferring them over the network, then linked them. This should preserve the dependency graph, but due to the hardware differences and running both stages on one machine, resource contention would obviously be quite different. Also note that my timings are individual runs (albeit ones which were quite stable) while Bun’s reported figures are medians. ↩︎

  3. A process can do substantial work internally, including running multiple threads, without launching anything else. The process timeline won’t show that parallelism. To see inside a process, we need tracing from the program itself, as Clang and LLD provide in the examples below. ↩︎

  4. LLVM emits OptModule from its legacy pass manager, which LLD uses for code generation. The inlining and other IR optimization passes can appear before it, so this bar is not the total time spent optimizing a module. ↩︎

  5. In the earlier Full-LTO linker replay, 26,825 OptFunction events with JSC symbols total about 209s. This is summed event time, not a measurement of JavaScriptCore’s entire contribution to the link. One example event takes 2.94s; its symbol demangles to JSC::JITThunks::initialize(JSC::VM&)↩︎

  6. Recording script. These timings are just for building Bun with the libraries already available; the WebKit and ICU rebuild happened beforehand and isn’t included. Of course, I could point buildprof at that build too, but that’s another rabbit hole… I did not rebuild a matching Full-LTO WebKit archive as a control, so I cannot attribute every second saved to the LTO setting alone. ↩︎

  7. buildprof currently supports compiler traces from Clang, LLD and nightly Rust. ↩︎

  8. eBPF tracing uses capabilities such as CAP_BPF and CAP_PERFMON, as described in the kernel’s capability definitions. ftrace provides separate tracing instances and PID filters, but these still need configuring and access to tracefs. ptrace also depends on the host’s security settings; containers may need additional permissions to allow tracing child processes. ↩︎

  9. Medians of five clean builds per mode on the same VM, with six build jobs. Measurement script↩︎

The Daily Front Page 19 of 30
Saturday, September 12, 2026 The Daily Front No. #260912 — Rust’s Impossible Value
article

Stabilizing Rust's Never Type

by cjd8·▲ 164 points·43 comments·lwn.net ↗
The type the language uses to mark a function that never returns.

A function's return type is supposed to indicate the kind of data that it produces. Rust's "never" type, which is denoted by an exclamation mark ("!"), is the type the language uses to mark a function that never returns and other places where a value can never occur. For a long time, the never type was used internally by the compiler, but was considered an unstable feature. On August 24, after more than two years of work, Rust-compiler-contributor "waffle" finally managed to stabilize the type. It took so long, in part, because it involved a small breaking change to previous Rust editions, which the compiler maintainers needed to ensure did not impact much real code.

(Note: Rust also uses exclamation marks to indicate calls to macros. The way the syntax is constructed, a place where it is valid to use the never type is not a valid place to put a macro invocation and vice versa.)

Why a never type?

There are two reasons that Rust has a never type, one practical and one philosophical. The practical reason is that it allows for more efficient generic code. For example, consider the FromStr trait in the standard library, which is used for types that can be instantiated from a string:

    trait FromStr: Sized {
        type Err;
        fn from_str(s: &str) -> Result<Self, Self::Err>;
    }

FromStr::from_str() either returns a converted result, or a custom error type. For example, attempting to convert "foo" into an integer will return a ParseIntError. But some types have an infallible conversion. For example, it is always possible to convert a string into a ByteString. That implementation of FromStr could set Err to be the never type. Then the compiler would know that the error branch of the returned Result is never present, and could optimize out all of the code that touches it or checks for it.

    impl FromStr for ByteString {
        type Err = !;
        fn from_str(s: &str) -> Result<Self, !> { ... }
        // Keeps the same generic interface,
        // but generates code equivalent to:
        // fn from_str(s: &str) -> Self { ... }
    }

The philosophical reason involves correct type inference. In Rust, constructs such as if statements and while loops are expressions; their results can be assigned to a variable. The compiler needs a type to infer for the result of an infinite loop, if the programmer writes one. That shouldn't come up often in real code, but it turns out to simplify type inference to be able to treat that case uniformly, rather than adding special rules to handle it.

In particular, the never type has a useful property for simplifying code: it automatically coerces to any other type. This sounds strange, but it is safe, since the never type represents the "result" of a computation that will never produce a value. So, anywhere that the code claims to have a value of the never type, the compiler knows that it can't possibly reach that code, and therefore it's safe to ignore it. This is a form of type-system-driven dead-code elimination.

For both of these reasons, Rust programmers have wanted to be able to use the never type in stable versions of the language. Making that happen required resolving a particularly thorny corner case.

Never fallback

Because of the way that conversions from the never type to other types are implemented, the compiler can sometimes end up in a situation where it cannot naively infer the concrete type of an expression. Consider this example, which defines an anonymous function (using ||, which is like Python or LISP's lambda) that never returns, and then calls it in a way that expects a concrete error type (using the ? operator):

    let function_that_never_returns = || { loop {} };
    function_that_never_returns()?;

That infinite loop is given a type of ! which is then implicitly converted to whatever the function is supposed to return. But since the function is defined locally and not given an explicit type, the compiler does not have sufficient information to say what that type is. The problem could be fixed by giving the function an explicit return type:

    let function_that_never_returns = || -> Foo { loop {} };

Since such an annotation would only be required in cases where the function cannot return anything, however, it would be a bit pointless to require the programmer to assign a fictitious type to it. So, the compiler includes a special rule: if, after all other type inference has been done, there is still an ambiguous type that cannot be determined, just assume that it should be the designated fallback type. Prior to the 2024 edition of Rust, that fallback type was () (the unit type, which has exactly one possible value). In the 2024 edition, the fallback type was changed to ! itself, essentially canceling out the implicit conversion. In the compiler internals, the never type still gets converted to an unknown type and then falls back, but from the programmer's perspective the behavior is identical to having the never type only undergo implicit conversion when required for the types to make sense.

That change of behavior was, technically, a breaking change. Type inference for some code could change, which could in turn cause compilation errors. That is the purpose of Rust's edition system: allow breaking changes in the front-end design of the language without breaking older code or requiring the whole ecosystem to update at once. In this case, however, there were reasons to want the new behavior backported to old editions.

Never infallible

For many years, the standard library has had an Infallible type to work around the unstable nature of the never type. It served the same semantic purpose as the never type, but did not have any special compiler support. Therefore, code using it would be technically correct but suboptimal (such as having an extra layer of tags in an enumeration or emitting dead code), because the optimizer would not always be able to remove references to Infallible. It was planned that, when the never type was eventually stabilized, Infallible would become a type alias for ! and all that old code would silently become more efficient. However, people pointed out a handful of ways that redefining Infallible had accidentally been made into a breaking change. Since ! has implicit conversions, changing the definition of Infallible could result in existing code needing additional type specifiers in order to type check.

Luckily, changing the definition of Infallible and changing the default fallback type, while both breaking changes, nearly cancel out. Any code that refers to the standard library's Infallible type by name would continue to work; it is only places where type inference is implicitly expected to produce Infallible that pose a risk of breaking existing code. With Rust's lack of implicit conversions in most cases, that will most often come up in places where the never type used to be implicitly converted to Infallible. If Infallible is made to be a type alias of the never type, then those places may experience never-type fallback, which would, in turn, change the inferred type and cause a compilation error if the never fallback type were not updated at the same time.

With both changes occurring simultaneously, the Rust maintainers believed that almost all existing Rust code would continue to compile — but "almost all" is not a reassuring qualifier when dealing with backward-incompatible changes. The Rust community does have a solution to this in the form of crater, which can download and compile all publicly available Rust libraries from crates.io in search of code that is broken by a compiler change.

Never say never

Waffle ran crater in April and found that, while there were 3,300 crates negatively impacted by the change, only seven were fully broken, with the rest broken by depending on old versions of libraries that had since been fixed. In the latter case, the problem would theoretically be fixable by releasing backported fixes for a handful of core libraries. This is not an accident; Rust has been emitting a warning whenever code triggers never-type fallback in a way that will break with the new change since 2024, so most libraries had plenty of time to update of their own initiative. The most common remaining error observed by crater is code that calls a generic function without enough type information for the compiler to pick a specific return type. Consider this function:

    fn foo<T: Default>() -> Result<T, Error> { ... }

It returns either a value of some caller-chosen type T that must implement the Default trait, or an error. If it is called without specifying a value for T, however, then type fallback can kick in:

    // No type specified.
    foo()?;

Previously, this would have made the compiler assume that T should be (), which implements Default, and so the code compiles. After this change (and on the 2024 edition), the compiler assumes that T should be !, which doesn't implement Default, and therefore causes a compilation error. The fix is to explicitly specify the type that foo() should return, either in the call or by pattern-matching assignment.

    foo::<()>()?;
    // or
    () = foo()?;

Even though it's not a complicated change, the Rust maintainers were not willing to break 3,300 crates. Waffle was asked to work with the maintainers of common libraries to backport simple changes like the above (making a new patch version, which many Rust build environments will pick up automatically), in order to reduce the number of libraries depending on broken dependencies. Several library authors were willing to make the backports, but some refused on the grounds that those old versions were past their end of life. Those maintainers pointed out that users could stay on an older version of Rust or update to the maintained version of the library. Even so, the successful backports addressed 1,553 of the failing crates.

After fixing a handful of related problems to reduce the number of broken crates even further, the Rust maintainers eventually agreed that even though there would still be some broken code it was worth making the change to simplify the language. So, starting in Rust 1.99, the never type will be stable and Infallible will be a type alias for the never type. Users who find that this breaks their code have a few options:

  • Stay on Rust version 1.98.
  • Update their dependencies to supported versions that include a fix for the problem.
  • Add a patch to explicitly specify the return types of affected function calls.

On the one hand, this is a breaking change, and people may see code that had remained stable and working suddenly fail to compile. That could be seen as a violation of Rust's commitment to backward compatibility. On the other hand, the problem is relatively rare, there are multiple simple ways to fix it, it has been warned about for years, and it has always been part of the plan for the language. Additionally, the Rust maintainers worked directly with the community to find and address the breakage, even going so far as to help backport fixes to long-dead versions of popular libraries. So, the whole process could also be seen as an affirmation of Rust's commitment to backward compatibility.

In the future, people learning the language will hopefully find the never type just a little less special. Either way, most users of Rust will probably not be affected at all, but never say "never".

The Daily Front Page 20 of 30
Saturday, September 12, 2026 The Daily Front No. #260912 — A Test for the Timing Bug
article

Testing Race Conditions

by alpaylan·▲ 64 points·3 comments·projectzero.google ↗
Many security bugs are race conditions.

Many security bugs are race conditions, where multi-threaded execution has to occur with the right interleaving for a negative effect to appear. This creates challenges for several use cases:

  • Confirming bug candidates that have been discovered manually or through static analysis.
  • Regression tests: After fixing a race condition bug, there is often no good way to write a regression test that reliably triggers the bug as part of a test suite.
  • Automatic bug discovery, such as fuzzing: It is hard for a fuzzer to exercise all interesting interleavings of concurrent operations, or reach code paths that are only exercised when operations are racing.

I mostly discover bugs by manually reading code. When I think I’ve found a bug, I normally write a test case to either prove or disprove that the bug exists. For race condition bugs, it can be hard to achieve either outcome. For Linux kernel bugs, I often resort to recompiling the kernel after adding conditional mdelay() calls (which spinloop for roughly the specified amount of time) in appropriate places; I usually make these conditional based on the name of the running thread, though sometimes more complex conditions are needed. On platforms that support DTrace (like macOS and Windows), it is possible to use DTrace probes that call chill() for similar effect, though the utility of this is limited as DTrace can only trace on non-inline function boundaries or explicit trace points, rather than on every instruction. Regardless of platform, this approach can be time consuming and can require trial and error to definitely determine whether code is buggy.

Additionally, in the Linux kernel, fixes for race condition bugs are often accompanied by hand-written ASCII diagrams showing problematic thread interleavings with call graphs and relevant memory accesses (for example, see this recent rt_spin_unlock UAF fix, or this recent jbd2 deadlock fix). It would be convenient to have developer tooling that can analyze potentially vulnerable code and show results in a similar representation.

Summary

I wrote tools for exploring possible interleavings of multi-threaded test cases for the Linux kernel:

  • A tool that automatically tests all possible A-B-A interleavings of a test case.
  • A terminal UI for manual exploration of possible interleavings.
  • A GUI for manual exploration of possible interleavings.

The kernel part of this is intended to also be usable for discovering race conditions via fuzzing, but userspace tooling for that still needs to be implemented.

The tools are available on GitHub under the name MAccConc, short for “Memory Access Concurrency”; see the README there for installation and usage instructions.

If you just want to see the tooling in action, skip to Demo: automatic testing.

If you’re just interested in the theory behind the tooling, read section Stable identifiers for memory accesses across runs: count-augmented stack traces.

Prior work

This project was inspired by discussions with Ned Williamson, whose sockfuzzer project involved exploration of concurrency bugs by using a custom scheduler that can reschedule at synchronization primitives to explore interleavings. See the conference talk slides and recording focused on the concurrency testing aspect of this.

My tooling is largely based on ideas similar to SKI, but SKI uses a different implementation: It records memory accesses and controls scheduling of vCPUs using a patched version of QEMU in TCG mode, and uses VM snapshots to explore different execution interleavings.

Discovering memory accesses that could contribute to race conditions (communication points)

As described in the SKI paper, interesting execution interleavings of a given multi-threaded test case can be discovered by tracing memory accesses of all threads and searching for pairs of accesses on two threads that could interact with each other - meaning, roughly, that at least one of them is a write operation, and they access overlapping memory ranges. The SKI paper calls such memory accesses communication points.

This requires some mechanism to collect memory access coverage. SKI did this by patching QEMU’s TCG mode; I am instead relying on ASAN instrumentation in “outline” mode (compiler backend flag asan-instrumentation-with-call-threshold=0, selected by CONFIG_KASAN_OUTLINE in the Linux kernel), which generates helper function calls on memory access. I believe that the kernel is the right place to collect this data because it would allow the kernel to also provide higher-level information about lock acquire/release events and such, though I have not implemented this at this time. Implementing this in the kernel also means that it would theoretically be possible to test on bare-metal hardware, rather than inside VMs.

Since Linux already has KCOV as a mechanism to feed basic block kernel coverage information to userspace, I decided to use the same mechanism to record information about memory accesses. An alternative would have been to use ftrace, which is oriented towards tracing use cases, and includes a function graph tracing mode built on fentry hooks and more complex output buffer management that is oriented towards use cases including system-wide data collection. I chose to use KCOV because of its simpler in-memory representation of trace data (which could become relevant for recovering trace data from crashed VMs); because it uses static always-on instrumentation rather than runtime-enabled instrumentation with near-zero overhead in disabled state; and because my impression is that KCOV is designed for higher-frequency trace events than ftrace.

Implementation detail: ASAN and TSAN

ASAN normally merges helper calls for subsequent memory accesses. To receive one callback per memory access, the kernel patches explicitly disable this compiler optimization using the asan-opt-same-temp backend flag.

ASAN is intended for identifying UAF, so it does not emit helper calls on direct stack memory access unless there is potential for out-of-bounds access. This means that some race conditions involving on-stack objects, such as wait queues, may not be detectable with this. ASAN also by default emits no helper calls for access to globals, but this optimization can be disabled using the asan-opt-globals backend flag.

An alternative would be to use TSAN instrumentation instead, which is designed for detecting data races and also provides information about access atomicity. The downside of TSAN instrumentation is that compilers do not support emitting both ASAN and TSAN hooks at the same time - so to still have working detection of memory safety violations (like UAF) while using TSAN hooks, it would be necessary to run the kernel’s ASAN implementation off of the TSAN hooks or change the compiler.

Implementation detail: KCOV and background work

Some race conditions involve background work, for example:

  • receive processing of loopback network packets
  • RCU callbacks

KCOV can optionally collect remote coverage for background work in some subsystems; however, in upstream Linux, most types of background work that would be interesting for me are not yet integrated with this mechanism, and remote coverage is currently mainly used for fuzzing subsystems that handle incoming data from devices, like bluetooth and USB.

Enabling this for other parts of the kernel should be relatively straightforward, and I have a draft patch for doing this for RCU callbacks.

Stable identifiers for memory accesses across runs: count-augmented stack traces

To test out different orderings of memory accesses, a way to stably identify interesting memory accesses across test case executions is needed. Identifying memory accesses based on the data address would not work if the data address was located in an object which is freshly allocated during each test case execution; and identifying memory accesses solely by instruction address would not work well if the memory access was in a function like memcpy() or spin_lock().

SKI solves this using VM state snapshots, so that each execution starts from the same global state.

I am instead identifying memory accesses with count-augmented stack traces, where each stack trace element essentially consists of a callee function address and a number indicating how many calls to this callee should be skipped in the calling stack frame.

An example of the semantics of a count-augmented stack trace would be something like: “On this thread, look at the second call to __x64_sys_recvfrom, then within that, the first call to __sys_recvfrom, then within that the first call to sock_recvmsg, then within that, the first call to unix_stream_recvmsg, then within that, the first call to unix_stream_read_generic, then within that, the second call to _raw_spin_unlock, and then within that, the first memory access at instruction address X”.

This unambiguously identifies a point in an execution trace, is independent of concrete data addresses, and is relatively stable with regards to changes in the control flow of irrelevant parts of the trace.

To make this work, KCOV must provide information about function entry/exit events so that when userspace is parsing KCOV coverage output, it can keep track of how the call stack changes. Doing this nicely requires compiler support as part of SanitizerCoverage; I landed an LLVM feature patch for this a few months ago (see documentation), which landed in the LLVM 23.1.0 release.

Forcing execution orderings with delay injection

To force specific execution orderings through KCOV, I implemented an ioctl KCOV_SET_DI using which userspace can request that actions (essentially wait/wake) are taken on memory accesses at specific count-augmented stack traces. (See documentation in my kernel branch.) Each action either sets one flag, or waits for one flag to be set, at a userspace-provided index in a shared array of flags. The possible action types are:

  • DI_STACK_WAKE_PRE: before the memory access, set flag N
  • DI_STACK_WAIT: before the memory access, spin-wait until flag N is set
  • DI_STACK_WAKE_POST: after the memory access, set flag N

With the same ioctl, userspace also configures an upper limit on spin-wait iterations.

Additionally, there are ioctls for userspace to directly interact with the same flags.

This API enables two different ways of using delay injection: constraint-style delay injection and fully-specified ordering.

Constraint-style delay injection (A-happens-before-B)

Userspace can set up a series of A-happens-before-B constraints, where each such constraint is implemented as a pair of actions in different threads that operate on the same flag:

  • DI_STACK_WAKE_POST for the access that should happen first
  • DI_STACK_WAIT for the access that should happen second

With this approach, the execution ordering is left partly non-deterministic. This is what the GUI and terminal UI tools currently implement.

An advantage is that this is somewhat more intuitive for simple cases; however, it requires recording timing information to show the user approximately in what order events happened, and it can make the execution trace more complicated. It also often requires more constraints than a fully specified ordering, and is more complicated to reason about.

Fully specified ordering (context-switch-style)

Userspace can decide on a specific ordering in which events should occur, by picking points at which execution should transfer from one context to another. For the simple case with two execution contexts, this requires that thread A starts running a syscall while thread B begins by spin-waiting on a flag; then when thread A reaches some count-augmented stack trace, thread A uses a combination of DI_STACK_WAKE_PRE and DI_STACK_WAIT to pause its own execution and let thread B continue; and later, thread B can do the same to switch back.

This is the approach I used for the automatic A-B-A interleaving tester.

Demo: automatic testing

I’ll explain more background below; but first, here are two shiny demos on a toy example!

This is an example of using the automatic A-B-A interleaving tester on this test case with concurrent dup(5) and close(5) calls:

#define _GNU_SOURCE
#include <errno.h>
#include <fcntl.h>
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
#include <unistd.h>

static int test_fd;
static int dup_res, dup_errno;

void test_setup(void) {
  test_fd = open("/", O_PATH);
}

void test_thread1(void) {
  dup_res = dup(test_fd);
  dup_errno = errno;
}

void test_thread2(void) {
  close(test_fd);
}

void test_end(void) {
  printf("dup(%d) = %d (%s)\n",
      test_fd,
      dup_res,
      dup_res == -1 ? strerror(dup_errno) : "success");
}

It discovers one ordering where dup(5) returns 5, which is working as intended but might be a somewhat surprising result:

sh-5.3# ./kcov-autorace testcase/demo-dup-vs-close.so
loading kallsyms
RCU state (excluded): base=ffffffff82970100 len=500
loading testcase
initializing kcov
collecting A-B coverage
dup(5) = 6 (success)
testing candidates
dup(5) = -1 (Bad file descriptor)
dup(5) = -1 (Bad file descriptor)
dup(5) = -1 (Bad file descriptor)
dup(5) = 5 (success)
dup(5) = 6 (success)
dup(5) = 6 (success)
dup(5) = 6 (success)
dup(5) = 6 (success)
dup(5) = 6 (success)
dup(5) = 6 (success)
dup(5) = 6 (success)
stats:  injection-failed:0  wait-timeout:7  reordered:4
sh-5.3#

Demo: GUI

And here is an example of me using the GUI on the same test case, using it to manually force an ordering where dup(7) returns 7.

First, I launch the GUI, then run the test case once in the guest:

sh-5.3# ./kcov-vsock-client testcase/demo-dup-vs-close.so
dup(7) = 8 (success)

At this point, no ordering constraints are enforced yet; dup() and close() are racing randomly. The GUI shows in what order execution happened:

This current view just shows function call graphs from both threads (thread 1 with black indent, thread 2 with red indent). The close() syscall happened to execute after dup() this time. Normal functions are shown in black; inline functions are shown in green, but only shown if they called a normal function (since “all inline functions” is not ticked).

Ticking “filter to communication points” shows a bunch of memory accesses in blue, which are communication points (as defined above, in short: reads from locations to which other threads write and writes to locations which other threads access; kfree() counts as a write operation). Each memory access line shows the type of access (Read/Write/Free), data address, access size, and the memory value before the access. Hovering over an access highlights all overlapping accesses in yellow.

Left-clicking on a memory access shows a view that is instead filtered to only show memory accesses overlapping the selected access. Note that this can show reads that were not identified as communication points (because all writes happen on the same thread).

Left-clicking a function name shows a source code view on the right, interspersed with trace data. Data values loaded by memory reads are shown in red (under the source line and column to which the compiler attributes the access); data writes are marked similarly with a red “WRITE”; memory accesses that are communication points are prefixed with “INTERFERENCE” in orange. Function calls are shown in blue.

By right-clicking on two memory accesses in the call graph view, it is possible to create an ordering constraint between the two accesses, such that the kernel will attempt to make the first selected access happen before the second selected access. Each ordering constraint is shown on the right side, represented as two count-augmented stack traces. Note that the last bottom element of the stack actually identifies a specific instruction, but the UI doesn’t really show this. Also, the count-augmented stack traces shown here do not include inline functions.

In this case, I have created one ordering constraint that orders the second file descriptor table access in __fget_files_rcu() (which is inlined into __fget_files()) before the file descriptor table entry removal in file_close_fd_locked() (which is inlined into file_close_fd()). This ensures that the file descriptor table lookup in dup() successfully looks up the file descriptor table entry before it is cleared by the concurrent close().

I have created another ordering constraint that orders the spin_unlock(&files->file_lock) in file_close_fd() before the spin_lock(&files->file_lock) in alloc_fd() so that the file descriptor table entry has been released by the time dup() searches for an unused entry.

In this view, ordering constraints have been specified, but the test case has not yet been run with this specified ordering.

(This view is filtered to show accesses to the files_struct::file_lock.)

Now, re-running the test case shows:

sh-5.3# ./kcov-vsock-client testcase/demo-dup-vs-close.so
dup(7) = 7 (success)

And the new trace appears in the UI, with brown “DELAY INJECTION” lines interspersed to show how the ordering constraints were applied.

Note that the UI shows the ordering of events based on timing information that is associated only with memory accesses; the placement for any event other than a memory access is inferred based on that. In views filtered by data accesses, function entry events are additionally only shown at the time of the first displayed non-function-entry event. For example, in the following screenshot, the first thread may have already entered get_unused_fd_flags() by the time file_close_fd() called spin_unlock(), even though the events are shown the other way around. However, memory accesses should be shown in approximately the right order; with the caveats that the order of memory accesses might be wrong if events happened at the same clock value, and that timing information is recorded by instrumentation that runs directly before the actual access. (Building the tool on fully specified orderings instead would avoid such caveats.)

(This view is filtered to show accesses to the file descriptor table entry.)

More documentation is available inside the GUI.

Implementation status

For LLVM: The required patch has landed in LLVM 23.1.0.

For the Linux kernel: The required patches are not yet in the upstream kernel. I am posting the Linux kernel patch series for upstream review around the same time as this blog post; a git branch with my patches is also available on github (with a few more patches that aren’t yet ready for upstreaming). If you want to test this tooling, you will need to use my kernel branch for now. (See the README in the tools repository for build instructions.)

My kernel patches are in a clean state; the userspace tooling is a bit more hacky, in particular the GUI implementation.

The command-line tooling can only handle two concurrent threads, while the GUI can handle additional execution contexts (with the kcov-vsock-client harness: background work launched by thread A).

I am looking forward to hearing if this is useful to others, and maybe even what tools others manage to build on top of this! Feel free to reach out to me (for example via email to maccconc-tooling@google.com).

Future work

Use fully specified orderings instead of constraint-style for manual tooling

The non-automatic tooling currently uses constraint-style delay injection; but as described above, fully-specified orderings have several advantages, including more deterministic behavior. I might change the GUI implementation to use fully-specified orderings instead in the future.

Type information for human-readable memory access traces

For reading memory access traces as a human, it might be helpful to provide information on the object types that are being accessed. One way to do this would be to follow what Microsoft’s debugging tools can do with CodeView debuginfo and use debuginfo to associate memory allocation function call sites with type information, then let the allocator track the call sites from which objects have been allocated.

I proposed to add such a feature to the DWARF standard, which has been accepted and is included in the current DWARF 6 draft (search for DW_AT_alloc_type), and added enough support to LLVM to make it work in the same cases where it already worked with CodeView; but so far that only works for C++ new calls, I did not land the changes necessary to make it work for malloc.

Making this work in the kernel would require infrastructure that either queries allocator metadata for every memory access record or provides an initial snapshot of heap allocator metadata across the system plus metadata about subsequent memory allocations.

Higher-level memory access feedback

One inefficiency in my current prototype is that userspace receives no information about the semantics of locking operations. If two threads each perform lots of memory accesses on an object while holding a lock protecting the object, this will generate a large number of potential communication points, but actually a locked section just represents one big communication point. It might be helpful if the kernel provided “lock acquired” and “lock about to be released” events.

But that might not be a very general approach, since impossible orderings caused by locking are not so different from impossible orderings caused by things like an object being initialized before it is published to a global pointer or such.

Detecting impossible orderings faster: Deadlock detection

In my current implementation, when an attempt is made to force an impossible ordering via delay injection, the result is that one thread spins/waits on a lock until another thread reaches the delay injection timeout, which is inefficient. It might help to have integration with lock debugging infrastructure that can detect such a semi-deadlock in simple cases and abort the test case faster.

Fuzzing: Building up test cases with potential communication points like Snowboard

Snowboard (a project that searches for concurrency bugs caused by interaction between fuzzer-generated single-threaded test cases) used recorded information about memory accesses in single-threaded test cases to identify which test cases could have interesting communication points when executed in parallel. It would be interesting to build something similar on top of this KCOV-based instrumentation.

It might also be interesting to use this for single-threaded test case creation: Start by collecting memory access coverage for individual system calls, then use that to determine which syscalls might interact with each other in interesting ways when executed in sequence, and build up longer system call sequences this way.

This would be easier using VM snapshots (like SKI), since my approach does not lead to stable data addresses across test case executions; but it would probably be possible by identifying memory locations that are different between test cases abstractly based on allocation sites, as long as allocation site information is available for all objects that are allocated per test case execution.

KCOV output to host-shared memory

My current tooling loses KCOV output if the kernel under test panics, so it can’t be used for displaying what happened when a kernel crash occurred.

For use cases where the kernel under test is a KVM guest, it might be useful to give the host direct access to the KCOV output buffer. One way to do this might be to use pages in a file on virtiofs with DAX as the KCOV output buffer, and allow writing KCOV output into userspace-provided pages.

Make zeroday hard.

The Daily Front Page 21 of 30
Saturday, September 12, 2026 The Daily Front No. #260912 — The Art of Keeping Feet Planted
article

Inverse Kinematics and Foot Locking

by airhangerf15·▲ 139 points·14 comments·theorangeduck.com ↗
Foot sliding is something that affects almost all animation systems.

One of my most requested blog posts has always been something on inverse kinematics and foot locking. That isn't surprising - foot sliding is something that affects almost all animation systems, and proper foot locking and inverse kinematics is something that is tricky to get right, yet can have a huge positive impact on the visual quality of the resulting animation.

But speak to any two animation programmers and they will likely give you two completely different answers as to the best way to solve foot-sliding. Undoubtedly this topic is much more of an art than a science - which might explain why it isn't always easy to find resources on it - and why it gets less attention in the academic sphere.

With that being said, I thought I would share a selection of little recipes related to solving foot-sliding. And while I'm certain these are not the proven-best solutions to this problem (and there are likely many games which do it better), they have served me well over the years, and so should make a good starting point for anyone getting into the topic.

For digestibility I've split the topic into a few parts. First I'll cover my method for solving a leg chain to position the toe at a desired location. Second, I'll cover my solution for locking the position of the toe during contacts at runtime using inertialization. Third, a technique for automatically annotating contacts in animation data. Fourth, an offline method for foot sliding removal for when you have access to the whole animation and cycles to spare. And finally I'll close with a few philosophical thoughts about the whole thing.

While I provide code through the article, the full source code for this article can be found here.

  1. Solving a Leg Chain
  2. Foot Locking
  3. Contact Times
  4. Offline Foot Locking
  5. Conclusion

Solving a Leg Chain

Overview

Here is the problem set-up. We have a chain of joints for the leg of a character, posed in a particular configuration, and we want to modify the local rotations such that the existing pose is preserved as much as possible, but where the toe ends up at some desired location:

setup

To do this we're going to perform the following steps:

  1. Compute the target heel location from the target toe location
  2. Solve the two-joint IK problem to position the heel at its target
  3. Rotate the heel joint to orient the toe towards the toe target
  4. Optionally rotate the toe-end to resolve any ground collisions

Find the Heel Target

Computing the desired heel location is easy - we simply compute the vector from the toe to the heel in the existing input pose and add this to the toe target.

*targetHeel = Vector3Add(*targetToe, Vector3Subtract(
    globalTransforms[heelBoneIndex].translation, 
    globalTransforms[toeBoneIndex].translation));

Which looks like this:

setup

2. Solve for the Heel Target

Once we have the heel target we can solve for the local rotations of the hip and knee joints such that it positions the leg chain correctly. To do this I'm going to use a slightly modified version of the two bone inverse kinematics code I shared a long time ago (see this page for the derivation of the QuaternionFromScaledAngleAxis function).

static inline Quaternion QuaternionExp(Vector3 v)
{
    float halfangle = sqrtf(v.x*v.x + v.y*v.y + v.z*v.z);

    if (halfangle < 1e-4f)
    {
        return QuaternionNormalize((Quaternion){ v.x, v.y, v.z, 1.0f });
    }
    else
    {
        float c = cosf(halfangle);
        float s = sinf(halfangle) / halfangle;
        return (Quaternion){ s * v.x, s * v.y, s * v.z, c };
    }
}

static inline Quaternion QuaternionFromScaledAngleAxis(Vector3 v)
{
    return QuaternionExp(Vector3Scale(v, 0.5f));
}

static inline void TwoBoneInverseKinematics(
    Quaternion *localHip,
    Quaternion *localKnee,
    Transform globalPelvis, 
    Transform globalHip, 
    Transform globalKnee, 
    Transform globalHeel, 
    Vector3 targetHeel, 
    Vector3 sideVector,
    float maxExtension,
    float softening)
{
    // Softly clamp the target based on the distance given by maxExtension 
    
    Vector3 targetClamp = targetHeel;
    float targetLength = Vector3Distance(targetHeel, globalHip.translation);

    if (targetLength > maxExtension - softening)
    {
        // Smoothly clamp when it gets within softening distance of the maxExtension
        
        float saturation = 1.0f - expf(
            -Max(targetLength - maxExtension + softening, 0.0f) / softening);
        
        targetClamp = Vector3Add(
            globalHip.translation, 
            Vector3Scale(Vector3Subtract(targetHeel, globalHip.translation), 
                (maxExtension - softening + softening * saturation) / targetLength)); 
    }
    
    // Compute the rotation axis based on vector perpendicular to the plane 
    // rotation which is closed to the provided knee side vector
    
    Vector3 axisDwn = Vector3Normalize(
        Vector3Subtract(globalHeel.translation, globalHip.translation));
        
    Vector3 axisFwd = Vector3Normalize(Vector3CrossProduct(axisDwn, sideVector));
    Vector3 axisRot = Vector3Normalize(Vector3CrossProduct(axisDwn, axisFwd));

    Vector3 a = globalHip.translation;
    Vector3 b = globalKnee.translation;
    Vector3 c = globalHeel.translation;
    Vector3 t = targetClamp;
    
    // Compute the change in rotation angle required using the cosine rule
    
    float lab = Vector3Distance(b, a);
    float lcb = Vector3Distance(b, c);
    float lat = Vector3Distance(t, a);
    float lca = Vector3Distance(a, c);

    float acab0 = acosf(Clamp(Vector3DotProduct(
        Vector3Scale(Vector3Subtract(c, a), 1.0f / lca), 
        Vector3Scale(Vector3Subtract(b, a), 1.0f / lab)), -1.0f, +1.0f));
        
    float babc0 = acosf(Clamp(Vector3DotProduct(
        Vector3Scale(Vector3Subtract(a, b), 1.0f / lab), 
        Vector3Scale(Vector3Subtract(c, b), 1.0f / lcb)), -1.0f, +1.0f));

    float acab1 = acosf(Clamp(
        (lab * lab + lat * lat - lcb * lcb) / (2.0 * lab * lat), -1.0f, +1.0f));
        
    float babc1 = acosf(Clamp(
        (lab * lab + lcb * lcb - lat * lat) / (2.0 * lab * lcb), -1.0f, +1.0f));
        
    // Compute the three world-space rotations needed to solve the two-bone ik problem
        
    Quaternion r0 = QuaternionFromScaledAngleAxis(Vector3Scale(axisRot, acab1 - acab0));
    Quaternion r1 = QuaternionFromScaledAngleAxis(Vector3Scale(axisRot, babc1 - babc0));
    Quaternion r2 = QuaternionNormalize(QuaternionBetween(
        Vector3Subtract(globalHeel.translation, globalHip.translation), 
        Vector3Subtract(targetClamp, globalHip.translation)));
    
    // Update the local space rotations for the hip and the knee
    
    *localHip = QuaternionMultiply(QuaternionMultiply(QuaternionMultiply(
        QuaternionInvert(globalPelvis.rotation), r2), r0), globalHip.rotation);
        
    *localKnee = QuaternionMultiply(QuaternionMultiply(
        QuaternionInvert(globalHip.rotation), r1), globalKnee.rotation);
}

The two additions here are simple. First, I let the user provide a max extension length maxExtension which I softly clamp the target towards based on the softening distance. This means that the limb only ever exponentially approaches the provided maxExtension, which helps prevent hyper-extension:

Vector3 targetClamp = targetHeel;
float targetLength = Vector3Distance(targetHeel, globalHip.translation);

if (targetLength > maxExtension - softening)
{
    // Smoothly clamp when it gets within softening distance of the maxExtension
    
    float saturation = 1.0f - expf(
        -Max(targetLength - maxExtension + softening, 0.0f) / softening);
    
    targetClamp = Vector3Add(
        globalHip.translation, 
        Vector3Scale(Vector3Subtract(targetHeel, globalHip.translation), 
            (maxExtension - softening + softening * saturation) / targetLength)); 
}

Second, I use the side vector of the knee joint to find a stable axis to rotate around. This prevents the need for a pole-vector or anything like that.

Vector3 axisDwn = Vector3Normalize(
    Vector3Subtract(globalHeel.translation, globalHip.translation));
    
Vector3 axisFwd = Vector3Normalize(Vector3CrossProduct(axisDwn, sideVector));
Vector3 axisRot = Vector3Normalize(Vector3CrossProduct(axisDwn, axisFwd));

Assuming you have two buffers localTransforms and globalTransforms filled with bone local and global transforms, you can call the function like this:

// Solve Two-Bone Inverse Kinematics to place heel

Vector3 sideVector = Vector3RotateByQuaternion(
    kneeSideVector, 
    globalTransforms[kneeBoneIndex].rotation);

float maxExtension = Vector3Distance(
    globalTransforms[hipBoneIndex].translation, 
    globalTransforms[heelBoneIndex].translation);

Quaternion modifiedHip, modifiedKnee;

TwoBoneInverseKinematics(
    &modifiedHip, 
    &modifiedKnee, 
    globalTransforms[pelvisBoneIndex], 
    globalTransforms[hipBoneIndex], 
    globalTransforms[kneeBoneIndex], 
    globalTransforms[heelBoneIndex], 
    *targetHeel, 
    sideVector,
    maxExtension,
    softening);
    
localTransforms[hipBoneIndex].rotation = modifiedHip;
localTransforms[kneeBoneIndex].rotation = modifiedKnee;

Where, for the geno character, kneeSideVector is just (Vector3){ 1.0f, 0.0f, 0.0f } and a typical value for softening might be 0.005f (meters).

Once completed this positions the heel at the target, but the local rotation we've applied means that the toe joint is no longer rotated toward its target.

3. Solve for the Toe Target

To solve for the toe target we can just perform a basic look-at using QuaternionBetween (For the curious, I wrote a bit about the derivation of this function here):

static inline Quaternion QuaternionBetween(Vector3 p, Vector3 q)
{
    Vector3 c = Vector3CrossProduct(p, q);

    Quaternion o = {
        c.x,
        c.y,
        c.z,
        sqrtf(Vector3DotProduct(p, p) * Vector3DotProduct(q, q)) + 
            Vector3DotProduct(p, q),
    };
    
    return QuaternionLength(o) < 1e-8f ?
        QuaternionFromAxisAngle((Vector3){ 1.0f, 0.0f, 0.0f }, PI) :
        QuaternionNormalize(o);
}

static inline Quaternion BoneOrientTowards(
    Transform boneParentTransform,
    Transform boneTransform,
    Transform boneChildTransform,
    Vector3 target)
{
    Quaternion desiredRotation = QuaternionMultiply(QuaternionNormalize(
        QuaternionBetween(
            Vector3Subtract(boneChildTransform.translation, boneTransform.translation),
            Vector3Subtract(target, boneTransform.translation))), 
            boneTransform.rotation);

    return QuaternionMultiply(
        QuaternionInvert(boneParentTransform.rotation), desiredRotation);
}

Which can be used as follows:

ForwardKinematics(globalTransforms, localTransforms, model);

localTransforms[heelBoneIndex].rotation = BoneOrientTowards(
    globalTransforms[kneeBoneIndex],
    globalTransforms[heelBoneIndex],
    globalTransforms[toeBoneIndex],
    *targetToe);

And gives the following behaviour:

4. Handle Ground-Plane collisions

A nice addition we can do is collide the targets with the ground plane. First, we need to find the heel and toe height in the bind pose (assume globalTransforms is filled with the bind pose transforms):

float heelMinHeight = globalTransforms[leftHeelBoneIndex].translation.y;
float toeMinHeight = globalTransforms[leftToeBoneIndex].translation.y;
float toeEndMinHeight = globalTransforms[leftToeEndBoneIndex].translation.y;

Then we clamp the targets to their respective heights during the leg solve.

targetToe->y = Max(targetToe->y, toeMinHeight);
targetHeel->y = Max(targetHeel->y, heelMinHeight);

When doing this we then might want to also re-orient the toe-end since if we clamp the bone height to prevent the end of the toe from intersecting the floor, the rotation should change with it.

ForwardKinematics(globalTransforms, localTransforms, model);

*targetToeEnd = globalTransforms[toeEndBoneIndex].translation;

if (enableHeightClamp)
{
    targetToeEnd->y = Max(targetToeEnd->y, toeEndMinHeight);
}

localTransforms[toeBoneIndex].rotation = BoneOrientTowards(
    globalTransforms[heelBoneIndex],
    globalTransforms[toeBoneIndex],
    globalTransforms[toeEndBoneIndex],
    *targetToeEnd);

Which gives the following result (here visualised with some vertical offset applied to the pelvis to make it clearer).

Conclusion

Putting everything together we get:

static inline void SolveLegChain(
    Model model,
    Transform* localTransforms,
    Transform* globalTransforms,
    Vector3 target,
    Vector3 *targetHeel,
    Vector3 *targetToe,
    Vector3 *targetToeEnd,
    int pelvisBoneIndex,
    int hipBoneIndex,
    int kneeBoneIndex,
    int heelBoneIndex,
    int toeBoneIndex,
    int toeEndBoneIndex,
    bool enableHeightClamp,
    bool enableHeelLookAt,
    bool enableToeLookAt,
    float heelMinHeight,
    float toeMinHeight,
    float toeEndMinHeight,
    float softening,
    Vector3 kneeSideVector)
{
    *targetToe = target;
    
    if (enableHeightClamp)
    {
        targetToe->y = Max(targetToe->y, toeMinHeight);
    }
    
    // Find the Heel Target Location

    *targetHeel = Vector3Add(*targetToe, Vector3Subtract(
        globalTransforms[heelBoneIndex].translation, 
        globalTransforms[toeBoneIndex].translation));

    if (enableHeightClamp)
    {
        targetHeel->y = Max(targetHeel->y, heelMinHeight);
    }
    
    // Solve Two-Bone Inverse Kinematics to place heel
    
    Vector3 sideVector = Vector3RotateByQuaternion(
        kneeSideVector, 
        globalTransforms[kneeBoneIndex].rotation);
    
    float maxExtension = Vector3Distance(
        globalTransforms[hipBoneIndex].translation, 
        globalTransforms[heelBoneIndex].translation);
    
    Quaternion modifiedHip, modifiedKnee;

    TwoBoneInverseKinematics(
        &modifiedHip, 
        &modifiedKnee, 
        globalTransforms[pelvisBoneIndex], 
        globalTransforms[hipBoneIndex], 
        globalTransforms[kneeBoneIndex], 
        globalTransforms[heelBoneIndex], 
        *targetHeel, 
        sideVector,
        maxExtension,
        softening);
        
    localTransforms[hipBoneIndex].rotation = modifiedHip;
    localTransforms[kneeBoneIndex].rotation = modifiedKnee;
    
    // Orient Toe towards Target
    
    if (enableHeelLookAt)
    {
        ForwardKinematics(globalTransforms, localTransforms, model);
        
        localTransforms[heelBoneIndex].rotation = BoneOrientTowards(
            globalTransforms[kneeBoneIndex],
            globalTransforms[heelBoneIndex],
            globalTransforms[toeBoneIndex],
            *targetToe);
    }
    
    // Orient Toe-End
    
    if (enableToeLookAt)
    {          
        ForwardKinematics(globalTransforms, localTransforms, model);
        
        *targetToeEnd = globalTransforms[toeEndBoneIndex].translation;
        
        if (enableHeightClamp)
        {
            targetToeEnd->y = Max(targetToeEnd->y, toeEndMinHeight);
        }

        localTransforms[toeBoneIndex].rotation = BoneOrientTowards(
            globalTransforms[heelBoneIndex],
            globalTransforms[toeBoneIndex],
            globalTransforms[toeEndBoneIndex],
            *targetToeEnd);
    }
    
    // Recompute final global transforms
    
    ForwardKinematics(globalTransforms, localTransforms, model);
}

You'll notice that I'm being lazy here and re-computing the full set of global transforms with the ForwardKinematics function each time we update the local transforms, but obviously in reality we only need to re-compute the global transforms for the bones which we're intending to use in the next stage.

Similarly, I'm clamping the heights assuming the ground plane is at zero here but if you have dynamic terrain you're going to need to do a raycast or simlar to get the real height of the ground under-foot.

But what this function ultimately gives us is a kind of black box where we can input an existing pose of the character and a new toe target, and always get something relatively sane as output. This greatly simplifies the problem of foot locking because now all we need to do is make sure our toe target is not sliding, and we can be relatively confident our results will look good.


Foot Locking

In the previous part we learned how we can solve a leg joint chain using some inverse kinematics and other tricks to place the toe joint at some target location.

In this part I'm going to show you my little recipe for locking the toe target during contact states. The way it works is very simple. When there is no contact active, we follow the toe position in the input source animation. When there is a contact active, we follow a static position on the floor which was the toe's location when the contact first became active. And we transition between these two sources using inertialization.

There are a bunch of variables we need to keep track of the state of things in the foot locking setup. For each contact point we'll need the following variables:

typedef struct FootLockingState
{
    Vector3 position;           // Current Position
    Vector3 velocity;           // Current Velocity
    Vector3 inputPosition;      // Input Source Position
    Vector3 inputVelocity;      // Input Source Velocity
    Vector3 offsetPosition;     // Inertialization Offset
    Vector3 offsetVelocity;     // Inertialization Offset Velocity
    float time;                 // Time since Inertialization transition
    Vector3 contact;            // Contact Location
    bool locked;                // If the contact is currently locked
    
} FootLockingState;

Then, at each frame, we can write a function which takes as input this state, as well as the toe location in the input pose, and updates the toe target, as follows:

void UpdateFootLockingState(
    FootLockingState* state, 
    Vector3 inputPosition, 
    bool inputContact, 
    float contactHeight, 
    float deltaTime,
    float unlockDistance,
    float lockDistance,
    float blendTime)
{    
    // Update Input State Position and Velocity via finite difference
    
    state->inputVelocity = Vector3Scale(
        Vector3Subtract(inputPosition, state->inputPosition), 
        1.0f / Max(deltaTime, 1e-8f));
        
    state->inputPosition = inputPosition;
    
    // Update Cubic Inertialization
    
    InertializeCubicUpdate(
        &state->position, 
        &state->velocity, 
        &state->time,
        state->locked ? state->contact : state->inputPosition,
        state->locked ? Vector3Zero() : state->inputVelocity,
        state->offsetPosition, 
        state->offsetVelocity,
        deltaTime,
        blendTime);
    
    // Check the distance of the input location from the locked state
    
    float inputDistance = Vector3Distance(state->position, state->inputPosition);
    
    if (!state->locked && inputContact && inputDistance < lockDistance)
    {
        // Lock if the input wants it the current state is within locking distanced
    
        state->locked = true;
        state->contact = state->inputPosition;
        state->contact.y = contactHeight;
        
        InertializeCubicTransition(
            &state->offsetPosition, 
            &state->offsetVelocity, 
            &state->time,
            state->inputPosition,
            state->inputVelocity,
            state->contact,
            Vector3Zero(),
            blendTime);
    }
    else if (state->locked && (!inputContact || inputDistance > unlockDistance))
    {
        // Unlock if the input wants to unlock or the distance is too far
    
        state->locked = false;
        
        InertializeCubicTransition(
            &state->offsetPosition, 
            &state->offsetVelocity, 
            &state->time,
            state->contact,
            Vector3Zero(),
            state->inputPosition,
            state->inputVelocity,
            blendTime);
    }
}

where InertializeCubicUpdate and InertializeCubicTransition are cubic inertialization functions defined as follows...

void InertializeCubicUpdate(
    Vector3* position, 
    Vector3* velocity, 
    float* time,
    Vector3 inputPosition,
    Vector3 inputVelocity,
    Vector3 offsetPosition,
    Vector3 offsetVelocity,
    float deltaTime,
    float blendTime)
{
    float t = Clamp((*time + deltaTime) / Max(blendTime, 1e-8f), 0.0f, 1.0f);
    float w0 = 2.0f * t * t * t - 3.0f * t * t + 1.0f;
    float w1 = (t * t * t - 2.0f * t * t + t) * blendTime;
    float w2 = (6.0f * t * t - 6.0f * t) / Max(blendTime, 1e-8f);
    float w3 = 3.0f * t * t - 4.0f * t + 1.0f;
    
    *position = Vector3Add(inputPosition, Vector3Add(
        Vector3Scale(offsetPosition, w0), Vector3Scale(offsetVelocity, w1)));
    *velocity = Vector3Add(inputVelocity, Vector3Add(
        Vector3Scale(offsetPosition, w2), Vector3Scale(offsetVelocity, w3)));
    *time = *time + deltaTime;
}

void InertializeCubicTransition(
    Vector3* offsetPosition, 
    Vector3* offsetVelocity, 
    float* time,
    Vector3 sourcePosition,
    Vector3 sourceVelocity,
    Vector3 destinationPosition,
    Vector3 destinationVelocity,
    float blendTime)
{
    float t = Clamp(*time / Max(blendTime, 1e-8f), 0.0f, 1.0f);
    float w0 = 2.0f * t * t * t - 3.0f * t * t + 1.0f;
    float w1 = (t * t * t - 2.0f * t * t + t) * blendTime;
    float w2 = (6.0f * t * t - 6.0f * t) / Max(blendTime, 1e-8f);
    float w3 = 3.0f * t * t - 4.0f * t + 1.0f;
  
    *offsetPosition = Vector3Subtract(Vector3Add(sourcePosition, 
        Vector3Add(Vector3Scale(*offsetPosition, w0), 
                   Vector3Scale(*offsetVelocity, w1))), destinationPosition);
    *offsetVelocity = Vector3Subtract(Vector3Add(sourceVelocity, 
        Vector3Add(Vector3Scale(*offsetPosition, w2), 
                   Vector3Scale(*offsetVelocity, w3))), destinationVelocity);
    *time = 0.0f;
}

The logic may look complex but it is fairly simple. If the input animation says the foot is locked (and the distance between the current output and the new input is low enough), we record the contact location and transition to following this as input. If the input animation says the foot is unlocked (or the distance between the output and the input is too large) we switch back to tracking the input location.

With this we get a toe target which we can smoothly toggle in and out of a locked state (shown here in yellow):

All that remains is to plug this target into the solver we made in the previous part, and we can force the toe into a locked position during contact times:

UpdateFootLockingState(
    &leftLockState, 
    globalTransforms[leftToeBoneIndex].translation, 
    testContacts.leftContacts[animationFrame] > contactThreshold,
    toeMinHeight,
    deltaTime,
    unlockDistance,
    lockDistance,
    lockBlendTime);

leftTarget = leftLockState.position;

if (enableInverseKinematics)
{
    SolveLegChain(
        genoModel,
        modifyTransforms,
        globalTransforms,
        leftTarget,
        &leftTargetHeel,
        &leftTargetToe,
        &leftTargetToeEnd,
        pelvisBoneIndex,
        leftHipBoneIndex,
        leftKneeBoneIndex,
        leftHeelBoneIndex,
        leftToeBoneIndex,
        leftToeEndBoneIndex,
        enableHeightClamp,
        enableHeelLookAt,
        enableToeLookAt,
        heelMinHeight,
        toeMinHeight,
        toeEndMinHeight,
        softening,
        (Vector3){ 1.0f, 0.0f, 0.0f });  
}

Which looks like this:

But how do we get contact time labels in the input animation? Well for that I have another little recipe...


Contact Times

To use the foot locking code we wrote in the previous part, we need to know in the input animation when the toe is meant to be in contact with the floor (and so therefore should not slide).

Typically, this is something we want to add to our animation data as additional meta-data. The only gold-standard for the annotation of such contact times is to manually label them - but we can use a few simple heuristics to get us about 90% of the way there.

The general idea is to look at the global velocity magnitude of the toe joints in the source animation:

toe velocity

Here we can see quite clearly when the velocity is low and the foot is in contact with the ground, which is good news - it looks like this signal should be fairly easy to threshold to get something that approximates binary contact labels:

toe velocity threshold

This is much more useful than toe height, as people tend to not lift their feet very much during locomotion and even during contact the height can go up and down a bit. This makes toe height much more difficult to threshold.

toe height

But foot height can still be a good sanity check, and avoid labelling contacts for when the foot is stationary but held in the air - so we can combine a thresholding of both of these signals if we want.

In most high quality animation data I find that between 0.5 m/s and 0.1 m/s tends to be a pretty reasonable velocity threshold and 0.1 m is a pretty reasonable height threshold (although that of course depends on your toe positioning on the skeleton).

There are a few post-processes I like to apply to the thresholded signal afterwards.

The first is a majority vote filter - a filter which slides over a short window of frames and takes a majority vote to decide the output for the frame centered at that window. This filter helps avoid single-frame activations or deactivations of the contact (in numpy I actually use the median filter which ends up doing the same thing for binary signals), where a width of 5 frames at 60Hz is a pretty good start.

majority vote filter

Sometimes I also like to apply Gaussian smoothing. This turns our binary signal into a continuous one, which allows us to threshold our contact activations at runtime in a way that gives us some buffer time in adjusting the sensitivity and how early or late we want the contact to kick in:

gaussian smoothing

This heutristic is far from fool-proof and I still find it difficult to always get contacts labelled cleanly during running animations. Part of the problem here is that the shoe and foot deformation can be relatively large for fast runs (the foot can move up and down a fair amount), and contact times can be relatively short.

At 30Hz contacts during a run can often be only 1 or 2 frames. Which means combined with the majority vote filter even a little bit of movement can mean you miss them. So keeping animation data at 60Hz really helps a lot here.

And if you are up-sampling from a 30Hz signal, velocities computed via cubic interpolation also tend to be more easy to threshold compared to those computed via linear interpolation. So automatic contact annotation is a good example of when it really pays to look after your data carefully.

In the next part I'll discuss an alternative method for resolving foot-sliding offline.


Offline foot locking

If we have access to the whole animation clip, and we have some cycles to spare, it's possible to take an entirely different approach to removing foot sliding than the one we used at runtime using inertialization.

I've found it is possible to get better results if we formulate the foot-sliding problem as a set of constraints, and then solve those constraints using a position-based-dynamics-like method. This sounds complicated, and understanding all of the theory behind why it works can be - but the actual implementation itself is very simple.

The general idea is this: let's compute the pelvis, and toe positions for every frame, and treat all of these positions as a set of connected particles. We're going to try and make the relative position of these particles between frames, as well as the relative positions of these particles within each frame, be similar to the input animation using soft spring-like connections. Then, we'll add additional harder connections that enforce that there is zero distance between frames for toe particles that are considered in-contact. This will effectively "bunch together" all of the toe particles during contact times, and distribute the effect of this adjustment nicely over the rest of the animation.

They way we'll enforce these constraints is simply by looping over the full animation many times and, for each constraint, slightly adjusting the particles it affects to try and make the constraint better satisfied. In the end we should get out a new set of pelvis and toe targets which don't exhibit foot sliding, but which preserve the existing animation, and which we can feed into our leg solver to adjust the local rotations.

First, let's compute the left toe, right toe, and pelvis locations in the input animation. These are the arrays we are going to modify to enforce the constraints.

Vector3* pelvisLocations = RL_CALLOC(animation.frameCount, sizeof(Vector3));
Vector3* leftToeLocations = RL_CALLOC(animation.frameCount, sizeof(Vector3));
Vector3* rightToeLocations = RL_CALLOC(animation.frameCount, sizeof(Vector3));

for (int i = 0; i < animation.frameCount; i++)
{
    pelvisLocations[i] = animation.framePoses[i][pelvisBoneIndex].translation;
    leftToeLocations[i] = animation.framePoses[i][leftToeBoneIndex].translation;
    rightToeLocations[i] = animation.framePoses[i][rightToeBoneIndex].translation;
}

Next we're going to iterate over the whole animation several times, and each time enforce our constraints:

float softFactor = 0.05f;
float hardFactor = 0.9f;

int iterations = 25000;

for (int iteration = 0; iteration < iterations; iteration++)
{
    for (int i = 0; i < animation.frameCount; i++)
    {
        // TODO: Enforce constraints...
    }
}

Inside the loop, we can formulate the inter-frame constraints like this:

// Left Leg Inter-Frame Constraints

if (i > 0)
{
    // Get the toe and hip positions in the input animation
    
    Vector3 restPrevToe = animation.framePoses[i - 1][leftToeBoneIndex].translation;
    Vector3 restCurrToe = animation.framePoses[i - 0][leftToeBoneIndex].translation;
    Vector3 restPrevHip = animation.framePoses[i - 1][pelvisBoneIndex].translation;
    Vector3 restCurrHip = animation.framePoses[i - 0][pelvisBoneIndex].translation;

    // Get the modified toe and hip positions

    Vector3 consPrevToe = leftToeLocations[i - 1];
    Vector3 consCurrToe = leftToeLocations[i - 0];
    Vector3 consPrevHip = pelvisLocations[i - 1];
    Vector3 consCurrHip = pelvisLocations[i - 0];

    // If the contact is active on both previous and current frames

    if (contacts.leftContacts[i - 1] > contactThreshold && 
        contacts.leftContacts[i - 0] > contactThreshold)
    {
        // Compute the mid-point between the current contacts and set the 
        // height to the ground height
        
        Vector3 toeTarget = Vector3Lerp(consPrevToe, consCurrToe, 0.5f);
        toeTarget.y = toeMinHeight;

        // Move the locations toward the mid-point
    
        leftToeLocations[i - 1] = Vector3Lerp(consPrevToe, toeTarget, hardFactor);
        leftToeLocations[i - 0] = Vector3Lerp(consCurrToe, toeTarget, hardFactor);
    }
    else
    {
        // Find the target toe locations relative to each other and remove
        // any ground penetration
        
        Vector3 prevToeTarget = Vector3Add(consCurrToe, 
            Vector3Subtract(restPrevToe, restCurrToe));
        Vector3 currToeTarget = Vector3Add(consPrevToe, 
            Vector3Subtract(restCurrToe, restPrevToe));
        prevToeTarget.y = Max(prevToeTarget.y, toeMinHeight);
        currToeTarget.y = Max(currToeTarget.y, toeMinHeight);

        // Softly move toward their targets to bias it toward the source animation

        leftToeLocations[i - 1] = Vector3Lerp(consPrevToe, prevToeTarget, softFactor);
        leftToeLocations[i - 0] = Vector3Lerp(consCurrToe, currToeTarget, softFactor);
    }
    
    // Softly move the pelvis toward the correct location from the source animation

    pelvisLocations[i - 1] = Vector3Lerp(consPrevHip, Vector3Add(
        consCurrHip, Vector3Subtract(restPrevHip, restCurrHip)), softFactor);
    pelvisLocations[i - 0] = Vector3Lerp(consCurrHip, Vector3Add(
        consPrevHip, Vector3Subtract(restCurrHip, restPrevHip)), softFactor);
}

And then we can enforce the in-frame constraints...

// Left Leg In-Frame Constraints

// Find the distance from the hip to the toe in the source animation

Vector3 restHip = animation.framePoses[i][pelvisBoneIndex].translation;
Vector3 restToe = animation.framePoses[i][leftToeBoneIndex].translation;
float restLength = Vector3Distance(restHip, restToe);

// Find the current direction from the hip to the toe

Vector3 currHip = pelvisLocations[i];
Vector3 currToe = leftToeLocations[i];
Vector3 currDirection = Vector3Normalize(Vector3Subtract(currHip, currToe));

// Softly enforce the hip-to-toe length from the source animation

pelvisLocations[i] = Vector3Lerp(currHip, Vector3Add(
    currToe, Vector3Scale(currDirection, +restLength)), softFactor);
leftToeLocations[i] = Vector3Lerp(currToe, Vector3Add(
    currHip, Vector3Scale(currDirection, -restLength)), softFactor);

We enforce the right toe constraints inside the loop in exactly the same way. Once we are done we can pass the new pelvis and toe targets to our leg solver and free what we have allocated.

Transform* localTransforms = RL_CALLOC(animation.boneCount, sizeof(Transform));
Transform* globalTransforms = RL_CALLOC(animation.boneCount, sizeof(Transform));

for (int i = 0; i < animation.frameCount; i++)
{
    BackwardKinematics(localTransforms, animation.framePoses[i], model);
    
    assert(pelvisBoneIndex == 0);
    localTransforms[pelvisBoneIndex].translation = pelvisLocations[i];
    
    ForwardKinematics(globalTransforms, localTransforms, model);
    
    Vector3 leftTargetHeel, leftTargetToe, leftTargetToeEnd;
    
    SolveLegChain(
        model,
        localTransforms,
        globalTransforms,
        leftToeLocations[i],
        &leftTargetHeel,
        &leftTargetToe,
        &leftTargetToeEnd,
        pelvisBoneIndex,
        leftHipBoneIndex,
        leftKneeBoneIndex,
        leftHeelBoneIndex,
        leftToeBoneIndex,
        leftToeEndBoneIndex,
        enableHeightClamp,
        enableHeelLookAt,
        enableToeLookAt,
        heelMinHeight,
        toeMinHeight,
        toeEndMinHeight,
        softening,
        leftKneeSideVector);
    
    Vector3 rightTargetHeel, rightTargetToe, rightTargetToeEnd;

    SolveLegChain(
        model,
        localTransforms,
        globalTransforms,
        rightToeLocations[i],
        &rightTargetHeel,
        &rightTargetToe,
        &rightTargetToeEnd,
        pelvisBoneIndex,
        rightHipBoneIndex,
        rightKneeBoneIndex,
        rightHeelBoneIndex,
        rightToeBoneIndex,
        rightToeEndBoneIndex,
        enableHeightClamp,
        enableHeelLookAt,
        enableToeLookAt,
        heelMinHeight,
        toeMinHeight,
        toeEndMinHeight,
        softening,
        rightKneeSideVector);
    
    ForwardKinematics(animation.framePoses[i], localTransforms, model);
}

RL_FREE(localTransforms);
RL_FREE(globalTransforms);

RL_FREE(pelvisLocations);
RL_FREE(leftToeLocations);
RL_FREE(rightToeLocations);

If we adjust the softFactor and hardFactor we can adjust the balance between how strongly we enforce the foot-sliding constraint vs how much we follow the source animation. And although the more iterations we apply the better - if we set this too high it can take quite a while to finish.

Here is a test where I've scaled up the root motion by factor of 1.25 to introduce foot sliding. First, here is what it looks like with no adjustments:

And here it is with the runtime inertialization-based removal:

Finally, here is what the offline method produces:

And here is another comparison, this time scaling the root motion down by 0.75. First, with no modification:

Then, with inertialization:

And with our offline method:

Because this method has access to the whole animation it does a nice job at distributing the correction and preserving the input animation, rather than just reacting to contacts as they come in.


Conclusion

In this final part I wanted to dive a bit deeper into what foot sliding actually is and share some tips and tricks and a bit of the philosophy on where all of these ideas came from.

First, where does foot sliding come from? Well foot sliding is effectively caused by (at least) two distinct and different situations:

  • When the root of the character moves in a way which does not match the animation data being played on the rest of the character (something I've written extensively about before). In fact, in this case it is actually the whole character which is sliding, not just the feet. We just notice this sliding a lot more on the feet since they are closest to the ground - but it is important to remember that it is really the whole animation which is wrong (and should be adjusted), not just the leg movement.
  • When we interpolate, blend, or modify local joint rotations on a skeleton structure it results in movement at the end of joint chains such as the feet which is not necessarily natural or realistic.

What we've seen so far are mostly examples of the first case, with scaled root velocity to make the character move at the incorrect speed, thus creating some "sliding", but the second case is just as common.

In both of these cases, due to some kind of modification of the animation data at runtime, the velocity of the feet in the source animation data is not matching the velocity of the feet we see visually in the game. An artefact which is particular apparent during contact phases.

And this brings us neatly onto my first philosophical point and common mistake when it comes to foot sliding:

Foot sliding is about velocity

As noted earlier, a common thing I see a lot of people do is try to automatically label foot contacts using the foot height and how close it is to the ground in the animation data. They're thinking about foot sliding as a violation of the physics and friction force of the ground.

The reason this doesn't work is because it misses the key point of foot sliding - which is that ultimately we want the velocity of the foot at runtime to match the velocity of the foot in the source data - the fact that we only try to do this during contact times is really more of a practical measure - if we tried to constrain the velocity at all times we wouldn't be able to keep up with the rest of the animation.

When we do something like scale up the root motion, this makes the feet move faster in the world - and this velocity error is obvious when we look at how much it is moving compared to the ground. And in fact, it becomes more obvious the larger the velocity is, no matter if the foot is in contact or not. So really, foot sliding is correcting an error in the velocity, and trying to do it at the point which is most obvious (when the foot is in contact), but this doesn't mean the flight phase of the foot is still correct!

Thinking about things in terms of friction and physics also is a mistake that brings me onto to the second important point:

The toe is what needs locking

If you look at locomotion data you will notice that 90% of the time it's the area of the foot attached to the toe joint which is actually in contact with the floor. During athletic motion people tend to land much more front-footed than we expect, which means the times where the heel is in contact in the floor but the toe isn't are fractionally small.

Additionally, when the heel does contact without the toe, it is rarely a stable pose. On the other hand, having the toe in contact without the heel is incredibly common and we often see pivoting around the toe contact, not to mention lifting and dropping the heel.

setup

You might have some bad animation where the heel is intersecting the floor, and feel tempted to "lock" the heel during this pose - but you really shouldn't - you will get much nicer looking results the more freedom you give to the heel to move.

So when we think about solving the IK problem we should basically forget about locking the heel and focus almost entirely on constraining the toe.

Which brings my onto the third important point

Inverse Kinematics is a modification, not a replacement

When people think of two-joint inverse kinematics they often think of a process which is going to replace the pose of three joints based on some procedural rules and (for example) a pole-vector:

setup

And this is usually how things are set up when doing rigging, as you can see above.

But inverse kinematics doesn't always have to be applied in that way - instead it can be applied as a minimal modification to an existing animation where we move joints to some target location, but preserve the rotations of the bones as much as possible. And this is how it works in my little recipe for foot locking - because we want to preserve as much as possible of the existing animation including all the nuances around the heel and toe relations, and subtle twists etc.

And there is a final important thing to keep in mind which I didn't touch on much...

Avoid the dinosaur

The final common issue I see with over-applying foot locking is pulling the character hips down to give the legs more extension room and avoid hyper-extension. But unfortunately this bending of the knees creates a t-rex effect:

If you examine real animation data you will see that the leg is often extremely close to hyper-extension. This is one thing that makes foot-locking so difficult.

setup

If you pull down the hips you wont get hyper-extension but your animation will look bad. In this case I feel that it is almost always better to just let the feet slide, which is why in my recipe I limit the extension so what is in the input animation.

So my final advice is this: a little bit of sliding is a lot better than breaking the source animation just so the motion is mechanically correct. In fact, stop thinking about it in terms of sliding (because the physics and friction metaphor can lead you down the wrong paths), and start thinking about it in terms of input motion velocity preservation.

I hope all of that has been insightful and perhaps seeded a few ideas of how to resolve foot-sliding in your own animation systems. All of the source code for this article can be found here. And as always, thanks for reading!

The Daily Front Page 22 of 30
Saturday, September 12, 2026 The Daily Front No. #260912 — The 8087’s Hidden Arithmetic
article

Microcode in Intel's 8087 floating-point chip: the scale instruction

by pwg·▲ 98 points·30 comments·righto.com ↗
Floating-point arithmetic was a mess.

Microcode in Intel's 8087 floating-point chip: the scale instruction

In the 1970s, floating-point arithmetic was a mess. Computer manufacturers had a dozen incompatible arithmetic standards. Moreover, floating-point systems were designed around hardware simplicity rather than mathematical rigor, leading to problems with numerical stability. This changed when Intel introduced the 8087 floating-point coprocessor chip in 1980, designed to be as accurate as possible, even in the corner cases. The 8087 became popular because it could be installed in the IBM PC, making floating-point operations up to 100 times faster in applications ranging from spreadsheets to CAD. But more importantly, the 8087 became the floating-point standard used by most computers today.

The 8087 implemented its instructions in complex low-level code called microcode. I'm part of a group, the Opcode Collective, that is reverse-engineering this microcode, and I've recently made some progress. In this post, I examine the microcode for one of the 8087's instructions—FSCALE—and describe how this microcode works. The FSCALE (Floating-point Scale) instruction provides a quick way to scale a number by a power of two, much faster than a multiplication. I figured that FSCALE was a simple, almost trivial instruction that would be straightforward to understand and explain. Spoiler: it is not simple. FSCALE uses over 140 micro-instructions and three levels of subroutine calls to handle many special cases. But the FSCALE microcode illustrates many interesting parts of the 8087, such as the shifter, the adder, and the exponent converter, and also reveals a hidden feature of the 8087, so hopefully you will find it interesting.

To explore the microcode, I opened up an 8087 chip and created a high-resolution image with a microscope. The large microcode ROM is in the center, holding the 1648 micro-instructions that control the chip. The microcode engine on the left steps through the microcode, handling jumps and subroutine calls. The bottom half of the chip is the "datapath", the circuitry that performs floating-point calculations; it is split into a 16-bit datapath for the number's exponent and a 64-bit datapath for the number's significand (also known as the fractional part).

Die of the Intel 8087 floating-point unit chip, with main functional blocks labeled. The die is 5mm×6mm. Click for a larger image.

Die of the Intel 8087 floating-point unit chip, with main functional blocks labeled. The die is 5mm×6mm. Click for a larger image.

Zooming in on the bottom part of the chip shows the datapath circuitry; I've highlighted the relevant parts below.1 The exponent ROM holds various constants. The exponent converter is a specialized circuit that examines exponents, detects special values, and converts between exponent formats.2 The shifter is a large component; it allows a 64-bit3 value to be shifted left or right by arbitrary amounts. (I wrote about the 8087's shifter circuitry here.) The adder is the heart of the 8087's calculations; it is used in a loop for multiplication, division, and square roots. The B register holds one input to the adder, while multiple sources can provide the other input. The sum register holds the adder's output. The eight stack registers and the temporary registers hold floating-point numbers.

A close-up of the 8087's datapath, showing functional blocks that are used by FSCALE.

A close-up of the 8087's datapath, showing functional blocks that are used by FSCALE.

Details of the 8087

In this section, I'll explain some features of the 8087 that are important for the FSCALE microcode. To use the 8087, a programmer stores values in its eight internal registers, organized as a stack. Each register holds an 80-bit floating-point number. To optimize performance, each value in the register stack has an associated "tag" value, which is mostly invisible to the programmer.4 A tag labels a value as valid, special, zero, or empty. A "normal" floating-point value is tagged as valid. If the floating-point value is infinity, Not a Number (NaN), or a denormalized value, then it is tagged as special. A zero value is tagged as zero. Finally, if a register is empty (e.g., its value has been popped off the stack), the register is tagged as empty.

The 8087 also has temporary registers that it uses internally: tmpA, tmpB, and tmpC. Like the stack registers, tmpA and tmpB are 80-bit registers, along with two tag bits. However, tmpC only holds a 64-bit significand.

The 8087 supports a variety of data types: floating-point numbers of various sizes, integers, and binary-coded decimal. But internally, everything is stored as an 80-bit floating-point number called a "temporary real"; for the rest of this article, I'll only be considering temporary real values. A number has three parts: the sign bit, the 15-bit exponent, and the 64-bit significand (the fractional part), In most cases, a floating-point number is represented by sign × significand × 2exponent. The significand is a 64-bit binary number of the form 1.bbb...: a leading 1, followed by the binary point (the binary equivalent of the decimal point) and the rest of the bits.5 What makes floating-point numbers useful is that their scope covers the incredibly small to the astronomically large, thanks to the exponent, which ranges from -16382 to 16383. One important detail is that the exponent is stored with a "bias" of 16383 added to it. Thus, the stored exponent is always positive, even if the real exponent is negative.6

The 80-bit temporary real format. The triangle indicates the binary point, analogous to the decimal point. From the Intel Numerics Supplement.

The 80-bit temporary real format. The triangle indicates the binary point, analogous to the decimal point. From the Intel Numerics Supplement.

The 8087 supports several types of numbers that are represented as special cases with special exponents, as shown below. Zero and infinity have both positive and negative values. "Not a Number" (NaN) represents values that don't make sense, such as 0/0 or sqrt(-1); NaN has a large number of representations, not a single value. The 8087 also supports denormalized and unnormalized values, which are extremely small values where the significand doesn't have a leading 1.

The encoding of special values. Based on Table S-31 in the Intel Numerics Supplement, but highly simplified. The "x" bits are arbitrary, as long as they don't conflict with another type.

The encoding of special values. Based on Table S-31 in the Intel Numerics Supplement, but highly simplified. The "x" bits are arbitrary, as long as they don't conflict with another type.

The 8087 has a complicated exception system with six types of exceptions to indicate if something went wrong with an arithmetic operation. The most serious is the "invalid operation", indicating that the operation does not make sense, such as 0/0 or ∞-∞. It also includes accesses to an empty register (stack overflow or underflow) or operations on a NaN value. The 8087 also has an overflow exception if a value is too large to store, an underflow exception if a value is too small, and a divide-by-zero exception (excluding 0/0). A denormalized operand exception indicates that the result is too small to store as a normal value, but can be stored as a denormalized value. Finally, a precision exception indicates that a value cannot be represented exactly and must be rounded. (Precision exceptions are very common; even 1/10 will yield one.)

The 8087 provides fine-grain control over each exception type, specified by bits in the control register. If an exception is unmasked, the 8087 sends an interrupt to the 8086 processor, which handles the problem in software, for instance by terminating the program or logging an error. Alternatively, the exception can be masked and the 8087 will continue execution as best it can. For instance, an invalid result will be replaced by NaN, while an overflow or divide-by-zero will be replaced by infinity. A precision exception will result in rounding. The point of masked exceptions is that calculations continue, yielding an answer that is as accurate as possible; in most cases, this is what the programmer wants.

These features make the 8087 flexible and provide accuracy, but they also make the microcode much more complicated, since the combinations of special cases need to be handled appropriately.

The 8087's microcode

Executing an 8087 instruction can require hundreds of internal steps to compute the result. These steps are implemented in microcode with micro-instructions that specify each step of the algorithm. (Keep in mind the two levels of instructions: the assembly language instructions used by a programmer and the undocumented low-level micro-instructions inside the chip.) The microcode ROM holds the 1648 micro-instructions that implement the 8087's instruction set. I'm working with the Opcode Collective to reverse-engineer the micro-instructions and fully understand the microcode (link).

The 8087's micro-instructions are complicated, with many corner cases and ad hoc functions, but I'll provide a simplified overview. Each micro-instruction consists of 16 bits, as shown below. The first three bits specify the micro-instruction's type, which controls the meaning of the remaining bits. The first type is a transfer operation, which transfers data from one internal register to another. The two fields specify the source and destination. The three remaining bits are used for various special cases. Next is a shift operation, which uses the barrel shifter to shift a value left or right. The third type of micro-instruction controls the adder (which can also subtract). The miscellaneous instructions include stack pointer operations, tag modification, exceptions, and subroutine return. The far jump and far call micro-instructions perform a jump or subroutine call to a target micro-address in a fixed list. The condition field allows conditional jumps/calls/returns based on numerous conditions, while the last bit inverts the condition. A local jump is a relative jump to a nearby micro-instruction.

Structure of an 8087 micro-instruction.

Structure of an 8087 micro-instruction.

The FSCALE microcode

When the 8087 starts executing an instruction, the instruction decoder circuitry determines the starting address of the microcode corresponding to the instruction. This 11-bit address is loaded into the microcode engine, which starts executing the microcode.7 The microcode for FSCALE (shown below) starts at decimal address 748.8

The idea behind FSCALE is straightforward: if you want to scale a floating-point number by 2N (for an integer N), you add N to the number's exponent. This allows you to multiply or divide by a power of two much faster than using the full floating-point multiplication operation. However, the microcode for FSCALE is unexpectedly complicated and uses several microcode subroutines. In brief, the microcode first checks for arguments that are zero and then handles other special arguments. It converts the scale argument to an integer and adds it to the exponent. Finally, it handles any overflow or underflow.

In more detail, the microcode routine starts by moving the first argument from the top of the stack (st(0)) to the tmpA temporary register. If the argument is zero, the routine immediately returns. (Thus, scaling 0 by anything—even NaN—will give a result of 0.) Next, the second value on the stack (the second argument) is moved to the tmpB temporary register. Likewise, the code returns if this value is 0, so scaling anything by 0 leaves the value unchanged.9 Next, a constant value is selected; selecting a constant and using it are two separate micro-instructions. (The 8087 has separate ROMs for 16-bit exponent constants and 67-bit significand constants; this one is an exponent constant.) In the normal case, execution jumps to address #0763, skipping the call to subroutine SPECIAL_TMPS.

FSCALE:
#0748 st(0) -> tmpA        Input argument from top of stack
#0749 jmp #0776 if tmpA:tag ZERO Bail if 0
#0750 stackPtr++
#0751 st(0) -> tmpB        Scale argument from stack(1)
#0752 stackPtr--
#0753 jmp #0776 if tmpB:tag ZERO Bail if 0
#0754 expconst 0x403e      Const 403e: exp shift to convert to int
#0755 jmp #0763 if not tmp empty/special/div
#0756 call SPECIAL_TMPS    Special handling
#0757 jmp #0762 if flag
#0758 jmp #0761 if not tmpB:tag SPECIAL
#0759 except:invalid       Invalid exception, use NaN
#0760 NaN -> tmpA
#0761 jmp #0776 if intr
#0762 jmp #0775 if expConv[0] Return tmpA if expConv set, otherwise continue
#0763 tmpB:exp -> Breg     Normal path
#0764 tmpB:sign,exp -> expConv ExpConv will test tmpB's sign
#0765 expConst -> tmpC     Const 403e
#0766 adder: tmpC - Breg cin=1 403e-exp is amount to shift to convert tmpB to int
#0767 sumreg:frac -> shiftcount Store in shifter control
#0768 shift tmpB:frac R count byte bit Perform the shift
#0769 shift R -> Breg      Breg holds scale argument as an int
#0770 jmp #0777 if neg     Negative Breg needs separate handling
#0771 adder: tmpA:exp + Breg cin=0 Add the scale to the exponent
#0772 sumreg:frac -> expConv Put result in expConv to check
#0773 sumreg:frac -> tmpA:exp Update exponent with sum
#0774 call NONNORMAL_RESULT if not exp normal Handle overflow/underflow
#0775 tmpA -> st(0)        Save result back to stack
#0776 RNI                  Done: Run Next Instruction
#0777 adder: tmpA:exp - Breg cin=1 Subtract Breg
#0778 jmp #0772            Continue processing

Continuing at #0763, the second argument is converted from a float to an integer, which takes a few steps. For example, suppose the argument is 9, which in floating point is 1.001×23. The significand bits 1000 are "left justified", but for an integer, these bits need to be "right justified" by shifting them to the right. In general, if the exponent is n, the significand is shifted right by 63-n bits. But recall that the exponent is biased by 16383. Thus, the significand must be shifted right by 63-(exp-16383) bits, that is 0x403e-exp bits. (This explains the constant 0x403e earlier in the microcode.)

Converting a float to an int by shifting.

Converting a float to an int by shifting.

In the microcode, the subtraction takes several steps. At #0763, the exponent of the second argument is moved to the B register, one of the inputs to the adder (completely different from tmpB).10 Next, the sign and exponent are moved to the exponent converter, a circuit that, among other things, tests for overflow. Next, the constant 0x403e (selected back at #0754) is moved to the tmpC register. At #0766, the adder is activated, subtracting the exponent from the constant.11 The adder puts the result into the sum register, and this value is copied to the shift count register, which controls the shifter. This value indicates how many bits the second argument must be shifted to convert it to an integer. At #0768, the shifter is activated to shift by the desired amount, using both the bit shift part and the byte shift part. As with the adder, activating the shifter and reading the result are separate micro-instructions; the result is put into the B register.

The core part of the FSCALE instruction is finally performed at #0771, adding the second argument to the first argument's exponent. The adder is activated to add the B register value (the scale) to the exponent, and the updated value is stored in tmpA's exponent. (Except if the scale factor is negative, it is subtracted via the #0777 path.)12 The value is also sent to the exponent converter circuit, which checks the exponent for overflow or underflow; if so, subroutine NONNORMAL_RESULT is called. But in the normal case, the updated value is copied from tmpA to the top-of-stack register st(0). Finally, RNI (Run Next Instruction) indicates that the microcode routine is done and the instruction is completed. Thus, even in the straightforward case, FSCALE takes about 22 micro-instructions.

Handling empty or special arguments

What happens if an argument accesses an empty stack location (i.e. stack underflow) or is a special value (infinity, denorm, NaN)? These cases are handled by a micro-subroutine that I'll call SPECIAL_TMPS15 because it processes special values in tmpA and/or tmpB. This subroutine is a general-purpose routine, used by basic arithmetic operations, FSCALE, FTST (test), and FPREM (partial remainder).

The control flow through SPECIAL_TMPS is rather convoluted since the code must prioritize issues if, say, one argument is empty and the other is a denorm. I'll just give a brief summary; see the footnote13 for details. First, the subroutine converts any denorms to unnorms. Then it checks for access to empty stack locations, raising an exception or interrupt if so. Then it checks the two arguments again. If either is NaN, an exception or interrupt is triggered. Otherwise, it returns a status indicating the type of arguments.

Unexpectedly, if both arguments are NaN, the code compares the two NaN values and returns the larger. This behavior may seem very weird, but it's a documented feature.14 You might think that NaN is a single value, but it's actually an enormous family of values. The idea was that the programmer could use different NaN values to signal where a problem occurs. For instance, you could put a different NaN in each location of an uninitialized array, so you could tell which position was accessed. For some reason, the designers of the 8087 decided that if you perform an operation with two different NaNs, the result is the larger one. Thus, the microcode needs code that detects if both operands are NaN and computes the larger, using a subtraction for the comparison (#1518).

SPECIAL_TMPS (J5):
#1484 call SPECIAL_VAL if tmpA:tag SPECIAL Handle special values in tmpA/tmpB
#1485 xchg tmp
#1486 call SPECIAL_VAL if tmpA:tag SPECIAL Handle tmpB special
#1487 xchg tmp
#1488 1 -> flag            Flag=1 by default
#1489 jmp #1500 if not tmp empty/special/div 0 -> expConv if tmps okay
#1490 1 -> expConv
#1491 jmp #1497 if not tmpA/B empty
#1492 except:invalid       Invalid if either empty
#1493 jmp #1525 if compare instruction No NaN for comparison
#1494 jmp #1511 if intr    Return if interrupt not masked
#1495 NaN -> tmpA          NaN if interrupt masked
#1496 return
#1497 jmp #1502 if tmpA:tag SPECIAL Special cases
#1498 jmp #1505 if tmpB:tag SPECIAL
#1499 0 -> flag            Div normal path:
#1500 zero -> expConv      Return flag 0, expConv 0
#1501 return
#1502 call SPECIAL_VAL     TmpA special
#1503 jmp #1512 if not flag Jump if NaN, fallthrough if infinity
#1504 jmp #1509 if not tmpB:tag SPECIAL
#1505 xchg tmp             TmpB special
#1506 call SPECIAL_VAL
#1507 xchg tmp
#1508 jmp #1521 if not flag Jump if NaN, return if infinity
#1509 0 -> flag            Clear flag, return
#1510 return
#1511 RNI                  End instruction with interrupt
#1512 jmp #1522 if not tmpB:tag SPECIAL TmpA NaN, now check tmpB
#1513 xchg tmp
#1514 call SPECIAL_VAL     Check tmpB
#1515 xchg tmp
#1516 jmp #1522 if flag    Jump if tmpB is not NaN
#1517 except:invalid       Invalid exception
#1518 tmpB:frac -> Breg    Both args are NaN, find larger
#1519 adder: tmpA:frac - Breg cin=1
#1520 jmp #1522 if adder sign See if tmpA < tmpB
#1521 tmpB -> tmpA         Take larger
#1522 except:invalid       Invalid exception
#1523 jmp #1525 if compare instruction No interrupt for comparison instruction
#1524 jmp #1511 if intr    End instruction with interrupt
#1525 1 -> flag            Return with flag set
#1526 return               End of J5

This subroutine makes heavy use of a helper subroutine, SPECIAL_VAL,16 that processes one argument. The helper converts a denormalized argument to an unnormalized argument, raising an exception or interrupt as appropriate. It also flags an input of infinity.

The hardware for the micro-instruction that exchanges tmpA and tmpB at #1485 is interesting. Instead of physically moving the values between the two registers, the micro-instruction toggles a flip-flop that exchanges the meaning of tmpA and tmpB. That is, if the flip-flop is set, a reference to tmpA goes to tmpB and vice versa. (This is a standard trick in microprocessors; the Intel 8080's XCHG instruction exchanges the DE and HL registers in a similar way. The Z80 uses the same trick for the EX and EXX instructions to exchange the regular register set with the secondary register set.)

The Intel 8087 chip is packaged in a 40-pin DIP (dual in-line package), as are the 8080 and Z80. This photo is here as a break from all the microcode.

The Intel 8087 chip is packaged in a 40-pin DIP (dual in-line package), as are the 8080 and Z80. This photo is here as a break from all the microcode.

Handling a non-normal result

If you take a very large number and scale it larger, you can end up with overflow. If you take a very small number and scale it smaller, you can end up with a denormalized number or underflow. This will trigger an overflow, denorm, or underflow exception, and an interrupt if unmasked. Moreover, the 8087 supports four rounding modes: round to nearest valid value, round down (toward -∞), round up (toward +∞), or round (chop) toward zero. Depending on the rounding mode, an overflow can result in either ∞ or the largest possible floating-point number. Similarly, an underflow can result in either zero or the smallest possible floating-point number. And depending on the infinity mode (affine or projective), infinity can be either signed or unsigned. Thus, the FSCALE microcode needs to handle many special cases for the result.

The subroutine to handle a non-normal result in tmpA is below. One interesting micro-instruction is update overflow/underflow exceptions, which triggers an exception if appropriate. For most exceptions, a micro-instruction triggers the exception (for example, except:precision at #0346). But for the overflow and underflow exceptions, the microcode delegates the task to hardware. Specifically, the 8087's "exponent converter" circuit examines the exponent to see if an overflow or underflow exists, based on the selected floating-point precision. The micro-instruction sets the overflow and underflow flags based on these values. Thus, a complex task is performed by a single microcode instruction, thanks to the hardware support of the exponent converter.

NONNORMAL_RESULT (J16):
#0318 return if tmpA:tag ZERO Handle non-normal result
#0319 update overflow/underflow exceptions Trigger exceptions if exp conv says to
#0320 expconst 0x6000      The interrupt bias constant 0x6000
#0321 jmp #0329 if not intr
#0322 expConst -> Breg     Interrupt path
#0323 jmp #0326 if neg
#0324 adder: tmpA:exp + Breg cin=0 Add bias for underflow
#0325 jmp #0327
#0326 adder: tmpA:exp - Breg cin=1 Subtract for bias overflow
#0327 sumreg:frac -> tmpA:exp New exponent to tmpA
#0328 return               Interrupt, so done
#0329 jmp #0344 if neg     Masked exception
#0330 tmpA:exp -> Breg     Underflow
#0331 adder: 1 - Breg cin=1 Amount to shift denormal
#0332 call CREATE_DENORM   Create a denormal
#0333 adder: zero + Breg cin=0, roundmode Add zero to round
#0334 call ADJUST_PRECISION Adjust to specified precision
#0335 jmp #0340 if Sum register is zero If zero, return +/- zero as appropriate
#0336 zero -> tmpA:exp     Denorm: exponent is 0
#0337 sumreg:frac -> tmpA:frac Save denorm fraction
#0338 special -> tmpA tag  Tag denom as special
#0339 return
#0340 tmpA sign -> sign latch Return +/- zero
#0341 zero -> tmpA
#0342 sign latch -> tmpA sign
#0343 return
#0344 NaN/Inf -> tmpA:exp  Overflow: maybe return infinity
#0345 tmpA:frac -> tmpB:frac Save tmpA frac in tmpB
#0346 except:precision     Set precision exception
#0347 Inf -> tmpA:frac     Put infinity in frac
#0348 special -> tmpA tag  Mark infinity as special
#0349 return if not round chop If rounding up, return infinity
#0350 1 -> Breg            Return max float: adjust down
#0351 adder: tmpA:exp - Breg cin=1
#0352 sumreg:frac -> tmpA:exp Exp=7fff-1=7ffe
#0353 adder: zero - Breg cin=1
#0354 sumreg:frac -> tmpA:frac Frac 0-1 = ff...ff
#0355 norm -> tmpA tag     Normal value
#0356 return if tmpB:frac[63] Return max float unless unnorm
#0357 tmpB:frac -> tmpA:frac Return original tmpA frac
#0358 return

The 8087 has interesting behavior if an overflow or underflow is unmasked and an interrupt occurs. The idea is to let the interrupt handler know what the exponent should have been. However, the proper value can't be used since it is too big or too small to fit in the exponent field (which is why the exception occurred). The solution is to add or subtract the constant 0x6000, resulting in an exponent that fits. The interrupt handler can subtract or add this constant to get the correct exponent. Lines #0322 to 0328 perform this addition or subtraction.

For a masked underflow, a denorm value is created by the subroutine CREATE_DENORM. The value is rounded to the specified precision by ADJUST_PRECISION. Finally, if the value is too small for a denorm, the value +0 or -0 is returned as appropriate.

For a masked overflow, the 8087 either returns Infinity or the largest-possible float, depending on the specified rounding mode. Infinity is represented by an exponent of all 1s, and a significand of 1000...; these values are loaded directly onto the bus by transistors. The maximum float, however, is computed: 1 is subtracted from the infinity exponent, and 1 is subtracted from a zero significand.

Helper subroutine: creating a denormal

One controversial feature of the 8087 is denormals, numbers that are smaller than "regular" floats. Recall that floating-point numbers have a significand with the first bit set to 1. But what happens if you hit the smallest possible exponent and want an even smaller number? The 8087 lets you break the rule that the significand starts with 1, producing smaller numbers known as denormalized numbers or denorms. Denorms significantly extend the range, providing numbers up to a factor of 2³ smaller. However, denorms don't have as much precision since the upper bits are "wasted". Moreover, calculations with denorms can be substantially slower because special handling is required.

Example of a normal number, reduced by a factor of 8, resulting in a denormal.

Example of a normal number, reduced by a factor of 8, resulting in a denormal.

The diagram above shows a normal number with the minimum possible exponent (-16382, which is 1 after biasing). Dividing the number by 8 (or scaling by -3) creates a denorm since the exponent can't be reduced any further. Instead, the significand is shifted 3 bits to the right. The exponent is replaced with the special value 0, indicating that the number is a denorm.

In the 8087, denorms are created by a microcode subroutine that I'll call CREATE_DENORM; it is used by many arithmetic operations, not just FSCALE. This subroutine takes a normal number and a shift amount. By shifting the normal number (as in the example above), it creates a denormalized number. The microcode (below) uses the exponent converter to check if the shift is 64 or more. If so, there will be nothing left after the shift, so zero is returned. Otherwise, the value is shifted to the right and the denorm is stored in the B register.

CREATE_DENORM (J20):
#0522 sumreg:frac -> expConv Create denorm
#0523 sumreg:frac -> shiftcount Number of bits to shift
#0524 jmp #0528 if exponent[6:14] == 0 Jump if < 64
#0525 zero -> Breg         No bits left, use zero
#0526 shift tmpA:frac L 0 bytes, 0 bits Run through shifter?
#0527 jmp #0532
#0528 shift tmpA:frac R count byte bit Shift right by the specified amount
#0529 shift R -> Breg      Result to Breg
#0530 shift tmpA:frac L ~count byte bit Now shift back for sticky test
#0531 NOP                  Wait for shifter
#0532 rounding(h) -> Breg[grs] Store the three rounding bits in the Breg
#0533 return

But why is the value then shifted to the left (#0530)? The purpose of this is to get the rounding bits. One of the principles of the 8087 is to get rounding correct, which is a lot harder than it seems. In order to decide how to round up a number, you need to keep track of an impossibly large number of bits. For instance, if you calculate 1 + 0 and round up, you get 1. But if you calculate, say, 1 + 2-10000 and round up, you get a float a bit higher than 1. The problem is how do you distinguish the two sums before rounding, without storing thousands of bits?

The trick is that the 8087 keeps three bits for use in rounding: the "guard" bit, the "round" bit, and the "sticky" bit. If you consider a "tail" of bits to the right of the significand, the guard bit is the most significant bit of the tail, followed by the round bit. The sticky bit is special: it is the OR of all the remaining bits in the tail, indicating if any of them are 1. Thus, 1 + 2-10000 has the sticky bit set, while 1 + 0 does not, so the two values can be rounded up differently. To generate the sticky bit, the 8087 uses a very large 64-bit NOR gate that tests the tail bits in parallel.

A diagram showing how the guard, round, and sticky bits are computed from a right shift. The numbers in this example are different from the previous example.

A diagram showing how the guard, round, and sticky bits are computed from a right shift. The numbers in this example are different from the previous example.

When a number is shifted to the right (e.g., when creating a denormal), bits are lost off the right. To generate the rounding bits, the value is shifted to the left, keeping all the tail bits that will eventually be discarded, and discarding the bits that will be in the final significand. The top two bits go into the guard and round bits, while the remaining bits are ORed together to generate the sticky bit from the rest.17 The diagram above is an example of this process. Suppose the value is being shifted to the right by 4 bits. The tail bits abcd (or at least d) will get lost in the shift. The rounding bits are computed by shifting the original significand to the right by 59 bits (the complement of 4). Bit 62 (a) becomes the new guard bit, bit 61 (b) becomes the new round bit, and the OR of the remaining 64 bits becomes the new sticky bit. (Note that the old guard, round, and sticky bits get ORed in too, so they aren't lost.) Merging the significand from the first shift with the rounding bits from the second shift produces the desired result.

Helper subroutine: adjusting precision

Although the 8087 supports three lengths of floats, it performs all calculations with 80-bit "temporary reals". At the end of an instruction, it converts the result to the desired length. (As a consequence, most instructions aren't any faster if you use a shorter float.) A microcode subroutine, which I call ADJUST_PRECISION, converts the result to the precision that is specified in the 8087's control word, using the specified rounding mode. This subroutine is used by most of the arithmetic instructions.

The 8087 supports three types of real numbers. From the Intel Numerics Supplement.

The 8087 supports three types of real numbers. From the Intel Numerics Supplement.

The first code path handles temporary reals (which have 64 bits of precision). The control word specifies one of four rounding modes. However, there are only two actions that can be taken for a particular significand: either round down (chop) or round up (chop and increment by 1). This decision is made by complicated logic circuits that examine the rounding bits, the rounding mode, and the sign to determine whether to round up or down. This simplifies the microcode but makes the hardware more complicated. The microcode performs a conditional return, returning if the significand doesn't need to be rounded up. Otherwise, the microcode increments the significand by adding 0 with a carry-in. It then checks for overflow, in which case it replaces the value with Infinity and sets a special flag.18

ADJUST_PRECISION (J11):
#0299 jmp #0306 if not precision64
#0300 return if not round up, update CC1 Update condition code, maybe return
#0301 adder: sumreg:frac + 0 cin=1 Add 1 to round up
#0302 return if not sumreg[64]
#0303 Inf -> sumreg:frac,sign Return infinity if overflow
#0304 2count++             Set special flag
#0305 return
#0306 23/52 -> shiftcount Short or long real: get appropriate shift
#0307 shift sumreg:frac,rnd L count byte bit sticky Shift to generate rounding bits
#0308 NOP                  Wait for shifter to complete
#0309 rounding(H) -> sumreg[grs] Store rounding bits
#0310 shift sumreg:frac R ~count byte bit Shift right to drop excess bits
#0311 shift R -> sumreg:frac
#0312 jmp #0314 if not round up, update CC1 Update condition code
#0313 adder: sumreg:frac + 0 cin=1 Round up if appropriate
#0314 shift sumreg:frac L ~count byte bit Shift left to realign
#0315 shift L -> sumreg:frac,sign
#0316 return if not sumreg[64] Return if not overflow
#0317 jmp #0303            Return infinity

The code is more complicated when returning a smaller precision (short real or long real), since the significand must be shortened. First, the code at #0306 loads the shifter with either 23 or 52, depending on the precision specified in the control word, and then shifts the value left. This produces the rounding bits as in the previous section. Next, the value is shifted to the right, shortening it to the desired length. As before, the significand is incremented or not, depending on whether it should be rounded up or not. Finally, the value is shifted back to the left, so the most significant bit of the significand is on the left. As before, if rounding up caused an overflow, infinity is returned.

One bizarre feature is that a jump with the "round up" conditional also has a side effect of updating the 8087's programmer-visible condition code register (CC1), indicating if the result was rounded up or down. That is, the 8087 has extra circuitry to detect this specific condition and load the value into the condition code latch. Strangely, the 8087 documentation doesn't describe this condition code action; Intel didn't document it until the 387SX floating-point chip in 1987.19

Conclusions

Floating-point has a long history before the 8087. For instance, the IBM System/360 mainframes (1964) supported 32-bit and 64-bit floating-point numbers. In 1977, AMD introduced the Am9511 floating-point chip, supporting 16- and 32-bit floating-point numbers, along with transcendental functions. What made the 8087 revolutionary is that it was carefully designed to be as mathematically accurate as possible, largely thanks to numerical expert William Kahan. (The 8087 led to the IEEE 754 Standard, now used by almost every computer and ending the anarchy of incompatible floating-point standards.)

The 8087 ended up extraordinarily complicated with three different sizes of floating-point numbers, four sizes of integers, four rounding modes, infinity modes, a collection of exceptions that could be masked or unmasked, denormalized and unnormalized numbers, signed and unsigned infinities, signed zeros, and a whole family of Not-a-Numbers. These features combine, yielding many corner cases. The 8087 deals with this complexity both through specialized circuits and through tangled microcode.

How complicated is the 8087? For users who didn't have an 8087 chip, Intel sold an 8087 Support Library that exactly emulated the 8087's instructions (but much slower). The emulator took 16K bytes of 8086 code, which was a lot when a full BASIC interpreter could fit in 8K. Another way of looking at this is that the hardware of the 8087 drastically reduced the amount of software required: the 8087 itself used 3.3K of microcode, compared to the 16K for the emulator in 8086 code.

I plan to continue reverse-engineering the 8087 microcode; for updates, follow me on Bluesky (@righto.com), Mastodon (@[email protected]), or RSS. I've been working on this with the members of the "Opcode Collective", especially Smartest Blob and Gloriouscow, who converted the ROM images to microcode data and extensively analyzed the contents. See the 8087 repository on GitHub for more.

The Daily Front Page 23 of 30
Saturday, September 12, 2026 The Daily Front No. #260912 — Fold Along the Seam
article

Designing for Dual Screen and Foldable Devices with CSS (2023)

by mooreds·▲ 60 points·20 comments·blog.stephaniestimac.com ↗
Let’s ignore the naysayers who vehemently don’t want to design for a new form factor.

With the announcement of Google's Pixel Fold, I thought its an apt time to revisit what technologies are available for designing and developing experiences that adapt to a foldable device.

Let's ignore the naysayers who vehemently don't want to design for a new form factor, while also ignoring the price of Goggle's first generation device and just focus on what's possible right now.

CSS Media Queries for Viewport Segments

There are two media queries for targeting your foldable device.

@media (horizontal-viewport-segments: <count>) { }
@media (vertical-viewport-segments: <count>) { }

The horizontal-viewport-segments query targets the device when it is being held like a book and the two screens are side-by-side.

And vertical-viewport-segments targets the device when the two screens are stacked on top of each other and fold like a laptop. I've found this mode useful when using my Surface Duo with PowerPoint. I can access my speaker notes on the bottom screen and my presentation is on the top screen. Pretty snazzy.

CSS Environment Variables

In order to get the geometry of each display, a handful of environment variables have been created for dual screen devices.

env(viewport-segment-width <x> <y>);
env(viewport-segment-height <x> <y>);
env(viewport-segment-top <x> <y>);
env(viewport-segment-left <x> <y>);
env(viewport-segment-bottom <x> <y>);
env(viewport-segment-right <x> <y>);

The x and y positions represent the two-dimensional grid created by hardware features that separate each viewport segment, with the coordinates 0,0 starting at the top-left segment.

alt: The environment variables laid out on each display screen with the integers for each display

Environment variables are cool because they will calculate dimensions by the device and make the adjustments for you meaning you don't have to create designs for every unique device out there. There are a number of foldable devices already and they all have different screen and hinge dimensions.

This set of environment variables also lets you place and position content across the two screens.

In my demo, I've built a recipe page that changes to a two column layout when on a dual screen device. I'm using CSS Grid and I can use the environment variables as grid column or row values.

@media (horizontal-viewport-segments: 2) { 

    .container {
        grid-template-columns: env(viewport-segment-width 0 0), 1fr;
    }

}

This line of code means that the first grid column will take up the entirety of left display when the device is held in the vertical position (like a book), and the second grid column will take up the remaining width of the display space available, which is the display on the right.

If you wanted to be more explicit, you could write:

@media (horizontal-viewport-segments: 2) { 

    .container {
        grid-template-columns: env(viewport-segment-width 0 0) env(viewport-segment-width: 1 0);
    }

}

Again, this means that each column takes up a display. Pretty sweet.

When should you include a dual screen mode in your product?

Back to the naysayers. When I first started talking about dual screen design, so many people reacted negatively. And honestly, no one is forcing you to create a website or app that adapts when it's spanned across two displays.

With the Surface Duo, you can utilize the browser in one display pane and it displays like a normal responsive website. Not every site or app does need a dual screen mode, but its best to find out what kind of devices your customers are using and whether or not a dual screen mode is worth the investment. As always, use data to guide your decisions.

The forecast for foldables shipped in 2023 is expected to be above 21 million devices (source). Compared to the mobile device market, its just a sliver, but that's still millions of foldables being used out there.

The other reason I'd create an app or website that adapts for a foldable device? If I wanted to stand out.

The Pixel Fold reinforces that foldables are here to stay, so why not delight your users?

Browser Support

The media queries and environment variables are currently supported in Microsoft Edge Stable.

Unfortunately, Chrome Stable doesn't appear to have the same support yet which seems like a huge miss for the Pixel Fold & Chrome teams. You can enable the experimental web platform flag under chrome://flags and test in Chrome with the Surface Duo emulator.

Closing

This article is just an overview and reminder that these capabilities exist for foldable web and app experiences. The space is ripe for innovation in design and layout on the web (and for native apps) as new foldable devices ship, but without the browser and OS support, companies won't be able to get devs and designers to take building foldable experiences seriously.

If your'e still interested exploring what's possible with dual screen layouts, I've a much more in depth article over on Smashing Magazine that walks through the recipe demo I made for the web. So download Microsoft Edge and happy building.

The Daily Front Page 24 of 30
Saturday, September 12, 2026 The Daily Front No. #260912 — The Mathematical Frontier
article

Navier-Stokes Announcement

by rvz·▲ 320 points·262 comments·claymath.org ↗

Navier-Stokes Announcement

C. Fukushima and J. Westerweel, Technical University of Delft, The Netherlands

The Millennium Prize Problems were unveiled by the Clay Mathematics Institute (CMI) at a meeting in Paris in 2000 to celebrate “The Universality of Mathematical Thought”.

These problems encapsulate some of the most difficult challenges that mathematicians were grappling with at the turn of the millennium. CMI’s purpose in articulating these problems and attaching a $1M prize to each of them was threefold: to elevate in the consciousness of the general public the fact that in mathematics the frontier is open, close at hand, and abounds with important unsolved problems; to emphasize the abiding value of working towards a solution of the deepest, most difficult problems; and to recognize achievements in mathematics of historic magnitude.

Each of the Millennium Prize Problems concerns a classical question of compelling origin that had defied all attempts to resolve it over many years. These are not arbitrary puzzles akin to fiendish crosswords. Rather, they are fundamental challenges that mark the frontier of human knowledge and challenge us to develop new structures and methods. They provide foci for the continuing struggle, across generations and cultures, to deepen our human understanding of mathematics and the universe that it describes. The deep innovations that are required to make significant progress on these problems open new vistas of possibility that typically reach far beyond the domain of the problem.

The Millennium Problem concerning the motion of fluids has all these attributes. The problem asks about the existence and smoothness of Navier-Stokes solutions in 3-dimensional Euclidean space. In recent years there has been an increasing sense of anticipation as breakthroughs in the surrounding field (some recognised by the Clay Research Award) have raised hopes that the Navier-Stokes problem might soon be resolved. The increasing ability of new technologies to accelerate mathematical research has heightened this sense of anticipation.

Today, CMI shares in the excitement of the global mathematical community as we contemplate the announcement that the Navier-Stokes problem has apparently been settled. We hope to see waves of new human understanding unleashed as the innovations behind this work are analysed and interrogated.

The Clay Mathematics Institute is dedicated to furthering the beauty, power and universality of mathematical thought. Curating the Millennium Prize Problems is one of the most important ways in which it pursues this goal. The rules governing the prizes describe the process for evaluating what has been achieved and for assigning credit. The process is deliberately unhurried, but we will provide updates.

The Clay Mathematics Institute

September 11, 2026

Image: Wikimedia Commons, C. Fukushima and J. Westerweel, Technical University of Delft, The Netherlands

The Daily Front Page 25 of 30
Saturday, September 12, 2026 The Daily Front No. #260912 — Small Questions, Big Appetites
article

Will There Be a 7G?

by Betelbuddy·▲ 87 points·152 comments·arxiv.org ↗

The transition from 5G to 6G is becoming concrete: the ITU-R IMT-2030 framework has established the high-level vision and capability set for 6G, while 3GPP Release 21 has defined the path toward the first 6G specifications. This raises a deliberately provocative question for the research and standards communities: will there be a 7G, and if so, what would justify it? This paper argues that 7G should not be treated as an inevitable numbering exercise or as a catalogue of more ambitious radio targets. Instead, its justification should depend on whether post-6G systems introduce needs or coordination problems that cannot be met by 6G/6G-Advanced, Wi-Fi, NTN, private cellular, neutral-host deployments, edge-cloud platforms, or complementary wireless and software-based systems. To support this assessment, the paper develops a readiness framework covering demand-led need, system-level discontinuity, coordination value, sustainability and circularity, trust, and geopolitical viability. It then applies the framework to candidate 7G discontinuities, including agentic network operation, RF-native computing, quantum-enabled interworking, policy-aware spectrum governance, grid-interactive infrastructure, outcome-assured services, and regionalized standards. The contribution is not a prediction of a fixed 7G architecture, but a structured basis for deciding whether 7G should become a distinct mobile generation, an extension of 6G evolution, or a broader post-6G infrastructure fabric.

The Daily Front Page 26 of 30
Saturday, September 12, 2026 The Daily Front No. #260912 — Small Questions, Big Appetites
article

Eating Fruit Skins

by surprisetalk·▲ 88 points·188 comments·pgadey.ca ↗

This summer, I learned that one can eat a whole apple. Every part of the apple is edible1. Before learning this, I would just eat the flesh and throw away the core. The core seemed hard and jagged and inedible. Turns out, you can eat the whole thing. I suppose I’ve always known that deer eat the whole apple. Why shouldn’t people eat the whole apple too2?

I tried it out a few times but initially hesitated to eat the stem. Once, I even ate everything but the stem and carried it around for a while. After a while I started to feel silly, carrying a little stem, and decided to eat the stem too. Indeed, the stem is edible too.

A few years back, I learned that one should eat the stems of strawberries. In my community, they’re considered medicine. A vital part of the fruit. When I started to ask around about this sort of thing, I learned that kiwi skins are also edible.


  1. Alex rightly points out that apple seeds do contain cyanide. Turns out, it’s not a huge risk for normal apple consumption. ↩︎
  2. This argument is obviously nonsense. There are probably a tonne of things that deer eat that people should not eat. I’m put in mind of a time when a friend said that horse chestnuts should be edible since horses presumably eat them. I thought about it and replied, “Maybe they’re used to induce vomiting in horses. Who knows?” It turns out, after a bit of Googling, that horse chestnuts make horses and people sick but deer can eat them. ↩︎
The Daily Front Page 27 of 30
Saturday, September 12, 2026 The Daily Front No. #260912 — Small Questions, Big Appetites
article

Apple iPod Engraver (2019)

by NaOH·▲ 178 points·40 comments·dunstanorchard.com ↗

I spent 2004–2006 working at Apple as a UI engineer for their online store. Part of my job was to prototype concepts that would add some interactive sparkle to the site.

One such project improved the “Personalize your iPod” page, where customers could submit two lines of text to be engraved on the back of the iPod they were ordering.

As originally designed the page offered no interaction beyond a plain form. I added a rotatable iPod, a live “engraved” preview of the customer’s text, and a highlight of any change in shipping times.

I animated the iPod by cycling through a series of JPEGs using JavaScript. The engraving was created using imagemagick, which took the user’s text and returned an “engraved” image to be overlaid on the rear of the iPod. The shipping highlight was achieved by switching CSS classes to change the background color, in the style of a classic yellow fade.

There’s an old working demo if you’d like to try it yourself.

We obviously have much better ways of doing such things today, but in 2005, with limited browser technology, those solutions seemed a little bit magical.

The Daily Front Page 28 of 30
Saturday, September 12, 2026 The Daily Front No. #260912 — Also on the Front Page
The Daily Front Page 29 of 30
Saturday, September 12, 2026 The Daily Front No. #260912 — Colophon

That's the Front for Today

Issue No. #260912 — Saturday, September 12, 2026 — went to press 2026-09-13 at 06:04 UTC.

About This Magazine

The Daily Front is a daily digital magazine assembled from the stories that reached the front page of Hacker News on Saturday, September 12, 2026. Headlines, points, and comment counts are recorded as they stood at press time. All articles remain the property of their original authors — every piece links back to its source and its discussion thread.

How It Was Made

Fetched, cleaned, and typeset by an automated pipeline. An editor model laid out the pages and chose the highlights; a second read a handful of the day's stories and briefed the cover illustrator — 33 model calls and 364k tokens in total. Set in Jacquard 12, Playfair Display, Source Serif 4, and IBM Plex Mono, all served via Google Fonts under the SIL Open Font License.

The Cover

The cover illustration was commissioned with this prompt:

At a quiet electronics workbench, an engineer pauses over a dismantled laptop processor, tracing the exposed neural-engine pathways with a fine probe while a small brass metronome ticks beside it. The chip rests on a heavy balance scale, deliberately weighted against a glowing accelerator module, as if speed is being restrained. Nearby, a strand of tiny metal links trails from the circuit board and vanishes one link at a time into the darkness, while a sealed hardware token and cooling fan sit untouched.

Brutalist 3D editorial render in matte clay, preserving the quiet electronics workbench: an engineer paused over a dismantled laptop processor, tracing exposed neural-engine pathways with a fine probe; a small brass metronome ticking beside it; the chip on a heavy balance scale deliberately weighted against a glowing accelerator module, restraining speed; tiny metal links trailing from the circuit board and disappearing one by one into darkness; a sealed hardware token and cooling fan untouched. Use crisp ambient occlusion, hard-edged simplified forms, and an asymmetric gallery-light composition with a deliberate palette of charcoal, bone white, oxidized brass, electric cyan, and restrained safety orange; isolate the weighted scale and processor in the primary light, with the vanishing links receding into a deep charcoal void.

Absolutely no text, letters, numbers, readable symbols, or logos anywhere in the image.

Production Ledger

StageModelCallsTokens InTokens Out
extractgpt-5.6-luna 29 231,354 104,083
layoutgpt-5.6-terra 1 18,448 2,160
covergpt-5.6-luna 2 1,704 381
covergpt-image-2 1 279 5,488

The Publisher

Published by Johnny.

Support the Press

If The Daily Front brightens your morning, consider supporting its publisher.

Credits & Contact

All content — articles, posts, comments, and the images within them — belongs to its original authors and is reproduced here to point readers back to the source. Full credit goes to those creators; every item links to its original and its Hacker News discussion.

If you are an author and would like your content removed from an issue, write to hi@johnnys.page and it will be taken down.

Feedback is always welcome at the same address: hi@johnnys.page.

Credit where credit is due.

Every page of this issue began as someone else's work — these are the original sources, linked in full.

  1. We must pace the frontier by apsec112 — darioamodei.com·HN discussion ↗
  2. Retrospectively Reverse-Engineering Apple's Neural Engine by zdw — eiln.github.io·HN discussion ↗
  3. A Mathematical Framework for Transformer Circuits (2021) by Bluestein — transformer-circuits.pub·HN discussion ↗
  4. Real-SWE: Benchmarking AI models on private, real-world, enterprise codebases by theanonymousone — withspecific.com·HN discussion ↗
  5. LRU is harder to beat than the KV-cache papers suggest by gauravapiscean — github.com·HN discussion ↗
  6. Fuck it, make it anyway by JayOtter — joelotter.com·HN discussion ↗
  7. The worst spam emails: iLands AI agent hustle by ColinWright — tedium.co·HN discussion ↗
  8. google.com/goto: Google's anti-scraping update by 1e1a — autom.dev·HN discussion ↗
  9. LG denies TV spying claims, says tracking and snooping concerns 'not true' by datakan — tomshardware.com·HN discussion ↗
  10. Android NAT-T keepalive offload bypasses VPN lockdown by mhitza — supuk.ch·HN discussion ↗
  11. How Trail of Bits helps verify the integrity of Signal chats by dgroshev — blog.trailofbits.com·HN discussion ↗
  12. I fixed a tractor using John Deere's self-repair service. Farmers aren't sold by sbulaev — wired.com·HN discussion ↗
  13. Make your first edit to OpenStreetMap by juliantigler — high5apps.github.io·HN discussion ↗
  14. Show HN: Bodily Oddities by vesterde — vester.si·HN discussion ↗
  15. Great Lakes sturgeon may be 400 years old:Scientists rethinking how to save them by bookofjoe — cbc.ca·HN discussion ↗
  16. Performance of WebAssembly Runtimes in 2026 by fagnerbrack — 00f.net·HN discussion ↗
  17. I made a build visualizer to understand Bun's compile times by lalitmaganti — lalitm.com·HN discussion ↗
  18. Stabilizing Rust's Never Type by cjd8 — lwn.net·HN discussion ↗
  19. Testing Race Conditions by alpaylan — projectzero.google·HN discussion ↗
  20. Inverse Kinematics and Foot Locking by airhangerf15 — theorangeduck.com·HN discussion ↗
  21. Microcode in Intel's 8087 floating-point chip: the scale instruction by pwg — righto.com·HN discussion ↗
  22. Designing for Dual Screen and Foldable Devices with CSS (2023) by mooreds — blog.stephaniestimac.com·HN discussion ↗
  23. Navier-Stokes Announcement by rvz — claymath.org·HN discussion ↗
  24. Will There Be a 7G? by Betelbuddy — arxiv.org·HN discussion ↗
  25. Eating Fruit Skins by surprisetalk — pgadey.ca·HN discussion ↗
  26. Apple iPod Engraver (2019) by NaOH — dunstanorchard.com·HN discussion ↗
  27. IKEA made a mod for Skyrim [video] by kegenaar — youtube.com·HN discussion ↗
  28. Nvidia is the central bank of AI by tolugenius — economist.com·HN discussion ↗
  29. Usenet rewind archive search engine by cstadler1869 — usenet-rewind.com·HN discussion ↗
  30. Linux Zoom client proactively reading everything written to X11 clipboard by encyclopedism — hachyderm.io·HN discussion ↗

Browse all issues in the archive →