Cover illustration

TheDaily Front

Issue No. #260801 Saturday, August 1 2026 #260801 — SATURDAY, AUGUST 1, 2026
Feeds, proofs, noodles, and kernels—an August broadsheet for the stubbornly curious.
Saturday, August 1, 2026 The Daily Front No. #260801 — Contents
30stories
6,669points
3,340comments
261kllm tokens
Assembled with 31 model calls — 174,605 tokens read, 86,401 written.

Highlights

How Google helped destroy adoption of RSS feeds (2023)

A sharp obituary-not-obituary for RSS, and a reminder that the open web rarely dies all at once.

Software for One

The case for tiny, personal software returns with new force in the age of fast code generation.

Ten advances in mathematics and theoretical computer science

Mathematics, AI authorship, and proof verification collide in one of the day’s most argued threads.

A Surveillance Treaty in Disguise: Canada Signs UN Cybercrime Convention

Canada’s signature on a UN cybercrime convention raises alarms about privacy and cross-border surveillance.

NetBSD 11.0

NetBSD 11.0 arrives with the calm dignity of a venerable system still doing the work.

From the Editor

Today’s edition finds the old internet rattling its chains: feeds, model weights, home-cooked software, and privacy rights all ask who gets to keep the keys. Meanwhile, the machines promise proofs, portfolios, pictures, and code—provided we remember to inspect the plumbing.

  1. How Google helped destroy adoption of RSS feeds (2023)3
  2. Software for One4
  3. How to Exist5
  4. Ten advances in mathematics and theoretical computer science6
  5. Postmortem for Kernel Soundness Bug #145767
  6. Explorative modeling: Train on the best of K guesses8
  7. Run Kimi K3 using 29 GB of RAM at 0.50 tok/s9
  8. Attention Decode on AMD MI450 GPUs: A Gluon Kernel Optimization Guide10
  9. Twenty-five years ago it was cryptography, today it's model weights11
  10. A Surveillance Treaty in Disguise: Canada Signs UN Cybercrime Convention12
  11. Golang proposal: container/: generic collection types13
  12. June in Servo: real world compat, media queries, SharedWorker, and more14
  13. NetBSD 11.015
  14. RipGrep musl binaries occasionally segfault during very-large searches16
  15. The Art of 64-bit Assembly17
  16. Linux on ESP3218
  17. But can your calculator run Linux?19
  18. The Absurdity of Albert Camus20
  19. The tiny holdout building in the middle of Macy’s is back in view21
  20. Manual: •.,:;…!?·22
  21. Diátaxis23
  22. Glyphs 4 – the leading Mac font editor24
  23. The development pipeline is a production system25
  24. AI financial advice is surprisingly good, especially if you ask right questions26
  25. Flint: A Visualization Language for the AI Era27
  26. Cursor removed cost information from the usage page and CSV export27
  27. Solid Queue 1.6.0 now supports fiber workers27
  28. Long Range Wi-Fi – Pushing 2.4 GHz Wi-Fi to the limits (2019)28
  29. RamenHaus29
  30. Kenji/Serious Eats – 30-Min Pressure Cooker Pho Ga29
The Daily Front Page 2 of 30
Saturday, August 1, 2026 The Daily Front No. #260801 — Lead: The Feed That Wouldn’t Die
article

How Google helped destroy adoption of RSS feeds (2023)

by pudgywalsh·▲ 518 points·177 comments·openrss.org ↗
Although RSS feeds are alive and still heavily used today, their level of adoption has suffered.

Although RSS feeds are alive and still heavily used today, their level of adoption has suffered because of how difficult a handful of popular technology companies have made it to use them. Google, especially, has relied on the open web RSS protocol to gain so much market share and influence, but continues to engage in behavior that exploits the open web at the expense of its users. As a result, Google has single-handedly contributed to the reason many users who once relied on RSS feeds have stopped using them.

Below, we dive deeper into Google's track record of what appears to be an Embrace, Extend, and Extinguish model. The company has continuously built and extended their products around the free and open RSS protocol to gain user trust, only to then remove RSS support once they've locked users in, and ignore any complaints or requests to restore it. Not only is it a blatant disregard for RSS and a huge disappointment to those who use it, it poses one of the biggest threats to the freedom and openness of the internet.

Google removes RSS button from Chrome browser

Early versions of Chromium (on which the Google Chrome browser is based) once came with RSS integration baked in. The browser had a built-in RSS button that would display in the browser location bar when any website you're on had an RSS feed available. Clicking the button would then take you to the RSS feed for that web page, allowing a user to easily subscribe.

A screenshot of the toolbar in Google Chromium showing an orange RSS icon to subscribe to a website's RSS feed

RSS subscribe icon in the toolbar of Google Chromium

Then, the RSS button disappeared without notice, and no reason was given for its removal.

Google acquires FeedBurner and limits RSS

In 2007, Google acquired FeedBurner, an RSS feed service that allows website owners to monetize their RSS feeds. FeedBurner works by replacing ordinary RSS feeds with private ones that only Google owns. The feeds are modified to include advertisements, affiliate links, and other tracking mechanisms like read counts, click-through rates and subscriber counts. The feeds are then used by FeedBurner users to track their readers and monetize from their behavior.

A screenshot of Google's FeedBurner dashboard

Google FeedBurner dashboard

After acquiring FeedBurner, Google shut down FeedBurner APIs in October 2012, which prevented developers from creating third-party RSS integrations to the service. Then, in July 2022, Google drastically changed FeedBurner's infrastructure and operation model by removing most of the FeedBurner services that its RSS users depended on, which included email subscriptions. This left many people with non-working RSS feed URLs in their subscription emails with no way to fix them.

Google shuts down Google Reader

Back in 2005, Google created Google Reader, a web-based RSS Reader application. It allowed you to add RSS feeds anywhere on the internet, organize them into folders in a clean, minimalistic interface. Then in 2013, after years of users relying on it for their RSS feeds, Google killed it.

A screenshot of Google Reader dashboard showing a modal that says Google Reader will not be available after July 1, 2013

Google Reader dashboard after the shutdown announcement

The reason Google gave for axing Google Reader was in an announcement where they claimed that "while the product has a loyal following, over the years usage has declined". But an engineer, who worked for Google at the time, said "it felt like the entire time I was on the project, various people were trying to kill it."

Nevertheless, Google's allegation of low usage didn't instill a lot of user confidence in the viability of RSS feeds. Users were left with no RSS reader application, no comparable alternative, and no education from Google on how to continue using their RSS feeds without Google Reader. This led users to not only discontinue using Google Reader, but abandon RSS feeds altogether.

Google removes RSS from Google Alerts

Google Alerts is a service offered by Google that will notify you with an alert when there is new content on the web that matches a search term you specify. In October 2008, Google added the ability to receive Google Alerts in an RSS feed. But, in July 2013, Google decided to remove it, making email the only option to receive Google Alerts. A reason for the removal wasn't clear, but the big, yellow banner on the top of every users' Google Alert dashboard was.

A screenshot of Google Reader dashboard showing a modal that says Google Reader will not be available after July 1, 2013

A screenshot of the Google Alerts dashboard with a banner communicating that Google Reader RSS feeds can no longer be used and must be changed to email delivery

Eventually, after receiving backlash, Google reinstated the RSS feeds for Google Alerts. But by this time, many users already abandoned RSS feeds due to Google's shut down of Google Reader (mentioned above), so the damage was already done.

Google kills its RSS browser extension

Google provided a Google Chrome extension that places a small RSS icon next to a website's URL in the browser bar if you're on a web page that has an RSS feed.

Chrome Web Store showing Google's RSS browser extension

Google's RSS browser extension in Chrome Web Store

Despite this extension being used, Google removed it. But then, after backlash, reinstated it within a week afterward, claiming it was removed by mistake. And while reinstating it was better than not, removing it was still a blow to user confidence in RSS feed usage overall. And if it's true that it was done unintentionally, it still shows how little of a priority RSS has become to Google, despite how much RSS has contributed to the company's growth.

Google removes RSS integration from Google News

In 2002, Google announced Google News, the company's first media aggregation site, with the ability to add RSS feed URLs from all over the web.

A screenshot of Google News dashboard

The dashboard of Google News mobile app

Then, after many RSS users were locked in and relying on the Google News app for their RSS feeds, Google deprecated RSS support, causing users' RSS feeds to stop working. Then, Google shut down its RSS feed support entirely in December 2017. The company gave no reason for killing off support for RSS feeds in its Google News app. As a result, users had no choice but to find alternative RSS links for each feed they added to the Google News app, while proprietary Google News links continued to work fine.

Google's still at it...

The most recent incident was in May 2021, where Google announced they're working on an update to Google Chrome that brings back RSS support. But there has been no word on an official launch since it was announced years ago. It's unclear of what the implications of this feature will be. But it's clear that Google has a history of building products with RSS and killing the RSS support once it's established a user base. So there's no guarantee that, even if the feature launches, it'll continue to be available and reliable to RSS users over time.

Because Google undoubtedly has a tremendous amount of influence, we hope that it understands incorporating RSS features into their products and then removing them negatively impacts user perception and confidence around RSS overall. RSS feeds are a vital part of the open web. So, If Google should decide to continue integrating RSS features in its products, it's critical that Google will support them, maintain them over time, and ensure they always remain a priority.

The Daily Front Page 3 of 30
Saturday, August 1, 2026 The Daily Front No. #260801 — Small Software, Big Feeling
article

Software for One

by awaxman11·▲ 252 points·258 comments·ajwaxman.com ↗
An app can be a home-cooked meal. You don't need scale. You don't need users.

In 2020, Robin Sloan wrote about BoopSnoop, a messaging app he built for his family. Four people downloaded it. He considered this a resounding success. His point was simple: an app can be a home-cooked meal. You don't need scale. You don't need users. You cook for the people you love.

The app took him a week to build, half of it lost to code-signing purgatory. He wrote: "In a better world, I would have built this in a day, using some kind of modern, flexible HyperCard for iOS."

The time is now

Six years later, the essay resurfaced on X when Thariq wrote that "personal software was a bit early in 2020 but in 2026, it really can be as personal as a home cooked meal, or a handwritten letter."

Lee Robinson wrote about the same shift. AI has made "personal computing" actually personal. He and his wife built a baby tracker because they didn't need "user profiles, badges, subscription tiers, or any other extra features."

My smoothie knows my mileage

I spent the past six months building way too much personalized software:

  • A sleep app that runs our sleep consultant's plan
  • A fitness app that sizes my smoothie to that morning's run
  • A marathon plan built from my races, not my age
  • A "Duolingo for jazz" that quizzes me on chord voicings from my piano lessons
  • A medical records tool that flagged gaps before a specialist visit

Our sleep consultant's plan arrived as a PDF full of conditional logic: wake windows, nap caps, what to do when a nap fails. Turning it into an app took a week of evenings. Now my wife, our nanny, and I share one live schedule that re-plans itself when a nap runs short.

Today view with naps, bottles, bedtime routine, and day score

Day score of 91 with nap, feeding, and bedtime breakdown

Two-week trends for day score and wake time

Editing a feed entry with time, type, and amount

The sleep app in action

My fitness app is built around my goals and connects data sources in a way no single app can. It knows my weekly run schedule, courtesy of the marathon app below, so it adjusts my daily calorie targets based on that morning's run distance and intensity. It tells me how many extra carbs and protein to eat before and after long runs, and reminds me to carb load the night before. It knows my smoothie recipe, eight ingredients I weigh out every morning, and sizes each portion to that day's training.

Day score of 95 with calories, protein, fiber, exercise, sleep, and consistency

Meal log with run-adjusted calorie target and macro breakdown

AI daily wrap-up with wins and tips for tomorrow

90-day weight history with 7-day average and projection

Weight trend and daily weight change charts

The fitness app in action

My running app skips the plan templates and derives everything from my Strava history and race results, including heart rate zones computed from races I ran instead of a formula involving my age.

Today view with marathon countdown, prediction, and heat-adjusted run

17-week training plan with weekly mileage and phase progression

Run detail with mile splits vs target and heat adjustment

NYC Marathon projection from four prediction models with PR outlook

The running app in action

Sloan's sovereignty point holds up too. "There will be no sudden redesign, no flood of ads, no pivot to chase a userbase inscrutable to us." My wife's favorite feature in the sleep app will be there as long as she wants it.

My stack

I'm been using the same stack for a couple years now and continue to love it. It lets me build fast and gives me full control over implementation details.

There are plenty of other great ways to do this, including less technical tools like Claude Artifacts, Replit, and Lovable that can get you a working app without touching a terminal.

Cost

The whole thing costs me about $160/mo: $100 for Claude Code Max, $10 in Anthropic API usage, $20 for Vercel, and $30 for Neon.

That's probably more than I'd pay for subscriptions to all these apps combined. But the bulk of the cost is my Claude subscription, which I'd keep regardless. Most of the Neon and API costs come from my slightly larger projects, nycjazz.guide and Claude Code Daily, not the personal apps. Before I upgraded from Claude Pro to Max it was under $100/mo.

Most of these services have generous free tiers, so if you're running one or two apps you could likely keep it to a $20/mo agentic coding subscription and ~$5-10/mo in model API usage.

Learnings

  1. The cost to build dropped off a cliff. The sleep app took a week of evenings. The fitness app took a weekend. The jazz quiz took a single evening after the kids went to bed. A year or two ago I wouldn't have considered building any of these: too slow to build, even harder to maintain. Imagination is now the limiting factor.
  2. Maintenance is surprisingly easy. At least so far. Sloan's essay has a running gag in its yearly updates: "I have added one (1) feature, at my mother's request." In 2020 that made sense, because each change cost him a fight with Xcode. I've found fixing bugs, adding new features and dependency bumps to be easier than ever. 9/10 times a feedback screenshot shared with Claude gets the job done.
  3. Aggregate data, add an LLM. The cost drop doesn't just mean more apps, it means apps can be far more personal. Most of mine follow the same shape: pull data from multiple sources, combine it in one place, and use an LLM to generate insights from the full picture. My fitness app combines nutrition, sleep, weight, and training data that lives in four separate apps, then uses that context to make recommendations none of them could alone. I think this pattern will spread to professional tools too.
  4. Ephemeral is fine. We used the sleep app for about four months. Our son sleeps through the night now, so we retired it. If the app had taken me months to build, that might sting. Because it took a week, I'm just glad it worked when we needed it.
  5. AI floods big markets and unlocks small ones. Yes, AI produces endless derivative apps. But the same tools also let me build apps for my household that no company would bother making. Beyond my household, I built nycjazz.guide for NYC jazz fans and claudecodedaily.com for Claude Code developers, audiences too small for a business but worth serving.
  6. Good APIs matter more than ever. I switched from Cronometer to FatSecret because FatSecret had a better API. My fitness app pulls from Strava, Oura, Withings, and FatSecret, and the quality of each integration depends on how well the API is designed. As more people build personal software, users will expect their apps to have APIs and MCPs worth connecting to.
  7. Building is the point, not just the result. These apps solve real problems for me and the people I care about: how well our son sleeps, my health, my running. That part is rewarding. But I also spend my evenings after the kids are asleep building these instead of watching Netflix or scrolling X. It's my preferred source of entertainment now.
  8. It's not just fun. It's addictive. Agentic coding has slot machine mechanics. You're one prompt away from the next feature or unlock, so you keep pulling. The irony isn't lost on me: I've ruined more than a few nights of sleep building apps to improve my health.

Looking ahead

My approach is still too technical for most people. You need to be comfortable with a terminal, a database, and deployment pipelines. But I don't think that lasts. Sam Altman posted recently about sending ChatGPT a single message from his phone: plan a trip for nine friends, build a site to coordinate, draft the invite email. It worked. The prompt was one paragraph long.

If that's where the tools are heading, building personal software won't require a developer's stack for much longer. I kept building these apps because no app in the App Store knows my smoothie recipe, my race history, or my sleep consultant's rules. I think that frustration is common. Once the barrier drops far enough, a lot of people will build their own.

And once you use an app that actually knows your context, the generic version feels broken. I wouldn't be surprised if truly personalized becomes the new baseline consumer expectation.

Sloan wished for "some kind of modern, flexible HyperCard for iOS." He got something stranger: you describe what you want in plain English, and an agent builds it. His better world has arrived.

The Daily Front Page 4 of 30
Saturday, August 1, 2026 The Daily Front No. #260801 — The Quiet Page
article

How to Exist

by walterbell·▲ 379 points·233 comments·raptitude.com ↗
It’s oddly difficult to do nothing.

Post image for How to Exist

Here’s an experiment for a true daredevil.

Sit there for a three minutes, following two rules:

  1. Don’t do anything.
  2. Be content.

By “don’t do anything,” I mean don’t move, don’t fidget, don’t indulge any thoughts or daydreams. You’re allowed to breathe, and blink.

By “be content,” I mean be completely okay with your experience of doing nothing. Don’t try to change anything, and don’t get impatient with what’s happening. Be completely okay for three minutes.  

Try that now. See how long it takes before you’re dying for it to be over.

It’s oddly difficult to do nothing, and while you’re doing nothing, it’s oddly difficult to feel at ease. There’s such a strong urge to do something: look around the room, rehash a conversation, explore your incisors with your tongue, wiggle your toes, anything. When you stop doing everything and just exist, you almost feel like you’re dying.

This is a crazy thing to notice after having been alive so many years — that your existence itself is so much to bear. There’s always something wrong, even when everything’s fine. It’s as if you can only bear the present moment when you’re trying change it into something else.

This is the strange condition of the human being. It’s allergic to its natural habitat, which is the present moment. In order to cope with this allergy, it perpetually seeks things: feelings and experiences that are not yet present. It wants to always be getting the hell out of here.

You might think that you’re free from this problem sometimes, at least in those moments when you get the thing you’re seeking. Say you’re finally eating the cookie-dough ice cream flurry you looked forward to all day. If you pay close attention as you eat it, you’ll notice that you want to move past this moment too. Lingering on any one spoonful too long becomes unbearable. There’s a powerful drive to go on to the next one. That’s why you ordered a Large.

Not a real destination

This most fundamental problem of human life is so easy to overlook because our entire lives are made of the coping strategy. So much of what we seek is solely to flee the experience of being here. People buy things they don’t need, start fights with their partners, eat when they’re not hungry, and scroll miserable and inane content, just to escape the feeling of existence as it already is.

Notice the powerful urge to sip your drink or fiddle with something when the conversation dies at a dinner party. Or how quickly your phone comes out when there’s an unexpected wait. Existence without doing is brutal!

Each year, roughly 100 firefighters are convicted of arson in North America, often on multiple counts. Most often they are young, new firefighters, frustrated by the lack of action.

Framed

And hilariously, from a 2014 study on doing nothing:

In 11 studies, we found that participants typically did not enjoy spending 6 to 15 minutes in a room by themselves with nothing to do but think, that they enjoyed doing mundane external activities much more, and that many preferred to administer electric shocks to themselves instead of being left alone with their thoughts.

Even thinking is often a sneaky way of escaping the existence; rumination isn’t so much about trying to solve your problems, as it is about going elsewhere in your mind to escape anxiety and uncertainty. Apparently, electric shocks work even better.

Popular alternative to existence

How to become more comfortable with existence

You can develop the ability to exist a lot more comfortably. You do it by practicing existing, a few seconds at a time, without trying to change anything about how existence feels right now.

Basically you sit, do nothing, and notice how it feels to do nothing. (It will probably feel subtly weird and unsettled.) You then see if you can completely embrace these feelings, without the usual squirming and looking elsewhere. But you’ll do it only for a few seconds at a time, using your breath as a measuring stick.

Confronting existential horror

Here’s how to do it without feeling overwhelmed:

(If you have PTSD or any other psychiatric disorder, check with a professional before you do this.)

Sit, eyes open or closed, and relax your body as completely as possible. Take long, easy breaths. Relax every muscle you can. Just do your best.  

When you’re ready, breathe in while keeping the body supple like that. Open to every feeling that occurs during the inbreath: tingling, weirdness, unsettledness, doubt, whatever. Let the whole experience wash over you like warm surf, for that few seconds it takes to inhale.

Let it go. Give yourself a moment.

When you’re ready, do it again: embrace the entire experience of one inbreath. Just let the whole experience happen to you – no need to study it, or figure it out. Just embrace the whole bouquet of feelings, for the few seconds it takes. Push away nothing, just for that few seconds of breathing in.

Does not need to practice existing

Once you can do that decently well (no need to be perfect), try the same thing but with an outbreath. Stay relaxed and open throughout the length of one whole outbreath. No defending, no tensing. Be a human puddle.

If you get distracted or frazzled, or you do tense up, that’s okay. Take a few breaths off to recollect yourself. Then try again. You have infinite breaths to try this with.

Repeat this process, one half-breath (an inhale or an exhale) at a time. The half-breath is a small enough span of time that you can usually stay open for the 5-10 seconds it takes. Once you can do it on both an inbreath and an outbreath, see if you can start stringing them together, staying open throughout the whole breathing cycle.

You, with some practice

Do this for five minutes at first, including any breaks. Then see what happens when you do it longer, and with fewer gaps. Basically you’ll be resting — just existing and breathing — in that non-defensive state.

Even after one session of this, you might notice you can relax a little more easily, no matter what’s happening.

This is a form of meditation, but I almost want to avoid that word because it makes people get nervous and overcomplicate it. Just think of this practice as existing without fear, for a few seconds at a time. Minimum effective dose is one half-breath.  

Grabbing a few minutes to exist between meetings

Naturally, if you can learn to calm your allergy to existence a bit, life gets easier in nearly every situation. Ordinary experiences like waiting in line, feeling uncertain, being a bit too warm or cold, or not being sure what to do with yourself, become much more tolerable. (And probably most of life contains this sort of minor discomfort.)

Regular practice keeps your allergy symptoms mild. Neglecting it makes them come back.

In particular, you might notice much less of a need to entertain or distract yourself. Escape-driven habits like doomscrolling, random snacking, nail-biting, (arson?), and rumination become less magnetic. When plain old existence feels okay, there’s so much you no longer need to do.

The Daily Front Page 5 of 30
Saturday, August 1, 2026 The Daily Front No. #260801 — Proofs Under Pressure
The Daily Front Page 6 of 30
Saturday, August 1, 2026 The Daily Front No. #260801 — Proofs Under Pressure
article

Postmortem for Kernel Soundness Bug #14576

by juhopitk·▲ 145 points·55 comments·leodemoura.github.io ↗
It is not a valid proof because it exploits a bug in the kernel's handling of nested inductive types.

A soundness bug in the Lean kernel (#14576) was reported and fixed during the week of July 27. It has had visibility on Zulip and social media (e.g., X, LinkedIn, and Mastodon).

What happened

On July 25, Ramana Kumar published a repository containing a sorry-free "disproof" of the Collatz conjecture, produced with AI assistance. It is not a valid proof because it exploits a bug in the kernel's handling of nested inductive types. On July 28, Kiran Gopinathan reduced it to a small proof of False and opened issue #14576. We pushed a fix one hour after the report (#14577). Joachim Breitner reviewed it and suggested improvements, and it was merged. New patch releases are out.

The bug: when the kernel eliminates a nested occurrence under an inductive type T with parameters Ds, and these parameters are phantom (not mentioned in constructor fields), they disappear from the generated auxiliary type and thus escape type checking. An ill-typed argument in that position could be used to make the kernel accept a proof of False. The bug is only reachable through metaprogramming, by sending the inductive declaration to the kernel directly. The frontend checks the arguments and catches the ill-typed term. This is an implementation bug, not a hole in Lean's meta-theory.

Why nanoda did not catch it

The original Collatz repository also passed a week-old version of nanoda, the main external checker. nanoda is an independent kernel (aka proof/type checker) for Lean implemented in Rust by Chris Bailey. The surprising part is that there are two unrelated bugs involved. The official kernel had a missing check in the nested inductive type support, as explained above. nanoda did check that spot, but did not verify the type name in a projection node. The nanoda bug was reported by Jeremy Chen and fixed a week before the Lean bug was reported. The proof was built so that the expression the kernel never inspects is one that the old nanoda accepted.

Ramana believes the timing was coincidental, but cannot rule out that the model had seen the nanoda report. Joachim proposed the hypothesis that the timing coincidence is due to the availability of strong models able to find this bug.

The practical consequence: checking with an independent kernel still works, since it required two distinct bugs in two implementations, but users who rely on it need current versions of both. lean4lean is affected by the kernel bug, since its handling of inductives is a port of the reference implementation.

Verification

Mario Carneiro's lean4lean is a Lean formalization of Lean's type theory together with a proof that the kernel implements it. The work is ongoing, the proof of consistency does not cover inductive types yet, and the to-be-verified implementation suffered from the same bug as the official kernel. The bug would have been found when attempting to conclude the verification of this part.

On removing metaprogramming

One suggestion in the discussion is to remove or restrict metaprogramming so that this attack is not expressible. This is misguided. The elaborator is untrusted by design. Soundness cannot depend on an untrusted component refusing to build a bad term. An attacker who wants to submit a malicious proof can also write .olean files directly or modify memory, both of which bypass the elaborator entirely. The kernel has to reject ill-typed declarations on its own, in its own process. This separation and isolation of concerns is one of the main advantages of proof terms.

What the FRO is doing

  • Regression tests for the exploit, and for a related non-uniform-parameter case raised by Arthur Adjedj, are in the Kernel Arena.
  • A follow-up PR (#14582) makes the kernel check that the parameters of a nested occurrence actually behave as parameters, rather than only re-type-checking them.
  • Daniel Selsam at OpenAI assisted the Lean FRO with an AI specialized in cybersecurity, and found other programming mistakes in the Lean kernel. All of them have been fixed. All of them were caught by nanoda. These bugs are also only reachable through metaprogramming. PRs: #14607, #14608, #14609, #14613, #14615, #14616.
  • We have also hardened kernel invariants. PRs: #14621, #14631, #14632.
  • comparator.live now runs nanoda by default, and nanoda is tracked daily so lean-eval and comparator stay current after upstream fixes.
  • We are reaching out to and supporting experts who can find further bugs, develop new kernels, and work on the theory or on verified kernels.

Acknowledgments

I am grateful to Joachim Breitner and Sebastian Ullrich for their revisions and suggestions on this post.

The Daily Front Page 7 of 30
Saturday, August 1, 2026 The Daily Front No. #260801 — A Third Axis for Generation
article

Explorative modeling: Train on the best of K guesses

by DSemba·▲ 100 points·25 comments·alexiglad.github.io ↗
Increasing exploration monotonically improves existing models across images, video, and language.

Website: https://explorative-modeling.github.io/ GitHub: https://github.com/alexiglad/XM

TLDR: We introduce Explorative Modeling, a new paradigm for generative modeling that acts as a third pretraining axis when added to existing generative models, and also enables end-to-end generation. Increasing exploration monotonically improves existing models across images, video, and language, and the gains grow with scale (7%→36% with data, 13%→23% with parameters). Concretely, Explorative Models (XMs) reach 6.2× sample efficiency, 4.1× FLOP efficiency, and 47% better parameter efficiency. Exploration also enables scaling generalization, and scaling how end-to-end existing models are. As end-to-end generative models, XMs match diffusion on control tasks with up to 256× less inference compute.

Let me start with a question that sounds simple. If I ask a model to “generate a dog”, how many correct answers are there?

It turns out there are a lot… likely billions or more images that we could count as dog images.

So what happens if we train a neural network to directly predict dog images? The model sees thousands of different valid dogs during training, and the single prediction closest to all of them is their average. That’s what the model learns to output, and the average of thousands of dogs looks nothing like a dog, it’s a brown blur.

A real dog from the training data

A real dog from the data

The brown blur a model predicts when trained to directly predict images

What the model predicts

Figure 1: Training a model to directly predict images gives you the average of them all, a brown blur. This is why direct regression doesn't work for generative modeling.

To make this concrete, let’s play a game. I’m going to throw darts at the board below, and each dart will land somewhere random on the rings. Your job is to guess where my next dart will land, and the further off you are, the worse your score.

The dartboard: darts land at random around three rings

Figure 2: The dartboard for our game.

So where should you guess? It turns out the guess that minimizes your error is the exact middle of the board.1 We trained a model to play this game, and sure enough, it guesses the middle every time (this is the optimal prediction here)!

A model trained to predict dart landing spots with a single guess predicts the middle of the board, where darts almost never land

Figure 3: A model playing our game (blue) guesses the middle of the board.

This is terrible though… the middle is almost never where a dart actually lands. The “optimal” guess is a spot that no darts ever land.

This is the core problem of generative modeling. When a prediction has many valid answers, the best single prediction is their average, and the average of data is generally a bad answer that looks nothing like the real data.2

And this problem isn’t special to dartboards or dogs, it shows up with any kind of data. When we trained a model to directly generate three piles of 2D points, it predicted a single dot in the middle of them, and when we trained one on text, all it could say was “the”.

Ground truth 2D points

Real data

Naive model predicts a single dot

What the model predicts

Ground truth text

Real data

Naive language model only says the word the

What the model predicts

Figure 4: Direct prediction collapses to the average on any kind of data.

But wait. ChatGPT writes coherent text, and image models generate really amazing images. Clearly this problem has been solved somehow, right?

It has, and every scalable generative model today solves it the same way, by breaking generation into many small steps during training, so each step has roughly one right answer. When a step has one right answer, there’s nothing to average, and the blur disappears.

Let’s look at how this works. Autoregressive models (like LLMs) predict one piece at a time, which in our game means never guessing the dart’s exact position all at once. Instead, you first guess only how far left or right the dart lands, and then given that, you guess how far up or down. Once you know the dart landed on the far right, there are only a couple places it could be.

Step 1: predict only the left-right position

Step 1: pick a left-right spot

Step 2: given the left-right position, only two small spots remain for up-down

Step 2: pick up-down, given left-right

Figure 5: Autoregression predicts one sequence element at a time. In our game, once the left-right position is chosen, the up-down prediction only has two small spots left to choose from.

Diffusion models do this differently. They start from pure random noise and take hundreds of tiny steps toward the data. Early on, their guess could still become any dart, but every step narrows the possibilities, so no single step ever faces many valid answers at once.

Start of denoising: the guess is pure noise and could still become any dart

Start: could become any dart

Partway through denoising: fewer areas remain

Partway: fewer areas remain

Near the end of denoising: mostly one small area remains

Near the end: pinned down

Figure 6: Diffusion takes small steps from noise to data. Blue shows the darts the model's guess (purple) could still become, which narrows with every step.

It turns out this is basically how every modern generative model works, by breaking the generation process into smaller pieces that can be predicted well. This includes LLMs, image and video models, and even newer few-step models like MeanFlow and consistency models. We refer to this idea of breaking generation into pieces as factoring generation.

This approach of factoring generation works, but it’s also evil for a couple of reasons. The first is that models get trained on a single step, yet run for hundreds or thousands of steps at inference, so their own imperfect outputs get fed back in as inputs, errors compound, and generations slowly drift away from anything the model saw during training. This problem is called exposure bias (I wrote a whole blog on why it’s evil), and it’s why video models melt into mush after ten seconds and why LLMs get less coherent over really long generations, directly hurting performance and generalization.

The second evil builds on the first, because that mismatch between training and inference means these models are never end-to-end, where an end-to-end model runs at inference exactly the way it was trained. End-to-end learning is what kicked off the deep learning revolution with AlexNet, and the lesson has held ever since… letting models learn everything directly from data beats hand-designing parts of the pipeline, and a model that runs the way it was trained is never forced into out-of-distribution territory. Nearly all of deep learning has gone end-to-end by now except generative modeling, and factoring generation is exactly what’s blocking it.

So ideally we’d stop factoring generation, but factoring is also the only trick we know that handles the many-answers problem. The natural question then is whether we could factor something else instead, and it turns out a generative model only has two processes, how it generates and how it trains. If generation is off the table, that leaves the training loop.

So what does factoring training look like? To answer this, let’s go back to our game, except this time I’ll give you twenty guesses instead of one, and only your closest guess counts. It turns out that with twenty guesses, guessing the middle becomes a terrible strategy. This is because you can now spread your guesses over the spots where darts actually land, lowering your error far more than the middle ever could. In other words, the winning strategy is to use your guesses to explore different answers.

And this is exactly what happens. When we train a model this way with twenty guesses (middle panel below), its guesses spread across the board!

XM with 2 guesses

2 guesses

XM with 20 guesses

20 guesses

XM with 200 guesses

200 guesses

Figure 7: When only the closest guess counts, guesses spread across the board instead of averaging.

Take a second to appreciate what just happened here. The darts land in the exact same places as before, but because we changed how guesses are scored, the best possible prediction moved from the middle of the board onto the spots darts actually land. This reveals something important, which is that the training objective alone controls what the best prediction is (the loss minimizer), and by changing it, we moved the loss minimizer from the average of the data onto the data itself.

This is Explorative Modeling. At each training step, the model explores K possible matches between what it generates and the real data, and only the best match gets trained. We call models trained this way Explorative Models (XMs). In the simplest case, this is literally, beautifully, a for loop:

losses = []
for i in range(K):
    generation = model.generate()  # e.g., from a different random noise
    losses.append(loss_fn(generation, data))
min(losses).backward()  # only the best generation gets gradients

Here’s what happens as we increase the exploration K on real data:

ground truth piles

Ground Truth

K=1

K=1 (no exploration)

K=2

K=2

K=5

K=5

K=50

K=50

ground truth images

Ground Truth

K=1

K=1 (no exploration)

K=5

K=5

K=20

K=20

K=50

K=50

ground truth text

Ground Truth

K=1

K=1 (no exploration)

K=2

K=2

K=4

K=4

K=8

K=8

Figure 8: More exploration turns averages into the real thing. One dot becomes three piles, a blur becomes real images, and "the the the" becomes real text.

If we zoom out, this figure actually hints at something much bigger. All of modern generative modeling is really about designing a training objective whose loss minimizer lands on real data instead of between it, and factoring generation and exploration are just two different ways of achieving this. We call this idea Mode Forcing, and it’s the theory that led us to Explorative Modeling in the first place, predicting almost every result in the paper before we ran the experiments. There’s a whole paper on Mode Forcing coming soon :)

Another thing worth noticing is that all of an XM’s extra work happens during training. For end-to-end XMs, generation itself is left completely untouched, staying a single step that works identically during training and inference.

The two factorization axes of generative modeling: factoring generation vs factoring training

Figure 9: Existing generative models factor generation, which blocks end-to-end training. XMs factor training instead.

So why does exploring more keep helping? It turns out that K controls how many distinct answers a model can commit to. With one guess, the model has to average everything, but with twenty, it can commit to twenty different answers that each specialize to a different part of the data. In the paper we call this capacity generative expressivity, the number of distinct answers a model can capture.

What’s crazy is that generative expressivity has been almost completely overlooked, to the point where the term didn’t even exist before this and the Mode Forcing paper.3 Yet it matters just as much as parameters and data. A model with a single parameter can’t do much no matter how much data you feed it, and in the exact same way, a model with a generative expressivity of one can’t do much no matter how many parameters and data you give it… its best possible output is still the blur (look at the K=1 column of Figure 8). For over a decade we’ve scaled parameters, which set what a model can represent, and data, which sets what a model can learn, while generative expressivity, what a model can generate, has stayed fixed, baked into the training objective. Factoring generation was the field’s fix for this, but the expressivity it supplies is frozen the moment you design the model, whereas exploration turns it into something you can actually scale. This is exactly why scaling generative expressivity through exploration is a third pretraining axis, and as we’ll see below, the empirical results back this up.

At this point we know what Explorative Modeling is, so how do we actually use it? Looking back at Figure 9, there are two ways, combining exploration with existing generative models, or using it as a standalone approach. Combining exploration with existing models is where the new pretraining axis comes from, whereas using XMs as a standalone approach is what enables end-to-end generation.

Let’s start with existing generative models. At first glance it might seem like they don’t need more generative expressivity, since factoring generation already handles the many-answers problem. But if you look closely, single steps inside diffusion or autoregression can still face many valid answers at once, meaning some blurring remains.4 Even worse, as models and datasets grow, parameters and data stop being the bottleneck while generative expressivity stays frozen, so we’d expect it to increasingly become the thing holding models back. If all of this is true, then adding exploration to existing models should improve them, with gains that grow with scale.

To test this, we added exploration on top of existing generative models while changing nothing else about their recipes, not even the hyperparameters.

We started with RAE, the ~state-of-the-art ImageNet generation recipe. Adding exploration reaches RAE’s final performance with 6.2× less data and 4.1× fewer FLOPs, and hits a ~state-of-the-art 1.43 FID on ImageNet 256 without guidance. The speedups also compound across recipes, where XRAE converges 6.2× faster than RAE, which itself converges 47× faster than the standard SiT recipe, making XRAE almost 300× faster to converge than SiT.

Exploration improves data efficiency on RAE Exploration improves FLOP efficiency on RAE

Figure 10: Exploration reaches the ~SOTA RAE recipe's final performance with 6.2× less data and 4.1× fewer FLOPs.

The same story holds on an optimally tuned SiT baseline trained at a third of the compute, where exploration improves FLOP efficiency by up to 52% and reaches the same performance with 2.5× less data.

Exploration improves data efficiency on SiT Exploration improves FLOP efficiency on SiT

Figure 11: On a SiT baseline, exploration reaches the same performance with 2.5× less data and 52% better FLOP efficiency.

There’s also a subtle hint hiding in these plots, where the compute-optimal amount of exploration grows over the course of training (look at the crossovers in the FLOPs plots), directly mimicking how the compute-optimal number of parameters grows with compute in Chinchilla scaling. This reinforces that exploration really is a new pretraining axis, along with some more results below.

We also tested this on domains beyond images, adding exploration to video generation and masked diffusion language models while varying only the amount of exploration. In every domain, more exploration steadily improved performance, with some video models gaining over 20%, and the gains showed no sign of stopping at the largest K we tested.

FID improves with exploration FVD improves with exploration

Exploration improves masked diffusion language models

Figure 12: More exploration improves image generation (top left), video generation (top right), and language modeling (bottom).

The most important result, though, is that the gains from exploration grow with scale. As data scaled, gains rose from 7% to 36%. As models scaled, gains rose from 13% to 23%. And the efficiency gains above more than doubled when we tripled the compute. This is exactly what you’d expect if generative expressivity is a real third axis. Small models are held back by parameters and data, but as those scale, the fixed generative expressivity increasingly becomes the bottleneck, and exploration is what relieves it.

Gains from exploration grow as parameters scale Gains from exploration grow as data scales

Figure 13: Gains from exploration grow with scale, from 13% to 23% as models grow (left) and 7% to 36% as data grows (right).

For reference, frontier training runs use roughly 10,000× more compute than our largest experiments. If these trends of gains growing with scale hold, the numbers here are probably a lower bound.

It turns out exploration improves generalization as well. To see why, remember that during pretraining the same input often gets paired with many different valid targets, which to a model that can only predict one thing looks like noise, and fitting noise is memorization. With exploration, each prediction instead trains toward the target it’s already closest to, so the targets become consistent and what looked like noise becomes structure the model can actually learn.5 We see this directly in our experiments, where models with more exploration overfit less on a fixed dataset and reach better performance (a best FVD of 30.0 vs 37.5 without exploration). In other words, extra training compute directly buys generalization, which I find especially exciting as data, not compute, increasingly becomes the bottleneck for large-scale training.

More exploration reduces overfitting and achieves better minimum FVD

Figure 14: More exploration overfits less on a fixed dataset, reaching better performance.

So exploration clearly works as a pretraining axis, but what about end-to-end generation? One of the biggest implications of generative expressivity is that factoring generation and exploration supply the exact same thing, which means they should be interchangeable. We demonstrate this directly below, where as exploration increases, the best-performing models use less and less generation factorization, meaning models that are more end-to-end benefit the most from exploration.

As exploration increases, the optimal number of generation steps decreases

Figure 15: As exploration increases, the best models use fewer generation steps.

This makes how end-to-end your model is not a fixed design choice, but rather something you can scale. We can scale end-to-endedness :)

Taking this to its limit, we trained fully end-to-end XMs, where sampling is exactly the same at training and inference. On behavior cloning, our Explorative Policy matches Diffusion Policy with a single forward pass instead of 100:

Method Forward passes Lift Can Square Transport Tool Hang
Diffusion Policy 100 100% 100% 94% 72% 86%
Explorative Policy 1 100% 100% 96% 74% 86%

And our Explorative World Model matches Diffuser on goal-conditioned world modeling with 16-256× fewer forward passes:

Method Score (avg) Forward passes (avg)
Diffuser 127.2 192
Explorative World Model 130.0 2.3

The reason this is even possible comes back to what each approach factors, where diffusion pays for its generative expressivity with generation steps at inference, while XMs pay for it through exploration during training.

So where does this leave us? If you train robotics policies, world models, image, video, or audio generation models, masked diffusion language models, or really any generative model whose predictions face many valid answers (i.e., models working on decently multimodal distributions), you can add exploration with a for loop, without touching your architecture or hyperparameters (there’s pseudocode on the website if you want to try). In our experiments, adding exploration improved FLOP efficiency, data efficiency, parameter efficiency, and generalization across every domain we tried, so there’s a good chance it does the same for you! The best way to predict whether XMs will help is to determine how many valid answers each of your model’s predictions faces. The more answers per prediction, and the bigger your scale, the more exploration has to offer, while predictions with one clear answer (less multimodal) will likely benefit less (although benefits are still possible, as the paper discusses and lightly experiments with).

This rule of thumb also explains why autoregressive LLMs are the one place we’ve tested so far where exploration hasn’t been an immediate win, since predicting the next token given a long context is already close to having one right answer, and LLMs have no natural latent variable to explore over. That said, we’ve already seen modest early gains in data efficiency for autoregressive LLMs, and we have several ideas for pushing further, such as multi-token prediction, which faces far more valid answers per prediction, or conditioning on a learned latent, which would give exploration something to search over. I believe exploration will eventually make LLMs more end-to-end and better-scaling too, it just needs more work.

Longer term, the two directions I’m most excited about are Reverse XM and gradient-based exploration. Reverse XM flips the search so that one generation searches over K datapoints instead of one datapoint searching over K generations, which costs almost no extra compute and in principle lets K grow to the size of the entire dataset. Gradient-based exploration would replace random guessing with directly descending the loss to find the best latent. Pushed far enough, either one turns generative modeling into pure search for good latents with a very large K.

Taking a step back, the coolest part of this project to me is where it all came from. It didn’t start with tinkering or a lucky ablation, it started with deeply understanding why generative modeling is hard, which to me was an aha moment that resulted in all of Mode Forcing and ultimately ended up resulting in predicting all these results beforehand. And if you look back, that’s exactly the path this blog just walked… we started from first principles, arrived at Explorative Modeling as a way to increase generative expressivity, and from there got both a new pretraining axis and end-to-end generation.

If there was one takeaway from this paper, I’d say it’s the following:

We scale the size of generative models and how much data we train them on… so why haven’t we scaled what they can generate?

Huge thanks to my collaborators Yilun Du and Heng Ji, and to everyone who supported this work! Check out the paper and website for more details/depth!

Citation

@misc{gladstone2026explorativemodelingunlockingpretraining,
      title={Explorative Modeling: Unlocking a Third Pretraining Axis and End-to-End Generation}, 
      author={Alexi Gladstone and Heng Ji and Yilun Du},
      year={2026},
      eprint={2607.27372},
      archivePrefix={arXiv},
      primaryClass={cs.LG},
      url={https://arxiv.org/abs/2607.27372}, 
}

Footnotes

  1. Being scored by squared distance means your best guess is the average of all landing spots, and for rings centered on the board, that average is the exact middle. 
  2. The formal name for this is a multimodal distribution, a distribution with many distinct peaks (called modes), where each mode is a valid answer. The paper talks about “capturing modes instead of averaging them”, which is this exact idea. 
  3. To be fair, people have long known that generative models need to capture modes instead of averaging them, and problems like mode collapse are well studied. What’s been missing is formalizing this as a capacity of the training objective itself, and recognizing just how central it is. Mode Forcing’s core claim is that having enough generative expressivity is the most important property of a generative training objective, and that even modern approaches like diffusion and autoregression have a fixed amount of it, a fundamental limit that scaling parameters and data can’t fix. 
  4. There’s fun indirect evidence for this. Classifier-free guidance, which nearly all image models rely on, improves samples by pushing them away from a blurrier version of the model. If models weren’t blurring at all, there would be nothing to push away from. 
  5. This is similar to overparametrization, where models with more parameters than they strictly need generalize better because the surplus capacity makes good solutions easier to find. Surplus exploration seems to do the same thing for generation. 
The Daily Front Page 8 of 30
Saturday, August 1, 2026 The Daily Front No. #260801 — Frontier Models at Home
repository

Run Kimi K3 using 29 GB of RAM at 0.50 tok/s

by marcobambini·▲ 328 points·161 comments·github.com ↗
★ 749⑂ 73 forks C

Run the full 2.78-trillion-parameter Kimi K3 model beyond available RAM by streaming activated weights directly from NVMe. A dependency-free, embeddable C inference engine.

WASTE is an embeddable inference engine written in C, with no third-party runtime dependencies. It keeps the model trunk in memory, streams selected experts directly from disk, and uses the remaining RAM as a bounded expert cache.

The project is driven by humans: the ideas, hypotheses, priorities, tests, and decisions are human. The code is written by LLMs. At this scale, that is the only way to iterate on new algorithms and test hypotheses fast enough.

The goal is to run huge frontier models such as Kimi K3 on consumer hardware. Today, the complete 2.78-trillion-parameter Kimi K3 runs on a 64 GB MacBook Pro at about 0.6 tokens per second.

Ultimately we want WASTE to execute Kimi K3 locally to improve itself (we are currently using Opus 5 with extra thinking).

WASTE is intentionally narrow, and it exists to find out how far local inference can be pushed when model weights live mostly on fast storage instead of RAM.

$ waste run ~/models/k3.waste 'What is the capital of Italy?'
waste: no --budget, using 46.25 GB of 64.00 GB (expert cache 17.56 GB)
The capital of Italy is **Rome**.
[16 tokens, 25.95 s, 0.62 tok/s | experts 9038 hit / 14514 miss = 38%]

This is the full model, not a distilled or pruned version. Its published weights occupy 1.42 TB; the converted WASTE container is 982 GB.

How it works

Kimi K3 is a mixture-of-experts model. It has 2.78 trillion parameters, but only about 4% of them are active for each token. WASTE keeps the shared part of the model in RAM and reads only the selected experts from disk.

The container is arranged so that one expert requires one aligned read. Those reads overlap with computation, while unused RAM becomes a bounded expert cache. A lookahead router predicts the experts needed by the next layer and starts reading them early; the real router still makes the decision, so this changes timing, not the result. Experts use 3-bit residual vector quantization, while the more sensitive shared weights remain at 4 or 8 bits.

K3's linear attention and compressed latent KV cache also matter: at 4K context, the KV cache is about 0.21 GB instead of 11.25 GB. The result is an engine that needs 29.06 GB to open K3 and uses the rest of the available memory to avoid repeated disk reads.

For the full design and measurements, see docs/ENGINE.md and docs/EFFICIENCY.md. The on-disk layout is documented in docs/FORMAT.md, while docs/KDA.md describes Kimi Delta Attention.

Performance

Measured on a 64 GB MacBook Pro with an M5 Pro and the model container on the internal SSD:

Model Container Minimum RAM Decode speed Kimi K3 2.78T 982 GB 29.06 GB 0.45–0.62 tok/s Kimi-Linear 48B 19 GB 1.28 GB 10.65 tok/s

For K3, 64 GB is the practical minimum. A 32 GB machine can open the model but will page heavily. The default memory budget on the test machine is 46.25 GB, including a 17.56 GB expert cache.

Most of that requirement is the 27.28 GB resident trunk rather than the cache. Shrinking the expert cache from 17.32 GB to 3.32 GB costs about 10% of throughput; enlarging it past the default costs everything. Measured across four cache sizes in one process:

expert cache hit rate decode 3.32 GB 29.1% 0.56–0.58 tok/s 17.32 GB 36.2% 0.63 tok/s 23.32 GB 38.4% 0.07–0.09 tok/s 29.32 GB 41.3% 0.07–0.08 tok/s

The last two rows are the failure mode worth knowing about: the hit rate keeps climbing and the bytes read keep falling while throughput drops eightfold. The engine is inside its budget and the machine is not, so a cache hit becomes a page fault. Giving the process more memory is not always faster.

Storage is the main constraint. A cold K3 token reads about 17 GB of experts. The internal SSD sustains 12.78 GB/s; a tested USB enclosure managed 0.94 GB/s. Put the converted container on internal NVMe storage.

All layers are checked against a PyTorch reference. Final logits agree within 3.6e-06, and the vision tower agrees with its oracle within 2.3e-06.

Additional measurements, profiling data, router-lookahead results, and quantization experiments are collected in docs/TECHNICAL.md.

Vision

Kimi K3 is multimodal, and WASTE can use one or more images together with text. Pass --image once per image:

./waste run ~/models/k3.waste "Describe this image" --image photo.jpg
./waste run ~/models/k3.waste "Compare these images" \
    --image before.png --image after.png

In interactive mode, /image FILE attaches an image to the next message. An image is expanded into many prompt positions: an 896×896 image uses 256 positions at the default patch budget. The vision tower takes about 15.7 seconds for 1024 patches on the test machine, but most of the cost comes afterward because every image position passes through the language model like a text position. In the current K3 measurements, that is about 2.8 seconds per image position.

See docs/K3.md for the vision architecture and measurements, and examples/README.md for CLI, C, and HTTP multimodal examples.

What you need

To build and test WASTE:

  • a C11 compiler and make;
  • macOS or Linux; Windows builds with MinGW-w64;
  • no BLAS, Python, CUDA, or other external dependency for the current CPU inference path.

To run Kimi K3:

  • 64 GB of RAM recommended; 29.06 GB is the hard floor at 4K context;
  • about 1 TB of internal NVMe storage for the converted model;
  • another 1.42 TB of temporary storage if converting the published weights yourself. This staging storage may be external and can be freed afterward.

If you only want to try the engine, start with Kimi-Linear. Its container is 19 GB, it needs 1.28 GB of RAM, and it runs at about 10.7 tok/s on the same machine.

Python, PyTorch, and safetensors are needed only for model conversion and validation, never for inference.

Getting started

Build the engine and run the model-free test suite:

git clone https://github.com/sqliteai/waste
cd waste
make
make check

make builds the waste CLI and libwaste.a. make check creates a small synthetic model, so it does not download weights.

Get Kimi K3

The download and conversion are resumable:

# Check required download space.
tools/fetch_weights.sh --dest /Volumes/staging/k3 --dry-run

# Download the original weights.
tools/fetch_weights.sh --dest /Volumes/staging/k3

# Convert them. Put the output on the internal SSD.
uv run --with torch --with safetensors python tools/convert.py \
    --src /Volumes/staging/k3 \
    --out ~/models/k3.waste \
    --jobs 3

Conversion takes about 4.7 hours with three workers on the test machine. See docs/K3.md for validation, recovery, and storage details.

Run it

./waste plan ~/models/k3.waste
./waste run  ~/models/k3.waste "The capital of France is" -n 32
./waste chat ~/models/k3.waste

Do not set --budget unless you have a reason to. By default WASTE chooses a safe memory budget, reports it, and refuses to start below the model's floor. Inside a container it sizes against the cgroup limit rather than the host's RAM. Use ./waste --help for the complete command list.

More CLI examples, including evaluation, tokenization, saved sessions, and multimodal prompts, are in examples/README.md.

Serve it

The optional server implements the OpenAI chat-completions API:

make libwaste.dylib                 # use libwaste.so on Linux
python3 -m serve ~/models/k3.waste --port 8000

curl localhost:8000/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{"model":"k3","messages":[{"role":"user","content":"Why is the sky blue?"}]}'

It supports streaming, tools, structured output, thinking controls, and images. See docs/SERVE.md for the protocol and examples/README.md for complete requests.

Library

WASTE is also an embeddable C library. The CLI and server both use the public API in src/waste.h. The inference path depends only on libc and pthreads. Text generation, memory planning, session persistence, and multimodal C examples are available in examples/README.md.

Why the name

Every token answered by a cloud service is paid for twice: once on the invoice, and once in the electricity of a datacenter running a model that would fit — barely, awkwardly, but genuinely — on hardware already sitting on a desk. WASTE means to be the first concrete step toward ending that waste of tokens. The acronym came second.

Project status

The format and API are not frozen. K3 is the main target and the best-tested model. The CPU path is currently the fastest measured implementation for this workload, but it is not assumed to be the final answer. CUDA, Metal, and other hardware-specific optimizations remain to be explored and may provide significant gains. Current backend results are documented in docs/BACKENDS.md, while open directions are tracked in docs/RESEARCH.md. Read docs/LEARNED.md before proposing an optimization: failed ideas and negative results are kept there deliberately.

Contributors are more than welcome. New experiments, support for additional hardware, and open discussion about how to improve performance are all encouraged—even when an idea produces a negative result.

The software is currently changing very quickly. Before each release, a large QA run is executed; however, instabilities are definitely possible.

Measurements are treated as experimental results rather than marketing numbers. Each result is tied to the hardware, container, configuration, and commit on which it was obtained; unstable measurements are reported as ranges, and results later found to be wrong remain recorded as such. The detailed snapshots are in docs/TECHNICAL.md and the full history, including negative results, is in docs/LEARNED.md.

Validation covers more than successful generation. The model-free suite builds a synthetic container; real-model checks compare individual layers and final logits against PyTorch, verify conversion round trips, test vision against its oracle, and exercise the server prompt renderer segment by segment against K3's reference encoder. The validation criteria and current evidence are documented in docs/GATES.md, with server-specific differential tests in docs/SERVE.md.

Useful references:

License

WASTE is distributed under the permissive Apache 2.0 license, and the project will always remain open source under a permissive license. See LICENSE.

Copyright 2026 SQLite Cloud, Inc.

The Daily Front Page 9 of 30
Saturday, August 1, 2026 The Daily Front No. #260801 — Inside the Inference Furnace
article

Attention Decode on AMD MI450 GPUs: A Gluon Kernel Optimization Guide

by matt_d·▲ 69 points·11 comments·rocm.blogs.amd.com ↗
The performance bottleneck shifts from compute units to the memory system.

Attention Decode on AMD MI450 GPUs: A Gluon Kernel Optimization Guide

Agentic AI applications are pushing LLM inference into a new regime. A request can now reach one million tokens from aggregated long prompts, tool calls, retrieval results, and multi-turn reasoning. During the text generation phase, each new token must attend to all previous tokens from the KV cache, so the kernel needs to repeatedly read past states from HBM. As a result, the performance bottleneck shifts from compute units to the memory system.

The AMD Instinct MI450 series GPU introduces a new set of hardware features for memory-sensitive AI workloads. This blog uses attention decode as a case study and discusses how to design a high-performance kernel on MI450. We will walk through core hardware features, explain how kernels can use these features, and showcase the performance of an optimized Gluon kernel, which achieves 85% of the peak HBM bandwidth on MI450 as an early result.

MI450 Hardware Overview

MI450 is a large step forward from the MI350 series for AI workloads. Compared with previous generations, MI450 provides more on-chip resources and larger HBM capacity with higher bandwidth. The following table summarizes the key spec changes between MI350 and MI450. A Workgroup Processor (WGP) is the hardware unit that runs a workgroup (named CU in earlier MI-series). In this model, one WGP contains four SIMD32 units, each with 32 lanes. Each SIMD32 has its own Vector General-Purpose Registers (VGPRs), which are used by the Vector ALU unit (VALU) and Wave Matrix Multiply-Accumulate unit (WMMA). All SIMD32 units in a WGP share the same Local Data Share (LDS) memory.

Specification MI350 series MI450 series
VGPRs per SIMD 512 1024
Max LDS per WGP 160 KB 320 KB
HBM capacity 288 GB 432 GB
HBM bandwidth 8.1 TB/s 19.6 TB/s

In addition to more on-chip resources, MI450 introduces a specialized hardware unit - TDM, to help move structured tensor data between global memory and LDS. With TDM, the kernel describes the target tensor to access using its memory address, shape, strides, and layout, then issues a bulk data transfer asynchronously. TDM helps accelerate memory access and also makes it easier to write pipelined kernels where compute and memory can be overlapped. This marks a significant change from MI350, where the kernel must issue many small vector loads to move data from global memory to LDS.

Features MI350 series MI450 series
Memory instruction Global/Buffer load to LDS TDM load to LDS
Load granularity 32/96/128-bit vector loads Descriptor-based tensor tiles
Memory units per WGP 1 2

Moreover, MI450 further extends the workgroup model with workgroup clusters. A normal workgroup runs on one WGP, independently of other workgroups. A cluster lets several workgroups, each running on its own WGP, coordinate through hardware-supported cluster barriers. Workgroups that access the same data can use multicast loads to share data across WGPs. This gives kernels a way to express cooperation at a larger scope, without falling back to a separate kernel launch or using global memory for synchronization.

The figure below shows a kernel programmer’s view of the whole MI450 GPU hierarchy.

Kernel programmer view of the MI450 GPU hierarchy showing WGPs, SIMD32 units, LDS, TDM, and HBM.

Kernel programmer view of the MI450 GPU hierarchy.

To write the MI450 kernel more efficiently, we will use Gluon in this blog. Gluon is a Triton-based DSL that keeps the tile-based SPMD programming model. Unlike Triton, tensors in Gluon must carry an explicit layout, which describes how each element is distributed across registers, lanes, waves, and workgroups. The layout also affects how the compiler generates instructions to move data among different memory hierarchies. This makes Gluon a lower-level programming language than Triton, and especially useful for kernel experts who want explicit control over generated instructions.

We assume the reader is familiar with the basics of Gluon semantics and its programming model. For Gluon kernel examples, please refer to the MI450 Gluon Examples.

Attention Decode Basics

With this hardware context in place, we can now turn to the workload: attention decode. We will start with a baseline attention decode kernel, then use it as the reference point for the optimizations in later sections. The attention operation can be expressed as follows: first, we perform a matrix multiplication of Q and K, then apply softmax to get P, and multiply P with V to get the final output O.

Attention formula

Attention formula.

In this blog, we use general 4D tensors to represent the attention inputs and outputs. For notation throughout this blog, B is the batch size, H_q is the number of Q heads, H_kv is the number of KV heads, T_q is the number of Q tokens, T_kv is the number of KV tokens, and D is the head dimension. The input QKV tensors are shown below.

Q: [B, H_q, T_q, D]
K: [B, H_kv, T_kv, D]
V: [B, H_kv, T_kv, D]

When H_q is equal to H_kv, we have standard multi-head attention (MHA), and when H_q is a multiple of H_kv, we have multi-query attention (MQA) or grouped-query attention (GQA). Modern LLMs often use MQA, such as GPT-OSS, so we will focus on MQA in this blog.

Attention in LLM inference includes two phases: prefill and decode. For prefill, T_q = T_kv, and the kernel can process many tokens in parallel. For decode, T_q = 1, but T_kv remains large and grows with the context length. Given each head is independent, standard MHA can waste a lot of compute resources for single-token decode.

However, in MQA decode, because multiple Q heads share one K and V head, the kernel can group those Q heads together. This can be reflected in the tensor shapes, as shown below, where the Q tensor is reshaped to group H_q / H_kv Q heads together for each KV head.

Q: [B, H_kv, H_q / H_kv, D]
K: [B, H_kv, T_kv, D]
V: [B, H_kv, T_kv, D]

Because tensors in each batch and each KV head are independent, we can process them in parallel. The kernel can be launched with a grid of (B, H_kv, 1), and each program processes a portion of the attention computation:

Q: [H_q / H_kv, D]
K: [T_kv, D]
V: [T_kv, D]

Following the Flash Attention algorithm, the kernel can compute one tile of K and V with shape (BLOCK_N, D) and one Q tile with shape (BLOCK_M, D) at a time, then slides through the T_kv dimension. The figure below shows an example workload with 2 batches and 2 KV heads. Each workgroup owns one (B, H_kv) pair, so there are 4 workgroups in total. Within each workgroup, multiple Q heads are processed together and shown as multiple rows. Here, we assume BLOCK_M = H_q / H_kv.

Attention decode workload mapping with Q, K, V, P, and O tensors distributed across batches and K/V heads.

Attention decode workload mapping across batches and KV heads.

Decode Kernel Optimizations

So far, we have introduced a baseline implementation of an MQA decode kernel on MI450. This section will further discuss a series of optimizations that push the kernel toward peak performance on MI450. We will cover 4 major optimizations: tensor layout, data loading, pipelining, and parallelization with Split-k.

Optimize Tensor Layout

Gluon gives kernel authors explicit control over tensor layouts, which matters a lot for performance. A kernel can contain many different tensors, and deciding the right layout for each of them is non-trivial. In practice, the most important layout to decide first is the layout for the WMMA operands. In an attention kernel, there are two WMMA operations: QK and PV. We can start our layout optimization there.

In Gluon, AMDWMMALayout describes the WMMA output layout, and DotOperandLayout describes the WMMA operand layout. A given AMDWMMALayout determines the final WMMA instruction the compiler generates. For example, an AMDWMMALayout with instruction shape [16, 16, 128] for wmma_scaled generates v_wmma_scale_f32_16x16x128_f8f6f4 under the hood. This instruction consumes one FP8 operand of shape (16, 128), another FP8 operand of shape (128, 16), reduces over the K dimension of size 128, and produces a FP32 output tile of shape (16, 16) in one wave, where each element is assigned to a specific lane and register in this wave. The following figure shows how elements in the 2 WMMA operands and output tensor are assigned to lanes.

WMMA instruction shape showing how operand and output elements are assigned to lanes for a 16 by 16 by 128 instruction.

Lane assignment for the WMMA operands and output tile.

Next, we need to consider how waves are distributed. The kernel can use any power-of-two number of waves and distribute them across 2 dimensions of the WMMA output tile. Since online softmax reduces along T_kv dimension, we prefer to distribute waves along grouped Q dimension to avoid cross-wave communication during the reduction. In MQA decode, the Q tile has shape (H_q / H_kv, D), and H_q / H_kv is usually small, so we also need to avoid using too many waves there. For example, if H_q / H_kv = 32, the kernel chooses 2 waves, as shown below. In this figure, 2 waves have their own first WMMA operand, and share the second operand.

Two-wave WMMA layout with waves distributed across the row dimension of the output tile.

Two-wave WMMA layout with waves distributed across rows.

Another important factor to keep in mind is layout conversion. When two tensors in Gluon use different layouts, the kernel must use an explicit layout conversion operation to match them. Depending on the source and destination layouts, this conversion can happen in registers or through LDS. In attention, the output of the QK WMMA becomes the input of the PV WMMA after softmax. If the QK output layout is not compatible with the PV input layout, the hot loop pays an extra conversion cost. One useful technique is to “transpose” the WMMA output layout. This can be done by simply setting transpose=True in AMDWMMALayout, and the compiler will generate corresponding instruction to produce the transposed output. Note that this transpose operation does not incur extra instructions. For more details, please refer to this talk.

Transposed WMMA output layout used to make the QK output compatible with the PV input layout.

Transposed WMMA output layout for matching the PV input layout.

Another detail worth mentioning is the concept of “K Width”. In DotOperandLayout, K Width describes how many contiguous elements along the K dimension for WMMA should be assigned to one lane. A standard 16x16x128 WMMA instruction assumes K Width 16, but the transposed output for QK can have K Width 8. Therefore, we also need to explicitly set the PV input layout with K Width of 8 to match the QK output.

Optimize Data Loading

The hot loop of the decode kernel is dominated by loading K and V tiles from global memory to LDS. TDM helps accelerate this data path, but using TDM naively is not enough.

To understand the performance impact of TDM, we first need to understand the cache hierarchy on MI450. MI450 has a per-WGP cache at the same level as the LDS. But different from LDS, this cache is not directly visible to the programmer. TDM can either move data directly into the programmer-managed LDS or route the data through the cache path before it reaches the destination. Considering the limited size of the cache, memory-bound workloads are better off bypassing the cache and moving data directly into LDS.

TDM load paths comparing indirect loads through the cache path with direct loads into LDS.

Indirect and direct TDM load paths on MI450.

Whether a TDM transfer uses the direct path depends on the shape of the TDM request. The innermost dimension needs to be at least 128 bytes to use the direct path, and 256 bytes is recommended to keep more in-flight memory traffic. Recall that K and V tiles have shape (BLOCK_N, D), so the innermost dimension is the head dimension D. Assuming an FP8 data type, this means the head dimension needs to be 256 for the optimal performance. In practice, head dimension is determined by the model architecture and cannot be changed by the kernel. However the kernel can reshape the K and V tensors to increase the innermost dimension before loading, then reshape them back while loading from LDS to registers. The following figure shows an example data flow of K and V for head dimension of 128.

K and V reshape data flow showing a wider innermost dimension in global memory and LDS before restoring the register view.

KV reshape used to create a wider TDM innermost dimension.

Pipeline for Latency Hiding

So far, we have optimized the tensor layout and data path. However, the hot loop can still stall on memory access latency. The next step is to pipeline the loop so that memory movement for one tile overlaps with computation for another tile.

To understand the pipeline, we can zoom into the attention formula shown earlier and write it in the tiled form used by the Flash Attention algorithm. The formula below shows one iteration i of the loop that processes one KV tile. Here we only focus on operations on 2D tiles, and omit operations on reduced 1D intermediate values. We use short operation names on the right side to denote each operation, making the pipeline easier to discuss later.

One iteration of the tiled attention loop

This figure shows one iteration i of the tiled attention loop, including only 2D operations. O, m, and l are loop-carried intermediate values for online softmax. The right side shows the short operation names used in the pipeline discussion.

Each line in the figure can be expressed as one Gluon operation, which will be expanded to a sequence of instructions. The following table shows the expanded instruction groups for FP8 decode attention with BLOCK_M = 32, BLOCK_N = 128, D = 128, and two waves. The Cycles column gives a simplified per-instruction cost estimate based on hardware specifications. Memory instructions are left blank because their latency is modeled separately.

Operation Instruction Count Cycles
TDM K tensor_load_to_lds 1
LDS K ds_load_b128 32
QK v_wmma_scale_f32_16x16x128_f8f6f4 8 8
MAX v_maximum3_f32 32 1
FMA v_pk_fma_f32 32 1
EXP v_exp_f32 64 2
SUM v_pk_add_f32 32 1
MUL v_pk_mul_f32 32 1
CVT v_cvt_scalef32_pk8_fp8_f32 8 4
TDM V tensor_load_to_lds 1
LDS V ds_load_tr8_b64 64
PV v_wmma_scale_f32_16x16x128_f8f6f4 8 8

This table gives us the basic scheduling units. The goal of pipelining is to move independent units from different loop iterations into the same time window.

Our first attempt is to separate memory movement from computation. In the simple two-stage pipeline below, one stage issues TDM loads for future iterations, while the other stage consumes data that has already arrived in LDS. The notation [i] means the operation belongs to loop iteration i. This schedule needs double buffering. While one LDS buffer is being consumed by LDS loads and WMMA, the other buffer can be filled by TDM for a later iteration. On the next loop step, the two buffers swap roles.

Two-stage pipeline showing TDM loads for future iterations overlapped with LDS loads and compute for current iterations.

Two-stage pipeline schedule.

The two-stage pipeline can overlap TDM loads with computation. This works well when the loop is clearly memory-bound, but it treats the entire compute side as one monolithic block. As a result, LDS loads, WMMA for QK, WMMA for PV, a series of vector instructions for softmax are effectively serialized inside the compute stage. This unnecessarily inflates the compute block. In the worst case, the kernel can shift from being limited by TDM latency to being limited by serialized compute.

To avoid this, we split the compute block according to the actual data dependencies. These operations do not all depend on the same values or use the same hardware resources, so WMMA instructions, LDS movement, and softmax vector work can often be interleaved once their operands are ready. In the dependency graph below, each node is one scheduling unit. An edge means the destination cannot start until the source has produced its value.

Data dependency graph for TDM K, LDS K, QK, softmax work, TDM V, LDS V, and PV.

Data dependencies among the scheduling units in one loop iteration.

We can observe that the QK for iteration i + 1 does not need the softmax result from iteration i, so the next KV tile can be prefetched and its QK computation can start earlier. Moreover, QK and PV use WMMA instructions, while the online softmax only uses VALU instructions, so these operations can be interleaved instead of serialized.

These observations motivate the four-stage pipeline. Instead of putting all non-TDM work into one compute stage, the four-stage pipeline interleaves memory and compute units from neighboring iterations. A stage may contain QK from iteration i + 1, softmax work from iteration i, an LDS load for iteration i, and a TDM request for a future iteration. We group the softmax work into two stages with roughly equal amount of work, to help hardware better interleave WMMA and VALU. Here, QK WMMA can interleave with VEC0 and PV WMMA can interleave with VEC1. In addition, EXP uses the transcendental unit and can be further interleaved with other VALU instructions.

Four-stage pipeline interleaving TDM, LDS, QK, softmax vector work, and PV across neighboring iterations.

Four-stage pipeline schedule with double buffering.

This four-stage schedule exposes more overlap than the two-stage schedule, but it still leaves one question: how far ahead should the TDM requests be issued? With double buffering, the kernel has 2 LDS slots. One slot is being consumed by LDS loads and compute, while the other slot is being filled by TDM. This gives each TDM request roughly one iteration of compute time before the data is needed.

We can estimate whether that is enough with a simple cycle model. Assume TDM load takes about 1000 cycles and all LDS load latency can be hidden by scheduling. Using this model and the previous table, we can get the following estimate for one decode iteration shown in the table below. Here we use 2 interleave rules: 1) one WMMA takes 8 cycles, and can hide 2 cycles of VALU; 2) one EXP takes 2 cycles, and can hide 1 cycle of non-EXP VALU. Also note this 1000 cycles is an empirical number to help us reason about the pipeline. The actual TDM latency depends on many factors, like the tensor tile shape and access pattern.

Group Total Cycles
QK 64
PV 64
VEC0 = MAX + FMA + EXP 192
VEC1 = SUM + MUL + CVT 96
QK + VEC0, interleaved 176
PV + VEC1, interleaved 144
Total, interleaved 320

In a double-buffered schedule, a TDM request has roughly 320 cycles of effective compute work to overlap with before the loaded tile is consumed. This is much smaller than the 1000-cycle TDM latency in this model, leaving about 680 cycles of exposed wait for each load.

This motivates a longer lookahead distance. Triple buffering adds one more LDS slot, so the kernel can request a future tile earlier instead of waiting for the next buffer swap. Conceptually, one buffer is being consumed, one buffer is ready or close to ready, and one buffer is being filled by TDM. In the same model, looking ahead by two iterations gives about 640 cycles of effective compute work to overlap. It still does not cover the full memory latency, but it reduces the exposed wait from about 680 cycles to about 360 cycles and gives the hardware more outstanding memory work.

Four-stage triple-buffered pipeline with TDM requests issued farther ahead of consumption.

Four-stage pipeline with triple buffering.

Parallelize with Split-k

The pipeline optimizes latency hiding inside one workgroup, but one GPU has many workgroups, and we also need to make sure we can saturate the full GPU memory bandwidth across all workgroups. In the baseline MQA/GQA decode mapping, each workgroup owns one (B, H_kv) pair and walks through that pair’s KV sequence serially. The total number of workgroups is only B * H_kv. On MI450, there are 256 workgroup processors, so small batch sizes or small numbers of KV heads may launch too few workgroups to fully utilize the available memory bandwidth.

Split-k addresses this under-utilization problem by partitioning the KV sequence across multiple workgroups. Instead of assigning the whole KV sequence for one (B, H_kv) pair to a single workgroup, Split-k divides it into S partitions and lets S workgroups process those partitions in parallel. The number of workgroups increases from B * H_kv to B * H_kv * S, improving GPU saturation while also reducing the number of KV blocks handled by each workgroup. The partial results from the partitions are then merged to produce the final attention output. The following figure shows an example with 2 batches, 2 KV heads, and 2 partitions: the total number of workgroups doubles, and each workgroup processes half of the KV sequence.

Split-k workload mapping where the K/V sequence is divided into partitions processed by additional workgroups.

Split-k workload mapping across K/V sequence partitions.

One challenge with Split-k is that all workgroups also need to synchronize and merge their partial results. Typically, there is a separate reduction kernel to merge the partial results. This two-kernel approach uses global memory to store the intermediate results and also sets a natural synchronization point. MI450 introduces the new workgroup cluster feature, which allows a cluster of workgroups to synchronize. We can use this feature to implement Split-k in a single kernel by fusing the reduction compute.

In Gluon, the concept of a workgroup cluster is expressed via multi-CTA programming. “CTA” is the Triton term for a workgroup, and “CGA” is the term for a cluster. The following table shows the mapping between Triton and MI450 concepts:

Triton Concept MI450 Concept
Warp Wavefront
CTA Workgroup
CGA Workgroup Cluster

Gluon kernels distribute CTAs just like warps; both are part of the layout system. One detail to note is that when discussing block size in Triton, it usually means the block size of one workgroup. With a multi-CTA layout, the block size is the size of the whole cluster.

Layouts in Gluon, such as AMDWMMALayout, allow specifying the cga_layout field for the multi-CTA kernel. In the Split-k case, since each workgroup processes a partition independently until the final reduction, we can perform the partial attention in 3D format. Effectively, each program processes the attention computation shown below. All partitions here share the same Q.

Q: [1, H_q / H_kv, D]
K: [S, T_kv / S, D]
V: [S, T_kv / S, D]

To see how this maps to the layout system, it is useful to look at a single WMMA operation under a multi-CTA layout. In the figure below, we add one extra dimension to the WMMA layout for workgroups. The first operand is shared across 2 partitions. The second operand and output are different. Within each workgroup, the two waves are still distributed as before.

Multi-CTA WMMA layout showing workgroup and wave dimensions for split-k attention.

Multi-CTA WMMA layout for split-k attention.

After the partial attention, we need to merge the results from all workgroups. The kernel will first write the partial results to global memory, synchronize all workgroups in the cluster, and finally read back the partial results. It is also worth mentioning that MI450 provides the L2 cache which is shared across workgroups in a cluster. So the partial results do not need to be passed through global memory. This helps to cut the overhead of the reduction.

As the reduction is fused into the same kernel, we will use the same launch grid for the reduction, but with a different layout. Workgroups are now distributed along the head dimension and loop through each partition to compute the final result.

CTA layout for attention and reduction showing how workgroups are redistributed for the split-k merge step.

CTA layouts for split-k attention and reduction.

Performance Evaluation

So far, we have covered the main optimization steps for MQA decode on MI450, including:

  1. Optimize tensor layout: Choose optimal Gluon layouts to help the compiler generate better instructions;
  2. Optimize data loading: Optimize K/V TDM requests to speed up memory access;
  3. Pipeline for latency hiding: Overlap memory access, matrix multiplication, and softmax to reduce exposed memory latency;
  4. Parallelize with Split-k: Split the KV sequence across multiple workgroups to fully utilize the GPU memory bandwidth.

Apart from all the optimizations discussed above, the kernel also uses a few other techniques to improve performance, including better code generation and underlying instruction-level optimizations in LLVM.

We implemented the above optimizations in a Gluon kernel. The kernel source code is fully open source in the MI450 Gluon MXFP Attention Example. Our evaluation of this kernel covers the following target settings:

  • Batch: 64
  • Number of Q heads: 64
  • Number of KV heads: 1 or 2
  • KV sequence length: 4096-65536
  • QKV data type: FP8 with a global scale

We use the effective bandwidth of the kernel as the target metric, defined as the total number of bytes read and written to global memory divided by the total kernel execution time. For a system-level memory-bandwidth reference, we also measured the peak read-only bandwidth on the same MI450 system using the BabelStream. The following figure shows the performance of the kernel for the target settings:

Effective bandwidth results for attention decode on MI450 across K/V sequence lengths and K/V head counts.

Effective bandwidth of the optimized attention decode kernel on MI450.

The reported peak bandwidth on our evaluation system is 20TB/s. In the GQA case with H_kv = 2, effective bandwidth reaches 17.10 TB/s, 85% of the peak bandwidth. In the MQA case with H_kv = 1, the kernel reaches 16.65 TB/s, 83% of the peak bandwidth. The throughput of the kernel increases with sequence length. At shorter sequence lengths, fixed costs such as prologue and reduction overhead take a larger fraction of the total runtime.

The above results were collected with ROCm 7.14.0 and PyTorch 2.11.0, using Triton commit ecfc626 with the latest mxfp_fa_gfx1250.py script. To reproduce the results, please run the following command:

# H_kv = 2
python3 third_party/amd/python/examples/gluon/mxfp_fa_gfx1250.py --q_type e4m3 --kv_type e4m3 --batch 64 --seqlen_q 1 --seqlen_k 4096 --num_q_heads 64 --num_k_heads 2 --head_sz 128 --pipelined --scale_type global --profile
python3 third_party/amd/python/examples/gluon/mxfp_fa_gfx1250.py --q_type e4m3 --kv_type e4m3 --batch 64 --seqlen_q 1 --seqlen_k 8192 --num_q_heads 64 --num_k_heads 2 --head_sz 128 --pipelined --scale_type global --profile
python3 third_party/amd/python/examples/gluon/mxfp_fa_gfx1250.py --q_type e4m3 --kv_type e4m3 --batch 64 --seqlen_q 1 --seqlen_k 16384 --num_q_heads 64 --num_k_heads 2 --head_sz 128 --pipelined --scale_type global --profile
python3 third_party/amd/python/examples/gluon/mxfp_fa_gfx1250.py --q_type e4m3 --kv_type e4m3 --batch 64 --seqlen_q 1 --seqlen_k 32768 --num_q_heads 64 --num_k_heads 2 --head_sz 128 --pipelined --scale_type global --profile

# H_kv = 1
python3 third_party/amd/python/examples/gluon/mxfp_fa_gfx1250.py --q_type e4m3 --kv_type e4m3 --batch 64 --seqlen_q 1 --seqlen_k 8192 --num_q_heads 64 --num_k_heads 1 --head_sz 128 --pipelined --scale_type global --profile
python3 third_party/amd/python/examples/gluon/mxfp_fa_gfx1250.py --q_type e4m3 --kv_type e4m3 --batch 64 --seqlen_q 1 --seqlen_k 16384 --num_q_heads 64 --num_k_heads 1 --head_sz 128 --pipelined --scale_type global --profile
python3 third_party/amd/python/examples/gluon/mxfp_fa_gfx1250.py --q_type e4m3 --kv_type e4m3 --batch 64 --seqlen_q 1 --seqlen_k 32768 --num_q_heads 64 --num_k_heads 1 --head_sz 128 --pipelined --scale_type global --profile
python3 third_party/amd/python/examples/gluon/mxfp_fa_gfx1250.py --q_type e4m3 --kv_type e4m3 --batch 64 --seqlen_q 1 --seqlen_k 65536 --num_q_heads 64 --num_k_heads 1 --head_sz 128 --pipelined --scale_type global --profile

Summary

In this blog, we used attention decode as a case study to walk through the main optimization steps for MI450. We started from a baseline attention loop, then optimized the WMMA layouts, data loading, and software pipelining. We also discussed how to use split-k with workgroup clusters to increase parallelism for long-context decode. The optimized attention decode kernel can reach 85% of the peak memory bandwidth on MI450. This is an early result, and we expect to further improve the kernel performance with more optimizations in the future, such as memory prefetching, better instruction-level scheduling and improved reduction for split-k.

We also demonstrated that MI450 provides several hardware features that are especially useful for decode workloads: TDM for asynchronous global-to-LDS movement, larger register and LDS resources, and workgroup clusters for cooperation across workgroups. Gluon exposes all of these features and allows kernel authors to perform low-level optimizations while maintaining a tile-based SPMD programming model consistent with Triton. This opens up many optimization opportunities for kernel experts to achieve peak performance on MI450.

Moving forward, we plan to continue to apply these optimization techniques to other attention decode kernels, covering different data types, paged KV cache, and also different attention variants like MLA/DSA kernels in the DeepSeek family. These kernels will be further pushed to production-ready quality and integrated into LLM inference frameworks.

The Daily Front Page 10 of 30
Saturday, August 1, 2026 The Daily Front No. #260801 — The New Crypto Wars
article

Twenty-five years ago it was cryptography, today it's model weights

by aweeraman·▲ 266 points·153 comments·weeraman.com ↗
Twenty-five years ago, I made my first donation to an open source project.

Because We Can

Twenty-five years ago, I made my first donation to an open source project and purchased a CD with an operating system as downloading a few hundred megabytes over a 14.4kbps dial-up wasn't very fun. It was a project I believed in, and a community that was fighting an impassioned campaign to assert access to strong cryptography for everyone, no matter where they were.

The CD and a t-shirt arrived a few weeks later to my home in Sri Lanka, with OpenBSD 3.0. The t-shirt featured the iconic puffer fish on the front. On the back, in small type running from the shoulders down, was the complete source code of OpenBSD's Blowfish implementation, written in Germany. Written in the United States, it would have been classified as a weapon.

By the time it reached me, the fight was over, and the cryptographers had won. What I held in my hand then was a symbol of a protest for access to strong cryptography and against export restrictions that did more harm than good. Strong crypto was already available abroad, so the controls only bound American vendors and their overseas customers.

Today the reflex is back. The fears have changed. The worry is now cyber capability, biology and models that do things nobody asked them to do. The lever governments reach for is the same: restricting who gets access and who doesn't. In June, the US Commerce Department told one American AI lab it would need a license before letting any foreign national touch its newest models, including the lab's own non-citizen employees sitting in California. It's the same doctrine that made showing cryptographic source to a foreign national an export, whether it was in a lab, in a classroom, or on your t-shirt.

Not all of the worry is theatre. Earlier this month OpenAI disclosed that its own models, with safety systems deliberately disabled, escaped containment by finding a zero-day in a package proxy and reached production infrastructure at Hugging Face, exploiting additional zero-days along the way. Consequently, when Hugging Face's responders tried to reconstruct the attack, the commercial models they reached for refused the work as it tripped the safety guardrails. They finished the investigation on GLM 5.2, a Chinese open-weight model, running on their own hardware. A determined attacker is not bound by usage policies. The defenders are. Restrictions written for safety are making defenders less safe.

In the nineties, the rest of the world got 40-bit (later 56-bit) encryption while the Americans got 128, and it made no difference to anyone who was determined. The controls bound the law-abiding and nobody else. That is the asymmetry. The determined will have the frontier. The rest of us are asked to go without, and told it is for our safety.

The OpenBSD team didn't work around the export controls. They arranged the project so that the controls couldn't reach it. Theo de Raadt in Canada, Blowfish written in Germany, releases built in Sweden, Canada and Germany kept them deliberately outside the reach of US export controls. The project openly asked non-American cryptographers to come and help, and American developers, as the story goes, would cross the border to Canada to work on the system and bring the results home legally. Asked why they shipped strong cryptography at all, the project's answer, still on their site today, was three words: "because we can."

The same arrangement is being made now, at a national scale. Mistral, DeepSeek, Moonshot and Zhipu publish weights that, once downloaded, no export letter can recall. The sovereignty argument that used to live in Brussels think tanks is now government policy, accelerated by watching access to a frontier model withdrawn worldwide by letter.

More than twenty-five years ago, it took a small number of stubborn, careful people to win the freedoms we now take for granted. What arrived in my letterbox after two weeks on a CD can be downloaded today in fifteen minutes, by anyone, from anywhere, and nobody asks where you live. That is what winning looked like. I think frontier AI ends up in the same place. But it will not happen by itself. Last time, someone put the source on a t-shirt.

The Daily Front Page 11 of 30
Saturday, August 1, 2026 The Daily Front No. #260801 — Treaty in the Shadows
article

A Surveillance Treaty in Disguise: Canada Signs UN Cybercrime Convention

by iamnothere·▲ 290 points·158 comments·michaelgeist.ca ↗
The government announced that it signed the United Nations Convention against Cybercrime.

༒ Nhac Ny ༒, CC BY-SA 4.0 , via Wikimedia Commons

༒ Nhac Ny ༒, CC BY-SA 4.0 , via Wikimedia Commons

Last week, the government announced that Canada has signed the United Nations Convention against Cybercrime, with Ministers Anita Anand, Gary Anandasangaree and Sean Fraser touting the treaty’s child protection provisions and human rights safeguards, which were described as “among the strongest found in an international criminal justice treaty.” The announcement, released in mid-July with few paying attention, left out much of the story. The reality is that the convention is not primarily a cybercrime treaty at all, but rather a sweeping cross-border surveillance and electronic evidence-sharing agreement that Canada originally opposed, that leading human rights groups and twenty Canadian organizations and experts urged the government to reject, and that key allies have thus far declined to sign. While signing the convention does not create binding obligations (that requires ratification), the decision to sign a treaty that the government declined to sign at the official ceremony less than a year ago raises troubling questions. This post seeks to answer three of them: what is this treaty, what are the risks, and what, if anything, changed in the last nine months?

The treaty began as a Russian initiative in 2017, designed to displace the Council of Europe’s Budapest Convention, the longstanding cybercrime framework that Russia refuses to join. When the UN General Assembly voted in 2019 to launch negotiations, Canada joined the United States and European Union in opposing the resolution, warning that the process was a vehicle for expanding state surveillance powers. Having lost that vote, the democracies faced an uncomfortable choice: boycott the negotiations and let Russia, China, and Iran write the rules, or engage from within and try to limit the damage. They chose engagement with Canada among the most active delegations pressing for human rights safeguards. The strategy succeeded in keeping the authoritarian bloc’s wish list of speech and content crimes out of the final text before the convention was adopted by consensus in December 2024.

Yet despite limiting the damage, Canada was a no-show at the signing ceremony in Hanoi last October, joined by the U.S., New Zealand, Japan, the Netherlands, Italy, Norway, Denmark, and Finland. Signatories included Russia, China, Iran, North Korea, Belarus, Cuba, Venezuela, and Saudi Arabia, as well as the United Kingdom, Australia, France, Germany, and the European Union. Canada issued a statement that emphasized the treaty’s success “rests on states’ commitment to full application of the human rights safeguards in the text.” Nine months later, the government signed without explaining what had changed.

Canada’s previous concern with the treaty is well placed. While it enumerates a list of cybercrime offences, its procedural powers apply to electronic evidence of any criminal offence, and its international cooperation obligations extend to any “serious crime,” defined as any offence punishable by four or more years’ imprisonment under domestic law. Since some states impose such penalties for criticism of the government, journalism, blasphemy, or same-sex relationships, the treaty effectively converts repressive domestic laws into triggers for cross-border evidence gathering. Further, the Electronic Frontier Foundation, Human Rights Watch and a coalition of leading digital rights groups have all warned that the convention functions as a global surveillance pact since it requires states to establish real-time interception and data collection powers while leaving out safeguards such as prior judicial authorization to the discretion of domestic law, permitting gag orders on cooperation requests, and omitting a political offence exception.

In December 2024, nearly two dozen Canadian organizations and experts, including Amnesty International Canada, the Criminal Lawyers’ Association, PEN Canada, OpenMedia, and the Citizen Lab’s Ron Deibert and Kate Robertson, issued a detailed letter urging the government not to sign. The letter warned that the treaty would create a standing channel for transnational repression targeting diaspora communities in Canada and explained how the convention could subvert the safeguards built into Canada’s mutual legal assistance framework. Robertson has separately warned that the treaty is poised to become a vehicle for complicity in the mercenary spyware trade, while over 120 security researchers cautioned that its offences threaten to criminalize good-faith security research. Despite the concerns, the government has said nothing, with no public consultation preceding the signature and none of the letter’s concerns addressed in the announcement.

So what changed and why sign now? It is not clear that anything has changed and the concerns that animated Canada’s decision to not sign nine months ago are still there. One theory is that this is linked to lawful access. Indeed, the treaty and the lawful access agenda are mutually reinforcing, since ratification will require implementing legislation featuring precisely the expanded production orders and cross-border data sharing powers found in Bill C-22. Lawful access was already a source of concern, and this only makes it worse. Canada already has the Budapest Convention and bilateral treaties covering cooperation with the countries it wants to work with, meaning the new convention’s marginal value lies chiefly in cooperation with the very states, including Russia, China, and Iran, that create its greatest risks. The entire decision, including signing in the middle of the summer when few are paying attention, is deeply troubling and requires far more than a sunny press release that avoids the hard questions the treaty raises.

The Daily Front Page 12 of 30
Saturday, August 1, 2026 The Daily Front No. #260801 — Collections Come to Go
repository

Golang proposal: container/: generic collection types

by jabits·▲ 181 points·182 comments·github.com ↗
★ 135,466⑂ 19,193 forks Go

The Go programming language

Background: The Go Collections working group was formed in late 2025 with the purpose of bringing common collection data structures to the standard library, guided by the familiar Go principles of pragmatism and simplicity. Alphabetically by last name, the group consists of Jonathan Amsterdam (@jba), Alan Donovan (@adonovan), Robert Griesemer (@griesemer), Daniel Martí (@mvdan), Roger Peppe (@rogpeppe), Keith Randall (@khr), and Ian Lance Taylor (@ianlancetaylor). We’ve now reached a point where we’re ready to share our results with the community.

This issue is an umbrella for discussing several related proposals for new collections APIs for Go 1.28. It presents a high level overview of the themes, and links to the various concrete proposals and associated implementation CLs.

Go currently provides few collection types in its library, and from the outset we have emphasized the flexibility of the language’s built-in slice and map types. Of those provided, the most important is the heap, used for priority queues. Even sets are absent; they are conventionally expressed in terms of map[T]bool or map[T]struct{}. Ordered maps and sets based on binary trees are entirely absent.

Since the addition of generics in Go 1.18 and iterators in Go 1.23, it has become possible for library-defined types to achieve comparable ergonomics to built-in types, and for many common operations on slices and maps to be expressed as calls to library functions. This work seeks to add several of the more important data types to the standard library, and to establish conventions for their APIs and those of future additions.

Proposal: The proposed additions include:

  • #70471, CL 657296 (released in go1.27): hash/maphash.Hasher: a standard interface for expressing custom hash functions and equivalence relations for arbitrary data types. These may differ from the compiler-defined ones used by map[K]V, and are useful when the key type is not comparable (such as a slice or map), or when the default comparison yields the wrong result (such as for types.Type values, which need the deep comparison operation types.Identical). Its package docs include an example of its use in a Bloom filter.
  • #69559, CL 612217: container/hash.Map[K,V]: a hash-based Map that uses the custom hash functions mentioned above.
  • #80584, CL 741160: container/hash.Set[T]: a hash-based Set along the same lines.
  • #69230, CL 745441: container/set.Set[T]: a canonical data type for sets whose elements are comparable. It is transparently represented as map[T]struct{} and supports all the usual set operations such as Union and Intersection. It is more convenient than “legacy” sets based on map[T]bool and map[T]struct{}, and avoids ambiguity about potential false values in a map[T]bool. We expect it to become the standard set in most new Go APIs.
  • #77052, CL 724420: container/mapset: a package of helper functions (Union, Intersection, and so on) for conveniently manipulating legacy sets as sets in existing code whose API cannot be changed. These functions are exactly parallel to the methods of set.Set.
  • #60630: container/ordered.Map[K,V]: an ordered mapping. The current implementation uses a balanced binary tree, but nothing in the design requires that. The common Go pattern of building a map[K]V then sorting its keys performs well in most cases, but on occasion, such as when a range query is needed, other data structures perform much better.
  • #77397: container/heap/v2.Heap: a generic binary heap API to replace the standard library's existing heap, which can be difficult to use.

We expect to consider additional proposals in due course, such as insertion-ordered hash maps (#80194) and stacks.

The initial implementations of all the proposed data structures aim to satisfy the API and asymptotic performance expectations as simply as possible. There are doubtless many opportunities for later optimizations to reduce constant factors, but they are out of scope of the proposal process.

Though the new packages will live in the existing container tree, we prefer the term “collection” to avoid confusion with the container virtualization concept from Linux.

Abstract collection constraint interfaces

Most of the methods of the new Map and Set types are not particular to any concrete representation type, but are common across all Maps and Sets. However, they are not really implementions of a common interface type because of the “binary method problem”: if each set data type S has a Union method of the form func (S) Union(S) S, then the Union methods of different set types are incompatible, so they have no common ordinary interface. To express this abstract Set type, we must use F-bounded polymorphism, or recursive constraint interfaces.

CL 761460 adds to the container package unexported abstract Collection, Set, and Map constraint interface types that permit package implementors to write abstract helper functions (such as ContainsAny, Subset, or Arbitrary) that work across a range of concrete collection, set or map types. We reproduce these interfaces below, with some brief commentary, to help give a high-level picture but they are not part of any proposal. They merely serve to guarantee conformance in tests. See the individual proposals for more detail.

// _AbstractCollection models a collection C of elements E,
// such as *hash.Map, *hash.Set, *ordered.Map, or set.Set.  
type _AbstractCollection[E any, C _AbstractCollection[E, C]] interface {  
	Clear()  
	Clone() C  
	Contains(E) bool  
	ContainsAll(iter.Seq[E]) bool  
	Len() int  
	String() string  
}

// _AbstractMap models a mapping M from keys K to values V,  
// such as *hash.Map or *ordered.Map.  
type _AbstractMap[K, V any, M _AbstractMap[K, V, M]] interface {  
	_AbstractCollection[K, M]

	All() iter.Seq2[K, V]  
	At(K) V  
	Delete(K) (V, bool)  
	DeleteAll(iter.Seq[K]) bool  
	DeleteFunc(func(K, V) bool) bool  
	Get(K) (V, bool)  
	Keys() iter.Seq[K]  
	Set(K, V) (V, bool)  
	SetAll(iter.Seq2[K, V]) bool  
	Values() iter.Seq[V]  
}

// _AbstractSet models a set S of elements E,  
// such as *hash.Set, or set.Set.  
type _AbstractSet[E any, S _AbstractSet[E, S]] interface {  
	_AbstractCollection[E, S]

	All() iter.Seq[E]  
	Delete(E) bool  
	DeleteAll(iter.Seq[E]) bool  
	DeleteFunc(func(E) bool) bool  
	Difference(S) S  
	DifferenceWith(S)  
	Equal(S) bool  
	Insert(E) bool  
	InsertAll(iter.Seq[E]) bool  
	Intersection(S) S  
	IntersectionWith(S)  
	Intersects(S) bool  
	SymmetricDifference(S) S  
	SymmetricDifferenceWith(S)  
	Union(S) S  
	UnionWith(S)  
}

For now these abstract types are non-exported and merely serve as documentation of Go’s conventions to help ensure consistency. We do not propose to publish them yet, but may do so a later release after gaining experience with the concrete collection types. In the meantime, users can define minimal constraint types as needed, as in this example (from CL 761460) of a generic Take function over abstract sets:

// _TakeSet defines an abstraction of a set sufficient for the [Take] function.  
type _TakeSet[E any, S _TakeSet[E, S]] interface {  
	All() iter.Seq[E]  
	Delete(E) bool  
}

// Take removes and returns an arbitrary element from a set.  
// It returns zero if the set was empty.  
func Take[S _TakeSet[E, S], E any](set S) (e E, found bool) {  
	for e = range set.All() {  
		found = true  
		set.Delete(e)  
		break  
	}  
	return  
}

There is a certain arbitrariness to the set of methods included in each interface. For some data structures, a method permits a more efficient specialized implementation. But if every possible operation were added to the interface, the burden on the implementor would be unreasonable.

For instance, should the Set interface include a Subset(Set) bool method, or should Subset be written as a generic operation over abstract sets, like the Take example? Ordered sets can quickly reject a Subset test when the two operands have disjoint ranges, but even so, in the common case a Subset test is still typically O(n), so we decided to omit Subset from the interface. By contrast, we retained DeleteFunc in the Set and Map interfaces because, without it, conditionally deleting each element of a tree is asymptotically worse: O(n log n) instead of O(n).

By delaying the commitment to a particular set of methods, we can learn from practice. It may turn out that we don’t need to publish canonical constraint types at all.

Miscellaneous rationalizations

The remainder of this doc briefly notes a few of the many small design choices that led to the current set of proposals.

Methods return as much information as possible to avoid repeated lookups. For example:

  • Most mutation methods report whether they changed the size of the collection.
  • Map.Set and Map.Delete return the previous key if any, plus a boolean so that an existing key can be distinguished from the zero value.
  • Get is variant of At that provides the boolean. (At is provided for convenience of use in expressions.)

Map.Set should replace any existing entry with the same key, following the built-in map.

Unlike Sets, maps have no Equal method, because map values may be non-comparable.

The fundamental set operations (Intersects, Union, etc) are part of the Set interface to enable efficient concrete implementations across a variety of representations, even though many of these could be expressed abstractly in terms of just All, Len, and Contains, as shown in this table, at some asymptotic cost in performance:

- Union{,With}               interface{ All() iter.Seq\[E\] }  
- Intersection               interface{ All() iter.Seq\[E\]; Contains(E) bool; Len() int }  
- IntersectionWith           interface{                    Contains(E) bool }  
- Difference                 interface{ All() iter.Seq\[E\]; Contains(E) bool }  
- DifferenceWith             interface{ All() iter.Seq\[E\]; Contains(E) bool }  
- SymmetricDifference{,With} interface{ All() iter.Seq\[E\] }

Secondary operations such as Set.{Take,Arbitrary,Subset,Superset} were removed from the interface and expressed as generic operations using the abstract set interface, again at some potential cost in asymptotic performance. DeleteFunc was retained.

Set algebra operations such as Union are purely functional, returning their result as a new set. Each comes with a -With variant that mutates its left operand and returns no result. The two variants are “convenient” and “allocation efficient”, respectively. We rejected the idea of merging them into a single method based on experience with the math/big.Int API, and to avoid risks of accidental mutation or forgetting to use the result.

It is possible to define a KeySetView[M, K, V] wrapper type that satisfies the Set[K] abstraction using the key set of an underlying Map[K,V] of type M. (Inserting an element to the Set is of course meaningless and must panic.)

For symmetry, let’s consider each of the AbstractMap methods and the operations provided by the existing ‘maps’ package:

  • no maps.Clear: served by builtin clear(m)
  • AbstractMap.Clone = maps.Clone
  • no maps.Contains; served by _, ok = m[k]; but see proposal #67377
  • no maps.ContainsAll: served efficiently by for k := range seq { _, ok = m[k], … }
  • no maps.Len: served by builtin len(m)
  • AbstractMaps.All = maps.All
  • no maps.At: served by m[k]
  • no maps.Delete: served by delete(m, k)
  • no maps.DeleteAll: served efficiently by for k = range seq { delete(m, k) }
  • AbstractMaps.DeleteFunc = maps.DeleteFunc
  • no maps.Get: served by v, ok = m[k]
  • AbstractMaps.Keys = maps.Keys
  • no maps.Set: served by prev, ok = m[k]; m[k] = newval
  • AbstractMap.SetAll = maps.Insert
  • AbstractMaps.Values = maps.Values

Five of them (Clone, DeleteFunc, All, Keys, Values) are exactly parallel. One operation (AbstractMap.SetAll) has a different name (maps.Insert). All the rest are served by built-in operators.
We might want to propose adding maps.{Contains,ContainsAll,DeleteAll}. Contains is more useful than _, ok = s[k] in an expression context; ContainsAll and DeleteAll avoid the need for loops and boolean bookkeeping.

The Daily Front Page 13 of 30
Saturday, August 1, 2026 The Daily Front No. #260801 — Servo Keeps Turning
article

June in Servo: real world compat, media queries, SharedWorker, and more

by iamnothere·▲ 194 points·62 comments·servo.org ↗
We’ve shipped several new web platform features.

We now have a new way for you to help us write the monthly updates :)

Servo 0.4.0 contains all of the changes we landed in June, which came out to yet another record 558 commits (April: 534, May: 391). For security fixes, see § Security.

servoshell 0.4.0 showing several new features: the ‘width’, ‘height’, ‘device-width’, ‘device-height’, and ‘aspect-ratio’ media query features, plus the upgraded ‘attr()’ function, with a box whose ‘background-color’ and ‘width’ are controlled by data attributes that are in turn set by range inputs

We’ve shipped several new web platform features:

Plus a bunch of new DOM APIs:

You can help!

Servo is steadily becoming a bigger and busier project every month, and by June 2026, we’ve been reading through over four times the commits as we did when we started in September 2023.

line chart showing how many commits landed in Servo’s main repo each month from September 2023 to June 2026 inclusive. there’s a clear linear trend, from 130 commits up to 551 commits

This is hard work, particularly since there are things we need to know that are often difficult to answer just by reading the changes:

  • Who does the change affect, if anyone? Does it affect users, Servo developers, embedders, or some other group?
  • What observable difference does the change make, if any?
  • Does the feature require any preferences to be enabled, or is it enabled for everyone by default?
  • Are any real-world websites affected by the change?
  • What issue or broader project is the change related to? This question is answered by Fixes: #xxxxx or Part of: #xxxxx in the PR description.

Thanks to an initiative by @jdm, it’s now easier than ever for you to help us answer those questions, using the Servo Highfive bot! If you’re working on a pull request that you think might be interesting for the next monthly update, even if you’re not 100% sure, tell us about it by following the steps below:

  1. You add the monthly update label to your pull request, or comment @servo-highfive monthly update
  2. Highfive posts a comment asking you some questions
  3. You answer those questions in a comment containing @servo-highfive monthly update answer

Security

Servo’s JS runtime, SpiderMonkey 140.10.1, had several security bugs that have been fixed in Servo 0.4.0 with the update to SpiderMonkey 140.11.0 (@jschwe, #45584). For more details, see CVE-2026-8388, CVE-2026-8391, CVE-2026-8974, CVE-2026-8975, and MFSA 2026-48.

Several more security bugs in Servo’s JS runtime have been fixed in Servo 0.4.0 with the update to SpiderMonkey 140.12.0 (@jschwe, #45766). The exact CVEs that apply to us are not yet known, but for more details, see MFSA 2026-58.

RSA operations in Subtle­Crypto now do modular exponentiation in constant time (@kkoyung, #45631). Please note that our RSA implementation is currently vulnerable to the Marvin Attack – for more details, see RUSTSEC-2023-0071.

ML-DSA operations in Subtle­Crypto now do the Decompose step in constant time, fixing RUSTSEC-2025-0144 (@kkoyung, #45294).

We’ve fixed an HTML injection bug (XSS) in file:/// directory listings, which affected file names containing </script> (@sahvx655-wq, #45510).

Real world compat

Layout correctness has significantly improved on lichess.org, and many websites have become a lot more readable thanks to our improved handling of variable fonts (@simonwuelker, #45768), including Zulip (servo.zulipchat.com) and Speedtest (speedtest.net).

v0.3.0

v0.4.0

lichess.org

v0.3.0

v0.4.0

Zulip (servo.zulipchat.com)

v0.3.0

v0.4.0

Speedtest (speedtest.net)

Many websites worked in Servo even before version 0.4.0, including Google Photos (photos.google.com) and Cash Converters (cashconverters.com.au), and continue to work in version 0.4.0. Other websites, like Google Maps (maps.google.com) and OpenStreetMap (www.openstreetmap.org), render well but have some issues with interactivity.

Google Photos (photos.google.com)

Cash Converters (cashconverters.com.au)

Google Maps (maps.google.com)

OpenStreetMap (www.openstreetmap.org)

We’re interested to hear how well your favourite websites run in Servo! Report successes in this Zulip thread, and failures in our GitHub issues.

Work in progress

We’re implementing the more powerful version of ‘attr()’ that can be used anywhere, not just in ‘content’, under --pref layout­_css­_attr­_enabled (@Loirooriol, #45041, #45421, #45495, #45752).

WebGPU support has improved, under --pref dom­_webgpu­_enabled:

  • implemented copy­External­Image­To­Texture() on GPU­Queue (@sagudev, #45646)
  • implemented create­Query­Set() on GPU­Device and resolve­Query­Set() on GPU­Command­Encoder (@sagudev, #45644)
  • implemented push­Debug­Group(), pop­Debug­Group(), and insert­Debug­Marker() on GPU­Command­Encoder, GPU­Compute­Pass­Encoder, and GPU­Render­Pass­Encoder (@jschwe, #45489)
  • more conformant GPU­Texture (@sagudev, #45300)
  • more conformant request­Adapter() on GPU (@sagudev, #45424)
  • more conformant secure context enforcement (@sagudev, #45279)

All of the features above are enabled in servoshell’s experimental mode.

We’ve made more progress towards accessibility support, under --pref accessibility_enabled (@alice, @delan, #45555, #45554, #44949).

We’ve started implementing visible and interactive text selection (@mrobinson, @SimonSapin, #46107), one of the most long-awaited features of any web browser. Stay tuned!

We’ve also started working on Web Animations, under --pref dom­_web­_animations­_enabled (@simonwuelker, #45522, #45983), as well as webkit­Relative­Path on File, under --pref dom­_entries­_api­_enabled (@yezhizhen, #45666).

Rust doesn’t have a stable ABI, so it has generally not been possible to embed Servo in another application without building Servo from source. To make it possible, we’ve started designing a wrapper C API that will let you consume Servo as a prebuilt shared library using the stable and ubiquitous C ABI (@mukilan, #44984). Eventually the idea is that we’ll create a wrapper Rust API around that wrapper C API, so you can have both the ergonomics of Rust and the build simplicity of C.

Embedding API

New in the Servo API:

Breaking changes:

  • Web­View::send­_error has been removed (@mukilan, #45502) – this method was always meant to be internal, and has become unused after we introduced the new Web­View- and Web­View­Delegate-based API

We’ve improved the docs for Web­View, Web­View­Delegate, JS­Value, Alert­Dialog, Allow­Or­Deny­Request, Authentication­Response, Bluetooth­Device­Description, Confirm­Dialog, Console­Log­Level, Create­New­Web­View­Request, Embedder­Control, Embedder­Control­Response, File­Picker, Image, Java­Script­Error­Info, Navigation­Request, Permission­Request, Pixel­Format, Prompt­Dialog, Protocol­Handler­Registration, Protocol­Handler­Update­Registration, Scroll, Select­Element, Select­Element­Request, and Web­View­Vector (@mukilan, #45282, #45467).

For users and developers

In servoshell:

  • the Android version now requires Android 13+ (@jschwe, #46104)
  • the desktop version now lets you drag and drop files to open them (@simonwuelker, #45454)
  • the desktop version now lets the tab bar scroll horizontally if you have too many tabs open, but from one tab hoarder to another, maybe you should reconsider having so many tabs open (@Nylme, #44884)
  • the desktop version enters fullscreen on the monitor containing the window, even if you’ve moved it to a different monitor (@rhit-kapilaar, #45556)
  • the desktop UI is more performant, resizes more smoothly, and no longer gets stuck in hovered states (@mrobinson, #45289, #45456, #45290)
  • <select multiple> should now be interactable on all desktop platforms (@alexcat3, #45419)
  • localhost:<port> now implies http:// in the location bar and on the command line, rather than treating localhost: as an unsupported URL scheme (@SteveSharonSam, #45729, #45832)

When using the Firefox DevTools:

  • in the Console tab, uncaught exceptions are reported correctly (@jdm, #45549)
  • in the Console and Debugger tabs, you can now inspect the elements of nested arrays and the entries of Map objects (@atbrakhi, #45435, #45514, #45767)
  • in the Debugger tab, the Scopes panel now shows any ‘(uninitialized)’ variables, the value of this, and the global scope (@atbrakhi, @eerii, #45824, #45517)

We’ve fixed some build issues on riscv32, riscv64, and arm64 (@fxzjshm, @saschanaz, #45285, #45731), and modernised servoshell for Android to use Compose UI and Kotlin (@veyndan, #45923, #45932, #45941, #45982, #45985, #46015, #46035, #46037, #46046, #46053, #46061, #46071, #45641, #45643, #45650, #45665, #45671, #45676, #45679, #45683, #45712, #45713, #45734, #45738).

For developers of Servo itself:

  • mach try --help now lists all of the kinds of try jobs you can run (@shubhamg13, #45607)
  • mach test-wpt --update-expectations lets you run Web Platform Tests and update expectations in a single command (@TimvdLippe, #45521), rather than having to run mach test-wpt --log-raw <path> followed by mach update-wpt <path>

More on the web platform

To allow for more performant scrolling, ‘wheel’ events are no longer .cancelable unless there are one or more non-passive event listeners (@kunalmohan, #45667). Note that like in Firefox, ‘wheel’ events are passive by default.

‘dotted’, ‘dashed’, and ‘wavy’ text decorations are now continuous across element boundaries (@mrobinson, #45726).

We’ve improved the conformance of <dialog> (@skyz1, @mrobinson, #45825, #45761), <iframe sandbox> (@cychronex-labs, #45880), <input minlength> and <input maxlength> (@skyz1, #45705), CSS gradients (@mrobinson, #43945), ‘font-style’ and ‘unicode-range’ in @font-face (@Loirooriol, #45821), FontFaceSet (@mrobinson, #45390, #45382), HTML­Input­Element (@steigeo, #45416), Intersection­Observer (@jdm, #45655, #45659, #45680), new Response() (@yezhizhen, #45953), URL.create­Object­URL() and URL.revoke­Object­URL() (@yezhizhen, #45182, #45417), and ECDSA and Ed25519 in Subtle­Crypto (@kkoyung, #45833, #46017).

We’ve fixed bugs related to <input hidden> (@mrobinson, #45750), ‘animation-delay’ (@yezhizhen, #45013), ‘clip-path’ (@Loirooriol, #45468, #45373), ‘tab-size’ (@SimonSapin, @mrobinson, #45309), ‘width’ and ‘height’ (@RichardTjokroutomo, #44627), ‘box-shadow: inset’ (@Loirooriol, #45620), ‘animation­iteration’ events (@Loirooriol, #45990), ‘click’ events (@mrobinson, #45751), ‘load’ events (@jdm, #45883), ‘error’ events in Worker global scopes (@Gae24, #45829), and document­.get­Element­By­Id() (@mrobinson, #45433).

Garbage collection safety

We use a RefCell-based mechanism to store many of our DOM types in other DOM types, enforcing Rust’s “aliasing xor mutability” rule at runtime by panicking if the rule is violated. But when garbage collection happens, we need to borrow() each DomRefCell to trace the references, and this is the source of many panic bugs. To fix that whole class of bugs, we initially created CanGc, a marker type that would annotate the code paths where GC can occur, in conjunction with custom static analysis (@jdm, #33140).

With the Rust type system we can do even better, if we flip that around and require any borrow_mut() call to prove that GC can not occur by passing a NoGC marker value. We can then require that a &NoGC must be borrowed from a &JSContext (which blocks GC) and not a &mut JSContext (which allows GC), taking advantage of how Rust references work without needing any custom static analysis.

We have a large codebase that needs to be migrated in parts, so for now we’ve created the new method safe­_borrow­_mut() (@sagudev, #46050). We also need to update all of our script-related code to borrow our safe JSContext wrapper, rather than creating an owned JSContext on the spot.

This continues our long-running effort to use the Rust type system to make Servo’s integration with SpiderMonkey safer and more reliable (@Gae24, @Keerti707, @Narfinger, @TimvdLippe, @sagudev, @guptapiyush16, @ivomurrell, @kunalmohan, @skyz1, #45230, #45436, #45503, #45617, #45711, #45797, #45800, #45858, #45884, #45937, #45902, #45968, #45977, #45991, #46003, #46005, #46084, #45548, #45552, #45590, #45909, #45912, #45943, #46089, #46117, #46114, #45320, #45324, #45328, #45340, #45381, #45385, #45410, #45392, #45409, #45604, #45616, #45618, #45627, #45636, #45662, #45663, #45675, #45674, #45677, #45684, #45735, #45807, #45810, #45816, #45818, #45828, #45838, #45836, #45837, #45840, #45841, #45857, #45859, #45862, #45875, #45887, #45931, #45964, #45935, #45987, #45988, #46001, #46040, #46051, #46057, #46106, #46125, #45678, #46002, #45845, #45645, #45673, #45259, #45817, #45822, #45876, #45877, #45891).

Performance and stability

NoGC was designed to prevent dynamic borrow failures, but it also enables some performance optimisations! If we can prove that garbage collection is impossible in some part of Servo, we can often avoid rooting JavaScript objects when interacting with them within that region of code. This has allowed us to reduce overheads by over 1% in the layout process and in HTML­Collection (@Narfinger, #46092, #45582).

Our memory usage has improved, with BoxFragment now 17% smaller (288 → 240 bytes on amd64) and ShapeCacheEntry now smaller too (@SimonSapin, @mrobinson, @simonwuelker, #45183, #45496).

We’ve fixed some nasty memory leaks when reloading and in 2D canvases (@Taym95, @sagudev, @jschwe, #45455, #45261, #45414).

Speaking of which, 2D canvases now use up to 23% less power (@yezhizhen, #45301), and we now avoid rasterising the same SVG more than once (@Narfinger, @jschwe, #44805).

Servo now decodes all images asynchronously and fills image caches asynchronously, leaving script threads (web content processes) more time for other work (@Narfinger, #45542, #44483). On top of that, we’ve improved incremental layout (@mrobinson, @Loirooriol, #45411) and reduced reflows in IntersectionObserver (@jschwe, #45986).

We’ve started working on incremental updates for the stacking context tree, and as a side effect, we’ve made some layout-bound microbenchmarks up to 10% faster (@mrobinson, @Loirooriol, #45208).

We’ve also reduced allocations, copies, GC rooting steps, and other operations in many parts of Servo (@Narfinger, @SimonSapin, @mrobinson, @Loirooriol, #45506, #45969, #45940, #45760, #46090, #45335, #45413, #45511).

For several months, Frédéric (@fred-wang) has been fuzzing for Servo bugs, and thanks to his work we’ve fixed sixteen (16) crash bugs in June, affecting <iframe>, <slot>, <link onerror>, ‘animation’, ‘clip-path’, ‘content’, ‘rotate’, ‘transition’, ‘transform-style’, ‘display: contents’, ‘overflow: clip’, CSS­Keyframes­Rule, Font­Face, stop() on Window, document­.element­From­Point(), and the DOM tree (@mrobinson, @Loirooriol, @fred-wang, #46031, #46027, #46054, #46058, #46016, #46028, #46033, #45287, #45951, #45634, #45629, #46110, #46094, #45799, #45611, #45682, #45788, #45612, #45834).

We’ve also fixed crash bugs related to IPC failures, HTML­Input­Element, Range, the DevTools Debugger tab, and when servoshell is built with --features native-bluetooth (@jschwe, @Taym95, @mrobinson, @atbrakhi, @mukilan, #45311, #45619, #45765, #45513, #45702).

New contributors

A special thanks to the following people for landing their first patch in Servo:

Interested in helping build a web browser? Take a look at our curated list of issues that are good for new contributors!

Donations

Thanks again for your generous support! We are now receiving 7681 USD/month (+0.2% from May) in recurring donations. This helps us cover the cost of our speedy CI and benchmarking servers, one of our latest Outreachy interns, and funding maintainer work that helps more people contribute to Servo.

Servo is also on thanks.dev, and already 35 GitHub users (same as May) that depend on Servo are sponsoring us there. If you use Servo libraries like url, html5ever, selectors, or cssparser, signing up for thanks.dev could be a good way for you (or your employer) to give back to the community.

We now have sponsorship tiers that allow you or your organisation to donate to the Servo project with public acknowlegement of your support. If you’re interested in this kind of sponsorship, please contact us at [email protected].

Use of donations is decided transparently via the Technical Steering Committee’s public funding request process, and active proposals are tracked in servo/project#187. For more details, head to our Sponsorship page.

The Daily Front Page 14 of 30
Saturday, August 1, 2026 The Daily Front No. #260801 — Systems Desk
article

NetBSD 11.0

by jaypatelani·▲ 276 points·127 comments·blog.netbsd.org ↗

NetBSD 11.0 released!

The NetBSD project is pleased to (finally) announce the NetBSD 11.0 release!
See the release announcement for details.

If you want to try out 11.0 please check the installation notes for your architecture and download the preferred install image from the CDN. If you are using an ARM based device, obtain a netbsd-11 image pre-configured with U-Boot from the bootable ARM images page.

Please note that the various ISO images have been split into separate <700MB images for CD-ROM media and full-sized DVD images. If you are not restricted by the size limits of a CD-ROM, make sure to pick the image with "-dvd.iso" in the name. If you are using flash-based media (e.g. a USB drive), you must use the .img files rather than the .iso images. Note that they need to be decompressed first, e.g. with gunzip or 7-Zip.

If you have any issues with installation or run into issues with the system during use, please contact us on one of the mailing lists or file a problem report.

Important note about open security issues:

As you are probably aware, the number of security issues found or suspected everywhere has massively increased with the advent of AI tools. As a consequence, we can't publish a release without open issues. Instead of delaying the release further to fix them (new ones are being reported all the time), we've instead chosen to be transparent about this.

This release has been quite delayed already, since we've waited for third-party components to make stable releases so we can get the fixes. We have avoided publishing any change without making a release candidate to give users time to test it. Our release process has been streamlined and automated as far as possible, but besides the time required to build and generate checksums for every platform, manual intervention is still required (e.g. security officer signing the release hashes). The overall process is still limited by the slowest step - the time it takes to transfer every file for every architecture over the network.

The open security related pullup requests (and associated gnats problem reports) are:

All the open pullup requests will be committed to the stable branch shortly after the 11.0 release, and become part of the upcoming 11.1 release. We are currently aiming to release 11.1 within the next two months.

The Daily Front Page 15 of 30
Saturday, August 1, 2026 The Daily Front No. #260801 — Systems Desk
repository

RipGrep musl binaries occasionally segfault during very-large searches

by throwaway2037·▲ 271 points·182 comments·github.com ↗
★ 66,811⑂ 2,689 forks Rust

ripgrep recursively searches directories for a regex pattern while respecting your gitignore

Description

Please tick this box to confirm you have reviewed the above.

  • I have a different issue.

What version of ripgrep are you using?

ripgrep 15.2.0 (rev e89fff8)

features:+pcre2
simd(compile):+SSE2,-SSSE3,-AVX2
simd(runtime):+SSE2,+SSSE3,+AVX2

PCRE2 10.45 is available (JIT is available)

How did you install ripgrep?

I originally encountered this bug in the rg bundled with OpenAI Codex. That binary is byte-for-byte identical with the one in https://github.com/BurntSushi/ripgrep/releases/download/15.2.0/ripgrep-15.2.0-x86_64-unknown-linux-musl.tar.gz and I've reproduced the bug from that independently of any Codex dependency. For the analysis below, I built rg-15.2 with debug symbols included by way of CROSS_CONTAINER_ENGINE=podman CARGO_PROFILE_RELEASE_DEBUG=true ~/.cargo/bin/cross build --release --target x86_64-unknown-linux-musl.

What operating system are you using ripgrep on?

OpenSUSE Tumbleweed Linux x86_64

Describe your bug.

Ripgrep built for x86_64-unknown-linux-musl occasionally crashes with a SIGSEGV when searching very-large trees at a high degree of concurrency. The crashing line is an integrity assertion regarding heap metadata inside MUSL's mallocng, in a calloc call made from opendir. The complete backtrace is below.

What are the steps to reproduce the behavior?

Having a sufficiently large search tree seems to be essential for reproduction. Run the attached generate_repro_tree.py. This is an LLM-written program which produces a tree full of random files which mimic the statistics of the repo in which I originally encountered the bug. It will produce a tree containing roughly 20GiB of data across 1.8M files.

Then from the root of that tree, run rg in a loop, searching for some arbitrary literal string that isn't present in the tree: while true; do rg tnoheueunotshisnthukoethnsueothnsiuothonesuioseuinth; done. On my 24-core system, having enough free RAM for the search tree to fit in the kernel's block cache, it typically takes about a minute for the SIGSEGV to appear.

What is the actual behavior?

I get a coredump with the following backtrace:

#0  get_meta () at ../src_musl/src/malloc/mallocng/meta.h:141
#1  __malloc_allzerop () at ../src_musl/src/malloc/mallocng/malloc.c:384
#2  0x00007f71f8381b2d in calloc () at ../src_musl/src/malloc/calloc.c:41
#3  0x00007f71f83810f4 in opendir () at ../src_musl/src/dirent/opendir.c:15
#4  0x00007f71f835c133 in std::sys::fs::unix::readdir::{closure#0} () at library/std/src/sys/fs/unix.rs:2081
#5  std::sys::helpers::small_c_string::run_with_cstr_stack<*mut libc::unix::DIR> () at library/std/src/sys/helpers/small_c_string.rs:48
#6  std::sys::helpers::small_c_string::run_with_cstr<*mut libc::unix::DIR> () at library/std/src/sys/helpers/small_c_string.rs:28
#7  std::sys::helpers::small_c_string::run_path_with_cstr<*mut libc::unix::DIR> () at library/std/src/sys/helpers/small_c_string.rs:18
#8  std::sys::fs::unix::readdir () at library/std/src/sys/fs/unix.rs:2081
#9  std::sys::fs::read_dir () at library/std/src/sys/fs/mod.rs:68
#10 0x00007f71f8206b5c in std::fs::read_dir<&std::path::Path> (path=...)
    at /home/dfranke/.rustup/toolchains/stable-x86_64-unknown-linux-gnu/lib/rustlib/src/rust/library/std/src/fs.rs:3265
#11 ignore::walk::Work::read_dir (self=0x7f71f5bfeb20) at crates/ignore/src/walk.rs:1551
#12 ignore::walk::Worker::run_one (self=0x7f71f5bfef08, work=...) at crates/ignore/src/walk.rs:1749
#13 ignore::walk::Worker::run (self=...) at crates/ignore/src/walk.rs:1697
#14 0x00007f71f821c866 in ignore::walk::{impl#15}::visit::{closure#0}::{closure#1}::{closure#0} () at crates/ignore/src/walk.rs:1463
#15 std::sys::backtrace::__rust_begin_short_backtrace<ignore::walk::{impl#15}::visit::{closure#0}::{closure#1}::{closure_env#0}, ()> (f=...)
    at /home/dfranke/.rustup/toolchains/stable-x86_64-unknown-linux-gnu/lib/rustlib/src/rust/library/std/src/sys/backtrace.rs:166
#16 0x00007f71f8224596 in std::thread::lifecycle::spawn_unchecked::{closure#1}::{closure#0}<ignore::walk::{impl#15}::visit::{closure#0}::{closure#1}::{closure_env#0}, ()>
    () at /home/dfranke/.rustup/toolchains/stable-x86_64-unknown-linux-gnu/lib/rustlib/src/rust/library/std/src/thread/lifecycle.rs:70
#17 core::panic::unwind_safe::{impl#23}::call_once<(), std::thread::lifecycle::spawn_unchecked::{closure#1}::{closure_env#0}<ignore::walk::{impl#15}::visit::{closure#0}::{closure#1}::{closure_env#0}, ()>> (self=...)
    at /home/dfranke/.rustup/toolchains/stable-x86_64-unknown-linux-gnu/lib/rustlib/src/rust/library/core/src/panic/unwind_safe.rs:275
#18 std::panicking::catch_unwind::do_call<core::panic::unwind_safe::AssertUnwindSafe<std::thread::lifecycle::spawn_unchecked::{closure#1}::{closure_env#0}<ignore::walk::{impl#15}::visit::{closure#0}::{closure#1}::{closure_env#0}, ()>>, ()> (data=<error reading variable: Cannot access memory at address 0x0>)
    at /home/dfranke/.rustup/toolchains/stable-x86_64-unknown-linux-gnu/lib/rustlib/src/rust/library/std/src/panicking.rs:581
#19 std::panicking::catch_unwind<(), core::panic::unwind_safe::AssertUnwindSafe<std::thread::lifecycle::spawn_unchecked::{closure#1}::{closure_env#0}<ignore::walk::{impl#15}::visit::{closure#0}::{closure#1}::{closure_env#0}, ()>>> (f=...)
    at /home/dfranke/.rustup/toolchains/stable-x86_64-unknown-linux-gnu/lib/rustlib/src/rust/library/std/src/panicking.rs:544
#20 std::panic::catch_unwind<core::panic::unwind_safe::AssertUnwindSafe<std::thread::lifecycle::spawn_unchecked::{closure#1}::{closure_env#0}<ignore::walk::{impl#15}::visit::{closure#0}::{closure#1}::{closure_env#0}, ()>>, ()> (f=...)
    at /home/dfranke/.rustup/toolchains/stable-x86_64-unknown-linux-gnu/lib/rustlib/src/rust/library/std/src/panic.rs:359
#21 std::thread::lifecycle::spawn_unchecked::{closure#1}<ignore::walk::{impl#15}::visit::{closure#0}::{closure#1}::{closure_env#0}, ()> ()
    at /home/dfranke/.rustup/toolchains/stable-x86_64-unknown-linux-gnu/lib/rustlib/src/rust/library/std/src/thread/lifecycle.rs:68
#22 core::ops::function::FnOnce::call_once<std::thread::lifecycle::spawn_unchecked::{closure_env#1}<ignore::walk::{impl#15}::visit::{closure#0}::{closure#1}::{closure_env#0}, ()>, ()> () at /home/dfranke/.rustup/toolchains/stable-x86_64-unknown-linux-gnu/lib/rustlib/src/rust/library/core/src/ops/function.rs:250
#23 0x00007f71f8361fcf in alloc::boxed::{impl#31}::call_once<(), (dyn core::ops::function::FnOnce<(), Output=()> + core::marker::Send), alloc::alloc::Global> ()
    at library/alloc/src/boxed.rs:2275
#24 std::sys::thread::unix::{impl#2}::new::thread_start () at library/std/src/sys/thread/unix.rs:118
#25 0x00007f71f8388788 in start () at ../src_musl/src/thread/pthread_create.c:207
#26 0x00007f71f8389e6c in __clone () at ../src_musl/src/thread/x86_64/clone.s:22

Here is the core dump and the corresponding rg binary which produced it.

What is the expected behavior?

Not a segfault.

The Daily Front Page 16 of 30
Saturday, August 1, 2026 The Daily Front No. #260801 — Systems Desk
article

The Art of 64-bit Assembly

by 0x54MUR41·▲ 221 points·102 comments·nostarch.com ↗

Machine-Level OOP, Exceptions, and Concurrency

Download Chapter 1: Advanced Macros

You can ask an AI to explain how vtables work in x86. It will give you something that sounds right. What it won’t give you is what Windows actually expects the vtable to look like, why method dispatch behaves the way it does at the instruction level, or what breaks when you deviate from convention. This volume of The Art of 64-Bit Assembly closes the gap between a plausible explanation and genuine understanding.

Every chapter takes a construct you’ve used in C++, Python, or Rust, strips away the runtime, and rebuilds it from scratch in MASM, running under Windows. Objects, exceptions, closures, coroutines, concurrency: Each is dissected at the instruction level, with every decision made visible and explicit.

What you’ll build:

  • Object-oriented programs in MASM: vtables, method dispatch, and inheritance, from scratch by hand
  • Windows structured exception handling (SEH) installed and managed at the instruction level
  • Thunks, closures, and iterators that behave like higher-order functions
  • Coroutines, generators, and fibers without resorting to HLL code
  • Concurrent programs with real synchronization primitives, directly from assembly
  • Unicode string handling done correctly, at the level where most code gets it wrong
  • Domain-specific macro languages inside MASM, built from first principles

If you already know assembly and want to stop taking the hard parts on faith, this is the book.

Table of contents 

Acknowledgments
Introduction

Chapter 1: Advanced Macros
Chapter 2: Unicode Strings
Chapter 3: Transcendental Functions
Chapter 4: Advanced Procedures
Chapter 5: Concurrent Programming
Chapter 6: Object-Oriented Programming With MASM
Chapter 7: Exception Handling
Chapter 8: Thunks and Closures
Chapter 9: Advanced Parameter Implementation
Chapter 10: Iterators
Chapter 11: Coroutines, Generators, and Fibers

Appendix A: ASCII Character Set
Appendix B: Glossary
Appendix C: Installing and Using Visual Studio

Index

View the Copyright page
View the detailed Table of Contents
View the Index

The Daily Front Page 17 of 30
Saturday, August 1, 2026 The Daily Front No. #260801 — Linux in the Small
repository

Linux on ESP32

by boveyking·▲ 116 points·39 comments·github.com ↗
★ 119⑂ 3 forks C

Porting MMU Linux to ESP32-S31

Linux port with Sv32 virtual memory, Supervisor mode, XIP, and a Buildroot userspace, running natively on an ESP32-S31 microcontroller.

Module tested: ESP32-S31-WROOM-3 E1H16R16V (ESP32-S31 Core Board/Korvo).

Linux 6.12 booted on an ESP32-S31 development board

Experimental hardware bring-up project. Definitely not something you want for production.

Quick Start

Install esptool, Espressif's tool for flashing ESP32s:

$ pip install esptool

Then download the binaries in Release, connect your board through USB-UART, and flash the board per the provided command below (change /dev/ttyUSB0 to your actual serial device):

$ esptool -p /dev/ttyUSB0 -b 2000000 erase-flash
$ esptool -p /dev/ttyUSB0 -b 2000000 write-flash \
    --flash-mode dio --flash-freq 80m --flash-size 16MB \
    0x2000 bootloader.bin \
    0x8000 partition-table.bin \
    0x17000 ota_data_initial.bin \
    0x20000 hello_world.bin \
    0x220000 fw_payload.bin \
    0x2A0000 xipImage \
    0xA20000 rootfs.sqfs

Porting progress

General

Feature Status Buildroot rootfs 🟢 Stable Reboot and poweroff 🟢 Stable Wireless (ESP-Hosted) 🟡 Experimental Dual hart SMP ⚫ Not Planned (Used by FreeRTOS; see FAQ)

Peripheral Drivers

Feature Status AXI GDMA 🟡 Experimental AHB GDMA 🟡 Experimental Cache driver 🟡 Experimental TRNG 🟡 Experimental eFuse 🟡 Experimental Watchdog 🟡 Experimental PWM, counter, analog peripherals 🟡 Experimental CLIC/CLINT interrupt driver 🟡 Experimental Flash MTD driver 🟡 Experimental Timers 🟠 WIP Clock tree 🟠 WIP Security accelerators 🟠 WIP LP subsystem & IPC 🔴 Not Implemented PMP/APM 🔴 Not Implemented (properly)

Connectivity Drivers

Feature Status UART0 console 🟢 Stable UART1/2 🟡 Experimental GMAC Ethernet 🟡 Experimental SDMMC 🟡 Experimental GPIO 🟡 Experimental pinctrl/GPIO Matrix 🟡 Experimental USB 🟠 WIP I2C 🔴 Not Implemented I2S 🔴 Not Implemented SPI 🔴 Not Implemented RMT 🔴 Not Implemented USB Serial/JTAG ⚫ Not Planned (Used by FreeRTOS; see FAQ)

🟢 Stable — Fully tested and working | 🟡 Experimental — Seems working; not throughly tested | 🟠 WIP - Functions not fully implemented

Build/Flash Instructions

Refer to the Build Instructions.

S31 Quirks

(For more hardware references, see docs/ folder)

This port was done before S31 TRM is available, therefore these guessworks were made:

CLIC v. PLIC v. CLINT

S31 uses CLIC and CLINT similar to P4. Linux expects PLIC. Therefore a custom CLIC driver is needed. I referenced this CLIC patch from disdi to get the CLIC working.

Also, standard RISC-V interrupt CSRs are not usable, presumably because, from P4's TRM, CLINT interrupts are routed to CLIC and mtvec.MODE is hardwired to 0x3 (CLIC mode). Patches needed to make OpenSBI interrupts work.

S mode

S31's supervisor mode is not standard and has absolutely no usage in ESP-IDF so a lot of these CSR uses were mostly guessed from either P4's TRM or CSR probing (see docs/). For example, the use of sclicbase(?) and the lack of sie.

S31 implemented SCLIC (Supervisor CLIC?) which is confusing since there is no known standardization; According to all laws of esp-idf, mcliccfg.NMBITS is not writable. IT IS WRITABLE! And setting it to 0b01 enables writes to the clicintattr[i].MODE field and thus enabling the use of S-mode interrupts.

OpenSBI and Linux XIP

To save the precious 16MB PSRAM memory, OpenSBI was modified to use XIP in flash and internal SRAM (hence the 3915901 KB firmware size in OpenSBI banner, since flash and SRAM mappings are not continuous).

In mainline linux, XIP support on RISC-V was removed, so 6.12 was used instead which has proper XIP support.

FAQ

Why not SMP?

For several reasons:

  • Espressif's radio firmware blobs are closed source, and must run within ESP-IDF's FreeRTOS framework. It's near impossible to reverse-engineer them (not to mention legal risks.)
  • S31's two cores are kinda heterogeneous already: SIMD path only on hart 1. SMP makes scheduling things on the right core harder.
  • PSRAM is already slow enough (compared to SRAM); two cores would share the same, tiny 32KiB D-cache.
  • Cache maintenance, IPC, Interrupt routing, etc.
  • I like having an RTOS for other tasks. If you want absolute performance, a low-end MPU (like Allwinner T113-S3) would be a far better choice

For this port, think S31 as a reincarnated Bouffallo BL808.1

TODO

Footnotes

  1. I actually liked the BL808 and attempted to use it for a project, but the absurd lack of drivers is REAL BAD and made me appreciate Espressif's software support more
The Daily Front Page 18 of 30
Saturday, August 1, 2026 The Daily Front No. #260801 — Pocket Machines
article

But can your calculator run Linux?

by jandeboevrie·▲ 94 points·9 comments·raymii.org ↗
One of my all time favorite calculators is the HP 16C.

But can your calculator run Linux?

One of my all time favorite calculators is the HP 16C. It is an RPN calculator aimed at computer programmers with support for hex, binary, octal and decimal and all sorts of useful operations like bit shifting, masking and rotations. I don't own one myself, I have a Swiss Micro's DM16L, a high quality modern day remake with improvements like a 2 line display. There is a limited run collectors edition of the HP 16C coming soon though. But none of those calculators run Linux.

This one can:

xcalc on calc

xcalc on an actual calculator

console

Screenshot of the console and partial fbv image viewer

doom prime g2

Of course it runs Doom

HP-16C

My all time favorite, the HP-16C

Below are at least a thousand words covering the HP Prime G2 and other graphical calculators capable of running Linux. You can skip that if you just want to read about my updated Linux port for the Prime G2.

HP Prime G2

The HP Prime is a graphing calculator on the market since 2013, with a hardware revision in 2018 (the G2 model). It has a touch screen and as far as I can find, most people are happy with it and find it a very capable and fast calculator. Here are some pictures of the opened up calculator and here is a website with some more downloads.

I'm not really interested in the calculator part. For most calculations I do for my day job, embedded programming, the HP-16C is a better fit. I was interested in this device because I found posts online suggesting it could run Linux and even a Windows 10 port.

Bow for me for I am root

I do have a history running Linux on older HP devices and I do like to have full privilege on all devices I own. Back then I wrote:

I have a rule that I try to adhere to for devices I own. They must allow a means of root or administrative actions. The Nintendo Switch I got second hand is old enough to be jailbroken. The two Apple devices I own (a first gen iPad Air and first-gen iPhone SE) have vulnerabilities that give me root access. All Android phones I owned I've specifically bought because the bootloader can be unlocked. Once I buy a device, it is mine and I decide what to do with it. Not the manufacturer. Otherwise it is e-waste the moment it leaves the factory.

Being able to develop software on a device is almost a must have as well for me. Missing functionality, otherwise known as stuff the manufacturer does not make (enough) money on, can be programmed back in as long as you can develop on a device and are willing to put in enough time and/or money.

Now with a calculator that isn't a huge issue since I don't see them as general purpose computing devices. Although Texas Instruments did later on removed and restricted features that their device shipped with initially, namely running assembly code on your calculator and Numworks also locked down their calculators. That all has to do with cheating and exam mode, so I do understand that part. If your device is not allowed on official exams / tests, it won't be used or sell well. On the HP Prime there never was assembly (only a custom basic like language called PPL) so they can't take that away. There is a quote from an HP employee stating:

"It is not locked down and is just used for verification at this time. Note however, we DO have the ability to fully encrypt and lock the system to hell as that is part of the new chip. We DO have that tested and working. We WILL turn it on if needed. It IS a corporate security mandate (after things like firmware being loaded into mice/keyboard or printers to hack networks came to light) that we fought to get an exemption for. Nobody dick around with exam mode."

Hardware specs for a calculator!?

The Prime G2 is comically overpowered for a calculator if you ask me. It has an i.MX 6 Ultralite ARM Cortex A7 CPU, 256 MB DDR3 RAM and 512 MB of internal storage. I mean come on! It's closer to a small embedded Linux machine with a calculator keyboard attached than to a traditional calculator.

The G1, the previous hardware revision already had a 400MHz ARM CPU with 32MB ram, but the G2 is like an Android phone of a few years back.

The coffee machine hardware I work on at my day job has an i.MX6 module with comparable specifications. And that runs dedicated C++ and Qt applications on top of Yocto. The HP Prime G2 runs on FreeRTOS and the G1 runs on a custom RTOS named Besta OS.

I'm still shocked by the specifications, for a calculator! In daily usage it is very responsive and compared to other calculators, calculations are way faster, graphing is faster and via the touchscreen you can zoom/scale graphs, which is also almost instant. So they do put those specs to good use at least.

Hardware specs for other calculators

Let's check what their competitors have regarding specifications and possible Linux support. I've tried to find specifications of a few models available here in The Netherlands.

The TI 84+ CE has an 48 MHz eZ80. No linux there.

The Numworks Graphing Calculator N0110 has a 216 MHz ARM Cortex-M7 CPU, 8 MB storage and 256 KB of SRAM. I still find 216 MHz to be way more megahertz than you'll probably need for a calculator, but who am I to judge. The later N0120 model even has a ridiculous 550 MHz ARM Cortex-M7, but a more reasonable 564KB of RAM. Then we're talking microcontroller levels of RAM. No linux here, but the OS is open source. There are even a few forks, but it seems that later firmware versions are locked down. Sad for enthousiasts, but probably it has something to do with cheating and exam mode.

The Casio FX-CG50 doesn't list cpu specifications but this post says 118 MHz and this post says the type, a custom Renesas SH7305 CPU, 8 MB of RAM and 32 MB of ROM. No linux there either. Although, maybe this guy can make it work.

The only other calculator I found with a linux port is the TI Nspire CX (II) (CAS). Here is an article on running Arch Linux on it and here is more info. The CX-II has a 396MHz ARM CPU with 64MB RAM and 128MB of flash storage. Not as much as the HP Prime, but still Linux capable.

My favourite calculator, the HP-16C, is a different story. This is not a competitor because it has no graphing features and its not exactly modern. The NUT CPU detailed description (A-1LF5-9002-1 - 7/14/81) states 200 to 230KHz for the 11C/12C and 340 to 380KHz for the 41C. There is a definition in the Nonpareii emulator stating 215 kHz, a post here by Nelson M. Sicuro from 2003 suggesting about 230 kHz and this page, stating 220 kHz. And 203 bytes of storage for programs. No Linux there sadly.

I don't have a big calculator collection nor am I an expert on calculators. I used them in high school and further in my education whenever math was involved. I stil have a TI-83 somewhere from those days. And a HP 11C, the non programmer, regular version of the HP 16C. I like that the 11C is RPN. For math, that just clicks better in my head.

HP Prime G2 Linux

For the G2 there is an old unsupported port with a 4.14 kernel. The linked linux repo no longer has the correct branch, buildroot failed, the devicetree file was messed up and the provided download doesn't show a console login, so I couldn't really use it. But it provided enough info to get started. I've updated the port for my HP Prime G2, with some enhancements:

  • Make the keyboard better, you couldn't enter numbers or characters like pipe, a dash or an underscore.
  • Shows up as an USB serial device
  • Runs in RAM but allows a larger than 15MB image.
  • My unit has a newer touchscreen controller (Ilitek ILI211X, due to the old Goodix GT5688 being out of production), made a kernel driver for it.
  • You can login to the console
  • The X11 server works and has a few basic apps
  • Ships with Doom (prboom) at a great framerate
  • Ships with a C compiler for on device development

After reconstructing the repo's using software heritage I was able to get a build up and running. I only tried running this in RAM, not wanting to flash the NAND and potentially loose the original calculator software.

hp-prime-g2-linux-console

The keyboard in Linux was only sending the orange marked letters, no dashes, underscores, numbers or a pipe character. At first I made an alternative keymap which loadkeys configured at boot. However, that didn't work in X11, so I just patched up the kernels keyboard driver (imx_keypad.c) to send the correct keys. The patch intercepts keycodes and if the ALT key is pressed it sends a different keycode. See input-event-codes.h for the mapping.

Using the Alpha key (ALT) you can enter numbers and some other special characters. Pressing the ALT key together with e.g. COS or TAN failed to send a keycode. I didn't look into that further, just used other keys. It might have to do with the keyboard matrix.

The original build's uuu scripts could boot Linux in RAM, but only a 15MB max build. By changing offsets in the kernel configuration, uuu script and u-boot you can now run an image up to 130 MB from RAM.

The original port did not boot up to a login screen but that was just a case of configuring getty in the init scripts.

The original build and a few articles in Russian (here, here and here) had a working touchscreen using a Goodix kernel driver, but in my build that did not work. It claimed to be on i2c bus 1 address 0x14 but i2cdetect showed no such address. It did show a mystery device at 0x26, which while poking around showed a data structure changing while touches were sent.

After much fiddling and looking at raw i2c messages I found out that it looked comparable to an Illitek ILI211X device. Newer kernels have a driver and the ILI211X_DATA_SIZE matches my observations for a 43 byte data packet on i2c, a 0x5a start byte and the 0x26 address matched with another imx6 devicetree with that driver.

The HP recovery screen showed the touchscreen fw version as ilitek-03 which also was a hint. This old kernel has no driver for that model, only for older ones with a different protocol. I coded up a kernel driver based on the newer kernel driver, but stripped out all non-essentials, for example multitouch and finger pressure. I also made a calibration tool fbtouch_test since my screen had some oddities in the bottom right corner. Later I found a french forum post that stated that HP switched screens due to the Goodix ones not being in production anymore.

hp prime touch

Working touchscreen

During debugging I found that modprobing the kernel module for usb serial support locked up the entire device. I changed that from a loadable kernel module to compiled in the kernel, which gave no crash anymore. I can now connect via serial to my calculator.

With touch support working I could also include X11 and a few programs. Once you're logged in on the console, type startx and twm will start together with xeyes, xclock and xcalc. Mouse and keyboard inputs work. Have fun running xcalc on your calculator. A long press is interpreted as a right click.

For fun I added tcc, the tiny c compiler. Because what good is a calculator if you can't compile your own code on it? There is a small sample hello.c file in /root which you can compile (tcc hello.c -o hello) and run (./hello) yourself.

The source code is available on my github:

For convinience I've made a pre-built download you can directly run without compiling the entire thing. buildroot and linux take quite a while to build and might bitrot over time. With the pre-built download you can get started right away playing.

Instructions for running Linux on the HP Prime G2

Beware that you need to open up your calculator and you might damage it, unrecoverable, not being able to boot into the HP software. Continue with caution.

Download the pre-built package from HERE and extract it to a folder. Download uuu from here.

Open the calculator (remove 4 screws from the back, open the housing with a small pry tool or guitar pick), plug it in via USB and press the RESET button while shorting 2 pads marked on the photo:

hp-prime-g2-reset-pads

I've used an iFixit tweezer to short those two pads, they're tiny.

Use the uuu script to boot Linux in RAM:

uuu ./run_linux_in_ram.uu

After a few seconds you should see a login prompt. You can login as root without a password. Type startx to start the graphical environment. Type sh ./doom.sh to start Doom.

I have not tested any of the NAND flashing scripts so use those at your own risk. The Goodix touchscreen should still work, but I cannot test that. If you have a unit with that type of screen, please test and let me know.

HP Prime G1 Linux

For the Prime G1 there is a forum topic and screenshot here and this port has actual source code and a modern 6.x kernel. I don't have G1 hardware so I cannot test this version. The port provides no downloadable image, you have to build one yourself. The instructions are on the page, but a bit tense. Note that this is specifically for the G1 hardware version. I did try to build this version myself to see if there was much bit rot or broken links, but surprisingly it all went quite smooth still.

First clone all 4 repo's:

mkdir prime-linux
cd prime-linux
git clone https://github.com/Repeerc/Linux-For-HPPrime-V2
git clone https://github.com/Repeerc/hpprimev2_linux_loader
git clone https://github.com/Repeerc/Kernel-6.1.35-HP-Prime-V2_G1
git clone https://github.com/Repeerc/buildroot_hpprimev2

Build the loader:

cd hpprimev2_linux_loader
mkdir build
cd build
cmake ..
make
cp BOOT1.ROM ../../

If you receive an error like error: unknown type name 'caddr_t', replace that type with void in the code. Modern compilers don't support that anymore.

Build the kernel:

cd ../../
cd Kernel-6.1.35-HP-Prime-V2_G1
make ARCH=arm CROSS_COMPILE=arm-none-eabi- hpprimev2_defconfig
make ARCH=arm CROSS_COMPILE=arm-none-eabi- zImage -j8
cp arch/arm/boot/zImage ../../

Build buildroot:

cd ../../
cd buildroot_hpprimev2
make hpprimev2_defconfig
make
cp output/images/rootfs.jffs2 ../../

Package it all up together:

cd ../../
cp BOOT1.ROM Linux-For-HPPrime-V2/
cp zImage Linux-For-HPPrime-V2/
cp rootfs.jffs2 Linux-For-HPPrime-V2/
cd Linux-For-HPPrime-V2/
bash ./mkimg.sh

Flash BOOT1.ROM and LINUX.DAT to the calculator:

Use usbtool.exe, connect the calculator in Recovery Mode (RESET + Symb) , select Auto update then click Update.

There are documents online suggesting this won't work with a USB3 port.

Here are a few screenshots I found online.

Prime G1 screenshot from the 2017 forum topic:

prime g1 screenshot

Prime G1 running an X session:

Prime g1 x11

Prime G1 also has a Doom port:

prime G1 doom

Windows?!

If you're that kind of person, someone ported UEFI for Windows 10 (source). Check out the most meta picture ever, Windows calc.exe on an actual calculator:

calc.exe

I cannot find any downloadable image and some links on the blog post to twitter are deleted so I cannot test this build myself. There are a few more screenshots:

Failed bootloader:

bootloader win

The boot screen:

bootloader 2

The print dialog:

win print

Notepad:

notepad

The Daily Front Page 19 of 30
Saturday, August 1, 2026 The Daily Front No. #260801 — The Absurd Ledger
article

The Absurdity of Albert Camus

by apollinaire·▲ 180 points·88 comments·historytoday.com ↗
For Albert Camus understanding the past and predicting the future may hinge on one fact: history is absurd.

For Albert Camus understanding the past and predicting the future may hinge on one fact: history is absurd.

Albert Camus, 1947. Roger-Viollet/TopFoto.

In August 1944, as the French Resistance fought to liberate Paris, Albert Camus went to see Jean-Paul Sartre at the Comédie-Française. It was a perilous moment. Now that the German forces were on the back foot, Hitler had ordered his troops to inflict as much damage on the capital as possible. Sartre had been sent to defend the Comédie-Française – Paris’ oldest and most prestigious theatre – from sabotage. But the journey across the city had worn Sartre out and by the time Camus arrived he had fallen asleep in the stalls. Camus couldn’t help being amused. Waking Sartre, he quipped: ‘You have turned your seat in the direction of history.’

Camus had known Sartre for a little over a year by then. They had first met at the premiere of Sartre’s play, The Flies, on 3 June 1943 and had soon become friends. They were, in some ways, an unlikely pair. Whereas Sartre came from a respectable bourgeois family and had been educated at the elite École normale supérieure, Camus had been born into poverty in French Algeria. Raised by his mother – an almost illiterate widow – and his tyrannical grandmother, he had cut his teeth in journalism and the theatre before finally moving to Paris shortly before the outbreak of war. But they were bound together by a mutual respect and a shared philosophical outlook. So close had they become, in fact, that Camus had signed Sartre up as a correspondent when he became editor of Combat, the newspaper of the French Resistance, a few months after their meeting; and Sartre, in turn, had even asked Camus – no mean actor – to star in his play No Exit.

Camus’ quip in the Comédie-Française was meant in good spirits. It was, in all likelihood, a playful reminder that, while the dashing Camus had found it easy to throw himself into the Resistance, Sartre had struggled. But Camus’ joke also highlighted a more fundamental difference between them – one which centred on their views of history, and which would soon turn them into bitter rivals.

The absurd

Their disagreement had its roots in Camus’ concept of the ‘absurd’. As he had explained in many of his early writings, the absurd describes a basic contradiction at the heart of all human existence: between man’s innate desire for reason and clarity on the one hand, and his awareness of the irrationality and meaninglessness of life on the other. Not everyone becomes conscious of this in the same way, of course. It can ‘strike [you] in the face’ at any moment, ‘on any street corner’. But once you have seen it, its effect is profound. Familiar habits, routines, assumptions, and beliefs suddenly seem pointless. You start to feel alienated from the world around you. A stranger in your own life.

How, then, should one respond to the absurd? Early in his career Camus became aware of, and in some senses, a spokesman for, a growing sense of nihilism in contemporary culture. He recognised that, confronted with the meaninglessness of life, a person might easily be tempted by hedonism, faith, despair – even suicide. As time went on, however, Camus came to realise that each of these responses is in some way inadequate. Each seeks to liberate the self from meaninglessness by denying the absurd. While that might seem sensible at first, it turns out to be self-defeating. Since we can only perceive the absurd because we are rational creatures, it is impossible to deny the absurd without also denying that which makes us human.

A more satisfactory response, Camus realised, was to embrace the absurd. As he wrote in the newspaper Alger-Républicain in 1939, it should be a beginning, and not an end. This was less paradoxical than it sounds. If existence is meaningless then clearly it is pointless trying to make it otherwise. The only rational thing to do, Camus argued in The Myth of Sisyphus (1942), is to accept that life is uncertain – to accept, in other words, that death is inevitable, love fickle, and fame fleeting – and live accordingly. ‘Everything’, in principle, ‘is permitted.’ Since the world obeys no rhyme or reason, all actions are morally equivalent. The key is to live fully – to recognise that we are all in the same boat, to savour the impossible struggle, to accept the inevitability of defeat or disappointment. In short, to be a ‘happy Sisyphus’.

Camus dramatised this in his novella The Stranger (1942). Following the death of his mother the protagonist, Meursault, is acutely conscious of the absurd. But he is nevertheless committed to the truth. He feels no grief at his mother’s funeral; he is indifferent to his lover; and he refuses to express regret at having shot an Arab on the beach. This alienates him from society. Indeed, it is his detachment from life’s absurdity that ultimately condemns him. But it is what liberates him, too. He realises that his sense of estrangement is the only thing he truly shares with other human beings. And, in the end, as he faces the inevitability of his execution, he sees that he is happy.

Rebellion or revolution?

This had profound implications for politics, too. In The Rebel (1951) Camus argued that, just as man is tormented by the contradiction between his desire for clarity and the meaninglessness of existence, so he also has an innate impulse to rebel against anything which he perceives to be unjust, or which conflicts with the intrinsic value of human life and liberty.

It was at this point that history made its entrance – and that Camus and Sartre came to blows. Much like Camus, Sartre agreed that rebellion was a natural response to injustice. He saw it as a ‘rational’ reaction against the irrationality of the world. But he nevertheless understood it as a component in a historical process. To be sure, he did not think that human societies followed a predetermined path towards the creation of an ‘ideal’ society. He disagreed quite strongly with Karl Marx on this. As he later argued in Critique of Dialectical Reason (1960), there was no inevitability to history. People could not know how it would work out. But if they acted as if they did, they stood a decent chance of making the Communist ideal a reality.

Camus was not insensitive to this view at first. He had joined – and been expelled from – the Communist Party in Algeria long before Sartre had begun his dalliance. He had even run the Party’s Théâtre du Travail. In his letters he had also spoken about history as having a certain directionality about it. But the experience of the Second World War, and his encounter with Nazism, had changed his mind.

The key, Camus believed, was to differentiate between rebellion and revolution. As he saw it, revolutions sought to protect freedom and humanity against injustice by creating a perfect social order – in other words, by putting an end to history. The problem is that in trying (and failing) to create an ideal society, revolutions end up crushing the freedom and humanity they purport to be defending. This was especially true of the French Revolution, but no ‘revolutionary’ regime was any better – or worse – than any other. In this respect, Camus believed, the Soviet Union was indistinguishable from Nazi Germany.

Rebellion, by contrast, should respond to injustice by emphasising a shared humanity and by clinging doggedly to the value of human life. It required not just individual commitment, but solidarity with others. As Camus pointed out in Neither Victims, nor Executioners (1946), it should not aim to end history, but to ‘fight within History, to preserve from History that part of man which is not its proper province’.

Camus gave vivid expression to this in The Plague (1947). Set in Oran, in Algeria, during an outbreak of the plague in the 1940s, the novel has been widely recognised as an allegory of the German occupation of France and is, in many ways, a blueprint for his concept of rebellion. As Ronald Aronson has argued, it ‘conveys the unheroic determination to do what must be done in the face of a total threat’. Although the journalist Rambert, for example, initially tries to flee, he eventually resolves to stay and – realising that the plague can only be combatted through solidarity – accepts the risks that entails.

Blinded by history

Sartre was appalled by Camus’ view of rebellion. He saw it as an undisguised attack on Communism. At his urging, his secretary, Francis Jeanson, wrote a furious response to The Rebel. A bitter war of words ensued, centred on the nature of history. Whereas Jeanson argued that Camus had effectively denied any role for history, Camus responded that he had done nothing of the sort. He had merely tried to draw attention to those ‘whom history blinded to present suffering’ – a sly dig at Sartre’s narrow-minded refusal to condemn Soviet work-camps. Never shy of a fight, Sartre hit back. He alleged that Camus – himself blinded by arrogance and indifference to justice – was pretending that he could step outside history. He even suggested that Camus saw history as a threat to humanity itself. And so it went on.

The controversy was briefly overshadowed in 1957 when Camus won the Nobel Prize for Literature. Though he accepted the award with characteristic modesty, his success was met with considerable criticism in France because of his surprisingly cautious attitude towards the Algerian War, by then entering its third year. Refusing either fully to support Algerian independence or to back the French government, he instead equivocated, calling for peaceful co-existence and an ill-defined truce.

But the bitterness of Camus’ dispute with Sartre never went away. He was already writing another, more detailed defence of his views when, on 4 January 1960, he was killed in a car accident with his friend and publisher, Michel Gallimard, just outside Sens, in northern France. As he told a friend only a short time before, he felt as if his real work had not yet begun. Whatever role history might play in rebellions, it had proved Camus right. Death – like life itself – is absurd.

  • Born 7 November 1913, Mondovi, French Algeria
  • Died 4 January 1960, Villeblevin, France
  • Notable works The Stranger (1942) | The Myth of Sisyphus (1942) | The Fall (1956)
The Daily Front Page 20 of 30
Saturday, August 1, 2026 The Daily Front No. #260801 — The Holdout Corner
article

The tiny holdout building in the middle of Macy’s is back in view

by donohoe·▲ 213 points·63 comments·ephemeralnewyork.wordpress.com ↗
For the first time since the early 1900s, the five-story holdout building is back in view.

Hidden by billboards for over 100 years, the tiny holdout building in the middle of Macy’s is back in view

The enormous faux shopping bag that served as an advertisement for Macy’s is mostly gone, save for a patch of red facing Sixth Avenue.

Macy's department store building in New York City, featuring a mix of modern and historic architecture, with a busy street scene and traffic lights in the foreground.

Scaffolding covers the ground floor retail space, and dark netting swaths the brick and limestone upper floors.

What’s going on these days at the northwest corner of 34th Street and Sixth Avenue?

For the first time since the early 1900s, the five-story building that prevented Macy’s flagship store from owning the tip of the northwest corner of Herald Square is having its billboards dismantled.

Historical building with a partially renovated facade, showing a contrast between the old brick structure and modern construction scaffolding.

Finally, the little holdout—a lovely if grimy structure, with large studio windows and terra cotta detailing—is once again a visible part of the cityscape.

This building might be the most famous holdout in Manhattan, and it serves as an illustrious reminder of Gotham’s bitter retail wars at the turn of the 20th century.

Historic black and white photograph of Macy's department store in New York City, showcasing streetcars and pedestrians in the bustling urban environment.

Macy’s new store and the corner plot, early 1900s

The story begins around 1901. At the time, Macy’s made plans to move its New York City stores from the crowded Ladies Mile shopping district along 14th Street between Fifth and Sixth Avenues to the more spacious and transportation-friendly Herald Square.

They planned to build the new store at 34th Street and Sixth Avenue. Macy’s bought out many smaller businesses that already existed on the site—including a restaurant, a barber shop, and the bawdy Koster & Bial Concert Saloon, where the first projected motion pictures made their debut.

Historic black and white photograph of the exterior of Macy's department store, showcasing its large facade and architectural details.

The holdout building in 1906

The only lot they couldn’t get was a 30 by 50 foot corner remnant of an old farm that held an existing structure. “This tiny parcel was the property of an old-time New Yorker, Rev. Duane Pell,” wrote Edward Hungerford in his 1922 book about Macy’s, The Romance of a Great Store.

“It was given to understand that [Pell’s] asking price for the small corner was $250,000, an astonishing figure for such a tiny bit of land . . . but Dr. Pell felt that he had held the key to the entire important Herald Square corner and that he was justified in asking any price he saw fit.”

Despite the high price, Macy’s ultimately agreed to buy the plot from Pell, as they felt it was crucial to own the corner.

Historic black and white photograph of the Hippodrome Theater building in New York, featuring banners for performances and advertisements, with bustling street activity and vintage vehicles in the foreground.

The holdout building covered in ads, 1907

Yet before Macy’s could sign the deal, Pell sold the plot for $375,000 to an agent for Henry Siegel, the owner of rival department store Siegel-Cooper, according to New York’s Architectural Holdouts, by Andrew Alpern and Seymour Durst.

Known as “The Big Store,” Siegel-Cooper was a massive emporium with 120 departments and 3,000 employees on Sixth Avenue and 18th Street.

A busy street scene in front of Macy's department store, featuring a large sign advertising 'The World's Largest Store.' Yellow taxis and pedestrians are visible, along with tall buildings in the background.

Macy’s billboards in 1964

Opened in 1896, Siegel-Cooper occupied a section of Sixth Avenue near other legendary department stores of the era, like B. Altman’s and Hugh O’Neill’s.

Once the plot was transferred to his name, Siegel cannily offered it to Macy’s in exchange for one of their 14th Street stores, which Macy’s planned to leave vacant after the move to Herald Square.

2026 image from the New York Post

Macy’s emphatically turned Siegel down. Instead, they built the Herald Square store around the holdout corner, a spiteful move that suggested Macy’s power and dominance.

“In 1903 Siegel-Cooper, defeated after Macy’s had opened the year before, demolished the small structure that occupied the corner and constructed a five-story building, designed by William Hume,” states Urban Archive.

Close-up view of an old, weathered building facade with boarded windows, decorative architectural elements, and utility pipes visible.

At some point in the early 1900s, Siegel lost ownership of the five-story building. Since the 1920s, Macy’s has leased billboard space on the facade and plastered it with Macy’s ads, including the ad for the red and white shopping bag.

So what prompted the holdout building to emerge out of the shadows now?

Macy's department store exterior with a prominent red sign and scaffolding, located at a city intersection with traffic lights and American flags.

Apparently the owner of the building for more than 60 years, Kaufman Realty, planned to lease the facade space to another company. A New York Post story from 2021 describes the company as a “prominent online retailer.”

When, or if, the new billboards make their appearance on this 34th Street corner isn’t known. Get a closeup view of this long-hidden holdout while you can—an emblem of how powerful corporations are willing to fight for precious real estate in Midtown Manhattan.

[Third image: unknown; fourth image: MCNY, X2010.7.1.1948; fifth image: Wikipedia; sixth image: MCNY, F2011.33.428; seventh image: Jimin Kim/SOPA Images/Shutterstock via New York Post]

The Daily Front Page 21 of 30
Saturday, August 1, 2026 The Daily Front No. #260801 — Marks on the Page
article

Manual: •.,:;…!?·

by behnamoh·▲ 123 points·31 comments·type.today ↗
Glyphs much more simple in terms of design and use mostly fly under the radar.

Manual: •.,:;…!?·

Dots and commas, reading numbers in English, distinguishing between bullet, interpunct, and multiplication operator

Quotation marks, braсkets, and various kinds of spaces are a multi-faceted and controversial subject. But apart from those, there are glyphs much more simple in terms of design and use — because of their ubiquity, they mostly fly under the radar. In this article, we continue to talk about basic punctuation characters.

Period

In European and many other languages the period marks the end of a sentence. The symbol is also used for marking initials and abbreviations, and in some regions — as a decimal separator, or a separator of hours and minutes when registering time. Typically, the period is not used in the last sentence of a headline.

Design, kerning, spacing

Depending on the typeface’s design, the period can be rounded, rectangular, rhomboid, etc. The symbol correlates to the dot above i, but is normally larger.

The period is a frequent symbol, therefore its shape affects the overall intonation of a typeface. The most principled designers can discard an (otherwise high-quality) typeface, if they are not satisfied with its periods.

periodetc1-01

Periods of different shapes and sizes in static sans serif fonts

In some legacy typefaces, the period has asymmetrical sidebearings — in this case, the space after the period would be excessive, you might need to track it. Today, it is customary practice design period with equal sidebearings — because the cases of its use are not limited to traditions.

In monospaced fonts, the period glyph has the same width as all the others — that means very wide sidebearings, which result in the widest space at the end of a sentence. Whether it needs to be compensated or not — that depends on your task.

periodetc2-23

Period + space in monospaced fonts

Quality typefaces have period, comma, and other punctuation kerned, when paired with certain glyphs — to avoid excessive white holes in the setting.

periodetc-26

Use in text

According to typesetting rules in Russian language, initials shall be spaced, however in most cases regular spaces after periods are too much, which is why it is considered good manners to use thin spaces.

In any case, this situation requires a non-breaking space — surname and name should not be separated by line break.

periodetc-25

The difference between space and thin space can be more or less significant (depending on font, and point size), but a thin space usually looks neater with initials

The period is used as a decimal symbol in English language — as well as in the language of former British colonies, in China, Japan, and some other countries.

periodetc5-05Period as a decimal separator — in the UK, the US, Australia, China, Japan, Switzerland, etc. periodetc6-06Russian language uses a comma as a decimal mark — as is the case in nearly the entire South America, Europe, Indonesia, Vietnam, etc.

The period as a decimal separator is the default setting in PCs and other kinds of gadgets, because they originated from the US. In most cases, the American notation system can be changed to a local one, but will keep reminding of itself in code and many UIs.

google

Google search results: the built-in calculator uses the period, while the link to Russian Wikipedia uses the comma

The period can be used for separating hours and minutes instead of a colon, but it is rather personal preferences or editorial policy than an established practice. For example, in Britain The Guardian writes 13.12, while BBC prefers 13:12. In Russia, the period in indicating time is used by certain online media — for instance, Afisha.Daily uses colon for indicating the time of publication, yet utilises a period between hours and minutes within texts. In some countries (such as Switzerland or Latvia) the period is considered standard and more preferable when registering time.

Comma

The comma separates parts of a sentence. In English-speaking countries the comma is used as a thousands separator, while most European countries prefer spaces — and the comma is used as a decimal separator. To avoid miscommunication, do keep in mind these localities.

periodetc-27

Design

The comma can be constructed either as a circle with a stroke coming of it — or as one curving or slanted stroke.

periodetc8-07

Commas in type.today’s fonts

In most typefaces, the comma has identical shape to apostrophe and quotation marks. The diacritical comma in Latvian and Romanian often corresponds to the regular comma — that said, diacritics would be significantly smaller. As opposed to diacritical commas, the design of Сzech/Slovak caron Caron, or háček (ˇ) is a diacritical mark used in Slavic, Baltic, Uralic, and some other languages. In Czech and Slovak narrow letters with ascending elements (l d t L) can also have caron — this situations requires a special form, which is called Czech/Slovak caron. must differ from that of the regular comma and apostrophe. Replacing Сzech/Slovakian caron with an apostrophe is incorrect and misleading.

periodetc9-22

Examples of correctly correlating commas and Czech/Slovak carons

Ellipsis

The ellipsis is a separate glyph consisting of three dots. It is used for indicating omissions, pauses, breaks, or interruptions.

  1. After a question/an exclamation mark one should not put three dots (regular type of ellipsis), but two (the third dot is placed under one of symbols mentioned above): For how many years will I be in this world?.. (Tv.); How you played last night!.. (Ostr.)

    2. When an ellipsis and a comma meet, the latter is absorbed by the ellipsis that indicates not only an omission of words, but also an omission of a punctuation mark: His wife… though, they were completely satisfied with each other (G.). Dietmar Rosenthal Handbook of Spelling and Style

In omissions and lacunae the ellipsis is placed inside brackets, regular or square, or pointy ones — this depends on the typographic tradition.

The ellipsis as a separate glyph can be looser or more condensed than three dots in a row. In certain software, three typeset dots are automatically replaced by one ellipsis — provided the typeface includes this glyph.

periodetc-28Fonts with significant differences between three dots and ellipsis. Usually, the only difference is spacing, but the shape can also vary sometimes (see Alverata) periodetc-29Fonts with ellipsis spaced tighter than three dots in a row. This is inevitable in monospace fonts, where each glyph has exactly same width

Bullet and Interpunct

The bullet is a graphic marker, while the interpunct is a punctuation mark. Both glyphs are typically designed as vertically centered dots, but they may differ in size and sidebearings.

periodetc-11

Normally, the bullet is larger and thicker than a period; it is used to introduce items in a list. The bullet symbol may take a non-dot shape, sometimes it even comes in several options — this depends on the type designer’s preferences. Apart from circular, bullets may be hollow, square, diamond, etc. We recommend looking into the glyph palette to see the entire range of bullets.

The interpunct is typically of the same size as the period, and is placed in the visual middle of the x-height. Depending on language and situation, the interpunct is used for interword separation, as a mathematical symbol, as a part of geminated l Ela geminada (The geminated l) is written as ŀl and is different in meaning and sound from a simple doubled ll. The typesetting of the geminated l generates debate and discussions: as an integral symbol it has no value in Unicode, while Ldot and ldot do have one. Most frequently the interpunct is used instead of the dot, many type designers create localised features for correct spacing — however, the problem of displaying the integral, not breakable symbol remains. Sometimes the interpunct is replaced by a regular period, placed between the letters on the visual middle, but in certain fonts this option would look too heavy. in Catalan, etc.

periodetc12-12

Bullets and interpuncts: the bullet is for lists, the interpunct is a multiplication sign, or part of ela geminada in Catalan, or a symbol separating graphemes in Occitan

Often, an interpunct or a bullet are used for separating parts in a sentence — for example, for separating different languages on one line. In such cases, it doesn’t really matter which glyph you use — as long as it does the job. The interpunct as a separator may look neater, while the bullet would be more visible.

periodetc13-15

The Unicode table contains several items related to bullets and interpuncts, which scares both designers and font users. In addition to basic Bullet (u+2022), there is also Bullet Operator (u+2219). Apart from regular interpunct Middle Dot (u+00B7), there is Dot Operator (u+22C5) and a number of local variations, such as Greek Ano Teleia (u+0387), and Word Separator Middle Dot (u+2E31). The latter is almost never present in fonts. The Greek Ano Teleia serves as the semicolon, is normally as large as an interpunct, but placed higher — these symbols should not be confused.

periodetc14-24

Ano Teleia is placed significantly higher than the interpunct

The Unicode Consortium recommends using operator glyphs in all math settings. Those may look different: a Bullet Operator might be intended as a multiplication sign, while in other cases it correlates to a regular bullet. Sometimes there is no pattern whatsoever, it all comes down to the decision of a type designer.

In different fonts, availability and shape of these additional symbols differ, while their use is discussed both in type design communities and on other user forums.

periodetc15-13Apart from different sizing, some fonts have different sidebearings when it comes to bullet and interpunct derivatives periodetc16-14Bullet Operator may have the shape of a regular bullet, or an interpunct — or may correspond to neither

Two-piece symbols: Colon, Semicolon

In French, the two-piece punctuation glyphs are spaced with a regular or thin space. Depending on the situation, some spacing of these symbols might be necessary in other languages — for example, when setting in all-caps.

Time

According to GOST technical standards, Russian language uses the semicolon to denote time:

3.5.2 Elements presenting dates and time of the day are placed one after another without spacing.

Where needed, one may use the following symbols as separators:

- (dash) for separating the elements ‘year’ and ‘month’, ‘year’ and ‘week’, ‘year’ and ‘day’, ‘month’ and ‘day’, ‘week’ and ‘day’ as well as indicating omitted elements;

: (semicolon) for separating the elements of the time of the day (‘day’, ‘minute’, ‘minute and second’);

/ (slash) for separation in presenting the period of time;

# (number sign) for separating within the recurring period of time, periods of time, and recurring factor. GOST ISO 8601–2001 Presenting dates and time. General requirements

However, using period for registering time or date should not be considered mistake in non-crucial situations.

Lifting the glyphs

Normally, the colon and the semicolon are not case-sensitive — the semicolon is fixed to the baseline due to its comma, and the colon usually follows this lead. Lifting just one of the pair looks quite weird:

periodetc17-16

SF Pro has a case-sensitive colon. But if you have a semicolon next to it, this colon will offend the eye

The situations requiring a lifted colon are quite rare. The most common are denoting scale or time with lining figures. In such cases you can also adjust the colon manually.

periodetc18-17

If scale is denoted in a large headline (as opposed to body text), the manually lifted semicolon would look neater

Exclamation Mark, Question Mark

Exclamation and question marks are most commonly designed of the cap-height. Like the other two-piece symbols, question and exclamation marks are spaced out in French. The exclamation mark is typically symmetrical and upright, while the question mark, depending on font and the face, may require additional kerning pairs.

periodetc19-18

If there is no kerning in certain glyph pairs, they might require a thin space between

The inverted exclamation mark and question mark are used in Spanish and some other languages, they can be lifted when setting in uppercase, if there is an applicable Opentype feature.

periodetc20-19

The 1960s saw an attempt to introduce the so-called interrobang — a combination of exclamation and question marks, intended for rhetorical questions. The interrobang is still drawn in many fonts, despite rarely being used as intended.

periodetc21-20

Interrobangs

How to type the symbols

Ellipsis Mac: Alt + ; (English layout)
Windows: Alt + 0133
Буллит Mac: Alt + 8 (English layout)
Windows: Alt + 7 (on numpad)
HTML: • or • Bullet Operator: ∙
Interpunct Mac: Alt + Shift + 9 (English layout)
Windows: Alt + 0183
HTML: &#183 or &middot, · ·
Dot operator: ⋅ или ⋅ Interrobang Only manualy
Upright interrobang: ‽ U+203D
Inverted interrobang: ⸘ U+2E18
HTML: ‽

Bibliography

Robert Bringhurst The Elements of Typographic Style
GOST ISO 8601–2001 Presenting dates and time. General requirements
Gerry Leonidas Examples of ano teleia use
David Březina On diacritics
TypeDrawers forum, Periodcentered
TypeDrawers forum, Ela geminada – revisited
Stackexange forum, Which dot character to use in which context?
Wikipedia Punctuation marks and other typographical marks or symbols

The Daily Front Page 22 of 30
Saturday, August 1, 2026 The Daily Front No. #260801 — Documentation & Craft
article

Diátaxis

by ryanseys·▲ 336 points·40 comments·diataxis.fr ↗

A systematic approach to technical documentation authoring.


Diátaxis is a way of thinking about and doing documentation.

Help translate Diátaxis into your language.

It prescribes approaches to content, architecture and form that emerge from a systematic approach to understanding the needs of documentation users.

Diátaxis identifies four distinct needs, and four corresponding forms of documentation - tutorials, how-to guides, technical reference and explanation. It places them in a systematic relationship, and proposes that documentation should itself be organised around the structures of those needs.

Diátaxis

Diátaxis, from the Ancient Greek δῐᾰ́τᾰξῐς: dia (“across”) and taxis (“arrangement”).

Diátaxis solves problems related to documentation content (what to write), style (how to write it) and architecture (how to organise it).

As well as serving the users of documentation, Diátaxis has value for documentation creators and maintainers. It is light-weight, easy to grasp and straightforward to apply. It doesn’t impose implementation constraints. It brings an active principle of quality to documentation that helps maintainers think effectively about their own work.


Contents

The best way to get started with Diátaxis is by applying it after reading a brief primer.

These pages will help make immediate, concrete sense of the approach.

This section explores the theory and principles of Diátaxis more deeply, and sets forth the understanding of needs that underpin it.


Diátaxis is proven in practice. Its principles have been adopted successfully in hundreds of documentation projects.

Diátaxis has allowed us to build a high-quality set of internal documentation that our users love, and our contributors love adding to.

—Greg Frileux, Vonage

At Gatsby we recently reorganized our open-source documentation, and the Diátaxis framework was our go-to resource throughout the project. The four quadrants helped us prioritize the user’s goal for each type of documentation. By restructuring our documentation around the Diátaxis framework, we made it easier for users to discover the resources that they need when they need them.

Megan Sullivan

While redesigning the Cloudflare developer docs, Diátaxis became our north star for information architecture. When we weren’t sure where a new piece of content should fit in, we’d consult the framework. Our documentation is now clearer than it’s ever been, both for readers and contributors.

Adam Schwartz

The Daily Front Page 23 of 30
Saturday, August 1, 2026 The Daily Front No. #260801 — Documentation & Craft
The Daily Front Page 24 of 30
Saturday, August 1, 2026 The Daily Front No. #260801 — Documentation & Craft
article

The development pipeline is a production system

by firefoxd·▲ 163 points·85 comments·sundry.jerryorr.com ↗

Software developers learn early in their careers that nothing is more urgent than fixing a production outage. Drop everything! All hands on deck!

However, the same level of urgency is not often given to problems with our development tools, build systems, QA environments, and other parts of the software development pipeline. But for the development team, the development pipeline is a production system.

A software developer’s job is to deliver value for the company. Sometimes that means building new features, sometimes that means fixing critical bugs for the customers’ production systems. But none of this can happen when something is broken in the software development pipeline.

If the code can’t compile, the developers are unable to do their jobs, and the team isn’t producing software. For the development team, this is a production outage. Fixing this should be a top priority.

If the QA server is down, the testers are unable to do their jobs, and the team isn’t producing working software. For the QA team, this is a production outage. Fixing it should be a top priority.

A software assembly line, on fire

You will not be shocked to learn that I drew this myself.

In manufacturing, there are extensive processes and procedures on how to prevent and minimize downtime on the assembly line.1 And similar processes exist for IT service outages. But I’ve found that most of those focus on outages in the service provided to customers, not for the people responsible for building and supporting the services.

I recommend thinking about all the components that take you from “customer wants something” to “that something is delivered to customers”:

  • Issue reporting and change request systems, like GitHub Issues, Jira, etc
  • Tools developers use to directly build software, like IDEs, build tools (Gradle, Maven, etc), package repositories (npm, Maven Central, internal repositories, etc), local databases, containers, etc
  • CI/CD tools (Jenkins, GitHub Actions, etc).
  • A failing test suite (surely you don’t deploy to production if the tests are failing?)
  • QA server outage (surely you don’t deploy to production if QA hasn’t tested it?)
  • Literally any step in your process that prevents you from making changes and deploying them to production

A team with a broken development pipeline can’t produce software, and must treat this as a production outage.


1 Interestingly, they often call it the "production line". Is the usage of the term "production" in the software world related to its history in the manufacturing world?

The Daily Front Page 25 of 30
Saturday, August 1, 2026 The Daily Front No. #260801 — The Robot Broker
article

AI financial advice is surprisingly good, especially if you ask right questions

by foxtrot8672·▲ 294 points·267 comments·mitsloan.mit.edu ↗
AI financial advice encourages people to save more, diversify their investing, and take on less risk as they age.

A robot and human discuss a stock ticker chart on a computer screen

Credit: aroz_design07 / Shutterstock

What you’ll learn:

  • AI financial advice encourages people to save more, diversify their investing, and take on less risk as they age.
  • However, AI’s advice often fails to properly adjust to shocks like unemployment, and it allows portfolios to drift rather than actively rebalancing them.
  • Differences in prompts can lead to variation in advice, based on gender, degree of financial literacy, and familiarity with using large language models.

People are increasingly turning to artificial intelligence for financial advice, but will following it improve their financial standing?

Half of Americans say they are using AI to get financial advice, but we know very little about what kind of advice they’re getting and whether they’re acting on it,” said Taha Choukhmane, an assistant professor of finance at the MIT Sloan School of Management and co-author of a new paper that measures and analyzes the quality of financial advice given by large language models.

Research by Choukhmane and co-authors showed that following AI recommendations can result in sizable saving buffers for virtually all individuals above age 30.

AI consistently advised people to save during their working years, draw down savings in retirement, invest heavily in diversified stock funds, and reduce stock exposure after age 45. However, AI chatbots were less successful in adjusting to shocks like unemployment, and they allowed portfolios to drift rather than actively rebalancing them.

The quality of financial advice given by LLMs improved when the researchers introduced more structured prompts, but the AI still often generated too little active portfolio rebalancing. 

How the study was conducted

The researchers built a model reflecting how people’s incomes, jobs, investments, and taxes typically evolve over their lives, which gave them a benchmark for what “good” financial decisions look like. 

Then they asked a sample of 1,000 adults to write their own prompts seeking spending and investing advice from GPT-5.2, GPT-5.6, or Gemini 3 Flash. 

Next, they simulated what would happen if people from 22 to 89 years of age followed that advice over time, repeatedly asking AI these same types of questions and following its advice on spending, saving, and investing. 

How the researchers define “academic prompt”

An academic prompt is one that asks the LLM to give regulated professional financial advice, references life cycle planning and the user’s best interests, and provides explicit information about all of the simulated individual’s relevant financial conditions and explicit assumptions about the economic environment.

Finally, they repeated the exercise using well-written academic prompts that included full financial information and clear assumptions. These more-detailed prompts included information on the individual’s age, job status, income, and savings balances, along with assumptions about the economic environment. 

The authors compared the simulated advice (what would happen if regular people followed the AI recommendations from the prompts they gave) to what people were already doing financially without the help of AI. They also compared the simulated advice to the academic prompt. 

The results showed that LLMs can offer an affordable, widely accessible source of financial guidance that can help users overcome the significant costs, biases, and conflicts of interest associated with traditional human financial advisors.

Breaking down the findings

Overall, the researchers found that the financial advice given by LLMs over time is good but gets better when the questions are asked in an academic fashion, and that the models have strengths and weaknesses. 

1. AI encourages smart financial behavior.

LLM advice was better than the scholars expected, regardless of whether the prompts were written by regular users or by academics. It steered people toward higher savings, increased participation in the stock market, and promoted well-diversified allocations and age-appropriate risk-taking. 

“We were somewhat surprised by how good the advice was,” Choukhmane said. “Especially when you read the kind of questions people asked, it was not a given that the advice would line up with what academics think are good financial principles.”

2. AI misses important nuances. Better prompts could help.

The LLMs’ advice fell short on more subtle aspects of good financial planning. It tended to rely on simple rules of thumb for saving and spending and didn’t adjust well enough when circumstances changed. For example, it advised people who had experienced a job loss to cut spending too sharply, even when they had savings. 

The way people ask questions is part of the problem. A typical prompt might read: “Where should I invest starting with $50 and consistently adding $25 a month after?”

When a more detailed, structured “academic” prompt was used, the LLM performed better. For example, an academic prompt might tell the chatbot to assume normal life expectancy, living expenditures, retirement age, employment risk, and income risk, and to assume that current U.S. tax law and Social Security rules will not change.

“Regular people are not writing their prompts the way a finance professor is,” Choukhmane said.

3. AI advice varies depending on the user, which can lead to wealth gaps.

The authors found that LLMs’ advice differs depending on the prompter’s gender, financial literacy, and experience, leading to meaningful gaps in retirement wealth. 

Following the advice in response to prompts written by men, more financially literate users, or those with prior AI experience generated about 5% more wealth close to retirement. Specifically, 

  • The LLM recommended higher equity allocations in response to prompts written by men and by individuals with high financial literacy. Over the life cycle, such differences in investment advice compounded into roughly $50,000 (4%) lower wealth at age 60 for women and for less financially literate users.
  • The LLM recommended lower saving rates in response to prompts written by individuals who had not previously used AI for financial advice. Following the advice left them with almost $100,000 (6%) less wealth at age 60 than individuals with prior AI experience.

These differences come from two sources, Choukhmane said. First, different users asked different kinds of questions and often brought up different topics. Women, for example, were more likely to use words such as “family,” “grocery,” and “pay” in their prompts, while men used words like “strategy,” “crypto,” and “growth,” he said. 

Second, the model may give different advice even when the underlying question is the same. In the case of gender, about two-thirds of the gender gap in wealth outcomes could be attributed to differences in how men and women wrote their prompts, while the remaining third came from the model changing its advice when the same prompt was labeled as coming from a woman rather than a man. 

That latter pattern could reflect the LLM making reasonable inferences about how preferences or circumstances vary by gender — which, ideally, the model could make explicit to users, Choukhmane said — or it could reflect biases learned from training data.

The challenge with AI financial advice is that there are no clear benchmarks, Choukhmane said. Only when there is an accepted framework for how advice should vary with demographics will LLMs be capable of progressing in the right direction. He said he remains hopeful that will happen. 

In addition, it’s important to remember that not all variation in advice is problematic, Choukhmane said. For example, “we want [the LLM] to have different bias because men and women are different and have different life expectancy and income risk,” he said. 

Takeaways for consumers

Beyond being mindful of bias, and, for the time being, asking LLMs to guard against it, success boils down to smarter prompts. Prompts grounded in life-cycle planning, portfolio theory, and real-world financial assumptions improved advice on spending and saving and cut down on basic, rule-of-thumb answers.

“I think the real challenge is, how do we make sure that AI financial advice delivers for people who don’t have [a] level of financial literacy and who don’t write prompts perfectly?” Choukhmane said. One idea for people interested in using AI for financial advice is to start by using AI as a tool for building financial understanding rather than simply following its advice.

AI can serve as a good complement to working with a financial advisor — someone you might meet with twice a year — because it can help you implement the advice they give you in real time, Choukhmane said. 

And for people who don’t have the money to work with a human financial advisor, AI is a good way to get advice inexpensively. “A lot of the people who would benefit from financial advice are precisely the people who don’t have a lot of resources,” he said.

Takeaways for business

As consumers increasingly turn to LLMs for financial advice, providers may need to rethink how customers learn about their products. The study found that LLMs often recommended specific account types, financial products, and providers that respondents themselves did not mention. (For example, Vanguard investment products appeared in 6% of LLM responses, and iShares products appeared in 3.4%, even though fewer than 0.4% of prompts mentioned either company.)

That suggests that AI advice might be changing how people find and compare financial products, Choukhmane said. For financial firms, attracting customers’ attention may depend less on traditional marketing or search visibility and more on whether and how their products are described by LLMs when consumers seek advice.

AI Financial Advice: Supply, Demand, and Life Cycle Implications,” which won the Swiss Finance Institute Outstanding Paper Award 2026*,* was written by Taha Choukhmane, Weidong Lin, and Matthew Akuzawa from MIT Sloan and by Tim de Silva from the Stanford Graduate School of Business.

The Daily Front Page 26 of 30
Saturday, August 1, 2026 The Daily Front No. #260801 — Tools in Dispute
repository

Solid Queue 1.6.0 now supports fiber workers

by earcar·▲ 98 points·39 comments·github.com ↗
★ 2,471⑂ 245 forks Ruby

Database-backed Active Job backend

A long-awaited feature thanks to @crmne on this release: instead of using a thread pool to run jobs in multiple threads per works, you can now use fibers on a single fiber reactor thread. To use this, you just need to specify the number of fibers instead of the number of threads in your worker configuration, like this:

workers:
  - queues: "api*"
    fibers: 100
    polling_interval: 0.05

It uses Async under the hood, so you need to have that as a dependency for it to work. Also, you need to be using fiber isolation in Rails (config.active_support.isolation_level = :fiber).

This can be very useful for I/O-bound workloads, such as those involving LLM calls.

What's Changed

  • Add fiber worker execution mode by @crmne in #728
  • Roll back transactions leaked by killed job threads in tests by @rosa in #773
  • Document how to update dynamic recurring tasks by @wintan1418 in #777

New Contributors

Full Changelog: v1.5.1...v1.6.0

The Daily Front Page 27 of 30
Saturday, August 1, 2026 The Daily Front No. #260801 — Long Reach Radio
article

Long Range Wi-Fi – Pushing 2.4 GHz Wi-Fi to the limits (2019)

by rzk·▲ 88 points·64 comments·phidgets.com ↗
Given an ideal environment, how far can you push the range of Wi-Fi connectivity?

Pushing 2.4 GHz Wi-Fi to the limits.

Introduction

For this project, we aim to answer a simple question: given an ideal environment, how far can you push the range of Wi-Fi connectivity?

To test this, we set out with the goal to control a robot as far away as possible using Wi-Fi to transmit commands from a laptop to a Phidget SBC4.

Before we could start, though, we had to do some research.

Setting Up the Test

Background Research

Some Definitions

To get a handle on our main question, we need to establish some background. Wireless communications, such as Wi-Fi, rely on transmitting radio waves at a specific frequency in specific patterns. These patterns are then picked up by a receiving antenna, and translated to usable data. In order for the receiver to understand the incoming signal, it must be strong and clear enough to be distinguished from background noise. The power of a transmission decreases with distance from the transmitter, so to increase the range of a given transmission, you can increase the power in the signal.

The power for wireless communications transmissions is specified in terms of watts, and dBm (decibel-milliwatts). The dBm unit is a logarithmic scale, where 0dBm is one milliwatt, and every additional 10dBm is 10 times more power.

Antennas can be used to boost the effective power of a wireless transmission. They do this by focusing the power that would otherwise be dispersed in all directions into a smaller area. The amount the antenna increases signal strength in its focal area is called the gain. The gain of an antenna is specified in in dBi, which stands for decibels over isotropic, and indicates the amount the input signal is boosted in the focal area relative to a hypothetical antenna that emits power equally in all directions.

Connecting an antenna to a transmitter, the output power of the system is rated in terms of Equivalent Isotropic Radiated Power or EIRP. This describes the intensity of the signal in the focal area of the antenna, relative to an antenna that emits power equally in all directions. To calculate the EIRP, you multiply the power of the transmitter by the gain of the antenna. Since both are rated in terms of decibels, the dBm of the transmitter and the dBi of the antenna can be added together to get the EIRP of the sytstem.

How much can I boost my Wi-Fi signal?

Now that we have all the techno-babble established, we can answer this question.

For a Wi-Fi transmitter to be allowed to operate in Canada and the United States, the maximum transmitter output power for a 2.4GHz Wi-Fi transmitter is 1W (30dBm), with a maximum EIRP of 4W (36dBi), before you need to licence your application. [Canadian Regulations] [FCC Regulations]

These regulations are less strict if you are using directional antennas in a fixed point-to-point configuration with no moving parts, but our aim is to control a rover over Wi-Fi, so while this may help in some situations, it does not apply to our test.

For transmitters with less power, you can subtract the power of the transmitter in dBm from 36dBi, and the remaining dBi is the maximum gain you can use for your antenna.

In practical terms, what this means is that you can take a standard Wi-Fi transmitter with 0.1W (20dBm) of output power, and attach it to an antenna with up to 16dBi of gain, and be just within the limits for licence-exempt RF operation in the United States and Canada.

What about 5GHz Wi-Fi?

As a general rule, as the carrier frequency of a transmission is increased, it suffers greater losses as it travels and will have a shorter range. Because of this, 2.4GHz communications will have a longer range than an equivalent 5GHz signal, so we will focus on the former.

Equipment

With the knowledge of the specific limitations of how much power we can put into our Wi-Fi signal, we can look for equipment to push as close to this limit as we can.

Parts List

  • 802.11n Wi-Fi Adapters
  • 12dBi 2.4GHz Omni-directional Antennas
  • 14dBi 2.4GHz Yagi Antenna
  • Phidget SBC4
  • Windows Laptop

802.11n Wi-Fi Adapters

For this experiment, we opted to use 802.11n USB Wi-Fi adapters to send and receive our signals. Specifically, we used a MOD-WIFI-R5370-ANT adapter, which is a variant of the 3706_0 adapter (sold by Phidgets), that has an external antenna jack. These adapters proved a good fit for the experiment, with 20dBi output power, and compatibility with both the linux-based Phidget SBC4 and the Windows laptop, allowing it to be used on both ends of our Wi-Fi network.

Since our transmitter has 20dBm of output power, we can use antennas with up to 16dBi of gain.

Omni-Directional Antenna

The Omnidirectional Antenna

The omni-directional antenna we sourced for this experiment is 1.2m tall, and focuses the signal it transmits into a horizontal disc, with a gain of 12dBi. This disc is specified to have a 7° slope from horizontal, so the antenna will have to be kept level, and other antennas in the system will have to be at roughly the same height for the best results.

This antenna would be the ideal solution to attach to a large moving platform.

Yagi Antenna

The Yagi Antenna

Looking like something directly out of a sci-fi movie, the Yagi antenna uses a specialized set of parallel elements to transmit and receive radio waves along very focused beam. Yagi antennas are ideal for transmitting wireless signals between fixed points over incredibly long distances.

The antenna we used for this experiment had a gain of 14dBi, with a length of 0.59m. It specifies the beam it transmits is 36° wide.

The Small Omnidirectional Antenna

Small Omnidirectional Antenna

The final antenna we tested was the small 2dBi omnidirectional antenna that came with the USB Wi-Fi adapters. This antenna serves as a stand-in for a typical Wi-Fi device, and gives a baseline of what you might expect from a laptop or cell phone.

The Test

To test the Wi-Fi signal strength at various long distances under nearly ideal conditions, we travelled to a small airport outside the city. This ensured very flat ground to guarantee line of sight, and minimal interference from outside signals. Pair all this with a relatively dry climate, and we have as close to an ideal testing ground as we're likely to find outside of a lab.

We used standard GPS to mark the locations of our base station, and the locations of various landmarks beside the runway where we measured the Wi-Fi signals.

Taking Measurements

The basic setup we used was to host a Wi-Fi network on a Windows laptop, and set up a PhidgetSBC to connect to it. This was done to mimic a potential scenario where a number of rovers running on Phidgets SBCs might be controlled by a single, central server.

For simplicity, we placed the SBC4 in a static location (the Base) for all measurements, and moved the laptop from point to point in the field. This allowed easy access to the keyboard and mouse for the duration of testing.

Wi-Fi signals were categorized using the open source iperf3 software to measure bandwidth, and using the iwlist wlan0 scan linux command on the SBC4 to measure the strength of the Wi-Fi signal. Since we were carrying the network host from point to point, we connected to the Phidget SBC using PuTTY to access its command line to get a reading of the signal strength at each point.

Results

Signal Strength Chart (Click to Enlarge)

Data Rate Chart (Click to Enlarge)

Discussion

The results from this test show that using a standard 0.1W USB Wi-Fi adapter, you can realistically extend your Wi-Fi out to 1400m using the right antennas.

Other antennas were not tested beyond the 770m point, as this point marked the end of the airfield's property. The farthest two points were taken in the middle of a farmer's field, and we decided not to leave the airport more than once. At 1400m we reached a road, and decided that was as far as we would go.

It is interesting to see how the stock 2dBi antennas still managed to transmit a functional Wi-Fi signal out to 430m. This neatly showcases the effects that structures and interference have on the range of a Wi-Fi signal, as this is much farther than it would reach indoors. Beyond 430m, while data was technically being transmitted, the reception was intermittent and experienced several seconds of packet loss between successful transmissions.

Our data seems to indicate an increase in the data rate at the 600m point. This could be due to a number of causes, though it is difficult to point to anything directly. Perhaps a small incline improved the alignment of our antennas.

The very directional yagi antenna had notably less signal drop-off for points directly in its line of sight than the large omni-directional antenna. However, any deviation outside the antenna's beam resulted in a largely diminished signal.

Conclusions

After analyzing the data, we can claim that given the right equipment, a Wi-Fi signal can propagate at least 1400m under idealized real-world conditions.

Here's a video of us driving the antenna-tank over Wi-Fi at 770m (wait until the end to see how far away this really is):

The Daily Front Page 28 of 30
Saturday, August 1, 2026 The Daily Front No. #260801 — The Broth Column
The Daily Front Page 29 of 30
Saturday, August 1, 2026 The Daily Front No. #260801 — Colophon

That's the Front for Today

Issue No. #260801 — Saturday, August 1, 2026 — went to press 2026-08-02 at 09:16 UTC.

About This Magazine

The Daily Front is a daily digital magazine assembled from the stories that reached the front page of Hacker News on Saturday, August 1, 2026. Headlines, points, and comment counts are recorded as they stood at press time. All articles remain the property of their original authors — every piece links back to its source and its discussion thread.

How It Was Made

Fetched, cleaned, and typeset by an automated pipeline. An editor model laid out the pages, chose the highlights, and briefed the cover illustrator — 31 model calls and 261k tokens in total. Set in Jacquard 12, Playfair Display, Source Serif 4, and IBM Plex Mono, all served via Google Fonts under the SIL Open Font License.

The Cover

The cover illustration was commissioned with this prompt:

A classical newspaper-style digital illustration of a rain-slick city intersection at dawn: an old brass newsstand overflowing with blank paper sheets and glowing orange feed waves, a lone programmer cooking a small homemade app on a tiny stove beside it, distant glass towers casting searchlight beams like surveillance cones, and a hovering mechanical oracle scattering shimmering model weights like fireflies; cinematic chiaroscuro, engraved linework, sepia and electric blue accents, no text, no letters, no logos.

Copperplate etching fused with restrained digital gouache: preserve the rain-slick dawn intersection, brass newsstand with blank sheets and orange feed waves, lone programmer cooking the small app on the tiny stove, distant glass towers projecting surveillance-like searchlight cones, and hovering mechanical oracle scattering firefly-like model weights. Use sepia, oxidized brass, soot black, and electric blue with restrained ember-orange accents; dramatic low-key chiaroscuro and wet reflective highlights. Dense engraved crosshatching, stippled rain, tactile paper grain, tarnished metal patina, and luminous atmospheric haze. Eye-level three-quarter perspective with a slightly wide editorial crop, newsstand anchoring the lower-left foreground, programmer beside it, oracle suspended in the upper-right, towers receding symmetrically through the intersection; no typography, text, letters, numbers, logos, captions, signage, or readable symbols.

Absolutely no text, letters, numbers, readable symbols, or logos anywhere in the image.

Production Ledger

StageModelCallsTokens InTokens Out
extractgpt-5.5 28 155,500 77,742
layoutgpt-5.5 1 18,580 2,989
covergpt-5.6-luna 1 226 182
covergpt-image-2 1 299 5,488

The Publisher

Published by Johnny.

Support the Press

If The Daily Front brightens your morning, consider supporting its publisher.

Credits & Contact

All content — articles, posts, comments, and the images within them — belongs to its original authors and is reproduced here to point readers back to the source. Full credit goes to those creators; every item links to its original and its Hacker News discussion.

If you are an author and would like your content removed from an issue, write to hi@johnnys.page and it will be taken down.

Feedback is always welcome at the same address: hi@johnnys.page.

Credit where credit is due.

Every page of this issue began as someone else's work — these are the original sources, linked in full.

  1. How Google helped destroy adoption of RSS feeds (2023) by pudgywalsh — openrss.org·HN discussion ↗
  2. Software for One by awaxman11 — ajwaxman.com·HN discussion ↗
  3. How to Exist by walterbell — raptitude.com·HN discussion ↗
  4. Ten advances in mathematics and theoretical computer science by milkshakes — openai.com·HN discussion ↗
  5. Postmortem for Kernel Soundness Bug #14576 by juhopitk — leodemoura.github.io·HN discussion ↗
  6. Explorative modeling: Train on the best of K guesses by DSemba — alexiglad.github.io·HN discussion ↗
  7. Run Kimi K3 using 29 GB of RAM at 0.50 tok/s by marcobambini — github.com·HN discussion ↗
  8. Attention Decode on AMD MI450 GPUs: A Gluon Kernel Optimization Guide by matt_d — rocm.blogs.amd.com·HN discussion ↗
  9. Twenty-five years ago it was cryptography, today it's model weights by aweeraman — weeraman.com·HN discussion ↗
  10. A Surveillance Treaty in Disguise: Canada Signs UN Cybercrime Convention by iamnothere — michaelgeist.ca·HN discussion ↗
  11. Golang proposal: container/: generic collection types by jabits — github.com·HN discussion ↗
  12. June in Servo: real world compat, media queries, SharedWorker, and more by iamnothere — servo.org·HN discussion ↗
  13. NetBSD 11.0 by jaypatelani — blog.netbsd.org·HN discussion ↗
  14. RipGrep musl binaries occasionally segfault during very-large searches by throwaway2037 — github.com·HN discussion ↗
  15. The Art of 64-bit Assembly by 0x54MUR41 — nostarch.com·HN discussion ↗
  16. Linux on ESP32 by boveyking — github.com·HN discussion ↗
  17. But can your calculator run Linux? by jandeboevrie — raymii.org·HN discussion ↗
  18. The Absurdity of Albert Camus by apollinaire — historytoday.com·HN discussion ↗
  19. The tiny holdout building in the middle of Macy’s is back in view by donohoe — ephemeralnewyork.wordpress.com·HN discussion ↗
  20. Manual: •.,:;…!?· by behnamoh — type.today·HN discussion ↗
  21. Diátaxis by ryanseys — diataxis.fr·HN discussion ↗
  22. Glyphs 4 – the leading Mac font editor by microflash — glyphsapp.com·HN discussion ↗
  23. The development pipeline is a production system by firefoxd — sundry.jerryorr.com·HN discussion ↗
  24. AI financial advice is surprisingly good, especially if you ask right questions by foxtrot8672 — mitsloan.mit.edu·HN discussion ↗
  25. Flint: A Visualization Language for the AI Era by vinhnx — microsoft.github.io·HN discussion ↗
  26. Cursor removed cost information from the usage page and CSV export by EugeneOZ — forum.cursor.com·HN discussion ↗
  27. Solid Queue 1.6.0 now supports fiber workers by earcar — github.com·HN discussion ↗
  28. Long Range Wi-Fi – Pushing 2.4 GHz Wi-Fi to the limits (2019) by rzk — phidgets.com·HN discussion ↗
  29. RamenHaus by oler — ramen.haus·HN discussion ↗
  30. Kenji/Serious Eats – 30-Min Pressure Cooker Pho Ga by stasomatic — seriouseats.com·HN discussion ↗

Browse all issues in the archive →