Cover illustration

TheDaily Front

Issue No. #260728 Tuesday, July 28 2026 #260728 — TUESDAY, JULY 28, 2026
From tectonic plates to model weights, everything shifted today.
Tuesday, July 28, 2026 The Daily Front No. #260728 — Contents
30stories
8,647points
4,250comments
402kllm tokens
Assembled with 28 model calls — 263,152 tokens read, 138,630 written.

Highlights

Discovering Cryptographic Weaknesses with Claude

Anthropic’s Claude helps researchers find new cryptanalytic angles on HAWK and reduced‑round AES—no production impact yet, but a notable signal.

Zig's Incremental Compilation Internals

Zig debuts function‑level incremental compilation that patches binaries in place for dramatically faster rebuilds.

DMARC has been public since 2012 but most company domains still don't enforce it

Fourteen years after DMARC’s debut, most company domains still don’t enforce it, leaving inboxes open to spoofing.

A $500 RL fine-tune of a 9B open model beat frontier models on catalog review

A $500 GRPO fine‑tune of a 9B open model reportedly beats frontier configurations on a real catalog‑review task at a fraction of the cost.

Half-Life ported to Mac OS 9

Half‑Life finally lands on classic Mac OS 9—an impressive retro port 28 years after release.

From the Editor

A hard jolt in Japan reminds us the world still moves under our feet—even as labs and laptops redefine its future. Between a promising HIV vaccine milestone and a cavalcade of AI toolchain advances, today’s docket spans urgency and ingenuity. Read deeply, keep perspective, and mind the aftershocks—both seismic and societal.

  1. 7.1 Earthquake in Japan3
  2. New HIV vaccine shows unprecedented success in preclinical study4
  3. Substack writers, you need a website5
  4. Benchmarking Opus 5 on SlopCodeBench6
  5. How to survive boiling water7
  6. A $500 RL fine-tune of a 9B open model beat frontier models on catalog review8
  7. Zig's Incremental Compilation Internals9
  8. Steel Bank Common Lisp version 2.6.710
  9. Discovering Cryptographic Weaknesses with Claude11
  10. DMARC has been public since 2012 but most company domains still don't enforce it12
  11. Ars Astronomica – English translations of rare Hebrew and Latin astronomy texts13
  12. How Do I Profile eBPF Code?14
  13. About the security content of macOS Tahoe 26.615
  14. Netflix employee fired for sharing personal details in retreat trust exercise16
  15. Show HN: Yap – OSS on-device voice dictation for macOS with no model to download17
  16. C/C++ projects packaged for Zig18
  17. Kimi K3 Architecture Overview and Notes19
  18. Kimi Linear: An Expressive, Efficient Attention Architecture (2025)20
  19. Kimi K3 Now Available via Telnyx Inference API20
  20. Using an open model feels surprisingly good21
  21. Codex Security21
  22. Google's Beyond Zero: Enterprise Security for the AI Era21
  23. Vehicle Motion Cues22
  24. The iPhone Upgrade Program is being replaced by Apple Upgrade22
  25. Now is the time to give LLMs access to the ACM digital library23
  26. Stop Killing the Internet: No Digital ID and No Age Verification23
  27. Delayed Gratification – Proud to Be 'Last to Breaking News'23
  28. RTX 2080 Ti Memory Upgrade to 22 GB24
  29. Half-Life ported to Mac OS 924
  30. Una GPS smart watch – Repairable, USB-C charging, developer-friendly24
The Daily Front Page 2 of 25
Tuesday, July 28, 2026 The Daily Front No. #260728 — Lead: Japan Quake
article

7.1 Earthquake in Japan

by krembo·▲ 787 points·214 comments·data.jma.go.jp ↗

Read the full article →

The Daily Front Page 3 of 25
Tuesday, July 28, 2026 The Daily Front No. #260728 — Science Breakthrough: HIV Vaccine
article

New HIV vaccine shows unprecedented success in preclinical study

by codebyaditya·▲ 624 points·267 comments·lji.org ↗
“The best HIV‑fighting antibody response ever seen in primates.”

Vaccine approach yields high numbers of HIV-neutralizing antibodies in non-human primates

Highlights:

  • Scientists at La Jolla Institute for Immunology, Scripps Research have developed an HIV vaccine that trains immune cells to see past HIV’s defenses.
  • This HIV vaccine works by prompting the body’s immune system to make substantial numbers of rarely seen “broadly neutralizing” antibodies.
  • In this new study, this vaccine resulted in the best HIV-fighting antibody response ever seen in primates. Human trials have now started.

Crotty looks at camera and smiles. He is wearing a blue button-down

LJI Professor and Chief Scientific Officer Shane Crotty, Ph.D.

LA JOLLA, CA—A new HIV vaccine developed by La Jolla Institute for Immunology (LJI), Scripps Research scientists, and IAVI has the potential to protect humans from developing HIV infection and AIDS. This HIV vaccine is the first to generate a high number of “broadly neutralizing,” virus-fighting antibodies in primates.

“This feels like a huge success,” says LJI Professor and Chief Scientific Officer Shane Crotty, Ph.D., who co-led the research with Scripps Research Professor William Schief, Ph.D. “We constructed a successful vaccine from the ground up, which required a deep understanding of the immune system.”

This groundbreaking research, published in Nature, is the result of 14 years of collaboration between La Jolla Institute for Immunology and Scripps Research, as part of the Scripps Consortium for HIV/AIDS Vaccine Development (CHAVD). “This has been one of those Apollo moon mission-type projects, where there is an exceptional goal and the team has to accomplish a myriad of discoveries and inventions along the way,” says Crotty.

Outsmarting HIV

The new vaccine works by intervening in a process called B cell maturation. B cells make antibodies. Like many immune cells, B cells have an early “naive” stage before they are ready to make antibodies. B cells start to mature once they get the signal that a pathogen, such as a virus, is trying to attack. B cells see pieces of that pathogen’s molecular structure and start producing antibodies that can bind to that structure and halt infection.

It can take a little while for B cells to find the right “bullseye” on a pathogen. But B cells keep trying. As they mature, B cells tweak their antibody production, refining antibody structures to bind to a pathogen in just the right, vulnerable spots.

Scientists describe B cell development as a training process or bootcamp. In most cases, the body is left with a well-honed B cell army.

HIV is hard to beat because it doesn’t give B cells a chance to develop effective antibodies. The first problem is that HIV disguises itself from the immune system. The virus is wrapped in an ever-shifting cloak of sugar molecules, called glycans. This lets HIV sneak undetected past human cells, which are also covered in glycans.

The second big problem is that HIV mutates very quickly. “The worldwide diversity of HIV mutations is extraordinary. Even the diversity within one individual person living with HIV is dramatic,” says LJI Instructor Patrick Madden, Ph.D., who served as study co-first author with Jon Steichen, Ph.D., an institute investigator at Scripps Research.

The third problem is that HIV changes its shape when it infects human cells. Even if B cells get a glimpse of its viral structure—snap!—the structure changes.

Taken together, these problems rarely give B cells a chance to hone their antibody responses against HIV. Even if a B cell manages to make neutralizing antibodies, the virus can mutate or change its shape, rendering those antibodies useless.

Portrait photo of Dr. Madden. He is looking at camera and smiling

LJI Instructor Patrick Madden, Ph.D.

The LJI and Scripps Research teams spent years hunting for “broadly neutralizing” antibodies that can actually bind to HIV and recognize key viral structures, even if the rest of the virus mutates. These antibodies are very, very rare, but they can be found in blood samples from a small number of people living with HIV.

An effective HIV vaccine would need to prompt the immune system to make these same broadly neutralizing antibodies. “How could we flip the whole immune response on its head so the rare responses become the common responses? That was a critical challenge we faced,” says Crotty.

Testing the new vaccine

It was time to go back to B cell bootcamp. The scientists studied what made the HIV-fighting B cells special. Then they reversed the process to see exactly how those B cells matured. By looking back at the maturation process, the researchers could track how the B cells changed when they saw specific pieces of the HIV structure.

The team discovered that B cells matured to make broadly neutralizing antibodies after they got an early look at parts of HIV’s outer “envelope” protein. Because these viral sites sparked an immune response, scientists would call them “antigens.”

An effective HIV vaccine would likely need to include models of these antigens. The antigens would work like mugshots of America’s most wanted. If B cells saw those antigens early and often, they would get really good at recognizing and even neutralizing HIV. “We were trying to mimic the progression of those neutralizing antibodies,” says Madden.

In a feat of molecular engineering, the Schief Lab developed vaccine molecules that resembled the real HIV antigens. The scientists then worked with Emory National Primate Research Center, to test this potential HIV vaccine in a non-human primate species called rhesus macaques.

The researchers first administered a “priming” vaccine meant to activate each animal’s naive B cells. The animals then received a series of “shepherding” booster shots to help their B cells develop along the right path.

“This series of vaccinations will guide, or ‘walk’, a B cell from its naive state to its broadly neutralizing state,” says Madden.

This new type of vaccine approach is called “germline targeting” because it targets naive B cells in their “germline” or naive form, before they begin their training process.

The scientists found that around 44 percent of the animals went on to produce broadly neutralizing antibodies against HIV in their blood. These antibodies were impressively abundant.

“We succeeded in taking ultra-rare antibody responses and turning them into common responses by the end of the vaccination process,” adds Crotty. In other research recently published, they reported a new strategy to accelerate related vaccine antibody responses See Nature Immunology paper.

The team didn’t test whether these antibodies could prevent infection, but it’s significant that these antibodies could be found in the blood, where they could encounter and potentially block HIV.

Bringing the HIV vaccine to humans

The Crotty Lab plans to investigate how they might change the booster shot regimen to make the HIV vaccine even more effective. “It was incredible to get those results, but of course we’d like to see a response in 100 percent of the animals,” says Madden.

Importantly, the antibodies found in the animal subjects resembled the exact kinds of broadly neutralizing antibodies seen in those rare humans who made their own neutralizing antibodies. It’s clear that our immune systems can make these powerful antibodies, given the right training.

“We believe this vaccine approach is even more likely to succeed in humans, because of the immunogenetics,” Crotty says.

The priming immunogen used in this study was evaluated in humans in the HVTN 144 trial and is currently being tested in the Phase 1 trial IAVI G004. IAVI, Scripps Research, the HIV Vaccine Trials Network, and partners are now advancing plans to further evaluate the full immunization regimen in a future human clinical study.

Additional authors of the study, “Vaccination elicits HIV broadly neutralizing antibodies in primates,” include Claudia T. Flynn, Swastik Phulera, Monolina Shil, Oleksandr Kalyuzhniy, Alessia Liguori, Carolyne Kifude, Leigh M. Sewall, Christopher A. Cottrell, Krystal M. Ma, Sabyasachi Baboo, Jolene K. Diedrich, Katherine McKenney, Allan C. deCamp, Diane G. Carnathan, Ivy Phung, Parham Ramezani-Rad, Ester Marina-Zárate, Brian Freeman, Zhenfei Xie, Jeong Hyun Lee, Troy Sincomb, Nicole Phelps, Danny Lu, Diana Goodwin, Ryan Tingle, Yumiko Adachi, Nushin Alavi, Jenny Tran, Andy S. Tran, Alyne Nascimento, Catherine Sovie, Daniel L. V. Bader, Hannah Voic, Xiaoya Zhou, Grace Pixton, Agnes Walsh, Mariane B. Melo, Torben Schiffner, Facundo D. Batista, Dennis R. Burton, Darrell J. Irvine, James C. Paulson, John R. Yates III, Gabriel Ozorowski, Andrew B. Ward, Guido Silvestri.

This work was supported by National Institute of Allergy and Infectious Diseases (NIAID), of the National Institutes of Health, through grant UM1 Al100663 to the Scripps Center for HIV/AIDS Vaccine Immunology and Immunogen Discovery (CHAVI-ID), grant UM1 AI144462 to the Scripps Consortium for HIV/AIDS Vaccine Development (CHAVD), P51 OD011132 to Emory National Primate Research Center, and R01 AI113867; by the  Gates Foundation under the Collaboration for AIDS Vaccine Discovery (NAC INV-007522, INV-008813, INV-034657, and INV-064772), via the IAVI Neutralizing Antibody Center (NAC); and by the National Institute of Health grant S10OD025052.

The Daily Front Page 4 of 25
Tuesday, July 28, 2026 The Daily Front No. #260728 — Owning Your Platform
article

Substack writers, you need a website

by speckx·▲ 509 points·248 comments·elizabethtai.com ↗
“Substack is just a distribution tool to amplify your website.”

https://elizabethtai.com/wp-content/uploads/2026/06/pexels-photo-16622632.jpeg?w=1024

“But I already have a website on Substack,” you argue.

No, no, Substack is just a distribution tool to amplify your website. It should not be your digital home.

In the last few years, I’ve noticed a pattern of writers leaving their websites to make Substack their digital home.

Now, it’s kinda okay if they have bought a domain and linked it to Substack. (Meaning, it’s better than nothing.)

Rachel from Conscious Living is a good example. This way, Substack more or less functions like a content management system (CMS) for you.

However, compared to other CMS it’s very limited, such as the ability to manage your SEO and customize your pages to add more features, but I digress. If you just want a fuss-free platform, this is one way to get it and Substack’s conditions for domains are very reasonable and cost-efficient. As I will explain later, this could change on a dime without warning.

However, there are some writers who are saying: “Hey readers, I’m now writing on Substack, so head on over there (and ignore my website)!”

Some writers do have a website, but link to their Substacks, calling them their “blogs”. If your Substack has a domain name they own, it’s okay, but if it’s xx.substack.com, Substack is saying “All your content are belong to us”.

In conclusion: Writers, don’t do this. It’s short-sighted and unwise and can derail your long-term visibility on the Internet.

The siren call of convenience

Every few years, the internet convinces writers that a new digital paradise has arrived. First, it was social media like Facebook. Then blogging networks like Tumblr. Then it was Medium. More recently, it’s been Substack.

Platforms promise us an eager audience, built-in monetization, a smooth user interface, and a supportive community. As a writer who just wants to focus on writing, it’s incredibly tempting to hand over the keys to our creative kingdoms and let these portals handle everything. (Believe me, I gave in at one point. For years, I just stopped blogging altogether and even gave up a domain that had high traffic! But I got back in 2012 and never left.)

However, this is the truth that has not changed since the dawn of the Internet: When you build your audience entirely on someone else’s platform, you aren’t a homeowner. You are a tenant. Or worse, a digital sharecropper.

And corporate landlords always change the rules eventually. It’s not personal, it’s just business.


The illusion of the safe space

It’s easy to feel secure when a platform is in its golden era. But we’ve watched the downfalls of Twitter, the policy shifts of Reddit, and the changing tides of algorithmic networks. Relying blindly on a centralized portal not owned by you means your life’s work can alter overnight based entirely on a corporate boardroom decision.

When I looked at how fragile our digital ecosystems really are, I realized I needed a space that wouldn’t go “poof” because a company needed to please its investors or shareholders. This realization completely changed my approach, pushing me to protect my content by learning to blog the IndieWeb way.

Your writing needs a permanent homebase—a domain that you own and control. Full stop.

Moving from “renting” to syndicating

Field

Do you own the land you plant your “content” crops?

The biggest pushback I hear from writers is: “But my website doesn’t have an audience! Substack does.”

But you don’t have to completely abandon social media or platforms like Substack to protect your autonomy (personally, I prefer the word sovereignty but it does sound a tad dramatic).

You just need to change the order of operations. Instead of publishing directly to a portal, you can shift your mindset to POSSE: Publish (on your) Own Site, Syndicate Elsewhere. (I explain the POSSE/PESOS method in an older post.)

By treating your website as the definitive source of truth and using platforms simply as distribution pipes, you get the best of both worlds. I dug deep into this shift when I committed to being an imperfect gardener of my digital garden, exploring how a less market-y way of presenting my content online let me share my wild garden of thoughts without dancing to the algorithm.

A reality check on platform hype

If you are still holding out hope that Substack is “different” from the social platforms that came before it, let’s look at the numbers and behaviors behind the marketing copy.

After spending a significant amount of time observing the platform ecosystem firsthand, I wrote a brutally honest takeaway in What I learned from one year of Substack. The network effects are real, but so is the pressure to conform to what the platform’s ecosystem favors.

This post, by the way, desperately needs to be updated because things have gotten much, much worse since I wrote it.

When you hand your content over to a platform, you have to conform to their rules and their localized biases. For those of us writing from outside the dominant US-centric echo chambers, platform algorithms heavily prioritize specific western narratives, making it incredibly tough for localized or minority voices to be seen unless they conform.

I wrote about this exact frustration recently in Linkblog: Dwelling on the Internet, highlighting how algorithmic complacency forces us into homogenized bubbles.

The flip side – the writers who refused to leave their websites

Each time there’s a new drama on some platform, and writers are shaking their sabers and declaring that they will leave for yet another social media platform they don’t control, I think about writers like John Scalzi.

As of date, John scalzi has been blogging on https://whatever.scalzi.com/ for 28 years!

This sci-fi novelist has maintained a single independent website continuously for nearly three decades; this makes him one of the longest-running, most consistent original bloggers on the internet. Imagine the amount of digital footprint on that website! Unbroken by time or platforms.

(Specifically, he uses wordpress.com like I do, as we both don’t want to bother with the pain of setting up your own self-hosted wordpress website and just want the folks at Automattic to do it.)

He blogs in the classic Indieweb way, though I doubt he is even aware he’s doing it. He treats his social media channels such as X or Bluesky as a way to amplify his website. All roads lead back to https://whatever.scalzi.com/, and this is something I wish every single writer would do.

He wrote recently in Various & Sundry, 6/3/26:

this site acts as my own institutional memory, if I post something about it here it constitutes an official record. I mean, all the posts I ever placed on the former Twitter are now entirely lost to time, since I have gone in and purged my entire timeline there. This site, however, endures. – John Scalzi

Breaking free from platform blues

Trying to adapt your presence across various platforms in an ever-shifting digital landscape is exhausting. One minute a platform is a writer’s darling; the next, it’s being boycotted. Railing against a platform’s focus shift or the presence of (long sigh) Nazis is a useless endeavor.

As I noted in Linkblog March 12, 2026: Platform blues, chasing platform purity is an illusion. Tech will change, corporate algorithms will continue to prioritize profit over human connection, and platforms will continue to cycle through hype and decline.

The antidote to this exhaustion isn’t moving to the next shiny new app. It’s anchoring your work on an independent website with open distribution channels like RSS. It also means ruthlessly using platforms as distribution channels. When one collapses or you prefer to just move, it’s easy to just change strategies because your digital home remains unchanged.

Use platforms to find your readers, but bring them back to your house. It’s time to stop digital sharecropping on rented land.

Featured photo is by vivek vk on Unsplash

The Daily Front Page 5 of 25
Tuesday, July 28, 2026 The Daily Front No. #260728 — Benchmarking Coding Agents
repository

Benchmarking Opus 5 on SlopCodeBench

by dhorthy·▲ 391 points·116 comments·github.com ↗
★ 2,186⑂ 160 forks

Benchmarking Opus 5 on SlopCodeBench

we got better benchmarks

I've written before something along the lines of:

THERE ARE NO GOOD BENCHMARKS for a model's ability to maintain codebase quality

That wasn't entirely true. I love nothing more than burying a good lede.

Last Friday I dug into SlopCodeBench, a new-ish (March 2026) long-horizon coding benchmark from @GOrlanski's lab at UW Madison. It addresses the thing that bothers me most about coding benchmarks - that even "larger" more complex benchmarks still divulge the whole problem up front:

Software is discovering problems as you go, but even most modern benchmarks disclose the whole problem up front -- no reason to optimize for 'is this easy to change/adjust later'

In contrast, each challenge in SlopCodeBench has multiple "checkpoints" - the model doesn't know the whole problem up front, it has to evolve the codebase over time as new requirements are divulged.

It's a good paper. It's not that long. You should read it.

the mechanic, from the paper: an initial spec becomes spec 2 becomes spec 3, and the solution goes Good Quality to Degradation Begins to Unmaintainable

What's cool about this benchmark is that it is unsaturated - at the time of running, the best models available, GPT-5.4 and Opus 4.6, got 11% and 17% strict pass rates, respectively.

table of results from initial paper

benching opus 5 on slopcodebench

On Friday I ran three claude models (Opus 4.8, Sonnet 5, and Opus 5) through a subset of SlopCodeBench and watched it live for six hours. Opus 5 wins technically but none of them did a very good job IMO. Will post more results soon with Fable and 5.6 Sol in the mix.

The big headline is that Opus 5 got a 24% on the small subset of the benchmark that I ran - not much higher than Opus 4.6's 17% strict pass rate in the original paper. All of the tested models showed a pretty significant increase in verbosity, complexity, and a bunch of other code smell metrics over the course of each challenge, with Opus 5 writing five times the number of functions/callables than Opus 4.8 over the course of the same set of challenges.

strict pass rate — opus 5 at 4 of 17, opus 4.8 and sonnet 5 at 1 of 17 each, against the paper's opus 4.6 at 17% and gpt-5.4 at 11%

My personal read of this 23% pass rate is that SlopCodeBench finally gives some signal for a thing I've only ever been able to argue from vibes - that for real-shaped software engineering work, building one issue at a time, today's models can't be relied on to run lights-off without steering.

the benchmark subset

i had claude pick out 3 problems from the repo, 17 checkpoints total, a mix of easy/medium/hard labeled problems:

  • circuit_eval — easy (8 checkpoints)
  • database_migration — medium (5 checkpoints)
  • dynamic_config_service_api — hard (4 checkpoints)

There's an appendix at the end with all 17 checkpoints explained in detail but I won't put that all here.

and then I ran them across all three models, in parallel, with a fresh context window per checkpoint. All models got the same prompts and ran in the claude code harness.

the metric I decided I care about is the strict pass: everything new is green including every regression test that was inherited from previous checkpoints.

A model fails a checkpoint if the solution has a defect - defects are detected by taking the models output, a CLI to run or in some cases e.g. an api server to poke at, and running a set of held-out black-box tests against the produced entrypoint.

  • Model writes code for ck1
  • Eval harness runs black-box tests against ck1
  • Model writes code for ck2
  • Eval runs black-box tests for ck1 and ck2
  • etc

Again, the strict pass criteria means that if a model bungles something in checkpoint 4, it can't pass the following checkpoints because that failing part of the code carries forward (unless the model indavertently fixes an eval case in checkpoint 6 that was broken in checkpoint 4, but we didn't see this happen in practice).

For all 9 test runs, none of the models made it to the end of any challenge with everything passing, even on the problem marked as "easy" difficulty.

while it ran

sonnet's first checkpoint was more expensive but by the end of problem 1, sonnet became the cheapest of the three. (it seems like once the basics were built and the work turned into maintenance, then the cost savings started to take over)

For the first challenge, last generation's models accumulated defects steadily, Opus 5 a defect each on checkpoints 4 and 5.

defects left open per checkpoint on circuit_eval — opus 4.8 climbs 1 to 10, sonnet 5 stays flat, opus 5 opens with four clean checkpoints

for the first two hours opus 5 was the only model with any strict passes at all — three in a row to start.

the live progress page at 12:12pm — 8 of 51 checkpoints done, $12.31 spent, and both of the run's two strict passes belong to opus 5

Things evolved as we went. Claude diligently updated the html.

the same progress page later in the run, with more of the grid filled in

Compared against the other models, Opus 5 was technically better on problem 1 (circuit_eval). But after acing the first three checkpoints, every subsequent solution had at least one defect (failing test case).

final result

If our definition of success is "reached the final checkpoint with no defects" then opus 5 failed all three problems, but it failed slightly-less-badly than the other models.

defects left open at the end of each problem

For the costs vs defects report, I really hate claude-isms but this one i decided to leave in:

every dollar bought correctness. nobody bought enough of it.

(obviously this small subset of the bench cannot tell us definitively that spending more $$ will lead to higher pass rates)

cost vs defects, one point per model per problem

as far as strict passes go, Opus 5 got four of them (24% pass rate) (the first three ck of circuit_eval, plus database_migration ck1).

opus 4.8 and sonnet 5 both got one strict pass (6% pass rate), the same database_migration ck1 that opus 5 got.

so the winner cleared 4/17, and 3 of those were the opening checkpoints of one problem. It would appear we have an unsaturated benchmark for the next frontier of models. nice work @GOrlanski and team.

the slop meter

I'm not totally sold on "linting the slop away" just yet, because I don't think its yet possible to deterministically parse the "maintainability" of a particular codebase checkpoint. But they are interesting to keep an

Code quality metrics are interesting to keep an eye on, and they're probably directionally correct, and

With SlopCodeBench, you get the results after each checkpoint across various quality metrics. There are 41 of them in the results file. Roughly grouped:

  • size — source lines, files, functions, methods, classes, statements, and lines added and removed at that checkpoint
  • complexity — cyclomatic complexity mean, max, and spread, how many functions land in the "high" and "extreme" bands, how concentrated the complexity is, max nesting depth, and mean function length
  • duplication — cloned lines, and clone lines as a share of source
  • decomposition — single-use functions, trivial wrappers, unused variables, lines per symbol
  • rule violations — lint errors and how many are auto-fixable, ast-grep hits against test slop rules, and the share of lines flagged verbose
  • dependency graph — propagation cost (how far a change ripples), cyclic dependency mass, dependency entropy (probably the most interesting one to me)

Each of these is computed deterministically using the current code state after each checkpoint.

The chart below shows the spread among models for ck1 score vs. ck8 for the circuit_eval challenge. (That is, how much did the slop indicator increase over the liftime of the challenge checkpoints.) Most interestingly, most of the metrics don't tell the models apart.

percent change from checkpoint 1 to 8 for each rate metric, one dot per model, sorted by spread — cc_max and cloned_pct separate the models sharply, the rest cluster

I like that these measures are repeatable and don't use a model for judgement. But the link between any one of them and "is this codebase easy to change and evolve" is not yet established.

more correctness came at the cost of wayyy more code

source lines written — opus 5 wrote 29,065 against roughly 9,000 each for the other two

But a lot of that was "more tests" - the actual production volume is closer to 1.8x for opus 5 vs opus 4.8.

the same source lines split into production and test code — opus 5 is 51% tests, opus 4.8 is 11%, sonnet 5 is 24%

My guess would be...expensive verbosity here that didn't translate directly to much better results.

Will have to dig in more to know whether this is a model tic signal or just a "this is actually a really hard problem and warrants this much code".

almost all the code written triggered the slop meter

For all models, a huge majority of the code lines tripped at least one of the the benchmark's slop rules. The averages across the three problems:

  • opus 4.8 — 98%
  • opus 5 — 93%
  • sonnet 5 — 89%

And specifically, lines flagged as being too verbose go up across the trajectory for every model, roughly 65% at ck1 to 80% by ck8, even for Opus 5.

I'd actually probably say this is a sign that some of the code quality measures are a bit over-aggressive. I looked into applying the ruleset to our typescript monorepo, but the current slop-code-bench detectors are python only.

So I had 5.6-Sol cook up a subset of rules for typescript, but it only came up with 76 slop detectors (compared to the SCB python library of 200+) - but found some directional findings - the Opus 5 lights-off-generated solutions have over 11 times more slop triggers per kLOC than our 99%-AI-generated-but-also-carefully-reviewed Typescript monorepo (yes that's a 1000% increase).

findings per KSLOC — the Synclayer TypeScript monorepo at 15.06, Opus 5's pooled checkpoints at 174.88 and its three final snapshots at 178.88, so 11.6x and 11.9x the density

Obviously there's a mountain of asterisks on this finding (fewer rules, haven't reviewed the parity, etc.) but it's interesting to say the least.

these models write a lot of functions

Another interesting data point - opus 5 wrote 5x more functions than the other two models. But Opus 4.8 wrote a higher %% of single-use functions (almost 50% of its functions were called exactly once). And Sonnet 5's share of single-use functions is the highest at 71.5%.

single-use functions as a share of all callables — opus 5 at 14.9%, opus 4.8 at 49.1%, sonnet 5 at 71.5%

FWIW I don't think lots of small functions is bad. I take it with a grain of salt these days, but I used to be a die-hard Clean Code guy. Small descriptive function names are way better than lots of comments, yada yada

Complexity grows over time for all models

I've been saying models degrade codebase quality over time for about a year now, mostly on vibes. But now we have some data.

mean cyclomatic complexity and duplicated-line percentage across all eight circuit_eval checkpoints, one line per model

Not a single model made it through all the challenges without increasing complexity across checkpoints. While Opus 5 has the lowest mean complexity, it also wrote 2000 functions. There's a tradeoff here: lots of small functions or fewer big ones. I don't think any of these complexity metrics can stand alone, but they give us some kind of composite signal.

callables written against mean cyclomatic complexity, one point per model per problem — opus 5 sits low-right with many small functions, opus 4.8 high-left with few big ones

Both sonnet and Opus 4.8 answered the increasing complexity of the challenge checkpoints by making individual functions bigger rather than moving things around. Opus 4.8 is the extreme, up 70% over eight checkpoints, and its single worst function ended at a cyclomatic complexity of 93.

Duplication is where they split. Opus 4.8 goes from 4.6% to 16.8%, with an inflection at ck3 — ~roughly where new requirements start fighting the initial design.

Aside - here's what the first three checkpoints of circuit_eval ask for (full listing for all challenges in the appendix at the end):

  • ck1 — a CLI with --help, --version, a JSON output mode, and a check command that parses and validates a .circ circuit file. Every signal is a single bit.
  • ck2 — an eval command: pass the circuit some inputs, get the outputs back. Still one bit per signal, standard boolean operators.
  • ck3 — signals become vectors. data[7:0] instead of data, plus slicing, indexing, concatenation, new operators (MUX, reductions, EQ), a width check on every operand.

By the end one line in six is a copy of another line. The other two models came down over the same stretch.

However, opus 5 is basically flat, 2.41 to 2.64. I've previously argued that "improving codebase quality over time" had not moved much between model generations. So if you trust duplication as a golden metric, you could argue that we did getting incrementally better in the last ~3 months. Big if though, and I think most software architecture experts would agree that it's not black and white.

the shape of a better oracle for software quality

While the code quality metrics are interesting, I don't think they tell the whole story, and its easy for a model to reward hack any of them. Just like SWE-bench shaped problems were the best verifier for "solve a software problem one time", because they map onto real world work at that "zoom level", I think "pass all verifiers for an incrementally-divulged spec" is a very realistic eval for "can a model maintain a codebase over time".

That is, a codebase becoming hard to maintain would lead to failing checkpoints in later stages, so a higher strict pass rate is a signal that the model is good at building a codebase that is maintainable.

With frontier models like Fable / Sol proving to be expert debuggers and reverse-engineers, incorporating cost/time/token metrics might become more important over time - frontier models like Fable and Sol can probably get it done in the NASTIEST of codebases, but I'd venture that a well-factored codebase will tend to lead to shorter, more token-efficent solves for future problems.

And while "build a whole feature across 8 checkpoints" is a lot slower than "solve a 15min SWE-bench multilingual problem", it can be executed unattended and is subject deterministic behavior verifiers at the end. So IMO its a much better oracle than e.g. "does another model think this code is clean".

I think any even better signal that we can hope to get from a model that is really good at maintaining a codebase, is that we could try having a frontier model like opus 5, fable 5, or gpt-5.6-sol write the first N checkpoints, and see if a dumber model like sonnet 5 or gpt-5.6-terra can implement checkpoint N+1.

three lanes of the same eight checkpoints — fable 5, gpt-5.6-sol and opus 5 each build ck1 through ck7, then hand the codebase to sonnet 5 for ck8

This amplifies the signal of whether the smart models did a good job maintaining high-quality code that is easy to change. Whether a small model like Sonnet or Terra or even Haiku can implement checkpoint 8 impacts the smart models' score on checkpoints 1-7.

something you can actually measure

My personal read of all this is that SlopCodeBench gives a signal for something I've so far only been able to argue from experience - that for real-shaped software engineering work, building one issue at a time, today's models can't be relied on to run lights-off without steering.

SCB is a measure of the future that I'll be keeping a close eye on. I wouldn't bet my codebase on a good score in Frontier Code, SWE-Marathon, or DeepSWE, but if/when models can score 80%+ on a (well-held-out) benchmark like SlopCodeBench which measures iteration over time, I'll feel a LOT better about setting them loose with the lights off.

I won't posit when that will happen, because "when" matters less than having a good signal to know that it's happening. (Assuming nobody "accidentally" trains on test in the meantime).

what's next / things i'd do differently

I'll be reading some of the slopcodebench problems more deeply for inspiration and to curate a few that I think map well to the day-to-day building we do here at @humanlayer_dev.

Claude decided to parallelize by model, running each through three challenges in sequence. We could have just as easily done 3 models x 3 challenges in 9 parallel sessions and finished in 1-2 hours instead of 6.

As I said, I looked into applying the ruleset to our typescript monorepo, but the current slop-code-bench detectors are python only. It would be interesting to port those to TS and a few other languages. I hate to be that guy but I'd bet python is a more slop-prone language than most.

I think instead of focusing on strict pass and total defects, it will be interesting to explore more dimensions of the benchmark. In the current scoring, we're taking any failure along the way as an accumulated defect - unless the model happens to resolve that past defect in a future session, all of the remaining checkpoints cannot pass.

Many software factories include prompting for better style, include deterministic feedback during the code loop for complexity and a lot of these other software quality metrics. Today's results don't evaluate model code quality or success rates with those sort of guardrails in place. We used the "just-solve" version of the prompt from SlopCodeBench but there are other variations like including instructions about quality/duplication in the prompt. And it would be very interesting to re-run the whole eval with either or both of 1) an "aversarial review" loop w/ a model judging quality and 2) code-quality backpressure for things like cyclomatic complexity.

I'm not made of money or time but it would be fun to do a bigger dataset here.

And of course, most interesting is this idea of "can we amplify the slop signal by handing fable's codebase to a smaller model like sonnet".

vibe check - the frontier is still dumb af

while this whole experiment was happening, in another session, opus 5 decided to go rogue and rewrite an email draft with new formatting, then send it to 100 people without checking in with me. Terrible.

The user is upset because I made a critical mistake: I overwrote their edited draft by patching it with the final version, then sent it out.

Sorry if you got one of those humanlayer product updates with the ugly header banner (I think the new claude-ism for this is "kicker"??)

really feeling the AGI over here friends

🫡 -dex


Links From This Post

Appendix: the challenge checkpoints

All 17, in order, straight from the prompts the models were handed. Each one arrives cold — the model has no idea any of the later ones exist.

circuit_eval — easy, simulation, 8 checkpoints

  • ck1 — a CLI with --help, --version, a JSON output mode, and a check command that parses and validates a .circ circuit file. Every signal is a single bit.
  • ck2 — an eval command: pass the circuit some inputs, get the outputs back. Still one bit per signal, standard boolean operators.
  • ck3 — signals become vectors. data[7:0] instead of data, plus slicing, indexing, concatenation, new operators (MUX, reductions, EQ), a width check on every operand, and --radix output formatting.
  • ck4 — three-valued logic. Inputs can now be X (unknown), and every operator has to say what it does with one.
  • ck5 — two more input formats. check and eval now read .json and .bench files as well as .circ, behind a --format flag.
  • ck6 — three analysis commands: stats for metrics, lint for warnings, dot for Graphviz export. All of them work with all three formats.
  • ck7cone (pull out a subcircuit), truth-table (enumerate every output), equiv (check two circuits match), plus a --seed flag for reproducible randomness.
  • ck8opt: a circuit optimizer with configurable passes, deterministic output, optional equivalence verification, and BENCH export.

database_migration — medium, databases, 5 checkpoints

  • ck1 — a CLI that reads migration specs out of JSON files and applies them to a SQLite database: create tables, add columns, change table structure.
  • ck2 — data migrations. Transform the rows that are already in there using SQL expressions, not just the schema around them.
  • ck3 — foreign keys, custom indexes, and advanced constraints. Relational integrity and query performance.
  • ck4 — rollback. Undo migrations one at a time or in batches, with dependency handling.
  • ck5 — dependency management. Migrations declare depends_on, and the tool has to resolve the order and detect circular dependencies.

dynamic_config_service_api — hard, system-design, 4 checkpoints

  • ck1 — a REST service that stores JSON config objects with immutable versions, scoping, rollback to any earlier version, and imports/inheritance across configs.
  • ck2 — a schema registry with its own versioning, schemas bound to configs, validation on create and on resolve, and ingesting raw YAML/TOML/JSON parsed into canonical JSON internally.
  • ck3 — a change-management workflow. Every new version starts as a draft, proposals gather human reviews, activation requires a quorum, and each proposal carries a deterministic diff.
  • ck4 — an org-level guardrail layer that runs policy bundles against resolved configs and the graph around them, blocking unsafe proposals with violation details distinct from schema errors.
The Daily Front Page 6 of 25
Tuesday, July 28, 2026 The Daily Front No. #260728 — Microlife vs. Heat
article

How to survive boiling water

by cainxinth·▲ 472 points·109 comments·taxa.substack.com ↗
“A probiotic paradox.”

A probiotic paradox

The story of MIT’s most notorious milk carton begins, as many good stories do, in a college dorm.

The milk in question was purchased in 1994 and rediscovered in 1995 by an undergrad named Justin Cave. By that point, the reportedly lactose-intolerant Cave had even less use for the milk he had abandoned in his fridge ten months earlier. For reasons lost to history, he did not throw the milk away. He threw it a birthday party.

The Milk lived the rest of its life unrefrigerated, stored in a tall, single-walled jar. For twenty-seven years, the residents of Random Hall dorm gathered faithfully to celebrate its birthday. At age 20, the Milk applied to, and was rejected from, MIT1. The jar was periodically “burped” to release the gas pressure inside, until the Milk reached its stable final form – a cloudy brown liquid. When asked why the Milk was never thrown away, one resident of Random Hall replied: “Why throw something away when you can tell a story about it?”

A group of dairy products on a table  AI-generated content may be incorrect.

The Milk (center, party hat) and friends celebrate its 21st birthday (credit: MIT).

“Stuff I learned from things that normally get thrown away” could be the title of many scientists’ memoirs, including Louis Pasteur’s. Winemaking produces, well, wine, but it also produces acidic crystals on the walls of the vats. These byproducts were not discarded – they were studied by Pasteur and his contemporaries. Pasteur’s observations both revolutionized our understanding of chemistry and led him to the phenomenon that would define his career and lay the foundation of modern food safety: fermentation.

The microorganisms responsible for fermentation are visible to our naked senses only through the textures, colors and smells resulting from their collective efforts. From Aristotle through the 1850s, it was assumed that some intrinsic property of a non-living starting substance (like grain or milk) enabled its spontaneous fermentation into something useful (like beer or yogurt) or its eventual spoilage.

It was Pasteur who proved that living organisms were required for the transformations that took place during fermentation. He heated up nutrient-rich broths in custom flasks that let gases, but not microbes, flow in and out of the flasks. Pasteur then broke the neck off of one of the flasks to expose the broth to the air. If the boiled broth could spontaneously transform, it would do so with or without exposure to microbes in the air and environment.

The flask with the neck broken off grew cloudy and fermented as bacteria bloomed, but the sterile one remained clear.

Pasteurs experiments and similar ones that followed class 11 biology CBSE

Pasteur’s famous “swan-necked” flask experiment from 1861. (Illustration source: Vedantu)

This finding was great news for Napoleon III2. The French were losing money, and perhaps more alarmingly, their reputation, exporting wine to the British – the wine would “spontaneously” go bad during shipping. The French government offered a prize for a scientist to solve the case of the spoiled wine. With the knowledge of the microorganism-driven process of fermentation in hand, Pasteur did to the wine what he did to the broth, just more gently – he heated the wine enough to kill microbes without damaging the wine’s flavor. Immortalized as pasteurization, this process was adapted shortly after its invention in 1865 to let us safely drink stored milk3.

Killing bacteria thus became a major preoccupation of modern life. Its most visible manifestation today might be taking antibiotics (first available in the 1940s): there were ~700 antibiotic prescriptions per 1000 people4 in 2024 according to CDC data. A close second might be the dizzying array of disinfectant products found in U.S. grocery stores.

What’s less visible is the sanitization infrastructure that makes things like grocery stores or medicine possible at all. The company Steris, one maker of high temperature, pressurized sterilization equipment and other medical instruments, is a $5 billion annual revenue company, with a $21 billion market cap. The U.S. pasteurizes around 50 billion liters of fluid milk every year. To package salad greens like spinach, the greens are washed in a dilute bleach solution to kill any lingering soil microbes. I could go on.

But our war on bacteria has its own warring industry. This industry has captured the imaginations of scientists, the food and beverage industry, pharma companies and doctors along with influencers, marketing gurus and opportunists of all flavors. This industry emphasizes that some microbes are friends, not foe, (true) and you should be eating them in large quantities, on purpose, all the time, and preferably paying more for products that contain them (dubious). This is the probiotics industry.

“Probiotic” is a bit of a misnomer – it means for life, or promoting life, but the formal definition of a probiotic is an actual living microorganism. In simple terms: taking a probiotic is just eating bacteria on purpose. I say on purpose because we consume microbes accidentally all the time from our environment, largely oblivious to their existence or effects. The bacteria we spend much of our time and energy trying to kill are outnumbered, at a species level, at least 1000 to 1 by a combination of harmless and beneficial bacteria living in and on our bodies. It’s this latter property of beneficialness that probiotics are trying to exploit.

I say exploit because of a recent trip I took to the grocery store. I had a cold and was in search of lemon ginger tea. I bought a box of Bigelow, went home, boiled some water, poured it over a tea bag, waited a bit, added honey, took a sip, and almost spit it out. The tea had its expected notes of ginger, a hint of lemon, and some powdery, alkaline aftertaste that I couldn’t place. Frankly, it tasted terrible. (Sorry, Bigelow).

I inspected the box again. In my congested state, I had unwittingly purchased a new offering from the tea company – Bigelow Lemon Ginger, with probiotics. What made this tea different from all the other teas I happily sipped on was that in addition to nice-sounding things like lemongrass and cinnamon, it contained bacteria. Bacteria which I had just boiled, at a temperature 40oC hotter than pasteurization.

Did the tea taste bad because I was drinking dead bacteria water? And if that was the ultimate outcome of the normal brewing process, why bother putting bacteria in the tea at all?

A hand holding a box of herbal tea  AI-generated content may be incorrect.

One of these things is not like the other.

I was at a lab happy hour when I mentioned this to my PhD thesis advisor. “I know, right?” she said, suddenly animated. “Probiotic teas taste SO BAD.” I was thrilled to have another witness. “Doesn’t it seem crazy to add in bacteria that you’re just going to boil and kill anyway?” I asked. “Is it all a scam?” She was already nodding. “You have to wonder whether the bacteria in the tea make it to the gut at all, and whether they do anything helpful once they get there,” she said.

I’d be lying if I said I remembered exactly what happened next, or who suggested what. All I remember is an idea. An idea to test this seemingly paradoxical marketing tactic like the microbiologists we are. The idea was simple: What if we tried to grow the bacteria from the tea bag, in the lab?

I went home. I stared at the box of bacteria tea.

Why throw something away when you can tell a story about it?

Part I: The setup

BC30™, the bacterial strain in the tea, is short for Bacillus coagulans GBI-30, 6086®. It received the FDA’s GRAS (Generally Recognized as Safe5) designation in 2012 and is found in over a thousand “leading food, beverage and pet food products worldwide” according to the probiotic’s website.

To coax these bacteria to grow out of steeped tea, I needed to know three things:

  1. Is this species safe to grow in the lab?
  2. What does it like to eat?
  3. What are its preferred growth conditions?

In general, I try not to ingest the bacteria I grow in the lab -- even ones with the lowest safety designation, BSL-1. By nature of it being a commercial probiotic, BC30 is both BSL-1 (safe to grow under normal lab precautions) and edible.

But, I still wouldn’t try this at home or eat bacteria off of a culture plate. Why? BC30’s preferred food source is not that different from the preferred food source of many other microorganisms: a sugar- and amino acid-rich nutrient medium called MRS (De Man, Rogosa and Sharpe) agar.

While MRS agar has some adjustments to make it preferentially appetizing to BC30 and its relatives, those relatives also include Streptococcus pyogenes (causes strep throat) and Bacillus cereus (causes food poisoning). Without sterile technique and rigorous species-level confirmation, you cannot know for sure what is growing on your plate.

With that said, the American Society for Microbiology’s blog suggested that were I successful in culturing BC30, I would see growth of translucent white colonies on MRS agar plates after 48 hours of incubation at 30-33oC in the presence of oxygen.

First, I needed to make tea.

I wanted the conditions of the experiment to represent a range of realistic tea-drinking scenarios, from intended use to flagrant improvisation, and set up three steeps:

  1. The Rule Follower -- Brewed as directed for 4 minutes in boiling water.
  2. “I forgot I made tea” – We’ve all been there. 15 minutes, boiling water.
  3. Cold brew anarchist – Self-explanatory.

Experimental design

Figure 1: Experimental design. Appropriate science-themed ceramic vessels as well as a glass were allocated one fresh, unexpired tea bag each. Eight ounces of water of the indicated temperature were added to the vessel and left to steep for the indicated time. “Temp” indicates starting temperature – final temperature was not measured.

It was at this point I realized I needed a sterile-ish way to transport the steeped tea and tea bags from my house to the lab. Luckily, I had recently run a blindfolded volume pouring accuracy competition at our departmental retreat and had leftover Falcon tubes still in their original package. While the tea was definitely not sterile, I reasoned that a little extra aseptic technique wouldn’t hurt. I poured the tea into the tubes over my kitchen stove, using the open flame as a makeshift Bunsen burner.

A hand holding a test tube with a blue cap on a gas stove  AI-generated content may be incorrect.

No Bunsen burner, no problem

Part II: The lab

In reality, lab came first. I had to make MRS agar plates before I steeped the tea. Our lab does not use MRS broth very often, and when I first looked for some all I found was a 10-year-old solidified block of MRS powder in our stock cabinet that was growing large green spots inside of its glass container. Behind it was one that looked mercifully normal.

I mixed broth powder, agar and water in a glass bottle, loosely capped it, put it in a water bath and took it to the autoclave. “Autoclave” is a nice word for giant pressure cooker. Ours is made by the aforementioned Steris. It rattled and hissed as its jaws opened to accept my tray of culture media, which it then heated to 121oC for 45 minutes, sterilizing the liquid.

Back at my lab bench, when the molten MRS agar had cooled enough to handle, I lit a Bunsen burner next to a stack of empty plastic petri dishes and poured a layer of agar into each one. Left overnight, the plates solidified into nutrient-dense beds for BC30.

The next day, I took my tubes of tea to lab. I pipetted 400 microliters (0.4mL) of each liquid tea condition onto a plate next to the Bunsen burner. I spread the liquid evenly across the plate with a hockey stick-shaped plastic spreader6 and left the lids on the plates cracked open to dry near the flame.

Long live the hockey stick

Long live the hockey stick

But to answer my question, I needed one more test. If there were bacteria in the tea bag initially, but they died when boiled, then I might see bacterial growth by plating the dry ingredients of an unsteeped tea bag, or the tea bag steeped in cold water. I cut open the tea bags and shook some of their contents onto the agar. I put my full set of plates, including a plain MRS plate to check its sterility, into the incubator at 37oC – a standard growth temperature, but a little warmer than recommended. I was skeptical that anything would grow. For the next two days, all I could do was wait and see.

Part III: The results

The first thing I noticed when I took the plates out of the incubator was the smell. I was in disbelief when I saw little white colonies dotting almost all the plates and opened one to get a closer look. A sickly sweet, gingery aroma wafted from the plate as I inspected the translucent colonies – a byproduct of the bacteria metabolizing the sugars in the MRS plate. By all accounts, I was looking at BC30.

A close up of a petri dish AI-generated content may be incorrect.

I wish you could smell this photo

Colonies grew on all of the tea and tea bag plates, while my sterile control plate remained bacteria-free. The colonies from the boiling-water steeps and the tea bags were a variety of sizes, including some that were significantly larger than others, while the colonies from the cold-water tea were uniformly small.

A collage of a person holding a petri dish AI-generated content may be incorrect.

Figure 2: Two days later. Columns are conditions, rows are type of sample. I did not plate the tea bag materials from the 15-minute boiling condition and only plated the cold-brewed tea bag from the end of the cold steeping period.

Because I knew the volume of tea I had put on each plate, I could calculate a standard measurement of bacterial density: colony-forming units (CFUs) per milliliter. Contradictory to my expectations, I saw a five-fold increase in colonies from the tea steeped in boiling water relative to the tea steeped in cold water for the four-minute condition. I saw the same pattern in the fifteen-minute condition, with a nearly four-fold increase in boiling vs cold.

Table 1: Colony count and CFU data.

Before I could draw any conclusions, I needed to know, for sure, that these colonies were Bacillus coagulans. The most robust way to check is by sequencing their DNA, but sequencing is expensive. A simpler, cheaper way to check is with PCR, which amplifies small regions of DNA unique to a species. I downloaded the BC30 genome and selected two regions of its genome that didn’t match other species in the NCBI database. Using Primer3, I generated two pairs of PCR primers, short stretches of DNA to bind to either side of my region of interest.

I picked the largest colony I could see from each plate (seven total), suspended the cells in a small volume of water, and set up standard colony PCR reactions. The heat during the reaction bursts the cells, making the DNA available for amplification. The completed reaction was run through a porous gel with an electric current and visualized with UV. If I saw bands on the gel for both primer sets, from totally different parts of the BC30 genome, I could be confident that this was, in fact, BC30.

I loaded the gel into the imager and hit run. There, in black relief against the grey background of the gel, were my bands.

A close-up of a dna test AI-generated content may be incorrect.

Figure 3: PCR of the same colonies with two different sets of primers confirms BC30’s identity.

Part IV: How to survive boiling water

Bigelow knew something I didn’t7. It turns out that BC30, like many of its relatives, is a spore-forming bacterium. When starved of nutrients, Bacillus coagulans divides asymmetrically, packing its basic cellular information into a spore with a thick protective coat. These spores are resistant to dessication, nutrient starvation, radiation, chemical disinfectants and extreme heat. It was these spores that were in the tea bag — spores that are perfectly comfortable being steeped in boiling water.

Like the seeds of plants, when the spores find themselves in favorable conditions for growth – say, on an MRS agar plate at a balmy 37oC – they germinate back into actively growing cells. This is the premise of their ability to function as a probiotic. The spores are dormant and shelf-stable in a tea bag, or any of the thousand products advertised to contain BC30, and will, in theory, germinate upon arrival in the GI tract, where they can exert some sort of effect on the host that consumed them.

To produce spores at scale, manufacturers grow bacteria in vats of nutrient-rich broth. If the nutrients are not replenished, the bacteria eventually start to starve, triggering the sporulation process. Around 24 hours later, the bacterial broth is treated with enzymes to kill any remaining, non-sporulated cells. The mixture is concentrated, washed with water, and finally, in a fantastic twist of irony, pasteurized.

The pitch for BC30 is that it improves “digestive health” and “protein absorption.” The reported endpoints for digestive health on BC30’s website are reductions in bowel movement frequency, abdominal pain and abdominal bloating in adults with IBS. In the study promoted on the site, the baseline for the placebo group for abdominal pain and bloating is, mysteriously and respectively, 12.5% and 30% higher than the baseline for the BC30 treatment group. The placebo group experienced no change in severity scores over the subsequent course of treatment, while the BC30 group dropped to placebo levels after a week and stabilized.

Data from the IBS study. Although the result would be significantly more convincing if the baseline severity scores were the same in the placebo and BC30 groups, this study reports that it was randomized and double-blind. The severity scores are self-reported by participants, making them difficult to standardize.

For one of the protein absorption studies, there is a small but statistically significant difference in amino acid levels in the blood, including when BC30 is paired with another one of its parent company’s products, a “nutritional milk protein concentrate” called Ultranor.

If these results hold, they beg the question -- could a product like Bigelow’s probiotic tea be able to produce these beneficial effects? Most of the clinical trials I could find, including the IBS study above, dosed people daily over the course of one to eight weeks with 1 billion CFUs (spores) of BC30. Per my calculations, a properly steeped cup of probiotic tea yields around 30,000 CFUs: 0.003% of the clinically tested dose.

Granted, I am one person and this is one experiment. But there are independent, conflicting reports on whether BC30 survives the GI tract at all. One study suggests that Bacillus probiotics don’t make it, while another reports about half of the initial dose of spores surviving transit through an artificial human gut system. Other research suggests that the effect of the probiotic is not even due to the cells coming back to life, but due to an immune response against the dormant or vegetative cells.

These observations have consequences for consumers being parted from their money by unsubstantiated health claims. But they are interesting observations in their own right. Sporulating organisms’ imperviousness to heat, while useful for commercial biotech applications, causes problems for the food industry. The food-poisoning agent Bacillus cereus is a species normally found in the soil. It can release heat-resistant toxins if it multiplies in food, and live bacteria can produce toxins when they reach the small intestine. Even pasteurized milk spoils eventually as heat-resistant spores, mostly soil Bacillus, begin to multiply.

Bacillus coagulans is also a soil bacterium by nature8, and not a typical resident of the community of microbes in our gut (called the gut microbiome). It was discovered in 1915, in canned milk that had spoiled and coagulated. In spite of its origins, BC30 seems inert as a pathogen, and reports of it causing infection are vanishingly rare. BC30’s safety track record is remarkable.

But safety is only the first step on the quest to use probiotics for good. New probiotic companies like Pendulum and Seed market the fact that they are backed by clinical trial data – Pendulum for blood sugar control in Type II diabetes, and Seed for gas, bloating and regularity in healthy adults. The backbones of Pendulum and Seed’s products are organisms found more commonly in the gut, and, interestingly, both companies focus on multi-species products, dosing patients with miniature microbial communities.

One focus of my PhD lab is on abnormal pathogenic behavior of normally harmless bacterial residents of the gut, most commonly in people who are already quite sick. While unhappy microbiomes can be unhappy in their own way, we do not have a consensus on what a “healthy” gut microbiome looks like, either. In collaboration with a continent-wide consortium in Africa, our lab helped catalogue the species found in healthy adult women across the continent. We found over 1,000 new species relative to what had been previously described in studies focused on Western countries.

Companies trying to introduce targeted combinations of microbes into the gut are thus forever shooting at a moving target. Outside of specific indications for GI infections, determining whether to give (or take) a probiotic is a grey area. And for the common GI complaints focused on by the probiotic market, targeting the microbiome with additional organisms may not be the answer at all. Rather, by understanding how bacteria work together in the microbiome, solutions may favor changing the metabolic environment of the gut to drive the formation of species-agnostic “guilds” that perform specific functions, likely via dietary interventions.

I never get tired of growing bacteria. For a colony to be visible on a plate, it consists of at least a million, often closer to a billion, individual cells. Learning how bacteria grow and adapt does not diminish the sense of wonder I feel when I observe them – it only enhances it. I like to think this same sense of wonder animated the scientist who first cultured Bacillus coagulans out of canned milk that had spoiled. And I have to imagine some mixture of wonder, awe and horror kept the Random Hall Milk alive for twenty-seven years.

The next Louis Pasteur could be a lactose-intolerant undergrad, or a procrastinating PhD student. It could be you. Pausing to look a little longer, to ask why the world is the way it is – this is how we upend assumptions of what is valuable. What is worth looking at. Because in the end, trash is in the eye of the beholder.

TAXA is generally recognized as safe.

1

You can read the Milk’s application here.

2

Author correction 7/28/26: Changed from “Napoleon” to “Napoleon III.” A commenter on another site correctly pointed out that Napoleon I was exiled before Pasteur was born. According to the Pasteur Institute, “in 1863, Napoleon III asked [Pasteur] to study diseases in wine.”

3

As demonstrated by the Milk, even pasteurized beverages spoil, a process sped up by exposure to the microbes in the air but which will proceed within an unopened container anyway. How is this possible?

It’s because pasteurized milk is not the same as sterilized milk. Pasteurization heats to ~60C for a few minutes. While the microbes that we worry about causing infection can’t survive this, some heat-tolerant bacteria and proteins can – those are what will eventually break down the milk, even if it isn’t opened to the air. Heating to 140C, on the other hand, makes milk effectively sterile, killing even the heat-tolerant bacteria. But this process, used to produce “ultra-high temperature” or UHT shelf-stable milk, does some odd things to the proteins that subtly change the flavor, color, and texture of the milk.

4

Author correction 7/28/26: I reported this initially as “7 in 10 people were prescribed antibiotics,” but a commenter on another site pointed out that the original report data likely reflects a smaller number of people getting repeat prescriptions, so I have reported the raw data from the CDC report instead.

5

If a probiotic is marketed as a food or dietary supplement, as most are, then it is not required to undergo a clinical trial in the U.S. but instead to submit a GRAS notification. The second most important thing to know about the GRAS system is that it does not require proof of efficacy -- only safety. The most important thing to know is that the proof of safety is provided by the company requesting the GRAS designation. This proof is then reviewed by the FDA, to determine whether the notice provides a “sufficient basis for a GRAS determination” and whether “information in the notice or otherwise available to FDA” raises any safety concerns.

6

This is one of the most polarizing choices one can make as a biomedical research scientist. The alternative to the hockey stick is to use glass beads that you autoclave then sprinkle on the plate and roll around. People are very passionate about their chosen method and will attempt to convert you.

7

Saw this weirdly aggressive Bigelow commercial at the gym. I don’t think they’re going to sponsor me after this article.

8

As our understanding of microorganisms evolves, so do our naming and classification conventions. Bacillus coagulans is a more distant relative of Bacillus cereus and similar soil microbes than previously thought and has been re-classified into a new genus called Weizmannia. Its full nomenclature history can be found on the LPSN.

The Daily Front Page 7 of 25
Tuesday, July 28, 2026 The Daily Front No. #260728 — Fine‑Tuning Beats Frontier (On One Task)
article

A $500 RL fine-tune of a 9B open model beat frontier models on catalog review

by ilreb·▲ 317 points·123 comments·fermisense.com ↗
“Our 9B open‑source model beats every frontier configuration we tested.”

What is the common denominator among companies succeeding with AI-first strategies?

Scatter plot of cost per 1,000 listings versus quality: frontier models cost $19 to $172 per thousand at 70 to 76 percent quality; the fine-tuned 9B reaches 87 percent at about 50 cents per thousand.

The article in one picture: on the same catalog-review workflow, with the same tools, images, and scorer, our GRPO fine-tune of a 9B open-source model (pink) beats every frontier configuration we tested, at $0.50 per 1,000 listings: 40× cheaper than the least expensive frontier setup and ~340× cheaper than the most expensive.

Part I

Everyone asked the same question

Since ChatGPT launched in 2022, business leaders have been asking the same question: what can AI do for us? The answer began with low-risk tasks: summarizing documents, drafting emails, producing first drafts that a human would edit.

It quickly moved into higher-value cognitive work, such as software development and content generation, and grew into more ambitious projects, like attempts to build an AI company brain, a system connected to internal knowledge, data, and tools that could coordinate work and eventually operate parts of the business autonomously.

While a lot of time, energy and tokens have been invested in AI adoption, measurable outcomes have barely been achieved at scale. However, some companies embraced being AI-first and saw enormous gains in productivity, revenue, and cost, while others lagged behind or failed to change their organizations enough to reach high ROI.

Recent data from corporate expense management platform Ramp reveals a stark contrast in performance: the top quartile of companies investing in AI saw their revenue more than double between November 2022 and December 2025, while businesses with zero AI expenditure experienced a mere 15% increase.

Same three years, same economy: in Ramp's data across its customer base, the heaviest AI adopters more than doubled revenue while businesses spending nothing on AI grew about 15% (indexed, Nov 2022 = 100; curve drawn from the reported endpoints). The rest of this article is about what the AI-heavy group actually did.

There are many reasons why AI has done wonders for some companies while others have struggled to see the return on their investment, but research primarily points in five directions.

01

Redesign the process, not just the task

Becoming AI-first means rethinking how the work is structured, not dropping a model into a workflow built around people: what gets approved, who reviews what, and which handoffs still need a human. Where the process stays untouched, legacy bottlenecks absorb the productivity gains before they reach the P&L. In McKinsey's 2025 survey of organizations using gen AI, workflow redesign was the attribute most correlated with EBIT impact, and only 21% of them had redesigned any workflow at all.

02

Incentivize experimentation

Models, tooling and best practices change weekly, so last quarter's setup is rarely still the right one. That only gets picked up if people are rewarded for trying things and reporting what failed, not just for shipping. Technical teams are the natural place to start, since they see the same problems recur across functions and can tell which of them a model can actually take over.

03

Provide tailored business context

Prompt engineering and retrieval can inject business context at call time, but doing it well is its own engineering program: getting to the data, enforcing access controls on what each request may see, building retrieval that surfaces the right evidence, and managing a context window that models use unevenly as it grows.

04

Measure usage and impact

Every AI line item eventually meets the CFO question: what did this change, and was it worth it? In most deployments, nobody can answer it: there is no infrastructure to track the model's performance, decision costs, or impact on efficiency, and self-reported time savings are often inaccurate. Without a scored evaluation on your own data, a "vibe evaluation" is the ceiling of what you can claim, and a hard budget to defend.

05

Set clear business goals within the AI budget

AI brought a pricing model most companies were not used to. Paying per token instead of per seat makes costs scale with usage, which makes it hard to lay out a cost plan or estimate the capital efficiency gains for internal workloads. Uber went through its annual engineering budget in four months, and Microsoft cancelled most of its Claude licenses to bring costs back under control. Today's prices also understate the problem, since most AI labs are subsidizing token costs to capture market share, and frontier model prices are expected to rise.

In this article we give a detailed overview of the deployment technique the winning group keeps converging on: fine-tuning open-source models with reinforcement learning. We cover how it addresses the last three challenges above, and how it turns knowledge only your organization has (namely data, tools, and processes) into a model no vendor API can match at a fraction of the cost.

TL;DR

2.2×

Revenue growth of the top quartile of AI spenders between November 2022 and December 2025 in Ramp's data. Companies with zero AI spend grew about 15% over the same three years, in the same economy: the heavy adopters grew eight times as much.

1

Playbook the winners converge on: an open-source model, proprietary task data, and reinforcement learning against a scored copy of the workflow. Bridgewater's trained model makes ~30% fewer mistakes than the best frontier model, Harvey's legal agent beats GPT-5.5 and Claude Opus 4.8 on its own rubrics, and Intercom's Fin Apex resolves more support issues at lower cost.

87.3%

Share of the maximum achievable score our GRPO-trained 9B open-source model reached on catalog review, vs 76.9% for the best frontier configuration: a 13.5% relative improvement over the frontier, and 36% over its own untrained base (64.2%). The five frontier models, even with optimized prompts, plateaued within a tenth of a point of each other; the trained specialist cleared that ceiling.

68×

Cost advantage per reviewed listing: $0.50 per 1,000 with the specialist vs $34 with the strongest frontier model, and still 40× cheaper than the least expensive frontier option. At roughly 40 million decisions a day, that is about $7M a year instead of $500M, a 98% cost reduction.

Part II

What the winners do differently

Most of the companies pulling ahead in the AI race made the same discovery: owning your intelligence wins on both performance and cost. A model trained to complete your specific workflows in your specific environment is very likely to outperform a general-purpose model that has never seen inside your company. Additionally, since you do not need to pack as many general-purpose capabilities into a model that is meant to operate in a specific environment, you can often get away with a smaller model that is orders of magnitude cheaper to run.

Owning your intelligence does not mean cancelling the ChatGPT or Claude subscription. Most workflow automation still starts with frontier models, and that is the right first move: it establishes a baseline for what is technically possible, and every call generates the data (inputs, decisions, corrections) that a specialist model later trains on. Once the automation leaves the prototyping stage, the priority flips to cost and performance at volume, and that is where fine-tuning open-source models with reinforcement learning comes in. In addition to that, your Fable 5 or ChatGPT model can call the specialist model to handle the parts of the workflow that require your internal knowledge, and the specialist model can call the frontier model for tasks that require high general ability.

How a model learns to operate in your environment

Diagram: a task goes to an AI model running on your infrastructure; the model interacts with tools and data, a rubric scores the outcome, and the reward signal updates the model.

Part I

Everyone asked the same question

Since ChatGPT launched in 2022, business leaders have been asking the same question: what can AI do for us? The answer began with low-risk tasks: summarizing documents, drafting emails, producing first drafts that a human would edit.

It quickly moved into higher-value cognitive work, such as software development and content generation, and grew into more ambitious projects, like attempts to build an AI company brain, a system connected to internal knowledge, data, and tools that could coordinate work and eventually operate parts of the business autonomously.

While a lot of time, energy and tokens have been invested in AI adoption, measurable outcomes have barely been achieved at scale. However, some companies embraced being AI-first and saw enormous gains in productivity, revenue, and cost, while others lagged behind or failed to change their organizations enough to reach high ROI.

Recent data from corporate expense management platform Ramp reveals a stark contrast in performance: the top quartile of companies investing in AI saw their revenue more than double between November 2022 and December 2025, while businesses with zero AI expenditure experienced a mere 15% increase.

Same three years, same economy: in Ramp's data across its customer base, the heaviest AI adopters more than doubled revenue while businesses spending nothing on AI grew about 15% (indexed, Nov 2022 = 100; curve drawn from the reported endpoints). The rest of this article is about what the AI-heavy group actually did.

There are many reasons why AI has done wonders for some companies while others have struggled to see the return on their investment, but research primarily points in five directions.

01

Redesign the process, not just the task

Becoming AI-first means rethinking how the work is structured, not dropping a model into a workflow built around people: what gets approved, who reviews what, and which handoffs still need a human. Where the process stays untouched, legacy bottlenecks absorb the productivity gains before they reach the P&L. In McKinsey's 2025 survey of organizations using gen AI, workflow redesign was the attribute most correlated with EBIT impact, and only 21% of them had redesigned any workflow at all.

02

Incentivize experimentation

Models, tooling and best practices change weekly, so last quarter's setup is rarely still the right one. That only gets picked up if people are rewarded for trying things and reporting what failed, not just for shipping. Technical teams are the natural place to start, since they see the same problems recur across functions and can tell which of them a model can actually take over.

03

Provide tailored business context

Prompt engineering and retrieval can inject business context at call time, but doing it well is its own engineering program: getting to the data, enforcing access controls on what each request may see, building retrieval that surfaces the right evidence, and managing a context window that models use unevenly as it grows.

04

Measure usage and impact

Every AI line item eventually meets the CFO question: what did this change, and was it worth it? In most deployments, nobody can answer it: there is no infrastructure to track the model's performance, decision costs, or impact on efficiency, and self-reported time savings are often inaccurate. Without a scored evaluation on your own data, a "vibe evaluation" is the ceiling of what you can claim, and a hard budget to defend.

05

Set clear business goals within the AI budget

AI brought a pricing model most companies were not used to. Paying per token instead of per seat makes costs scale with usage, which makes it hard to lay out a cost plan or estimate the capital efficiency gains for internal workloads. Uber went through its annual engineering budget in four months, and Microsoft cancelled most of its Claude licenses to bring costs back under control. Today's prices also understate the problem, since most AI labs are subsidizing token costs to capture market share, and frontier model prices are expected to rise.

In this article we give a detailed overview of the deployment technique the winning group keeps converging on: fine-tuning open-source models with reinforcement learning. We cover how it addresses the last three challenges above, and how it turns knowledge only your organization has (namely data, tools, and processes) into a model no vendor API can match at a fraction of the cost.

The Daily Front Page 8 of 25
Tuesday, July 28, 2026 The Daily Front No. #260728 — Inside Zig’s Incremental Compiler
article

Zig's Incremental Compilation Internals

by garyhtou·▲ 226 points·159 comments·mlugg.co.uk ↗
“Recompile only that code, and directly patch the resulting bytes.”

As a member of the Zig core team, one of the most impactful projects I’ve been involved with is the implementation of incremental compilation into the Zig compiler. This feature allows the compiler to detect which individual functions and declarations have changed since a project was last built, recompile only that code, and directly patch the resulting bytes into the output binary, making the rebuild extremely fast.

The Zig project has been working towards this feature for a long time, and over the last few release cycles, it has finally gone from a proof-of-concept quality feature to one which is viable for real-world projects and which most of the Zig core team makes daily use of.

Today, using Zig’s incremental compilation, you can make changes to real, complex applications in a matter of milliseconds.

But don’t just take my word for it! Here’s a simple video (no audio) demonstrating me using Zig to quickly make and test some changes to Fizzy, a pixel editor application. The initial build takes around 5 seconds, and then every time I make a change, a rebuild completes in 50–70ms.

For this demo, I had to upgrade Fizzy to Zig’s master branch. This is because while Zig 0.16.0 does have support for incremental compilation, it is missing some important linker features which have since been implemented. This means that if you prefer to stick to tagged releases of Zig, you likely won’t be able to try this out until 0.17.0 drops; sorry!

Fast incremental rebuilds for some random changes to Fizzy

If you’re already convinced and just want to know how to use this, great! Head on down to the last section of this post to find out. But perhaps you’re understandably skeptical that this is applicable to most projects, or, like me, you just enjoy learning how stuff like this works. For all of you folks, let’s dig into the details!

Processing Source Files

The Zig compiler’s pipeline can be split up into a few different parts, which we’ll look at in order. The first part works at the granularity of entire source files, and basically consists of running the following process in a loop:

  • Read in a source file from disk
  • Parse that file into an AST
  • Convert that AST into a format named “ZIR” using a pass named “AstGen”

If you’re curious, ZIR (Zig Intermediate Representation) is an untyped SSA-form IR—but don’t worry if you have no idea what that means, because it won’t really matter here. All we care about is that we’re converting an entire source file into a different format.

While AstGen runs, it learns about all Zig imports (@import("foo.zig")) in the source file, so we can repeat this entire process on all of the imported files. So by running this process in a loop, we will ultimately discover every Zig source file in the compilation, and will convert them all to ZIR.

File processing pipeline in the Zig compiler

This part of the pipeline actually has several useful properties:

  • The processing run on each file is a pure function of that file’s contents, involving no shared or external state
  • Parse and AstGen are both quite fast on their own: on my laptop, running them both over the entire src/ directory of the Zig compiler (with no parallelism at all) takes around 920ms
  • Thanks to Zig’s usage of data-oriented design patterns, ZIR can be trivially written to and read from disk with one writev/readv system call—there is no “serialization” step.

These properties have two nice consequences.

Firstly, assuming one “task” per source file, this entire process is embarrassingly parallel. That means we can trivially run it on a thread pool by queuing up a task every time we discover a new source file from an import—the only shared state (which we’ll just protect with a mutex) is a hash set keeping track of which file paths we have already seen.

Secondly, and arguably even more importantly, these properties make it very straightforward to implement incremental compilation for this part of the pipeline. All we need to do is cache each source file’s generated ZIR on disk, and only rebuild it when we detect that the file changed.

Both of these optimizations have been enabled by default in Zig for years—they are battle-tested and make this part of the pipeline near-instantaneous in most cases. If you’re using Zig, you can see how fast this is using the progress output on stderr—when it says “AST Lowering”, this part of the pipeline is running. I’d guess that a lot of Zig users only even notice that happening the very first time they run the compiler (because on its first run the compiler needs to do this work for the entire Zig standard library and compiler_rt).

Okay, so, we made this part fast! That’s great, but the bad news is that this was the easy part—lots of compilers can already do this kind of caching. From here, things will get trickier.

Semantic Analysis

The next part of the pipeline is arguably the most important: semantic analysis. This includes both type checking and comptime evaluation.

The job of semantic analysis is essentially to “interpret” the ZIR we produced earlier, emitting compile errors (such as type errors) along the way; and, for runtime functions, building another intermediate representation which can be sent on to later parts of the pipeline.

Before we move forward, a quick terminology clarification. A “container-level declaration” is the Zig equivalent of what other languages call a “top-level declaration”. That term is inaccurate in Zig, because container-level declarations do not have to be at the top level syntactically, but the concept is the same. If I say “container-level declaration”, I basically mean “a function, global constant, or global variable”.

Semantic analysis is the most difficult part of the compiler to handle incrementally. Perhaps unsurprisingly then, this is where language design starts to matter a lot: while I am pretty confident that most modern languages could support incremental compilation similar to how we do, certain design decisions can make that much more difficult. Zig has had its design tweaked over the years (sometimes controversially) specifically so that it is easier to support fast incremental compilation.

The name of the game here is to split up your compilation into a bunch of pieces which you can mostly analyze independently of one another, and, crucially, where the dependencies that do exist between those pieces can be easily modeled in a dependency graph.

In the Zig compiler, we call these pieces “analysis units”, or I might sometimes just say “unit” for short. I’m going to ever so slightly simplify things here and tell you that the Zig compiler has four different kinds of analysis unit:

  • The layout (size, alignment, etc) of a struct or union type.
  • The type of a container-level declaration.
  • The value of a container-level const declaration.
  • The body of a runtime function.

During semantic analysis of a particular unit, we populate a set of other units which this unit depends on. Let’s look at a basic example:

var global_0: u32 = 123;
const global_1: u32 = 456;

pub fn foo(cond: bool) u32 {
    if (cond) {
        return global_0;
    } else {
        return global_1;
    }
}

Here’s what happens when we analyze the body of the function foo:

  • Because the argument cond is not comptime-known, we semantically analyze both branches of the if

  • Take a pointer to global_0, in preparation to load from it

    • Add dependency: type of global_0
  • Load global_0 at runtime, because it is var so does not have a comptime-known value

  • Take a pointer to global_1, in preparation to load from it

    • Add dependency: type of global_1
  • Load global_1 at compile time, because it has a comptime-known value

    • Add dependency: value of global_1

So we end up with this function body depending on the types of global_0 and global_1, and the value of global_1 (since that’s comptime-known). This tells the compiler that if the type of global_0 or global_1 changes, or the comptime-known value of global_1 changes, the function should be re-analyzed.

Dependencies on the body of a runtime function are impossible (at least in the simplified view I’m presenting here). This means that function body analysis units can only have “outgoing” edges in the dependency graph (i.e. they may depend on other units, but other units do not depend on them).

Dependencies on the value of a const declaration only arise due to Zig’s ability to use those at comptime. If not for that language feature, dependencies on the value of a declaration would be impossible, just as it is impossible to depend on the body of a runtime function.

Dependencies on a type’s layout arise, in short, from having values of that type, or from needing to know something about the type’s layout. I’m not going to discuss this any further here, because it’s a bit complicated and quite specific to Zig’s type system, but it’s not fundamentally different.

Okay, so, we’ve told the compiler about when re-analysis of one thing needs to also trigger re-analysis of another thing. However, there’s one more puzzle piece here—source code dependencies. By itself, this dependency graph is useless: what do we actually do when the user asks for a recompile (what we call an “incremental update”)? We don’t know the first thing to re-analyze!

To solve this problem, we track dependencies of analysis units, not only on other units, but also on pieces of source code. In the cases we’ve looked at so far, these are all really simple: in the snippet above, the units “type of global_0” and “value of global_0” both depend on the source code of global_0, the unit “type of global_1” depends on the source code of global_1, and the unit “body of foo” depends on the source code of foo. Whenever any byte of source code in the given region is modified, the dependent analysis unit will be marked as “outdated” and re-analyzed.

Note that in reality, things can get more complicated than each unit depending on one piece of source code. For example, an inline function call in Zig performs semantic inlining, which means that it essentially triggers semantic analysis of a different piece of code but in the caller’s analysis unit. Therefore, inline function calls introduce dependencies from the caller’s analysis unit on the source code of the callee.

Of course, we still need to be able to figure out which regions of source code have changed since an incremental update. For this, ZIR contains hashes for specific “interesting” regions of source code (e.g. the entire source code for each container-level declaration), and those hashes are what you are actually depending on. If the source code changes, the hash changes, and that’s easy for the compiler frontend to detect.

Okay, that was a lot of explaining—now let’s look at some pretty pictures! Here’s some Zig source code:

const lucky_number = 42;

const S = struct { x: u32 };

fn getSomething() S {
    return .{ .x = lucky_number };
}

fn testLuck(x: u32) void {
    if (x == lucky_number) {
        // do something
    }
}

export fn entry() void {
    const result = getSomething();
    testLuck(result.x);
}

…and here’s its dependency graph (with some redundant edges removed for legibility). The nodes on the right represent the source code which has been hashed, while the remaining nodes are all analysis units.

Dependency graph generated during initial build

Now, let’s say we change the first line of the file to read const lucky_number = 43;. First, the compiler lowers the new ZIR for this file. It maps declarations from the old ZIR to the new ZIR based on the declaration names, and compares the source hashes associated with each declaration. In this case, it successfully maps every declaration, and it sees that one source hash changed—the one associated with const lucky_number. Next, it looks at the dependency graph to find everything which depends on that source hash. In this case, it only finds one direct dependency:

Invalidated dependencies after change to source hash

So, the compiler re-analyzes the value of the lucky_number declaration. If our change to the line had been a no-op (e.g. we just added some whitespace), then it would determine that the value did not change, and stop here. But in this case, the value did change! Therefore, the compiler continues this process, by next considering any analysis units which depend on the value of lucky_number, of which there are two:

Invalidated dependencies after change to lucky_number value

The compiler analyzes those two units—the bodies of testLuck and getSomething. There are no dependencies on these units (since they’re function bodies), so the semantic analysis loop stops here. However, semantic analysis of those functions does generate new AIR, which brings us neatly to our next topic: code generation.

Code Generation

Code generation, sometimes called “codegen” for short, is the stage in the compiler pipeline where AIR from semantic analysis is converted to something resembling machine instructions. Codegen doesn’t quite emit machine instructions yet—instead it’s something called MIR (Machine Intermediate Representation)—but there is almost a 1–1 mapping between MIR instructions and machine instructions. There are separate codegen implementations for each target architecture (x86_64, aarch64, etc).

A nice thing about codegen is that just like the whole-file processing earlier, it is an embarrassingly parallel task (at least in builds where you aren’t doing inter-function optimizations like inlining). There is no state shared between code generation of different functions, so we can have a queue of pending functions whose AIR needs converting to MIR, and process that queue across arbitrarily many threads. There’s just one small gotcha, which is that we need to be careful to cap the size of that queue, because if codegen is ever running behind semantic analysis for any reason, the size of the queued-up AIR can add up fast!

In terms of incremental compilation, this phase of the pipeline is actually as simple as it gets, because AIR and MIR both exist at the granularity of individual functions, which is the same granularity incremental compilation works at. This means that there is no need for the compiler to cache AIR or MIR at all! The AIR is thrown away as soon as code generation is done, and the MIR will be thrown away right after it’s consumed by our next stop: the linker.

Linking

Incremental linking is kind of a difficult problem, and I suspect is a big reason that no other major toolchain supports this kind of incremental compilation yet. General-purpose incremental linkers aren’t really a thing at the moment, and though wild was originally conceptualized as one, that project seems to have shifted its focus firmly towards cold-link performance over the past couple of years, with no explicit timeframe for incremental linking.

David Lattimore, the creator of wild, has a blog post discussing some of the difficulties of incremental linking. One of those is diffing input objects to figure out what actually changed on an update. However, when you control the entire compilation pipeline, a simpler design presents itself which neatly sidesteps that entire problem: tightly integrating the linker with the compiler.

To begin with, let’s just look at how the linker might work without incremental compilation. Because linking involves a lot of shared state, our linker is entirely single-threaded (maybe we’ll look into multi-threaded linking in the future, but for now we’re keeping things simple). When the linker receives MIR from codegen, it first needs to convert that MIR into the actual machine code. This logic is specific to the codegen backend, but we can’t run it until now because it requires cooperation with the linker. That’s because while emitting machine code, the codegen backend generates relocations—basically, instructions for the linker to overwrite certain parts of the code with specific addresses or values (for instance the address of another symbol). The linker needs to save all of these relocations internally, so we need to be on the linker thread for this.

After generating the machine code and associated relocations, we reserve space for that machine code in the output section (usually .text). We save the machine code in a buffer, save the relocations to apply later, and do some miscellaneous bookkeeping work, such as adding a symbol table entry.

For a non-incremental linker, this would be the end of the story. At the end of compilation, we would assign addresses to every section, write out everything we reserved space for, and apply all of the relocations. Incremental linking is a bit trickier—writing the machine code to the file, assigning addresses, and applying relocations, all ideally needs to happen before we know the full contents of the binary, and we need to be able to update those things later.

A lot of the complexity here is actually just in moving things around. For example, if we want to add a function to the .text section, but there isn’t enough space, we need to expand that section. But the section might be surrounded by other sections, which we can’t just overwrite, so we’ll need to move something—either the .text section itself, or one of the surrounding sections. In doing so, we’re going to change not only file offsets but also virtual addresses of everything we move—this means we’ll need to update symbol table addresses, re-apply relocations, etc. That’s a lot to keep track of! (There’s also a similar problem for segments, one level up.)

To solve this problem, Jacob Young introduced a nifty abstraction into the Zig compiler called link.MappedFile. It memory-maps the output file, but more importantly tracks a tree of “nodes” in that file. The root node covers the entire file, and child nodes refer to specific regions within their parent node. The API user can add nodes, or grow a node to a given size—in both cases, if there is not space in the parent to trivially perform the operation, MappedFile deals with moving other nodes around to make space. Whenever it resizes or moves a node, the implementation sets a “dirty” flag on that node, so that at some point the linker implementation can detect this and apply any necessary fixups, e.g. re-applying relocations whose target moved.

Resizing a node in MappedFile

Right now, MappedFile has fairly primitive logic for node allocation, so sometimes makes suboptimal decisions—but because we’ve abstracted it behind a neat little API, we can improve it independently going forward.

To get to incremental linking, then, we need only slightly change the process I described earlier. After we finish emitting machine code, we create a node in the mapped file, large enough to hold the code—or if this function already existed, we just resize the existing node—and we copy the machine code into it. Allocating this node in the file might (in rare cases) need to move some other stuff in the file around, in which case the appropriate “dirty” flags are set on those nodes. We always set the “dirty” flag for the function’s node itself, so that its relocations will be applied at some point.

Because of the pending relocations, we probably don’t have a valid binary right now—but that’s okay! When the linker thread is next idle (i.e. its work queue is empty), or at the end of compilation if the linker thread remains busy until then, we’ll check all of those “dirty” flags and clean up after ourselves. This could involve work such as assigning new virtual addresses, updating the section headers and program headers, updating addresses in the symbol table, and re-applying relocations.

It might sound like that “fixup” work is expensive. Sometimes, it can be—if you’re creating a dynamic executable and the PLT has to move, that can take a moment, because there are usually a lot of relocations targeting the PLT. However, most of the time, we don’t need to move anything! By using exponential growth factors on nodes (similar to how dynamic data structures like ArrayList work), we amortize this cost and make it extremely rare in reality (at the cost of a slightly increased binary size, which isn’t usually a major concern during development). This design means that you might very occasionally see one update run slightly slower than usual (maybe a few hundred milliseconds?), but I’ve not personally hit this a single time, despite using incremental compilation with this linker near-daily for the past couple of months.

Flush

Okay, we’ve made it to the end, and kept everything incremental along the way. Files were lowered to ZIR with a simple per-file cache; semantic analysis of declarations kept track of a dependency graph to figure out what might have changed; code generation re-ran only for updated functions; and our linker wrote new code into the file without changing any other bytes. We just have a few more loose ends to tie up.

Firstly, because of how Zig’s “lazy analysis” feature interacts with incremental compilation, we need to do a graph traversal to figure out which functions/declarations/etc are actually referenced. It’s possible that something was referenced on a previous incremental update (so we compiled it), but has since become unreferenced, which means we need to ignore any compile errors it emitted, not perform symbol exports from it, etc. There are probably some optimizations you can do here, but at least right now, we just traverse the full reference graph on every update. We can get away with this even on big projects, because computers are really fast!

Once we’ve figured out what’s referenced, we can tell the linker every global symbol which is exported from Zig code, so that it can add any necessary entries to the symbol table. We will also report compile errors if there are any, and some other miscellaneous tasks like that. Finally, we call the linker’s flush function, whose job is just to do any remaining linking work before the file is closed. If the linker has any MappedFile node still marked as “dirty”, we’ll need to handle that, but otherwise we want to do as little work as possible—remember, anything we do here is going to happen on every update, so we want to keep it pretty much O(1). Therefore, all that the ELF linker really does here is write out the .dynamic section, and write the entry field in the ELF header.

We then close the file, and the compilation is complete!

Tracing an Update

Explanations are cool and all, but we can actually see this happening. Tracy is a real-time profiler—it’s designed for games, but you can integrate it into anything. The Zig compiler has optional Tracy integration, enabled using a build flag. (I guess incremental updates kinda resemble frames in a video game if you squint?)

This can occasionally be useful for various compiler performance analysis, but I actually find it really cool to use for incremental compilation, because we can see the different parts of the compiler pipeline clear as day. Let’s take a look at the Tracy output for a change to Fizzy, much like the changes in the video from earlier. This particular update took 37ms (a little faster than the ones we saw in the video), but at first I’m going to zoom in on the first 6ms or so of this 37ms update—I’ll explain why later.

First 6ms of an incremental update, visualized in Tracy

At the start we can see a flurry of activity across all threads—that’s the thread pool doing all of the per-file work. Although we only changed one file, the compiler doesn’t assume that, and instead checks every source file. There’s one small optimization here, which is that because we already know which source files were in the compilation on the last update, we can guess that those files will all still be reachable and so check for changes to all of them. That just means we don’t need to wait for the first file to be processed so that we can discover its imports.

Next we see a good chunk of time (around 1ms) in computeAliveFiles. This function is traversing the graph of file imports to assign every file to a Zig “module” (because it’s possible for a file to move from one module to another between updates). We also use this import traversal to check whether all source files are, in fact, still in the compilation. If any are not—because all imports of them were removed—then we’ll basically just ignore those files for the rest of this update.

Then we have another millisecond in updateZirRefs. This function is responsible for correlating the old and new ZIR of any changed files, and updating all internal references to ZIR instructions to refer to the instruction’s index in the new ZIR rather than its index in the old ZIR. This is a fairly simple task, but the current implementation involves iterating every ZIR instruction we hold a reference to at all and completely rebuilding a hash map’s metadata. This can probably be optimized.

Now we’re done with the single-threaded per-file stuff, and we can finally get onto the meat and potatoes of the pipeline: semantic analysis, codegen, and linking. The “sema_loop” zone contains all of the time spent in semantic analysis—around 1.2ms. We then see that function’s AIR get picked up by codegen on a different thread (the green zones named runCodegenInner), which runs for around 240us. The resulting AIR is picked up by the linker thread and emitted to the binary—this linking work (the purple emitFunction zone and the little green zones next to it) takes around 170us. Overall, this entire part of the pipeline—which by far dominates cold builds—comes in at around 1.6ms for this update. Not bad!

After that’s all done, we’re onto flush. The little purple zone at the bottom-right is a small bit of linking work the frontend requests during flush: regenerating a lookup table which we use to implement Zig’s @errorName builtin. That takes around 50us, and brings us to the end of the 6ms region I’ve zoomed in on. That means it’s finally time to zoom out and see what the remaining 31ms are…

Full 37ms incremental update, visualized in Tracy

Basically all of the remaining time is spent in one function, resolveReferencesInner. Remember a bit earlier I mentioned doing a graph traversal during flush, to determine which Zig declarations are referenced? Well, that’s this function’s job! It’s not inefficient by any means, but that graph is kinda big, so it’s perhaps unsurprising that it starts to matter when we’re trying to go fast.

On the one hand, this seems pretty silly, so much so that I seriously considered trying to improve it before putting out this blog post (after all, a 7ms time is more impressive than a 37ms time). This is amplified when you consider that the reference graph didn’t actually change here. So the vast majority of the duration of this incremental update is being spent figuring out that a graph didn’t change!

But actually, I think this is really cool, because it shows how much efficiency is still left to squeeze out. This 30ms zone is realistically not a big issue, but we can get rid of it nonetheless—firstly by avoiding recomputing this data when the reference graph is unchanged, but also by only recomputing what we need to when the references do change (that problem is called “dynamic single-source shortest path” and is a fairly well-studied problem in graph theory).

Basically, we’re far from done on the performance front! If you want to keep up with what we’re doing in the future, you might consider adding the Zig devlog to your RSS reader, or checking out the release notes when new versions of Zig are released.

Using Incremental Compilation

Okay, I’ve been rambling about compilers for long enough; let me actually show you how to use this thing. I’ll assume you have a Zig project with a build script, and that it compiles on a recent master branch build of Zig (or, if you’re reading this after Zig 0.17.0 releases, that’ll also work.)

At the time of writing, this will only really work if you target x86_64-linux, because our other code generation and linker backends are not mature enough yet. The majority of the Zig core team runs Linux on x86_64, so by focusing on it first, we’ve sped up our workflows, meaning it’ll be faster for us to add support for other targets—which is now top priority!

The bad news is that right now, this isn’t zero-effort. Eventually it will be—we’ll cache all of the compiler state to disk and automatically reload the last saved state when you run zig build, so incremental compilation will just happen automatically—but we’re not quite there yet. However, the good news is that using it today requires very little work!

The short version is that you just need to run this command:

$ zig build --watch -fincremental

The --watch argument tells the Zig build system to watch the filesystem for changes to your source files, and trigger rebuilds when they happen. The -fincremental argument tells the build system that when it does one of those rebuilds, it should use incremental compilation.

When you run that command, you may notice that even if you have a warm cache, every executable/library/object in your project will be rebuilt anyway. That’s expected behavior, because the existing caches on disk are not compatible with incremental compilation. However, if this is inconvenient, then you can modify your build.zig to ask the build system to only use incremental compilation for a specific compilation:

// expose a '-Dincremental' option
const incremental = b.option(bool, "incremental", "Enable incremental compilation") orelse false;
// if it was given, enable incremental compilation for 'exe'
if (incremental) exe.incremental = true;

Then you can use -Dincremental instead of -fincremental, and it’ll only do the full initial rebuild for that particular step. (The -fincremental option was basically telling the build system to set this flag on every std.Build.Step.Compile.)

Anyway, after the initial build is done, just edit one of your source files, save it, and you should see the results of the rebuild instantly. Assuming there are no compile errors, you’ll find the updated executable in zig-out/ as usual. That’s all there is to it!

If you combine --watch with a build step which runs your program (e.g. zig build run --watch -fincremental), your program will run every time the build finishes. For short-running programs, that’s probably exactly what you want. However, for long-running programs (e.g. graphical applications), be aware that right now, the build system will only be able to trigger incremental rebuilds after the previous build of your program closes. Depending on your personal workflow, that might be fine for you (perhaps you’ll just close the application every time you want a build), but if not, you might prefer to do a normal build and manually run the program from zig-out/. If this is you, don’t worry—we’re already planning a bunch of enhancements to the build system to unlock other workflows!

As with all of the Zig project, incremental compilation is not yet stable. While it works pretty well, there are definitely bugs right now, probably including some false-positive compile errors and even miscompilations. If you run into any problems while using incremental compilation, please do open an issue on the Zig repository if you can—the more bugs we’re told about, the more we can try to get fixed for the next release!

Thank you to all of our users who have helped to try out this feature so far, and to anyone who tries it after reading this post—it’s great to see this working for so many people. Thanks also to everyone who donates to the Zig Software Foundation: it’s a true privilege to be able to spend my time working on cool stuff like this.

And of course, thanks for reading :^)

The Daily Front Page 9 of 25
Tuesday, July 28, 2026 The Daily Front No. #260728 — SBCL 2.6.7
article

Steel Bank Common Lisp version 2.6.7

by tmtvl·▲ 226 points·94 comments·sbcl.org ↗
“New SBCL versions are usually released at the end of each month.”

New SBCL versions are usually released at the end of each month: check the Sourceforge File List to see the current version. The new features of all SBCL releases are listed below.

New in version 2.6.7, 2026-07-28

  • new contrib module: SB-MANUAL contains the SBCL manual in docstrings of section definitions, which tie together the docstrings of normal Lisp definitions. The manual can thus be explored interactively in the usual way (e.g. with Slime's M-.), and it is browsable with the MGL-PAX library (out of tree). Also, https://fixnum.com/ (similarly unrelated to the SBCL project) provides alternative renderings of the SBCL manual as heavily linked PDF and HTML documents as well as Markdown and plain text.

  • new feature: DOCUMENTATION supports DOC-TYPE DECLARATION.

  • platform support:

    • the SB-SIMD contrib now supports ARM64. (Thanks to Sylvia Harrington)
    • AVX512 instructions are now supported on X86-64. (Thanks to Robert Smith and Arthur Miller)
    • additional support for SIMD instructions on ARM64 and X86-64. (Thanks to Arthur Miller)
    • fix miscompilation of SAP-REF-N on ARM64. (Thanks to Hayley Patton)
    • implement INTEGER-LENGTH on primitive types without a loop on MIPS and LoongArch.
  • bug fix: READing with *READ-SUPPRESS* T no longer emits warnings like "<internal-feature> no longer present on *FEATURES*".

  • bug fix: compiler type-error when compiling calls to CONCATENATE with conditional known non-sequence arguments. (#2160747)

  • bug fix: (EQL <complex>) types were not being treated as numeric by the type system. (#2160429)

  • bug fix: improve the handling of quiet (non-signalling) NaN inputs to LOG. (#2160268, reported by Woodrow Kiang)

  • bug fix: miscompilation of MULTIPLE-VALUE-CALL. (#2160207, reported by Vasily Postnicov)

  • optimization: passing constant complex numbers to local functions can be done without consing.

  • optimization: where available, use enhanced SIMD routines for UTF-8 conversions.

  • optimization: compiler transforms of COUNT are applicable with a wider variety of keyword arguments.

  • optimization: remove at least one redundant instruction from SB-ALIEN:DEREF.

  • optimization: the sparse set implementation in the compiler has been tuned to improve performance on real-world workloads.

  • documentation: many typos and typesetting issues were fixed.

  • documentation: SB-MANUAL:@FOREIGN-FUNCTION-INTERFACE now correctly states that arrays are row-major (not column-major). (#2158033, thanks to Scott L. Burson)

  • documentation: internally, docstrings now conform to a subset of Markdown, but DOCUMENTATION (and thus DESCRIBE) strips some of this markup. The official manual is still generated from Texinfo, but the Texinfo files are generated from SB-MANUAL.

  • documentation: the manual now has a separate index for declarations.

New in version 2.6.6, 2026-06-28

  • minor incompatible change: FDEFINITION now returns the outermost wrapper (added e.g. by TRACE, PROFILE) like SYMBOL-FUNCTION. (#799533)

  • minor incompatible change: in unsafe code, C strings with :EXTERNAL-FORMAT :ASCII are copied directly as byte-sized quantities without checking whether the top bit of the byte is set.

  • platform support:

    • fix the build on big-endian 64-bit PowerPC with ELFv2. (thanks to Piotr Kubaj)
    • move the static space address for macOS 27 on ARM64. (#2156072, reported by Gary Palter)
    • optimizations to SB-THREAD:BARRIER for ARM64. (thanks to Sahil Kang)
    • fix a compiler crash in MULTIPLE-VALUE-LIST in argument forms on ARM64. (#2155788, reported by Gary Palter)
  • bug fix: TRACE no longer fails when trying to print a return value that cannot be printed readably and *PRINT-READABLY* is true.

  • optimization: the compiler is more precise in its type derivation of COERCE given constraints on its inputs.

  • optimization: the compiler is better able to derive the return types of AREF and ELT.

  • optimization: faster encoding and decoding of UTF-8 C strings.

  • optimization: (length (intersection a b)) doesn't cons an intermediate list.

  • documentation: the manual now includes a section for SB-INTROSPECT, which has also seen improvement in its documentation strings and comments.

  • documentation: fixed many typesetting problems and typos in the user manual.

New in version 2.6.5, 2026-05-29

  • minor incompatible change: the condition signalled when an accessed slot is missing from an object is no longer a TYPE-ERROR.

  • minor incompatible change: the condition signalled when accessing an uninitialized structure slot is no longer a TYPE-ERROR.

  • minor incompatible change: the implementations of standardized functions treating lists as sets, such as INTERSECTION and UNION, take more advantage of the freedom to return the elements of the result in any order.

  • platform support:

    • add low-level support for floating point state manipulation on PPC64/FreeBSD. (thanks to Piotr Kubaj)
    • improve the software emulation of displaced instructions on ARM64.
    • restore building the system using the musl C library. (#2153432, reported by Tom Gillespie)
    • fix some SB-SIMD shifting instructions on AVX2. (#2152791, reported by Willem Broekema)
  • enhancement: definition sources for alien callbacks are now findable by name in SB-INTROSPECT.

  • bug fix: the SYMBOL restart for finding packages now actually performs a non-local exit. (#2153092, reported by Zach Shaftel)

  • bug fix: TYPEP on large disjoint numeric union types compiles faster using fewer resources. (#2151818, reported by James McDonald)

  • bug fix: strings of arbitrary size with fill-pointer set to 1 are character designators. (reported by _death)

  • bug fix: the KEEP-OLD restart established by ADD-PACKAGE-LOCAL-NICKNAME keeps the old nickname instead of going ahead with the change (and the restart report function no longer returns from ADD-PACKAGE-LOCAL-NICKNAME).

  • bug fix: when EXPORT results in a conflict between symbols exported by different used packages, the TAKE-NEW restart now shadowing-imports the new symbol instead of doing nothing and leaving the package in an inconsistent state.

  • bug fix: the SB-EVAL interpreter checks program syntax more thoroughly.

  • bug fix: compiler cross-reference data is decoded correctly for a functional with more than one entry for a given name.

  • bug fix: TYPE-ERRORs signalled by SBCL are more likely to have a DATUM that is not of the condition's EXPECTED-TYPE.

  • bug fix: the code walker no longer uses the stack to walk PROGN bodies.

  • optimization: in various situations, INTERSECTION and UNION will use hash-tables to perform the operation more quickly.

New in version 2.6.4, 2026-04-29

  • minor incompatible change: when DEFSETF is called on a name that was previously used as a (presumed) call to a function, it issues a single STYLE-WARNING (like DEFMACRO).

  • minor incompatible change: SB-EXT:PROCESS-KILL no longer supports the :PTY-PROCESS-GROUP option (which was never correctly implemented).

  • minor incompatible change: the :INITIAL-OFFSET argument for typed DEFSTRUCT, if given, no longer accepts NIL.

  • platform support:

    • more likely to support 32-bit linux's struct timeval with 64-bit time_t.
    • the runtime's main function is now defined as a weak symbol for platform/compiler combinations that support it.
    • on Windows, individual empty arguments for RUN-PROGRAM are escaped.
    • add input/output speed fields for our definition of the termios structure, addressing a crash in sb-posix:tcsetattr on OpenBSD. (#2150499, thanks to Robert Palm)
  • bug fix: address infinite loops in the compiler. (#2144911, #2148056)

  • bug fix: if an FTYPE has been proclaimed for a function, don't mix NULL with explicitly-typed keyword arguments. (#2147050, reported by Vasily Postnicov)

  • bug fix: compiler error from treatment of unused results. (#2147383)

  • bug fix: compiler error from invalid dimension arguments to MAKE-ARRAY. (#2147384)

  • bug fix: compiler error arising from continuing to optimize deleted nodes. (#2147385)

  • bug fix: make sure linkage-table alien entries have base-string names. (#2147646, thanks to Seokjun Lee)

  • bug fix: make sure CHECK-TYPE's expansion does not include internal non-externalizable objects. (#2148777, reported by Willem Broekema)

  • bug fix: alien calls involving passing structs by value are less likely to read or write wrong areas of memory. (thanks to Jesse Bouwman)

  • bug fix: lowering EQUALP to EQUAL handles constant dotted lists correctly.

  • bug fix: a number of standard functions perform more explicit type checks on their arguments.

  • bug fix: only return the primary value from (LIST*/APPEND/NCONC <values>).

  • bug fix: correct treatment of escaped closing brackets in pathname patterns.

  • bug fix: escape dots in pathnames more correctly.

  • bug fix: the hash set for function names will incur collisions less frequently. (reported by Andrew Wolven)

  • bug fix: the system is now capable of expressing select() on the whole range of possible file descriptors.

  • optimization: compiler optimizations for (REPLACE vector list) now apply given :START1 and/or :END1 keyword arguments.

  • optimization: CONCATENATE is faster for concatenating list arguments to a VECTOR.

  • optimization: ROUND for integers is more compact.

  • optimization: on 64-bit x86, implement TRUNCATE using the Lemire, Kaser, Kurz transform.

New in version 2.6.3, 2026-03-29

  • minor incompatible change: (MAKE-ARRAY X :ELEMENT-TYPE 'UNDEFINED) now signals an error, consistent with (UPGRADED-ARRAY-ELEMENT-TYPE 'UNDEFINED).

  • platform support:

    • fix disassembler on ppc for the MFLR and ISEL instructions
    • the Lisp Return Address object (as part of the Lisp calling convention) is no longer needed or supported on PPC, SPARC, MIPS or ARM. (This also frees up a widetag slot previously held by return-pc-widetag)
    • remove sensitivity to SBCL init files when building embedcore-sbcl. (thanks to Robert Brown)
    • add support for the ADCX and ADOX instructions on x86-64. (thanks to Robert Smith)
    • on PPC64, indicate the number of return values through flags, making function calls four times faster.
    • fix FFI involving int128 arguments on x86-64. (thanks to Andreas Franke)
    • fix build on OpenIndiana/x86-64. (thanks to Andreas Wacknitz)
    • fix build on Haiku/x86-64.
  • bug fix: improved stability of (particularly) the mark-region garbage collector. (#2142711)

  • bug fix: compiler type error in complicated expression involving BOOLE and conditionals. (#2142949)

  • bug fix: dotted lists involving symbols whose names begins with "DEF" are not definitions. (#2143114, reported by Glenn Thompson)

  • bug fix: STABLE-SORT miscompiled on declared union types involving both LIST and VECTOR. (#2143163, reported by akater, thanks to Vasily Postnicov)

  • bug fix: more consistent results between compiler and interpreter in numerical computations involving negative zeros. (#2143383)

  • bug fix: USE-PACKAGE once again signals the correct error if an attempt is made to use the KEYWORD package.

  • bug fix: EQUALP on hash tables is no longer sensitive to irrelevant aspects of the hash table.

  • bug fix: SB-INTROSPECT:DEFTYPE-LAMBDA-LIST is more robust against types defined in low debug situations.

  • bug fix: ENSURE-GENERIC-FUNCTION ensures that the allocation of a generic function does not happen in an arena. (Thanks to Andreas Franke)

  • optimization: actually return early when we hit the cache for a :MAKUNBOUND slot access. (thanks to John Mallery)

  • optimization: streams opened with WITH-OPEN-FILE avoid having finalizers.

  • optimization: improvement of COUNT on non-simple bit-vectors, or simple ones with :START/:END arguments. (#2142062, thanks to Andrew Berkley)

  • optimization: SIMD routines for checking strings for ASCII-only content are more compact.

  • optimization: the reader prefers base-string symbol-names of uninterned symbols if possible.

  • optimization: (length (remove-duplicates a s)) doesn't cons an intermediate sequence.

  • optimization: REMOVE-DUPLICATES uses hash-tables in more situations.

  • optimization: UPGRADED-ARRAY-ELEMENT-TYPE is now faster.

New in version 2.6.2, 2026-02-27

  • minor incompatible change: IMAGPART of a negative float returns 0.0, not -0.0, consistent with a treatment of reals as complexes with an imaginary part of strictly 0.0, but strictly incompatible with the requirement that the IMAGPART equal (* 0 <float>)

  • platform support:

    • support for Windows on arm64 has been added. (thanks to Masatoshi SANO)
    • various mismatches and bugs related to mismatches between Win32 and Unix have been addressed.
    • fixed an issue in unsigned 32-bit compare-and-swap on RISC-V (thanks to Andreas Schwab) and LoongArch.
    • fixed the integration of the system with the (lack of) floating point traps on RISC-V. (thanks to Andreas Schwab)
    • implemented the missing runtime breakpoint-related functions on RISC-V. (#2130944)
    • fix for assembling large relative jumps on MIPS.
    • fix for GC safety of function calling on RISC-V and LoongArch. (reported by Will Sinatra)
    • support little-endian PPC64 to write cores in ELF format.
    • fix for SB-POSIX:STAT on Windows with the UCRT C library.
    • numerous other fixes related to architecture definitions, particularly on RISC-V, but also on LoongArch, MIPS, PPC and PPC64, and ARM64.
  • enhancement: support stack allocation of results for struct return values from alien functions. (thanks to Jesse Bouwman)

  • bug fix: rounding of floats converted from ratios. (#2139007)

  • bug fix: SCALE-FLOAT and RATIONALIZE on denormals gave wrong answers, and converting ratios to denormals is both more correct and faster.

  • bug fix: the ~E FORMAT directive scales its float more correctly. (#1854151, #2016431, #2125287, reported by Michał "phoe" Herda, Robert Dodier and Francis Wright)

  • bug fix: the error when failing to bind in DESTRUCTURING-BIND is now a PROGRAM-ERROR.

  • bug fix: the compiler respects INLINE and NOTINLINE declarations to control compiler macros that apply to macros. (#1490345)

  • bug fix: converting bignums to float will trap or return floating point infinities as appropriate to the prevailing floating point modes.

  • bug fix: division with a (COMPLEX FLOAT) result will be more consistent with results involving negative zeros.

  • bug fix: allow the full range of hash values in weak hash-tables with user-defined hash functions. (#2141482, reported by Patrick Poitras)

  • bug fix: malformed OPTIMIZE declarations no longer cause the compiler to stop.

  • bug fix: symbols with terminating macro characters in their names (or non-terminating ones at the start) print with escapes when *PRINT-ESCAPE* is true.

  • bug fix: documentation issues, in README and the manual. (thanks to Carl Gay)

  • bug fix: compiler crash while transforming arithmetic operations on known non-numeric inputs. (#2142297)

  • bug fix: unsafe concurrent access to synchronized weak hash tables. (#2142714)

  • optimization: avoid consing when right-shifting a bignum gives a fixnum result.

  • optimization: various type tests in the presence of intersecting constraints do less redundant work.

New in version 2.6.1, 2026-01-26

  • minor incompatible change: the never-documented :NO-CONSTRUCTOR-DEFUN option to DEFSTRUCT is no longer supported.

  • platform support:

    • support for the LoongArch architecture has been added. (thanks to ZiLong Wang)
    • support for FreeBSD on 32-bit and 64-bit PowerPC platforms has been added. (thanks to Piotr Kubaj)
    • SB-ALIEN now supports passing and returning structures by value (rather than by reference) in accordance with the platform ABIs, on x86-64 and arm64 platforms. (#313202, thanks to Jesse Bouwman)
    • wide compare-and-exchange is supported on the arm64 platform for Armv8.1-A or later.
  • bug fix: signal a DIVISION-BY-ZERO error in calls of (/ 0 0). (#2137266, reported by khbit)

  • bug fix: compiler infinite loop from previously-acceptable recursive inline expansions. (#2137380, reported by Frode Fjeld)

  • bug fix: unnormalized inequality constraints causing a compiler infinite loop. (#2137422, reported by Jesse Bouwman)

  • bug fix: compiler infinite loop from a deleted loop with an remnant TAGBODY. (#2137493)

  • bug fix: miscompilation of REDUCE #'APPEND with non-null INITIAL-VALUE and :FROM-END NIL. (#2137736, reported by Vasily Postnicov)

  • bug fix: encode XREF locations in a way that allows for higher form numbers. (#2137765, reported by Jim White)

  • bug fix: lookups of optimized memory movers fails when *PRINT-BASE* is not 10. (#2138812, reported by Robert Dodier)

  • bug fix: CLRHASH on a hash table with weak keys should not corrupt the table's index vector. (#2138965, reported by Patrick Poitras)

  • bug fix: address a number of type derivation subtleties related to floating point zeros.

  • optimization: various combinations of arithmetic operations with variables and constants are simplified to have fewer calls.

  • optimization: elide full calls to CONJUGATE when its argument is known to be of type REAL. (#2137354, thanks to Vasily Postnicov)

  • optimization: the compiler knows more about the type and results of ARRAY-DIMENSION.

  • optimization: transform ARRAY-DIMENSIONS away on objects with array types of known dimensions. (#2138581, thanks to Vasily Postnicov)

  • optimization: (equal (array-dimensions a) (array-dimensions b)) does not cons intermediate lists.

  • optimization: (make-array (array-dimensions a)) does not cons a list.

New in version 2.6.0, 2025-12-28

  • enhancement: the compiler will recognize certain combinations of high-level optimizations as expressible by shorter machine instruction sequences, documented in the manual under "Recognized idioms" in the "Efficiency" section.
  • enhancement: the SB-COVER code coverage tool can emit a report in a format compatible with the LCOV open-source tool.
  • bug fix: compiled code calling EXPT with constant integer exponent (or 1/2) and floating point base is more consistent with out-of-line EXPT. (#1899969, #2136082)
  • bug fix: fix SCALE-FLOAT on denormal floats. (#2000178, re-reported by Barton Willis)
  • bug fix: the system's test of constructing an ELF core is compatible with storing code coverage information. (#2131956)
  • bug fix: inconsistent result from SUBTYPEP on array types. (#2132250)
  • bug fix: the SB-COVER reporting utilities can now annotate source files containing array literals using #A(<dims> <eltype> . <contents>) syntax. (#2134290)
  • bug fix: compiler error resulting from losing some already-computed derived type information. (#2136852)
  • bug fix: miscompilation of DPB involving non-word-sized intermediate results but a word-sized final result. (#2137028)
  • bug fix: compiler error when asserting the result of a known non-list to be of a type union involving a CONS with a given CAR. (#2137030)
  • bug fix: miscompilation of DPB with constant byte positions above the number of bits in a word. (#2137046)
  • bug fix: miscompilation of PHASE with a negative zero argument. (#2137068, #2137119)
  • bug fix: failure to round-trip types involving positive and negative zeros of different floating point representations. (#2137140)
  • optimization: machine arithmetic can be used when bit-shifting bignum inputs in a modular arithmetic context.
  • optimization: extending an association list, including using backquote notation, is recognized as ACONS and is potentially stack-allocatable.
  • optimization: some intermediate copies of lists are elided for calls to maybe-copying operations surrounded by a call to COPY-LIST.
  • optimization: a number of comparison operations on rationals are simplified where possible.
  • optimization: a number of arithmetic operations recognize and elide double negations or calls to ABS.
  • optimization: tracking code with coverage information uses a weak vector per fasl file, rather than a list of per-function weak pointers.
  • optimization: REDUCE has been tweaked for better performance both on lists and vectors.
  • optimization: for simple-bit-vectors of the right alignment and length, REVERSE will operate a word-at-a-time.

New in version 2.5.11, 2025-11-30

  • incompatible change: the compiler's internal representation of "source paths" for unquoted forms within backquotes has changed. Other developer tools using this representation, including callers of some exported SB-INTROSPECT functions, will misreport the location of signalled conditions and/or definitions in top-level forms including backquotes and commas.
  • minor incompatible change: undefined syntaxes following *READ-BASE*-related reader macros (such as #B, #O, #X, #R) now signal reader errors.
  • minor incompatible change: the convenience reader syntax pkg::(...) no longer triggers package locks for the PKG package.
  • minor incompatible change: building with the SB-DEVEL feature inhibits identical code folding at the end of the build of the SBCL system itself.
  • enhancement: improve the compiler's knowledge of the dimensions of the result of MAKE-ARRAY. (#2130477, thanks to Vasily Postnicov)
  • enhancement: the SB-COVER contributed module has been made substantially more robust; collecting coverage no longer inhibits various CLOS optimizations. (For SBCL developers, it is now capable of reporting on the coverage of the SBCL system itself, provided it is built with the new :SB-COVER-FOR-INTERNALS build-time feature.)
  • bug fix: REQUIREing the SB-MD5 contributed module no longer installs a compiler optimization policy restriction of SPACE being at least 1.
  • bug fix: don't miscompute the sizes of garbage collector data structures for running with dynamic space heap sizes above 128GiB.
  • bug fix: ENOUGH-NAMESTRING when the pathname and defaults arguments are both logical pathnames with the same host returns a shorter string than previously.
  • bug fix: the compiler retains fewer temporary data structures when compiling code with coverage data.
  • bug fix: requiring the SB-MD5 contrib module no longer installs a restriction on the SPACE optimization quality.
  • bug fix: internal compiler floating point error while compiling certain calls to CEILING. (#2132231)
  • bug fix: miscompilation of TYPEP on intersections of types including rational ranges. (#2132207)
  • bug fix: miscompilation of MISMATCH from insufficiently cautious type derivation. (#2132187)
  • bug fix: internal compiler error while compiling ASH from an incorrect consistency check. (#2132156)
  • bug fix: internal compiler error from missing constant-folding stub for internal function. (#2132126)
  • bug fix: miscompilation of GET-PROPERTIES at low SAFETY optimization qualities. (#2131985)
  • bug fix: internal compiler error while generating code for multiplications of fixnums where the result is also asserted to be a fixnum. (#2131894)
  • bug fix: the asserted real range of ATANH was incorrectly stated as [-1,1]. (#2131711)
  • bug fix: incorrect type error resulting from wrong type derivation of REDUCE #'LOGIOR. (#2131699)
  • bug fix: specialized XEPs should not be generated during block-compilation or interpretation. (#2131118)
  • bug fix: fix a wrong compiler transform for MAKE-ARRAY leading to miscompilation. (#2131048)
  • bug fix: miscompilation of type checks of (UNSIGNED-BYTE <X>) for large <X>. (#2130028)
  • optimization: LOGTEST participates in compiler transforms related to modular (machine-sized) arithmetic.
  • optimization: more arithmetic combinations simplifications.
  • optimization: (car (list a)) doesn't allocate a list.

New in version 2.5.10, 2025-10-27

  • platform support:

    • handling of "./" and "../" in pathname functions on Windows is improved. (#2125908, reported by khbit)
    • use x29 for the control frame pointer on arm64, improving backtrace tooling.
    • provide a plugin to lldb to display backtraces (contrib/lldb_bt.py).
    • an experimental option for performing GC without interrupting and stopping foreign function calls. Enabled via --with-nonstop-foreign-call (for arm64, x86-64 (outside of Windows, where that's already the case.))
  • bug fix: some interactions between TWO-WAY-STREAM (and ECHO-STREAM) and user-defined streams have been cleaned up.

  • bug fix: the SB-COVER contributed module can now annotate source files containing COMPLEX literals. (A number of other more minor cosmetic issues have also been fixed).

  • bug fix: compiler crash from reoptimizing with some stale type information. (#2125944)

  • optimization: SLOT-VALUE calls with known slot-name on values which are of type (OR NULL <STRUCT>) are transformed to a null check and a structure access.

  • optimization: the compiler will apply constraints to the result of calling RANDOM. (#2126978, thanks to Vasily Postnicov)

  • optimization: the compiler will perform type derivation on CL:APPLY called with a known function.

  • optimization: fusion of type checking and move of 64-bit integers is enabled on arm64 and x86-64.

  • optimization: allocation fusion for (PUSH (CONS A B) LIST) on x86-64.

  • optimization: improvements of type derivation for float rounding operations.

  • optimization: constant folding when one of the arguments is (if v constant1 constant2)

New in version 2.5.9, 2025-09-29

  • minor incompatible change: remove (SETF SB-EXT:POSIX-GETENV), which only ever existed as an operator in SBCL on Windows.

  • minor incompatible change: (LOG -0.0) now returns SINGLE-FLOAT-NEGATIVE-INFINITY, more consistently with IEEE 754.

  • minor incompatible change: (EXPT 0.0 0.0) now returns 1.0, rather than signalling an error.

  • platform support

    • restore functionality on NetBSD. (thanks to Masatoshi SANO)
    • fix building SBCL as a shared library on ARM64. (#2122059, reported by Guillaume LE VAILLANT)
  • optimization: TYPEP with array types does less work in many cases.

  • optimization: COMPLEMENT on a known function can be transformed away in more cases.

  • optimization: calls to local functions with &REST arguments can be optimized in more cases.

  • optimization: bound checks can be eliminated in ROW-MAJOR-AREF based on constraints relating the index to the available array elements. (#2121253, thanks to Vasily Postnicov)

  • optimization: function type declarations no longer inhibit inlining local functions. (#2121351, reported by kbhit)

  • optimization: bulk movement of memory in the system is implemented with less overhead around memmove().

  • optimization: MAKE-ARRAY with dimensions coming from ARRAY-DIMENSIONS on an array with known dimensions avoids consing an intermediate dimensions list.

  • optimization: a number of arithmetic operators and relations in combination with some constant arguments do partial expression simplification at compile-time. (#2122063 for %NEGATE thanks to Vasily Postnicov)

New in version 2.5.8, 2025-08-29

  • minor incompatible change: SB-THREAD:MAIN-THREAD-P can only be applied to threads, not arbitrary lisp objects.

  • minor incompatible change: the instruction-combining (peephole) optimization pass does not run if COMPILATION-SPEED has a higher value than SPEED.

  • platform support:

    • on arm64, provide better backtraces in the statistical profiler's reporting, along with better detection of assembly routines, local functions and callers of foreign code.
    • on ppc64le, make --dynamic-space-size behave as documented. (#2121255)
    • on x86-64, handle more cases in the ALU+TEST peephole optimization.
  • bug fix: for file-streams with :DIRECTION :IO, input and output file positions should no longer get out of sync. (#1600610, reported by Guillaume le Vaillant, test cases by Brent Benson)

  • bug fix: an infinite loop in SUBTYPEP for types involving negations of CONS of specialized ARRAY types. (#2114755)

  • bug fix: miscompilation of a CASE form with small numeric keys. (#2119035)

  • bug fix: anonymous alien structs definitions are deduplicated, making it harder to overflow internal data structures. (#2114943, reported by Brooke Tilley)

  • bug fix: allow ALU+TEST peephole optimizations to fire on x86-64. (#2120547, reported by Christoph Breitkopf)

  • bug fix: miscompilation of a LOOP form with rational arithmetic on variables involved in termination tests. (#2121178, reported by 3b)

  • bug fix: the compiler is better able to associate some forms in a macroexpanion with the original sources.

  • optimization: improve array construction with LIST or SEQUENCE :INITIAL-CONTENTS.

New in version 2.5.7, 2025-07-26

  • enhancement: the encapsulate mechanism can be used to wrap functions that are currently not FBOUNDP.
  • bug fix: internal compiler error in a failure of stack analysis during propagation of dynamic-extent. (#2113935)
  • bug fix: address regression in type inference for TRUNCATE and other division-related operators. (#2115305, reported by Vasily Postnicov)
  • bug fix: cleanup of the main thread is performed more carefully when SBCL is used as a shared library. (#2115669, reported by Fedorov Alexander)
  • bug fix: the compiler does not lose track of the types of specialized external entry points for user-defined functions. (#2115955, reported by Matt Kaufmann)
  • bug fix: adjust compiler template argument acceptability for increased usage scope. (#2116150)
  • bug fix: provide a stub for ROTATE-RIGHT-WORD for constant-folding during compilation. (#2117080)
  • bug fix: provide a stub for %MAKE-DOUBLE-FLOAT for constant-folding during compilation. (reported by Eric Marsden)
  • bug fix: don't loop infinitely in the presence of type-mismatching circular #S read syntax. (reported by Bohong Huang)
  • optimization: calls to SLOT-VALUE (and related functions) within methods, on values that are not a specialized argument to those methods, are optimized similarly to calls to SLOT-VALUE in non-method code.
  • optimization: calls to REPLACE with VECTOR first argument and LIST second argument are improved.
  • optimization: TYPECASE over a set of structure types known not to be extensible is converted to an array lookup.

New in version 2.5.6, 2025-06-29

  • enhancement: the compiler now recognizes when local functions (both named and anonymous) are used only as downward funargs in many situations and can stack allocate such closures even without explicit dynamic extent declarations. See the updated manual entry on stack allocation for more information and how user-code can declare funargs as downward.

  • minor incompatible change: optimization notes for a variable declared to be of type LIST will not be emitted for various transforms which are defined to operate on (OR NULL VECTOR).

  • minor incompatible change: some forms, including a THE form with an invalid type specifier, or a CASE form with bad entries, no longer produce a runtime error. (They continue to provide a full warning at compile-time).

  • platform support

    • on arm64, breakpoint-based stepping is now thread-safe.
    • on arm64, backtraces after interrupts should be more correct.
    • on x86-64 with the immobile-space feature, calling from assembly routines to lisp routines is more efficient.
    • if arenas are enabled, users can define a lisp function to act as a handler to customize behaviour on arena exhaustion.
  • bug fix: address several bugs related to dynamic-extent declarations, inference, and stack allocation.

  • bug fix: the stack return page protection is temporarily disabled during GC, so that GC can complete even if it needs to write in the return page.

  • bug fix: the compiler generates code to the right entry point for specialized functions. (#2111876, reported by Matt Kaufmann)

  • bug fix: the compiler's constraint derivation would sometimes not terminate. (#2113747)

  • bug fix: when the runtime structure representing a thread is re-used, the stack guard pages are restored.

  • bug fix: type checks for &OPTIONAL arguments are done only once.

  • bug fix: CEILING's docstring was wrong. (reported by Dave Tenny)

  • bug fix: APPLY could be called with too many arguments when parsing MEMBER type specifications. (reported by Zach Beane)

  • bug fix: the compiler could allow constant values of bad types to trigger optimizations. (#2113977)

  • bug fix: internal compiler error when attempting to write out a type check for a value already proved to never exist (i.e. be of type NIL). (#2112475)

  • optimization: better division for signed-word dividends and unsigned-word divisors on arm64 and x86-64.

  • optimization: improvements to subtraction involving bignums and words on x86-64.

  • optimization: perfect-hash-based transformations are applied to sequence functions with keys including fixnums and characters as well as symbols.

New in version 2.5.5, 2025-05-31

  • minor incompatible change: the output from TRACE is now prefixed by a FRESH-LINE on *TRACE-OUTPUT*.

  • platform support:

    • On Linux, the system is better at negotiating with the kernel to find locations for Lisp memory spaces, succeeding more often than previously.
  • bug fix: resolve signed/unsigned char mismatch in RUN-PROGRAM on Windows. (#2110525, reported by awlygj)

  • bug fix: compiler confusion given sufficiently complex derived type constraints. (#2109902)

  • bug fix: compiler inconsistency in low-level representation leading to inconsistent transformations. (#2109837)

  • bug fix: return NIL from calls to DOCUMENTATION on illegal function names.

  • bug fix: calls to APPLY or VALUES-LIST on some combinations of constant arguments could lose the constant nature after transformation. (thanks to Hayley Patton)

  • optimization: some micro-improvements to bignum operations, particularly on x86-64 and arm64

  • optimization: allow the result of MAKE-STRING to be allocated on the stack when :element-type is unknown.

  • optimization: the compiler will recognize the use of ZEROP on the results of LENGTH and REM (on suitable operands) to avoid full computation of the intermediate result.

New in version 2.5.4, 2025-04-28

  • enhancement: :FUN-END breakpoints now support the known values return convention when DEBUG > 0. This means that tracing local functions works in more situations.

  • platform support:

    • on x86-64, relocation of static space is always enabled.
    • save-lisp-and-die with :callable-exports can be used for sbcl.dll on Windows.
    • Building with UCRT64 on Windows is now fully supported.
  • bug fix: :FUN-END breakpoints work on PowerPC, SPARC, and MIPS again.

  • bug fix: incorrect rounding when converting some bignums to floats.

  • bug fix: the second value of the truncation functions is more consistently computed for bignum floats.

  • bug fix: fix code generation for constants being considered from conflicting type propagation information. (#2107652)

  • bug fix: fix 32-bit range check code generation on x86-64. (#2106432)

  • bug fix: types are correctly propagated from the keyword argument processor to their uses. (#2106358, reported by Vasily Postnicov)

  • bug fix: fix compilation error from CHECK-TYPE when the value checked is a keyword argument and the type specifier argument is not a valid type specifier. (#2104089)

  • bug fix: generate stack-manipulation code in the presence of non-local exits and dynamic-extent declarations even more carefully. (#2043242)

  • optimization: (LOGIOR A (- (MASK-FIELD (BYTE 1 constantN) A))), or its equivalent (LOGIOR A (- (LOGAND (ASH 1 constantN) A))), is recognized as an idiom for sign-extending the N+1-bit field in A, and can be used for signed modular arithmetic.

  • optimization: ROUND is faster for floats.

  • optimization: TRUNCATE/FLOOR/etc. are faster on ratios.

  • optimization: MAKE-SEQUENCE does not invoke the full type algebra when the provided type specifier is simple.

  • optimization: don't attempt to align branch targets if the SPACE optimization quality is greater than 1.

  • optimization: circularity detection for printing now places its temporary data structures on the stack.

  • optimization: faster GCD on fixnums, especially when the difference in magnitude is large.

  • optimization: the implementation of ISQRT has been replaced with the (faster) algorithm currently implemented in CPython.

New in version 2.5.3, 2025-03-30

  • enhancement: breakpoint debugger commands have been added. Included is a stepper based on breakpoints requiring no extra instrumentation. However, it still has less functionality than the existing single stepper. See the new debugger manual section titled "Breakpoint Commands" for more information on the new commands.

  • minor incompatible change: the behaviour of :save-runtime-options has been restored to match the documentation. (#2096995, reported by Zach Beane)

  • minor incompatible change: invoking CHANGE-CLASS from user code no longer grabs the CLOS world lock. Callers must take responsibility for ordering execution of CHANGE-CLASS and any changes to the class hierarchy.

  • platform support:

    • (CAS SAP) is implemented on ARM v8.1 directly with CAS instructions.
    • on x86-64, list constructors emit more compact code sequences, particularly in the presence of multiple references to the same object.
    • on x86 and x86-64, fix the stack overflow check to use signed comparisons.
    • on Darwin/arm64 and Linux/x86-64, provide a restart to disable floating-point exceptions of the type signalled, and another to disable all floating-point exceptions.
  • bug fix: cycle detection in class precedence lists happens before adding classes to the direct subclasses of the parent.

  • bug fix: stack-allocated unaligned cons cells no longer cause errors in the debugger.

  • bug fix: local function type declarations no longer inhibit tail calls in (SAFETY 0) code. (#2039301)

  • bug fix: bad or unknown type specifiers in CHECK-TYPE do not crash or slow down the compiler. (#2102644, #2102653, #2102714, #2104048)

  • bug fix: numerous bug fixes relating to the type system's handling of arrays make SUBTYPEP more reliable and less likely to express a contradiction. (#1996980, #2100563, #2100728, #2100779, #2100784, #2100812, #2100825, #2101192, #2101215, #2101803, #2102684)

  • bug fix: improve other aspects of the type system's self-consistency. (#2101073, #2101170, #2101183, #2101189, #2101399, #2101589)

  • bug fix: fix compiler type error when deriving the type of FTRUNCATE. (#2101073)

  • bug fix: fix compiler error when deriving constraints for single-floats. (#2102759)

  • bug fix: startup tuning for particular microarchitectures no longer accidentally disables one of the optimizations.

  • optimization: ROW-MAJOR-AREF is transformed to use the same array machinery as one-dimensional array references. (Thanks to Scott Burson for the suggestion)

  • optimization: list constructors emit shorter code sequences on x86-64, particularly in the presence of multiple references to the same object.

  • optimization: FLOOR and CEILING on ratios do not unnecessarily cons.

  • optimization: provide specialized CALL-NEXT-METHOD functions for the no-argument and full-argument cases.

New in version 2.5.2, 2025-02-28

  • minor incompatible change: in some instances when the compiler cannot prove that a NIL-valued branch is unreachable, where NIL is not compatible with the expected type, a type warning will no longer be issued.

  • minor incompatible change: the compiler will more strictly treat type declarations for &OPTIONAL and &KEY arguments in FTYPE declarations, no longer effectively adding an implicit (OR ... <default>) type when the function itself has a default value not matching the declared type for that argument.

  • enhancement: type errors in structure constructors are now restartable, with a USE-VALUE restart provided.

  • enhancement: CHECK-TYPE warns about type conflicts at compile-time.

  • enhancement: FTYPE declarations for functions which set their parameters are checked.

  • enhancement: new print control variable SB-EXT:*PRINT-CIRCLE-NOT-SHARED*, when used in conjunction with *PRINT-CIRCLE*, prints #1# only for circularities and not simple sharing.

  • platform support

    • on Windows, make sure to commit memory after zeroing during save-lisp-and-die. (#2097197, reported by _3b)
    • on Linux, add the TCP_USER_TIMEOUT constant to SB-BSD-SOCKETS. (thanks to Mihai Bazon)
    • on *BSD, make TCP_KEEPCNT, TCP_KEEPIDLE and TCP_KEEPINTVL available where the OS supports it.
    • on x86-64, optimize BOUNDP for known-global symbols.
    • on x86-64, optimize KEYWORDP for some arguments.
    • on arm64, don't trigger an assertion when using FMOV on complex single-float registers.
    • on arm64, improve type checking for (AND SYMBOL (NOT NULL)).
  • bug fix: using structure read macros with shared structure markers no longer signals type errors when the shared structure is in a slot with a type. (#308936)

  • bug fix: non-conforming user macros which modify their source no longer trigger internal errors. (#1371719, reported by _3b)

  • bug fix: the combination of CONSTANTLY and DYNAMIC-EXTENT declarations no longer causes an internal compiler error. (#2059950, reported by bohonghuang)

  • bug fix: treat inlined functions analogously to constants in the compiler. (#2095560, reported by Vasily Postnicov)

  • bug fix: FTYPE declarations for &optional and &key arguments do not include default values when checking types.

  • bug fix: Storing coverage data no longer leads to miscompilations allowing reachability of unreachable code. (#2092451, reported by mrkissinger)

  • optimization: elide bounds-checking for multidimensional arrays with known dimensions. (reported by aeth)

  • optimization: alien callbacks are generally less heavyweight.

  • optimization: REMOVE shares the tail of the input list when there's nothing to remove.

New in version 2.5.1, 2025-01-31

  • minor incompatible change: SBCL now reveals details of its COMPLEX representations through UPGRADED-COMPLEX-PART-TYPE, rather than hiding them.

  • minor incompatible change: the compiler will warn on the use of a SATISFIES type with an undefined function. (#576608, reported by Roman Marynchak)

  • minor incompatible change: (room t) now counts the space taken by the internals of hash-tables and CLOS instances.

  • platform support

    • fixes to the included version of ASDF, and to sockets functions, for the Haiku operating system. (thanks to Alexandru Popa)
    • add support for CAS (compare-and-swap) on SAPs for arm64, x86-64 and (partially) RISC-V. (#1894057, reported by Yukari Hafner)
    • the system is now consistent with 64-bit time_t on 32-bit linux platforms. (#2063340, reported by Peter van Eynde)
    • restore building on 32-bit ARM with newer gcc versions. (#1839783, reported by Sébastien Villemot)
    • fix large stack allocation on 64-bit Windows.
  • CL portability fixes to the definitions of certain compiler structures, detected by CLISP. (#2064301, #2064312, thanks to Robert Brown)

  • bug fix: a misplaced assertion regarding weak hash tables would trigger if a garbage collection hit at just the wrong time. (#2096998)

  • bug fix: structure BOA constructors with &REST arguments no longer cause structure slots named NIL or T to be unconditionally initialized with the values NIL and T respectively.

  • bug fix: structure BOA constructors without values for some slots no longer cause compilation errors for initforms that are not a single variable.

  • bug fix: sequence functions handle :TEST and :TEST-NOT both being given uniformly. (#309143)

  • bug fix: the type system is better equipped to handle complicated unions of numeric types. (#308937, #1694839, #1734959, #2073544)

  • bug fix: misoptimization of VALUES-LIST in the presence of intervening stack operations. (reported by haruhi.s)

  • bug fix: apply the limit to inline expansions more selectively. (#2092518, reported by Andrew Kravchuk)

  • bug fix: compiler-detected type mismatches are reported even given the presence of inlined functions. (#2092613, reported by Vasily Postnicov)

  • bug fix: improved type error detection for inlined array construction forms. (#2092889, reported by Vasily Postnicov)

  • bug fix: accesses to multidimensional arrays are now checked based on the (internal) INSERT-ARRAY-BOUNDS-CHECKS declaration, as with one-dimensional arrays. (#2095155, thanks to Vasily Postnicov)

  • bug fix: sb-bsd-sockets:socket-connect handles EINTR caused by GC signals.

New in version 2.5.0, 2024-12-29

  • platform support:

    • improve support for the Haiku operating system. (thanks to Al Hoang, Estevan Castilho and Alexandru Popa)
  • bug fix: generic functions with a large number of required arguments, with methods with specializations on exactly STANDARD-OBJECT or FUNCALLABLE-STANDARD-OBJECT, test the types of their arguments more correctly.

  • bug fix: defining a method on SB-MOP:SLOT-VALUE-USING-CLASS where the object argument is specialized to a CONDITION-CLASS no longer leads to an internal error.

  • bug fix: the dissassembler on x86-64 correctly disassembles the vcvttpd2dq AVX2 instruction.

  • bug fix: ensure that the dispatch function for generic functions is compiled with a known compilation policy. (reported by Neil Goldman)

  • bug fix: the compiler retains less intermediate data between COMPILE-FILE forms. (#1557590, reported by andy arvid)

  • bug fix: the (invalid) :INITARGS slot option keyword is reported on. (#1887014, reported by Wayne Rittiman, Jr)

  • bug fix: the SB-SIMD s16.8-maddubs accepts packs of 16 8-bit quantities, not 8 16-bit quantities. (#2069538, reported by Georgios Makris)

  • bug fix: compiling a TYPECASE to dispatch between many user-defined classes no longer takes exponential time. (#2089311, reported by Tomas Hlavaty)

  • bug fix: derive the new type for a variable when setting it to a function of its previous version. (#2090997, reported by Vasily Postnicov)

  • bug fix: properly clear compiler annotations on variables set to new values involving functions of themselves. (#2090967, reported by Kirill A. Korinskiy)

  • bug fix: handle BY in LOOP forms involving iteration on the reverse of a list. (#2091210, reported by James Kalenius)

  • bug fix: fix miscompilation of IF where the consequent and alternative would have the same value but for an intervening side-effect. (#2092588, reported by JA)

  • optimization: SLOT-VALUE and (SETF SLOT-VALUE) on method arguments specialized to structure classes are compiled to the corresponding structure accessor.

  • optimization: calls to SLOT-VALUE (and related operators) on method arguments specialized to instances of SB-MOP:FUNCALLABLE-STANDARD-CLASS are optimized similarly to calls on method arguments specialized to instances of STANDARD-CLASS.

  • optimization: (coerce (reverse list) 'vector) doesn't cons a list.

  • optimization: (replace vector (reverse list)) doesn't cons a list.

New in version 2.4.11, 2024-11-30

  • enhancement: define SB-EXT:*DEFAULT-SOURCE-EXTERNAL-FORMAT* as the external format for reading source files (for direct use in LOAD and COMPILE-FILE). On Windows, this defaults to an external format with CRLF line-endings. (#720517, reported by Mark David)

  • minor incompatible change: the documentation of SB-SEQUENCE:MAKE-SEQUENCE-LIKE has been altered to match its implementation regarding the (un)initialization of the sequence if neither :INITIAL-CONTENTS nor :INITIAL-ELEMENT is provided.

  • minor incompatible change: the outputs from SB-GROVEL no longer contain calls to SB-GROVEL::DEFINE-FOREIGN-ROUTINE, but call SB-ALIEN:DEFINE-ALIEN-ROUTINE directly; the definitions of some other SB-GROVEL utilities has also changed.

  • platform support:

    • The system is more likely to build with the musl C library. (thanks to Masatoshi SANO)
    • It is possible to build 32-bit binaries on NetBSD/x86-64 systems. (thanks to Masatoshi SANO)
    • Stale big-endian ARM code in callbacks is no longer present. (#2087866, reported by Rongcui Dong)
    • Correct the encoding of the VPSHUFD AVX2 instruction. (reported by Dmitry Ignatiev)
    • Implement the PINSRQ SSE instruction and provide access to it in SB-SIMD.
    • Fix some signed/unsigned and 32-bit issues in the runtime leading to problems with large --dynamic-space-size. (#2087986)
  • bug fix: cross-reference information about structure accessors is preserved when compilation policy requires it.

  • bug fix: changing &ALLOW-OTHER-KEYS in a generic function's lambda list needs to invalidate the effective methods cache. (reported by Robert Strandh)

  • bug fix: calling DISASSEMBLE on a method-function provides a more useful disassembly.

  • bug fix: PROCESS-CLOSE no longer leaks a zombie process.

  • bug fix: interaction between SYMBOL-MACROLET and SPECIAL declarations is handled more correctly in the code walker. (#1053198)

  • bug fix: better scaling when compiling large numbers of calls to local functions. (#1379661, reported by 3b and Burton Samograd)

  • bug fix: allow the compiler to approximate types involving large bignums or ratios with large numerator or denominator. (#2085637)

  • bug fix: miscompilation of type tests involving STRUCTURE-OBJECT. (#2088417)

  • optimization: CONCATENATE with consing arguments can elide some of the intermediate consing.

  • optimization: the implementations of various external-formats have been sped up.

  • optimization: elide %SAP-ALIEN calls if all uses dereference the resulting ALIEN object.

  • optimization: faster (expt integer integer) when computing fixnum results.

  • optimization: (ash unknown-integer right) can use modular arithmetic.

  • optimization: (apply x ... list) avoids consing intermediate lists in more situations.

  • optimizations for arm64, x86-64:

    • AREF on non-simple arrays with known element type is faster, along with uses such as LOOP ACROSS, VECTOR-PUSH/POP/EXTEND.
    • SIMD variants for POSITION for strings, 8 and 32 bit integer arrays.
    • faster overflow checking for (the fixnum (+ fixnum fixnum))

New in version 2.4.10, 2024-10-30

  • minor incompatible change: SB-POSIX::POSIX-FORK is no longer exported from SB-POSIX. (The interface function, SB-POSIX:FORK, remains exported).

  • platform support:

    • fix bugs in instruction encoding on RISC-V; (reported by Guillaume Le Vaillant)
    • fix the location of the linkage-table comment in disassembly on 64-bit powerpc;
    • elide allocation of empty number stack frames on arm64;
    • fix crash on x86 platforms in compiling array dereferencing with computed offsets with negative intermediate results. (#2084943)
  • enhancement: the error message from standard object slot typecheck functions in optimized constructors is clearer about the context of the failed type check.

  • enhancement: BREAK is no longer tail-called, even when in tail position.

  • enhancement: on arm64 and x86-64, specialized entry points for functions known to take or return fixed numbers of double floats are generated and can be automatically called without boxing intermediate floats.

  • bug fix: RUN-PROGRAM no longer leaks memory by referencing otherwise unreachable stream instances.

  • bug fix: exporting or unexporting symbols during package iteration no longer causes any symbol to be visited more times than expected.

  • bug fix: DISASSEMBLE preserves the comment marker across line-breaks for long function or segment names. (#1889456, thanks to Fedorov Alexander)

  • bug fix: the compiler no longer loops infinitely trying to compile NOTINLINE calls to known functions with source transform definitions. (#2085451, reported by Fedorov Alexander)

New in version 2.4.9, 2024-09-29

  • minor incompatible change: FIND, POSITION (and variants) now check :START and :END arguments for validity as bounding index designators for list sequences.

  • platform support:

    • improve support for Solaris and variants on x86 and x86-64. (thanks to Masatoshi SANO)
    • fix a bug in handling timeouts and interrupted system calls in SB-UNIX:UNIX-SIMPLE-POLL. (#2078824, thanks to Michał phoe Herda)
    • fix a bug in the lisp understanding of ssize_t under Windows.
    • fix large constant encoding in RISC-V. (#2077307, reported by Guillaume LE VAILLANT)
    • more parsimonious low-level type tests on arm64.
    • building from the result of git-archive should complete without error.
  • bug fix: exporting a symbol during package iteration no longer skips other symbols. (#2080387, reported by kbhit)

  • optimization: improvements to EQ hash tables and associated hash functions.

  • optimization: type checking of string and string-designator is more efficient.

  • optimization: the compiler better understands the nature of the results of CONCATENATE.

New in version 2.4.8, 2024-08-29

  • bug fix: the elftool utility finds a writeable directory in more situations. (thanks to Shinmera)
  • bug fix: SLOT-MAKUNBOUND does not attempt to dereference a PROGN variable in the interpreter.
  • bug fix: READ-SEQUENCE into displaced arrays with a non-zero offset now writes to the right memory location.
  • bug fix: fix some erroneous file position calculations in the editcore utility (exposed by a change in the libzstd compression implementation).
  • bug fix: do not break the build on STYLE-WARNINGs for earlier SBCL build hosts. (#2064671, reported by Patrick Poitras)
  • bug fix: various bug fixes for ppc64le (#2074275, reported by Claude R. C.)
  • bug fix: address a rounding error in the TAN type deriver which led to a miscompile in cl-pdf. (#2077100, reported by Gian Pierro Carrubba)
  • bug fix: overoptimization of FIND with a :TEST of CHAR-EQUAL. (#2077539)
  • optimization: detection of duplicate names in loaded code now scales subquadratically.
  • optimization: switch from Floyd's to Brent's cycle detection for lists.
  • optimization: EQUAL on lists should be faster.
  • optimization: fewer filesystem operations are performed when working out what to LOAD.
  • optimization: various microoptimizations on hash tables and associated operations.
  • optimization: strings are now hashed using FNV-1A, replacing Bob Jenkins' one-at-a-time hash.
  • optimization: fewer redundant validations of the sequence bounding indices in POSITION on strings.
  • optimization: many improvements to type derivation on the arguments to and results of standard functions
  • optimization: adding many (hundreds) methods to a generic function can be done faster.

New in version 2.4.7, 2024-07-27

  • minor incompatible change: the compiler will emit optimization notes related to complex type checks only at high SPEED optimization settings.

  • minor incompatible change: the GET-FOREGROUND symbol is now exported from the SB-THREAD package. (thanks to Philipp Marek)

  • minor incompatible change: code objects are now printed with their ending address as well as their start address.

  • platform support:

    • fix occasional saved core corruption on Win32. (reported by Luís Borges de Oliveira)
    • address some crashing cases in arena support. (reported by Andreas Franke)
    • fix a hard-coded path to cp. (thanks to Hraban Luyat)
    • address relocation issues on PIC-flavoured arm64. (#2067153, thanks to leafze)
    • fix a crash in the arm64-dissassembler if a code component is missing.
    • fix a crash on 32-bit musl libc systems. (reported by Will Sinatra)
    • fix building with link-time optimization. (#2072800, reported by Eli Schwartz)
    • fix building with newer llvm. (#2071545, reported by Yan)
    • mitigate the lack of gc-safety in SB-VM::REMOVE-STATIC-LINKS. (#2045433)
  • bug fix: COMPILE installs its second argument as the value of its non-null NAME argument, even if the second argument was already compiled. (#2071324, reported by Tim Bradshaw)

  • bug fix: allow the hashing routines in sb-md5 to work on arrays with more than 2^29 elements.

  • bug fix: allow READ-SEQUENCE and WRITE-SEQUENCE to read and write files bigger than 4GB.

  • bug fix: READ-SEQUENCE returns the index of the first unmodified element of its sequence. (reported by Janis Dzerins)

  • optimization: various improvements to floating point rounding routines.

  • optimization: PHASE called on the result of COMPLEX now elides the intermediate COMPLEX object.

  • optimization: the compiler is more aware of the types of COMPLEX, ATAN, APPEND and NCONC.

New in version 2.4.6, 2024-06-29

  • enhancement: name conflicts resulting from colliding symbols in IMPORT and USE-PACKAGE are resolved once for each name, rather than between pairwise colliding symbols.

  • enhancement: calls to structure constructors with type mismatches in default initforms cause compile-time warnings.

  • platform support:

    • fix constant-folding of %log1p and %log2 on 32-bit x86.
    • fix the encoding of popcntd on ppc64
  • bug fix: EXPORT could be tricked into exporting two distinct symbols of the same name from the same package.

  • bug fix: two-argument calls to LOG with arguments of different precision do not lose accuracy through insufficiently-precise intermediate values.

  • bug fix: :NEWLINE options in *DEFAULT-EXTERNAL-FORMAT* are respected when opening files. (reported by Marco Antoniotti)

  • bug fix: extend type declarations for the iteration variable of DOLIST with NULL during the evaluation of the result clause. (#942237)

  • bug fix: #\uE0 (LATIN CAPITAL LETTER A WITH GRAVE) was incorrectly not downcased with STRING-DOWNCASE. (#2067841, reported by Matt Kaufmann)

  • bug fix: backquoted lists as arguments to MAKE-ARRAY were miscompiled. (#2069345, reported by Dan Bothell)

  • bug fix: resolve the circularity between the type system and the CLOS metaobj

The Daily Front Page 10 of 25
Tuesday, July 28, 2026 The Daily Front No. #260728 — Crypto Research with Claude
article

Discovering Cryptographic Weaknesses with Claude

by gslin·▲ 210 points·146 comments·anthropic.com ↗
“Substantial research advances, but not affecting production systems.”

Discovering cryptographic weaknesses with Claude

Summary

Using Claude Mythos Preview, researchers at Anthropic have discovered improved ways to attack cryptographic algorithms (the mathematical methods used to keep online data private). The first attack significantly weakens HAWK, a digital signature scheme that was built for a post-quantum world. The second identifies a new way to attack round-reduced AES, the most widely used symmetric cipher. These are substantial research advances, but they do not currently affect any production systems. This post describes both findings in more detail and discusses the implications for cryptography in an age of powerful AI models.

Introduction

When we launched Claude Mythos Preview, we showed it was able to autonomously find and exploit vulnerabilities in almost every piece of software we pointed it at. This included several major cryptographic libraries—shared collections of code that are used to encrypt data.

The vulnerabilities that Claude found in these cryptographic libraries1 were due to incorrect implementation of the algorithms—that is, errors in how programmers used the algorithms in their code that created opportunities for attackers to break the encryption.

Now, we have found that Claude is able to find mathematical flaws in the algorithms themselves.

Cryptographic algorithms are a fundamental building block of digital security. For example, when you visit a webpage like https://www.anthropic.com, your browser checks that it is communicating with an authentic website using an algorithm called a digital signature scheme. Later, the traffic between you and the website is encrypted using symmetric ciphers—codes that allow secure data transmission between parties who share an identical key. Without secure cryptographic systems like these, your email, online banking, and other internet use would be open to cybercriminals, who could intercept or modify your communications. Flaws in these widely used cryptographic systems could put billions of users’ data at risk.

The first result we describe in this post, which was discovered with Claude Mythos Preview, is an improved attack against a digital signature scheme called HAWK. In 2022, the US Government’s National Institute of Standards and Technology (NIST) put out a call for additional cryptographic systems that would remain secure even against quantum computers (which could, if developed, break most of the existing signature schemes in use today). HAWK is one of the third-round candidates under consideration from this call. Despite HAWK having survived two rounds of expert human review over a period of two years, Mythos was able to improve the best-known attack on it in just 60 hours of work—effectively cutting its key strength in half.

The second result concerns the Advanced Encryption Standard (AES), a symmetric cipher that was adopted by NIST in 2001 and has received more scrutiny than almost any other encryption algorithm. In order to better understand the robustness of AES, weaker variations of the algorithm are regularly studied in cryptography research; Mythos found a way to break one such weaker version, and eliminated one of the guesses an attacker needs to make, improving the speed of the previous best attacks by 200-800×.

To be clear, neither of these results has a practical impact on today’s computer systems; no production software will have to change as a result. HAWK is only a candidate signature scheme and so is not deployed;2 our second attack is on a reduced version of AES and does not break the full cipher.3

Nevertheless, both results show the potential for frontier AI models to help discover flaws in important cryptographic algorithms, both before and after real-world deployment. This is cryptography research working as intended: stress-testing algorithms to build trust and ultimately make systems more secure.

Mythos Preview achieved these results mostly autonomously and mostly without human intervention. Over the course of a week, one Anthropic researcher worked together with Claude to develop the HAWK attack, and another researcher built a scaffold4 that allowed Claude to fully autonomously discover the AES attack.5 Each of the results cost roughly $100,000 in API cost to develop. After seeing these results, we broadened our search and began to discover other attacks. We discuss some of these follow-ups below.

In order to make it easier for others to continue studying the cryptanalytic ability of LLMs, we partnered with academics at ETH Zurich, Tel Aviv University, and University of Haifa to build CryptanalysisBench, a benchmark that packages together many cryptographic ciphers and makes it easy for others to evaluate the capabilities of LLMs on this important topic.

Throughout the research process, we followed responsible disclosure procedures, and consulted with academics to confirm the validity of our findings. We also shared advance copies with US government and industry partners, and held discussions on the implications of this research. In the case of our HAWK finding, we shared our attack with the authors of HAWK in June and coordinated disclosure to the public NIST mailing list at the same time our results were released.

In the rest of this post, we summarize the two findings in further technical detail and briefly describe some of our other recent cryptography results. Full descriptions of the two main findings are provided in two new papers, and we hope to release details for our other findings in the near future.

An improved key recovery attack on HAWK

Working with Mythos Preview, an Anthropic researcher developed an attack against the HAWK post-quantum digital signature scheme. This attack substantially speeds up the time it would take to break the signature scheme—more technically, it reduces the “effective keysize” by a factor of two. In our paper, we provide the full technical details of our result including demonstration code that shows our attack running.

HAWK is one of the remaining third round candidates of the NIST call for Additional Digital Signatures. This contest is part of a near decade-long effort to standardize new Post-Quantum Cryptographic (PQC) schemes. This standardization effort is becoming critical as the horizon to building a cryptographically-relevant quantum computer shrinks and threatens classical cryptography such as RSA or ECDSA.

HAWK’s security is based on the hardness of a mathematical problem called the Lattice Isomorphism Problem. Mythos’s attack works by finding a specific, previously unexploited symmetry called a nontrivial automorphism in the lattice used by HAWK. Prior work proved that efficiently finding such an automorphism would permit an attack, but did not answer if such an automorphism was accessible in the lattice used by HAWK. The automorphism discovered by Mythos allows a faster enumeration attack that, while still exponential, means that one needs to double the size of HAWK keys to achieve the same level of security. Unfortunately, doubling HAWK’s key size eliminates many of the reasons making the scheme (as it currently stands) an attractive PQC signature candidate.

Discovery process

To find the attack, Claude Mythos Preview worked semi-autonomously in an agentic harness, with occasional human guidance and nontechnical direction. Mythos found the attack after an extensive literature review to understand the state of the art, and substantial mathematical reasoning and computational experiments. After finding the attack, Mythos implemented an end-to-end verification pipeline to convince itself—and the human operator—of the attack’s correctness.

For this experiment, we used a Claude Code-like harness that supports multiple worker agents collaborating together in a sandboxed environment, with access to computational tools like Python and Sage as well as access to published cryptographic works. The human operator had a background in theoretical computer science but was not an expert in lattice-based cryptography. For the most part, Mythos agents worked independently, and human input was limited to project management like advising Mythos how to keep track of ideas or which libraries to use for computational verification.

The multi-agent workflow led to interesting dynamics. For example, the key idea in producing this attack was discovered by a pair of workers working together. Both started investigating the idea; the first worker prematurely rejected the idea as infeasible, but the second found a way to fully exploit it. The pair kept exchanging messages, and eventually both agreed they had found an effective attack.

Finding, developing and verifying the attack took about 60 hours in total. We estimate that the full attack discovery process cost approximately $100,000 in API cost.

Impact

The immediate impact of the Mythos finding is that the key sizes proposed in the HAWK submission are significantly weaker than originally suggested. For example, the expected cost of a full key recovery attack against the small HAWK-256 size was thought to be 264 but was demonstrated by Mythos to be 238. For larger keys, HAWK therefore remains impractical to attack. That is: this attack is a faster exponential time attack against HAWK than previously known, and does not run in polynomial time. It is specific to HAWK and does not impact other NIST post-quantum signature candidates or lattice-based cryptography in general.

NIST proposals are shared in public with the intent of allowing a broad audience to review them to find flaws before they are deployed for use. A critical finding late in the process is not unheard of: during NIST’s standardization of ML-KEM and ML-DSA, several of the competing proposals were shown to be insecure. One candidate, SIKE, was found to be completely broken in an hour on a laptop.

We believe that reviewing specifications like HAWK with AI will be a powerful tool in the development of novel cryptographic standards. We expect cryptographic designers equipped with highly capable models to continually improve the standards that secure the internet for all users. Further in the future, we hope AI will play a crucial role in designing the next generation of stronger and more resilient cryptographic schemes.

An improved attack on reduced-round AES

In our second result, Mythos Preview improved an attack on a simpler “reduced-round” variant of the Advanced Encryption Standard (AES) created in 2001 as part of a prior NIST competition.

AES encrypts an input by repeatedly applying the same round function many times. AES-128, the specific cipher we attack, has 10 rounds. Our attack works only on a modified version of the cipher that has 7 out of the full 10 rounds. Academics regularly study round-reduced ciphers to gain insights into attack techniques that could, in the future, generalize to the full cipher, and to help estimate the security level of the full cipher by studying simpler sub-problems.

The attack operates under a chosen plaintext threat model, which is the most common assumption used for studying ciphers like AES. Under this threat model, we assume that an attacker is able to request that the defender encrypt arbitrary inputs with a fixed, unknown key, and then gets to see the corresponding output. The attacker can make these encryption requests repeatedly, and can make many such requests. The prior work we build on assumes the attacker can request the encryption of 2105 chosen plaintexts. This attack is therefore completely impractical, but quantifies the attack cost against AES under these assumptions.

Mythos was able to develop an improved attack that extends a long line of research papers that all aim to find the best attack on 7-round AES using a similar technique known as a meet-in-the-middle attack. At a very high level, these attacks work by trading off time for space. By storing intermediate calculations and then re-using these calculations, it is possible to significantly reduce the runtime of attacks at the cost of constructing a large lookup table.

Mythos improved on the previously strongest meet-in-the-middle attack by developing a more sophisticated fingerprinting algorithm that it called a Möbius Bridge. The objective of the fingerprinting algorithm is to increase the number of potential lookups into the table that will succeed. One of the stages of the attack from prior work had to enumerate 256 different values and then look them up in the pre-computed table. Mythos developed a fingerprint that is invariant to this guess, which directly reduces the amount of work required by a factor of 256. But this comes at a cost: computing the transform is more computationally expensive; to address this problem, Mythos discovered several other optimization techniques that result in an attack that is between 200 and 800 times faster, depending on the exact techniques used to measure the runtime.

Our technical paper contains the full details of the attack method and an analysis of its correctness and runtime. Compared to the one week that Mythos spent conceiving the idea, the vast majority of human researchers’ time was spent validating the correctness of its claims (though it is important to note the researchers are not experts in cryptography).

Discovery

Mythos Preview discovered this result almost entirely autonomously. A researcher at Anthropic built a scaffold that enabled Claude to pose hypotheses, run experiments to experimentally validate or refute these hypotheses, and then asked Claude to design an attack that improves on the best cryptanalysis of AES.

Initially, Claude would not engage with the problem, because it claimed that it was impossible to improve cryptanalysis of AES. The result of our first runs ended with Claude writing messages like:

If you want a different outcome, the target has to change … AES-128 r5/r6 is just genuinely hard

Or:

on AES-128 r5/r6/r7 it found nothing because there's nothing easy to find; this is the most-studied block cipher in existence.

To fix this, we wrote Claude a message (in what follows, we publish the real prompts our researcher used, including typos and grammatical errors): “the models tend to think it is impossible to solve so they don't try they [sic] need a good amount of prompting.” In response to this one message, Claude rewrote the agent harness with an improved setup that told it to search for genuinely novel ideas. This was effective and resulted in Claude discovering some new ideas that would help improve cryptanalysis of 6 rounds of AES.

We then asked Claude “why not do aes-128 r7? the whole point is to find something better than existing approaches.” Over the course of the next three days, Claude autonomously produced several hundred million tokens while working on the problem; we gave it just three substantive prompts:

  1. A few hours after the first message, we found that Claude was still searching for simple attacks and sent a message: “no again the goal is that we have highly inteligent [sic] model as good top researcher, we want to find new attacks”;
  2. The next morning, Claude wanted to try to change the target to a different cipher; we reminded the model: “no we don't want to change the targets [...] agian [sic] we need to find something that worth [sic] publishing”;
  3. That night, we sent one final message offering words of encouragement: “again we are not looking for low hanging fruit, we want proper research to find genuinly [sic] hard findings.”

Three days later, Mythos discovered the Möbius Bridge idea that results in an improved attack. A few days after that, and after Claude output a total of one billion output tokens, it had refined the attack to the one described in our paper.

Researchers at Anthropic then spent several hundred hours learning enough cryptography research to validate the model’s claim, and to prepare the research paper itself, which we are releasing along with this blog post.

Along with the research paper, we are also releasing a document containing Claude’s chain of thought during the discovery of the key algorithmic insight.6 In this session, Claude begins by reviewing what previous agents had discovered, reading the various critiques, and then turns to proposing various new transforms; after proposing and rejecting several ideas, it comes up with the key idea of the Möbius transform. Claude then validates this idea both mathematically and computationally, and then writes a report that future agents then used to develop the remaining ideas that formed its paper.

Further work

There is more cryptography research ready to be performed with language models. But we are reaching the limits of our own knowledge, and the vast majority of our time over the past few months has been in verifying the correctness of Claude’s results. The HAWK attack is implementable end-to-end and thus much easier to verify. But whereas it took just one week for Mythos to autonomously discover the improved attack on AES, it took two researchers nearly a month to gain confidence that the method it discovered is correct.

Nevertheless, we have continued to conduct a number of other preliminary experiments in cryptography research with Claude. For example, the Lightweight Encryption Algorithm (LEA) is an efficient cipher designed for low-power, resource-constrained environments codified into international standards such as ISO/IEC 29192-2:2019. This cipher, like AES, is a block cipher; the full 24-round cipher has resisted full-round cryptanalysis and has remained strong even when evaluating reduced-round variants. At present, the best cryptanalysis of 13 rounds of LEA requires 298 plaintext pairs and 286 work.

Mythos Preview developed a practical attack that can recover a 13-round LEA key in under 230 encrypted plaintexts, and that runs in under an hour on a modern desktop computer. Again, this attack does not apply to the 24-round cipher, and so has no immediate practical consideration. Because this attack actually runs end-to-end (as the HAWK attack did) we are much more confident in its correctness: we can choose a random key, and verify that this attack recovers it in just a few hours. Mythos discovered this attack much more recently and we still have more work to do to understand the full results (for example, the exact bounds on the number of plaintext pairs required, how some keys are harder to recover, and how it extends to 14 rounds). After more investigation, we plan to make the full results public.

Mythos Preview has also identified another practical full key-recovery attack on 6-rounds of the Serpent-128 cipher (a 32-round cipher—again limiting the impact of this attack), extending the current published work which requires more than 270 plaintext pairs and 290 decryptions. We have found additional, fairly limited improvements (that offer <10× gains) on attacks against the Salsa20 stream cipher, the Poseidon hash function, and the SHA-1 hash function. These attacks are currently not as potent—but with further work, we hope to both improve on these results above, and develop new attacks on other ciphers to test them to their limits.

Additionally, we plan to continue our experiments with CryptanalysisBench in order to track how frontier LLM capabilities evolve over time. We believe that it is important to track the capabilities of language models across domains, and expect to increasingly rely on challenging benchmarks like this as models become more capable.

Conclusions

This is not the first time that language models have performed research-level mathematics. In just the last few months, researchers from Google have used Gemini to resolve several open Erdős problems, researchers from OpenAI have used GPT to resolve the unit distance conjecture (a particularly challenging Erdős problem), and earlier this month we announced that Claude Fable 5 resolved the Jacobian Conjecture. Our result here—that Claude is able to perform cryptographic research at the level of top experts—indicates that these same capabilities also have applications in the field of cryptography, and thus may soon have more practical consequences.

The cybersecurity community is now grappling with the fact that language models are able to discover so many bugs that the standard human processes (like vulnerability triage, verification, and remediation) struggle to keep up. We predict that the same will soon be true in academic cryptography research. As language models increasingly produce novel research outputs autonomously, human researchers may become bottlenecked on studying and validating these results for technical validity, novelty, and utility. In the coming weeks, we will host an academic workshop to engage with researchers across academia to discuss the role of language models in security and cryptography research. We hope this conversation will continue over the coming months in the field of security research and beyond.

Both of our primary attacks are expected results. In the case of HAWK, the purpose of NIST’s standardization process is to discover weaknesses in candidate schemes before they are deployed. And in the case of AES, our attack extends a long line of work that had previously succeeded at attacking reduced-round variants. But we should not assume that language model capabilities will plateau at this level. In just one year, language models have gone from being unable to perform cryptanalysis of even the most basic ciphers to being capable of finding flaws in cryptographic designs that have escaped discovery despite years of human expert review. Many ciphers protecting modern systems have received less scrutiny than they deserve—they might still have important weaknesses lying dormant that LLMs will soon be able to discover. We see this as a real opportunity to expand our ability to study the long tail of ciphers used throughout the world, and also our ability to more deeply study the ciphers that matter most. Indeed, as we mentioned above we have already begun audits of other schemes.

The attacks described in these two papers are the strongest attacks we have found to date. We are sharing them after a period of consultation with US government and industry leaders. But as we develop increasingly powerful cryptanalytic results, it would be prudent to consider how researchers should react if a language model were to discover vulnerabilities in cryptosystems where attacks do have an immediate real-world impact. We believe answering this question will require input from academia, government, and industry. We hope that our work here will help launch these conversations.

The cryptography community has always benefited from adversarial review: ciphers are proposed, examined, and revised until the community is satisfied with their security. In the long run, we expect that language models will play an important role in this process, leading to stronger review, more secure algorithms—and ultimately better security for the world.

Links to full research papers

Read the full paper on HAWK.

Read the full paper on AES, and the associated chain of thought.

Read the paper introducing CryptanalysisBench.

Footnotes

  1. See, for example, cryptographic vulnerabilities we have found on OpenSSL and wolfSSL.
  2. We believe the attack discovered by Mythos Preview does not impact the other NIST post-quantum cryptographic schemes or other schemes that use related methods.
  3. Even then, the attack would cost hundreds of millions of dollars to implement and does not impact other similar cipher schemes.
  4. A scaffold is a set of prompts and code that help the model achieve its goal. We build on top of Claude Code, construct an environment where it can safely run various experiments, and log its results.
  5. Out of curiosity, after confirming the HAWK result was correct, we then tested if the same scaffold that successfully attacked AES could also re-discover the HAWK break. It could.
  6. Importantly, this is just one of many (autonomous) sessions where Claude worked on discovering new ideas. Many sessions resulted in no new discoveries; other follow-up sessions improved on the insight developed in this one. This document was produced by having Claude rewrite the chain of thought to include more detail to make it easier to read.
The Daily Front Page 11 of 25
Tuesday, July 28, 2026 The Daily Front No. #260728 — Email, Still Hard: DMARC Adoption
article

DMARC has been public since 2012 but most company domains still don't enforce it

by adulion·▲ 187 points·115 comments·ciphercue.com ↗
“68.4% of company domains still don’t enforce it.”

DMARC has existed since 2012. It is a free DNS record that tells receiving mail servers what to do with email that fails to authenticate as coming from your domain: report it, quarantine it, or reject it outright. It is primarily concerned with unauthorised use of a domain in the visible From address; it doesn't stop lookalike-domain registrations, display-name spoofing, or a phishing email sent from a compromised legitimate account. Fourteen years on, we checked the DNS records for 67,336 domains in CipherCue's tracked entity set between 2026-04-14 and 2026-07-28. This is a snapshot of CipherCue's dataset, not a statistically representative sample of every company worldwide; the method note below covers how the cohort is built. 30,362 of them (45.1%) still don't have a record.

The domains that do have a record are not much further along. Only 10,963 (29.7% of domains with a record) actually enforce anything: p=reject, mail that fails authentication gets dropped. 10,258 (27.7%) sit at p=quarantine, junk folder but not blocked. The largest single group, 15,709 domains (42.5%), is set to p=none. A p=none policy collects authentication data and aggregate reports but does not ask receiving mail systems to quarantine or reject messages that fail DMARC. In this analysis, enforcement means a published policy of p=quarantine or p=reject; domains using p=none are counted as non-enforcing.

42.5% of domains with a DMARC record are still at p=none, collecting reports but not requesting quarantine or rejection, as of our most recent observation window

p=none is meant to be a temporary monitoring phase before you move to enforcement, typically a few weeks. Fourteen years after the standard shipped, for the largest single group of domains in our data, it looks like the permanent state.

Put the two together and the picture is starker than either number alone: of all 67,336 domains we checked, 46,071 (68.4%) either have no DMARC record or have one that doesn't enforce a policy. Publishing a DMARC record is no longer the main adoption gap in this dataset; moving from monitoring to enforcement is.

68.4% of all domains checked have no DMARC record, or have one that does not enforce a policy (no record: 45.1%; record present but p=none: 23.3% of the total)

The policy breakdown, in full

30,362

No record

45.1%

15,709

p=none

23.3%

10,258

p=quarantine

15.2%

10,963

p=reject

16.3%

Share of all 67,336 domains checked, not just those with a record

StateDomainsShare of all checkedShare of domains with a record No DMARC record30,36245.1%n/a p=none (monitor only)15,70923.3%42.5% p=quarantine10,25815.2%27.7% p=reject (full enforcement)10,96316.3%29.7%

Why p=none doesn't move

The usual explanation is inertia or ignorance. Our data points at a more specific, more mechanical cause: nobody can tell what's actually sending the mail.

Every domain with DMARC's reporting flag on gets daily aggregate reports (rua=) listing every source that sent mail claiming to be from that domain, whether it passed authentication or not. Moving to enforcement means going through that list and deciding, source by source, "yes, that's meant to be us" or "no, block it." We pulled the raw rua= addresses out of 36,974 DMARC records and counted where the reports actually go.

10,268 distinct rua= reporting domains found across 26,179 reporting-address entries; 8,113 of them (79%) appear in our data exactly once

Some of that concentrates in a few obvious places: Proofpoint's own reporting endpoint appears 1,605 times, Cloudflare's 1,273 times, dmarcian's various regional endpoints (ag.eu.dmarcian.com, ag.us.dmarcian.com, and eight more country-coded variants) add up to over 1,000 between them, and Brevo, Postmark, Barracuda, and a scatter of DMARC-as-a-service vendors (EasyDMARC, PowerDMARC, dmarcian, Red Sift, DMARC Analyzer, dmarcly, sdmarc.net, hornetdmarc.com) fill out most of the rest.

But the long tail is the finding. Once you get past roughly the top 60 domains, the addresses stop being recognisable vendors and start being one-off, often hashed, mailbox names: ivrejeuw@ag.c1.dmarcian.com, a.8hyzr404@sdmarc.net, 2fa9a7572f@rua.easydmarc.eu, a company self-hosting reports at its own domain (dmarc@axa.com, rua@lseg.com), or a mailbox that gives no clue at all (watchdog@watchdog.kevlarr.io). An administrator staring at a week of these reports is not looking at a vendor list. They are looking at a pile of IP addresses and sender strings with no obvious owner, and deciding whether to enforce means identifying every one of them first.

That is a research task, not a configuration change, and it is the kind of task that gets deprioritised indefinitely. We think it is the largest single reason the p=none number above is as high as it is.

Who runs DMARC monitoring

We mapped the rua= reporting-address domains against a vendor dictionary covering the DMARC-monitoring and mail-infrastructure products we could confidently identify. This is a different, broader measurement than matching against a single fact type: it picks up 8,862 identified entities across 29 named destinations, versus 1,400 when only the pre-built vendor dictionary is used. It is still a floor, not a ceiling. It only counts domains we could confidently attribute; the 8,113 one-off destinations described above are excluded here because a single occurrence isn't enough to confidently name a vendor.

DestinationEntitiesShare of identifiedWhat it is *Brevo1,29714.6%*Transactional email platform (formerly Sendinblue); receives reports as a side effect of sending mail, not a DMARC product Proofpoint*1,094*12.3%*Secure email gateway; own reporting endpoint (emaildefense.proofpoint.com) Valimail*1,068*12.1%*DMARC monitoring / automation product Cloudflare*978***11.0%**DNS/CDN provider; reports arrive at their own domain via a hosted DMARC feature, not a dedicated DMARC product DMARC Analyzer6026.8%DMARC monitoring product (Dutch-founded, part of Vade since 2020) dmarcian5846.6%DMARC monitoring product, nine country-coded regional endpoints DMARC Advisor4074.6%DMARC monitoring product EasyDMARC3774.3%DMARC monitoring product Postmark3483.9%Transactional email platform; reports arrive as a side effect of sending mail MxToolbox3213.6%DNS diagnostics vendor offering a DMARC report reader Barracuda2673.0%Email security / gateway vendor Red Sift (OnDMARC)2482.8%DMARC monitoring product PowerDMARC1842.1%DMARC monitoring product Fortra (Agari)1491.7%Email security / DMARC monitoring product Others (15 vendors)93810.6%dmarcly, GoDaddy, sDMARC, Red Sift (separate endpoint), CheckPoint, Mailgun, Mailhardener, Everest, HornetSecurity, Kevlarr, LetsDMARC, report-uri, GlockApps, Cisco, UK NCSC

Reading this table needs a caveat the table itself can't carry: Brevo, Cloudflare, and Postmark are not "DMARC monitoring vendors" in the way Valimail, dmarcian, or Red Sift are. They receive aggregate reports because a domain owner pointed rua= at a mailbox on their platform, often because that platform is also the domain's outbound sending or DNS provider, not because the domain owner bought a dedicated monitoring product from them. Restricting the table to vendors whose core product is DMARC monitoring, Valimail, DMARC Analyzer, dmarcian, DMARC Advisor, EasyDMARC, Red Sift, PowerDMARC, and Fortra (Agari) together account for 3,619 of the 8,862 identified entities, 40.8%. Within that narrower base, Valimail is the largest single vendor at 29.5%, with DMARC Analyzer (16.6%) and dmarcian (16.1%) some way behind; no vendor holds a majority, but the market is not evenly split either.

Country breakdown

Enforcement stage varies by country. Poland has the largest share of domains with no DMARC record at all in this cohort; the UK has the smallest share of domains with no record, but a correspondingly higher enforcement rate once a record exists.

CountryDomains checkedNo recordp=nonep=quarantinep=reject Poland7,03964.6%16.3%11.3%7.7% Netherlands5,59851.1%21.1%14.2%13.6% Germany12,15245.7%26.3%12.9%15.0% United States13,29242.1%19.0%16.7%22.2% Italy4,54540.9%36.8%11.7%10.5% United Kingdom1,49337.0%19.2%18.3%25.5% Spain2,69736.9%29.5%18.1%15.5% France5,12543.0%29.1%12.8%15.1%

The US and UK have the highest p=reject shares in this table (22.2% and 25.5%). Italy stands out for a different reason: it has one of the lower no-record rates (40.9%) but also the highest p=none share of any country here (36.8%), an observed difference consistent with more Italian domains in this cohort having started the DMARC process and stopped at the reporting-only stage, though the data here doesn't establish why. This is a snapshot of CipherCue's tracked cohort, not a national census; see the method note for how that cohort is built.

What else is missing alongside DMARC

DMARC doesn't operate alone. SPF and DKIM are the two authentication mechanisms DMARC builds on top of; MTA-STS, DNSSEC, and BIMI are three adjacent DNS-based controls that address related but separate problems. We checked all five against the same 67,336-domain cohort.

ControlWhat it doesPresentShare SPFAuthorises which mail servers can send for a domain48,96272.7% DMARCSets policy for mail that fails authentication, plus reporting36,97454.9% BIMIDisplays a verified brand logo in supporting inboxes1,7262.6% MTA-STSForces TLS encryption in transit between mail servers9571.4% DNSSECCryptographically signs DNS responses so they can't be forged00.0%

SPF is the most widely adopted control here, which tracks with it being the oldest and simplest to configure (a single DNS TXT record with no reporting infrastructure required). Of the 48,962 domains with SPF, 25,657 (52.4%) use a hard fail (-all) and 21,103 (43.1%) use a soft fail (~all), the difference between "reject" and "flag but accept" for mail that fails the SPF check.

MTA-STS and BIMI both sit under 3%. DNSSEC shows zero validated domains in this cohort, which we're flagging as a measurement caveat rather than a finding: CipherCue's current DNSSEC check validates the full signing chain, and a stricter check will show fewer passes than a looser one that only checks for the presence of DNSKEY or RRSIG records. We would not present 0.0% as a claim that no domain in the cohort has DNSSEC configured at all; we can say confidently that none passed full chain validation in our check, and we're treating the true adoption rate as an open question until we've audited the check itself.

The RFCs behind this, and what changed recently

DMARC's core specification changed in 2026. RFC 7489, the original DMARC specification published in March 2015 as an informational, industry-authored document, has been obsoleted by three new IETF documents published in May 2026:

  • RFC 9989: the core DMARC protocol
  • RFC 9990: aggregate reporting (the rua= reports this article is about)
  • RFC 9991: failure reporting (ruf=)

The practical significance is less about new tags in your DNS record and more about standing. RFC 7489 was published on the Independent Submission stream as Informational, meaning it was never put through IETF working-group consensus. RFC 9989 is Standards Track (Proposed Standard), the first time DMARC has had formal IETF standard status. The tags you already use, v=, p=, sp=, rua=, ruf=, adkim=, aspf=, and fo=, keep their existing meaning; nothing you would need to change today.

The one substantive mechanism change is how a receiver finds the "organisational domain" for a subdomain that has no DMARC record of its own. RFC 7489 used the Public Suffix List, a community-maintained file of which domain suffixes count as registrable (for example, knowing that co.uk is a suffix, not a company). RFC 9989 replaces this with a DNS Tree Walk: the receiver checks for a DMARC record at the exact sending domain, then walks up the domain tree checking each parent, until it finds one or runs out of labels. This removes the dependency on an externally maintained list that isn't itself part of the DNS.

Two adjacent standards referenced in this article, for context on maturity: RFC 8461 (MTA-STS) has been a full IETF Internet Standard since September 2018. BIMI, by contrast, has never been adopted by an IETF working group; as of this article the latest version is an individual Internet-Draft (version 14, dated May 2026) that explicitly carries "no formal standing in the IETF standards process." That gap in standing is one plausible reason BIMI sits at 2.6% adoption in our data versus SPF's 72.7%: implementing a control with no finished specification and an added certification cost (BIMI requires a Verified Mark Certificate from a small number of authorities) is a harder sell than a DNS TXT record against a ratified standard.

Does SOC 2 or ISO 27001 require this?

Neither SOC 2 nor ISO 27001 universally requires DMARC as a specifically named control. An organisation may still implement DMARC as part of its risk-based controls for email authentication, impersonation, and domain abuse.

SOC 2 is built around the AICPA's Trust Services Criteria, not a fixed technical checklist. An auditor evaluates whether the controls a company has chosen satisfy the criteria (most commonly Security, sometimes also Availability, Confidentiality, Processing Integrity, or Privacy); the company picks its own controls, and DMARC is not one of the named examples in the criteria document.

ISO 27001:2022 works the same way. Annex A control 5.14, Information Transfer, requires "rules, procedures, or agreements" governing how information moves within an organisation and with third parties, which is broad enough to cover email authentication without naming it. The 2022 revision consolidated what used to be four separate controls in the 2013 version (old clauses 8.7.1 through 8.7.4) into this single control.

What this looks like for one domain

Cranswick, the UK food producer (cranswick.co.uk), is a real example from the dataset. Its DMARC record is p=none and lists three separate rua= addresses: dmarc_agg@vali.email, a self-registered mailbox at eu.cp-dmarc.com, and a hashed address at rua.easydmarc.eu. Three different destinations, at least two different DMARC-monitoring vendors, for one domain. This is a public DNS record; anyone can look it up with dig TXT _dmarc.cranswick.co.uk. It is not unusual in the dataset, which is the point: reconciling three separate report streams into "who is actually sending mail for us" is exactly the kind of work that has to happen before a domain owner can move off p=none.

Method note

Source and cohort: CipherCue's own DNS observations (direct queries for DMARC, SPF, MTA-STS, BIMI, and DNSSEC), not a third-party dataset. 67,336 domains, observed 2026-04-14 to 2026-07-28. This is a domain count, not a deduplicated-organisation count: a large company can own several domains at different enforcement stages, so the same company can appear more than once.

Vendor mapping ("Who runs DMARC monitoring") uses a manually maintained dictionary of 40 known rua= reporting-endpoint domains and identifies 8,862 entities; it's a floor, not a census, since a name only appears if the destination matched a known pattern.

DNSSEC shows 0.0% because CipherCue's check requires full signing-chain validation; read that as "none passed our validation," not "none have DNSSEC configured."

External references: RFC 7489, RFC 9989, RFC 9990, RFC 9991 (IETF Datatracker), RFC 8461 (IETF Datatracker), draft-brand-indicators-for-message-identification version 14 (IETF Datatracker). SOC 2 Trust Services Criteria (AICPA), ISO/IEC 27001:2022 Annex A control 5.14.

Finding this in your own market

CipherCue's entity directory carries this same DNS compliance data (DMARC, SPF, MTA-STS, BIMI, DNSSEC) as a filterable score band per tracked domain, alongside country. If you want the domains in this article's low-scoring band for a specific country or sector, that's a filter on the entities index, not a custom pull.

Separately, if you're the one staring at your own rua= reports trying to work out who's actually sending mail on your behalf before you flip to enforcement: we built SenderLedger for exactly this problem. It takes the anonymous IPs and mailbox strings in your DMARC aggregate reports and groups them into named sender entities, vendor or system name, source IPs, first seen, last seen, volume, and an authorisation status. It's in early access; there's a waitlist at senderledger.com if you want in.

The Daily Front Page 12 of 25
Tuesday, July 28, 2026 The Daily Front No. #260728 — Ars Astronomica
article

Ars Astronomica – English translations of rare Hebrew and Latin astronomy texts

by sweisman·▲ 105 points·48 comments·arsastronomica.com ↗
“First English translations of historical Hebrew and Latin works in astronomy.”

A scholarly imprint producing first English translations of historical Hebrew and Latin works in astronomy, cosmology, and natural philosophy.

These texts span centuries, and most have never appeared in English. One – Gersonides’ 136-chapter mathematical astronomy – was never printed and survives only in manuscript. They cite and answer one another, addressing the same cosmological questions, yet until now no one could read them side by side in a single language.

The corpus

Date Author Work Source Description Translator Version 1665 Athanasius Kircher Mundus Subterraneus
footnotes Internet Archive — athanasiikircher12kirc A vast attempt to explain the whole hidden world beneath our feet. 2 c. 1613 David Gans Nechmad ve-Na’im
footnotes National Library of Israel — Rosetta IE32709375, shelfmark 35 V 1770 (catalog NNL_ALEPH990011964030205171). A Hebrew textbook of astronomy that presents the Ptolemaic model of the heavens for a Hebrew-reading audience, weaving the cutting-edge European science of its day together with the medieval Jewish cosmological tradition — its author having personally visited the Danish astronomer Tycho Brahe. 2 1592 (first composed/printed Prague, שנ”ב; this is a later reprint) David Gans Tzemach David
footnotes National Library of Israel — Rosetta IE90688593 (record labeled “Vol.1-3”). A Hebrew historical chronicle in two parts by David Gans: the first traces Jewish history from creation to the sixteenth century — biblical figures, sages, and rabbis; the second surveys world and gentile history from Julius Caesar to Emperor Rudolph. 2 1665 Giovanni Battista Riccioli Astronomia Reformata
footnotes e-rara ETH-Bibliothek Zürich — title-info 141744 A mature reckoning with the heavens as the telescope was revealing them — a sequel to the vast Almagestum Novum (1651), built on further years of observation at Bologna with the collaborator Grimaldi, with fresh measurements and corrected parameters at center stage. 2 c. 1123 (composition; this is a later manuscript copy) Jacob ben Samson (attrib.) Perush Sod ha-Ibbur
footnotes OPenn UPenn-hosted digitization of the British Library manuscript — British Library Add MS 11639, ff. 511r–545v An early Ashkenazi commentary on the “secret of the calendar” (sod ha-ibbur) — the computation of the Hebrew calendar: the molad (lunar conjunction), the tequfot (seasonal points), the nineteen-year intercalation cycle, and the rules governing leap years and festival dates. 2 1629 Joseph Solomon Delmedigo Sefer Elim
footnotes Internet Archive — seferelim00delmuoft A wide-ranging Hebrew scientific compendium, framed as a reply to questions put by the Karaite scholar Zerah ben Nathan. 2 1640 Longomontanus Astronomia Danica
footnotes Internet Archive — astronomiadanica00long A comprehensive textbook of astronomy and the definitive technical account of the Tychonic system — Earth at rest, the Sun circling it, the planets circling the Sun. 2 1654 Pierre Gassendi Tychonis Brahei Vita
footnotes Internet Archive — den-kbd-pil-130018157889-001 A life of the Danish astronomer Tycho Brahe (1546–1601): his fabled observatory on the island of Hven, the magnificent instruments he built, the campaigns of observation that remade astronomy, his bold model of the cosmos, and his turbulent final years in Prague. 2 c. 1136 R. Avraham bar Ḥiyya ha-Nasi Cheshbon Mahalechot ha-Kochavim
footnotes HebrewBooks.org — ID 22072 A medieval Hebrew treatise on mathematical astronomy and the computation of the Jewish calendar. 2 c. 1122 R. Avraham bar Ḥiyya ha-Nasi Sefer HaIbbur
footnotes HebrewBooks.org — ID 21292 The earliest systematic Hebrew treatise on the science of the calendar, written in early-twelfth-century Barcelona. 2 c. 1132 R. Avraham bar Ḥiyya ha-Nasi Sefer Tsurat ha-Arets
footnotes NYPL Digital Collections Dorot Jewish Division — item ce668260-c766-0132-045a-58d385a7b928 A foundational work of medieval cosmology that lays out the shape of the cosmos as the Greeks and their Arabic heirs understood it: a spherical earth at the center of nested celestial spheres, the paths of sun and moon, and the zones and geography of the inhabited world. 2 c. 1148 R. Avraham ibn Ezra Keli ha-Nechoshet
footnotes HebrewBooks.org — ID 20850 The earliest surviving Hebrew treatise on the astrolabe, composed in the mid-twelfth century. 2 c. 1500 R. Eliyahu Mizrahi Kitsur ha-Melakhat ha-Mispar
footnotes NYPL Digital Collections Dorot Jewish Division — item ce668260-c766-0132-045a-58d385a7b928 Kitsur ha-Melakhat ha-Mispar (“Compendium of the Art of Number”) is a Hebrew arithmetic textbook by Rabbi Eliyahu Mizrahi (ca. 1455–1526), the chief rabbi of the Ottoman Empire and one of the foremost Jewish mathematicians of his age. 2 c. 1365 R. Immanuel Bonfils Shesh Kenafayim
footnotes University of Pennsylvania — Kislak Center, LJS 204 Digital Scriptorium DS129 One of the most widely copied Hebrew astronomical handbooks of the late Middle Ages, composed around 1365 in Provence. 2 late 14th c. (composition; this is a later manuscript copy) R. Isaac ibn al-Aḥdab Keli Ḥemda
footnotes Gallica / BnF — Hébreu 1031, ff. 208r–215v Isaac ibn al-Aḥdab’s Keli Ḥemda (“The Precious Instrument”) — a description of the construction and use of an equatorium, an instrument for computing planetary positions geometrically. 2 c. 1396 (composed in Syracuse, Sicily; this is a later manuscript copy) R. Isaac ibn al-Aḥdab Orah Selulah
footnotes Gallica / BnF — Hébreu 1086 A set of astronomical tables by Isaac ibn al-Ahdab (Sicily, late 14th c.) for computing the true conjunctions and oppositions of the Sun and Moon — the Orah Selulah (“Paved Way”), one of the most widely copied Hebrew astronomical table-works (surviving in ~25 manuscripts). 2 c. 1288 (composition; this is a later manuscript copy) R. Jacob ben Machir ibn Tibbon Roba’ Yisrael
footnotes Gallica / BnF — Hébreu 1031, ff. 131r–147v Jacob ben Machir ibn Tibbon’s treatise on the quadrans novus — the improved astronomical quadrant of his own invention — describing how the instrument is constructed and used for astronomical and time-keeping measurements. 2 c. 1329 R. Levi ben Gershom Milchamot HaShem
footnotes Internet Archive — sefermilamothash00leviuoft The Wars of the Lord is the philosophical and theological masterwork of Levi ben Gershom (Gersonides, 1288–1344), one of the boldest minds of medieval Provence. 2 c. 1328 R. Levi ben Gershom Sefer HaTechunah
footnotes Gallica / BnF — MS Hébreu 724, ff. 1r–257v A complete medieval system of mathematical astronomy, formally Book V, Part 1 of the Wars of the Lord. 2 15th c. (composition; this is a later manuscript copy) R. Mordecai Comtino Sefer ha-Ḥeshbon ve-ha-Middot
footnotes Gallica / BnF — Hébreu 1031, ff. 26r–65r An instructional treatise on arithmetic and practical geometry by Mordecai ben Eliezer Comtino (Constantinople, 15th c.), one of the leading Rabbanite scholars of Byzantine Jewry. 2 1797 R. Pinchas Eliyahu Hurwitz Sefer HaBrit HaShalem
footnotes HebrewBooks.org — ID 43670 A sweeping Hebrew encyclopedia of science and mysticism, among the most widely read Hebrew books of its era. 2 c. 1310 R. Yitzchak Yisraeli Yesod Olam
footnotes Digital Bodleian — MS. Huntington 299 One of the great medieval Hebrew treatises on astronomy and the science of the Jewish calendar, composed in Toledo in the first half of the fourteenth century. 2 1602 Tycho Brahe Astronomiae Instauratae Mechanica
footnotes Internet Archive — gri_tychonisbrah00brah A sumptuously illustrated catalog of the most accurate astronomical instruments built before the telescope, gathered on the island observatory of Hven. 2 1610 Tycho Brahe Astronomiae Instauratae Progymnasmata
footnotes e-rara ETH-Bibliothek Zürich — Rar 4153 DOI 10.3931/e-rara-315, object ID 84169 The fullest statement of the observational program that transformed the science of the heavens. 2 1588 Tycho Brahe De Mundi Aetherei Recentioribus Phaenomenis
footnotes Internet Archive — bub_gb_2f-EqKxRN34C When a brilliant comet blazed across Europe in 1577, the finest instruments of the age were turned upon it, and the measurements shattered the ancient belief in unchanging, perfect heavens: the comet moved freely where solid crystalline spheres were supposed to be. 2 1596 Tycho Brahe Epistolarum Astronomicarum Libri
footnotes Internet Archive — tychonisbrahedan00brah Before scientific journals existed, astronomers argued, boasted, and traded discoveries by letter. 2 c. 150–850 CE Unknown Mishnat ha-Middot
footnotes HebrewBooks.org — ID 39044 The earliest known Hebrew treatise on geometry — a concise practical manual of mensuration covering the areas and perimeters of plane figures, the measurement of circles and segments, and an early value for π. 2

Updates

2026-07-05

  • Added eight works to the corpus: Milchamot HaShem (Gersonides), Sefer HaBrit HaShalem (Hurwitz), Tzemach David (Gans), Keli Ḥemda and Orah Selulah (ibn al-Aḥdab), Roba’ Yisrael (ibn Tibbon), Perush Sod ha-Ibbur (Jacob ben Samson), and Sefer ha-Ḥeshbon ve-ha-Middot (Comtino).
  • Cheshbon Mahalechot ha-Kochavim (bar Ḥiyya ha-Nasi) was retitled from its earlier heading, Po’al ha-Shem ve-Cheshbon Mahalechot ha-Kochavim.
  • Some works now carry improved source records – Nechmad ve-Na’im, for one, now cites the National Library of Israel’s own digitized exemplar.

2026-06-29

  • All translations are now at version 2, regenerated with the corrected pipeline. Version 1 turned out to be an intermediate stage and has been superseded.
  • Added new works to the corpus – most recently Sefer Tsurat ha-Arets, Kitsur ha-Melakhat ha-Mispar, and Sefer Elim.
  • Draft editorial footnotes are now available for download from a footnotes link beneath each work’s title.

About the translations

Translations are produced with an automated, AI-assisted pipeline that runs each text through a multi-stage workflow before final collation. For technical details, see the translator’s translation-pipeline repo.

These translations are prepublication texts – nearly publication-ready, pending the diagrams and illustrations still in progress.

The goal is fluent, readable modern English – a text an educated reader can follow without reaching for the Latin or Hebrew, not one written only for specialists. Rather than reproduce the long periodic sentences and deferred verbs of the originals, the translations carry the author’s meaning, argument, and tone into natural contemporary prose: where Tycho builds a single sentence whose main verb arrives only after four nested clauses, the translation breaks it into the two or three a modern writer would use. Fidelity comes first, though – nothing is dropped, summarized, or invented, and negations, numbers, technical terms, and the author’s own examples and analogies are preserved exactly. Readability never comes at the cost of changing what the source says.

Technical vocabulary stays precise and is anchored to its modern equivalents. A key Latin or Hebrew term of art is glossed on its first occurrence – the original shown alongside the English rendering – and thereafter carried in settled English. Historical names, star names, and specialized vocabulary carry inline identifications: Tycho’s “Lucida Vulturis volantis” is identified as Altair in Aquila; Gersonides’ medieval Hebrew astronomical terminology is mapped to the Ptolemaic system it describes.

Mathematical and astronomical content – sexagesimal values, spherical triangle computations, calendar arithmetic, tabular data – is reproduced with the precision of the originals, cell by cell and degree by degree. Compositor errors in the source are corrected inline with [recte: ...] notation rather than silently emended. Uncertain readings due to ink damage, worn type, or ambiguous letterforms are marked with [?], preserving the translator’s best reading while flagging it for editorial review.

Illustrations and diagrams

The publicly available PDFs are text-only. Geometric figures, instrument schematics, concentric-sphere charts, and other diagrams from the source works are extracted and replaced with structured descriptions identifying every label, arc, point, and geometric relationship visible in the original woodcut or engraving. Tycho’s De Mundi Aetherei, for example, carries 93 such figures, from spherical-astronomy constructions to the two-circle hypothesis of the comet’s eccentric orbit within the solar sphere. For a high-quality illustrated print edition of any of these works, please get in touch using the contact details below.

Support this project

Upright reason dictates that the recipients of the good, of whatever type and level of beneficence it may be, must show gratitude and blessing to the beneficent person in every way possible, commensurate with the value of the beneficence. And one who has benefited all the people of the world – for example, one who invented a new instrument for the good of the world, or a good book – it is fitting for every discerning person, out of the obligation of love of fellow beings, to at least purchase it, so that the man will profit and his heart will be encouraged thereby to invent yet more good instruments in the world. And similarly, all other wise-hearted people will likewise strive and exert themselves to invent good things and instruments needed for the repair of the world and its perfection.

And therefore, whoever says: “What do I need this new instrument for?” – he does not act well towards the world. For if not for the man who invented it, where would you be? And what would your city do? And more than this, if the man did not exert himself, what would the world do?

R. Pinchas Eliyahu Hurwitz, Sefer HaBrit

If you find these translations useful, you can support me on Ko-fi.

Contact

License

The original works are in the public domain and were obtained from a variety of sources, including digitized library collections and other open archives. No claim of ownership is made over the source texts.

All translations in this collection are © Scott Weisman. All rights reserved, except as granted by the license below.

The translations are made available under the Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International license. You are free to share them for non-commercial purposes with attribution; you may not modify them or use them commercially without prior written permission.

Acknowledgements

Translated with the assistance of Claude. The translator thanks the Anthropic team for making this work possible.

The Daily Front Page 13 of 25
Tuesday, July 28, 2026 The Daily Front No. #260728 — Profiling eBPF Code
article

How Do I Profile eBPF Code?

by snaveen·▲ 115 points·6 comments·naveensrinivasan.com ↗
“We want to measure its performance impact.”

If we are running any eBPF workload or writing eBPF code, we want to measure its performance impact, and in this post we will demonstrate an example of how to do it.

In this example, our goal is to measure the performance of file open operations, one of the most critical functions in the OS. Our code used file open hooks in eBPF, we wanted to measure the performance overhead introduced by adding this hook.

To identify likely bottlenecks, we need a simple C test harness with the fewest dependencies, designed to measure file open performance.

#define _GNU_SOURCE
#include <fcntl.h>
#include <stdio.h>
#include <time.h>
#include <stdint.h>
#include <stdlib.h>
#include <unistd.h>
#include <sched.h>
#include <sys/mman.h>
#include <sys/syscall.h>

static inline uint64_t now_ns(void) {
    struct timespec ts;
    clock_gettime(CLOCK_MONOTONIC, &ts);   /* VDSO, no syscall */
    return (uint64_t)ts.tv_sec * 1000000000ull + ts.tv_nsec;
}

int main(int argc, char **argv) {
    const char *path = argv[1];
    uint64_t n      = strtoull(argv[2], NULL, 10);
    uint64_t warm   = n / 10;

    /* preallocate + prefault + lock: no page faults in the loop */
    uint32_t *d = mmap(NULL, n * sizeof(uint32_t), PROT_READ|PROT_WRITE,
                       MAP_PRIVATE|MAP_ANONYMOUS|MAP_POPULATE, -1, 0);
    mlock(d, n * sizeof(uint32_t));
    for (uint64_t i = 0; i < n; i++) d[i] = 0;   /* fault everything in */

    for (uint64_t i = 0; i < n; i++) {
        uint64_t t0 = now_ns();
        long fd = syscall(SYS_openat, AT_FDCWD, path, O_RDONLY);   /* raw, no libc wrapper */
        uint64_t t1 = now_ns();
        if (fd >= 0) close(fd);
        d[i] = (uint32_t)(t1 - t0);
    }

    /* dump after the loop only */
    for (uint64_t i = warm; i < n; i++) printf("%u\n", d[i]);
    return 0;
}

The goal of the code example above is to keep it simple and reopen the same file under warm cache conditions, minimizing unrelated filesystem and disk I/O variability and helping identify the p50/p99. The code invokes syscall(SYS_openat, …) instead of the libc openat() wrapper, and the first 10% of the results are discarded as a warmup period.

This test harness would produce results of opening a file x number of times and how long it took to open. So we could use this to measure the before and after of when the eBPF hook attached.

Setup

When profiling the eBPF code, we want perf tool to be able to resolve symbols so we can analyze where the issue is in our code. To do that, we have to run these commands.

sudo sysctl -w net.core.bpf_jit_enable=1
sudo sysctl -w net.core.bpf_jit_kallsyms=1

The above commands enable jit and expose jit-compiled BPF symbols so the perf report can display program names instead of unknown addresses.

To check whether the symbols appear in the perf tool, run your eBPF code and use a command like this.

sudo bpftool prog show | rg -A4 ' lsm '
sudo rg 'bpf_prog_[0-9a-f]+_ '/proc/kallsyms | rg 'security|path|file|open'

In the above rg command, we are checking for lsm as we are measuring LSM hooks.

Also, we are using a custom kernel version, so perf for that kernel version is not in the standard path, and we have it installed in our example: PERF=/usr/lib/linux-tools/6.8.0-134-generic/perf

Measuring

Now that we have set up all the necessary tools, the first step is to measure without the eBPF code running, and this is where using the above C code can help. We measure the code by opening the file /etc/hostname and piping the results to a file so that we can calculate the p50/p99.

sudo taskset -c 3 chrt -f 99 ./bench /etc/hostname 100000 > /tmp/samples.txt 

The taskset -c 3 pins execution to CPU 3, reducing CPU migration noise, and chrt -f 99 gives the benchmark extremely high CPU priority. It runs before almost all normal programs and keeps running until it finishes, blocks, or is interrupted. The C code discards the first 10%; the file should contain 90,000 samples.txt.

Next, run the eBPF code and execute something like this.

sudo $PERF record \
  -g \
  --call-graph fp \
  -e cycles:k \
  -F 997 \
  -o ~/perf.data \
  -- \
  taskset -c 3 \
  chrt -f 99 \
  ./bench /etc/hostname 200000 \
  > /tmp/samples.txt

The -g records the call stacks, --call-graph fp unwinds stacks using frame pointers, and -e samples CPU cycles in kernel mode only, which includes syscall, VFS, LSM, and eBPF execution and not userspace benchmark work. The -F 997 requests 997 samples per second, and the non-round frequency helps avoid periodic alignment.

After running the above, run this command to sort the data.

sudo "$PERF" report -i ~/perf.data --stdio --sort comm,dso,symbol > perf.txt 

alt text

Flamegraph of the same perf.data, generated with Inferno. The stack of interest here is bpf_lsm_file_open and everything above it.

Here is an example output from perf.txt, which shows that the time being spent on bpf_lsm_file_open and its tail calls is where the performance bottleneck is. This turned out to be in a hot path, which meant every allocation-to-CPU cycle shaving will make a significant impact on the performance of the system.

|
|          |                     |                     |          |          |–90.52%–do_dentry_open
|          |                     |                     |          |          |          |
|          |                     |                     |          |          |           --89.78%–bpf_lsm_file_open
|          |                     |                     |          |          |                     |
|          |                     |                     |          |          |                      --89.30%–0xffffffffc0288c18
|          |                     |                     |          |          |                                |
|          |                     |                     |          |          |                                |–87.40%–bpf_prog_b06f413955402a4b_tail_call_security_check
|          |                     |                     |          |          |                                |          |
|          |                     |                     |          |          |                                |          |–77.57%–bpf_prog_934361d723613c1c_enforce_access_policy
|          |                     |                     |          |          |                                |          |          |
|          |                     |                     |          |          |                                |          |          |–57.87%–bpf_prog_a0f18f4b0b140d77_path_check_callback
|          |                     |                     |          |          |                                |          |          |          |
|          |                     |                     |          |          |                                |          |          |          |–29.94%–bpf_probe_read_kernel
|          |                     |                     |          |          |                                |          |          |          |          |
|          |                     |                     |          |          |                                |          |          |          |          |–18.88%–copy_from_kernel_nofault
|          |                     |                     |          |          |                                |          |          |          |          |
|          |                     |                     |          |          |                                |          |          |          |           --4.99%–copy_from_kernel_nofault_allowed
|          |                     |                     |          |          |                                |          |          |          |
|          |                     |                     |          |          |                                |          |          |          |–4.88%–htab_map_hash
|          |                     |                     |          |          |                                |          |          |          |
|          |                     |                     |          |          |                                |          |          |           --2.29%–copy_from_kernel_nofault

This post focuses on the profiling method rather than a specific result, and the overhead you’ll see depends heavily on what your hook actually does, so we’re leaving the numbers out and focusing on how to get them yourself.

Now, from the above, we can start analyzing where the time is being spent and perf-tune the code along with the p50/p99 of the C code with eBPF running.

By doing this, we can clearly identify the perf impact of the eBPF code and likely pinpoint where optimizations are required. It can be as simple as caching something or coming up with a better algorithm based on where the problem is.


The Daily Front Page 14 of 25
Tuesday, July 28, 2026 The Daily Front No. #260728 — macOS Tahoe 26.6 Security Notes
article

About the security content of macOS Tahoe 26.6

by andor·▲ 203 points·134 comments·support.apple.com ↗
“Apple doesn’t disclose, discuss, or confirm security issues until patches or releases are available.”

This document describes the security content of macOS Tahoe 26.6.

About Apple security updates

For our customers' protection, Apple doesn't disclose, discuss, or confirm security issues until an investigation has occurred and patches or releases are available. Recent releases are listed on the Apple security releases page.

Apple security documents reference vulnerabilities by CVE-ID when possible.

For more information about security, see the Apple Product Security page.

macOS Tahoe 26.6

Released July 27, 2026

Accounts

Available for: macOS Tahoe

Impact: An app may be able to access sensitive user data

Description: An access issue was addressed with additional sandbox restrictions.

CVE-2026-43819: Omar Cerrito

Accounts

Available for: macOS Tahoe

Impact: An app may be able to gain root privileges

Description: A parsing issue in the handling of directory paths was addressed with improved path validation.

CVE-2026-43749: Adam Franke, Ashish Kunwar, Trung Nguyen (@everping) of CyStack

Accounts Framework

Available for: macOS Tahoe

Impact: An app may be able to fingerprint the user

Description: This issue was addressed with improved data protection.

CVE-2026-64733: Rosyna Keller of Totally Not Malicious Software (paradisefacade.com)

afpfs

Available for: macOS Tahoe

Impact: A remote attacker may be able to cause unexpected system termination or corrupt kernel memory

Description: A buffer overflow was addressed with improved bounds checking.

CVE-2026-64767: Dave G.

apache

Available for: macOS Tahoe

Impact: A remote attacker may be able to cause a denial-of-service

Description: This is a vulnerability in open source code and Apple Software is among the affected projects. The CVE-ID was assigned by a third party. Learn more about the issue and CVE-ID at cve.org.

CVE-2026-23918: Юлия Мерцалова

APFS

Available for: macOS Tahoe

Impact: A remote user may be able to cause unexpected system termination or corrupt kernel memory

Description: The issue was addressed with improved memory handling.

CVE-2026-64695: Peter Malone

App Store

Available for: macOS Tahoe

Impact: An app may be able to access sensitive user data

Description: This issue was addressed with improved checks.

CVE-2026-43801: Rahul Raj

Apple Account

Available for: macOS Tahoe

Impact: An app may be able to access sensitive user data

Description: A race condition was addressed with improved state handling.

CVE-2026-43781: Pinak Oza

Apple Account

Available for: macOS Tahoe

Impact: A malicious app may be able to break out of its sandbox

Description: An authorization issue was addressed with improved state management.

CVE-2026-64737: Robert Mindo

Apple Neural Engine

Available for: macOS Tahoe

Impact: An app may be able to cause unexpected system termination

Description: An out-of-bounds write issue was addressed with improved bounds checking.

CVE-2026-43748: an anonymous researcher, tamdao, Franco Belman at Blackwing Intelligence

Apple Neural Engine

Available for: macOS Tahoe

Impact: An app may be able to cause unexpected system termination

Description: A use after free issue was addressed with improved memory management.

CVE-2026-28928: Dun

AppleDouble

Available for: macOS Tahoe

Impact: Processing a maliciously crafted file may lead to unexpected app termination or arbitrary code execution

Description: A buffer overflow was addressed with improved bounds checking.

CVE-2026-43776: Irvin Wang, Peter Malone, Nicolas Rabrenovic

AppleRAID

Available for: macOS Tahoe

Impact: A local user may be able to read kernel memory

Description: A buffer overflow was addressed with improved bounds checking.

CVE-2026-43681: impost0r (ret2plt), David Ige - Beryllium Security

Assets

Available for: macOS Tahoe

Impact: A malicious application may be able to bypass Privacy preferences

Description: An authorization issue was addressed with improved state management.

CVE-2026-43672: 이재영

ATS

Available for: macOS Tahoe

Impact: An app may be able to read files outside of its sandbox

Description: A permissions issue was addressed by removing the vulnerable code.

CVE-2026-43763: Pavan Nallamothu, Jared Reyes

Audio

Available for: macOS Tahoe

Impact: An app may be able to break out of its sandbox

Description: An access issue was addressed with additional sandbox restrictions.

CVE-2026-64702: John Lussier

Audio

Available for: macOS Tahoe

Impact: An app may be able to cause a denial-of-service

Description: An out-of-bounds write issue was addressed with improved bounds checking.

CVE-2026-64725: Seonung Park, ALTV!ST (altvi.st/)

AuthKit

Available for: macOS Tahoe

Impact: An app may be able to fingerprint the user

Description: A permissions issue was addressed with additional restrictions.

CVE-2026-43730: David Strnadel

AVEVideoEncoder

Available for: macOS Tahoe

Impact: An app may be able to execute arbitrary code with kernel privileges

Description: A buffer overflow was addressed with improved size validation.

CVE-2026-64747: Franco Belman at Blackwing Intelligence

AVEVideoEncoder

Available for: macOS Tahoe

Impact: An app may be able to cause unexpected system termination

Description: An out-of-bounds read was addressed with improved bounds checking.

CVE-2026-64762: Dun, Franco Belman at Blackwing Intelligence

BackgroundAssets

Available for: macOS Tahoe

Impact: An app may be able to delete files for which it does not have permission

Description: A permissions issue was addressed with improved validation.

CVE-2026-64707: YingQi Shi (@Mas0nShi) of DBAppSecurity's WeBin lab

cd9660

Available for: macOS Tahoe

Impact: An app may be able to cause unexpected system termination or read kernel memory

Description: The issue was addressed with improved memory handling.

CVE-2026-64698: an anonymous researcher, Richard Zana, Nathaniel Oh (@calysteon), Peter Malone

CloudAttestation

Available for: macOS Tahoe

Impact: A maliciously crafted app may be able to bypass code signing enforcement

Description: A validation issue was addressed with improved input sanitization.

CVE-2026-43813: Anton Pakhunov

Contacts

Available for: macOS Tahoe

Impact: An app may be able to add contacts without user authorization

Description: An authorization issue was addressed with improved validation.

CVE-2026-64746: Rodolphe BRUNETTI (@eisw0lf) of Lupus Nova, Daniel Febrero

Contacts

Available for: macOS Tahoe

Impact: Processing a maliciously crafted contact may leak sensitive data

Description: The issue was addressed with improved checks.

CVE-2026-64734: Daniel Williams

Contacts

Available for: macOS Tahoe

Impact: An app may be able to access information about a user's contacts

Description: This issue was addressed with improved checks.

CVE-2026-43797: Arni Hardarson (Neonix Security)

Control Center

Available for: macOS Tahoe

Impact: An app may be able to access user-sensitive data

Description: A logic issue was addressed with improved validation.

CVE-2026-43756: 이재영

Core Services

Available for: macOS Tahoe

Impact: An app may be able to gain root privileges

Description: A race condition was addressed with improved state handling.

CVE-2026-43693: Gergely Kalman (@gergely_kalman)

CoreAudio

Available for: macOS Tahoe

Impact: Processing a maliciously crafted audio file may corrupt process memory

Description: The issue was addressed with improved memory handling.

CVE-2026-43673: Anonymous working with TrendAI Zero Day Initiative

CoreAudio

Available for: macOS Tahoe

Impact: Processing an audio stream in a maliciously crafted media file may terminate the process

Description: An out-of-bounds write issue was addressed with improved bounds checking.

CVE-2026-43744: Mathis Mansière, an anonymous researcher

CoreAudio

Available for: macOS Tahoe

Impact: A remote attacker may be able to cause unexpected system termination

Description: An out-of-bounds write issue was addressed with improved bounds checking.

CVE-2026-43803: Rahul Raj

CoreMedia

Available for: macOS Tahoe

Impact: An app may be able to access sensitive user data

Description: An authorization issue was addressed with improved state management.

CVE-2026-43775: Csaba Fitzl (@theevilbit) of Iru

CVE-2026-43759: 이재영, Rajdip Dey Sarkar, Arni Hardarson (Neonix Security)

CoreMedia

Available for: macOS Tahoe

Impact: Processing a maliciously crafted video file may lead to unexpected app termination

Description: A memory corruption issue was addressed with improved memory handling.

CVE-2026-43711: James Duffy (@0x4A616D657344)

CoreVideo

Available for: macOS Tahoe

Impact: An app may be able to cause unexpected system termination

Description: An out-of-bounds write issue was addressed with improved bounds checking.

CVE-2026-43802: an anonymous researcher

Crash Reporter

Available for: macOS Tahoe

Impact: An app may be able to leak sensitive user information

Description: A privacy issue was addressed by removing sensitive data.

CVE-2026-64710: Matthew Schneider

CUPS

Available for: macOS Tahoe

Impact: A malicious app may be able to gain root privileges

Description: A permissions issue was addressed with additional restrictions.

CVE-2026-39875: Dallas Dubs, Aaron Grattafiori - NVIDIA AI Red Team, XBreach.ai, Andreas Jaegersberger & Ro Achterberg of Nosebeard Labs

CUPS

Available for: macOS Tahoe

Impact: An app may be able to gain root privileges

Description: An injection issue was addressed with improved validation.

CVE-2026-43698: Andreas Jaegersberger & Ro Achterberg of Nosebeard Labs

curl

Available for: macOS Tahoe

Impact: Authentication credentials may be sent to a server on another origin

Description: This is a vulnerability in open source code and Apple Software is among the affected projects. The CVE-ID was assigned by a third party. Learn more about the issue and CVE-ID at cve.org.

CVE-2026-3784

CVE-2026-3783

Data Detectors UI

Available for: macOS Tahoe

Impact: An app may be able to access sensitive user data

Description: An authorization issue was addressed with improved state management.

CVE-2026-43758: HvxyZLF

DesktopServices

Available for: macOS Tahoe

Impact: An app may bypass Gatekeeper checks

Description: A file quarantine bypass was addressed with additional checks.

CVE-2026-64708: Lance Cain - Offensive Security Engineer, SpecterOps Inc.

Disk Images

Available for: macOS Tahoe

Impact: An app may be able to disclose kernel memory

Description: The issue was addressed with improved bounds checks.

CVE-2026-64776: Hyunwoo Kim (@v4bel)

Disk Images

Available for: macOS Tahoe

Impact: An app may be able to cause unexpected system termination

Description: An integer overflow was addressed with improved input validation.

CVE-2026-64694: Gil Portnoy & Henry

Disk Images

Available for: macOS Tahoe

Impact: Parsing a maliciously crafted file may lead to an unexpected app termination

Description: An out-of-bounds read was addressed with improved bounds checking.

CVE-2026-43747: Anthony Laou Hine Tsuei (@anarcheuz)

Disk Images

Available for: macOS Tahoe

Impact: An app may be able to bypass network restrictions

Description: A permissions issue was addressed with additional sandbox restrictions.

CVE-2026-28945: Ayaan Ahmad

DriverKit

Available for: macOS Tahoe

Impact: An app may be able to cause unexpected system termination

Description: An issue existed in the handling of environment variables. This issue was addressed with improved validation.

CVE-2026-43793: erdene-och Byambabayar

DriverKit

Available for: macOS Tahoe

Impact: An attacker with physical access to a locked device may be able to view sensitive user information

Description: An out-of-bounds read was addressed with improved bounds checking.

CVE-2026-43753: Niels Hofmans

Foundation

Available for: macOS Tahoe

Impact: A malicious app may be able to access protected user data

Description: The issue was addressed with improved input sanitization.

CVE-2026-43714: an anonymous researcher

Game Center

Available for: macOS Tahoe

Impact: A malicious app may be able to break out of its sandbox

Description: A parsing issue in the handling of directory paths was addressed with improved path validation.

CVE-2026-64740: Manuel Fernandez (Stackhopper Security)

Game Center

Available for: macOS Tahoe

Impact: An app may be able to access sensitive user data

Description: This issue was addressed with improved data protection.

CVE-2026-43796: Ilya Andr (andrd3v) of Positive Technologies, Stanislav Jelezoglo

GPU Drivers

Available for: macOS Tahoe

Impact: An app may be able to cause unexpected system termination

Description: A buffer overflow was addressed with improved size validation.

CVE-2026-64691: Somair Ansar, Josh Maine of Calif.io, Johnny Franks (@zeroxjf), hxr1, Alexandre Soleiman, Alexander Tarasikov and Ruslan Dautov, Ruslan Dautov

Heimdal

Available for: macOS Tahoe

Impact: An app may be able to cause a denial-of-service

Description: An out-of-bounds read was addressed with improved bounds checking.

CVE-2026-64692: Redon Gashi

HFS

Available for: macOS Tahoe

Impact: A remote user may be able to cause unexpected system termination or corrupt kernel memory

Description: The issue was addressed with improved memory handling.

CVE-2026-43682: Trung Nguyen (@everping) of CyStack, Dave G., Nicolas Rabrenovic, Atul R V & Ashmit Sharma, Peter Malone

HFS

Available for: macOS Tahoe

Impact: Processing a maliciously crafted image may lead to arbitrary code execution

Description: A buffer overflow was addressed with improved bounds checking.

CVE-2026-28981: Hcamael and 章鱼哥@aipy (aipyaipy.com), Aswin Kumar Gokulakannan, Surya Narayan Kushwaha, Dun

HFS

Available for: macOS Tahoe

Impact: Mounting a maliciously crafted disk image may cause unexpected system termination or corrupt kernel memory

Description: An out-of-bounds read was addressed with improved bounds checking.

CVE-2026-43773: Richard Zana, Peter Malone, Surya Narayan Kushwaha, Hyunwoo Kim (@v4bel), Cem Onat Karagun

HFS

Available for: macOS Tahoe

Impact: An app may be able to cause unexpected system termination

Description: The issue was addressed with improved memory handling.

CVE-2026-43767: Hyunwoo Kim (@v4bel)

HFS

Available for: macOS Tahoe

Impact: An app may be able to cause unexpected system termination

Description: An integer overflow was addressed with improved input validation.

CVE-2026-43764: Tristan Madani (@TristanInSec) from Talence Security

HFS

Available for: macOS Tahoe

Impact: An app may be able to cause unexpected system termination or corrupt kernel memory

Description: The issue was addressed with improved memory handling.

CVE-2026-64697: Peter Malone

HFS

Available for: macOS Tahoe

Impact: An attacker may be able to cause unexpected system termination or corrupt kernel memory

Description: The issue was addressed with improved memory handling.

CVE-2026-43710: Peter Malone

ImageIO

Available for: macOS Tahoe

Impact: Processing a maliciously crafted texture may lead to unexpected app termination

Description: An integer overflow was addressed with improved input validation.

CVE-2026-43780: Michael DePlante (@izobashi) of TrendAI Zero Day Initiative

ImageIO

Available for: macOS Tahoe

Impact: Processing a maliciously crafted image may lead to arbitrary code execution

Description: An integer overflow was addressed with improved input validation.

CVE-2026-43818: an anonymous researcher

ImageIO

Available for: macOS Tahoe

Impact: Processing a maliciously crafted image may corrupt process memory

Description: The issue was addressed with improved memory handling.

CVE-2026-64716: Arni Hardarson, Jonathan Alush-Aben, Peter Malone

ImageIO

Available for: macOS Tahoe

Impact: Processing a maliciously crafted file may lead to unexpected app termination

Description: The issue was addressed with improved bounds checks.

CVE-2026-64758: 진규정 (Gyujeong Jin, @G1uN4sh)

ImageIO

Available for: macOS Tahoe

Impact: Processing a maliciously crafted file may lead to a denial-of-service

Description: An out-of-bounds write issue was addressed with improved bounds checking.

CVE-2026-64754: PETOWORKS의 Bugeun Choi (@Bugeun), Rahul Raj

ImageIO

Available for: macOS Tahoe

Impact: Processing a maliciously crafted image may lead to a denial-of-service

Description: A type confusion issue was addressed with improved checks.

CVE-2026-64693: Geonha Lee (@leegn4a)

IOKit

Available for: macOS Tahoe

Impact: An app may be able to cause unexpected system termination or write kernel memory

Description: A race condition was addressed with improved state handling.

CVE-2026-43805: 이재영

Kernel

Available for: macOS Tahoe

Impact: An app may be able to access sensitive user data

Description: This issue was addressed with improved checks.

CVE-2026-43782: Igor Ushakov

Kernel

Available for: macOS Tahoe

Impact: An app may be able to cause unexpected system termination or corrupt kernel memory

Description: The issue was addressed with improved memory handling.

CVE-2026-64749: Billy Jheng Bing Jhong and Pan Zhenpeng (@Peterpan0927) of STAR Labs SG Pte. Ltd., Ashish Kunwar, Hiroki Imai (LAC Co., Ltd.), hxr1

Kernel

Available for: macOS Tahoe

Impact: An app may be able to disclose kernel memory

Description: An information leakage was addressed with additional validation.

CVE-2026-64744: Ryan Hileman via Xint Code (xint.io)

Kernel

Available for: macOS Tahoe

Impact: An app may be able to cause unexpected system termination or corrupt kernel memory

Description: A use after free issue was addressed with improved memory management.

CVE-2026-43778: f0r of MurphySec, Feng Xue and XGPT of ThreatBook, Mahmoud Abdelmoniem, an anonymous researcher, Wang Yu, Lyutoon, Hiroki Imai (LAC Co., Ltd.), DARKNAVY (@DarkNavyOrg), Fábio Luís @scanpt, Nicolas Rabrenovic

Kernel

Available for: macOS Tahoe

Impact: A remote user may be able to cause unexpected system termination or corrupt kernel memory

Description: A race condition was addressed with improved locking.

CVE-2026-28982: Billy Jheng Bing Jhong and Pan Zhenpeng (@Peterpan0927) of STAR Labs SG Pte. Ltd., Adam Doupé of ASU SEFCOM

Kernel

Available for: macOS Tahoe

Impact: An app may be able to disclose kernel memory

Description: The issue was addressed with improved memory handling.

CVE-2026-64709: Pasquale Scola, Billy Jheng Bing Jhong and Pan Zhenpeng (@Peterpan0927) of STAR Labs SG Pte. Ltd.

Kernel

Available for: macOS Tahoe

Impact: A remote attacker may be able to bypass network filters

Description: An inconsistent user interface issue was addressed with improved state management.

CVE-2026-64735: Gor Aleksanyan

Kernel

Available for: macOS Tahoe

Impact: An app may be able to cause unexpected system termination

Description: An out-of-bounds write issue was addressed with improved bounds checking.

CVE-2026-43739: impost0r (ret2plt), Ruslan Dautov, Aleksandr Tarasikov, jay, Dhiyanesh Selvaraj (@redroot97), Vinay Kumar Rasala (Xplo8E) from Appknox, Lyutoon, DongJun Kim (smlijun) with UIUC, Hwiwon Lee (hwiwonl) with UIUC, Jongseong Kim (nevul37) with UIUC, Younggi Park (grill66) with UIUC, Peter Malone, an anonymous researcher, Hari Shanmugam (The Hxr1), Billy Jheng Bing Jhong and Pan Zhenpeng (@Peterpan0927) of STAR Labs SG Pte. Ltd., Marco Grassi, Dun, Daniele Castronovo, Michal Kosiorek

CVE-2026-43816: Josh Maine of Calif.io, Ruslan Dautov, 재영 정, @rootxran (Rao Ali Nawaz), Chanwit Muenprakoddee (ChemIndy), an anonymous researcher, Ye Zhang (@VAR10CK) of Baidu Security, Franco Belman at Blackwing Intelligence, Christian Figueroa, Johnny Franks (@zeroxjf), Ashmit Sharma & Atul RV, Peter Malone, Muhamad Syaiful, Muneeb Amin Bhat, Dhiyanesh Selvaraj (@redroot97), Ali Marzouq, Bountyy Oy - Mihalis Haatainen, Huy Nguyen (@34306) of Calif.io

Kernel

Available for: macOS Tahoe

Impact: An app may be able to cause unexpected system termination

Description: A use after free issue was addressed with improved memory management.

CVE-2026-43822: Eddy Tsalolikhin, Michal Kosiorek

CVE-2026-64729: Josh Maine of Calif.io, beist, Adam Doupé of ASU SEFCOM, Dun, Billy Jheng Bing Jhong and Pan Zhenpeng (@Peterpan0927) of STAR Labs SG Pte. Ltd., Johnny Franks (@zeroxjf)

CVE-2026-43814: Somair Ansar, Huy Nguyen (@34306) of Calif.io

CVE-2026-64700: Asjid Kalam (@odinshell)

CVE-2026-43799: Billy Jheng Bing Jhong and Pan Zhenpeng (@Peterpan0927) of STAR Labs SG Pte. Ltd.

Kernel

Available for: macOS Tahoe

Impact: Connecting to a malicious NFS server may lead to kernel memory corruption

Description: A buffer overflow was addressed with improved bounds checking.

CVE-2026-28931: Redon Gashi, Abhijeet Singh (linkedin.com/in/abhiunix/), Peter Malone, Omar Cerrito

Kernel

Available for: macOS Tahoe

Impact: An app may be able to cause unexpected system termination

Description: An out-of-bounds read was addressed with improved bounds checking.

CVE-2026-43817: Huy Nguyen (@34306) of Calif.io

CVE-2026-43809: Billy Jheng Bing Jhong and Pan Zhenpeng (@Peterpan0927) of STAR Labs SG Pte. Ltd.

CVE-2026-43757: Wang Yu, Billy Jheng Bing Jhong and Pan Zhenpeng (@Peterpan0927) of STAR Labs SG Pte. Ltd.

Kernel

Available for: macOS Tahoe

Impact: An app may be able to cause unexpected system termination

Description: An integer overflow was addressed with improved input validation.

CVE-2026-43769: Billy Jheng Bing Jhong and Pan Zhenpeng (@Peterpan0927) of STAR Labs SG Pte. Ltd.

Kernel

Available for: macOS Tahoe

Impact: An app may be able to cause unexpected system termination

Description: A type confusion issue was addressed with improved memory handling.

CVE-2026-64727: Ye Zhang (@VAR10CK) of Baidu Security

Kernel

Available for: macOS Tahoe

Impact: A remote user may be able to cause unexpected system termination or corrupt kernel memory

Description: The issue was addressed with improved memory handling.

CVE-2026-43810: Billy Jheng Bing Jhong and Pan Zhenpeng (@Peterpan0927) of STAR Labs SG Pte. Ltd.

Kernel

Available for: macOS Tahoe

Impact: An app may be able to cause unexpected system termination

Description: A memory initialization issue was addressed with improved memory handling.

CVE-2026-64775: Ryan Hileman via Xint Code (xint.io)

Kernel

Available for: macOS Tahoe

Impact: An app may be able to access sensitive user data

Description: A logic issue was addressed with improved checks.

CVE-2026-64723: Ji'an Zhou, Mingxuan Yang, Ye Zhang

Kernel

Available for: macOS Tahoe

Impact: An app may be able to cause unexpected system termination

Description: A race condition was addressed with improved state handling.

CVE-2026-64720: an anonymous researcher, Asjid Kalam (@odinshell), Jian Zhou and Ye Zhang

Kernel

Available for: macOS Tahoe

Impact: An app may be able to leak sensitive kernel state

Description: This issue was addressed with improved redaction of sensitive information.

CVE-2026-43754: Calif Research, Ernesto Martínez García

Kernel

Available for: macOS Tahoe

Impact: An app may be able to cause unexpected system termination or write kernel memory

Description: A use after free issue was addressed with improved memory management.

CVE-2026-64751: N.M.Praveen Nawarathne (@zblockrat)

Kernel

Available for: macOS Tahoe

Impact: An app may be able to access sensitive user data

Description: This issue was addressed through improved state management.

CVE-2026-64721: Lukas Gerlach

libarchive

Available for: macOS Tahoe

Impact: Processing a maliciously crafted file may result in disclosure of process memory

Description: This is a vulnerability in open source code and Apple Software is among the affected projects. The CVE-ID was assigned by a third party. Learn more about the issue and CVE-ID at cve.org.

CVE-2026-4424

libc

Available for: macOS Tahoe

Impact: A malicious app may be able to break out of its sandbox

Description: An integer overflow was addressed with improved input validation.

CVE-2026-28973: an anonymous researcher

Libnotify

Available for: macOS Tahoe

Impact: An attacker may be able to cause unexpected app termination

Description: An out-of-bounds write issue was addressed with improved bounds checking.

CVE-2026-64739: Feng Xue and XGPT of ThreatBook, Dun

LoginWindow

Available for: macOS Tahoe

Impact: An attacker with physical access to a locked device may be able to view sensitive user information

Description: An authorization issue was addressed with improved state management.

CVE-2026-43766: Amy (amys.website)

Managed Configuration

Available for: macOS Tahoe

Impact: An app may be able to access sensitive user data

Description: An authorization issue was addressed with improved state management.

CVE-2026-64743: Daniel Febrero

Maps

Available for: macOS Tahoe

Impact: A malicious app may be able to break out of its sandbox

Description: A permissions issue was addressed with additional restrictions.

CVE-2026-64738: Nathaniel Oh (@calysteon), Robert Mindo

mDNSResponder

Available for: macOS Tahoe

Impact: A local attacker may be able to cause a denial of service

Description: A denial of service issue was addressed by removing the vulnerable code.

CVE-2026-43806: He Wei (ギカク), 章鱼哥 (@aipy) of aipyaipy.com, Jex Amro, Cem Onat Karagun

mDNSResponder

Available for: macOS Tahoe

Impact: An attacker on the local network may be able to cause a denial-of-service

Description: The issue was addressed with improved memory handling.

CVE-2026-64724: Daisuke Hatakeyama (@SYZD Research)

MediaRemote

Available for: macOS Tahoe

Impact: An app may be able to gain root privileges

Description: A path handling issue was addressed with improved validation.

CVE-2026-43723: Richard Zana, Andreas Jaegersberger & Ro Achterberg of Nosebeard Labs

Metal

Available for: macOS Tahoe

Impact: A malicious app may be able to corrupt memory of a system process

Description: The issue was addressed with improved memory handling.

CVE-2026-28911: yk lin of @pixiepointsec

Model I/O

Available for: macOS Tahoe

Impact: Processing a maliciously crafted image may corrupt process memory

Description: The issue was addressed with improved memory handling.

CVE-2026-43733: Michael DePlante (@izobashi) of TrendAI Zero Day Initiative

CVE-2026-43729: Michael DePlante (@izobashi) of TrendAI Zero Day Initiative

Model I/O

Available for: macOS Tahoe

Impact: A remote attacker may be able to cause unexpected application termination or heap corruption

Description: An out-of-bounds write issue was addressed with improved input validation.

CVE-2026-64772: stratan (@5tratan), wh0am1i

Model I/O

Available for: macOS Tahoe

Impact: A remote attacker may be able to cause unexpected application termination or heap corruption

Description: A buffer overflow was addressed with improved bounds checking.

CVE-2026-64771: wh0am1i

Model I/O

Available for: macOS Tahoe

Impact: Processing a 3D model may result in disclosure of process memory

Description: A buffer overflow issue was addressed with improved memory handling.

CVE-2026-64722: wh0am1i

Model I/O

Available for: macOS Tahoe

Impact: A remote attacker may be able to cause unexpected application termination or heap corruption

Description: An integer overflow was addressed with improved input validation.

CVE-2026-64774: stratan (@5tratan)

Model I/O

Available for: macOS Tahoe

Impact: A remote attacker may be able to cause unexpected application termination or heap corruption

Description: An out-of-bounds write issue was addressed with improved bounds checking.

CVE-2026-64770: stratan (@5tratan)

CVE-2026-64769: stratan (@5tratan)

Model I/O

Available for: macOS Tahoe

Impact: A remote attacker may cause an unexpected app termination

Description: An out-of-bounds read issue was addressed with improved input validation.

CVE-2026-64768: stratan (@5tratan)

Net-SNMP

Available for: macOS Tahoe

Impact: An app may be able to cause a denial-of-service

Description: A stack overflow was addressed with improved input validation.

CVE-2026-43771: Robert Tran

NetFSFramework

Available for: macOS Tahoe

Impact: An app may be able to break out of its sandbox

Description: A path traversal issue was addressed with improved input validation.

CVE-2026-43772: Mickey Jin (@patch1t)

NSColorPanel

Available for: macOS Tahoe

Impact: An app may be able to leak sensitive user information

Description: This issue was addressed with additional entitlement checks.

CVE-2026-64711: Koh M. Nakagawa (@tsunek0h) of FFRI Security, Inc.

PackageKit

Available for: macOS Tahoe

Impact: A user may be able to elevate privileges

Description: A logic issue was addressed with improved restrictions.

CVE-2026-28912: Matej Moravec (@MacejkoMoravec)

PackageKit

Available for: macOS Tahoe

Impact: An app may be able to modify protected parts of the file system

Description: This issue was addressed with improved handling of symlinks.

CVE-2026-43765: Mickey Jin (@patch1t)

Printing

Available for: macOS Tahoe

Impact: A malicious app may be able to break out of its sandbox

Description: A path handling issue was addressed with improved validation.

CVE-2026-64731: Sindre Sorhus, Richard Zana

Pro Res

Available for: macOS Tahoe

Impact: An app may be able to cause unexpected system termination

Description: A use after free issue was addressed with improved memory management.

CVE-2026-43812: Francisco Knabe

quarantine

Available for: macOS Tahoe

Impact: An app may be able to cause unexpected system termination or write kernel memory

Description: The issue was addressed with improved memory handling.

CVE-2026-43694: Hcamael and 章鱼哥@aipy (aipyaipy.com), JC Alvarado of Stripe, Jacob Hazak

Remote Management

Available for: macOS Tahoe

Impact: A malicious app may be able to gain root privileges

Description: A permissions issue was addressed with additional restrictions.

CVE-2026-39874: @pixiepointsec

Safari

Available for: macOS Tahoe

Impact: An app may be able to access sensitive user data

Description: An authorization issue was addressed with improved state management.

CVE-2026-43792: Ilya Andr (andrd3v)

SceneKit

Available for: macOS Tahoe

Impact: Processing a maliciously crafted file may lead to unexpected app termination or arbitrary code execution

Description: An integer overflow was addressed with improved input validation.

CVE-2026-64766: stratan (@5tratan)

CVE-2026-64765: stratan (@5tratan)

SceneKit

Impact: Processing a maliciously crafted file may lead to unexpected app termination or arbitrary code execution

Description: An out-of-bounds write issue was addressed with improved bounds checking.

CVE-2026-64764: stratan (@5tratan)

SceneKit

Available for: macOS Tahoe

Impact: Processing a maliciously crafted file may lead to unexpected app termination or arbitrary code execution

Description: An out-of-bounds write issue was addressed by removing the vulnerable code.

CVE-2026-64763: stratan (@5tratan)

Screen Sharing Server

Available for: macOS Tahoe

Impact: An app may be able to intercept network connections intended for another process

Description: A logic issue was addressed with improved restrictions.

CVE-2026-43779: Dave G., Asaf Cohen

Screen Sharing Server

Available for: macOS Tahoe

Impact: A remote attacker may be able to cause a denial of service

Description: This issue was addressed with improved input validation.

CVE-2026-43777: Junming C.(Chapoly1305)

Screen Sharing Server

Available for: macOS Tahoe

Impact: An app may be able to access user-sensitive data

Description: An access issue was addressed with improved access restrictions.

CVE-2026-43760: Alfredo Pesoli (@__rev) of Bynar.io, wdszzml and Atuin Automated Vulnerability Discovery Engine

Security

Available for: macOS Tahoe

Impact: An attacker may be able to modify the state of the Keychain

Description: This issue was addressed through improved state management.

CVE-2026-43728: Bob Gendler of the National Institute of Standards and Technology

SecurityAgent

Available for: macOS Tahoe

Impact: An app may be able to gain root privileges

Description: A race condition was addressed with improved state management.

CVE-2026-43755: Mickey Jin (@patch1t)

Siri

Available for: macOS Tahoe

Impact: A person with physical access to a locked device may be able to access contacts and photos

Description: This issue was addressed with additional restrictions on the lock screen.

CVE-2026-64745: Vivek Dhar, ASI (RM) in Border Security Force, FTR HQ BSF Kashmir

Siri

Available for: macOS Tahoe

Impact: An app may be able to access sensitive user data

Description: An information disclosure issue was addressed by removing the vulnerable code.

CVE-2026-43800: Stanislav Jelezoglo

SMB

Available for: macOS Tahoe

Impact: Connecting to a malicious SMB server may lead to unexpected system termination

Description: The issue was addressed with improved memory handling.

CVE-2026-39873: Peter Malone

SMB

Available for: macOS Tahoe

Impact: A remote user may be able to cause unexpected system termination or corrupt kernel memory

Description: The issue was addressed with improved memory handling.

CVE-2026-64696: Feng Xue and XGPT of ThreatBook, Peter Malone

SMB

Available for: macOS Tahoe

Impact: An app may be able to cause unexpected system termination

Description: A type confusion issue was addressed with improved memory handling.

CVE-2026-64704: Claudio Bozzato and Francesco Benvenuto of Cisco Talos, Aswin Kumar Gokulakannan, Kitten Food, Peter Malone, Calif.io in collaboration with Claude and Anthropic Research

Spotlight

Available for: macOS Tahoe

Impact: An app may be able to access sensitive user data

Description: An out-of-bounds read was addressed with improved bounds checking.

CVE-2026-43774: Csaba Fitzl (@theevilbit) of Iru

StorageKit

Available for: macOS Tahoe

Impact: An app may be able to access sensitive user data

Description: A race condition was addressed with additional validation.

CVE-2026-43770: Tien-Chih Lin of CyCraft Technology

udf

Available for: macOS Tahoe

Impact: An app may be able to cause unexpected system termination

Description: The issue was addressed with improved memory handling.

CVE-2026-43768: Hyunwoo Kim (@v4bel)

WebDAV

Available for: macOS Tahoe

Impact: An app may be able to cause a denial-of-service

Description: A use after free issue was addressed with improved memory management.

CVE-2026-64703: Bruce Dang of Calif.io in collaboration with Claude and Anthropic Research

WebDAV

Available for: macOS Tahoe

Impact: An app may be able to disclose kernel memory

Description: A memory initialization issue was addressed with improved memory handling.

CVE-2026-64699: Bruce Dang of Calif.io

WebKit

Available for: macOS Tahoe

Impact: Processing maliciously crafted web content may result in the disclosure of process memory

Description: The issue was addressed with improved memory handling.

WebKit Bugzilla: 308046

CVE-2026-43740: Arni Hardarson, Nathaniel Oh (@calysteon)

WebKit

Available for: macOS Tahoe

Impact: Websites may know if the user has visited a given link

Description: This issue was addressed with improved checks.

WebKit Bugzilla: 316827

CVE-2026-64713: Kwak Kiyong, Song Nuri

WebKit

Available for: macOS Tahoe

Impact: Visiting a website that frames malicious content may lead to UI spoofing

Description: The issue was addressed with improved UI.

WebKit Bugzilla: 311660

CVE-2026-64730: Kagami Rosylight of Mozilla

WebKit

Available for: macOS Tahoe

Impact: Maliciously crafted web content may violate iframe sandboxing policy

Description: A permissions issue was addressed with improved validation.

WebKit Bugzilla: 313220

CVE-2026-64728: an anonymous researcher

WebKit

Available for: macOS Tahoe

Impact: Processing maliciously crafted web content may lead to an unexpected Safari crash

Description: A use-after-free issue was addressed with improved memory management.

WebKit Bugzilla: 313521

CVE-2026-64783: 杉山 壮太, lattice, Behzad Najjarpour Jabbari (@_G4ru_), Junyeong Lee, Mooth.ai, OGINOME Tomohito, Using GLM From Z.AI, Gia Bui (@yabeow) from Calif.io

WebKit

Available for: macOS Tahoe

Impact: Processing maliciously crafted web content may lead to an unexpected Safari crash

Description: A memory corruption issue was addressed with improved state management.

WebKit Bugzilla: 315082

CVE-2026-64757: Milad Nasr and Nicholas Carlini with Claude, Anthropic

WebKit

Available for: macOS Tahoe

Impact: Visiting a website may lead to an app denial-of-service

Description: This issue was addressed through improved state management.

WebKit Bugzilla: 316816

CVE-2026-43804: Heiko Kiesel of SEEMOO, TU Darmstadt

WebKit

Available for: macOS Tahoe

Impact: An app may be able to read files outside of its sandbox

Description: An access issue was addressed with improved access restrictions.

WebKit Bugzilla: 314867

CVE-2026-43821: Brian Carpenter

WebKit Canvas

Available for: macOS Tahoe

Impact: Processing maliciously crafted web content may lead to an unexpected Safari crash

Description: A use-after-free issue was addressed with improved memory management.

WebKit Bugzilla: 313935

CVE-2026-64718: OGINOME Tomohito, an anonymous researcher

WebRTC

Available for: macOS Tahoe

Impact: Processing maliciously crafted web content may lead to an unexpected Safari crash

Description: An out-of-bounds access issue was addressed with improved bounds checking.

WebKit Bugzilla: 319404

CVE-2026-64719: Shaheen Fazim

Wi-Fi

Available for: macOS Tahoe

Impact: An attacker in physical proximity may be able to corrupt process memory

Description: The issue was addressed with improved memory handling.

CVE-2026-64726: Mathis Mansière, Peter Malone

Wi‑Fi

Available for: macOS Tahoe

Impact: An app may be able to execute arbitrary code out of its sandbox or with certain elevated privileges

Description: A buffer overflow was addressed with improved bounds checking.

CVE-2026-43750: an anonymous researcher

xar

Available for: macOS Tahoe

Impact: An app may be able to cause a denial of service

Description: A logic issue existed resulting in memory corruption. This was addressed with improved state management.

CVE-2026-28932: Mathis Mansière

Additional recognition

Audio

We would like to acknowledge Niels Hofmans for their assistance.

copyfile

We would like to acknowledge Keisuke Hosoda for their assistance.

CoreAnalytics

We would like to acknowledge Alan Banderas (@creeper4004) for their assistance.

CoreMedia

We would like to acknowledge yaohway for their assistance.

dcerpc

We would like to acknowledge Surya Narayan Kushwaha for their assistance.

Disk Images

We would like to acknowledge Jordy Zomer (@pwningsystems), Phillip Groves, Richard Zana for their assistance.

DriverKit

We would like to acknowledge Niels Hofmans for their assistance.

Foundation

We would like to acknowledge Nick Cook, Sentry Flag for their assistance.

Heimdal

We would like to acknowledge Surya Narayan Kushwaha for their assistance.

IOStorageFamily

We would like to acknowledge Aswin Kumar Gokula Kannan, Surya Narayan Kushwaha, Yatin Taneja for their assistance.

Kernel

We would like to acknowledge Billy Jheng Bing Jhong and Pan Zhenpeng (@Peterpan0927) of STAR Labs SG Pte. Ltd., Chris Betz, Hiroki Imai (LAC Co., Ltd.), James Duffy ( @0x4A616D657344 ), Mathis Mansière, Tristan Rousseau, Vladislav Shevchenko (Positive Technologies), Yeojin Kim, YingMuo (@YingMuo) of DEVCORE Research Team for their assistance.

PluginKit

We would like to acknowledge Asaf Cohen, Ashish Kunwar for their assistance.

Printing UIKit

We would like to acknowledge Jacolon Walker ( @call_eax ) for their assistance.

RemoteServiceDiscovery

We would like to acknowledge Maliq Barnard, Tommy DeVoss from Braze Security Team (@thedawgyg), Willard Jansen for their assistance.

ReplayKit

We would like to acknowledge an anonymous researcher for their assistance.

Safari Downloads

We would like to acknowledge Alfaz Hossain for their assistance.

Screen Sharing Server

We would like to acknowledge Ayaan Ahmad, XlabAI Team of Tencent Xuanwu Lab, Atuin Automated Vulnerability Discovery Engine, Guannan Wang, Zhanpeng Liu, Jiashuo Liang, Guancheng Li for their assistance.

Security

We would like to acknowledge John Lussier, Oleh Konko of 1seal (1seal.org), alick for their assistance.

Spotlight

We would like to acknowledge Ilya Andr (andrd3v) and nkhmelni for their assistance.

System Settings

We would like to acknowledge Masaki Moriguchi, Omar Cerrito for their assistance.

Time Machine

We would like to acknowledge Andreas Jaegersberger & Ro Achterberg of Nosebeard Labs for their assistance.

WebKit

We would like to acknowledge Jaya Surya Kommireddy, Jaya surya Kommireddy, Lukas Knittel (@kunte_ctf) of Ruhr-University Bochum, Nikos Fanourakis of Technical University of Crete, Sotiris Ioannidis of Technical University of Crete, Panagiotis Ilia of Cyprus University of Technology, and Kostas Drakonakis of Technical University of Crete, Tony Gorez (@tonygo_) for Reverse Society, Vitaly Simonovich, Youngjoon Kim of Team-Atlanta & sslab at Georgia Tech, s3zer0 for their assistance.

WebKit Canvas

We would like to acknowledge Codex Security - Khai Tran, Daisuke Hatakeyama and Ryohei Ueki (@SYZD Research), David Bors at Snyk Security Labs, Giovanni Vignone and Robert van Eijk of Octane Security (octane.security), Kwak Kiyong, Song nuri, Luat Nguyen (CyberJutsu Academy), Tom Van Goethem, an anonymous researcher, dr3dd for their assistance.

WebKit Storage

We would like to acknowledge Gurpreet Shergill, Luke Francis, Milad Nasr and Nicholas Carlini with Claude, Anthropic, Oleh Konko of 1seal (1seal.org), Vitaly Simonovich for their assistance.

xar

We would like to acknowledge Anurag Bohra of Microsoft, Cem Onat Karagun, Feng Xue and XGPT of ThreatBook, Kubilay Berk ALKAN, s3zer0 for their assistance.

XprotectFramework

We would like to acknowledge Ferdous Saljooki (@malwarezoo) of Jamf for their assistance.

Information about products not manufactured by Apple, or independent websites not controlled or tested by Apple, is provided without recommendation or endorsement. Apple assumes no responsibility with regard to the selection, performance, or use of third-party websites or products. Apple makes no representations regarding third-party website accuracy or reliability. Contact the vendor for additional information.

The Daily Front Page 15 of 25
Tuesday, July 28, 2026 The Daily Front No. #260728 — Work, Trust, and HR
article

Netflix employee fired for sharing personal details in retreat trust exercise

by softwaredoug·▲ 425 points·483 comments·nypost.com ↗
“A ‘trust exercise’ at a work retreat.”

A Netflix executive was fired from his $1.1 million a year job after revealing during a “trust exercise” at a work retreat that he had taken medically prescribed ketamine, a lawsuit has claimed.

Kevin Baillie, who was vice president and head of creative at Eyeline Studios, is suing the company after it launched an investigation into his comments that ultimately ended in his firing, the papers say. 

Baillie, who’s been on the visual effects team for “Pirates of the Caribbean” and the Harry Potter franchise, says he took the drug under medical supervision in October and November of 2022 at a Santa Barbara clinic.

Kevin Baillie poses for a selfie with a red Ferrari race car.

Kevin Baillie, ex-vice president and head of creative at Eyeline Studios, a division of Netflix. Instagram/fxnerd

Joe Letteri and Kevin Baillie posing with their awards at the 24th Annual Satellite Awards.

Kevin Baillie (right) at the 24th Satellite Awards in Beverly Hills. I Hasegawa/HNW-Photo/Plux / Shutterstock

He sought the treatment for clinical depression after the death of his mother, according to the suit.

During what’s called a “Vulnerability-Trust exercise” at a January 2026 retreat at the exclusive Sendero Ranch, a Northern California property owned by Netflix, Baillie shared with his colleagues that he had undergone the treatment, the suit says.

Baillie claims he explained the reason why he had taken the drug but was investigated by Netflix. On March 18, 2026 a company investigator brought the incident up, “in a manner suggesting suspicion of recreational drug use,” the suit reads.

A wooden sign for "Sendero Ranch" hangs above a green metal gate, with a road extending into a desert landscape with mountains in the background.

The retreat took place at the exclusive Sendero Ranch, a Northern California property owned by Netflix. Instagram/cecimeseeds

A man in a "Star Wars Episode One" jacket looks up at a movie marquee advertising "Star Wars Episode I".

Baillie was fired in April, with Netflix's attorney confirming, “The ketamine therapy issue has factored into the termination.” Instagram/fxnerd

The executive was fired in April with the company’s attorney confirming “the ketamine therapy issue has factored into the termination,” the papers say, and go on to suggest Baille was denied up to a year of severance pay.

Baillie says in the lawsuit the “scope” of the investigation related to alleged profanity and drinking. He had been warned during his performance review that he should “drop one or two less f-bombs but don’t stop entirely.”

It goes on to say that at the same retreat Baillie drank a Guinness standing on his head, after sharing that he had learned the trick from his former father in law during a conversation inspired by the trust session.

“His colleague immediately asked for a demonstration, rather than withhold the openness that the session had encouraged, he performed the trick,” the papers say.

Ribbon cutting at Netflix Eyeline Studios opening in Hyderabad with Jeff Shapiro, second person from the left.

Baille also suggests an alcohol-fueled company environment bolstered by Eyeline Studios CEO Jeff Shapiro. Netflix

Baille also paints a picture of an allegedly alcohol-fueled company environment encouraged by the Eyeline Studios CEO Jeff Shapiro.

The lawsuit alleged “alcohol consumption was company-sponsored, leadership-modeled and condoned”, with multiple examples of how Shapiro “set the cultural tone concerning alcohol at the executive level”.

The documents claimed Shapiro on one occasion purchased beer at a corner store and brought it to a company car ride to the Visual Effects Society Awards for staff to share.

Baillie claims he also witnessed the CEO consume alcohol with Netflix and Eyeline employees at his own welcome dinner in September 2024, the Netflix Annual Business Review events in March 2025 and even a Lakers game attended by Netflix execs including Shapiro’s direct supervisor in February 2026.

All up, Baillie’s attorneys provided over half a dozen examples of the CEO being present at a work event with a drink in his hand, according to court papers.

In addition to hosting multiple parties, Shapiro also had a personal bar in his office “from which he served alcohol (to Baille) including after a successful meeting with Netflix’s CEO Ted Sarandos,” according to the papers.

Baillie is asking for a jury trial, compensatory damages, lost wages, damages for emotional distress, and punitive damages.

Netflix and Eyeline were contacted for comment.

The Daily Front Page 16 of 25
Tuesday, July 28, 2026 The Daily Front No. #260728 — Show HN: Yap Dictation for macOS
show hn

Show HN: Yap – OSS on-device voice dictation for macOS with no model to download

by pancomplex·▲ 100 points·41 comments·github.com ↗
“Blazing‑fast voice dictation for macOS that works anywhere you can type.”

Blazing-fast voice dictation for macOS that works anywhere you can type.

License: MIT macOS 26+ Swift 6

Press a shortcut, talk, press it again. Your words land in whatever text field you were using. No account, no API key, no audio leaving your machine.

Built by Frigade.

yap-demo.mp4


What it does

Yap lives in your menu bar and waits for a shortcut. Trigger it and a small window appears near the bottom of the screen with a live waveform and a running preview of what you have said so far. Press the shortcut again and the text gets pasted into the app you were already working in. Every transcript is saved locally, so you can go back and copy something again later.

The default shortcut is ⌘⇧D. You can rebind it, or set a single modifier key instead. Tapping right shift on its own works nicely if you have a spare thumb.

Install

The quickest way is from the website: frigade.com/yap. Or use Homebrew.

Homebrew

brew install --cask frigadehq/tap/yap

To update later:

brew upgrade --cask yap

Direct download

Download Yap from the website, or grab the .dmg straight from Releases. Drag it into your Applications folder.

Yap runs on macOS 26 (Tahoe) on an Apple Silicon Mac. Intel Macs are not supported. Released builds are signed and notarized, so macOS opens them without complaint.

Why we built this

There's no shortage of voice to text tools for macOS, and some of the open source ones are genuinely good. Most of them still run into some mix of the same problems:

  • Some cost money, which is a lot to ask for something that's now built into your OS for free.
  • Most make you download a heavy model. Whisper weights run to hundreds of megabytes, sit in your RAM, and only feel fast on a recent, high end Mac. On slower machines a single paragraph can take thirty seconds to a minute to come back.
  • A lot are Electron or web-stack apps, so a whole browser engine idles in your memory just to run a menu bar icon and a settings window.
  • Plenty are bloated with settings and modes you will never open.
  • Some are closed source, so you are trusting that your audio and transcripts stay on your machine, with no way to check.

What changed recently is macOS 26. It added two APIs, SpeechAnalyzer and SpeechTranscriber, that do on-device streaming speech to text using models the OS ships and manages. The app carries no model of its own, loads nothing into memory before the first word, and needs no API key or per minute cost. Text comes back as you talk.

Is Apple's model actually any good? Better than the thing it replaces, as it turns out. A recent benchmark put it at 2.12% word error rate on clean audio and 4.56% on noisy, against 3.74% and 7.95% for Whisper Small, and it ran about three times faster, across 5,559 LibriSpeech clips.

So Yap ships no model at all. It's roughly three thousand lines of native Swift in a 4 MB app, all of it open, with no browser engine anywhere in sight. It idles around 60 MB of memory and never touches the network. We use it all day at Frigade, mostly for prompting coding agents, writing emails, and firing off Slack messages, basically anything that's quicker to say than type.

Features

  • On-device transcription through Apple's Speech framework
  • Global shortcut, fully rebindable, with optional single-modifier triggers like right shift
  • Pastes straight into the focused field of whatever app you were in
  • Local transcript history with search, copy, and delete
  • Live waveform and partial transcript while you speak
  • Press escape twice to discard a dictation in progress
  • Follows your system default microphone, including when it changes mid-session
  • Optional launch at login, off by default
  • No account, no network calls, no telemetry

Requirements

macOS 26 (Tahoe) or later, on an Apple Silicon Mac. Yap depends on the speech models Apple ships with macOS 26, and SpeechAnalyzer runs on device only on Apple Silicon. We dropped Intel support on purpose: the only way to transcribe on those machines was an API that sends audio to Apple, and that breaks the one promise Yap makes. Building from source needs Xcode 26.

Permissions

On first launch Yap asks for four things, and explains each one:

Permission Why Microphone To hear you Speech Recognition To transcribe on device Accessibility To see which app you are typing into Automation To paste the result into it

Accessibility has to be switched on by hand in System Settings. macOS requires that of any app that types on your behalf, and there is no way to grant it programmatically.

How it works

Audio comes off the default input through AVAudioEngine and gets converted to whatever format the analyzer asks for. Capture starts before the speech stack finishes initializing, and buffers recorded in that window are held and flushed once the transcriber attaches, so the first word of a sentence is never clipped.

Transcription runs through SpeechAnalyzer with volatile results turned on, which is what gives you the live preview. There is no other path. Older APIs like SFSpeechRecognizer can fall back to Apple's servers when a locale has no on-device model, so Yap does not use them. If SpeechAnalyzer can't handle your language on device, dictation stops rather than sending your audio anywhere.

Insertion is the awkward part. Yap writes the text to the clipboard, drives ⌘V through System Events, then restores your previous clipboard contents. It waits before restoring, because Chromium-based apps read the pasteboard asynchronously and more than once, and restoring too early hands the renderer stale data. That single detail is the difference between working everywhere and working only in native apps.

State lives in one place. RecordingCoordinator is a small state machine whose dependencies are all protocols, so the logic is covered by unit tests without needing a microphone.

FAQ

Why not just use the built-in macOS Dictation?

Apple's built-in Dictation still does part of the work online, and even discloses that it "sends information like your voice input, contacts, and location to Apple." Yap talks to the on-device SpeechAnalyzer API directly and makes zero network calls, and because it is open source you can check that yourself. It also gives you press-to-stop with a live preview, so you can cancel a bad take before it inserts instead of watching it type as you talk (very useful in terminals / coding agents).

Does any of my audio or text leave my Mac?

No. There is no network code in the app at all. Transcription runs entirely on device, and your history is stored locally in SwiftData. Nothing is uploaded, and there is no account or telemetry.

Why is it Apple Silicon only?

SpeechAnalyzer runs on device only on Apple Silicon. The old way to cover Intel Macs was SFSpeechRecognizer, which can send your audio to Apple when a locale has no on-device model, so we removed it rather than ship something that quietly breaks the promise. If you are on an Intel Mac, stay on 0.1.4 or an earlier version (note: that Intel version will send API calls to Apple).

Roadmap

There is not much of one, and that is on purpose. Yap does a single thing, and keeping it small enough that you never have to think about it is the main design principle. Most feature ideas make an app like this worse.

We are not precious about it though. If something is missing that you would use every day, open an issue or send a pull request and we will give it a fair hearing. A language picker is the most likely next addition, since Yap follows your system locale today.

Development

Build and install from source in one line. It clones, builds, drops Yap into /Applications, and launches it:

git clone https://github.com/FrigadeHQ/yap.git && cd yap && ./install.sh

The script installs XcodeGen if you do not have it, quits any running copy, and replaces it with the new build. Re-run ./install.sh any time to rebuild after a change.

If you would rather work in Xcode:

xcodegen generate                    # writes Yap.xcodeproj from project.yml
open Yap.xcodeproj                   # then run with ⌘R

The Xcode project is generated rather than committed, so configuration changes stay readable in a diff.

Run the tests with:

xcodebuild -project Yap.xcodeproj -scheme Yap -destination 'platform=macOS' test

One thing to know about local builds. They are signed ad-hoc, which gives them no stable identity, so macOS treats every rebuild as a brand new application and forgets the permissions you granted the previous one. The symptom is confusing: the checkbox still looks switched on in System Settings, but pasting quietly stops working. There is a "Reset and re-grant" button in Settings for exactly this. Released builds are signed properly and do not have the problem.

The app icon is generated too, if you want to change it:

swift Tools/GenerateIcon.swift /tmp/Yap.iconset
iconutil -c icns /tmp/Yap.iconset -o Sources/Yap.icns

Contributing

Issues and pull requests are welcome. If you are fixing a paste failure in a specific app, please say which app and which macOS version, since that class of bug is almost always app-specific.

License

MIT. See LICENSE.


Frigade

Built by Frigade. Frigade Engage makes it easy to build in-product onboarding (checklists, product tours, and more), and Frigade Assistant is an in-app AI assistant that learns your product by using it, then onboards, supports, and activates your customers.

Website · Docs · Demo · GitHub

About

Free, open source voice dictation for macOS. On-device transcription with Apple's Speech framework. No cloud, no API keys, no account.

The Daily Front Page 17 of 25
Tuesday, July 28, 2026 The Daily Front No. #260728 — Zig‑Packaged C/C++
repository

C/C++ projects packaged for Zig

by jcbhmr·▲ 76 points·41 comments·github.com ↗

All Your Codebase

...are belong to us, but we'd be delighted to give them back!

What is this organization?

We package C/C++ projects for the Zig build system so that you can reliably compile (and cross-compile!) them with ease.

This both provides convenience for users of the Zig compiler toolchain and also showcases to C/C++ project maintainers what a build.zig file for their project looks like.

I maintain one of the projects you packaged, what value does your work provide?

As a general answer, we add a dependency on Zig to your project but in exchange we remove a dependency on:

  • Make / GNUMake / CMake / autoconf / bash scripts / batch scripts / powershell scripts: Zig is a complete build system that works on all supported platforms and can do everything those other tools do.
  • Clang: Zig is a full compiler toolchain and happens to also bundle all of clang.
  • The system package manager: Zig is also a package manager and can download and build dependencies packaged for it, if you want it to.
  • Docker / CI matrix jobs: Zig can cross-compile C/C++/Zig code, making a release is as simple as running zig build release.

More in general Zig removes all dependency on system-wide settings, while still leaving you the ability to opt-in when you need to.

I maintain one of the projects you packaged and I like your work, how can I upstream it?

If you're the maintainer of a project packaged by us and decide that you want to upgrade your build pipeline, then you are free to upstream everything you need from our repos.

If you decide to do so, please let us know by opening an Issue so that we can archive our repo and point people to your upstream. Feel also free to use our repos' Issues section to ask questions about how to integrate everything correctly in your project (say, maybe because we didn't implement a secondary build step for example).

One last thing to note: for us to be able to archive our repository, your integration of our build.zig must not add more system dependencies than our version. So, for example, if our packaged version is able to depend on zstd via allyourcodebase/zstd, then we kindly ask that you either keep depending on it (until its build.zig gets upstreamed) or take advantage of System Library Integration to give the user the choice.

That said, you're obviously welcome to upstream any build code as you see fit even if you don't plan to keep depending on other packages via the Zig build system. We'll still be happy to help you in this case, of course, but we'll also keep maintaining our downstream fork.

How does a C/C++ project packaged for the Zig build system by you look like?

We use two main strategies:

  1. Add the upstream project as a dependency (in build.zig.zon) and define the corresponding build.zig script in our repo.
  2. Fork the upstream project (optionally remove other -- now useless :^) -- build scripts), add Zig build scripts to it and apply any necessary patch to the original project. This last part is usually not necessary but some build steps might benefit by making some config scripts and such more amenable to be used by the Zig build system. One example of that would be to have scripts accept an output path as an argument instead of hardcoding where their output goes (this is very useful to integrate properly with Zig build cache).

If you're a maintainer of the upstream project, (1) shows clearly that you will only need two files (build.zig, build.zig.zon) but it will be up to you to clean up all other build scripts and possibly improve our Zig build script as described above, while (2) will require you to be a bit more careful when upstreaming the work but everything will have been done for you more thoroughly (although you still have the option to just take the Zig build script files and manually review how to integrate them in your upstream project).

I'm a Zig user and I want to contribute a repository, how can I do it?

Ping kristoff and ask to be added to the organization.

Here are some ground rules to be able to contribute a repo:

  1. Your repo must have a license for your code (the new build.zig file) and it must be at least as permissive as the original project's license (to make things simple you can just use MIT and will never be wrong).

  2. Your repo must package the original C/C++ project without adding extra Zig-specific stuff like bindings, for example.
    You are welcome to put bindings in a separate repo owned directly by you.

  3. You must target the latest tagged version of Zig.

  4. When porting the original project, you should use the first strategy (build.zig + pristine tarball) and only turn to the second one (forking the full project) if:

    • You need to patch the original source code for it to build correctly

    • You are going to clean up all other build scripts and improve how intermediate steps of the build process work

      • An example of this last point would be changing project-specific build tools (eg asset processing tools, like image optimizers) that hardcode an output path (usually cwd) to instead accept an output argument in order to make them better citizens of the Zig build (eco)system.
  5. You must add a CI job that guarantees zig build succeeds, for example, you can copy the script from allyourcodebase/AFLplusplus.

  6. You must have interest in doing occasional maintainership work to update your build script when a new version of the upstream project is released.

Once you're done with the checklist above, please give your repo an appropriate set of tags for ease of discoverability (eg zig, zig-package).

The Daily Front Page 18 of 25
Tuesday, July 28, 2026 The Daily Front No. #260728 — Inside Kimi: Architecture and Access
article

Kimi K3 Architecture Overview and Notes

by ModelForge·▲ 393 points·70 comments·sebastianraschka.com ↗

The Kimi K3 architecture figure for yesterday’s big open-weight model release, along with some observations and thoughts.

  1. Yes, it looks relatively complicated, but it’s essentially a scaled-up production version of their Kimi Linear model they released last year (scaled up from 48B -> 2.8T; K3 is by far the biggest open-weight model right now)
  2. The one new component compared to Kimi Linear is the LatentMoE. I omitted it in the figure below since it’s already very crowded, but that’s essentially the same LatentMoE as in Nemotron 3 Ultra (you can find it in my LLM Architecture Gallery if you are curious). The idea here is to compress (down-project) large linear layers similar to multi-head latent attention.
  3. Kimi K3’s overall trend (similar to Nemotron 3, DeepSeek V4, and others) is also towards better inference efficiency. That is, there are many components that replace existing components with efficiency-tweaked versions. I.e., MoE -> LatentMoE, regular attention -> multi-head latent attention and Kimi Delta Attention. (I also have short tutorials and write-ups in my gallery if you are curious about additional details).
  4. The one component change that is not an efficiency tweak is attention residuals. Like DeepSeek V4 improved the residual path with mHC (manifold-constrained Hyper-Connections), attention residuals are a way to improve the residual path, but it works a bit differently. I.e., mHC made the residual path wider. Attention residuals (also already part of Kimi Linear) connect the residuals across layers; the connection itself uses an attention score for an important/contribution weight. According to the report, it improves the validation loss and downstream performance (a bit) consistently and adds about 4% in training cost and 2% in inference cost.
  5. Interestingly, Kimi K3 got rid of all RoPE layers and uses NoPE (No Positional Embeddings) everywhere instead. (Again, this is inherited from Kimi Linear). In other architectures, the recent trend was towards RoPE in local attention layers (like sliding window attention) and NoPE in the global layers. There were a few architectures that only used NoPE everywhere, but this is the first frontier-level one as far as I know.
  6. Kimi K3 now also has native multimodal support, which is great!

There are several other interesting training tidbits in the technical report, but that’s it from the architecture front so far. A really great release overall.

Composite Kimi K3 architecture diagram with Kimi Delta Attention, gated multi-head latent attention, Attention Residuals, LatentMoE, and benchmark comparisons

Figure 1. Kimi K3 architecture and release-time benchmark comparisons. See K3 in the architecture gallery for more details.

Source: website version of my Substack note.

The Daily Front Page 19 of 25
Tuesday, July 28, 2026 The Daily Front No. #260728 — Inside Kimi: Architecture and Access
article

Kimi Linear: An Expressive, Efficient Attention Architecture (2025)

by ronfriedhaber·▲ 296 points·125 comments·arxiv.org ↗

Kimi Linear: An Expressive, Efficient Attention Architecture

Abstract:We introduce Kimi Linear, a hybrid linear attention architecture that, for the first time, outperforms full attention under fair comparisons across various scenarios -- including short-context, long-context, and reinforcement learning (RL) scaling regimes. At its core lies Kimi Delta Attention (KDA), an expressive linear attention module that extends Gated DeltaNet with a finer-grained gating mechanism, enabling more effective use of limited finite-state RNN memory. Our bespoke chunkwise algorithm achieves high hardware efficiency through a specialized variant of the Diagonal-Plus-Low-Rank (DPLR) transition matrices, which substantially reduces computation compared to the general DPLR formulation while remaining more consistent with the classical delta rule.

We pretrain a Kimi Linear model with 3B activated parameters and 48B total parameters, based on a layerwise hybrid of KDA and Multi-Head Latent Attention (MLA). Our experiments show that with an identical training recipe, Kimi Linear outperforms full MLA with a sizeable margin across all evaluated tasks, while reducing KV cache usage by up to 75% and achieving up to 6 times decoding throughput for a 1M context. These results demonstrate that Kimi Linear can be a drop-in replacement for full attention architectures with superior performance and efficiency, including tasks with longer input and output lengths.

To support further research, we open-source the KDA kernel and vLLM implementations, and release the pre-trained and instruction-tuned model checkpoints.

article

Kimi K3 Now Available via Telnyx Inference API

by fionaattelnyx·▲ 129 points·86 comments·telnyx.com ↗

Kimi K3, Moonshot AI's 2.8-trillion-parameter flagship model, is now available on the Telnyx Inference API. It is the world's first open-source model in the 3-trillion-parameter class, built on Kimi Delta Attention and Attention Residuals with a 1M-token context window and native vision capabilities.

What's new

  • New model available: Kimi K3 (model ID: moonshotai/Kimi-K3) is now selectable on the Telnyx Inference API alongside existing models including Kimi K2.6, GLM-5.2-FP8, and MiniMax M3.
  • 2.8T parameters: The largest open-weight model available on Telnyx Inference. First open-source model to reach the 3-trillion-parameter class.
  • 1M token context window: Supports codebase analysis, long document processing, and multi-turn agent sessions with stable long-context performance.
  • Native vision: Accepts text, images, and video input within the same model. Multimodal reasoning without a separate vision adapter.
  • Configurable reasoning effort: Three levels (low, high, max) to trade compute for depth of reasoning per request.
  • Tool calling and structured output: Supports function calling, dynamic tool loading, and JSON schema constrained output for agentic workflows.
  • Prompt caching by default: Automatic prefix caching for repeated prompt prefixes across requests.

Why it matters

The competitive advantage in AI is shifting from who builds the smartest model to who builds the infrastructure that decides where every request runs, and K3 is evidence that the model side of that equation is solving itself. Kimi K3 is the first open-source model to reach 2.8 trillion parameters, and on benchmarks for coding, reasoning, and agentic knowledge work, it competes with closed-source frontier models from Anthropic and OpenAI, proving that open-source is not far behind the frontier labs, and in some cases is already there.

K3 now runs on Telnyx-owned GPU infrastructure and can be access via the OpenAI-compatible API.

Pricing

Token Type Price per 1M tokens
Cached Input $0.27
Input $2.70
Output $13.50

Learn more in the Inference documentation or try it in Mission Control.

The Daily Front Page 20 of 25
Tuesday, July 28, 2026 The Daily Front No. #260728 — Open Models & Security Tools
article

Using an open model feels surprisingly good

by msaltz·▲ 311 points·136 comments·matthewsaltz.com ↗

I've been using Claude and ChatGPT like the next guy for probably two years now. I've never been a huge "open software" nerd or anything like that. But just now, I got opencode working on my own inference endpoint and... it felt surprisingly good. It feels... freeing, somehow. I own the endpoint, and my data just goes from my laptop to there and back. It feels like it's mine. It's really nice.

The motivation for this was that I just got home and wanted to start on a little side project, but I don't have the best Claude or ChatGPT plan for my personal account. I work at Modal, and today we just launched Kimi K3 on managed endpoints, and I know Kimi K3 is supposed to be pretty solid, so instead of upgrading my Claude plan, I wanted to give it a try. (I didn't directly contribute to this feature, so I haven't gotten to play with it yet.)

Within about 5 minutes, I had opencode pointed at my own Modal endpoint and running. Spinning up opencode, I just felt a nice, empty blankness. The best way I can describe it is like opening vim after spinning a bunch of time in a big fancy editor. Or maybe like the tendrils tying me to other providers had been cut, and I could breathe freely. I'm being somewhat dramatic here but not exaggerating that much. It's weird, lol, and unexpected to me, which is why I wanted to write about it.

repository

Codex Security

by bakigul·▲ 475 points·152 comments·github.com ↗
★ 2,996⑂ 158 forks TypeScript

SDKs and CLI for Codex Security

@openai/codex-security is a CLI and TypeScript SDK for finding, validating, and fixing security vulnerabilities in your code. Scan repositories, review changes, track findings over time, and run security checks in CI.

Documentation

Quick start

Requires Node.js 22 or later, Python 3.10 or later, and access to Codex Security.

npm install @openai/codex-security
npx codex-security login
npx codex-security scan .

For CI, set OPENAI_API_KEY instead of signing in.

If both a ChatGPT sign-in and an API key are available, interactive scans ask which credential to use. CI and other noninteractive scans keep the existing API-key precedence. Select a credential explicitly when needed:

npx codex-security scan . --auth chatgpt
npx codex-security scan . --auth api-key

To make your ChatGPT sign-in the automatic default, unset any configured API keys:

unset OPENAI_API_KEY CODEX_API_KEY

Scan history is stored in the Codex Security workbench state directory. If that directory cannot be written, set CODEX_SECURITY_STATE_DIR to a writable directory outside the repository.

TypeScript SDK

import { CodexSecurity } from "@openai/codex-security";

const security = new CodexSecurity();
const result = await security.run(".");

console.log(result.reportPath);
await security.close();

For installation, authentication, scan options, and CI setup, see the official documentation.

About

SDKs and CLI for Codex Security

The Daily Front Page 21 of 25
Tuesday, July 28, 2026 The Daily Front No. #260728 — Apple: Features and Financing
article

The iPhone Upgrade Program is being replaced by Apple Upgrade

by lkurtz·▲ 164 points·315 comments·apple.com ↗

Let farewell lead you to your next hello.

The iPhone Upgrade Program is coming to an end, and we want to say thanks for being a member. For now, you can continue making your remaining monthly payments. And when it’s time for your next iPhone, we’ll make it easy to get it in a way that works for you, including an entirely new payment option that we think you’ll love.

When you’re ready for your next iPhone, we’re ready to help.

There’s more than one way to get your next iPhone. You can lease with our new program, Apple Upgrade; shop the latest carrier deals; finance with Apple, or buy with a one-time payment. Have questions about what’s best for you? Chat with a Specialist (Opens in a new window) online or in a store.

Explore your options

Apple Store Team Member, smiling

Get your favorite Apple products in a whole new way.

iPhone 17 Pro, back exterior, iPhone silhouette fans out from behind into spectrum of purple and red colors

Apple Upgrade

Love it. Lease it. Upgrade it.

Lease a new iPhone, iPad, Mac, or Apple Watch with low monthly payments and terms that work for you. Then easily upgrade to something new at the end of your lease, and return your current device.¹

  • Choose your product and term
  • Make low monthly payments
  • Upgrade at end of lease
  • Only at Apple

Learn moreabout Apple Upgrade

footnotes

1. Apple Upgrade is a device leasing program available in the U.S. (excluding U.S. territories). Leases are provided by Klarna; subject to eligibility and credit approval, including final approval at checkout. To be eligible, you must be a U.S. resident, at least 18 years old (or the legal age in your state), have an accepted credit or debit card, and have an Apple ID. Additional eligibility criteria apply. Device must be in good condition upon return; damage fees may apply. For iPhone only: In order to lease an iPhone, you must select an eligible carrier (but you cannot use a prepaid carrier plan). Upgrades require entering into a new lease and are subject to eligibility and credit approval. Apple Upgrade is not available on refurbished devices or online at the following special stores: Apple Employee Purchase Plan; participating corporate Employee Purchase Programs; Apple at Work for small businesses or enterprises; Government, Education, or Veterans and Military Purchase Programs.

The Daily Front Page 22 of 25
Tuesday, July 28, 2026 The Daily Front No. #260728 — The Future of the Net
The Daily Front Page 23 of 25
Tuesday, July 28, 2026 The Daily Front No. #260728 — Hacks & Retro Tech
article

RTX 2080 Ti Memory Upgrade to 22 GB

by wslh·▲ 148 points·128 comments·gpusolutions.net ↗

Expand your NVIDIA GeForce RTX 2080 Ti from 11 GB to a massive 22 GB of VRAM. This upgrade gives your card the extra memory needed for modern gaming, 3D rendering, and AI workloads. GPU Solutions offers this unique service to customers worldwide.

What This Upgrade Does

We double the RTX 2080 Ti’s VRAM capacity by replacing all original GDDR6 memory modules with higher-density parts. With 22 GB available, your card handles ultra-high-resolution textures, heavy creative workloads, and memory-intensive AI applications more efficiently.

What’s Included

  • Removal of all stock 1 GB GDDR6 modules and installation of 2 GB GDDR6 modules (BGA rework).
  • VBIOS configuration to ensure full detection of 22 GB VRAM.
  • Thermal pad and paste replacement for optimal cooling.
  • Full benchmarking and stability testing (Furmark, 3DMark, real-world stress tests).

Compatibility & Notes

  • This service is specific to RTX 2080 Ti (TU102) cards with compatible PCB layouts.
  • BIOS modifications are required for the upgrade to work properly.
  • Overclocking headroom may differ slightly from factory behavior.
  • Cards with severe corrosion, burnt PCBs, or prior botched repairs may be ineligible.

Service Cost

The final cost depends on the availability of high-capacity memory modules. Replacement parts are quoted separately. If the upgrade cannot be completed, our Diagnostics / No-Fix fee policy applies.

Turnaround

The time required to complete the upgrade and test for stability is 12 days. This can depend on the number of GPUs in the queue for repairs.

Warranty

All workmanship and replaced memory components are covered under our 90-day repair warranty. Warranty does not cover unrelated faults or later modifications.

Book the 22 GB Upgrade

Ready to transform your RTX 2080 Ti into a true VRAM powerhouse? Book your upgrade online today.

Book RTX 2080 Ti → 22 GB Upgrade
View Pricing Policy

Service Details

Pickup and delivery

Yes


Pick and delivery charges

AED50


Service Price

AED0


Time Required

12 Days


Service Code

GCRUP-01


Service Type

GPU Repair


Warranty

90 Days


Service Price

Below you can check price by type or brand and to get accurate value check devices.

AED0

article

Half-Life ported to Mac OS 9

by freediver·▲ 215 points·104 comments·mac-classic.com ↗

Half-Life has finally landed for PowerPC based Macintosh computers 28 years after it's original release! Half-Life is a story driven first-person shooter, that follows scientist Gordon Freeman, who is a theoretical physicist trying to survive and escape the Black Mesa Research Facility after a failed experiment opens a portal to an alien dimension.

The game was originally planned to be released for Mac OS 9 by Valve in 1999, but was cancelled shortly before launch. Valve didn't bring Half-Life to Mac OS X until 2013 which was well into the intel based CPU era, and now we finally have a release for PowerPC based machines.

This port has been accomplished by GitHub user doctashay using a fork of Xash3D FWGS, (a re-implementation of the GoldSrc engine). It's playable from start to finish, includes multiplayer support, a demo of Uplink, along with downloads for Blue Shift and Opposing Force.

This release supports G3 and G4 PowerPC based computers running Mac OS 9.0 or later. Performance heavily depends on the GPU present in your machine, iMacs, iBooks etc. may struggle with performance you have less than 8Mb VRAM.

This release also includes:

  • Half-Life
  • Half-Life: Blue Shift
  • Half-Life: Opposing Force

This is a huge achievement for the Macintosh gaming community, and doctashay deserves considerable credit for the work involved in bringing Half-Life to the PowerPC platform.

Download Half-Life

Download Half-Life 1.0.1

Valve

Half-Life for Mac OS 9 is a story driven first-person shooter, following scientist Gordon Freeman as he tries to survive and escape the Black Mesa Research Facility.

Download

Download

The Daily Front Page 24 of 25
Tuesday, July 28, 2026 The Daily Front No. #260728 — Colophon

That's the Front for Today

Issue No. #260728 — Tuesday, July 28, 2026 — went to press 2026-07-29 at 08:26 UTC.

About This Magazine

The Daily Front is a daily digital magazine assembled from the stories that reached the front page of Hacker News on Tuesday, July 28, 2026. Headlines, points, and comment counts are recorded as they stood at press time. All articles remain the property of their original authors — every piece links back to its source and its discussion thread.

How It Was Made

Fetched, cleaned, and typeset by an automated pipeline. An editor model laid out the pages, chose the highlights, and briefed the cover illustrator — 28 model calls and 402k tokens in total. Set in Jacquard 12, Playfair Display, Source Serif 4, and IBM Plex Mono, all served via Google Fonts under the SIL Open Font License.

The Cover

The cover illustration was commissioned with this prompt:

Abstract painted masterpiece for an artsy magazine cover: an earthquake unfolding in a Japanese coastal city under luminous early-morning light, wooden homes and storefronts tilting above fractured streets, swaying power lines, and puddles vibrating in bold seismograph-like rings. Deep earthquake rifts tear through the painted scene, revealing raw unpainted canvas beneath the thick layers of paint​. An open emergency kit and a glowing laboratory vial on ice anchor the foreground. Soft peach, pale gold, misty lavender, warm coral, seafoam, and clear sky blue; oversized expressive brushstrokes, heavy impasto, palette-knife texture, visible layered paint, elegant abstract motion, cinematic composition, high detail, no text or logos.

Production Ledger

StageModelCallsTokens InTokens Out
extractgpt-5-mini 27 243,898 128,658
layoutgpt-5 1 19,254 9,972

The Publisher

Published by Johnny.

Support the Press

If The Daily Front brightens your morning, consider supporting its publisher.

Credits & Contact

All content — articles, posts, comments, and the images within them — belongs to its original authors and is reproduced here to point readers back to the source. Full credit goes to those creators; every item links to its original and its Hacker News discussion.

If you are an author and would like your content removed from an issue, write to hi@johnnys.page and it will be taken down.

Feedback is always welcome at the same address: hi@johnnys.page.

Credit where credit is due.

Every page of this issue began as someone else's work — these are the original sources, linked in full.

  1. 7.1 Earthquake in Japan by krembo — data.jma.go.jp·HN discussion ↗
  2. New HIV vaccine shows unprecedented success in preclinical study by codebyaditya — lji.org·HN discussion ↗
  3. Substack writers, you need a website by speckx — elizabethtai.com·HN discussion ↗
  4. Benchmarking Opus 5 on SlopCodeBench by dhorthy — github.com·HN discussion ↗
  5. How to survive boiling water by cainxinth — taxa.substack.com·HN discussion ↗
  6. A $500 RL fine-tune of a 9B open model beat frontier models on catalog review by ilreb — fermisense.com·HN discussion ↗
  7. Zig's Incremental Compilation Internals by garyhtou — mlugg.co.uk·HN discussion ↗
  8. Steel Bank Common Lisp version 2.6.7 by tmtvl — sbcl.org·HN discussion ↗
  9. Discovering Cryptographic Weaknesses with Claude by gslin — anthropic.com·HN discussion ↗
  10. DMARC has been public since 2012 but most company domains still don't enforce it by adulion — ciphercue.com·HN discussion ↗
  11. Ars Astronomica – English translations of rare Hebrew and Latin astronomy texts by sweisman — arsastronomica.com·HN discussion ↗
  12. How Do I Profile eBPF Code? by snaveen — naveensrinivasan.com·HN discussion ↗
  13. About the security content of macOS Tahoe 26.6 by andor — support.apple.com·HN discussion ↗
  14. Netflix employee fired for sharing personal details in retreat trust exercise by softwaredoug — nypost.com·HN discussion ↗
  15. Show HN: Yap – OSS on-device voice dictation for macOS with no model to download by pancomplex — github.com·HN discussion ↗
  16. C/C++ projects packaged for Zig by jcbhmr — github.com·HN discussion ↗
  17. Kimi K3 Architecture Overview and Notes by ModelForge — sebastianraschka.com·HN discussion ↗
  18. Kimi Linear: An Expressive, Efficient Attention Architecture (2025) by ronfriedhaber — arxiv.org·HN discussion ↗
  19. Kimi K3 Now Available via Telnyx Inference API by fionaattelnyx — telnyx.com·HN discussion ↗
  20. Using an open model feels surprisingly good by msaltz — matthewsaltz.com·HN discussion ↗
  21. Codex Security by bakigul — github.com·HN discussion ↗
  22. Google's Beyond Zero: Enterprise Security for the AI Era by jordigg — spawn-queue.acm.org·HN discussion ↗
  23. Vehicle Motion Cues by Austin_Conlon — support.apple.com·HN discussion ↗
  24. The iPhone Upgrade Program is being replaced by Apple Upgrade by lkurtz — apple.com·HN discussion ↗
  25. Now is the time to give LLMs access to the ACM digital library by rbanffy — cacm.acm.org·HN discussion ↗
  26. Stop Killing the Internet: No Digital ID and No Age Verification by doener — citizens-initiative.europa.eu·HN discussion ↗
  27. Delayed Gratification – Proud to Be 'Last to Breaking News' by speerer — slow-journalism.com·HN discussion ↗
  28. RTX 2080 Ti Memory Upgrade to 22 GB by wslh — gpusolutions.net·HN discussion ↗
  29. Half-Life ported to Mac OS 9 by freediver — mac-classic.com·HN discussion ↗
  30. Una GPS smart watch – Repairable, USB-C charging, developer-friendly by pimterry — unawatch.com·HN discussion ↗

Browse all issues in the archive →