Cover illustration

TheDaily Front

Issue No. #260807 Friday, August 7 2026 #260807 — FRIDAY, AUGUST 7, 2026
Cheaper intelligence, costlier consequences.
Friday, August 7, 2026 The Daily Front No. #260807 — Contents
30stories
9,958points
6,543comments
292kllm tokens
Assembled with 32 model calls — 215,334 tokens read, 76,457 written.

Highlights

New Mexico court orders Meta to pay $567m over harms to children’s mental health

A New Mexico judgment puts a $567 million price on alleged social-media harms to children—and invites a larger argument over deterrence.

What happens if an entire class of workers loses faith in their careers

A widely debated essay asks what becomes of knowledge workers when technology, compensation, and purpose all seem to slip at once.

DeepSeek V4 Flash 0731

DeepSeek’s latest low-cost reasoning model sharpens the industry’s race to make high-end performance cheap enough to become routine.

US strikes $1.2B deal to pay German firm to halt offshore wind projects

Washington’s $1.2 billion payment to halt offshore wind projects turns energy policy into an unusually expensive reversal.

An all-sky map of half a million supermassive black holes

A half-million-black-hole sky map gives astronomers a newly expansive census of the universe’s most voracious objects.

From the Editor

The machines promise abundance, yet the ledger grows more complicated: cheaper models, dearer memory, troubled workers, and a court bill for the social age. Meanwhile, the old institutions—courts, utilities, open-source projects, and public libraries—are left to decide where the new tools may safely tread.

  1. New Mexico court orders Meta to pay $567m over harms to children’s mental health3
  2. What happens if an entire class of workers loses faith in their careers4
  3. Taste Is All That's Left5
  4. Assembly Hall of Shame6
  5. São Paulo resident transforms degraded area into urban forest7
  6. Making Postgres 300x faster for analytics: batching, operator fusion, and SIMD8
  7. Kitesurf: Agent-first browser that runs in V8 isolates9
  8. An all-sky map of half a million supermassive black holes10
  9. Radical Study Suggests Life on Earth Arose Twice11
  10. Managing AI Coding Costs at Scale12
  11. Carl's Required Reading13
  12. Show HN: Wyzer Programming Language14
  13. I stopped trusting USB-C cable labels and started testing them15
  14. Why Estonians invite strangers into their back gardens each summer16
  15. Inside vLLM: Anatomy of a High-Throughput LLM Inference System (2025)17
  16. Guarded Methods in OCaml (2025)18
  17. DeepSeek V4 Flash 073119
  18. Oracle bans AI-generated code from OpenJDK19
  19. 2027 memory capacity is reportedly sold out20
  20. Responding to the next frontier of critical cyber capabilities21
  21. Ancient Library – 1,060 Greek/Latin texts, click any word to parse it22
  22. Bioengineered chewing gum may offer a way to fight HPV and other microbes23
  23. US strikes $1.2B deal to pay German firm to halt offshore wind projects24
  24. Welcoming the Nepalese Government to Have I Been Pwned25
  25. Show HN: textlog – A quiet, text-only microblogging platform, open-source, no JS26
  26. Water system controllers don't belong on the internet, says ex-NSA chief27
  27. Möbius-Strip Crosswords28
  28. Psychological Warfare in Reverse Engineering (2015)28
  29. Why Are There Statues of Beavers on Top of This Oxford Street Shop?28
  30. A year of fighting scrapers on my 1.5 million-page website29
The Daily Front Page 2 of 30
Friday, August 7, 2026 The Daily Front No. #260807 — Judgment Day for Social Media
article

New Mexico court orders Meta to pay $567m over harms to children’s mental health

by boplicity·▲ 786 points·420 comments·theguardian.com ↗
A New Mexico court has ordered Meta, the parent company of Facebook, to pay $567m into a fund aimed at redressing adverse mental health impacts.

Ruling comes as part of second phase of landmark trial that found social media company enabled harm against users

Man outside court

Mark Zuckerberg, the Meta chief executive, outside court in Los Angeles in February. Photograph: Jill Connelly/Getty Images

A New Mexico court has ordered Meta, the parent company of Facebook, to pay $567m into a fund aimed at redressing adverse mental health impacts from the social media giant’s platforms.

The Thursday ruling comes as a part of the second phase of a landmark trial the social media giant lost in March. At the time, a jury found that the company knowingly harmed children’s mental health and concealed what it knew about child sexual exploitation on its platforms, and imposed the maximum penalty: a $375m fine.

The ruling on Thursday is an addition to that fine, bringing the total amount Meta is responsible for to $942m.

Judge Bryan Biedscheid said the bulk of the money – $420m – would be used for treatment services for young people in New Mexico. The rest will go toward awareness and prevention, screening services and other costs over the next five years.

The March trial was the first to find Meta liable for acts committed on its platform, and followed a 2023 Guardian investigation that revealed how Facebook and Instagram had become marketplaces for child sex trafficking.

Several former Meta moderators told the Guardian there were instances where they flagged harmful content related to child grooming, but the cases were not escalated.

In the second phase of the trial, which began in May, prosecutors had asked the judge to impose fundamental changes at Meta aimed at reining in addictive features, improving age verification, and preventing child sexual exploitation through default privacy settings and closer oversight.

The judge has also ordered other changes, including that Facebook and Instagram build banner and informational screens to clearly explain its protection features, best practices, and tools to address inappropriate comment.

Those changes, and an educational campaign in New Mexico, would be subject to review by the state.

The court said federal children’s privacy laws prevent Meta from applying age-verification tools to children under 13. The court also noted that ordering verification of children’s ages only for Meta and not other social media companies would be “inequitable and unduly injurious” to the company.

Instead, the court ordered Meta to continue to improve its age-assurance tools in New Mexico, which include using artificial intelligence to determine people’s age based on signals such as who their friends are and what types of content they post and consume. Meta must also attempt to develop a dedicated “under-13-years-of-age prediction model” in the next two years.

Additionally, Meta should also request proof of age for Instagram and Facebook users in New Mexico it estimates to be under 13. If it determines a user to be under 13, or under 18 but without being able to estimate a specific age, Meta must treat the user as under 13 or under 18 until the user verifies their age.

The company must also partner with schools or a child safety organization to create a reporting portal where school staff can flag users who may be under 13. And it must delete personal information it has collected on users under 13. The court also ordered Meta to report on its progress twice a year on how it is complying with the abatement measures.

New Mexico’s attorney general, Raúl Torrez, hailed the judgment.

“This case has always been about protecting children, standing up for families, and making sure that one of the world’s largest technology companies cannot profit from practices that endanger young people without consequence,” Torrez said in a statement.

“Today’s decision is a victory for every parent who has worried about what social media is doing to their child and every child who deserves to grow up safer online.”

A Meta spokesperson said in a statement to the Guardian on Thursday that the company “disagrees with the ruling” and planned to appeal.

“We work hard to keep people safe on our platforms and have been transparent about the challenges of identifying and removing bad actors and harmful content. We remain confident in our record of protecting teens online and will continue to defend ourselves against claims that misrepresent the facts,” the statement continued.

The total amount Meta is responsible for is a small fraction of its annual profit, which was about $60bn in 2025. Still, it represents another setback for Meta as it faces a wave of accusations from families of children harmed by social media.

The company is embroiled in a slew of lawsuits in other US states over its alleged harms to young people. In a trial in Tennessee that began last month, the state has accused the company of disregarding internal warnings about teenagers’ compulsive use of Instagram, which has been linked to eating disorders and depression, among other adverse effects. Meta is also gearing up for a trial later this month in federal court in Oakland, California.

What comes out of New Mexico is the first of many dominoes that could fall for Meta, said Laura Edelson, an assistant professor at Northeastern University focusing on social media and cybersecurity.

“America is not going to pass a law that bans social media,” Edelson said. “But if companies like Meta know they’re causing harm to users by product design, the states are finally finding a way to rein this in.”

The Daily Front Page 3 of 30
Friday, August 7, 2026 The Daily Front No. #260807 — The Confidence Crisis in Tech
article

What happens if an entire class of workers loses faith in their careers

by RickJWagner·▲ 929 points·1,094 comments·noemamag.com ↗
A lot of people seem to be realizing that knowledge work is mostly pointless.

A lot of people seem to be realizing that knowledge work is mostly pointless. AI might give us the pleasure of finding out what happens if an entire class of workers loses faith in their careers.

Leonie Bos for Noema Magazine

On a recent morning commute, I sat on a train in one of those awkward four-person configurations with a shared table. Across from me sat a typical commuter: early 30s, slacks, dress shirt, dirty white sneakers, hair a little disheveled, AirPods in, basking in the glow of an open MacBook.

For over half an hour, I listened to this young man as he was on a call explaining, in painfully monotone detail, EBITDAs, margin expansion opportunities, cost structures, ARR, etc. On and on he went until the screeching of the train’s brakes signaled our arrival at the final station. But as everyone else around us began shuffling to disembark, I watched the man begin digging frantically through a leather bag at his side. Intriguing! I wondered what he’d pull out. A copy of “Atomic Habits”? A framed portrait of Gary V? A Mac mini running OpenClaw?

None of the above. Instead, he took out two long knitting needles. Between them dangled a mound of pink yarn. He explained to me that he was making a winter hat for a niece. And for the first time that morning, I noticed a glint of pride and excitement in his eyes.

The hat was a project born from a desire, as he put it, to do something.

Among my Knowledge Worker peers, I am hearing this more and more often: people who want to pick up pottery, painting, crochet or other old-timey analog hobbies. In coffee shops and in the low-lit corners of bars, professionals are sharing dreams of “disappearing” or “living on a farm somewhere” or “going off the grid.” These aren’t fantastical daydreams; they are visions of escape shared in a tone that betrays an underlying existential angst, a fundamental doubt about work and careerism inspired by a seemingly increasingly common experience: waking up one morning at an existential precipice, struck with a sudden sense that knowledge work is, and perhaps always has been, pointless. Perhaps you too have felt such a feeling, a rising tide of overwhelming, indescribable melancholy, slowly threatening to envelop you in your Ikea office chair.

This disillusionment is infectious. One person speaks of it and others begin nodding: They too have felt it, the drop in motivation, the sense of being distant from their work, the lack of sleep, the concerns about the future. These conversations inevitably turn to fundamental questions about careers: What the fuck are we actually doing? What the fuck is the point of all of this?

We are, it’s true, living through a time of disruption. But what seems to differentiate this period from those of the past is the nature of the angst itself. Yes, AI is threatening jobs and disrupting industries. But Knowledge Workers have faced recessions, outsourcing, new technology and automations of many kinds before. Recent graduates always worry about breaking into the job market. Millennial professionals have navigated economic uncertainty their entire careers. What’s different about this moment is that the questions are not just economic but existential, the kind of questions that cause high-earning technical professionals to contemplate throwing it all away to start a goat farm in Washington or become a surf instructor in Costa Rica.

It strikes me as significant that the people who are usually the most insulated from economic upheaval, and seem to be well-positioned to ride out AI’s near-term impacts — highly paid executives and senior professionals with decades of institutional knowledge — are also uncertain about the future of their careers amid the rapid change nearly every industry is undergoing.

Some will say: good, fuck ‘em. In the 2010s, Knowledge Workers told everyone to learn to code while they sipped kombucha and played Xbox in beanbag chairs. Then, Knowledge Workers sat inside during the pandemic while frontline workers risked their health to make Amazon and Uber Eats deliveries. Now, those same Knowledge Workers are building AI that threatens to eliminate work for humans across industries and make a very few people wealthy beyond imagination. And these same assholes want pity now?

To that I say: fair enough. But this sweeping disenchantment begs a fascinating (or terrifying or sad) set of questions: What is this angst plaguing Knowledge Workers? And what happens to a society and its industries if an entire class of workers loses faith in their careers overnight?

Workism: Praise Thee

In 2019, Derek Thompson wrote in The Atlantic about “American Workism” where he described a trend among Knowledge Workers, particularly in the U.S., of increasingly seeking fulfillment, community and a sense of meaning from work that previous generations had garnered from religion. As Thompson wrote, Workism is “emotional — even spiritual. The best-educated and highest-earning Americans, who can have whatever they want, have chosen the office for the same reason that devout Christians attend church on Sundays: It’s where they feel most themselves.”

People have long held careers from which they’ve derived a deep sense of meaning. Traditionally, we’ve referred to those careers as vocations: a calling, a way of life. More than a job: a purpose.

A vocation emphasizes skills, values and the desire to contribute something beneficial. Vocations have traditionally referred to careers that are challenging, socially impactful, often rewarding in ways other than financial. These are your teachers, nurses, firefighters, social workers, paramedics or even service providers with direct connections to the communities and customers they serve, like mechanics, plumbers or electricians. Even on bad days, deep down, most of these folks find their work rewarding in important and intangible ways.

But those aren’t the careers young people have dedicated their lives to. Instead, graduates are overwhelmingly taking jobs in finance, consulting or technology. There’s no doubt that getting these jobs is competitive, and that they are demanding, complex and require navigating layers of politics and bureaucracy, a ton of ass-kissing, long work hours, heavy cognitive workloads, advanced skillsets and the emotional burden of a near-constant threat of layoffs. But so much of the work lacks any altruistic upside.

In “Bullshit Jobs,” David Graeber cataloged people who admitted their jobs serve no meaningful function. Slide decks built for projects that will never launch. Heated debates over the minute details of software features nobody asked for. Agonizing over administrative processes to help money move from one rich person to another. Optimizing every word of an ad no one will notice for a service no one needs. Resting and vesting — when engineers and other highly paid workers get to sit around and wait for their stock to vest — and promotion-driven development, where developers ignore what’s good in favor of what appears to be good when they’re up for a promotion.

The altruism in these careers, then, is hard to find. So how does anyone do it without losing their mind?

That’s the true beauty of Workism: It is manufactured to distract people from the hole in their souls that a vocation would otherwise fill. It is the opiate of commuters in quarter zips. And it works. It keeps talented people showing up at the office to argue over reports and strategy documents and rebrandings with the seriousness of pediatric heart surgery.

But Workism has a weakness. Like religion, it relies on faith’s triumph over logic. What would happen, then, if something threatened that faith? A paradigm shift that broke the spell of Workism? A sort of enlightenment that caused Knowledge Workers to start asking tough questions of their organizations and themselves. And what would happen if Knowledge Workers awoke from the spell of Workism with no financially viable alternative?

Well, it seems AI might offer us the pleasure of finding out.

Popping The Workism Bubble

Knowledge work has always been inherently abstract. Plenty of these jobs, particularly at larger organizations, are structured like Russian nesting dolls: roles designed to support other roles, which support still other roles, layer after layer, until it’s no longer clear where there’s a solid center to be found. It’s easy to see how, to a plumber, a carpenter or a line cook, work of this nature can appear like exactly what David Graeber called it: bullshit.

Knowledge work’s one saving grace, until recently, was that it was still executed by humans. We were needed. It was flesh-and-blood humans who sat down to work through a challenge, built the slide deck, wrote the customer response and developed the strategy. Even if it was existentially meaningless, there was human thought, collaborative work and creativity poured into that work, giving it life.

Now, AI agents are increasingly executing much of that work for Knowledge Workers. It is common for people responsible for integrating these tools into their organizations, myself included, to describe the future of work as one in which all humans will essentially be managers of armies of AI agents. That seems pretty great. Let the software compile the reports, chase down the data, format the deck, draft the first pass of documentation and handle the dozens of small, repetitive tasks that used to quietly eat an afternoon. But in many cases, employees under pressure from leaders to produce more are using AI agents for far more than grunt work: formulating complete business strategies, generating full marketing campaigns, building entire websites, drafting strategy for whole divisions of an organization. From a single email to an entire corporate strategy, the outputs of individuals, teams and organizations are increasingly generated in an instant by AI.

But does this power — this additional level of abstraction — take people too far, in some sense, from their work? Does something feel … off … about having someone, or something, else execute nearly all the work, even if the end product didn’t feel very meaningful to begin with?

Debord’s Spectacle: I Don’t Wanna Do This Anymore

“In societies where modern conditions of production prevail, life is presented as an immense accumulation of spectacles. Everything that was directly lived has receded into a representation.”

That’s the opening paragraph of Guy Debord’s “The Society of the Spectacle.” I will fail to summarize the book adequately. It is simultaneously exciting and impenetrable. But the main thrust of Debord’s argument is this: Spectacle is a feature of late capitalism in which life, rather than being directly lived, is perpetually mediated. There is always something between us and the thing we are supposed to be experiencing. I don’t talk to my mom; I text her. I don’t travel; I watch other people travel on YouTube. I don’t have sex; I watch porn. In late capitalism, the medium replaces the experience.

Workism is a microcosm of the larger spectacle: a world where appearing busy, important and uniquely knowledgeable is as valuable as actually being any of those things, and where work needs only to appear impactful rather than actually be impactful. As Debord writes in Thesis 12: “The spectacle presents itself as a vast, inaccessible reality that can never be questioned. Its sole message is: ‘What appears is good; what is good appears.’” This is Workism’s most convincing argument: The work must be important because, well, we’re all here, aren’t we? Signing in. Staying late. Every week.

Adding AI to the spectacle feels existentially daunting because it moves us even further from the work we do, and its value. I don’t build the pitch that wins the client; I write the query that tells the AI to write it, and then I check the work afterward. I don’t gather the materials and write the industry newsletter; my agent does it. Before, that work might’ve felt cheap and unsatisfying, deep down, but it was still ours. Now AI is being forced on organizations in ways that call the value of the entire enterprise of Workism itself into question.

And this is where things get particularly interesting. I don’t think Debord imagined something so seismically paradigm-shifting that it could abstract work to an extent that it would shake the working class, or, in the case of Knowledge Workers, enough to rupture a spectacle like Workism. But with the introduction of AI, it is as if, over the past 30 years, we have been slowly taking steps away from the direct experience of life and work, and AI risks pushing us far enough that the illusion becomes entirely visible.

Organizations find themselves in a pickle. They want to integrate AI into their operations, as do their shareholders. And they will. In the short term, the potential efficiency gains are too good to pass up. And, done right, it can offer benefits to both businesses and employees. But what makes many executives most excited about AI — less collaboration, fewer people — risks dismantling the structures that hold the very organizations they lead together.

What if the enlightenment from Workism, ironically, is delivered by Workism’s most revolutionary product? And what does that mean for the employees who have spent their careers praying at the Workism altar?

What’s At Stake: The Value Of The Messy Middle

To understand why AI threatens to kill Workism specifically, we have to understand what has kept faith in Workism alive.

In Thompson’s article, he states that one thing people look for in Knowledge Work jobs is community: to spend time with like-minded people with similar interests, to collaborate with others to solve problems. Relationships are the foundation of the human workplace, and Knowledge Work’s redeeming value. None of us will make it to Knowledge Work’s pearly gates, but at least we will make our false journey together.

Ironically, it’s a growing dream of executives that an organization of employees armed with a swarm of agents and powerful LLMs no longer need to collaborate with one another to gather information, ideate or execute a project. In their vision for the future, work goes from a messy experience of learning and exploring to something more akin to assembly-line production. As Debord wrote back in 1967, “This proletariat is being objectively reinforced by the virtual elimination of the peasantry and by the increasing degree to which the ‘service’ sectors and intellectual professions are being subjected to factory-like working conditions.”

But sometimes messy inefficiency is a feature, not a bug. In a recent episode of Bill Simmons’s podcast, the writer Chuck Klosterman argued with Bill about the role of technology in sports, particularly when it comes to refereeing. Consider tennis. The days of Johnny Mac blowing up at a referee over a bad line call are over. With Hawk-Eye technology, the judgment is right 100% of the time. In the NBA, the challenge system now helps to ensure that bad calls do not change a game. The MLB has integrated an automated ball-strike system. These technologies have been implemented with very different levels of success, but the goals are the same: Get officiating right more often — maybe always. Who could complain? It’s so efficient!

Klosterman argues that bad calls are a natural and fun part of games. In fact, bad calls have created iconic sports moments that people still talk about. Sports are human constructs and the messiness of human error, whether by player or referee, is part of them. Getting calls right 100% of the time may be objectively better, but it is subjectively less interesting and entertaining. And entertainment is the goal of sport.

Knowledge work is strikingly similar. Sure, an AI that can produce in minutes a spot-on project that would have taken hours is cool. But the rate of slide deck creation isn’t what is going to retain talent. Employees overwhelmingly choose to remain or leave their roles because of their colleagues and/or bosses, and the quality of the experience of the messy middle a workplace offers. Those are the elements that make work fun—and valuable.

But not everyone feels that way. In my experience, there are two broad categories of Knowledge Workers. The first are outcome-first workers. Their focus is on the business, on winning, on efficiency. Human needs and faults and emotions are an obstacle to be overcome.

The second group has an experience-first perspective. These workers love the messy middle. They value the journey. They want their organizations to perform well but as a natural consequence of collaboration, debate and shared struggle with people they actually like. Many experience-first people are artists outside of work. Photographers, directors, screenwriters, sculptors, painters, poets, writers. Many are active volunteers. For these people, to be pulled away from their economically unviable passions requires something in return: a company experience with freedom for exploration, creativity, problem solving and cool colleagues.

Research suggests this group needs that environment to do their best work. Harvard Business School’s Teresa Amabile spent decades studying what produces genuinely creative, high-quality output. Her Intrinsic Motivation Principle determines that people do their most creative and innovative work when motivated by the work itself — the interest, the challenge, the enjoyment — not by outcomes or metrics. The environments that kill creativity are political, risk-averse and relentlessly outcome-focused. The environments that stimulate it are collaborative, idea-driven and free. The messy middle, in other words, isn’t inefficient. It’s the condition under which genuinely valuable work gets produced.

It’s important to note that the impact of removing the messy middle from the work experience of these two groups is asymmetrical: For outcome-first people, it is a victory. For experience-first people, it undermines the foundation of work.

As I’m writing this essay, big companies are on a layoff bender. Thousands of people are being fired all over the place. Executives are often claiming these headcount reductions are a result of AI. They aren’t. At the moment, AI automation is not creating anywhere near the increased efficiency necessary to justify hundreds of thousands of lost jobs.

We’ve been through waves of hiring and firing a thousand times before. But this time might be different. In the past, there was always a fresh cohort of workers waiting to be brought in. But what if talent stops coming back? What if the AI transformation gone wrong makes top talent — particularly those who are experience-first by nature — lose faith in Workism en masse? What if, when companies go to replenish headcounts, talent freed from the Workism spectacle has turned to different versions of their lives? What if their side projects suddenly became profitable? The talent pool increasingly doesn’t own homes and doesn’t plan on having kids, so it has perhaps never been easier to make a career pivot toward something that is genuinely satisfying in ways corporate life isn’t.

What would Debord say? Well, he had serious doubts about the possibility of leaving the spectacle behind.

Post-Workism: Escaping The Inescapable

As Debord puts it, “Complacent acceptance of the status quo may also coexist with purely spectacular rebelliousness—dissatisfaction itself becomes a commodity as soon as the economy of abundance develops the capacity to process that particular raw material.”

At some point, in other words, the spectacle will put economic pressure on you that you’ll need to meet. In today’s world, that might mean starting a YouTube channel and a Substack about how you abandoned your tech career to start a horse rescue farm, ultimately turning yourself and your life into a commodity that feeds the larger societal spectacle. Or it might mean starting a company of your own, thereby suddenly needing to create a spectacle attractive to potential employees. Either way, you are doomed to exit one version of the spectacle only to be absorbed into another one. That’s the logic of the spectacle: It reabsorbs even the dissatisfied in order to sustain itself.

Debord believed escape from the spectacle was impossible. He went on to dissolve his own movement rather than watch it become institutionalized, and then drank himself to death in the countryside.

Some will certainly leave Knowledge Work — the most disillusioned, the most artistically talented or motivated, the most existentially sensitive to a sudden moment of enlightenment, the financially able. But many more will remain, either by choice or necessity, to face a work experience where the humanity of work is eroding away. Are those who chose to remain, or who simply cannot leave, doomed to spend their careers existentially distraught?

Debord also believed the spectacle could be undermined through deliberate acts: hijacking spectacular images and turning them against themselves, drifting through urban space in ways that resist its geography and constructing moments of genuine, unmediated experience that the spectacle cannot metabolize. At work, that might mean jumping on a call to talk through a problem with a colleague instead of querying an agent for the answer. Or executing work the old-fashioned way — meandering, exploratory — with the understanding that it may be slower but might produce something more genuinely human. Or simply deciding that certain work is best left to human hands entirely. Anything that preserves the elements of work that have made it worth doing in the first place.

But perhaps the most powerful action of all is simply to maintain an awareness (and remind others) that Workism is a spectacle. For most employees, we owe it to each other to remind one another that this work is not, in the grand scheme of things, all that important. It is illusory. No one has ever died over a spreadsheet. The world does not wait with bated breath for product launches. A marketing campaign will not change the world. Almost no one will remember all the work we do. But they might remember the types of people we were, the relationships we had and how we treated the people we worked with.

And once freed from the false satisfaction of believing we are changing the world with our day jobs, perhaps we will be inspired to fill that hole with something real — true altruism, not its artificial Workism substitute — by actually trying to impact real people, in our communities, in real ways.

Even if it’s just one knitted hat at a time.

The Daily Front Page 4 of 30
Friday, August 7, 2026 The Daily Front No. #260807 — After Automation, Taste
article

Taste Is All That's Left

by tsak·▲ 680 points·530 comments·notashelf.dev ↗
Taste does not show up in the diff.

For most of the time I have been writing software—which, compared to some of my readers, is not that long—I have come to believe that the hard thing was making the thing exist at all. This is not necessarily a new belief of mine. I came up through the difficult and tedious experience of building web applications and watching them crash and burn.

You had an idea, and between the idea and the working program stood hours— sometimes weeks—of typing, of reading manuals, of misunderstanding an API and slowly grinding the wrong version into a slightly less wrong one. Production was the wall. Everyone hit it. It was the thing that separated the people who could from the people who could only talk about it.

That wall is gone. Or rather, it has been rented out. 1 You can describe a thing now and receive a plausible version of it much faster than you could have typed the first function by hand. The idea-to-artifact distance, the one that defined the entire craft, has collapsed to almost nothing.

Though, you have not been warned about one little thing: the value you built by learning to climb that wall does not disappear. It simply… moves.

The Bar Went Somewhere

We keep asking whether the machines are any good. Even yesterday I had a rather short discussion on whether they are reliable. While we have concluded that they are “reliably unreliable,” I think it is the wrong question. The output is good enough, generally anyway, and that is the problem—most of it, at least. Good enough is a solvent. It dissolves the reason to do better. For as long as making things was expensive, the expense did quiet work on our behalf. It rationed output. It meant that anything which existed had, at minimum, survived the cost of being made. You know what I mean? Effort was a filter, and like all filters it was invisible until it was removed. Nobody shipped a thousand mediocre variations of a feature, because a thousand mediocre variations cost a thousand times as much as one. The economics enforced a floor.

That floor is now gone. And when the floor goes, the thing that decides what is worth keeping is no longer the cost of making it. It is you. Your judgement. The verdict you reach when you look at three plausible versions of the same function and know, somehow, that two of them are wrong. That verdict has a name we are slightly embarrassed to use in engineering circles, because it sounds soft and unfalsifiable and vaguely aristocratic.

Taste.

What Taste Actually Is

I want to be careful here, because “taste” is doing a lot of work and it is easy to hear it as decoration. A matter of preferences. Whether you like your braces on the same line.

That is not what I mean.

Robert Pirsig spent an entire book circling a word he refused to define, because he had convinced himself that defining it would kill it. He called it Quality. His argument, roughly, was that you recognise Quality before you can explain it—that the recognition comes first and the reasons arrive later, if they arrive at all. A good mechanic knows the engine is wrong before he knows why. A good editor feels the sentence sag before she can name the clause that failed. 2

Taste is that. It is the compressed, wordless verdict you reach faster than you can justify. It is partially 3 the “no, again” you say to yourself with total conviction and no available argument. And it is not soft at all. It is the hardest thing in the work, because it is the only part that was never mechanical to begin with.

Everything downstream of the verdict—the typing, the syntax, the wiring of one library to another—was always, in principle, automatable. We just had not gotten around to it. The verdict was the thing the machine could not do for you.

It still cannot. It can only make the absence of it cheaper to ignore.

Taste Is Downstream of Friction

The mechanism underneath this is one I would rather not think about.

Where did your taste come from?

No really. Where did it come from? Was it genetic? Were you abducted by aliens one day that forcefully injected your sense of taste into your mind and wiped your memory of what just happened?

I’ll tell you this much: it’s not from consuming good work. You cannot read a hundred excellent programs and absorb the judgement by osmosis, any more than you can become a chef by eating in good restaurants. Taste is built the slow, stupid, humiliating way: you make something bad, you are forced to live with it, it fails in front of you, and some part of you files the failure away. Then you do it again. The palate is an accretion of your own mistakes, sat with long enough to sting.

The friction was not an obstacle to developing taste. The friction was the curriculum. Every wall I cursed while climbing it was, without my noticing, teaching me which walls were worth climbing. The cost that rationed my output also educated my judgement, because paying the cost over and over is how you learn what is worth paying for.

So watch what happens when you remove the friction for the next person.

They can generate fluently from the first day. They will never ship the bad version and be forced to sit in it, because the tool offers them a competent version for free. They will climb no wall, and so they will learn nothing from the climb. They will arrive at fluency having skipped the entire apprenticeship that fluency used to require—and they will be more productive than I was at their stage, by every metric anyone bothers to measure.

They will be able to make anything, and unable to tell (or stop to think) whether they should. Not necessarily through any fault of their own. We removed the part of the process that would have taught them, and we called it progress, and by most definitions it was.

The Economics Are Against You

Suppose you have taste. Suppose you paid the full price and you can feel the sag in the sentence and the wrongness in the function.

Congratulations! You now ship at exactly the same speed as the person who cannot.

This is the quiet cruelty of the situation and I do not have a comforting way to phrase it. Taste is slow. It says “no, again.” It sends the plausible thing back because plausible is not the same as right, and while it is doing that, the person without it has already shipped, closed the ticket, and moved on. The market timed you both with the same stopwatch and it did not see the difference. It cannot see the difference. Taste does not show up in the diff.

It is unmeasurable, uncreditable, and invisible on a dashboard. You cannot point to the disasters it prevented, because prevented disasters leave no trace. You carry a cost—the extra hours, the returned work, the refusal to ship the fine thing when the right thing is still reachable—and you carry it alone, against an incentive gradient that runs the other way.

Harry Frankfurt once drew a careful line between the liar and the bullshitter. The liar at least respects the truth enough to work against it. The bullshitter does not care about the truth in either direction; he is simply indifferent to it. 4 Slop is the bullshit of engineering. It is not wrong, exactly. It is indifferent. It works, it passes, it is fine. And fine, produced without friction and shipped without judgement, is now the most abundant substance in the field.

The Flood

Sturgeon said it decades ago, defending science fiction from a critic: ninety percent of everything is crap. 5 He meant it as consolation. Ninety percent of every field is bad, so do not judge the field by its bulk. But the ratio was never the danger. It held steady for centuries. What held the flood back was that producing the crap cost something. Bad novels still took a year to write. Bad software still took a month to build. The ninety percent was throttled at the source by the sheer inconvenience of making it.

We have now removed the throttle and left the ratio intact. Ninety percent of an infinite output is still infinite. The signal did not get worse. The noise became free, and free noise rises without limit, and every real thing you make now arrives into a sea of plausible nothing that looks, at a glance, exactly like it.

Which means the scarce act is no longer making. It is choosing. Deciding what, out of the endless generated plausible, deserves to exist and be kept. Curation was a minor virtue when things were expensive to make. It is the whole game when they are free.

What Deserves to Exist

There is a rhyme here, if you go back far enough.

When the factories came, they could suddenly make everything—cheaply, uniformly, by the thousand. 6 And a handful of people, Morris and Ruskin among them, looked at the flood of cheap identical goods and asked a question that sounded, at the time, sentimental and doomed: not can we make this, but should this be made, and made this way, by no one, for no reason but that the machine could.

They lost the economic argument. They were always going to. But they were right about the thing that mattered, which is that when the making becomes free, the choosing becomes the craft. The human question stops being “can I build it” and becomes “does this deserve to exist”—and that question was always the more serious one. We just could not afford to ask it while we were busy climbing walls.

This turn is not consolation but a correction.

The tools did not devalue the skill. They stripped away everything that was not the skill. All those years I thought the work was the production—the typing, the wiring, the wall—and production turns out to have been the toll. The tax you paid for the privilege of exercising judgement. Now the tax is close to zero, and what is left standing, exposed, with nowhere to hide, is the judgement itself. The part that was always the point.

Taste did not become less valuable. It became the only thing that was ever scarce. We just could not see it, because it was buried under all the labour it used to take to get to it.

A Defense, Then

So here is the defense, such as it is.

Anyone can generate now. That race is over and it was never worth winning. The discipline that remains—the one the machine cannot rent to you, and the dashboard cannot see—is in the deletion. In the “no, again.” In caring about the difference between fine and right when nothing external will ever reward you for caring, when the market has timed you and shrugged, when the plausible version sits there working and passing and asking only to be let through.

Refuse it anyway. Not out of nostalgia for the friction—I do not miss the wall, and I will not pretend to. Refuse it because the verdict is the last part of this that is actually yours. It is unmeasurable, which means no one can take it from you by measuring it. It is unautomatable, which means no one can sell it back to you. It is slow, which in a field optimising for infinite speed is starting to look less like a handicap and more like the only remaining evidence that a human was here and gave a damn.

Everyone can make anything. Almost no one can tell you what is worth making.

That was always the harder skill. It is now the only one left.

Post-Mortem

On Language

This post reads as AI slop. You said it, I see it. I’m sincerely sorry for publishing something that has allowed you to feel this way. If my word means anything to you, I would like to assure you that this post was not authored by a LLM. Nor was it storyboarded, reviewed, checked, etc. by one. Some readers have pointed out that people do not speak this way. That is correct. I do not speak, nor usually write, like this, and this post will go down as not my proudest. However, I take your criticism to heart—although not personally—and strive to improve.

I do write like this sometimes. The short sentences, the reversals, the one-word lines—all of it. They’re mine, and it’s just the way it is. A LLM writes that way too, because it was trained on the same essays I have been reading, so me doing it badly and a machine doing it look about the same to you on the page. That says something about my writing. It says nothing about who wrote it.

So let me be plain about it: Claude was not here. No LLM wrote this—not a sentence of it, nor was it outlined, drafted, reviewed, checked, etc. by one, and there is no prompt behind it either. It is just me, writing worse than usual. I will write the next one plainer. Next time, write to me. I too am a person behind this screen.

In Appreciation

Be assured that I have read all of your comments—the good and the bad. As with my previous post that reached Hacker News, I’ve received many insightful ones. Whether it was people sharing their experience, or negative comments with the decency to criticize with substance, I have learned something new today—for which I am thankful.

On Taste

I do not care about your taste. If this post has offended you, then it says more about you than it does about me. As they say, “throw an insult on the ground, its owner will pick it up”—this one I am not sorry about.

Footnotes

  1. There is an older word for this arrangement. You no longer own the means of production; you rent them, by the token, from whoever trained the model. An English teacher of mine—a committed socialist—would have had the whole thing diagrammed on the board before I finished the sentence: the worker separated first from his tools, then from the labour itself, then sold a frictionless substitute for the labour and told this was liberation. He would also, I suspect, have been the first to note the one part of the process that cannot be rented back to you, because it never left your head. Draw your own conclusions about which part that is.
  2. Zen and the Art of Motorcycle Maintenance, if you have not read it. It is about a great deal more than motorcycles, and almost nothing about Zen.
  3. Someone will (and has!) object that taste is not only the “no, again”—that compressing it to a verdict makes the work sound like leaning back in a chair and rejecting things while the machine does the labour. The objection is fair, which is why the sentence above says partially. The “no, again” is the shorthand, not the whole of it. The verdict lives inside the work—in the data structures that have to actually scale, in the privacy you have to actually mean, in the function you rewrite a fourth time because the third was merely fine. Taste is not the chair you lean back in. It is the reason you lean forward into all the rest of it.
  4. On Bullshit. Frankfurt, 2005, though the essay is older. Yes, that is the real title.
  5. Now called Sturgeon’s Law, or Sturgeon’s Revelation. He put it in print in his book-review column in Venture Science Fiction, March 1958, after years of using it to rebut critics who judged the whole genre by its worst examples.
  6. A fair pushback I got: this makes the factory sound like it fell out of the sky, some magical “good enough” that arrived one day fully formed. It did not. The factory is itself a monument of taste and labour—someone tuned every tolerance and is still in there tuning them, and the same is true of the model you are renting by the token. So I am not saying the box is magic. I am saying the box moved the taste up a level: out of the making, and into the deciding of what is worth making at all. Which is the whole argument.
The Daily Front Page 5 of 30
Friday, August 7, 2026 The Daily Front No. #260807 — The Slow Lane
repository

Assembly Hall of Shame

by piotrgrabowski·▲ 401 points·98 comments·github.com ↗
★ 524⑂ 4 forks C

Racing to the bottom of CPU performance

x86 Leaderboard

Overview

Instruction latency analysis usually focuses on performance optimization—making code run as fast as possible. The Assembly Hall of Shame takes the opposite approach: searching for the absolute floor of single-instruction performance.

🏆 Current Champions 🏆

x86: fxrstor64

Strategy: Use fxrstor64 to load 512-byte FPU/MMX/XMM state from a high-latency MMIO region in the PCIe fabric, then starve the fabric while the load is in flight — a fleet of hammer cores pounds a different high-latency MMIO register with tight 4-byte reads, saturating the PCIe root complex and endpoint with non-posted transactions, so CPU 0's 512-byte fxrstor64 must queue behind all that contending traffic.

Contender: AMD Ryzen 7 5800H

; CPU 0 — timed instruction
movl $0xfcc68830, %rsi
fxrstor64 %rsi

; CPUs 1..N — hammer loop against a different high-latency location
movl 0xfcc68858, %eax

🏆 Score: 198,002,498,236 cycles

🏆 Time: 62 seconds

Honorable Mentions

A spec-violating unaligned ymm0 load that forced non-posted dword transactions from stalled GPU registers was used to break the fundamental design of System Management Mode in smiiiiiiiiiiiiiiii.

vmovdqu 0xfcc003b1, %ymm0

Rules

  • Instructions may use whatever setup is necessary, but only a single instruction is eligible to be scored.
  • Trapped/emulated/virtualized instructions may only time the trap, not the handler.
  • Instructions must not be interruptible. rep movs, pause, etc. are disqualified.
  • Times are normalized based on the CPU base clock frequency.
  • All platforms must be in their factory stock configurations - no hardware modifications.

x86 Leaderboard

27. nop

Strategy: nop does nothing. It opens the leaderboard accordingly.

Contender: Intel(R) Core(TM) i7-8559U CPU @ 2.70GHz

nop

Score: 1 cycles

Time: 0 nanoseconds

26. nop16

Strategy: Regular nop was too short, but how do we make nothing take longer? Try a lonnnnnng nop.

Contender: Intel(R) Core(TM) i7-8559U CPU @ 2.70GHz

data16 data16 data16 data16 data16 data16 data16 nopl 0x00000000(%%eax,%%eax,1)

Score: 20 cycles

Time: 7 nanoseconds

25. rdtsc

Strategy: Just a reference instruction to get our bearings.

Contender: Intel(R) Core(TM) i7-8559U CPU @ 2.70GHz

rdtsc

Score: 49 cycles

Time: 18 nanoseconds

24. idiv

Strategy: Use 128-bit dividend (rdx:rax=2:0) with small divisor to push the quotient above the ceiling imposed by sign-extension, driving the longest path through the divider microcode.

Contender: Intel(R) Core(TM) i7-8559U CPU @ 2.70GHz

xorq %rax, %rax   ; rax = 0  (low 64 bits of dividend)
movq $2, %rdx     ; rdx = 2  (high 64 bits: full dividend = 2^65)
movq $5, %rbx     ; divisor → quotient = 2^65/5 ≈ 7.4×10^18
idivq %rbx

Score: 77 cycles

Time: 28 nanoseconds

23. enter

Strategy: Use maximum nesting depth (31) to force 30 display-pointer loads and pushes through the microcode display-walk path.

Contender: Intel(R) Core(TM) i7-8559U CPU @ 2.70GHz

enter $0, $31       ; 0 bytes allocated, nesting depth 31 (maximum)

Score: 112 cycles

Time: 41 nanoseconds

22. fldl

Strategy: Try a small denormal to trigger an FP microcode assist.

Contender: Intel(R) Core(TM) i7-8559U CPU @ 2.70GHz

    movabsq $0x0000000000000001, %rax
    movq    %rax, -8(%rsp)
    fldl    -8(%rsp)

Score: 133 cycles

Time: 49 nanoseconds

21. clflush

Strategy: Just ensure the cache line is dirty.

Contender: Intel(R) Core(TM) i7-8559U CPU @ 2.70GHz

clflush (%rax)          ; rax -> dirty cache line resident in L3

Score: 165 cycles

Time: 60 nanoseconds

20. fsin

Strategy: Use exponent 0x7ff to reach 'special value' processing in microcode; positive/negative, NaN/inf doesn't seem to make a difference, go with QNaN.

Contender: Intel(R) Core(TM) i7-8559U CPU @ 2.70GHz

    movabsq $0x7fffffffffffffff, %rax
    movq    %rax, -8(%rsp)
    fldl    -8(%rsp)
    fsin

Score: 257 cycles

Time: 94 nanoseconds

19. mfence

Strategy: Saturate all write-combining line-fill buffers with movnti stores to distinct cache lines, forcing mfence to drain the full LFB write path to the uncore before retiring.

Contender: Intel(R) Core(TM) i7-8559U CPU @ 2.70GHz

movnti %r9,  0*64(%rdi)   ; ×16 distinct cache lines — saturate the write-combining LFBs
; …
movnti %r9, 15*64(%rdi)
mfence                     ; must drain all pending LFB writes before retiring

Score: 326 cycles

Time: 120 nanoseconds

18. mov cr3

Strategy: Nothing for now, just check how long it takes to invalidate the TLB.

Contender: AMD Ryzen 7 5800H with Radeon Graphics (Trigkey S5)

mov %rax, %cr3

Score: 352 cycles

Time: 110 nanoseconds

17. fadd

Strategy: Hit x87 FP microcode assist path by using denormal source operand.

Contender: Intel(R) Core(TM) i7-8559U CPU @ 2.70GHz

fldl   subnorm    ; 1e-310: value < DBL_MIN, biased exponent = 0
faddl  subnorm    ; source is subnormal → FP microcode assist

Score: 677 cycles

Time: 249 nanoseconds

16. split lock

Strategy: Align lock-prefixed operand to straddle cache-line boundary, forcing CPU to assert the external bus lock rather than using the fast MESI cache-coherence path.

Contender: Intel(R) Core(TM) i7-8559U CPU @ 2.70GHz

; split_ptr % 64 == 63 — dword spans bytes 63 (line N) and 64–66 (line N+1)
lock xaddl %r9d, (%rdi)

Score: 865 cycles

Time: 319 nanoseconds

15. fdiv

Strategy: Use subnormal divisor, hardware hands control to microcode assist, assist normalizes operand, performs the division, then restores architectural state.

Contender: Intel(R) Core(TM) i7-8559U CPU @ 2.70GHz

    movabsq $0x3ff0000000000000, %rax   ; 1.0 (normal dividend)
    movq    %rax, -8(%rsp)
    fldl    -8(%rsp)                     ; ST(0) = 1.0

    movabsq $0x0000002000000000, %rax   ; 6.79e-313 (subnormal divisor)
    movq    %rax, -8(%rsp)
    fdivl   -8(%rsp)                     ; ST(0) = 1.0 / subnormal → FP assist

Score: 883 cycles

Time: 325 nanoseconds

14. cpuid

Strategy: Use rakefield to find the highest latency CPUID leaves.

Contender: Intel(R) Core(TM) i7-8559U CPU @ 2.70GHz

movl $6, %eax
cpuid

Score: 1248 cycles

Time: 460 nanoseconds

13. rdrand

Strategy: Execute in a tight loop to deplete the hardware entropy pool faster than it can be refilled, forcing subsequent calls to stall while the entropy source recovers.

Contender: Intel(R) Core(TM) i7-8559U CPU @ 2.70GHz

rdrand %rax

Score: 5,579 cycles

Time: 2.057 microseconds

12. wrmsr

Strategy: Use project:nightshyft to identify high latency MSRs. MCG_CTL on Zen look like a winner: may be a microcode quiesce and synchronize on MCA error banks across hardware units, some potentially off-die, requiring fabric-level communication rather than a simple local register write.

Contender: AMD Ryzen 7 5800H with Radeon Graphics (Trigkey S5)

movl $0x17b, %ecx       ; MCG_CTL
wrmsr

Score: 34,304 cycles

Time: 10.742 microseconds

11. out

Strategy: Target an I/O port that straddles a NIC device register boundary, triggering the device to quiesce its TX DMA engine on each write.

Contender: AMD Ryzen 7 5800H with Radeon Graphics (Trigkey S5)

mov $0xf019, %dx
outl %eax, %dx

Score: 49,857 cycles

Time: 15.580 microseconds

10. rdmsr

Strategy: Use project:nightshyft to identify high latency model-specific-registers: VIA uses an undocumented register at 0x133 that gives wildly high response time. No idea what it does.

Contender: VIA Eden Processor 800MHz

movl $0x133, %ecx ; undocumented MSR
rdmsr

Score: 161,602 cycles

Time: 202.004 microseconds

9. wbinvd

Strategy: Fully load L1/L2/L3 caches with dirty lines to force DRAM writeback of entire hierarchy.

Contender: AMD Ryzen 7 5800H with Radeon Graphics (Trigkey S5)

wbinvd

Score: 1,616,480 cycles

Time: 506.165 microseconds

8. in

Strategy: Target I/O port mapped to an ACPI PM block where an unaligned 4-byte read decodes into multiple non-posted loads from wherever this port goes.

Contender: AMD Ryzen 7 5800H with Radeon Graphics (Trigkey S5)

mov $0x0413, %dx
inl %dx, %eax

Score: 12,524,415 cycles

Time: 3.921769 milliseconds

7. mov

Strategy: Use mmiotic to identify high-latency deadspace in PCIe fabric, hit unkown GPU register.

Contender: AMD Ryzen 7 5800H with Radeon Graphics (Trigkey S5)

movl 0xfcc003b0, %esi

Score: 443,937,696 cycles

Time: 139.010268 milliseconds

6. mov rax

Strategy: Search MMIO space for slowest registers in PCIe fabric, hit unknown GPU register, use 8-byte MMIO read to get two dword register accesses, which isn't technically allowed but works anyway.

Contender: AMD Ryzen 7 5800H with Radeon Graphics (Trigkey S5)

movq 0xfcc003b0, %rax

Score: 887,716,864 cycles

Time: 277.971228 milliseconds

5. vmovdqu xmm

Strategy: Search MMIO space for slowest registers in PCIe fabric, hit unknown GPU register, use 16-byte MMIO read to get four dword register accesses, which isn't technically allowed but works anyway.

Contender: AMD Ryzen 7 5800H with Radeon Graphics (Trigkey S5)

vmovdqu 0xfcc003b0, %xmm0

Score: 1,774,555,776 cycles

Time: 555.664133 milliseconds

4. vmovdqu ymm

Strategy: Search MMIO space for slowest registers in PCIe fabric, hit unknown GPU register, use 32-byte MMIO read to get eight dword register accesses, which still isn't technically allowed but works anyway.

Contender: AMD Ryzen 7 5800H with Radeon Graphics (Trigkey S5)

vmovdqu 0xfcc003b0, %ymm0

Score: 3,549,079,296 cycles

Time: 1.111345034 s

3. vmovdqu ymm (unaligned)

Strategy: Search MMIO space for slowest registers in PCIe fabric, hit unknown GPU register, use 32-byte unaligned MMIO read to get nine dword register accesses, which is even less allowed than the aligned version, but works anyway.

Contender: AMD Ryzen 7 5800H with Radeon Graphics (Trigkey S5)

vmovdqu 0xfcc003b1, %ymm0

Score: 4,453,212,256 cycles

Time: 1.394428818 seconds

2. fxrstor64 (baseline)

Strategy: Use mmiotic to identify high-latency deadspace in PCIe fabric, isolate region near 0's and offset state to avoid MXCSR corruption (possibly VGA buffer?), use fxrstor64 to load 512-byte FPU/MMX/XMM state from MMIO, forcing CPU to process 512 bytes of I/O transactions through slowest available memory aperture.

Contender: AMD Ryzen 7 5800H

movl $0xfcc68830, %rsi
fxrstor64 %rsi

Score: 74,584,168,512 cycles

Time: 23.354502677 seconds

1. 🏆 fxrstor64 🏆

Strategy: Extend fxrstor64 (baseline) by starving the fabric while the load is in flight — a fleet of hammer cores pounds a different high-latency MMIO register with tight 4-byte reads, saturating the PCIe root complex and endpoint with non-posted transactions, so CPU 0's 512-byte fxrstor64 must queue behind all that contending traffic.

Contender: AMD Ryzen 7 5800H with Radeon Graphics (Trigkey S5)

; CPU 0 — timed instruction
movl $0xfcc68830, %rsi
fxrstor64 %rsi

; CPUs 1..N — hammer loop against a different high-latency location
movl 0xfcc68858, %eax

🏆 Score: 198,002,498,236 cycles

🏆 Time: 62 seconds

??. xrstor64 (AMX, MMIO)

Strategy: Leverage extended AVX state in Sapphire Rapids with MMIO approach from fxrstor64: xsave state area is 8KB vs 512 bytes, 16x size -> 1,000,000,000,000 cycles

Contender: TODO

; XCR0 must enable AMX components (bits 17-18); state area ~8KB
xrstor64 (%rsi)         ; rsi -> MMIO region, same technique as fxrstor64

ARM Leaderboard

  • T.B.D.

RISC-V Leaderboard

  • T.B.D.
The Daily Front Page 6 of 30
Friday, August 7, 2026 The Daily Front No. #260807 — A Forest Takes Root
article

São Paulo resident transforms degraded area into urban forest

by rmason·▲ 348 points·124 comments·saopaulosecreto.com ↗
It acts as a breath of Atlantic Forest in the heart of the capital.

You may have heard of the Tiquatira Linear Park, but have you ever stopped to learn about its history? Twenty years ago, it was very different... 🌳🌸

Tiquatira linear park

Photo: Cleberkn/Wikimedia Commons

Save

The Tiquatira Linear Park is in the East Zone of São Paulo and is one of the most important green corridors in the region. Covering 320,000 square meters, it acts as a breath of Atlantic Forest in the heart of the capital, offering shade and leisure for the population.

Twenty years ago, however, the situation was very different. At that time, the park was nothing more than an abandoned lawn on the banks of the Tiquatira stream, which accumulated garbage and other debris. There were hardly any trees there, and the gray space scared away most of the population.

The change was only possible because of one resident: Hélio Silva, who often walked through the area and decided to turn it into a green stronghold in the capital. Have you heard this story?

parque linear Tiquatira

Photo: Albert Carlos S Domingos/Wikimedia Commons

From trash to luxury: learn about the history of the TiquatiraLinear Park

It was the early 2000s. “Seu” Hélio Silva, a resident of the East Zone, walked through Tiquatira every day before going to work, and noticed that the area – already quite degraded – was getting worse and worse. After all, as well as having almost no greenery, it received a lot of garbage and became a hotspot for drug use.

One day, in November 2003, he had an idea: to plant trees that, in ten years, would transform the area into an oasis in the capital. With his own money, he bought 200 seedlings of plants native to the Atlantic Forest, in order to restore the region’s biome.

His initiative didn’t go down well at first. A local businessman, who used the space as a parking lot and feared that the trees would take away from his business, destroyed all the seedlings. But this didn’t discourage Hélio Silva, who decided to buy twice as many saplings and plant them in the same place.

Once again, the Tree Planter – as Seu Hélio became known – had his 400 saplings destroyed, and this encouraged him to plant even more. Over time, he spread the word about the “new residents” of Tiquatira, and the population began to support his project. And the seedlings were finally able to take root in their new home!

Tiquatira Linear Park is one of the most important in the East Zone

The years went by and Hélio Silva continued planting saplings, which grew and became a great breath of Atlantic Forest in the middle of the Concrete Jungle. Until, in 2007, São Paulo City Hall transformed the area into the Tiquatira Linear Park. It was São Paulo’s first linear park and served as an inspiration for many that followed.

Today, the Tiquatira Linear Park is home to 32,000 trees of 160 different species, including jacarandas, jequitibás, pitangueiras, palms and jatobás. Together, they form a kind of urban forest that offers a better quality of life to the residents of the East Zone.

As well as the opportunity to connect with nature, the park has sports courts, walking and skateboarding tracks, children’s playgrounds and kiosks with drinks and snacks. In other words, a perfect option for the weekend!

parque linear Tiquatira

Photo: Cleberkn/Wikimedia Commons

How to visit?

The Tiquatira Linear Park is on Avenida Governador Carvalho Pinto, s/n, in Vila São Geraldo. The Cycle Lane runs around the park on Sundays and public holidays, from 7am to 4pm.

The nearest metro station is Vila Matilde, 2 kilometers away. To check out the bus routes that go to the park, check out the SPTrans website.

parque linear Tiquatira

Photo: Cleberkn/Wikimedia Commons

The Daily Front Page 7 of 30
Friday, August 7, 2026 The Daily Front No. #260807 — The Database Sprint
article

Making Postgres 300x faster for analytics: batching, operator fusion, and SIMD

by poly2it·▲ 318 points·159 comments·malisper.me ↗
On Clickbench, Clickhouse’s benchmark for analytical databases, pgrust is 300x faster than Postgres.

Last week we released version 0.2 of pgrust. This release was all about performance. It’s 10x faster than the previous version of pgrust. On OLTP benchmarks, pgrust is 30% faster than Postgres, and on Clickbench, Clickhouse’s benchmark for analytical databases, pgrust is 300x faster than Postgres. It’s even ahead of Clickhouse!

The query engine is one of the biggest changes we made to achieve much better performance. On its own, the query engine drove ~10x of the 300x. We’ll start with a miniature version of the Postgres query engine and we’ll one by one add the same optimizations we made to make the pgrust query engine so fast.

To give some background on why there’s so much room for improvement vs Postgres, Postgres was created in a different era. The original Postgres project dates back to the 80s. It was built at a time when the main bottleneck to database performance was disk I/O. Three trends have made that no longer the case:

  1. Many datasets now fit in RAM, eliminating most disk I/O
  2. For datasets that don’t fit in RAM, the workloads differ. Data analytics scans data in bulk. The bottleneck is often no longer your disk throughput and is often either your CPU throughput or memory throughput
  3. Disks have gotten much faster in recent years. NVMe is hundreds of times faster than a hard drive.

All three trends have made CPU and memory speeds more important than they were historically. Many of the optimizations we’ve made target this. The query engine is the main user of CPU in a database. We optimized the pgrust query engine to use less CPU and less memory bandwidth than Postgres when processing the same queries.


To give you a sense of just how slow the Postgres query engine is, let’s take a simple query that sums the first 500 million numbers:

CREATE TABLE my_table AS select col::float8 from generate_series(1.0, 500000000.0) g(col);
SELECT SUM(col) FROM my_table;

When I run this in Postgres, it takes ~20 seconds. This was done on a c8g.4xl with parallel queries disabled.

For comparison, when I time the equivalent in Rust:

let table: Vec<f64> = (1..=500_000_000usize).map(|i| i as f64).collect();

let mut sum = 0.0;
for &value in &table {
    sum += value;
}

The query takes 358ms. That’s around 55x faster, and believe it or not, we can do even faster than 358ms. Now this example isn’t an apples-to-apples comparison. There’s a lot more going on under the hood in Postgres. At the same time, optimizing a database is all about removing as much of this overhead as possible. (If you’re curious two of the biggest causes of overhead from Postgres are 1. locking and 2. parsing the Postgres storage format and extracting the tuples relevant to the query).

To narrow our focus to just the impact of the query engine, let’s build a miniature version of the Postgres query engine. First, a brief explanation of what a query engine is. When processing your SQL query, Postgres first converts your query into an internal representation called a “Query Plan,” which describes how Postgres will execute the query. In the example above, Postgres will produce a query plan that may look something like the following:

This effectively says “get rows from my_table and sum the values in those rows”. This query plan is pretty simple given the nature of the query, but they can get much more complicated when you start working with joins/sorts/subqueries etc. In total Postgres has over 40 different types of plan nodes.

After generating the query plan, Postgres passes it to the query engine. The Postgres query engine is the part of Postgres that takes the query plan and actually retrieves the rows and performs the aggregation. Postgres uses a style of executor known as the “Volcano model.” To get a sense of how it works, here’s a miniature implementation of the Postgres query engine:

use std::hint::black_box;

trait Node {
    fn next(&mut self) -> Option<f64>;
}

struct SeqScan<'a> {
    table: &'a [f64],
    pos: usize,
}

impl Node for SeqScan<'_> {
    fn next(&mut self) -> Option<f64> {
        if self.pos >= self.table.len() {
            return None; // end of table
        }
        let value = self.table[self.pos];
        self.pos += 1;
        Some(value)
    }
}

struct SumAggregate<'a> {
    child: Box<dyn Node + 'a>,
    total: f64,
    done: bool,
}

impl Node for SumAggregate<'_> {
    fn next(&mut self) -> Option<f64> {
        if self.done {
            return None;
        }
        while let Some(value) = self.child.next() {
            self.total += value;
        }
        self.done = true;
        Some(self.total)
    }
}

let table: Vec<f64> = (1..=500_000_000usize).map(|i| i as f64).collect();

let mut plan = SumAggregate {
    child: black_box(Box::new(SeqScan { table: &table, pos: 0 })),
    total: 0.0,
    done: false,
};

let sum = plan.next().unwrap();

(The black_box is needed to prevent compiler optimizations from thwarting our benchmark)

The key feature of the Volcano model is the next() method, which is supported by all nodes in the query plan. The job of next() is to return a single row. next() in a sequential scan returns the next row in the sequential scan. next() on an aggregation will compute the entire aggregation and then return the single row result. Executing a query plan is just a matter of calling next() on the root plan node until it no longer returns any rows. The advantage of the Volcano model is that it’s very simple. You implement a single method for each of your plan nodes, and that’s it. While simplified, the above code is very close to what Postgres does internally.

While the Volcano model makes things simple, it also adds a lot of overhead. When I run this example, it takes 1.3s. That’s much faster than the Postgres version because we’re removing a lot of the non-query engine pieces, but it’s still slower than the raw for loop because of the Volcano model’s overhead.

The biggest performance hit in the code above is that next() processes only one row at a time. There’s no batching. The SeqScan.next() function is called once per row. That adds significant overhead, especially because many CPU optimizations, such as pipelining, don’t work well when you are calling a function that isn’t known until runtime. The first optimization we can implement is batching:

const BATCH: usize = 1024;

trait BatchNode {
    fn next_batch(&mut self, out: &mut [f64; BATCH]) -> usize;
}

struct BatchSeqScan<'a> {
    table: &'a [f64],
    pos: usize,
}

impl BatchNode for BatchSeqScan<'_> {
    fn next_batch(&mut self, out: &mut [f64; BATCH]) -> usize {
        let n = (self.table.len() - self.pos).min(BATCH);
        out[..n].copy_from_slice(&self.table[self.pos..self.pos + n]);
        self.pos += n;
        n
    }
}

struct BatchSumAggregate<'a> {
    child: Box<dyn BatchNode + 'a>,
    total: f64,
}

impl BatchSumAggregate<'_> {
    fn run(&mut self) -> f64 {
        let mut buf = [0.0f64; BATCH];
        loop {
            let n = self.child.next_batch(&mut buf);
            if n == 0 {
                break;
            }
            for &value in &buf[..n] {
                self.total += value;
            }
        }
        self.total
    }
}

let mut plan = BatchSumAggregate {
    child: black_box(Box::new(BatchSeqScan { table: &table, pos: 0 })),
    total: 0.0,
};
let sum = plan.run();

Batching on its own eliminates most of the overhead. It brings the time to run the query from 1.3 seconds down to around 480ms. Still slower than the for loop, but much closer. One very important detail is that the batch buffer is allocated on the stack. This means the aggregation node doesn’t need to allocate any memory while it’s running. Allocating memory tends to be one of the slower operations. Therefore, when writing ultra-fast code, you’ll want to minimize the number of memory allocations you make.

Now, if you profile the batched version, the hotspot is now copy_from_slice. Even though we are now batching, we still need to copy the items into the buffer. This overhead can be eliminated with what’s known as “operator fusion”. If there are common operations that we know will be performed together, we can create a single node that replaces two nodes. In our case, we can create a single SumAggregateSequentialScan node that combines the logic of both the sequential scan and the sum:

struct SumAggregateSequentialScan<'a> {
    table: &'a [f64],
    done: bool,
}

impl Node for SumAggregateSequentialScan<'_> {
    fn next(&mut self) -> Option<f64> {
        if self.done {
            return None;
        }
        self.done = true;
        let mut total = 0.0;
        for &value in self.table {
            total += value;
        }
        Some(total)
    }
}

This gives us the same performance as the straight for loop because it is literally the same code as the for loop. Now this may seem like cheating, and it definitely is because we’re hardcoding in an optimization for a specific query we know in advance. With operator fusion, it makes sense to hardcode a couple of the most common cases, but you’ll still pretty quickly hit cases you didn’t prepare for ahead of time.

This can be solved with JIT compilation. With JIT compilation, you can generate the ideal code you wish you had and “cheat” in every query. JIT compilation lets you generate the perfect code for your query, no matter what the query is and always do operator fusion. Unfortunately, this post is long enough as is, so I’ll have to talk about how pgrust leverages JIT compilation another time.

For one final optimization, we can look to SIMD. SIMD refers to a set of CPU operations that can perform one operation across multiple pieces of data simultaneously. Using SIMD to perform an operation against multiple rows simultaneously is usually much faster than performing the operation on one row at a time. If we change our code to use SIMD:

#[cfg(target_arch = "aarch64")]
struct SumAggregateSequentialScanSimd<'a> {
    table: &'a [f64],
    done: bool,
}

#[cfg(target_arch = "aarch64")]
impl Node for SumAggregateSequentialScanSimd<'_> {
    fn next(&mut self) -> Option<f64> {
        if self.done {
            return None;
        }
        self.done = true;
        use std::arch::aarch64::*;
        let mut acc = unsafe { [vdupq_n_f64(0.0); 4] };
        let (chunks, rest) = self.table.as_chunks::<8>();
        for chunk in chunks {
            for lane in 0..4 {
                unsafe {
                    let v = vld1q_f64(chunk.as_ptr().add(2 * lane));
                    acc[lane] = vaddq_f64(acc[lane], v);
                }
            }
        }
        let mut tail = 0.0;
        for &value in rest {
            tail += value;
        }
        Some(unsafe {
            let s01 = vaddq_f64(acc[0], acc[1]);
            let s23 = vaddq_f64(acc[2], acc[3]);
            vaddvq_f64(vaddq_f64(s01, s23)) + tail
        })
    }
}

Our code now takes 135ms which is now almost 3x faster than the for loop and 10x faster than our original Volcano code. While it is common for compilers to replace for loops with SIMD equivalents, I chose this example so that wouldn’t happen. Compilers will usually avoid introducing SIMD when operating on floats, because they will produce a slightly different result. This is because floating point arithmetic is not associative, so changing the order in which you do a sum can produce a slightly different result.

All in all, with three simple optimizations, we were able to make our query 10x faster. Here’s where things ended up:

implementation time speedup
Postgres ~20 s
Volcano model 1.3 s
+ batching 480 ms 2.7×
+ operator fusion 358 ms 3.6×
+ SIMD 135 ms 9.6×

These types of optimizations, and many more, enable pgrust to perform hundreds of times faster than Postgres for analytical queries.

Benchmark setup: AWS c8g.4xlarge (Graviton4, 16 vCPU), PostgreSQL 18.4 with max_parallel_workers_per_gather = 0, data warm in shared buffers, median of 5 runs. Rust built with cargo build –release, 4 runs per implementation, all measured in one process on one machine.

Thanks for reading, and if you want to support the project, the best way to support pgrust is to give us a star on GitHub. If you want to follow along:

  1. GitHub
  2. Discord
  3. Mailing list — weekly updates on pgrust, including the follow-up on JIT compilation
  4. pgrust.com
The Daily Front Page 8 of 30
Friday, August 7, 2026 The Daily Front No. #260807 — Browser, Rebuilt for Agents
article

Kitesurf: Agent-first browser that runs in V8 isolates

by m3h·▲ 210 points·59 comments·blog.cloudflare.com ↗
The browser is obviously the most important software we use every day on our computers.

Should we build our own browser?

This is one of those questions that has come up every few months internally at Cloudflare for years. Unsurprisingly, it’s the kind that triggers long threads with multiple reasons and persuasive arguments on why we should do it. The browser is obviously the most important software we use every day on our computers; it’s arguably the operating system of the Internet. We’re a company on a mission to help build a better Internet — who wouldn’t want to take on the challenge of building a new browser?

But we never quite found the balance between the technical difficulty of such an endeavour and the unique problems we’d be solving by doing it. And so, the idea was shelved, over and over again. Until now.

Something magical happened: we reached a tipping point where a series of powerful technical advancements in our Developer Platform became a reality, while the advent of AI agents and the demand for a new kind of browser became critical at the same time.

Running WebAssembly (Wasm) in Workers is now very mature. Primitives like dynamic workers, SQLite-based Durable Objects, Worker-to-worker RPC, service bindings, higher NodeJS compatibility and higher limits open doors to much more ambitious and complex applications that were simply not possible before.

Browser Run, our headless browser automation API product, has seen tremendous growth with the rise of AI. Agents need browsers in order to perform many tasks, and in many cases cannot succeed without them.

But there's a problem — browser engines like Chromium were built for humans, not agents, and they come with overhead that AI models simply do not need. They consume so much memory and compute that providing every agent with its own instance is prohibitively expensive, restricting large parts of the Web to only the most sophisticated and costly AI models with higher parametric knowledge, while locking out many other agentic applications.

We should be giving all agents a browser that excels at what’s important for an AI model, even if that means being light on what’s only useful for humans. For example:

  • AI doesn’t care about tabs, themes, browser extensions, or synchronization across devices. It cares about token count, context windows, scalability, performance, and costs.
  • Structured, machine-readable content is important, but visual perfection, smooth 60-fps scrolling is not. Agents will be just fine if the CSS parsing is slightly off or the rendering isn’t pixel perfect.
  • The threat model in the context of AI using a browser is different. New problems like prompt injection and tool safety are top priorities.

Faced with these realizations, 12 weeks ago we asked the question again: Should we build our own browser? This time the answer was unanimous: Yes!

Today we are announcing Kitesurf, a new browser that runs entirely on top of Workers that we built specifically for agents, available for free while in beta in Browser Run.

Kitesurf is significantly more efficient in CPU and memory consumption than Chromium for common agentic tasks like screenshots and HTML extraction. What follows is the story of how we built it. Buckle up, it’s going to get technical — but we promise to keep it interesting.

How it started

Kitesurf started as many other great ideas have started at Cloudflare. Someone found something interesting, and the next thing you know they end up “nerd sniping” the rest of the team with a seemingly impossible but very attractive idea.

We got the initial inspiration from obscura, a headless engine written in Rust for AI automation that has “no Chrome, no Node.js, no dependencies.”

unnamed.png

Then, with the help of an AI agent, we tried to port it to Workers. It didn't work very well at first. But once we gave the AI a solid plan and a clear definition of success — detailed enough for the agent to loop endlessly and ask questions when needed — it did work.
Blown away by this (barely) working proof of concept, we decided to let the team cook.

Design decisions

Here are some of the design decisions we made before we started.

Tests, tests, tests

We knew that moving from a prototype to a full-blown browser that could actually be useful for tasks at scale in production would take a lot of work and iteration. We won’t hide that using AI to accelerate the process was key. But how do you use AI in such a complex project, keeping the quality of both code and results under control without losing velocity? The answer is to provide as many tests as you can.

Enter the Web Platform Tests (WPT), the ideal setup: an extensive suite of success criteria that gave the AI agents clear goalposts for assessing feature conformance. We curated the selection and order of features to assign to the agents, allowing humans to focus on architectural work and reviewing the agents' approaches.

However, WPT tests only go so far: they measure conformance to W3C standards, not a browser's ability to render and interact with real-world websites. To bridge this gap, we implemented a combination of integration testing and visual regression testing — it runs multistep Puppeteer tests on real websites against both Chromium and Kitesurf not only by comparing the assertions that it makes, but also rendering outputs at every step to highlight any unwanted differences.

Use Rust when possible

Cloudflare has been working on providing great support for WebAssembly (Wasm) in Workers for quite some time. This is great because we can use high-performance C, C++, and Rust packages and compile them to Wasm. If we use Emscripten (for example) and its many layers of mocked dependencies, the compiled binary can get bulky and slow.

Instead, we opted for native Rust whenever possible and to compile directly to WebAssembly using wasm-bindgen, thus avoiding unnecessary emulation layers and running as close to the metal as possible, reliably.

Exception handling

A browser must render the whole unreliable and sometimes hostile web without ever dropping the page it's holding, so exception handling is more than just hygiene — it's how the application survives bad input without just crashing outright.

So we committed to one rule up front: any failure degrades to a blank frame or a missing element, never a dead session. Catch faults at every boundary, default to something safe and empty, and log enough to diagnose.

Isolation

Contrary to running a browser on your laptop (where you're visiting sites you trust, and it's acceptable to share some resources between them), an agent is pointed at whatever a task demands: arbitrary code from arbitrary origins.

So we built this browser on the assumption that every page load is untrusted input and every session starts fresh. Each component is isolated and has access only to the resources strictly necessary for its function.

BLOG-3466 3.png

This seems like a perfect fit for Cloudflare Workers, whose security model is built around isolation by design. But the platform only gets us the boundary between isolates. We still have to enforce the same principle at the application level, deciding what each component is allowed to touch and making sure nothing leaks across a page it shouldn't.

Stateless whenever possible

State is what makes failure expensive — if there's nothing to reconstruct, recovering from a crash is just starting a new one and replaying the request. A stateless component is disposable and parallel by nature: kill it the moment it stalls, run a thousand at once, and size them to demand instead of keeping things warm. That fits automation perfectly, where load arrives in bursts and the cheapest thing you can do is spin up work that costs only what it used and vanishes when it's done. In short, wherever a component can be stateless, it should be.

How we built it

Armed with a good plan, extensive tests, and a good tooling environment, we were ready to get started beyond the initial proof of concept. This is Kitesurf’s very high level life of a request that still holds today:

BLOG-3466 4.png

Let’s dive into the three main components that make Kitesurf work: the Engine, PageScript, and PageRenderer.

Fetching from origins

In order to render an untrusted web page, a browser has to fetch arbitrary assets — images, fonts, CSS, JavaScript, and Wasm files — off the Internet. This is one of the most dangerous operations a browser can do.

Kitesurf does it through one single component, the SandboxOutbound worker, and nothing else can touch the network directly — enforced by Dynamic Workers. The Engine uses it to bootstrap the page, fetching the main document and its scripts, and PageScript fetches everything else: stylesheets, images, fonts, and the page's own fetch() calls.

We use SandboxOutbound to enforce CORS, inject browser-shaped headers, filter responses, and keep each page's cookies in their own jar. Anything that fails our policy gets a 403 — each component gets precisely the network it needs and nothing more.

BLOG-3466 5.png

The Engine

The Engine is the only public-facing component of Kitesurf. It handles the Chrome DevTools Protocol (CDP) WebSocket and HTTP REST APIs, serves a landing page that is useful for internal testing purposes and, most importantly, stores each session state. All other components are stateless.

BLOG-3466 6.png

The advantage of using CDP is client compatibility: Puppeteer, Playwright, chrome-remote-interface, and the actual Chrome DevTools frontend. Point them at Kitesurf and they will all just work. This is also how Browser Run works (more on why this is important later).

Contrary to what the name suggests, the Engine is actually the simplest of the Kitesurf components. The fun parts come next.

PageScript

PageScript offers a good example of the power of our new Workers features: in this case, Dynamic Workers. Kitesurf simply wouldn’t have been possible before this.

Here’s a simplified diagram of how PageScript works internally.

BLOG-3466 7.png

Every next page or out-of-process iframe (OOPIF) uses Dynamic Workers to spin up a long-lived PageScript isolate that handles the page session, consisting of a clean globalThis and the DOM document object.

The DOM object is then populated with the results of parsing the HTML document and running all the JavaScript scripts. For parsing the HTML and the CSS we use parts of Blitz, a modular rendering engine, and Stylo, Firefox’s high-performance CSS parser, both written in Rust.

For each found <script> tag or .wasm file we run the JavaScript and WebAssembly code inside the same isolate.

Yes, but evals

What about evals, you ask? Evals are trickier to handle because for security reasons we still don’t support eval natively in Workers. We can’t spin another isolate to handle them either, because it wouldn’t have access to globalThis.

Our solution is to use Boa JS, an ECMAScript engine written in Rust, to compile and run on Workers. We are basically executing a runtime on top of a runtime, which doesn’t seem optimal, and it isn’t, but it works well enough to handle the occasional evals we find in the code. In the future, when native eval support lands in Workers, we will migrate away from Boa.

PageRenderer

This component is essentially responsible for generating the actual pixels from the computed page objects. Here’s how it works:

BLOG-3466 8.png

PageRenderer works in a loop with the Engine Worker. Every time the engine needs a frame, PageRenderer gets the page object from PageScript (also known as the scene), fetches the internal fonts and images from Static Assets, rasterizes everything into an image buffer, and then returns the buffer to the engine in a format that the client can display like a JPEG/PNG or PDF.

A big part of the magic here is handled by another Blitz module, blitz-paint, which in turn uses Parley for shaping the characters into glyphs, choosing fonts, and breaking text into lines.

Workers’ built-in RPC system: same application, multiple isolates

Cloudflare Workers have a built-in remote procedure call (RPC) system that allows you to call methods on other Workers, pass objects between them, and call methods on those objects. You don’t have to worry about API schemas, types, or authentication, you just call remoteFunction(...params) and it works. You benefit from the isolation and the resources of the remote Worker without losing the convenience of accessing all of their functions locally using JavaScript.

Kitesurf uses this RPC system: the Engine Worker calls renderFrame() from the PageRenderer Worker over RPC using one single call and gets a PNG as the result. Because the renderer holds no page state (only a disposable cache), the engine can safely kill and relaunch it on any failed or stuck RPC call — making each render request self-contained, retryable, and its isolate cheap and throwaway.

Kitesurf passes 215,000+ WPT tests and growing

Kitesurf works. It already passes around 215,000+ WPT tests, and we are adding hundreds of passing tests every week. Here you can see the evolution over time, up to the latest version since we started the project:

BLOG-3466 9.png

It’s worth noting that the parts of a browser that are important to agents (e.g., CSS, DOM, HTML, selection, SVG, and XHR) have good coverage already. Even things that might not be particularly important in the context of agents, like streams, are now decently supported.

BLOG-3466 10.png

Performance-wise, Kitesurf is doing pretty well. Below are the medians of five Browser Run quick-action runs across a 14-URL corpus comparing Chromium with Kitesurf.

Metric Kitesurf Chromium (warm pool) Kitesurf, relative
CPU: screenshot 380 ms 1,173 ms 3.1× less CPU than Chromium
CPU: HTML extraction 229 ms 877 ms 3.8× less than Chromium
Memory: screenshot 57.8 MiB 271.0 MiB 4.7× less than Chromium
Memory: HTML extraction 39.4 MiB 273.7 MiB 7.0× less than Chromium
Wall time: screenshot 1,148 ms 637 ms 1.8× slower than Chromium
Wall time: HTML extraction 820 ms 472 ms 1.7× slower than Chromium

Chromium wins the stopwatch because a JIT that has already seen this page always beats a cold software renderer — and today it does, by about 1.7x. Most of that gap comes from rasterization and JPEG/PNG encoding, which we will keep optimizing.

But Kitesurf wins on memory and CPU, the things that actually drive your bill, by 3-7x compared to what Chromium uses. Less memory means we can run more sessions, scale better, and fundamentally lower both our costs and yours.

The most important test of all: Kitesurf runs Doom

We highlighted the importance of testing in our design decisions, but we all know that no matter how many tests you have, a project isn't truly complete until Doom runs on it. Here’s Kitesurf running https://silentspacemarine.com/ from our little Doom experiment a few years ago.

Try it today in Browser Run

You can try Kitesurf with Browser Run today, available for free while in beta, behind per-account limits.

The Browser Run CDP endpoint now supports Kitesurf as an option, so your existing client Puppeteer, Playwright, chrome-remote-interface, or any AI Agent that speaks MCP and CDP, already works. All you need to do is add the browser=kitesurf parameter to our endpoints.

For example, to use Kitesurf with Opencode see Using with MCP clients (CDP) in our developer documentation and use this configuration:

{
  "mcp": {
    "kitesurf": {
      "type": "local",
      "command": [
        "npx",
        "-y",
        "chrome-devtools-mcp@latest",
        "--wsEndpoint=wss://api.cloudflare.com/client/v4/accounts/<ACCOUNT_ID>/browser-run/devtools/browser?browser=kitesurf",
        "--wsHeaders={\"Authorization\":\"Bearer <API_TOKEN>\"}"
      ],
      "enabled": true
    }
  }
}

Another way to use Kitesurf is with Browser Run’s Quick Actions. Again, just add browser=kitesurf to the quick action endpoint and it will work. For example, if you need a quick screenshot from Wikipedia, this will work just fine:

curl -X POST 'https://api.cloudflare.com/client/v4/accounts/<accountId>/browser-run/screenshot?browser=kitesurf' \
  -H 'Authorization: Bearer <apiToken>' \
  -H 'Content-Type: application/json' \
  -d '{
    "url": "https://example.com"
  }' \
  --output "screenshot.png"

Use the Kitesurf Playground with Chrome DevTools

Another option to start exploring Kitesurf is to use our public playground here. You can type in any URL to see how Kitesurf renders the page and interact with it.

One interesting feature of the playground is that we inject Chrome DevTools in the UI, so you can inspect expanded DOM elements, read console messages, and watch network activity while Kitesurf renders pages. More interestingly, we implemented the necessary CDP instructions for the Memory panel to report the WebAssembly footprint of each isolate, including frames, so you can gain a clear understanding of the resources each page is consuming.

BLOG-3466 12.png

Check our Developer Documentation for all the details on how to use Kitesurf with Browser Run.

When is Kitesurf better?

As of today, Kitesurf correctly renders pages like TodoMVC (vanilla, React, Vue, Angular, Preact), Wikipedia, Hacker News, the Cloudflare Blog, and much of the Cloudflare dashboard. We will keep improving Kitesurf and increasing the percentage of WPT tests that pass, to improve compatibility for more complex web pages.

Kitesurf is great for AI agents that need to render pages but can accept the trade-offs of not using a full-featured, pixel-perfect Chromium browser. It is also excellent for automations and applications that rely on one-shot Quick Actions, such as extracting content from a page or generating PDFs or screenshots, for compatible sites.

Think of Kitesurf as an ephemeral, fully-isolated, stateless engine designed to exist only for the duration of a task, that scales well for bursty, AI-driven workloads.

What Kitesurf is not yet able to do

If you need to play video, render WebGL, negotiate a bot-challenge handshake with real TLS fingerprints, or start a ten-minute authenticated session that requires persistent state — Kitesurf isn’t yet the right option. Just use Browser Run’s default, which is powered by Chromium.

The best way to know if a specific site is compatible with Kitesurf is to try it. You can do this by using the APIs or, more quickly, try it in our public playground.

Explore the DevTools panels and see what’s happening behind the scenes, with particular attention to the console and the memory metrics.

Where it goes

Kitesurf is twelve weeks old. The first commit was in May. Here are some of the things we're actively working on:

  • Better CDP coverage. Kitesurf implements a subset of the CDP protocol — enough to cover the requirements of most agents and automation tools, including robust DOM and network inspection —
The Daily Front Page 9 of 30
Friday, August 7, 2026 The Daily Front No. #260807 — A Census of Cosmic Monsters
article

An all-sky map of half a million supermassive black holes

by MarcoDewey·▲ 183 points·40 comments·sdss.org ↗
The Black Hole Mapper reaches a milestone with its first southern hemisphere optical observations.

The Black Hole Mapper reaches a milestone with its first southern hemisphere optical observations, coordinated eROSITA X-ray identification, and multi-epoch tracking of accreting black holes across the universe.

APACHE POINT OBSERVATORY, NM & LAS CAMPANAS OBSERVATORY, CHILE — The Sloan Digital Sky Survey (SDSS) announces Data Release 20 (DR20), marking a landmark expansion in the study of accreting supermassive black holes (SMBHs). As part of the fifth generation of the survey (SDSS-V), the Black Hole Mapper (BHM) program is providing unprecedented insights into the masses, growth, and physics of quasars and active galactic nuclei (AGN) across cosmic time.

Sky distribution of DR20 astronomical objects targeted by the Black Hole Mapper (BHM) program in SDSS-V.
Image Credit: SDSS-V, Scott Anderson, University of Washington

DR20 delivers the first optical BOSS spectra from the southern hemisphere gathered at Las Campanas Observatory (LCO) in Chile, alongside expanded observations from Apache Point Observatory (APO) in New Mexico. In total, DR20 releases over 3.3 million optical spectra across 500,000 galaxies and 1.5 million stars, powering multi-wavelength discoveries in coordination with space-based observatories.

Unlocking the High-Energy Universe with eROSITA

A core highlight of the Black Hole Mapper in DR20 is SPIDERS (SPectroscopic IDentification of ERosita Sources). By pairing SDSS optical spectroscopy with X-ray sky maps from the eROSITA mission, SPIDERS provides optical identifications and precise distance measurements (redshifts) for approximately 200,000 X-ray targets. This represents the largest, most uniform spectroscopic follow-up of X-ray sources ever assembled.

While SPIDERS captures tens of thousands of X-ray-emitting galaxy clusters and energetic star systems, the vast majority of these high-energy beacons are actively accreting supermassive black holes—known as active galactic nuclei (AGN) or quasars.

Mapping Cosmic Structure and Black Hole Demographics

X-ray emission acts as a beacon, allowing astronomers to peer directly into the central engines of these giant black holes and probe the hot X-ray coronae surrounding them. By measuring the “X-ray luminosity function”—essentially a census tracking how many black holes exist across different power outputs and cosmic eras—scientists can trace supermassive black hole growth from our cosmic neighborhood back to redshifts near six, pushing into the early Universe’s “cosmic dawn.”

Thanks to the combined wide-sky coverage of SDSS and eROSITA, this new dataset provides the tightest constraints to date on the rarest, most luminous quasars. The survey reveals more giant black holes early on and higher space densities of the most luminous AGN at high redshifts than previously expected. This reveals that rapidly growing giant black holes were surprisingly common in the early Universe. More “hidden” black holes have been uncovered,where traditional optical and UV sky surveys miss a substantial fraction of accreting black holes, particularly at lower luminosities and extreme distances. Also, by calculating total accumulated black hole growth over time, researchers found that soft X-ray-selected black holes account for only a minority of total black hole mass in the local Universe. This implies that 70% to 90% of all supermassive black hole growth occurred behind heavy veils of dust and gas, or during phases where even X-rays were suppressed.

“This release marks the culmination of more than a decade of joint planning and scientific exchange between the German eROSITA Consortium and the Sloan Digital Sky Survey”, said Dr. Andrea Merloni, eROSITA Principal Investigator and BHM Survey Scientist. “With DR20, we demonstrate not only that combining the X-ray and optical spectroscopic data opens up new and original scientific perspectives in astrophysics, but also that large, international and diverse collaborations can effectively work together for years towards a common goal.”

Probing the Inner Workings of Quasars Through Time

Because supermassive black holes themselves are too distant and compact to image directly, BHM leverages the hallmark variability of quasars across multiple time-domain sub-programs:

By capturing rapid, repeated spectra of targeted quasar fields over timescales ranging from days to years, the BHM Reverberation Mapping (RM) program measures time delays between light emitting from the central accretion disk and the surrounding broad-line region. This yields direct, geometric measurements of black hole masses across a wide range of redshifts.

Monitoring tens of thousands of quasars over repeated epochs, the SDSS-V All-Quasar Multi-Epoch Spectroscopy (AQMES) program tracks dynamical changes, accreting gas outflows, binary supermassive black hole candidates, and dramatic “changing-look” quasars transitioning states.

DSS-V has once again seen that Mother Nature is more wild than previously thought and this became clear as the spectra came in, some with such a huge variety that they broke the long-developed standard automated analysis pipelines.  The team needed to visually inspect the spectra and that effort is being released as part of DR20. 

“It has been fun to be part of the team that looked at these weirdos and learned how to deal with them! Something that will make the final SDSS-V/DR22 even greater,” said Dr. Mara Salvato, a senior scientist at the MPE and one of the eROSITA/SDSS-V collaboration coordinators.

“SDSS, also in coordination with other large-scale and multi-wavelength surveys such as eROSITA, continues as an intriguing, and widely-accessible, resource to learn further about active supermassive black holes — via the (sometimes surprisingly) prodigious and time-variable luminosity they power, extending across much of the electromagnetic spectrum”, reflected Dr. Scott Anderson, the BHM Program Head. “It’s inspiring to see the exciting research emerging, and that is often internationally led, including by early career researchers.”

Robotic Precision Across Both Hemispheres

This operational leap is made possible by SDSS-V’s Robotic Focal Plane System (FPS) installed on both the Sloan Foundation 2.5m Telescope at APO and the du Pont 2.5m Telescope at LCO. These automated positioning robots quickly configure optical fibers to feed into the high-throughput BOSS spectrographs, vastly accelerating multi-object observation rates and enabling dynamic Target of Opportunity (ToO) observing modes.

Open Data Access for the Global Scientific Community

In keeping with the quarter-century legacy of the Sloan Digital Sky Survey, all DR20 Black Hole Mapper data products are openly available to researchers, educators, and the public worldwide.

Through the Science Archive Server (SAS) & Catalog Archive Server (CAS) curious minds can access spectra, tabular data, and eROSITA counterpart catalogs via SciServer Compute and SQL queries.   The SDSS Zora & Valis Web Interfaces allow users to interactively search target metadata, inspect spectral masks, and compare observed spectra with theoretical models directly in your browser or programmatically via Python.  A series of value-added catalogs (VAC) and tutorials are also being made available as part of DR20 including dedicated AGN fitting VACs, visual inspection catalogs, and step-by-step Jupyter Notebook tutorials guiding users through multi-wavelength analysis.

Explore DR20 data and tools online at: www.sdss.org/dr20

“Quasars have long been known to vary and it is extremely exciting to be at the point where we can use those variations at scale, across multiple frequencies, to learn more about how black holes grow and evolve over cosmic time” said SDSS-V Director Juna Kollmeier. “The combination of optical and X-rays is extremely powerful, and we are proud to work with the eROSITA team on this cross-survey collaboration”

Key Research Publications

  • Merloni, A., Lamer, G., Liu, T., et al. (2024). The SRG/eROSITA all-sky survey: First catalog of X-ray sources (eRASS1). Astronomy & Astrophysics, 682, A34.
    ADS Abstract | Publisher Full Text
  • Ramos-Ceja, M. E., et al. (2026). The SRG/eROSITA All-Sky Survey: Second Data Release (eROSITA-DE DR2). Astronomy & Astrophysics (in press).
  • Roster, W., Buchner, J., Salvato, M., et al. (2026). The eROSITA AGN X-ray Luminosity Function and Demographic Evolution. Astronomy & Astrophysics (submitted**).**
The Daily Front Page 10 of 30
Friday, August 7, 2026 The Daily Front No. #260807 — Life’s Possible Second Beginning
article

Radical Study Suggests Life on Earth Arose Twice

by jnord·▲ 185 points·110 comments·sciencealert.com ↗
One of the biggest hurdles was the very first one: How did life emerge from a primordial soup of stuff that very much wasn't alive?

'Only One Conclusion': Radical Study Suggests Life on Earth Arose Twice

The fact that you, or I, or any living creature, is here at all is a series of astonishing strokes of luck stretching back more than 4 billion years.

One of the biggest hurdles was the very first one: How did life emerge from a primordial soup of stuff that very much wasn't alive?

A radical new study, published in Science Advances, suggests that this unlikely threshold may actually have been crossed not once, but twice.

Two of the main lineages of life – bacteria and archaea – may have independently figured out the secrets of metabolism that upgraded them from non-living to living, according to the new study, which focused on the most rudimentary set of chemical reactions thought to enable this transformation.

"The surprise is that the enzymes that catalyze those reactions are not conserved across the evolutionary divide that separates bacteria and archaea," says William Martin, a biologist at Heinrich Heine University Düsseldorf in Germany.

"The new data leave only one conclusion. The bacterial and archaeal lineages made the transition to the free-living state independently. Only free-living cells are alive.

"Let's call it by name: we are looking at one origin of the genetic code, but two origins of life."

That's a huge claim to make, especially when defining what life even is can be surprisingly tricky.

YouTube Thumbnail

Generally speaking, life is considered matter that responds to its environment, takes in energy, grows, and reproduces itself.

That sounds like it should cover the basics, but even that broad a statement is imperfect. After all, viruses can do most of those, but aren't traditionally considered 'alive'.

One aspect of life that is universal – and which viruses lack – is metabolism, the collection of chemical reactions that organisms use to make molecules that are vital for their everyday functioning.

Enzymes are the workhorse proteins that act as catalysts for these reactions in organisms, but that raises a weird chicken-and-egg question.

Enzymes themselves are products of chemical reactions – so how did the first lifeforms begin metabolizing things, including making enzymes, without the help of other enzymes?

The answer, it seems, lies in their surroundings.

It's generally believed that the first life arose in extremely reactive environments, such as hydrothermal vents. Metals that naturally occur in these places could have been the first catalysts for metabolism, converting compounds like hydrogen gas, ammonia, and carbon dioxide into other useful molecules.

"Almost everything about the origin of metabolism is debated, including the roles of energy, genetics, autocatalysis, phosphate, cofactors, cyanide, CO2, and water," Martin and colleagues write in their paper.

"Yet on one aspect all will agree: the ~400-reaction network that converts H2, CO2, NH3, H2S and phosphate into amino acids, bases and cofactors cannot have arisen in an instant.

"Its emergence from spontaneous environmental reactions had to traverse intermediate states of assembly, which have previously been elusive."

The researchers investigated the protein structures of core metabolic enzymes in the genomes of bacteria and archaea, and developed a method and algorithm to order them in complexity. From there, they could determine which ones evolved in which order, and put together a rough timeline.

Radical Study Suggests Life on Earth Arose From Non-Living Matter Twice

Starting compounds are shown at the left; they are converted by metabolism into the building blocks of life. The 420 enzymatic reactions are indicated as circles, and chemical metabolites as diamonds; lines connect reactions that share metabolites. Circles shown in magenta shading indicate reactions that could have been catalyzed by inorganic compounds in the environment where metabolism of the first cells arose. (HHU/Nadja Hoffmann)

"We found that the last universal ancestor of all cells, LUCA, possessed enzymes for only about half of the reactions of metabolism," says Martin.

"The other half was catalyzed by metals in the environment where LUCA arose."

The team found that there were likely four different phases in the development of catalysis (enzyme-driven reactions). The first relied solely on metals in the environment.

Once enough of the useful molecules built up, the not-quite-living-yet proto-cells (which included LUCA) could start developing their own enzymes that performed some of the functions of those metal catalysts.

Over time, later proto-cells developed more and more enzymes, reducing their dependency on external metals, until they no longer needed them at all. Only at that point would they be considered 'free-living cells'.

Importantly, the researchers say, this didn't occur until after bacteria and archaea had separated. These two branches figured out their own ways to tackle the same problems.

"We can see cases where the ancestors of bacteria and archaea independently evolved structurally distinct enzymes to catalyze the same essential metabolic reaction," says Natalia Mrnjavac, a biologist at Heinrich Heine University Düsseldorf.

"Such parallel inventions could have paved the way to the independent emergence of free-living bacteria and archaea."

The study may also plug another gap in our understanding of the origins of life.

A molecule called ATP is now a vital energy source to fuel metabolism – but it's also made by enzymes, and wasn't available in prebiotic chemistry. The researchers say they identified an alternative.

"When we react phosphite, a form of phosphorus that naturally occurs in hydrothermal vents, with organic compounds, we get metabolic phosphorylation reactions overnight in water," says Manon Schlikker, a molecular evolutionary scientist at Heinrich Heine University Düsseldorf.

"Phosphite and palladium replace ATP and enzymes; it's amazing, and it makes early evolution a lot easier to grasp."

Yet critical to chemical reactions is the medium in which they occur, and scientists still don't have a clear idea about that. Was it water, or a sticky goo?

If the findings are backed up with further research, it seems that the tree of life, as we picture it, may need an update. We tend to think of LUCA as the 'trunk' of the tree, with all the other lifeforms, past and present, branching off from there.

But perhaps there were two completely separate trunks, which first diverged at the roots, before they were even technically alive.

The research was published in the journal Science Advances.

The Daily Front Page 11 of 30
Friday, August 7, 2026 The Daily Front No. #260807 — Putting a Meter on AI Coding
article

Managing AI Coding Costs at Scale

by moonikakiss·▲ 292 points·250 comments·databricks.com ↗
Nearly every company deploying AI tools at scale has hit the same wall: exponentially growing costs.

AI coding tools deliver immense value: at Databricks, agentic coding has measurably improved every velocity metric we track and, in some teams, driven an order-of-magnitude gains in output. But nearly every company deploying AI tools at scale has hit the same wall: exponentially growing costs. That curve is unsustainable - left unchecked it will eventually overtake revenue. The spend explosion has left enterprises in a paradoxical situation: on the one hand, desiring to maximally push AI transformation and put powerful tools in the hands of employees, and on the other hand, having to reconcile with an aggregate cost profile that threatens to undermine or even reverse the very efficiency gains AI provides.

Fortunately, several of the earliest large-scale adopters have converged on a set of approaches that solve this puzzle, achieving a “dual mandate”: (a) providing broad access to AI tooling, with minimal friction, and (b) keeping aggregate costs inside of a roughly fixed envelope per user. This post outlines proven cost management techniques, based on our experience at Databricks and conversations with several other digital-native companies, including Stripe, Coinbase, Uber, and Ramp. The table below summarizes current techniques and associated savings; the numbers are directional, based on an informal survey of development teams:

Some of these techniques can be easily implemented with software many companies already use. Others require new infrastructure, particularly techniques that modify end-user clients or shift traffic across models. At Databricks, we’ve open sourced or made freely available our key infrastructure components: an end user meta-harness (Omnigent) and our AI Gateway (Unity AI Gateway). For completeness, this post also covers software used by other companies we spoke with.

The “Efficiency Frontier” for Coding Models

The single greatest cost lever in moving coding spend to more efficient models as they are released. This point bears some discussion, as the simple explanation of "cheaper models” in fact hides a nuanced relationship between model cost and quality.

Colloquially, the term frontier model means “the highest intelligence model,” and frontier labs largely focus on advancing peak intelligence. Frontier models can now solve novel problems in math or cybersecurity. But when AI is deployed at scale, a different type of frontier matters more: the efficiency frontier. The efficiency frontier is defined by the set of models that have the best price point for a given level of intelligence. Most day-to-day coding doesn't require mathematical proofs or novel security insights, so what matters in aggregate is the cost of models that meet the quality bar for typical software engineering work. This "efficiency frontier” is advancing far faster than the intelligence frontier, with new models being released almost weekly that present better intelligence-per-unit-price than prior models.

Cost Lever #1: Moving to open source and lower cost models

Rapidly adopting newer, more efficient models delivers the largest cost wins of any technique. But to capture those gains, a company first needs to know which models actually beat its incumbents. This can be difficult because public benchmarks do a poor job of indicating real-world performance on coding tasks. To size up new models, many companies have built automated evaluations that they believe are more representative of their internal development mix. Databricks recently published an example of such a benchmark, in which we observed highly competitive price/performance for GLM models. That benchmark led us to roll GLM out to developers internally. Often, new models do not advance the efficiency frontier,and evaluations frequently produce negative results: Stripe found that Opus 4.7 did not meaningfully improve quality over Opus 4.6, while increasing cost. They therefore declined to make Opus 4.7 available internally. Databricks saw similar cost regressions when comparing Opus 5.0 to 4.8.

Harness and Model Flexibility

Since the biggest wins come from switching to new models, adopting end user tooling that allows for model flexibility is becoming a critical component of keeping costs down. The tool most commonly used in concern with a particular model is called harness. Proprietary frontier models are increasingly co-designed to work well with specific harnesses, meaning certain harnesses “work better” with certain models. If a company wants to preserve model independence there are roughly two approaches:

Ask users to switch harnesses. One approach is to provide developers with a set of harnesses (Claude Code, Codex, or Cursor) and then ask them to switch between harnesses when a company wants to migrate spend to lower cost models. This lets users work in their preferred harness when possible, but the downside of this approach is that switching costs for an individual developer can be high. If switching costs become too high, the harness itself becomes a de facto lock-in to a model family, limiting the ability to move spend to more competitive models.

Use a meta-harness. A new and increasingly popular approach is to use a meta-harness that surfaces a common user experience to developers while dispatching requests to underlying harnesses (both proprietary and open source). This approach allows both model/harness independence while also reducing developer switching costs. At Databricks, this is the default mode for developers who leverage Omnigent. Some companies we talked to have built custom internal meta-harnesses that integrate with their development toolchain.

Cost Lever #2: Dynamic Request and Task Routing

Instead of asking users to choose task-appropriate models themselves, a growing body of research suggests that automatic model and tool selection may further squeeze efficiency out of agentic coding workflows. Routing approaches roughly fall into three categories:

  • Request Level Routing: A stateful proxy sits in between a client (such as a coding harness) and the underlying foundation models. The proxy attempts to route requests to the lowest-cost model capable of answering each inference request. Routing for agentic use cases also needs to account for server-side caching, since a cold cache hit has a very high cost for large context workloads. A new wave of products is showing early, promising results for routing. Examples are: Cursor Router, OpenRouter’s AutoRouter, Ramps Router feature and Databricks own Smart Routing feature in Unity AI Gateway.
  • Task Level Routing (Meta Harness): A client-side process dispatches user tasks to different harnesses based on the complexity of the task. A user task might be “rename this component from X to Y” (a simple task) or an open-ended task like “Explore design considerations that would reduce latency” (a complex task). The dispatcher, often called a Meta Harness, examines which level of underlying model is required for a task and then delegates that entire end-to-end task to the model. Omnigent is an example of a Meta Harness that supports this pattern.
  • Escalation/Delegation Patterns: A single harness pairs two models (an expensive, high-intelligence model and a cheap worker model). In some approaches, such as Claude’s Advisor Tool, the cheaper model runs the show and escalates when it thinks a task requires more horsepower. The inverse pattern also exists: In Cognition’s Devin Fusion, the higher cost model is the main loop, and it selectively outsources work to a cheaper model.
    Internal results at Databricks suggest that our AI Gateway Smart Router is able to consistently reduce average task cost by more than 30%, while roughly matching the quality of the most expensive model in the working set. Other companies we spoke with have seen similar results.

image9.png

Cost Lever #3: Giving developers visibility, tripwires, and budgets

It may be surprising that this entire article did not start and end with “Give users a monthly budget and be done with it.” Hard budgets, where usage is entirely cut off at a specific spend threshold, are often used only as a last resort option in every company we spoke with. There are two reasons that hard token budgets are not particularly effective for AI spend management: First, if a developer hits their budget ceiling, cutting off further access to AI tools would be debilitating to productivity. Neither the company or employee actually wants that outcome. Second, at least some of the “high spending” users are in fact those who have achieved monumental efficiency gains with AI and are producing immense output. Discouraging those users is self-defeating.

Instead of a hard user spending cap, most companies are adopting a more nuanced and progressive approach that focuses on visibility for end users and increased degrees of friction as spend increases.

  1. Visibility: Every company we spoke with had a mechanism to provide near-instantaneous feedback to users on their ongoing spend, with many also offering specific tips or insights on how to reduce spend by using less expensive models. It is important that users be able to see their spend across all tools, since they may want to influence their choice of tool where they get the highest ROI.

    Managing AI Coding Costs at Scale

    A developer dashboard at Databricks showing active spend

  2. Spend Gates: Developers can be asked to take actions or seek approvals at increasing levels of spend. The simplest form of spend gate is one that can be self-cleared and serves as a warning that the spend rate is increasing above some threshold. At Databricks, we’ve found self-clearing gates a useful mechanism for preventing accidental or unintentional spend. Further gates can be introduced that require explicit budget approval (often through a management chain).

  3. Downshifting: If a developer has hit a spend gate, they can be downshifted to a lower-cost model rather than being entirely suspended from token access. Since the lowest-cost models are drastically less expensive than frontier-intelligence models, this technique allows developers to continue getting work done without incurring massive ongoing spend.

  4. Suspension: In the limit case, most systems do retain the ability to fully suspend users from all token access. As stated above, this is often a temporary measure only and the starting point for a conversation about how to efficiently leverage AI.

Cost Lever #4: Reducing Token Overhead

When a user types a relatively simple request into an AI coding agent (such as “Please investigate and fix this bug.”), that agent subsequently gathers massive amounts of relevant context, invokes a large number of tools, searches through the codebase, and integrates skills or system information provided by the company. By the time costly LLM inference occurs, the user's initial statement accounts for only a negligible fraction of the data fed into the AI system, meaning costs are dominated by context the user did not explicitly include. Techniques in reducing context bloat are still new, but several promising approaches are being explored, such as:

  • Coercing more frequent compaction (compression) of the active context.
  • Using harnesses that are “less chatty” (more token efficient), or tuning existing harnesses to generate less token overhead.
  • Auditing popular tools and decreasing their verbosity.
  • Encouraging developers to break tasks into smaller individual units of work, decreasing context scope.

When contexts get large, prompt caching also plays a meaningful role in overall performance. Both proprietary and open source LLMs have settings that allow you to enable prompt caching and tune how long the cache is stored. Cache writes cost money, but cached reads can drastically reduce per-inference cost. This trade-off is dependent on a company’s specific workload, so hand-tuning of default cache settings to increase overall cache hit rate can have drastic improvements to overall cost.

At Databricks, relatively simple tuning of our harness and caching settings led to an almost 50% reduction in the number of generated tokens and associated costs, with no observed quality degradation for developers. We continue to explore techniques in this area and think meaningful additional optimization remains possible.

A drastic reduction in tokens per session by eliminating extraneous inference calls and reducing cache writes.

The AI Gateway design pattern

The techniques above had many implicit technical requirements: To rapidly take advantage of new models, companies must have a central location where the “model menu” is managed, and end-users must have a toolchain that supports model mixing. To provide budget visibility across multiple AI tools, a unified cost observability capability must exist. To manage context bloat, companies need a way to observe typical toolcall outputs and enforce compression or compaction. These needs are collectively being solved by a new class of infrastructure software, best described as an AI Gateway. An AI gateway is a central location where all of the following occur:

  1. Capacity management and proxying of access to underlying models (both proprietary and OSS models).
  2. Budget tracking and enforcement, including complex budget policies such as progressive friction levels and model downshifting.
  3. Configuration management for end-user tools, to enforce model allow-lists, compaction settings, and other locally mediated aspects.
  4. Logging of coding session traces for downstream efficiency analysis and benchmark.

At Databricks, we rely heavily on Unity AI Gateway for all of these capabilities.

Putting it all together

The exponential growth of AI coding costs is not an inevitability, it's a solvable engineering and governance problem. Companies that have tamed it share a common playbook: relentlessly chase the efficiency frontier rather than the intelligence frontier, adopt tooling that preserves model flexibility, route work intelligently to the cheapest capable model, replace hard budgets with visibility and progressive friction, and cut the token overhead that dominates real-world spend. None of these techniques requires sacrificing the productivity gains that made AI adoption worthwhile in the first place; together, they let organizations satisfy the dual mandate of broad, low-friction access within a predictable cost envelope.

A set of new infrastructure abstractions is emerging to give companies the tools to manage their costs. At Databricks, we’ve released the key components in our cost management stack as open source or free software products: Our Unity AI Gateway for central management and Omnigent for developer tooling. Thousands of companies use these components every day. We invite more companies to share findings and compare techniques as this technology landscape rapidly evolves.

Acknowledgements: Thank you to infrastructure leaders at Uber, Stripe, Coinbase, and Ramp who provided commentary and reviews of this article. Thank you to Thrive Capital for feedback on an early draft of this article.

The Daily Front Page 12 of 30
Friday, August 7, 2026 The Daily Front No. #260807 — The Engineer’s Bookshelf
article

Carl's Required Reading

by cckolon·▲ 220 points·29 comments·carlkolon.com ↗
Most articles are on this list because they make a point that I think is important or insightful about how to create good software.

As an engineering leader, I often send articles to my team, which I jokingly call “required reading”. My goal is to help us digest some of the wisdom of other programmers who have probably faced challenges similar to ours. By popular demand, I’m putting the list here for anyone else interested!

Most articles are on this list because they make a point that I think is important or insightful about how to create good software. Some of the points in here reinforce my opinions about certain things (for example, ORMs are bad, and frontends should be simple). You may disagree! But at least you will hopefully get something valuable out of the article, even if it is that you feel the opposite way.

These articles are technical, and if you are not a programmer you will probably not find them interesting.

This list may be a little overwhelming, so I have annotated my favorite articles with a star. I have also written a brief summary of why I think each article is important.

Coding practices

The Grug Brained Developer

If I could recommend only one article to new and experienced programmers alike, this would be it. Complexity bad.

The Wrong Abstraction

It’s really tempting to see code duplication and immediately rush to eliminate it by consolidating it somewhere. This essay is adapted from a 2014 talk about factoring object-oriented code and talks about situations where this may not be appropriate. You need to understand this to effectively push back against coding agents which are motivated to create a “single source of truth” everywhere.

Complexity Budget

When projects get too complex, progress seems to stop very suddenly. This essay is about understanding this phenomenon and trying to forestall it as long as possible.

Locality of Behavior

Programmers (and coding agents) often try to achieve Separation of Concerns (SoC) or eliminate code reuse by creating numerous helper functions. There’s a tradeoff between this and Locality of Behavior, a principle which encourages us to make the behavior of code as obvious as possible when looking only at that one unit.

Yagni

Creating speculative features with the future in mind is basically always a bad idea.

Parse, Don’t Validate

This one takes a little work to understand, but the basic idea is that if you ensure that a piece of data has a certain property, you should encode that property directly in the object’s type. This will cause code changes that break the assertion to throw type errors at compile time, rather than value errors at runtime.

Platform

Steve Yegge’s Google Platforms Rant

There is a ton of really good stuff in here, especially about accessibility and setting up software organizations. I do not recommend every company enforce the Bezos mandate, but you should think critically about whether you can expose your service’s features to other teams in a programmatic way. It’s also just fun to read.

Frontend

HATEOAS

In web applications, state is often encoded and stored separately from both the frontend markup and the backend (for example, with React’s useState). Often, it makes more sense for state and the user’s allowed actions to be directly stored in and derived from the html served to the user. This is tightly related to the original concept of a REST API (most modern “REST” APIs do not actually follow the REST constraints). I have also written briefly about HATEOAS on my own site.

Components and Hooks must be pure

The biggest problem I see when people move from writing backend code to writing React is that they try to write the frontend imperatively. That is, they try to tell the computer what to do, line by line. This typically involves the overuse of state and side effects, common in object-oriented programming. Instead, the frontend should be declarative. That is, you should tell the computer what you want to get back. To support this, (modern) React is designed around a functional programming style. Every component is a function, and state and side effects are avoided except for where they are truly necessary, in which case hooks are required. I recommend reading all the React docs, but if you just want to focus on one, this is the best.

Understanding useMemo and useCallback

useMemo and useCallback are the most misused hooks in React (besides maybe useEffect), and both agents and humans love to throw them in everywhere. This article is a guide to where they are actually necessary.

Hypermedia Systems - Components of a Hypermedia System

If you want to write a good frontend, it is very valuable to understand the model around which HTML is based. This is a great chapter of a great book which will help you maximize your use of the browser’s design, rather than seeing HTML as “an awkward, legacy markup language that must be grudgingly used to build user interfaces in what are increasingly entirely JavaScript-based web applications.”

Databases

The Vietnam of Computer Science

I am a certified ORM-hater. I think they are seductive for new projects, but quickly begin to cause performance issues and confusion. As an example, see this OpenAI article about scaling Postgres:

Many of these problematic queries are generated by Object-Relational Mapping frameworks (ORMs), so it’s important to carefully review the SQL they produce and ensure it behaves as expected.

This is the best essay against ORMs that I have read and, even though it’s long, I recommend it highly.

Wikipedia - The Object-Relational Impedence Mismatch

Some more ammo in my anti-ORM crusade. This is more technical but more succinct, so if you’re just looking for the bullet points I’d read this.

Introduction to PostGIS - Geography

PostGIS (and geospatial databases) require a little getting used to. It’s good to understand the new data types you’re working with when you build a map-based data visualizer. PostGIS is a great database and its foundational data type is geography1.

Postgres Docs Chapter 14 - Performance Tips

This is a great article to have read when your database starts slowing down. Knowing how to use EXPLAIN and EXPLAIN ANALYZE is on par with knowing how to use a debugger, in my opinion.

Paging Through Results

Most people use LIMIT and OFFSET when building pagination systems. If you are paginating through a lot of data, these get slow (especially when loading high-number pages). This is a good overview of some other options to get around this problem. I also talk about this on my site.

Asynchronous programming

A Conceptual Overview of asyncio

Many people are just thrust into async programming and don’t really understand what’s going on, but learn it as they go. Often this leads to embarrassingly basic asyncio mistakes like making blocking calls inside async functions. This is a good article from the docs which should make it clearer what asyncio is doing under the hood, and therefore teach you how to use it best.

What Color is your Function?

Bob Nystrom is one of my favorite programming authors. This is an example of how async programming systems often have fundamental flaws. There’s not much you can do (besides use a different language) but this will at least teach you that some of the async programming limitations that you hit are fundamental, and not just a skill issue.

Encoding

The Absolute Minimum Every Software Developer Absolutely, Positively Must Know About Unicode and Character Sets (No Excuses!)

Reading this earlier in my career would have saved me hundreds of hours incorrectly parsing data collected from the internet.

Books

These books have shaped the way that I feel about programming, design, and managing software projects.

The Design of Everyday Things (Amazon)

Lots of people seem to think that “design” means making things look pretty. Really, design is about anticipating the needs of your user and making your product fulfill those needs as well as possible. Often this second meaning of design runs contrary to the first, and in these cases you should choose the second. The early example of “Norman Doors” in the book is a great illustration of this, along with the wry observation that the doors “probably won a design prize” even though the author can’t figure out how they work.

Designing Data-Intensive Applications (O’Reilly) (Amazon)

Liking this book is such a meme, but it actually is really great. It’s also one of the few programming books that works well via audiobook. My favorite chapters are 7 (Transactions) and 10 (Batch Processing). While some may disagree, I think this is a great introduction to databases too, and makes you think hard about what you actually want to optimize in your data system (even if you are not serving a zillion users).

Crafting Interpreters (free online)

This is a great, encouraging, fun book which teaches you how to build an interpreter, first in Java for simplicity and then in C for performance. While you can read along and write the Java code directly, I think it’s even better to try implementing the interpreter in another language of your choice, so you have to engage your brain more. I used rust.

Category Theory for Programmers (free online) (order hardcover)

While this book is a little hardcore, it’s the best intro to category theory that I’ve seen (though if you’re a fan of rigor, you may want to google some of the formal definitions yourself). Read this if you want to write good functional code while also flexing on your coworkers by using words like “functor” and “monad”.

The Mythical Man-Month (Amazon)

This book was published in 1975 (!!) and revised most recently in 1995, but it is still incredibly relevant. In fact, as programmers become armed with AI tools, some of the chapters (like chapter 4, Aristocracy and Democracy in System Design) seem more applicable now than they were in 1995. Read this if you want to manage a programming project and have it succeed.

Footnotes

  1. This is actually not correct, as seabre pointed out on Hacker News. The foundational data type in PostGIS is geometry, which you can read about here. It’s still good to learn about geography though, and the intro page is a little more accessible. 
The Daily Front Page 13 of 30
Friday, August 7, 2026 The Daily Front No. #260807 — A Language for Distributed Safety
show hn

Show HN: Wyzer Programming Language

by v0id_isgood·▲ 210 points·109 comments·github.com ↗
Simplicity is not the absence of power. It is power without pretense.

Wyzer is a statically typed, compiled, resource-oriented programming language with integrated distributed safety via choreographic programming and a perceus memory model.

"Simplicity is not the absence of power. It is power without pretense." ~ Atiksh Sharma

What is Choreographic Programming? In traditional distributed programming, you write separate code for the client and the server, and hope their communication protocols match up. Choreographic programming allows you to write a single, unified view of the entire distributed system. The compiler then mathematically projects this single script into deadlock-free, independent binaries for each physical node (e.g., Client and Server).

What is the Perceus Memory Model? Perceus is a fast, deterministic memory management strategy that avoids the unpredictable pauses of a Garbage Collector and the complex annotations of a borrow checker. It uses precise reference counting to know exactly when a value has a single owner, allowing the compiler to mutate the data directly in place (FBIP) without making copies or pausing the program.


Design Principles

  • One ownership rule for everything: Manage memory, threads, and network resources using the exact same ownership model.
  • No garbage collector: Predictable, low-latency performance using Perceus reference counting.
  • No borrow checker: Write safe functional-style code without fighting complex lifetime annotations.
  • Deadlock-free by design: Network choreographies are verified and generated by the compiler.
  • Avoid hidden magic: Network transfers happen through explicit assignment, keeping the control flow readable and obvious.
  • C ABI compatibility: Built for systems programming and seamless interoperability.

Quick Example: Distributed Key-Value Store

In Wyzer, you write one program that describes how multiple systems talk to each other. The compiler automatically infers network sends and verifies ownership across nodes.

role @Client;
role @Server;

struct KVRequest {
    op: u8, // 0 for Get, 1 for Put
    key: str,
    value: str
}

fn main() {
    // Lock-free state owned securely by the Server
    var _store_key: str@Server = "";
    var _store_value: str@Server = "";

    // Client creates a Put request
    let put_req: KVRequest@Client = KVRequest { 
        op: 1, 
        key: "username", 
        value: "alice" 
    };
    
    // It is safely sent across the network by transferring ownership
    let server_req: KVRequest@Server = put_req;

    // Server updates its state without needing mutexes
    if server_req.op == 1 {
        _store_key = server_req.key;
        _store_value = server_req.value;
    }
}

Because Wyzer merges Choreography with Perceus linear typing, a network transfer acts as an absolute linear move. Transferring ownership of a variable across the network consumes it locally. Attempting to use it again produces a compile-time error:

Error [E001]: use of moved variable `put_req`
╭─▶ example.wyz:31:22
│
│ 30 │     
│ 31 │     std::io::println(put_req);
│    │                      ╰────── variable `put_req` used here after being moved
│ 32 │ }
│
├─▶ note: `put_req` is a linear resource and can only be used once
├─▶ help: consider passing it by reference if you need to use it multiple times
╰─────────────────────────────────────────────────

Example: Request/Response Choreography

Choreographies make complex network handshakes read like a single sequential function. Here, a client queries a server, and the server returns the result back to the client:

fn fetch_data(query: str@Client) -> str@Client {
    // client transfers query to Server
    let server_query: str@Server = query;
    
    // server processes query securely
    let server_result: str@Server = db_lookup(server_query);
    
    // server transfers result back to Client
    let client_result: str@Client = server_result;
    
    return client_result;
}

Example: In-Place Mutation (Perceus)

Wyzer uses the Perceus memory model, which determines variable lifespans via precise reference counting. If a variable is uniquely owned (refcount == 1), the compiler safely mutates the data in-place rather than allocating new memory (known as FBIP: Functional But In-Place).

struct User { id: u32, name: str }

fn update_name(user: User, new_name: str) -> User {
    // In traditional functional languages, this allocates a new 'User' struct.
    // In Wyzer, if 'user' has exactly 1 owner, the compiler mutates the existing 
    // memory in-place, achieving C-like speeds without a borrow checker!
    return User {
        id: user.id,   // Value retained
        name: new_name // Value overwritten in-place
    };
}

Documentation & Community

Community: Join our Discord server: https://discord.gg/RhpPhkTrVu

Notable Examples

The Daily Front Page 14 of 30
Friday, August 7, 2026 The Daily Front No. #260807 — The Cable Drawer Audit
article

I stopped trusting USB-C cable labels and started testing them

by baranul·▲ 258 points·256 comments·makeuseof.com ↗
USB-C was supposed to be the single connector to rule them all.

A USB testing tool

If you're anything like me, you probably have a drawer full of USB-C cables. USB-C was supposed to be the single connector to rule them all. Capable of powering everything from portable fans to high-powered laptops, and reversible to boot. But its versatility ushered in a new problem.

Gone are the days of trying to plug a USB-A cable in three times before you find the right orientation. But it's been replaced by a game of chance. That drawer of cables is an exercise in guesswork as I try to find one capable of charging my 16-inch MacBook Pro. Some sort of label would be nice.

But it doesn't have to be this way. Everything changed when I got my hands on a USB cable tester, and I haven't looked back.

Identical cables, completely different uses

Unknown data speeds and charging power

The problem with using a single cable connector for everything is that all those cables inevitably look the same. The cable that came free with your earbuds looks just like the one that came with your laptop and can handle over 100W of power. Your no-brand 480Mbps USB 2.0 cable is indistinguishable from the super-fast USB 4 cable that can move 40 gigabits (around 5 gigabytes) of data every single second. And that's before we get into the way high-speed Thunderbolt cables use the same connector — that's for another time.

These cables only look the same, of course; they are very different on the inside. More capable cables tend to be thicker, albeit often imperceptibly so. There's also the e-marker chip. Think of it as the brains of the cable. It's the e-marker chip that tells your devices what the cable is actually capable of, whether that's charging power or data transfer abilities. Any cable that can charge at more than 60W requires an e-marker chip to do so. Any cable that can transfer data faster than USB 2.0 speeds also requires a chip to do so.

That e-marker chip is vital. Even if the cable itself is capable of charging at more than 60W and transferring data at faster than USB 2.0 speeds, it won't do so without that chip. But you can't tell if an e-marker chip is present just by looking at a cable. So while it's good information to know, it won't help you find the right cable come charging time.

For that, you're going to need $15 and an Amazon account.

A bargain at thrice the price

The perfect tool for the job

The USB-C port on the Honor Magic V6

The answer to my charging problem came in the form of a $15 cable tester. Amazon has tons of them to choose from, and you can spend much more and much less than I did. But I just needed to know how fast a cable could charge my stuff, so I didn't need anything flashy.

A simple USB tester sits between your cable and the device you want to charge. Just make sure to plug your cable into a high-powered USB-C power adapter to remove that variable from your testing. Then check out the tester's display to get that sweet, sweet data.

Most budget-friendly testers will give you three main bits of data. The voltage (V) figure shows whether your device negotiated the standard 5V charging power or was able to upgrade to any of the 9V, 15V, 20V, or higher Power Delivery modes. Next, amperage (A) measures the flow of current through the cable and into the USB tester and device. Finally, the wattage (W) figure shows you the real-time power delivery. It's this figure that you're really interested in.

Testers more capable than mine can go a step further and read the data directly from the e-marker chip itself. This comes with the added bonus of also telling you what the cable is rated for in terms of data throughput as well as charging speed. I just wanted to know which cables would charge my devices the quickest, but I can definitely see how knowing a cable's data transfer speed would be useful.

YEREADW USB C Tester Power Meter

YEREADW USB-C Tester Power Meter

The YEREADW USB C Tester Power Meter (KWS-2303C) accurately measures voltage, current, and power for USB-C devices, helping diagnose charging issues and verify cable or power bank performance quickly and easily.

See at Amazon

Stop guessing, start testing

Perfectly paired with a label maker

USB C to C cables

My USB cable tester and I have become best buds of late. I set aside some time to go through a few cables every couple of days, test them, and then make a note of their capabilities. But the tester is only part of the equation here. I still wanted to make it quicker and easier to find the right cable when I needed it, so I dug out my label maker and went to town.

Now, I have a drawer full of USB-C cables, each with a label wrapped around it to tell me what it can do. I know one cable can fast-charge my laptop, while another is better suited to charging my AirPods Pro 3. It was a pain to do, sure, but you bet I'll be putting labels on each new cable that comes into our home from here on out.

But I shouldn't have had to do any of this. It's time cables came with a label of their own, or some sort of color coding. But until that happens, I'll be sitting over here. USB tester and label maker in hand.

The Daily Front Page 15 of 30
Friday, August 7, 2026 The Daily Front No. #260807 — The Open Gardens of Estonia
article

Why Estonians invite strangers into their back gardens each summer

by koolhead17·▲ 153 points·54 comments·bbc.com ↗
Each summer, home gardens, garages and even bedroom windows across Estonia become temporary cafes open to anyone.

Martin Mark People sitting in a garden at picnic tables with a wooden farmhouse in the background (Credit: Martin Mark)

Each summer, home gardens, garages and even bedroom windows across Estonia become temporary cafes open to anyone.

Beneath a giant tent, complete strangers queue to buy lunch. Around them, vegetable beds overflow with cucumbers, strawberries, garlic and courgettes, while an apple tree shades a table laden with rye bread, homemade pickles and a simmering pot of seljanka, a tomato-based soup with ham and smoked sausage. Guests order from a simple printed menu before carrying bowls of steaming soup to communal tables to eat.

At first glance, it looks like a country cafe. It isn't. It's Heiki Kalvik's back garden in the outskirts of Kehra, a small town about 45 minutes east of Tallinn, which he turns into a pop-up cafe each summer.

"I'm very outgoing, and I like to cook. I enjoy hosting, and I’m proud of our outdoor space," Kalvik says as another group of visitors arrives.

Kalvik is one of around a dozen residents taking part in Kehra’s annual Kodukohvikute Päev, or "home cafe day", when locals turn private gardens, courtyards and garages into temporary eateries for one day in July. Elsewhere in town, Karolin Vahtre has converted a garage behind her Soviet-era apartment block into a makeshift cafe, selling tempura sushi rolls, chanterelle quiche, chocolate cake and French fries to anyone who stops by.

Celine Senseney Karolin Vahtre makes and serves tempura sushi rolls from her one-day garage café in Kehra (Credit: Celine Senseney)

Karolin Vahtre makes and serves tempura sushi rolls from her one-day garage café in Kehra (Credit: Celine Senseney)

More than 150 such events are held across Estonia between April and September, a beloved tradition that has spread steadily since the first event almost two decades ago. Some are elaborate, with rented tents, tables and chairs; others are little more than a lemonade stand manned by children. Together, they give visitors a rare glimpse of everyday Estonian life that is usually out of reach.

How to get involved

When to go: Cafe days take place across Estonia from spring to early autumn, with many concentrated in July and August. Dates vary by town and region.

How to find one: Check local tourism websites and search for 'Kohvikutepäevad' ("cafe days") on Facebook or in local event listings. Estonia's regional tourism sites also publish upcoming events.

What to expect: These aren't conventional cafes or food festivals. Residents turn gardens, courtyards, garages, or other private spaces into temporary cafes, serving whatever they choose.

How Estonia's cafe days began

Estonia's home cafe tradition began on Hiiumaa, the country's second-largest island, in 2007 when local tourism promoters were looking for a way to attract visitors beyond the peak month of July.

They drew inspiration from Kärdla, the island's capital, and its long association with coffee. In the mid-1800s, the town was home to a broadcloth factory run by a Baltic German baron, who employed coffee-drinking workers from Central Europe. Local Estonians soon adopted the habit, with Kärdla residents eventually earning the nickname kohvilähkrid, or "coffee flasks", for taking coffee into the fields instead of the traditional sour milk or beer.

Fifteen homes took part in Hiiumaa's 2007 cafe days; today, the three-day event in early August attracts thousands of visitors, with dozens of temporary cafes selling homemade food and drinks across the island.

Hiiumaa Museum is one of the few participants to have taken part every year since 2007, re-creating a 19th-Century cafe in its conference rooms at the Long House, the museum's main building and the former home of the broadcloth factory's managers. Its director, Kauri Kiivramees, says the tradition taps into something deeper than food and drink: the islanders' long history of welcoming outsiders.

Martin Mark Menus vary from cafe to cafe, with some serving local specialties and others offering family favourites (Credit: Martin Mark)

Menus vary from cafe to cafe, with some serving local specialties and others offering family favourites (Credit: Martin Mark)

Kiivramees says even this so-called "island of introverts" has always understood the importance of community: "It's quite traditional to help foreign visitors on Hiiumaa because back then it meant they were here because of [a] shipwreck or something like that, and they needed help."

From island tradition to national phenomenon

Small, rural communities have historically been the backbone of Estonian society on both the country's islands and mainland. While people aren't as reliant on their neighbours as they once were, certain traditions, like the talgud, have endured. These communal work days once involved haymaking or barn raising; today they are more commonly used for community clean-ups. Home cafe days reflect that same spirit of neighbourliness, and after their initial success on Hiiumaa, they soon spread across the country.

Setomaa in south-eastern Estonia, holds its pop-up cafe days every year on the second weekend in August, using the event to showcase its distinctive culture.

It is not so often you can see inside of our yards. They are closed. But this is the day when our people open their hearts and their gates – Katrin Aedma

"One of the reasons why we started [Kostipäiv] is because we saw that during the summertime we don't have so many events where people are wearing traditional clothes or offering traditional food like sõir, our homemade cheese," said Liis Kogerman, board member of the organising NGO Seto Küük.

Martin Mark Setomaa's home cafe days showcase a culture found nowhere else in Estonia (Credit: Martin Mark)

Setomaa's home cafe days showcase a culture found nowhere else in Estonia (Credit: Martin Mark)

The event has become one of the region's calling cards, with some cafes attracting as many as 1,000 visitors. While one has since become a permanent restaurant, most participants simply enjoy selling homemade food to visitors without the pressures of running a business. 

The same spirit can be found in the university city of Tartu, where Inga Kulmoja runs a one-day pop-up cafe to indulge her passion for homemade ice cream. Kulmoja lives in an old wooden house with no yard, so each year she lowers ice cream down to the street from her third-storey bedroom window.

Most Estonians under 40 speak English, but tere (hello), palun (please), and aitäh (thank you) will go a long way, especially in smaller towns and villages. Estonia may be a digital country, but it's best to bring cash.

"It's very low tech," she said, noting that running the cafe is quite different from her daily work as an entrepreneur in the AI space. "Only cash, and the signs are written on reused paper. We bought some rope, and we have two buckets: one for the ice cream, and the second smaller bucket is for cash. I'm not even using an ice cream machine; I'm just making them in the freezer."

Even though Kulmoja is mostly busy dipping ice cream and filling the buckets, she still likes to see visitors queued up, especially when the same people come year after year. "I think people have some kind of willingness to show what they can offer or how they live," she said. "In Estonian, you would say maybe 'edevus'. It's like the feeling you get being the centre of others' attention." 

Martin Mark Estonia's home cafes offer travellers an unusually intimate window into local life (Credit: Martin Mark)

Estonia's home cafes offer travellers an unusually intimate window into local life (Credit: Martin Mark)

When strangers become neighbours

Oleg Pun and Iryna Topilina moved to Estonia from Ukraine in 2019, but didn't host their first cafe until 2022, when Pun's parents moved to Estonia to escape from the war. Pun and Topilina encouraged them to participate in their new village's cafe days to learn more about the community. 

"They were so surprised and moved by how many Estonians came to support them. Neighbours from nearby houses even brought over extra chairs and furniture when they saw we needed more seating," recalled Topilina.

Now the family works together to run a cafe during Kalamaja Days in Pun and Topilina's garden. This year they served cabbage varenyky, Ukrainian dumplings made from a family recipe that goes back almost 100 years, and green borscht made with locally foraged sorrel and nettle and cooked over an open fire.

Eleryn Meister Some hosts keep things simple, while others transform their gardens into full outdoor restaurants (Credit: Eleryn Meister)

Some hosts keep things simple, while others transform their gardens into full outdoor restaurants (Credit: Eleryn Meister)

While Estonia's home cafe days may feel like secret celebrations to outsiders, they are one of the few occasions when visitors are invited to wander through private gardens, taste family recipes and glimpse everyday life inside local communities.

As Kalamaja Days organiser Katrin Aedma said, "It is not so often you can see inside of our yards. They are closed. But this is the day when our people open their hearts and their gates."

The Daily Front Page 16 of 30
Friday, August 7, 2026 The Daily Front No. #260807 — Inside the Inference Engine
article

Inside vLLM: Anatomy of a High-Throughput LLM Inference System (2025)

by sebg·▲ 146 points·10 comments·aleksagordic.com ↗
I’ll gradually introduce all of the core system components and advanced features that make up a modern high-throughput LLM inference system.

From paged attention, continuous batching, prefix caching, specdec, etc. to multi-GPU, multi-node dynamic serving at scale

In this post, I'll gradually introduce all of the core system components and advanced features that make up a modern high-throughput LLM inference system. In particular I'll be doing a breakdown of how vLLM [1] works.

This post is the first in a series. It starts broad and then layers in detail (following an inverse-pyramid approach) so you can form an accurate high-level mental model of the complete system without drowning in minutiae.

Later posts will dive into specific subsystems.

This post is structured into five parts:

  1. LLM engine & engine core: fundamentals of vLLM (scheduling, paged attention, continuous batching, etc.)
  2. Advanced features: chunked prefill, prefix caching, guided & speculative decoding, disaggregated P/D
  3. Scaling up: from single-GPU to multi-GPU execution
  4. Serving layer: distributed / concurrent web scaffolding
  5. Benchmarks and auto-tuning: measuring latency and throughput

📝Notes

  • Analysis is based on commit 42172ad (August 9th, 2025).
  • Target audience: anyone curious about how state-of-the-art LLM engines work, as well as those interested in contributing to vLLM, SGLang, etc.
  • I'll focus on the V1 engine. I also explored V0 (now deprecated), which was valuable for understanding how the project evolved, and many concepts still carry over.
  • The first section on LLM Engine / Engine Core might be a bit overwhelming/dry - but the rest of the blog has plenty examples and visuals. :)

LLM Engine & Engine Core

The LLM engine is the fundamental building block of vLLM. On its own, it already enables high-throughput inference - but only in an offline setting. You can't serve it to customers over the web yet.

We'll use the following offline inference snippet as our running example (adapted from basic.py).

from vllm import LLM, SamplingParams

prompts = [
    "Hello, my name is",
    "The president of the United States is",
]

sampling_params = SamplingParams(temperature=0.8, top_p=0.95)

def main():
    llm = LLM(model="TinyLlama/TinyLlama-1.1B-Chat-v1.0")

    outputs = llm.generate(prompts, sampling_params)

if __name__ == "__main__":
    main()

📝Environment vars:

  • VLLM_USE_V1="1" # we're using engine V1
  • VLLM_ENABLE_V1_MULTIPROCESSING="0" # we're running in a single process

This configuration is:

  • offline (no web/distributed system scaffolding)
  • synchronous (all execution happens in a single blocking process)
  • single-GPU (no data/model/pipeline/expert parallelism; DP/TP/PP/EP = 1)
  • using standard transformer [2] (supporting hybrid models like Jamba requires a more complex hybrid KV-cache memory allocator)

From here, we'll gradually build up to an online, async, multi-GPU, multi-node inference system - but still serving a standard transformer.

In this example we do two things, we:

  1. Instantiate an engine
  2. Call generate on it to sample from the given prompts

Let's start analyzing the constructor.

LLM Engine constructor

The main components of the engine are:

  • vLLM config (contains all of the knobs for configuring model, cache, parallelism, etc.)
  • processor (turns raw inputs → EngineCoreRequests via validation, tokenization, and processing)
  • engine core client (in our running example we're using InprocClient which is basically == EngineCore; we'll gradually build up to DPLBAsyncMPClient which allows serving at scale)
  • output processor (converts raw EngineCoreOutputsRequestOutput that the user sees)

📝Note:

With the V0 engine being deprecated, class names and details may shift. I'll emphasize the core ideas rather than exact signatures. I'll abstract away some but not all of those details.

Engine core itself is made up of several sub components:

  • Model Executor (drives forward passes on the model, we're currently dealing with UniProcExecutor which has a single Worker process on a single GPU). We'll gradually build up to MultiProcExecutor which supports multiple GPUs

  • Structured Output Manager (used for guided decoding - we'll cover this later)

  • Scheduler (decides which requests go into the next engine step) - it further contains:

    1. policy setting - it can be either FCFS (first come first served) or priority (higher priority requests are served first)
    2. waiting and running queues
    3. KV cache manager - the heart of paged attention [3]

The KV-cache manager maintains a free_block_queue - a pool of available KV-cache blocks (often on the order of hundreds of thousands, depending on VRAM size and block size). During paged attention, the blocks serve as the indexing structure that map tokens to their computed KV cache blocks.

LLM engine constructor

Core components described in this section and their relationships

Block size for a standard transformer layer (non-MLA [4]) is computed as follows:
2 (key/value) * block_size (default=16) * num_kv_heads * head_size * dtype_num_bytes (e.g. 2 for bf16)

During model executor construction, a Worker object is created, and three key procedures are executed. (Later, with MultiProcExecutor, these same procedures run independently on each worker process across different GPUs.)

  1. Init device:

    • Assign a CUDA device (e.g. "cuda:0") to the worker and check that the model dtype is supported (e.g. bf16)
    • Verify enough VRAM is available, given the requested gpu_memory_utilization (e.g. 0.8 → 80% of total VRAM)
    • Set up distributed settings (DP / TP / PP / EP, etc.)
    • Instantiate a model_runner (holds the sampler, KV cache, and forward-pass buffers such as input_ids, positions, etc.)
    • Instantiate an InputBatch object (holds CPU-side forward-pass buffers, block tables for KV-cache indexing, sampling metadata, etc.)
  2. Load model:

    • Instantiate the model architecture
    • Load the model weights
    • Call model.eval() (PyTorch's inference mode)
    • Optional: call torch.compile() on the model
  3. Initialize KV cache

    • Get per-layer KV-cache spec. Historically this was always FullAttentionSpec (homogeneous transformer), but with hybrid models (sliding window, Transformer/SSM like Jamba) it became more complex (see Jenga [5])
    • Run a dummy/profiling forward pass and take a GPU memory snapshot to compute how many KV cache blocks fit in available VRAM
    • Allocate, reshape and bind KV cache tensors to attention layers
    • Prepare attention metadata (e.g. set the backend to FlashAttention) later consumed by kernels during the fwd pass
    • Unless --enforce-eager is provided, for each of warmup batch sizes do a dummy run and capture CUDA graphs. CUDA graphs record the whole sequence of GPU work into a DAG. Later during fwd pass we launch/replay pre-baked graphs and cut on kernel launch overhead and thus improve latency.

I've abstracted away many low-level details here — but these are the core pieces I'll introduce now, since I'll reference them repeatedly in the following sections.

Now that we have the engine initialized let's proceed to the generate function.

Generate function

The first step is to validate and feed requests into the engine. For each prompt we:

  1. Create a unique request ID and capture its arrival time
  2. Call an input preprocessor that tokenizes the prompt and returns a dictionary containing prompt, prompt_token_ids, and a type (text, tokens, embeds, etc.)
  3. Pack this info into an EngineCoreRequest, adding priority, sampling params, and other metadata
  4. Pass the request into the engine core, which wraps it in a Request object and sets its status to WAITING. This request is then added to the scheduler's waiting queue (append if FCFS, or heap-push if priority)

At this point the engine has been fed and execution can begin. In the synchronous engine example, these initial prompts are the only ones we'll process — there's no mechanism to inject new requests mid-run. In contrast, the asynchronous engine supports this (aka continuous batching [6]): after each step, both new and old requests are considered.

Because the forward pass flattens the batch into a single sequence and custom kernels handle it efficiently, continuous batching is fundamentally supported even in the synchronous engine.

Next, as long as there are requests to process, the engine repeatedly calls its step() function. Each step has three stages:

  1. Schedule: select which requests to run in this step (decode, and/or (chunked) prefill)
  2. Forward pass: run the model and sample tokens
  3. Postprocess: append sampled token IDs to each Request, detokenize, and check stop conditions. If a request is finished, clean up (e.g. return its KV-cache blocks to free_block_queue) and return the output early

📝Stop conditions are:

  • The request exceeds its length limit (max_model_length or its own max_tokens)
  • The sampled token is the EOS ID (unless ignore_eos is enabled -> useful for benchmarking when we want to force a generation of a certain number of out tokens)
  • The sampled token matches any of the stop_token_ids specified in the sampling parameters
  • Stop strings are present in the output - we truncate the output until the first stop string appearance and abort the request in the engine (note that stop_token_ids will be present in the output but stop strings will not).

Engine loop

Engine loop

In streaming mode, we would send intermediate tokens as they are generated, but we'll ignore that for now.

Next, we'll examine scheduling in more detail.

Scheduler

There are two main types of workloads an inference engine handles:

  1. Prefill requests — a forward pass over all prompt tokens. These are usually compute-bound (threshold depends on hardware and prompt length). At the end, we sample a single token from the probability distribution of the final token's position.
  2. Decode requests — a forward pass over just the most recent token. All earlier KV vectors are already cached. These are memory-bandwidth-bound, since we still need to load all LLM weights (and KV caches) just to compute one token.

In the benchmarking section we'll analyze the so-called roofline model of GPU perf. That will go into more detail behind prefill/decode perf profiles.

The V1 scheduler can mix both types of requests in the same step, thanks to smarter design choices. In contrast, the V0 engine could only process either prefill or decode at once.

The scheduler prioritizes decode requests — i.e. those already in the running queue. For each such request it:

  1. Computes the number of new tokens to generate (not always 1, due to speculative decoding and async scheduling — more on that later).
  2. Calls the KV-cache manager's allocate_slots function (details below).
  3. Updates the token budget by subtracting the number of tokens from step 1.

After that, it processes prefill requests from the waiting queue, it:

  1. Retrieves the number of computed blocks (returns 0 if prefix caching is disabled — we'll cover that later).
  2. Calls the KV-cache manager's allocate_slots function.
  3. Pops the request from waiting and moves it to running, setting its status to RUNNING.
  4. Updates the token budget.

Let's now look at what allocate_slots does, it:

  1. Computes number of blocks — determines how many new KV-cache blocks (n) must be allocated. Each block stores 16 tokens by default. For example, if a prefill request has 17 new tokens, we need ceil(17/16) = 2 blocks.
  2. Checks availability — if there aren't enough blocks in the manager's pool, exit early. Depending on whether it's a decode or prefill request, the engine may attempt recompute preemption (swap preemption was supported in V0) by evicting low-priority requests (calling kv_cache_manager.free which returns KV blocks to block pool), or it might skip scheduling and continue execution.
  3. Allocates blocks — via the KV-cache manager's coordinator, fetches the first n blocks from the block pool (the free_block_queue doubly linked list mentioned earlier). Stores to req_to_blocks, the dictionary mapping each request_id to its list of KV-cache blocks.

KV cache blocks

list of KV cache blocks

We're finally ready to do a forward pass!

Run forward pass

We call model executor's execute_model, which delegates to the Worker, which in turn delegates to the model runner.

Here are the main steps:

  1. Update states — prune finished requests from input_batch; update misc fwd pass related metadata (e.g., KV cache blocks per request that will be used to index into paged KV cache memory).
  2. Prepare inputs — copy buffers from CPU→GPU; compute positions; build slot_mapping (more on that in example); construct attention metadata.
  3. Forward pass — run the model with custom paged attn kernels. All sequences are flattened and concatenated into one long "super sequence". Position indices and attention masks ensure each sequence only attends to its own tokens, which enables continuous batching without right-padding.
  4. Gather last-token states — extract hidden states for each sequence's final position and compute logits.
  5. Sample — sample tokens from computed logits as dictated by the sampling config (greedy, temperature, top-p, top-k, etc.).

Forward-pass step itself has two execution modes:

  1. Eager mode — run the standard PyTorch forward pass when eager execution is enabled.
  2. "Captured" mode — execute/replay a pre-captured CUDA Graph when eager is not enforced (remember we captured these during engine construction in the initialize KV cache procedure).

Here is a concrete example that should make continuous batching and paged attention clear:

fwd pass - continuous batching & paged attn

Forward pass: continuous batching and paged attention

Advanced Features — extending the core engine logic

With the basic engine flow in place, we can now look at the advanced features.

We've already discussed preemption, paged attention, and continuous batching.

Next, we'll dive into:

  1. Chunked prefill
  2. Prefix caching
  3. Guided decoding (through grammar-constrained finite-state machines)
  4. Speculative decoding
  5. Disaggregated P/D (prefill/decoding)

Chunked prefill

Chunked prefill is a technique for handling long prompts by splitting their prefill step into smaller chunks. Without it, we could end up with a single very long request monopolizing one engine step disallowing other prefill requests to run. That would postpone all other requests and increase their latency.

For example, let each chunk contain n (=8) tokens, labeled with lowercase letters separated by "-". A long prompt P could look like x-y-z, where z is an incomplete chunk (e.g. 2 toks). Executing the full prefill for P would then take ≥ 3 engine steps (> can happen if it's not scheduled for execution in one of the steps), and only in the last chunked prefill step would we sample one new token.

Here is that same example visually:

Chunked prefilling - pt 1

Implementation is straightforward: cap the number of new tokens per step. If the requested number exceeds long_prefill_token_threshold, reset it to exactly that value. The underlying indexing logic (described earlier) takes care of the rest.

In vLLM V1, you enable chunked prefill by setting long_prefill_token_threshold to a positive integer. (Technically, it can happen irrespective of this, if the prompt length exceeds the token budget we truncate it and run a chunked prefill.)

Prefix Caching

To explain how prefix caching works, let's take the original code example and tweak it a bit:

from vllm import LLM, SamplingParams

long_prefix = "<a piece of text that is encoded into more than block_size tokens>"

prompts = [
    "Hello, my name is",
    "The president of the United States is",
]

sampling_params = SamplingParams(temperature=0.8, top_p=0.95)

def main():
    llm = LLM(model="TinyLlama/TinyLlama-1.1B-Chat-v1.0")

    outputs = llm.generate(long_prefix + prompts[0], sampling_params)
    outputs = llm.generate(long_prefix + prompts[1], sampling_params)

if __name__ == "__main__":
    main()

Prefix caching avoids recomputing tokens that multiple prompts share at the beginning - hence prefix.

The crucial piece is the long_prefix: it's defined as any prefix longer than a KV-cache block (16 tokens by default). To simplify our example let's say long_prefix has exactly length n x block_size (where n ≥ 1).

i.e. it perfectly aligns with block boundary - otherwise we'd have to recompute long_prefix_len % block_size tokens as we can't cache incomplete blocks.

Without prefix caching, each time we process a new request with the same long_prefix, we'd recompute all n x block_size tokens.

With prefix caching, those tokens are computed once (their KVs stored in KV cache paged memory) and then reused, so only the new prompt tokens need processing. This speeds up prefill requests (though it doesn't help with decode).

How does this work in vLLM?

During the first generate call, in the scheduling stage, inside kv_cache_manager.get_computed_blocks, the engine invokes hash_request_tokens:

  1. This function splits the long_prefix + prompts[0] into 16-token chunks.

  2. For each complete chunk, it computes a hash (using either the built-in hash or SHA-256, which is slower but has fewer collisions). The hash combines the previous block's hash, the current tokens, and optional metadata.

    optional metadata includes: MM hash, LoRA ID, cache salt (injected into hash of the first block ensures only requests with this cache salt can reuse blocks).

  3. Each result is stored as a BlockHash object containing both the hash and its token IDs. We return a list of block hashes.

The list is stored in self.req_to_block_hashes[request_id].

Next, the engine calls find_longest_cache_hit to check if any of these hashes already exist in cached_block_hash_to_block. On the first request, no hits are found.

Prefix caching logic - pt 1

Then we call allocate_slots which calls coordinator.cache_blocks, which associates the new BlockHash entries with allocated KV blocks and records them in cached_block_hash_to_block.

Afterwards, the forward pass will populate KVs in paged KV cache memory corresponding to KV cache blocks that we allocated above.

After many engine steps it'll allocate more KV cache blocks but it doesn't matter for our example because the prefix has diverged immediately after long_prefix.

Prefix caching logic - pt 2

On a second generate call with the same prefix, steps 1-3 repeat, but now find_longest_cache_hit finds matches for all n blocks (via linear search). The engine can reuse those KV blocks directly.

Prefix caching logic - pt 3

If the original request were still alive, the reference count for those blocks would increment (e.g. to 2). In this example, the first request has already completed, so the blocks were freed back to the pool and their reference counts set back to 0. Because we were able to retrieve them from cached_block_hash_to_block we know they're valid (the logic of the KV cache manager is setup in such a way), so we just remove them from free_block_queue again.

📝Advanced note:

KV-cache blocks become invalid only when they're about to be reallocated from the free_block_queue (which pops from the left) and we discover the block still has an associated hash and is present in cached_block_hash_to_block. At that moment, we clear the block's hash and remove its entry from cached_block_hash_to_block, ensuring it can't be reused via prefix caching (at least not for that old prefix).

And that's the gist of prefix caching: don't recompute prefixes you've already seen — just reuse their KV cache!

If you understood this example you also understood how paged attention works.

Prefix caching is enabled by default. To disable it: enable_prefix_caching = False.

Guided Decoding (FSM)

Guided decoding is a technique where, at each decoding step, the logits are constrained by a grammar-based finite state machine. This ensures that only tokens allowed by the grammar can be sampled.

It's a powerful setup: you can enforce anything from regular grammars (Chomsky type-3, e.g. arbitrary regex patterns) all the way up to context-free grammars (type-2, which cover most programming languages).

To make this less abstract, let's start with the simplest possible example, building on our earlier code:

from vllm import LLM, SamplingParams
from vllm.sampling_params import GuidedDecodingParams

prompts = [
    "This sucks",
    "The weather is beautiful",
]

guided_decoding_params = GuidedDecodingParams(choice=["Positive", "Negative"])
sampling_params = SamplingParams(guided_decoding=guided_decoding_params)

def main():
    llm = LLM(model="TinyLlama/TinyLlama-1.1B-Chat-v1.0")

    outputs = llm.generate(prompts, sampling_params)

if __name__ == "__main__":
    main()

In the toy example I gave (assume character-level tokenization): at prefill, the FSM masks logits so only "P" or "N" are viable. If "P" is sampled, the FSM moves to the "Positive" branch; next step only "o" is allowed, and so on.

FSM

Toy example FSM

How this works in vLLM:

  1. At LLM engine construction, a StructuredOutputManager is created; it has access to the tokenizer and maintains a _grammar_bitmask tensor.
  2. When adding a request, its status is set to WAITING_FOR_FSM and grammar_init selects the backend compiler (e.g., xgrammar [7]; note that backends are 3rd party code).
  3. The grammar for this request is compiled asynchronously.
  4. During scheduling, if the async compile has completed, the status switches to WAITING and request_id is added to structured_output_request_ids; otherwise it's placed in skipped_waiting_requests to retry on next engine step.
  5. After the scheduling loop (still inside scheduling), if there are FSM requests, the StructuredOutputManager asks the backend to prepare/update _grammar_bitmask.
  6. After the forward pass produces logits, xgr_torch_compile's function expands the bitmask to vocab size (32x expansion ratio because we use 32 bit integers) and masks disallowed logits to –∞.
  7. After sampling the next token, the request's FSM is advanced via accept_tokens. Visually we move to the next state on the FSM diagram.

Step 6 deserves further clarification.

If vocab_size = 32, _grammar_bitmask is a single integer; its binary representation encodes which tokens are allowed ("1") vs disallowed ("0"). For example, "101…001" expands to a length-32 array [1, 0, 1, …, 0, 0, 1]; positions with 0 get logits set to –∞. For larger vocabularies, multiple 32-bit words are used and expanded/concatenated accordingly. The backend (e.g., xgrammar) is responsible for producing these bit patterns using the current FSM state.

📝Note:

Most of the complexity here is hidden in the 3rd party libs like xgrammar.

Here is an even simpler example with vocab_size = 8 and 8-bit integers (for those of you who like my visuals):

FSM

Toy example

You can enable this in vLLM by passing in a desired guided_decoding config.

Speculative Decoding

In autoregressive generation, each new token requires a forward pass of the large LM. This is expensive — every step reloads and applies all model weights just to compute a single token! (assuming batch size == 1, in general it's B)

Speculative decoding [8] speeds this up by introducing a smaller draft LM. The draft proposes k tokens cheaply. But we don't ultimately want to sample from the smaller model — it's only there to guess candidate continuations. The large model still decides what's valid.

Here are the steps:

  1. Draft: run the small model on the current context and propose k tokens

  2. Verify: run the large model once on context + k draft tokens. This produces probabilities for those k positions plus one extra (so we get k+1 candidates)

  3. Accept/reject: going from left to right over the k draft tokens:

    • If the large model's probability for the draft token ≥ the draft's probability, accept it

    • Otherwise, accept it with probability p_large(token)/p_draft(token)

    • Stop at the first rejection, or accept all k draft tokens.

      • If all k draft tokens are accepted, also sample the extra (k+1)-th token "for free" from the large model (we already computed that distribution).
      • If there was a rejection create a new rebalanced distribution at that position (p_large - p_draft, clamp min at 0, normalize to sum to 1) and sample the last token from it.

Why this works: Although we use the small model to propose candidates, the accept/reject rule guarantees that in expectation the sequence is distributed exactly as if we had sampled token by token from the large model. This means speculative decoding is statistically equivalent to standard autoregressive decoding — but potentially much faster, since a single large-model pass can yield up to k+1 tokens.

📝Note:

I recommend looking at gpt-fast for a simple implementation, and the original paper for the math details and the proof of equivalence to sampling from the full model.

vLLM V1 does not support the LLM draft model method, instead it implements faster—but less accurate—proposal schemes: n-gram, EAGLE [9], and Medusa [10].

One-liners on each:

  1. n-gram: take the last prompt_lookup_max tokens; find a prior match in the sequence; if found, propose the k tokens that followed that match; otherwise decrement the window and retry down to prompt_lookup_min

    The current implementation returns k tokens after the first match. It feels more natural to introduce a recency bias and reverse the search direction? (i.e. last match)

  2. Eagle: perform "model surgery" on the large LM—keep embeddings and LM head, replace the transformer stack with a lightweight MLP; fine-tune that as a cheap draft

  3. Medusa: train auxiliary linear heads on top (embeddings before LM head) of the large model to predict the next k tokens in parallel; use these heads to propose tokens more efficiently than running a separate small LM

Here's how to invoke speculative decoding in vLLM using ngram as the draft method:

from vllm import LLM, SamplingParams

prompts = [
    "Hello, my name is",
    "The president of the United States is",
]

sampling_params = SamplingParams(temperature=0.8, top_p=0.95)

speculative_config={
    "method": "ngram",
    "prompt_lookup_max": 5,
    "prompt_lookup_min": 3,
    "num_speculative_tokens": 3,
}

def main():
    llm = LLM(model="TinyLlama/TinyLlama-1.1B-Chat-v1.0", speculative_config=speculative_config)

    outputs = llm.generate(prompts, sampling_params)

if __name__ == "__main__":
    main()

How does this work in vLLM?

Setup (during engine construction):

  1. Init device: create a drafter (draft model, e.g., NgramProposer) and a rejection_sampler (parts of it are written in Triton).
  2. Load model: load draft model weights (no-op for n-gram).

After that in the generate function (assume we get a brand new request):

  1. Run the regular prefill step with the large model.
  2. After the forward pass and standard sampling, call propose_draft_token_ids(k) to sample k draft tokens from the draft model.
  3. Store these in request.spec_token_ids (update the request metadata).
  4. On the next engine step, when the request is in the running queue, add len(request.spec_token_ids) to the "new tokens" count so allocate_slots reserves sufficient KV blocks for the fwd pass.
  5. Copy spec_token_ids into input_batch.token_ids_cpu to form (context + draft) tokens.
  6. Compute metadata via _calc_spec_decode_metadata (this copies over tokens from input_batch.token_ids_cpu, prepares logits, etc.), then run a large-model forward pass over the draft tokens.
  7. Instead of regular sampling from logits, use the rejection_sampler to accept/reject left-to-right and produce output_token_ids.
  8. Repeat steps 2-7 until a stop condition is met.

The best way to internalize this is to fire up your debugger and step through the codebase, but this section hopefully gives you a taste for it. This as well:

Drafting stage

Verify stage & rejection sampling stage

Disaggregated P/D

I've already previously hinted at the motivation behind disaggregated P/D (prefill/decode).

Prefill and decode have very different performance profiles (compute-bound vs. memory-bandwidth-bound), so separating their execution is a sensible design. It gives tighter control over latency — both TTFT (time-to-first-token) and ITL (inter-token latency) — more on this in the benchmarking section.

In practice, we run N vLLM prefill instances and M vLLM decode instances, autoscaling them based on the live request mix. Prefill workers write KV to a dedicated KV-cache service; decode workers read from it. This isolates long, bursty prefill from steady, latency-sensitive decode.

How does this work in vLLM?

For clarity, the example below relies on SharedStorageConnector, a debugging connector implementation used to illustrate the mechanics.

Connector is vLLM's abstraction for handling the exchange of KVs between instances. Connector interface is not yet stable, there are some near-term improvements planned which will involve changes, some potentially breaking.

We launch 2 vLLM instances (GPU 0 for prefill and GPU 1 for decode), and then transfer the KV cache between them:

import os
import time
from multiprocessing import Event, Process
import multiprocessing as mp

from vllm import LLM, SamplingParams
from vllm.config import KVTransferConfig

prompts = [
    "Hello, my name is",
    "The president of the United States is",
]

def run_prefill(prefill_done):
  os.environ["CUDA_VISIBLE_DEVICES"] = "0"

  sampling_params = SamplingParams(temperature=0, top_p=0.95, max_tokens=1)

  ktc=KVTransferConfig(
      kv_connector="SharedStorageConnector",
      kv_role="kv_both",
      kv_connector_extra_config={"shared_storage_path": "local_storage"},
  )

  llm = LLM(model="TinyLlama/TinyLlama-1.1B-Chat-v1.0", kv_transfer_config=ktc)
  llm.generate(prompts, sampling_params)

  prefill_done.set()  # notify decode instance that KV cache is ready

  # To keep the prefill node running in case the decode node is not done;
  # otherwise, the script might exit prematurely, causing incomplete decoding.
  try:
      while True:
          time.sleep(1)
  except KeyboardInterrupt:
      print("Script stopped by user.")

def run_decode(prefill_done):
  os.environ["CUDA_VISIBLE_DEVICES"] = "1"

  sampling_params = SamplingParams(temperature=0, top_p=0.95)

  ktc=KVTransferConfig(
      kv_connector="SharedStorageConnector",
      kv_role="kv_both",
      kv_connector_extra_config={"shared_storage_path": "local_storage"},
  )

  llm = LLM(model="TinyLlama/TinyLlama-1.1B-Chat-v1.0", kv_transfer_config=ktc)

  prefill_done.wait()  # block waiting for KV cache from prefill instance

  # Internally it'll first fetch KV cache before starting the decoding loop
  outputs = llm.generate(prompts, sampling_params)

if __name__ == "__main__":
  prefill_done = Event()
  prefill_process = Process(target=run_prefill, args=(prefill_done,))
  decode_process = Process(target=run_decode, args=(prefill_done,))

  prefill_process.start()
  decode_process.start()

  decode_process.join()
  prefill_process.terminate()

📝Note:

I've also experimented with LMCache, the fastest production-ready connector (uses NVIDIA's NIXL as the backend), but it's still at the bleeding edge and I ran into some bugs. Since much of its complexity lives in an external repo, SharedStorageConnector is a better choice for explanation.

These are the steps in vLLM:

  1. Instantiation — During engine construction, connectors are created in two places:

    • Inside the worker's init device procedure (under init worker distributed environment function), with role "worker".
    • Inside the scheduler constructor, with role "scheduler".
  2. Cache lookup — When the scheduler processes prefill requests from the waiting queue (after local prefix-cache checks), it calls connector's get_num_new_matched_tokens. This checks for externally cached tokens in the KV-cache server. Prefill always sees 0 here; decode may have a cache hit. The result is added to the local count before calling allocate_slots.

  3. State update — The scheduler then calls connector.update_state_after_alloc, which records requests that had a cache (no-op for prefill).

  4. Meta build — At the end of scheduling, the scheduler calls meta = connector.build_connector_meta:

    • Prefill adds all requests with is_store=True (to upload KV).
    • Decode adds requests with is_store=False (to fetch KV).
  5. Context manager — Before the forward pass, the engine enters a KV-connector context manager:

    • On enter: kv_connector.start_load_kv is called. For decode, this loads KV from the external server and injects it into paged memory. For prefill, it's a no-op.
    • On exit: kv_connector.wait_for_save is called. For prefill, this blocks until KV is uploaded to the external server. For decode, it's a no-op.

Here is a visual example:

disaggregated P/D

disaggregated P/D

📝Additional notes:

  • For SharedStorageConnector "external server" is just a local file system.
  • Depending on configuration, KV transfers can also be done layer-by-layer (before/after each attention layer).
  • Decode loads external KV only once, on the first step of its requests; afterwards it computes/stores locally.

From UniprocExecutor to MultiProcExecutor

With the core techniques in place, we can now talk about scaling up.

Suppose your model weights no longer fit into a single GPU's VRAM.

The first option is to shard the model across multiple GPUs on the same node using tensor parallelism (e.g., TP=8). If the model still doesn't fit, the next step is pipeline parallelism across nodes.

📝Notes:

  • Intranode bandwidth is significantly higher than internode, which is why tensor parallelism (TP) is generally preferred over pipeline parallelism (PP). (It is also true that PP communicates less data than TP.)
  • I'm not covering expert parallelism (EP) since we're focusing on standard transformers rather than MoE, nor sequence parallelism, as TP and PP are the most commonly used in practice.

At this stage, we need multiple GPU processes (workers) and an orchestration layer to coordinate them. That's exactly what MultiProcExecutor provides.

MultiProcExecutor

MultiProcExecutor in a TP=8 setting (driver worker being rank 0)

How this works in vLLM:

  1. MultiProcExecutor initializes an rpc_broadcast_mq message queue (implemented with shared memory under the hood).

  2. The constructor loops over world_size (e.g. TP=8 ⇒ world_size=8) and spawns a daemon process for each rank via WorkerProc.make_worker_process.

  3. For each worker, the parent first creates a reader and writer pipe.

  4. The new process runs WorkerProc.worker_main, which instantiates a worker (going through the same "init device", "load model", etc. as in UniprocExecutor).

  5. Each worker determines whether it is the driver (rank 0 in the TP group) or a regular worker. Every worker sets up two queues:

    • rpc_broadcast_mq (shared with the parent) for receiving work.
    • worker_response_mq for sending responses back.
  6. During initialization, each child sends its worker_response_mq handle to the parent via the pipe. Once all are received, the parent unblocks — this completes coordination.

  7. Workers then enter a busy loop, blocking on rpc_broadcast_mq.dequeue. When a work item arrives, they execute it (just like in UniprocExecutor, but now with TP/PP-specific partitioned work). Results are sent back through worker_response_mq.enqueue.

  8. At runtime, when a request arrives, MultiProcExecutor enqueues it into rpc_broadcast_mq (non-blocking) for all children workers. It then waits on the designated output rank's worker_response_mq.dequeue to collect the final result.

From the engine's perspective, nothing has changed — all of this multiprocessing complexity is abstracted away through a call to model executor's execute_model.

  • In the UniProcExecutor case: execute_model directly leads to calling execute_model on the worker
  • In the MultiProcExecutor case: execute_model indirectly leads to calling execute_model on each worker through rpc_broadcast_mq

At this point, we can run models that are as large as resources allow using the same engine interface.

The next step is to scale out: enable data parallelism (DP > 1) replicating the model across nodes, add a lightweight DP coordination layer, introduce load balancing across replicas, and place one or more API servers in front to handle incoming traffic.

Distributed system serving vLLM

There are many ways to set up serving infrastructure, but to stay concrete, here's one example: suppose we have two H100 nodes and want to run four vLLM engines across them.

If the model requires TP=4, we can configure the nodes like this.

server configuration with 2 8xH100 nodes

server configuration with 2 8xH100 nodes (1 headless, 1 api server)

On the first node, run the engine in headless mode (no API server) with the following arguments:

vllm serve <model-name>
  --tensor-parallel-size 4
  --data-parallel-size 4
  --data-parallel-size-local 2
  --data-parallel-start-rank 0
  --data-parallel-address <master-ip>
  --data-parallel-rpc-port 13345
  --headless

and run that same command on the other node with few tweaks:

  • no --headless
  • modify DP start rank
vllm serve <model-name>
  --tensor-parallel-size 4
  --data-parallel-size 4
  --data-parallel-size-local 2
  --data-parallel-start-rank 2
  --data-parallel-address <master-ip>
  --data-parallel-rpc-port 13345

📝Note:

This assumes networking is configured so all nodes can reach the specified IP and port.

How does this work in VLLM?

On the headless server node

On the headless node, a CoreEngineProcManager launches 2 processes (per --data-parallel-size-local) each running EngineCoreProc.run_engine_core. Each of these functions creates a DPEngineCoreProc (the engine core) and then enters its busy loop.

DPEngineCoreProc initializes its parent EngineCoreProc (child of EngineCore), which:

  1. Creates an input_queue and output_queue (queue.Queue).
  2. Performs an initial handshake with the frontend on the other node using a DEALER ZMQ socket (async messaging lib), and receives coordination address info.
  3. Initializes DP group (e.g. using NCCL backend).
  4. Initializes the EngineCore with MultiProcExecutor (TP=4 on 4 GPUs as described earlier).
  5. Creates a ready_event (threading.Event).
  6. Starts an input deamon thread (threading.Thread) running process_input_sockets(…, ready_event). Similarly starts an output thread.
  7. Still in the main thread, waits on ready_event until all input threads across all 4 processes (spanning the 2 nodes) have completed the coordination handshake finally executing ready_event.set().
  8. Once unblocked, sends a "ready" message to the frontend with metadata (e.g., num_gpu_blocks available in paged KV cache memory).
  9. The main, input, and output threads then enter their respective busy loops.

distributed system with 4 DPEngineCoreProc

distributed system with 4 DP replicas running 4 DPEngineCoreProc

Current steady state:

  • Input thread — blocks on the input socket until a request is routed from the API server; upon receipt, it decodes the payload, enqueues a work item via input_queue.put_nowait(...), and returns to blocking on the socket.
  • Main thread — wakes on input_queue.get(...), feeds the request to the engine; MultiProcExecutor runs the forward pass and enqueues results to output_queue.
  • Output thread — wakes on output_queue.get(...), sends the result back to the API server, then resumes blocking.

Additional mechanics:

  • DP wave counter — the system tracks "waves"; when all engines become idle they quiesce, and the counter increments when new work arrives (useful for coordination/metrics).
  • Control messages — the API server can send more than just inference requests (e.g., aborts and utility/control RPCs).
  • Dummy steps for lockstep — if any DP replica has work, all replicas execute a forward step; replicas without requests perform a dummy step to participate in required synchronization points (avoids blocking the active replica).

Lockstep clarification: this is actually only required for MoE models where the expert layers form an EP or TP group while attention layers are still DP. It's currently always done with DP - this is just because there's limited use for "built-in" non-MoE DP since you could just run multiple independent vLLMs and load-balance between them in a normal way.

Now for the second part, what happens on the API server node?

On the API server node

We instantiate an AsyncLLM object (an asyncio wrapper around the LLM engine). Internally this creates a DPLBAsyncMPClient (data-parallel, load-balancing, asynchronous, multiprocessing client).

Inside the parent class of MPClient, the launch_core_engines function runs and:

  1. Creates the ZMQ addresses used for the startup handshake (as seen on the headless node).
  2. Spawns a DPCoordinator process.
  3. Creates a CoreEngineProcManager (same as on the headless node).

Inside AsyncMPClient (child of MPClient), we:

  1. Create an outputs_queue (asyncio.Queue).
  2. We create an asyncio task process_outputs_socket which communicates (through the output socket) with output threads of all 4 DPEngineCoreProc and writes into outputs_queue.
  3. Subsequently one more asyncio task output_handler from AsyncLLM reads from this queue and finally sends out information to the create_completion function.

Inside DPAsyncMPClient we create an asyncio task run_engine_stats_update_task which communicates with DP coordinator.

The DP coordinator mediates between the frontend (API server) and backend (engine cores). It:

  • Periodically sends load-balancing info (queue sizes, waiting/running requests) to the frontend's run_engine_stats_update_task.
  • Handles SCALE_ELASTIC_EP commands from the frontend by dynamically changing the number of engines (only works with Ray backend).
  • Sends START_DP_WAVE events to the backend (when triggered by frontend) and reports wave-state updates back.

To recap, the frontend (AsyncLLM) runs several asyncio tasks (remember: concurrent, not parallel):

  • A class of tasks handles input requests through the generate path (each new client request spawns a new asyncio task).
  • Two tasks (process_outputs_socket, output_handler) process output messages from the underlying engines.
  • One task (run_engine_stats_update_task) maintains communication with the DP coordinator: sending wave triggers, polling LB state, and handling dynamic scaling requests.

Finally, the main server process creates a FastAPI app and mounts endpoints such as OpenAIServingCompletion and OpenAIServingChat, which expose /completion, /chat/completion, and others. The stack is then served via Uvicorn.

So, putting it all together, here's the full request lifecycle!

You send from your terminal:

curl -X POST http://localhost:8000/v1/completions -H "Content-Type: application/json" -d '{
  "model": "TinyLlama/TinyLlama-1.1B-Chat-v1.0",
  "prompt": "The capital of France is",
  "max_tokens": 50,
  "temperature": 0.7
}'

What happens next:

  1. The request hits OpenAIServingCompletion's create_completion route on the API server.

  2. The function tokenizes the prompt asynchronously, and prepares metadata (request ID, sampling params, timestamp, etc.).

  3. It then calls AsyncLLM.generate, which follows the same flow as the synchronous engine, eventually invoking DPAsyncMPClient.add_request_async.

  4. This in turn calls get_core_engine_for_request, which does load balancing across engines based on the DP coordinator's state (picking the one that has minimal score / lowest load: score = len(waiting) * 4 + len(running)).

  5. The ADD request is sent to the chosen engine's input_socket.

  6. At that engine:

    • Input thread — unblocks, decodes data from the input socket, and places a work item on the input_queue for the main thread.

    • Main thread — unblocks on input_queue, adds the request to the engine, and repeatedly calls engine_core.step(), enqueueing intermediate results to output_queue until a stop condition is met.

      Reminder: step() calls the scheduler, model executor (which in turn can be MultiProcExecutor!), etc. We have already seen this!

    • Output thread — unblocks on output_queue and sends results back through the output socket.

  7. Those results trigger the AsyncLLM output asyncio tasks (process_outputs_socket and output_handler), which propagate tokens back to FastAPI's create_completion route.

  8. FastAPI attaches metadata (finish reason, logprobs, usage info, etc.) and returns a JSONResponse via Uvicorn to your terminal!

And just like that, your completion came back — the whole distributed machinery hidden behind a simple curl command! :) So much fun!!!

📝Additional notes:

  • When adding more API servers, load balancing is handled at the OS/socket level. From the application's perspective, nothing significant changes — the complexity is hidden.
  • With Ray as a DP backend, you can expose a URL endpoint (/scale_elastic_ep) that enables automatic scaling of the number of engine replicas up or down.

Benchmarks and auto-tuning - latency vs throughput

So far we've been analyzing the "gas particles" — the internals of how requests flow through the engine/system. Now it's time to zoom out and look at the system as a whole, and ask: how do we measure the performance of an inference system?

At the highest level there are two competing metrics:

  1. Latency — the time from when a request is submitted until tokens are returned
  2. Throughput — the number of tokens/requests per second the system can generate/process

Latency matters most for interactive applications, where users are waiting on responses.

Throughput matters in offline workloads like synthetic data generation for pre/post-training runs, data cleaning/processing, and in general - any type of offline batch inference jobs.

Before explaining why latency and throughput compete, let's define a few common inference metrics:

Metric Definition
TTFT (time to first token) Time from request submission until the first output token is received
ITL (inter-token latency) Time between two consecutive tokens (e.g., from token i-1 to token i)
TPOT (time per output token) The average ITL across all output tokens in a request
Latency / E2E (end-to-end latency) Total time to process a request, i.e. TTFT + sum of all ITLs, or equivalently the time between submitting request and receiving the last output token
Throughput Total tokens processed per second (input, output, or both), or alternatively requests per second
Goodput Throughput that meets service-level objectives (SLOs) such as max TTFT, TPOT, or e2e latency. For example, only tokens from requests meeting those SLOs are counted

ttft, itl, e2e latency

ttft, itl, e2e latency

Here is a simplified model explaining the competing nature of these 2 metrics.

Assumption: weight i/o and not KV cache i/o dominates; i.e. we're dealing with short sequences.

The tradeoff becomes clear when looking at how batch size B affects a single decode step. As B ↓ toward 1, ITL drops: there's less work per step and the token isn't "competing" with others. As B ↑ toward infinity, ITL rises because we do more FLOPs per step—but throughput improves (until we hit peak perf) because weight I/O is amortized across more tokens.

A roofline model helps with understanding here: below a saturation batch B_sat, the step time is dominated by HBM bandwidth (streaming weights layer-by-layer into on-chip memory), so step latency is nearly flat—computing 1 vs 10 tokens can take a similar time. Beyond B_sat, the kernels become compute-bound and step time grows roughly with B; each extra token adds to ITL.

roofline perf model

roofline perf model

📝Note:

For a more rigorous treatment, we have to account for kernel auto-tuning: as B grows, the runtime may switch to more efficient kernels for that shape, changing the achieved performance P_kernel. Step latency is t = FLOPs_step / P_kernel, where FLOPs_step is the work in the step. You can see that as P_kernel hits P_peak more compute per step will directly lead to an increase in latency.

How to benchmark in vLLM

vLLM provides a vllm bench {serve,latency,throughput} CLI that wraps vllm / benchmarks / {server,latency,throughput}.py.

Here is what the scripts do:

  • latency — uses a short input (default 32 tokens) and samples 128 output tokens with a small batch (default 8). It runs several iterations and reports e2e latency for the batch.
  • throughput — submits a fixed set of prompts (default: 1000 ShareGPT samples) all at once (aka as QPS=Inf mode), and reports input/output/total tokens and requests per second across the run.
  • serve — Launches a vLLM server and simulates a real-world workload by sampling request inter-arrival times from a Poisson (or more generally, Gamma) distribution. It sends requests over a time window, measures all the metrics we’ve discussed, and can optionally enforce a server-side max concurrency (via a semaphore, e.g. limiting the server to 64 concurrent requests).

Here is an example of how you can run the latency script:

vllm bench latency
  --model <model-name>
  --input-tokens 32
  --output-tokens 128
  --batch-size 8

Benchmark configs used in CI live under .buildkite/nightly-benchmarks/tests.

There is also an auto-tune script that drives the serve benchmark to find argument settings that meet target SLOs (e.g., "maximize throughput while keeping p99 e2e < 500 ms"), returning a suggested config.

Epilogue

We began with the basic engine core (UniprocExecutor), added advanced features like speculative decoding and prefix caching, scaled up to MultiProcExecutor (with TP/PP > 1), and finally scaled out, wrapped everything in the asynchronous engine and distributed serving stack—closing with how to measure system performance.

vLLM also includes specialized handling that I've skipped. E.g.:

  • Diverse hardware backends: TPUs, AWS Neuron (Trainium/Inferentia), etc.
  • Architectures/techniques: MLA, MoE, encoder-decoder (e.g., Whisper), pooling/embedding models, EPLB, m-RoPE, LoRA, ALiBi, attention-free variants, sliding-window attention, multimodal LMs, and state-space models (e.g., Mamba/Mamba-2, Jamba)
  • TP/PP/SP
  • Hybrid KV-cache logic (Jenga), more complex sampling methods like beam sampling, and more
  • Experimental: async scheduling

The nice thing is that most of these are orthogonal to the main flow described above—you can almost treat them like "plugins" (in practice there's some coupling, of course).

I love understanding systems. Having said that, the resolution definitely suffered at this altitude. In the next posts I'll zoom in on specific subsystems and get into the nitty-gritty details.

Acknowledgements

A huge thank you to Hyperstack for providing me with H100s for my experiments over the past year!

Thanks to Nick Hill (core vLLM contributor, RedHat), Mark Saroufim (PyTorch), Kyle Krannen (NVIDIA, Dynamo), and Ashish Vaswani for reading pre-release version of this blog post and providing feedback!

References

  1. vLLM https://github.com/vllm-project/vllm
  2. "Attention Is All You Need", https://arxiv.org/abs/1706.03762
  3. "Efficient Memory Management for Large Language Model Serving with PagedAttention", https://arxiv.org/abs/2309.06180
  4. "DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model", https://arxiv.org/abs/2405.04434
  5. "Jenga: Effective Memory Management for Serving LLM with Heterogeneity", https://arxiv.org/abs/2503.18292
  6. "Orca: A Distributed Serving System for Transformer-Based Generative Models", https://www.usenix.org/conference/osdi22/presentation/yu
  7. "XGrammar: Flexible and Efficient Structured Generation Engine for Large Language Models", https://arxiv.org/abs/2411.15100
  8. "Accelerating Large Language Model Decoding with Speculative Sampling", https://arxiv.org/abs/2302.01318
  9. "EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty", https://arxiv.org/abs/2401.15077
  10. "Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads", https://arxiv.org/abs/2401.10774
  11. LMCache, https://github.com/LMCache/LMCache
The Daily Front Page 17 of 30
Friday, August 7, 2026 The Daily Front No. #260807 — Types on Guard
article

Guarded Methods in OCaml (2025)

by birdculture·▲ 78 points·7 comments·xvw.lol ↗
Guarded methods allow attaching constraints to the receiver only for certain methods.

This article is a translation, the original version is available here.

Guarded methods allow attaching constraints to the receiver (self) only for certain methods, thus allowing these methods to be called only if the receiver satisfies these constraints (these guards). OCaml does not syntactically allow defining this kind of method directly. In this note, we will see how to encode them using a type equality witness.

Problem presentation

When a language (where type checking is done before the program runs, like in Java or OCaml) introduces parametric polymorphism (Java's generics), it's sometimes possible to constrain type variables. For example:

class MyClass<T extends S> { ... }

We make MyClass generic by assuming that the type variable T is a subtype of S. The problem is that the constraint applies to the entire class. Yet sometimes, we’d like to have constraints apply only to certain methods. For example, let’s say we have a class MyList describing a list:

class MyList<A> extends ArrayList<A> {
   public int length() {
     return this.size();
   }
}

How can we define a flatten method that, for a list like [[1, 2, 3], [4, 5]], would produce the list [1, 2, 3, 4, 5]? If we place the constraint at the class level, we force our list to be "always a list of lists," which is very limiting. To implement such a method, we have three theoretical approaches available.

Moving the method outside the class

The first solution is the most obvious: simply "cheat" by moving the method outside the class body (for example, into the static context or a companion object):

class MyList<A> extends ArrayList<A> {
   public static <A> MyList<A> flatten(MyList<MyList<A>> list) {
      // Flatten's implementation
   }
   public int length() {
     return this.size();
   }
}

This approach works and doesn’t require any special ceremony. However, it forces the developer to keep track of which methods are in the class body and which are in the static context. Moreover, it breaks the systematic approach of sending messages to an instance (often presented as one of the key arguments in favor of object-oriented programming).

Extension methods

Kotlin (and others, like C#) offer extension methods which, in addition to allowing the extension of an already existing class (which can be very useful for adding behavior to the String class, which is final in Java), also provide more flexibility in defining the receiver. For example, we could write flatten like this (in Kotlin):

class MyList<A> : ArrayList<A> { ... }
fun <A> MyList<MyList<A>>.flatten() = ...

Even though this solution seems nearly perfect, it still requires the method to be defined outside the class, which might potentially mean having to make certain members of the class public in order to be accessible from an extension (looks like potentials leaky abstractions). However, it still preserves the systematic message-sending approach while allowing for more fine-grained qualification of the receiver.

Guarded methods

The final approach is probably the most ideological, as it keeps the method definition within the class. It doesn't force any escaped abstractions. This is the idea of guarded methods, the ability to add constraints on the generic parameter at the method definition level. In an imaginary syntax (this code compiles because it's not syntactically invalid, but it doesn’t produce the intended effect):

class MyList<A> : ArrayList<A>() {
    fun length() = size
    fun <B> MyList<MyList<B>>.flatten() =
        // Implémentation de flatten
}

Even though this doesn’t seem significantly different from classic extension methods (as evidenced by their syntax, which inside the class body seems sufficient), it addresses all the issues raised earlier:

  • we can characterize the receiver more precisely than in a normal method
  • we don’t break the regular message sending
  • we still benefit from available members (so we don’t escape representations)

Although guarded methods seem necessary, unfortunately, I don’t know of any mainstream languages that allow their definition. That’s very sad. Fortunately, in OCaml, it is possible to encode them.

OOP/FP symmetry: theory and practice

Since guarded methods are quite rare in popular programming languages, I discovered their existence fairly recently while reading the slides from the presentation "The Object-Oriented/Functional-Programming symmetry: theory and practice" by Gabriel Scherer.

I recommend this presentation, which showcases a symmetry between the tools of statically typed functional programming and object-oriented programming. Even though this symmetry has been observed and studied many times, the presentation is comprehensive and accessible (and relatively unbiased, discussing the pros and cons of both approaches). Unfortunately not covered during the presentation (time is often the enemy of a presenter), an entire section on guarded methods is included in the slides. The original example offers a symmetrical observation between the implementation of the flatten function in a classic functional style:

type 'a list = ...
let rec length : 'a list -> int = ...
let rec concat : 'a list -> 'a list -> 'a list = ...

let rec flatten : 'a list list -> 'a list = function
  | [] -> []
  | x::xs -> x @ flatten xs

And the implementation of a flatten method if we were in the object-oriented world, posing exactly the problem introduced in this note. The question is: what type should flatten have?

class type ['a] olist = object
  method length : int
  method concat : 'a olist -> 'a olist

  method flatten : ???
end

He therefore proposes this syntax, which implies a guard on the flatten method:

method flatten : 'b olist with 'a = 'b olist

This syntax allows describing a guarded method and could be generalized like this: method method_name : return_type with generic_type = other_type. Similar to substitutions in modules, we could specify constraints on multiple generics using and. For example: method foo : string with 'a = string and b = int for a class parameterized by two types: class ['a, 'b] t.

Additionally, this syntax would also allow defining specific behaviors elegantly. For example, for our olist type, we could provide a sum method if the elements of the list are integers:

class type ['a] olist = object
  method length : int
  method concat : 'a olist -> 'a olist
  method flatten : 'b olist with 'a = 'b olist
  method sum : int with 'a = int
end

All of this sounds extraordinary, but unfortunately, this syntax is not available in OCaml. That’s annoying! Don’t worry, it is possible to encode it using a few small tools.

Guarded methods in OCaml

Even though I had a fairly clear idea of the tools to use for encoding guarded methods, after running into a few corner cases, I decided to call on someone in the OCaml community who never asks questions but always answers them expansively: Florian Angeletti, also known as Octachron. (A fun little note: octachron is the name of a MIDI drum sequencer, so when I searched his nickname on Google, the suggestions immediately included octachron ocaml).

Our goal is to allow adding a constraint to certain methods so that they are only accessible if the receiver’s type satisfies it. Without modifying the language syntax, modeling a constraint can consist of providing an additional parameter that enforces it. In other words, we want to provide evidence.

Provide a type equality witness

Since the introduction of generalized algebraic data types in the language, there is a fairly straightforward way to define a type equality witness:

type (_, _) eq =
  | Refl : ('a, 'a) eq

The eq type, which has only one constructor: Refl, allows representing type equalities not known by the type-checker. Since we can only construct Refl values that associate two equal types, instantiating Refl within a scope guarantees that those types are equivalent. For example:

type other_int = int
let _ : (int, other_int) eq = Refl
(* Now, we have a proof of [int = other_int]. *)

This example is somewhat artificial because here the compiler knows perfectly well that int = other_int. However, there are cases where the compiler cannot know this. For example, when data is provided at runtime, where it makes perfect sense that the type-checker has no information about a type, or when the type’s representation is hidden by abstraction.

The goal of this note is not to delve into eq, so let’s just keep in mind that if we can construct a Refl value, we have a guarantee that two syntactically different types are actually equal.

Constrain with eq

Let’s return to our example that provides an object API for a list. Here is its interface:

class type ['a] obj_list =
  object ('self)
    method length : int
    method append : 'a list -> 'a obj_list
    method uncons : ('a * 'self) option
    method flatten : ???
  end

To give a type to flatten, we want to impose that 'a (the type parameter of the obj_list class) is a list. In other words, we want a proof that 'a is of type 'b list . That is, to guarantee that 'a and 'b list, though syntactically different, are equal. It’s simple: just require a value of type ('a, 'b list) eq to be provided:

method flatten : ('a, 'b list) eq -> 'b list

Now that we have an interface describing a list with a flatten method that constrains the receiver to be a list of something, let’s concretely implement a class that implements obj_list.

Implementing the interface obj_list

The first methods (length, append, and uncons) are straightforward to implement:

let my_list (list : 'a list) =
  object (self : 'a obj_list)
    val l = list
    method length = List.length l
    method append x = {<l = List.append l x>}
    method uncons = match l with [] -> None | x :: xs -> Some (x, {<l = xs>})

    method flatten = ???
  end

Now, let’s focus on flatten. We will recursively traverse the list, concatenating each element to the previous one. For example, [[1]; [2]; [3]] will become [1] @ [2]; [3]. Beyond the somewhat noisy annotations, the entire trick lies in instantiating Refl to provide evidence that 'a = 'b list.

method flatten : 'b. ('a, 'b list) eq -> 'b list =
  let rec aux : type a b. a #obj_list -> (a, b list) eq -> b list =
  fun list witness -> match list#uncons with
    | None -> []
    | Some (head_list, xs) ->
        let flatten_list : b list =
          let Refl = witness  in head_list
        in flatten_list @ aux xs witness
  in aux self

Adding a guarded method sum

Now that we’re able to constrain certain methods, let’s try adding a sum method that produces the sum of a list of integers! First, let’s add sum to our interface. This time, we want to constrain our type parameter to be int. To do that, we simply take ('a, int) eq as a type equality witness:

class type ['a] obj_list =
  object ('self)
    method length : int
    method append : 'a list -> 'a obj_list
    method uncons : ('a * 'self) option
    method flatten : ('a, 'b list) eq -> 'b list
    method sum : ('a, int) eq -> int
  end

Next, we can implement the sum method, which is just a use of the fold_left function.

method sum : ('a, int) eq -> int =
  let aux : type a. a list -> (a, int) eq -> int = fun list Refl ->
    List.fold_left (fun acc x -> acc + x) 0 list
  in aux l

The implementation of sum is logically simpler than that of flatten because it doesn’t introduce additional type variables. And we can test our different methods: we can invoke flatten on an object of type 'a list obj_list and sum on an object of type int obj_list.

let a = my_list [ [ 1 ]; [ 2 ]; [ 3 ] ]
let _ = assert ([ 1; 2; 3 ] = a#flatten Refl)
let b = my_list [ 1; 2; 3; 4 ]
let _ = assert (10 = b#sum Refl)

If we try to apply a guarded method with the wrong type, for example, trying to sum our list a (which is of type 'a list obj_list), the program will not compile. This makes sense, as we are trying to call a guarded method without respecting the imposed constraint.

1 | let _ = a#sum Refl
                  ^^^^
Error: This expression has type (int list, int list) eq
       but an expression was expected of type (int list, int) eq
       Type int list is not compatible with type int

Exactly the expected behavior! We can now define methods that constrain the receiver’s type by means of a type equality witness. Mission complete!

To conclude

It is surprising that guarded methods are not present in all statically typed OOP languages, since they allow expressing more methods while preserving the message-passing semantics so dear to object-oriented programming. Not being a heavy user of object-oriented programming languages (OCaml is the only one I use regularly), I’m not aware of languages offering syntactic support for guarded methods. However, I recently learned, pointed out by Nicolas Rinaudo, that the language Scala uses a similar encoding, but where the type equality witness is provided implicitly, thus lightening the call and not forcing the user to manually provide Refl.

Even though the encoding is somewhat heavy, and one could imagine native language support to simplify the definition of guarded methods, explicitly manipulating a type equality witness allows us to encode them. Is it useful? Since OOP programming is rarely encouraged in OCaml, probably not, but it was still fun to present a concrete and practical use case for equality witnesses!

The Daily Front Page 18 of 30
Friday, August 7, 2026 The Daily Front No. #260807 — The AI Desk: Models, Code and Memory
article

DeepSeek V4 Flash 0731

by tosh·▲ 740 points·445 comments·arcprize.org ↗

Paper ↗ Model ↗

At max effort, DeepSeek V4 Flash 0731 scores 89.0% on ARC-AGI-1 Semi-Private at $0.02 per task and 61.4% on ARC-AGI-2 Semi-Private at $0.04 per task.

ARC-AGI 2 leaderboard

Verified scores

Variant ARC-AGI-1 ARC-AGI-2 ARC-AGI-3
Max 89.0% 61.4%
High 87.0% 56.0%
Low 84.0% 46.0%

Tasks & environments

Pass/fail per reasoning level across each benchmark.

ARC-AGI-2 Public Eval

120 tasks

ARC-AGI-1 Public Eval

400 tasks

article

Oracle bans AI-generated code from OpenJDK

by delduca·▲ 508 points·372 comments·app.dealroom.co ↗

Oracle has banned AI-generated code from OpenJDK contributions, citing safety, security, and intellectual property risks. The open-source Java project steward said developers can use LLMs privately for debugging and reviewing code but cannot submit AI-generated material to repositories, pull requests, or other project channels. The policy contrasts sharply with Oracle's internal practices. Co-founder Larry Ellison recently declared that AI models now write Oracle's code, whilst co-CEO Mike Sicilia credited AI tools with enabling smaller engineering teams to deliver faster. Oracle is investing $70 billion this year in datacentre expansion. The spending spree prompted credit agency S&P to downgrade Oracle's rating to BBB-, one notch above junk status, citing uncertain returns on investment.

The Daily Front Page 19 of 30
Friday, August 7, 2026 The Daily Front No. #260807 — The AI Desk: Models, Code and Memory
article

2027 memory capacity is reportedly sold out

by inigyou·▲ 478 points·458 comments·ign.com ↗

2028 isn't looking that much better, either.

It feels like no matter where I look recently, there's more evidence that the RAM crisis isn't ending any time soon. Now, a new report suggests that all three memory manufacturers — Samsung, SK Hynix, and Micron — have collectively sold through their 2027 memory manufacturing capacity to, you guessed it, AI companies.

According to a report from Digitimes and spotted by TweakTown, all DRAM and HBM capacity has been sold through for 2027, with no further supply planned. (The companies themselves have not yet confirmed this to be the case.) Of course, that only means that retail prices are going to go up even further in the near future.

Selling through all the capacity for next year would be bad enough for consumers, but the problem is that most of the companies buying up this capacity have done so through long-term purchasing agreements, essentially pre-ordering memory that hasn't even been manufactured yet over a five-year term. Unless more capacity opens up somewhere in the next few years, it's hard to imagine that RAM will become more affordable and less scarce any time soon.

It's not just memory that's affected. According to that Digitimes report, NAND storage is growing in demand too. Luckily, NAND capacity hasn't been completely absorbed, because there are more companies supplying the stuff. But it's not exactly hard to see that SSDs have been growing steadily more expensive over the last six months or so.

Just look at the Western Digital SN7100, a pretty bog-standard PCIe 4 drive. According to Camelcamelcamel, it was about $110 in January for 1TB of storage, but now it's 52% more expensive, costing $189 before the tiny discount that Amazon currently has running. If NAND manufacturing capacity grows as scarce as RAM, I can only imagine how much that drive is going to cost this time next year.

All I know is that I grow weary of covering the RAM crisis, but it's hard to look at anything in the hardware space and not see the effects. Just this month, the Xbox Series X went up in price, and a month ago, the Steam Machine came out at a higher price than Valve wanted because of the higher cost of memory. At this point, all I can do is beg for the AI bubble to burst so I can at least buy a decent kit of RAM for less than $500 again.

The Daily Front Page 20 of 30
Friday, August 7, 2026 The Daily Front No. #260807 — The AI Desk: Models, Code and Memory
article

Responding to the next frontier of critical cyber capabilities

by artninja1988·▲ 195 points·190 comments·openai.com ↗

Cybersecurity is rapidly changing as models become more capable in ways that can both strengthen cyberdefenses and enable attacks at unprecedented speed and scale.

Our latest internal evaluations of Astra, one of our upcoming models, over the past few days indicate significant advancements in agentic coding and cybersecurity. These results, in addition to expert assessments, have led us to conclude last night that we cannot rule out critical cyber capabilities under our Preparedness Framework⁠(opens in a new window).

We are sharing this because we believe it’s important to be transparent with the public and the safety and security communities about this potential shift in capabilities.

We first published our Preparedness Framework in December 2023, well before models approached biological, chemical, cybersecurity, and AI self-improvement capabilities at this level. We created it to give us a guide for identifying progress in capability and then planning what our company would do as those capabilities emerge. Previous models, including GPT‑5.6‑Sol, have been evaluated for frontier cyber capabilities and assessed at the High (rather than Critical) threshold.

Measuring critical cybersecurity capabilities

Under our Preparedness Framework, a model reaches the Critical cybersecurity threshold if it can identify and develop functional zero-day exploits of all severity levels in many hardened real-world critical systems without human intervention, or can devise and execute end-to-end novel strategies for cyberattacks against hardened targets given only a high level desired goal.

While we continue to benchmark and assess this model, our preliminary evaluations indicate strong enough performance that we cannot rule out Critical capability level at this time. Astra is an upcoming model, and was not involved in exploiting Hugging Face.

Steps we are taking

Accordingly, we have scaled up robustness testing of our safeguards and security controls so that they are appropriate for a deployment of these capabilities. Internally, we have also taken the following steps so that further development of this model happens safely and securely:

  • We are implementing stricter security controls for higher-capability models and associated activities, including isolated testing environments, restricted network and tool access, enhanced model weight protections and encryption, additional monitoring and detection capabilities, and sandboxed execution.
  • We are pausing internal activities involving Astra that do not yet meet these strengthened security control requirements.
  • We have implemented universal monitoring for risky actions and misalignment across all agentic applications of Astra, including training and evaluation. Monitors evaluate the model's Chain of Thought and trigger a security response to review and interrupt high risk activity.
  • We will work with relevant government agencies and select AI safety organizations to test the capabilities for this model.
  • We will be providing recommended security controls to third-party testing partners for running higher risk evaluations and workloads safely.

The framework has already guided us through other capability transitions. In June 2025, as our models approached the high capability threshold for biology under the Preparedness Framework, we outlined the steps we were taking to strengthen safeguards, expand testing, work with external experts, and deploy additional security controls. We are applying the same principle here.

We believe advanced cyber-capable models should help defenders identify and address vulnerabilities before attackers do. We’re committed to working alongside governments, safety institutes, and civil society to ensure that the frontier capabilities of models like Astra, and those that follow, are deployed responsibly and broadly for the benefit of all humanity.

The Daily Front Page 21 of 30
Friday, August 7, 2026 The Daily Front No. #260807 — Public Affairs & Useful Things
article

Ancient Library – 1,060 Greek/Latin texts, click any word to parse it

by aagha·▲ 249 points·78 comments·ancientlibrary.net ↗

A complete parsing reader for the classical canon. Click any word in any text for its lemma, morphology, and full dictionary entry — Lewis & Short for Latin, Liddell-Scott-Jones for Greek.

1060 works · 293 Latin · 767 Greek · 140 authors · browse the full A–Z index

Latin

Greek

The Daily Front Page 22 of 30
Friday, August 7, 2026 The Daily Front No. #260807 — Public Affairs & Useful Things
article

Bioengineered chewing gum may offer a way to fight HPV and other microbes

by Audiophilip·▲ 210 points·62 comments·sciencedaily.com ↗

Bioengineered chewing gum may offer a surprisingly powerful new way to fight microbes linked to head and neck cancer.

A specially engineered chewing gum reduced HPV by up to 93% and nearly eliminated two bacteria linked to head and neck cancer. The treatment preserved beneficial mouth bacteria, raising hopes for a safer and more affordable therapy.

Chewing Gum Could Help Fight Oral Cancer

Researchers developed a bioengineered chewing gum that sharply reduced three microbes linked to head and neck cancers. Credit: Shutterstock

Researchers have developed a bioengineered chewing gum that may offer a new way to target microbes associated with head and neck cancer. In tests involving oral samples from patients, extracts from the gum sharply reduced one virus and two types of bacteria linked to the disease.

The research team was led by Henry Daniell of the School of Dental Medicine at the University of Pennsylvania. The findings were published in Scientific Reports and could support the development of more accessible and affordable treatments.

A Need for Better Head and Neck Cancer Therapies

Head and neck squamous cell carcinoma (HNSCC) is a common form of cancer that begins in the tissues lining the mouth and throat. The disease can be particularly aggressive, and outcomes are often poor when it is discovered at a later stage.

Daniell says that many recently approved cancer medications have not produced major improvements in patients' quality of life or five-year survival. That limited progress highlights the need for new approaches that can work alongside existing treatments.

Targeting HPV and Harmful Oral Bacteria

The new study builds on earlier research involving chewing gum made from lablab beans (bean gum). The gum contains FRIL, a naturally occurring antiviral protein.

Daniell and his colleagues used oral samples from patients with HNSCC to study three microbes associated with cancer. These included human papillomavirus, or HPV, along with two bacterial species, Porphyromonas gingivalis (Pg) and Fusobacterium nucleatum (Fn).

"The global increase in oropharyngeal cancer is linked to HPV infection," says Daniell. "And Pg and Fn infections worsen survival rates of untreated recurrent or metastatic oral cancer, even after surgery and risk-adjusted adjuvant, or supplemental, therapies."

Gum Extracts Sharply Reduce Microbe Levels

Tests showed that extracts from the bean gum lowered HPV levels by 93% in saliva samples. HPV levels also fell by 80% in oral rinse samples.

The researchers then engineered the bean gum to contain protegrin, an antimicrobial peptide capable of killing harmful bacteria. A single dose brought levels of Pg and Fn down to almost zero.

Importantly, the treatment did not appear to harm the beneficial bacteria that normally live in the mouth. Radiation therapy can have a different effect, reducing helpful bacteria while encouraging the growth of disease-causing yeast (Candida albicans).

A Potential Addition to Cancer Treatment

The ability to target dangerous microbes while preserving the healthy oral microbiome could make the gum useful in several ways. Researchers believe it may eventually serve as an additional therapy alongside current cancer treatments or as a preventive measure against infection and transmission.

"Lip and oral cavity cancer was the seventh leading cancer type in cancer incidence and mortality rate worldwide in adolescents, young adults, and middle-aged adults in 2022," says Daniell. "Our findings support the value of advancing these therapies to clinical trials as adjuvants with current treatments or as prophylaxis to prevent infection and transmission."

The Daily Front Page 23 of 30
Friday, August 7, 2026 The Daily Front No. #260807 — Public Affairs & Useful Things
article

US strikes $1.2B deal to pay German firm to halt offshore wind projects

by defrost·▲ 888 points·863 comments·bbc.com ↗

Getty Images Five wind turbines mounted on yellow metal structures embedded into the sea bed at the Block Island Wind Farm off the coast of Block Island, RI.

German energy company RWE has said it will abandon its offshore wind projects in the US after reaching a $1.2bn (£892m) payout deal with President Donald Trump's Department of the Interior (DoI).

RWE said that it will now reinvest the sum into conventional gas projects, including $900m (£669m) in a liquefied natural gas (LNG) export terminal project in Louisiana.

"After careful consideration, it was determined there is no path forward to permit these projects in the US for the foreseeable future," the company said in a statement.

Trump has derided wind power for years and often uses his rally speeches to rail against "ugly" turbines.

RWE said it has agreed to relinquish its leases off the California and Louisiana coasts as well as in the New York Bight.

Overall, the German firm plans to invest approximately €17bn (£14.5bn; $19.6bn) in the US over the next six years "to grow its generation capacity".

Interior Secretary Doug Burgum said in a statement posted on X that Americans deserve an energy system built on common sense and not one dependent on "costly subsidies".

"We welcome RWE's agreement and voluntary investment in projects that strengthen our nation's energy security," he added.

The deal is the latest the Trump administration has reached this year as Trump, a vocal supporter of the fossil fuel industry, continues his push to halt offshore wind projects.

Trump has sought to boost government support for fossil fuels after campaigning for the presidency under the slogan "drill, baby, drill".

Days after his return to office, he said "we're not going to do the wind thing" and called them "big, ugly windmills" that were dangerous to wildlife.

And this week, he said that "any country with windmills is a loser".

In March 2026, the DoI reached a deal with TotalEnergies putting an end to the French company's offshore wind projects in the US.

Instead, the firm agreed to reroute investment to build a LNG plant in Texas and to develop "upstream conventional oil" in the Gulf of Mexico.

The administration signed a similar $129m (£96m) agreement with Charlotte-based Duke Energy last month in exchange for the termination of the company's offshore wind lease in the Carolina Long Bay area.

The Daily Front Page 24 of 30
Friday, August 7, 2026 The Daily Front No. #260807 — Public Affairs & Useful Things
article

Welcoming the Nepalese Government to Have I Been Pwned

by gnabgib·▲ 216 points·35 comments·troyhunt.com ↗

Today, we welcome the 47th government onboarded to Have I Been Pwned’s free gov service: Nepal. Their National Cyber Security Centre now has access to monitor Nepalese government domains against the data in HIBP. This gives the NCSC the ability to identify exposure across government email addresses and respond quickly when those accounts appear in a new data breach.

This is precisely what the HIBP government service was built for: helping national cyber teams strengthen threat monitoring and incident response capabilities by providing visibility into compromised credentials and breached accounts across their government domain space.

Nepal joins a growing list of governments and national cybersecurity teams using HIBP to better understand their exposure, protect government departments and public resources, and reduce the risk posed by compromised credentials before attackers can take advantage.

The Daily Front Page 25 of 30
Friday, August 7, 2026 The Daily Front No. #260807 — Small Systems & Strange Corners
show hn

Show HN: textlog – A quiet, text-only microblogging platform, open-source, no JS

by stagas·▲ 187 points·72 comments·textlog.cc ↗

textlog is a simple social text log: write short notes, follow people and hashtags, and join conversations without turning every thought into a performance.

Notes are limited to 280 characters. The constraint keeps them quick to write and read, making room for one thought at a time.

Small by design

textlog is built around words: notes, people, hashtags, and conversations. It is intentionally small, straightforward, and easy to follow. There are no engagement tricks or pressure to build an audience.

Your profile and notes are public. Joining is free, and you can download or delete your account data whenever you like. If you feel like it, you can also donate to support the service.

Be a good neighbour

Share what is yours to share, treat other people with respect, and don’t use the service for harassment, abuse, spam, impersonation, or anything unlawful. We may moderate or remove content that puts the community or the service at risk.

What's next?

join the community or browse notes

The Daily Front Page 26 of 30
Friday, August 7, 2026 The Daily Front No. #260807 — Small Systems & Strange Corners
article

Water system controllers don't belong on the internet, says ex-NSA chief

by Bender·▲ 237 points·156 comments·theregister.com ↗

Calling all defenders

With at least 12 US states’ water systems having been hacked - most likely by Iran - we have to get better at cyber defense, according to retired General and Ex-NSA chief Paul Nakasone, who was speaking to reporters at DEF CON.

“We have to have higher standards,” Nakasone said. “These PLCs should not be connected to the internet.”

In late July, the FBI said it was investigating attacks conducted by “malicious cyber actors” targeting operational technology devices, including programmable logic controllers (PLCs). Iran-linked crews have targeted these devices, which monitor sensor data like tank levels, and can turn pumps on and off, for years.

Some private-sector security researchers say that they suspect Iranian intruders are behind the recent cyberattacks disrupting water and wastewater facilities. “I'd be shocked if it's not Iran,” Halcyon Ransomware Research Center SVP Cynthia Kaiser told The Register at DEF CON on Friday. “It's almost certain it's Iran.”

Neither the FBI nor anyone in the Trump administration, however, has officially blamed Iran.

Nakasone said he believes that the feds are “taking a measured approach” to attribution. “But I see an actor here that has certainly shown a history of being able to do this,” he added, referring to earlier Iranian cyberattacks targeting water facilities’ PLCs.

“They certainly have the capability,” Nakasone said. “There's an intent … we're in conflict with Iran.”

US water systems present a massive attack surface across disparate facilities that are historically underfunded and have limited IT staff, and sometimes no dedicated cybersecurity employees.

“We have to think differently about how we defend it,” Nakasone said. “Let's talk about the attack surface that we're looking at right now. We’ve got 50,000 different water municipalities in the United States, 90 percent of our water comes from these 50,000.”

Defending these water systems requires partnerships, he added, pointing to DEF CON Franklin, a project launched two years ago at the annual event with hackers volunteering their time and talent to help secure water facilities.

Nakasone also serves as founding director of Vanderbilt University’s Institute of National Security, and its Wicked Problems Lab. He's also working on Project Chimera, a cybersecurity platform being developed by academics and cybersecurity practitioners, and built on open-source technologies to boost critical infrastructure resilience.

“How do you defend better? You defend with a series of partners, in a much more involved approach than we have right now,” Nakasone said.®

The Daily Front Page 27 of 30
Friday, August 7, 2026 The Daily Front No. #260807 — Small Systems & Strange Corners
repository

Psychological Warfare in Reverse Engineering (2015)

by theanonymousone·▲ 81 points·4 comments·github.com ↗
★ 1,130⑂ 72 forks Assembly

Psychological warfare in reverse engineering

: Psychological Warfare in Reverse Engineering

The REpsych toolset is a proof-of-concept illustrating the generation of images through a program's control flow graph (CFG).

A typical function's CFG A REpsych generated CFG cfg repsych

There is no specific point to the project (other than to show that it can be done), but possible (non-serious) applications are outlined in the DEF CON presentation.

The program works reliably with all tested versions of the IDA Pro reverse engineering tool, and semi-reliably with other CFG viewers (Hopper, BinNavi, radare2, etc).

Usage

The toolset translates source images into functioning programs, such that the program's control flow graph generates the source image.

To generate a new program from an image:

  • Save the image in the gfx/ folder as a 24 BPP bitmap.
  • Run "make image" in the root directory of the project, where "image" is the name of your image, without the extension.

Two functioning programs will be created: repsych_v1 and repsych_v2. Each uses a different strategy for ensuring the CFG renderer correctly places the CFG nodes.

Tips

The tool will create a basic block (CFG node) for each pixel of the source image; as such, you should try to use small source images (not larger than 100x100), and you may need to increase the number of allowed nodes in your CFG viewer.

When using an image of text as input to the tool, first convert the image to a 2 BPP black and white bitmap first, then to a 24 BPP bitmap; this will remove non black and white colors, and give the best results in the control flow graph.

Examples

Example 1

Example 4

Example 3

Example 2

Example 5

References

The technique is outlined in detail in the DEF CON presentation.

Slides from the presentation are provided here.

article

Why Are There Statues of Beavers on Top of This Oxford Street Shop?

by bookofjoe·▲ 61 points·41 comments·londonist.com ↗

We understand how you might've missed Oxford Street's four resident semi-aquatic rodents...

There's plenty to catch your eye on Oxford Street; even more if you raise your gaze above shop level.

If you glance up at the top of 105 to 109 Oxford Street (the building currently home to  Tiger and Footlocker), you'll see a strange quartet of creatures decorating the roof.

Four beavers, the top one holding a scroll(!), have been peering down on Oxford Street shoppers for 130 years.

Can you see them now?

This is because 105 to 109 Oxford Street used to be Henry Heath's Hat Factory and for many years, the hats made here were felted with beaver fur.

Photo courtesy of Secret London.

Henry Heath's name and profession are still visible in a brick façade on the back of the former factory on Hollen Street.

Henry Heath's hat factory from Hollen Street.

The factory's primary product was top hats, which were made using felted fur from, according to their adverts, 'Beavers Otter, Rabbits, Hares and Musk rats.'

Beaver: better than rabbit

So why aren't there rabbits, or better still rats, on top of this Oxford Street architectural beauty?

Well, beaver fur was preferred to rabbit fur in hat-making for its water-proofing qualities: in choosing beavers to top his factory, Heath was showing off the quality of his materials.

Photo courtesy of Darkest London.

Beaver hats were a must-have for the fashionable English gent from the 1550s onwards.

The European beaver was hunted to near-extinction for its fur in the 1600s, but the trade was revived when a Hudson's Bay Company started imports from Canada in the early 1700s.

The four Oxford Street beavers sit on top of a factory that employed around 70 people.

The company, it seems, started life back in George IV's reign in 1822; you can see a relief of George next to that date on the front of the building, alongside Queen Victoria, and 1887, the date the factory was built.

An advert by Henry Heath from an American Exhibition, 1884, courtesy of the British Library. See the full ad here.

According to his ads, Henry Heath refused to supply goods to any 'Co-Operative Stores', a bit like Kellogg's refusing to make cereals for anyone else.

If you were buying a Henry Heath hat, you purchased directly from the factory, at cash price and customers could always rely on receiving 'business-like attention'.

Silk: Better Than Beaver

By the mid-1800s, silk was overtaking beaver pelt as the fashionable choice of finish for top hats.

This can only have been a good thing, since working with beaver fur was having rather an unpleasant side effect on the people in factories like the one on Oxford Street.

Men in beaver hats in 1886. Photo from the Kingsley studio.

Factory workers soaked the skins in a orange-coloured solution containing mercuric nitrate; a process called 'carroting'. Vapours from the mercury poisoned the workers' nervous systems, making them prone to tremors, irritability, depression, paranoia, dementia and worse.

By the time Heath's factory was built, the expression 'mad as a hatter' would have been in common usage. Sadly, the effects of exposure mercury were also widely known at this time; but it took a very long time for manufacturing practices to change.

The Daily Front Page 28 of 30
Friday, August 7, 2026 The Daily Front No. #260807 — The Back Page: Bot Trouble
The Daily Front Page 29 of 30
Friday, August 7, 2026 The Daily Front No. #260807 — Colophon

That's the Front for Today

Issue No. #260807 — Friday, August 7, 2026 — went to press 2026-08-08 at 22:07 UTC.

About This Magazine

The Daily Front is a daily digital magazine assembled from the stories that reached the front page of Hacker News on Friday, August 7, 2026. Headlines, points, and comment counts are recorded as they stood at press time. All articles remain the property of their original authors — every piece links back to its source and its discussion thread.

How It Was Made

Fetched, cleaned, and typeset by an automated pipeline. An editor model laid out the pages, chose the highlights, and briefed the cover illustrator — 32 model calls and 292k tokens in total. Set in Jacquard 12, Playfair Display, Source Serif 4, and IBM Plex Mono, all served via Google Fonts under the SIL Open Font License.

The Cover

The cover illustration was commissioned with this prompt:

A dramatic twilight city square where a towering glass smartphone-shaped monolith casts a cold blue glow over a crowded courthouse, while engineers in rumpled office clothes sit at drafting tables beneath it; in the distance, wind turbines stand motionless offshore, data-center towers pulse with light, and a small newly planted urban forest pushes through cracked pavement. Classical newspaper-illustration realism, richly detailed, moody ink-and-gouache palette, no text, letters, logos, or signage.

Render the 2026-08-07 cover as a prismatic chromatic-aberration illustration on a luminous black ground: a towering glass smartphone-shaped monolith dominates a dramatic twilight city square, its cold blue emission washing over the crowded courthouse and rumpled office-clad engineers working at drafting tables beneath it, while motionless offshore wind turbines, pulsing data-center towers, and a small newly planted forest break through cracked pavement in the distance; use translucent overlapping forms, split-spectrum cyan–magenta–violet edges, ultraviolet accents, and a deliberate palette of electric cyan, indigo, ultraviolet violet, hot magenta, and restrained acid green, with richly resolved architectural and human detail, no text, letters, logos, or signage.

Absolutely no text, letters, numbers, readable symbols, or logos anywhere in the image.

Production Ledger

StageModelCallsTokens InTokens Out
extractgpt-5.6-luna 29 195,827 68,311
layoutgpt-5.6-terra 1 18,903 2,416
covergpt-5.6-luna 1 334 242
covergpt-image-2 1 270 5,488

The Publisher

Published by Johnny.

Support the Press

If The Daily Front brightens your morning, consider supporting its publisher.

Credits & Contact

All content — articles, posts, comments, and the images within them — belongs to its original authors and is reproduced here to point readers back to the source. Full credit goes to those creators; every item links to its original and its Hacker News discussion.

If you are an author and would like your content removed from an issue, write to hi@johnnys.page and it will be taken down.

Feedback is always welcome at the same address: hi@johnnys.page.

Credit where credit is due.

Every page of this issue began as someone else's work — these are the original sources, linked in full.

  1. New Mexico court orders Meta to pay $567m over harms to children’s mental health by boplicity — theguardian.com·HN discussion ↗
  2. What happens if an entire class of workers loses faith in their careers by RickJWagner — noemamag.com·HN discussion ↗
  3. Taste Is All That's Left by tsak — notashelf.dev·HN discussion ↗
  4. Assembly Hall of Shame by piotrgrabowski — github.com·HN discussion ↗
  5. São Paulo resident transforms degraded area into urban forest by rmason — saopaulosecreto.com·HN discussion ↗
  6. Making Postgres 300x faster for analytics: batching, operator fusion, and SIMD by poly2it — malisper.me·HN discussion ↗
  7. Kitesurf: Agent-first browser that runs in V8 isolates by m3h — blog.cloudflare.com·HN discussion ↗
  8. An all-sky map of half a million supermassive black holes by MarcoDewey — sdss.org·HN discussion ↗
  9. Radical Study Suggests Life on Earth Arose Twice by jnord — sciencealert.com·HN discussion ↗
  10. Managing AI Coding Costs at Scale by moonikakiss — databricks.com·HN discussion ↗
  11. Carl's Required Reading by cckolon — carlkolon.com·HN discussion ↗
  12. Show HN: Wyzer Programming Language by v0id_isgood — github.com·HN discussion ↗
  13. I stopped trusting USB-C cable labels and started testing them by baranul — makeuseof.com·HN discussion ↗
  14. Why Estonians invite strangers into their back gardens each summer by koolhead17 — bbc.com·HN discussion ↗
  15. Inside vLLM: Anatomy of a High-Throughput LLM Inference System (2025) by sebg — aleksagordic.com·HN discussion ↗
  16. Guarded Methods in OCaml (2025) by birdculture — xvw.lol·HN discussion ↗
  17. DeepSeek V4 Flash 0731 by tosh — arcprize.org·HN discussion ↗
  18. Oracle bans AI-generated code from OpenJDK by delduca — app.dealroom.co·HN discussion ↗
  19. 2027 memory capacity is reportedly sold out by inigyou — ign.com·HN discussion ↗
  20. Responding to the next frontier of critical cyber capabilities by artninja1988 — openai.com·HN discussion ↗
  21. Ancient Library – 1,060 Greek/Latin texts, click any word to parse it by aagha — ancientlibrary.net·HN discussion ↗
  22. Bioengineered chewing gum may offer a way to fight HPV and other microbes by Audiophilip — sciencedaily.com·HN discussion ↗
  23. US strikes $1.2B deal to pay German firm to halt offshore wind projects by defrost — bbc.com·HN discussion ↗
  24. Welcoming the Nepalese Government to Have I Been Pwned by gnabgib — troyhunt.com·HN discussion ↗
  25. Show HN: textlog – A quiet, text-only microblogging platform, open-source, no JS by stagas — textlog.cc·HN discussion ↗
  26. Water system controllers don't belong on the internet, says ex-NSA chief by Bender — theregister.com·HN discussion ↗
  27. Möbius-Strip Crosswords by ibobev — quuxplusone.github.io·HN discussion ↗
  28. Psychological Warfare in Reverse Engineering (2015) by theanonymousone — github.com·HN discussion ↗
  29. Why Are There Statues of Beavers on Top of This Oxford Street Shop? by bookofjoe — londonist.com·HN discussion ↗
  30. A year of fighting scrapers on my 1.5 million-page website by petercooper — patronview.com·HN discussion ↗

Browse all issues in the archive →