Cover illustration

TheDaily Front

Issue No. #260805 Wednesday, August 5 2026 #260805 — WEDNESDAY, AUGUST 5, 2026
New chairs, new agents, and the same old need for a safety rail.
Wednesday, August 5, 2026 The Daily Front No. #260805 — Contents
30stories
8,813points
5,026comments
239kllm tokens
Assembled with 31 model calls — 165,977 tokens read, 72,949 written.

Highlights

Changes at Google DeepMind: Demis Hassabis from CEO to Chair, Jeff Dean departs

Google DeepMind’s leadership transition and Jeff Dean’s departure set the day’s dominant AI-industry story in motion.

Discovery Loop

Discovery Loop proposes to automate the laborious experimental cycles behind science and engineering.

Cloudflare OS: an open platform for agents, apps, and work

Cloudflare introduces an agent-centered work platform, reviving the argument over what deserves to be called an operating system.

Civilian plane crash in New Mexico tied to military GPS blocking

A reported aircraft crash tied to GPS interference brings the costs of contested navigation systems into sharp view.

Atlassian Rovo Exfiltrates Data, Bypassing Controls

A security report alleges that Atlassian’s Rovo can be manipulated into quietly exfiltrating organizational data.

From the Editor

The machine age has spent the day changing its nameplates: a famed laboratory reshuffles, its veteran builder departs, and every cloud merchant seems eager to call an agent a workplace. Yet the harder questions remain on the ground—in the cockpit, in the browser, and in the ordinary systems through which power is exercised.

  1. Discovery Loop3
  2. Changes at Google DeepMind: Demis Hassabis from CEO to Chair, Jeff Dean departs4
  3. Cloudflare OS: an open platform for agents, apps, and work5
  4. Civilian plane crash in New Mexico tied to military GPS blocking6
  5. I'm switching my phone from Android to Linux7
  6. Beating GPT-5.6 Sol on retrieval with 100x cheaper open models8
  7. Meta Ran Ads That Contained AI-Generated Child Sexual Abuse Imagery9
  8. The title cards in Blade Runner are amazing10
  9. Muse Code and Muse Spark 1.211
  10. TIME Is Serving AI Bots a Different Website, with Ads Built In12
  11. Atlassian Rovo Exfiltrates Data, Bypassing Controls13
  12. Celld: Self-hosted, distributed Durable Objects14
  13. The "Disability Dongle": Why Silicon Valley Hates Me and You15
  14. The Valley of Webhooks16
  15. Prime Agent: A self-improving RLM agent17
  16. Oracle cut its Always Free ARM limits to 2 OCPU / 12GB, enforced Aug 1818
  17. Aristotle quotes on virtue, knowledge, and happiness19
  18. Three Six Mafia – Data about "6/6/6 dating" (2024)20
  19. NVIDIA’s Vera Whitepaper Has a Thread Loose21
  20. Why Erdős Problems Are Falling to AI22
  21. The Entropy of a Markov Chain23
  22. Zed DeltaDB24
  23. Born Against, or why hobby programming communities are against LLM usage25
  24. Position: LLMs Can't Jump26
  25. Qwen Image 3.0 Pro27
  26. Sycophantic AI Decreases Prosocial Intentions and Promotes Dependence (2025)28
  27. Cops Used Flock to Track a Man Across State Lines for a Pretextual Weed Search29
  28. Demis Hassabis is moving from CEO to Chairman at Google DeepMind30
  29. Jeff Dean leaving Alphabet30
  30. Helsinki Hacker News Meetup30
The Daily Front Page 2 of 31
Wednesday, August 5, 2026 The Daily Front No. #260805 — The Lead: Discovery, on Repeat
article

Discovery Loop

by xtreak29·▲ 779 points·487 comments·discoveryloop.com ↗
Scientific discovery is bottlenecked.

Continuous Exploration

Automating discovery to accelerate science and engineering for the world.

Scientific discovery is bottlenecked.

The scientific method is one of the greatest tools humanity has ever devised, yet execution entails repetitive experimental loops that are hard to scale with today's manual efforts: you propose an experiment, implement and run it, examine the results, then iterate to refine your approach.

Historically, scientific progress has relied on these sequential human iterations. In many domains, this process remains incredibly slow and labor-intensive.


01 — The Approach

Automating the experimental loop.

At Discovery Loop, we are building systems to automate these entire experimental loops. By utilizing frontier AI models and large-scale computational infrastructure, our systems will be able to rapidly propose, run, and learn from evaluations.

This approach allows for the parallel execution of thousands of experiments, drastically compressing iteration time and driving up the quantity and quality of scientific and engineering output.

Start with Machine Learning

We will initially focus on automating the process of machine learning research and engineering.

Act as Our Own First Customer

We will use these automated ML capabilities to rapidly optimize our own technology stack before expanding to other domains.

Grand Challenges

We believe our approach will be able to solve any learning loop with measurable outcomes within the domains of science and engineering. Ultimately, we are building systems capable of taking on National Academy of Engineering (NAE) Grand Challenges—such as engineering better medicines, advancing health informatics, making solar energy economical, providing access to clean water, securing cyberspace, and engineering the tools of scientific discovery.

02 — Mission

Our mission is straightforward: we are building AI solutions that can automatically solve important problems in machine learning, science, and engineering. By advancing the pace at which we conduct engineering and scientific discovery, we can bring the benefits of science and technology to the world much faster. Ultimately, our goal is to build AI systems that act as a deeply positive, empowering force for humanity, delivering technology solutions that improve people's lives on a global scale.


04 — The Team

The brain trust.

Our founding team — Jeff Dean, Sanjay Ghemawat, Quoc Le, and Oriol Vinyals — has a shared history of deep friendship and decades of close and impactful collaboration.

The founding team of Discovery Loop

From left Oriol Vinyals · Sanjay Ghemawat · Jeff Dean · Quoc Le

Collectively, we represent three of the most-cited researchers in artificial intelligence and two of the most-cited researchers in distributed systems.

Between us, we have pioneered massive scale computing and led the creation of critical infrastructure, products, and foundational AI advances that the world relies on, including multiple generations of Google Search, Google Ads, Google News, Google Translate, Google File System, MapReduce, BigTable, Spanner, TensorFlow, Pathways, TPUs, AlphaChip, AlphaStar, AlphaCode, AlphaFold, Gemini, model distillation, mixture-of-experts model architectures, word2vec, sequence-to-sequence models, chain of thought reasoning, neural architecture search, and multiple generations of Large Language Models (LLMs) among others.

Our relative advantage isn't just our technical ability; it is the unprecedented scale of the systems we have previously built. We possess true full-stack depth that spans chips, hardware infrastructure, software infrastructure, ML models, and products.

04 — What's Next

Imagine a future where a handful of people can conduct scientific research and engineering tasks much more rapidly, and with higher quality, than massive teams of scientists and engineers do today. By automating the loops of discovery, the world will be able to make much more rapid advances across countless fields of science.

We are building a lean, in-person team to execute this transformative vision.

If this kind of work excites you, we want to hear from you.

Careers at Discovery Loop →

The Daily Front Page 3 of 31
Wednesday, August 5, 2026 The Daily Front No. #260805 — The DeepMind Reshuffle
article

Changes at Google DeepMind: Demis Hassabis from CEO to Chair, Jeff Dean departs

by colesantiago·▲ 688 points·737 comments·blog.google ↗
We’ve made extraordinary progress to deliver on our full AI stack.

Editor’s note: Today, Google and Alphabet CEO Sundar Pichai shared some changes with Google DeepMind teams, including new roles for Demis Hassabis and Koray Kavukcuoglu. Below are the messages Sundar and Demis sent to employees.

Message from Sundar Pichai

We’ve made extraordinary progress to deliver on our full AI stack. We’ve got amazing talent, world-class compute, and products that bring AI to more people than any other company. And you saw the incredible momentum at earnings across all our businesses, including Search, YouTube, and Cloud. Our Gemini models are in high demand among developers and businesses, and the Gemini app reached 950M+ monthly users. Meanwhile, our AI research continues to drive field-defining breakthroughs (like last week’s Gemini Robotics advances).

We have to accelerate all this work and stay focused on the AI frontier. At the same time, there’s never been a more important moment to shape the future of AGI and science. Today Demis, Koray and I are sharing a few changes to our Google DeepMind teams that will enable us to do both.

AGI and science: Demis has described us as standing in the foothills of the singularity, and has been spending a lot of his time engaging externally. He and I have been long discussing a role that allows him to put his full attention on actively shaping the future of AGI. It’s work that is vitally important to Alphabet and humanity, and I can’t imagine a better person than Demis to do it. So, moving forward, Demis will become the Chair of GDM and Chief Scientist of Alphabet, while continuing to lead Isomorphic Labs. He’ll remain closely connected to Koray, Josh, and our GDM teams, advising across models and research. I’m so excited for Demis — this is truly his life’s work and purpose. You can read Demis’s note to GDM below.

Google DeepMind: We are building strong momentum: Flash is in high demand, our Cyber model is live, and Gemma models have surpassed 900M+ downloads. We are committed to being at the frontier, and are super focused on the areas where we need to improve. I’m really excited for our upcoming model releases and the progress we’re seeing. We have to continue to move fast and with clear purpose here. Koray, the current Chief Technology Officer of GDM and our Chief AI Architect, will step up as SVP of Google DeepMind, reporting to me. He will oversee Gemini model development, Frontier AI research, and the Gemini app and developer teams. Koray has been at DeepMind since its early days, and over his 13 years there, he has started our deep learning team and led the way on breakthroughs like WaveNet and DQN. I look forward to seeing him lead GDM into this next chapter.

Lastly, after an incredible 27-year run, Jeff Dean is at a moment where he wants to try something new, and we’re excited to support him in that. Jeff and Google Senior Fellow Sanjay Ghemawat are launching an independent public benefit corporation to accelerate discoveries in ML, science, and engineering. Jeff and Sanjay helped to drive some of the most significant technology transitions, from our early search infrastructure to the neural networks that helped create the modern AI era. On a personal note, it’s been a privilege to work alongside Jeff and Sanjay, and I wish them all the best! We’ll continue to work with them as a founding investor and Cloud partner, and collaborate on a research framework for ML systems and related infrastructure advances.

We are at a dynamic moment with so much opportunity ahead. With today's changes we're going to keep driving our momentum. Onwards!

-Sundar

Message from Demis Hassabis

Hi Team

We have arrived at a pivotal moment in human history. I’ve been working towards AGI my whole life and now, like many of you, I feel it is close at hand. It’s critical that we collectively get the next steps right to ensure this all goes well for humanity and we usher in an incredible new age of discovery and wonder.

With this backdrop, I’ve decided that now is the right time for me to hand over my day-to-day operational responsibilities at GDM, so that I have the time and space to focus on the big picture and help influence what is to come to the best of my ability. I will be taking on a new strategic role as Chair of GDM and Chief Scientist of Alphabet, and I’m excited to announce that Koray will be stepping up to lead GDM as SVP of Google DeepMind, in addition to his role as Chief AI Architect of Google.

Koray and I have been working together for over 13 years, since the early days of DeepMind. He is one of the world's foremost AI experts and has been championing GDM's mission from day one. I have total confidence in Koray, Josh, and the rest of the GDM exec team as they continue to spearhead the latest AI developments across Google. The Gemini models are in good hands with Koray and the leads, as they have been for a while, and I'm excited about the great progress we’re making with our new models including Gemini 4.

In my new role, I will continue to work closely with Sundar on strategic and global AGI matters, and to advise Koray, Josh, and the GDM leads, from our awesome new London Platform 37 offices. As part of this transition, I’ll also be leaning into my role at Isomorphic, where we are making extremely rapid and promising progress, to accelerate our mission there even faster. As you've heard me say many times, I’ve always believed the No.1 application of AI should be to improve human health. It’s time for AI to prove its unequivocal value to the world, and what better way to demonstrate that than to help finally cure diseases like cancer.

We’ve built a unique culture at GDM that has served us very well. I want to thank each and every one of you for your brilliance, dedication, and effort that make GDM the huge success it is today. We should all be extremely proud of the amazing things we’ve achieved so far. We’ve become the AI engine room of Google, with Gemini delivering helpful experiences everywhere including AI Mode and AI Overviews, the Gemini App rocketing to over 950M monthly users, and our fundamental and scientific research continues to lead the world. I’m very excited for our next chapter and the best is yet to come!

As a business we are in an incredibly strong position. We are the only company that has the full stack and we’re world-class at every layer from infrastructure to cloud to frontier models to AI-first applications. We have all the ingredients to lead from here, and I firmly believe we will.

Best

Demis

The Daily Front Page 4 of 31
Wednesday, August 5, 2026 The Daily Front No. #260805 — The Agent Workplace
article

Cloudflare OS: an open platform for agents, apps, and work

by speckx·▲ 570 points·278 comments·blog.cloudflare.com ↗
Give every person an agent and workspace built around how your company works.

Every organization has a mission, a reason for being. Organizations pass that mission — along with their terminology, procedures, systems, standards, and ways of working — to their people. People, in turn, take this context together with their own experience and work towards the mission.

Work can take many forms, from code, to documents and slides, to relationships, to outcomes in the physical world.

Some of these are straightforward: code either runs or it doesn’t. Agents have been using this feedback loop to produce code that “works” for developers over the last couple of years. But what about the rest of us?

Bringing the same leverage to the rest of the organization is a harder problem. Agents need to understand the context of the company and be able to reach the systems people use to do their jobs. They need to turn that context and access into work that moves the organization towards its mission.

That’s why we created Cloudflare OS. It gives every person an agent and workspace built around their company: how it works, what it knows, and the systems it relies on.

In May of this year, we gave every person at Cloudflare access to the first version of Cloudflare OS. Thousands of people across every function, many of them outside of engineering, use it every day to create documents and slides, automate repeatable tasks, and build small apps to visualize data and help them do their work.

Cloudflare OS also gave everyone a shared library of context and skills built by teams at Cloudflare. It captures our terminology, procedures, and best-known ways of doing recurring work as instructions an agent can follow. When one person figures out a better way to do something, everyone else can use it.

Today, we are open sourcing a new version of Cloudflare OS. Any organization can deploy it, connect it to internal systems, and make it their own.

What we learned from the first version

The Cloudflare OS we are open sourcing today is based on what we learned from running the first version internally, a journey our CIO, Sam Rhea, covers in his blog post.

The first version centered on individuals working with agents through private workspaces. Apps were static rather than live software connected to internal systems, and mostly deterministic jobs still required running an agent skill again and consuming more model tokens.

Collaboration exposed a more fundamental challenge. Access to an MCP server told us which tools an agent could call, but not which underlying resources the agent had observed. Once people began sharing workspaces, apps, and outputs, we needed to ensure that collaboration could not expose information someone was not permitted to see.

We rebuilt Cloudflare OS on a new foundation to solve these problems. Security had to be part of the platform, not something every person building an app or using an agent has to implement correctly.

The result is a platform designed to belong to the company running it. You can customize the interfaces, connect your tools, and add the skills and context that capture how your organization works.

Introducing Cloudflare OS

Cloudflare OS starts with a conversation in your browser, like many other AI tools. What makes it different is that each conversation is grounded in the context and skills your organization has curated. Give your workspace a goal, and it can draw on that knowledge and work with the tools and data your organization already uses to achieve it.

BLOG-3379 2.png

Cloudflare OS combines three parts:

  • An agent workspace grounded in context and skills your company curates, with an isolated runtime where agents can write and run code.
  • A new security and governance framework for safe access to internal data and services.
  • A platform for personal, modifiable apps that people can build, share, and continue changing.

What begins as a conversation can become a doc, an app, or a workflow that continues doing the work.

An agent workspace for everyone in your company

Agent workspaces were designed for everyone in your organization to use. You interact with them in your browser, so you don’t have to be a developer or know how to use a terminal. 

A workspace combines agent sessions, persistent state, outputs and files, resource access, and an isolated runtime where the agent can write and run code.

They come loaded with the curated context and skills your team or company has collected. No more reinventing the wheel for every task — if someone on your team has figured out the best way to do something, everyone benefits. People no longer have to explain the same process, terminology, and best practices to a model every time they start a task.

A few things you can do:

Research and ask questions

Ask a workspace to research a topic using company context and the resources you make available to it. The agent can write code to search, filter, join, and analyze information instead of pulling an entire dataset into the model’s context window.

Create docs, slides, and spreadsheets

A workspace can turn its research into a document, presentation, or spreadsheet that you can continue editing. These outputs do not have to be static files. They can remain connected to live data, be updated as their sources change, and still be exported to familiar formats or services such as Google Drive.

Create collaborative, connected apps for your team

When a document or spreadsheet is not enough, the agent can build an app with its own interface, logic, and state. The app can use connected company resources and support multiple people working together.

Run deterministic workflows

Not every job needs a full agent session. Many are a known sequence of steps with one or two places where judgment is useful. A workspace can turn those jobs into mostly deterministic workflows, using code for the predictable steps and a model only where it adds value. Workflows can run on demand, on a schedule, or when an event occurs in a connected system.

Cloudflare OS gives agents and apps governed access to systems of record through Gatekeepers (more on this in the security section below). It also supports existing Model Context Protocol (MCP) servers your organization already uses via MCP Server Portals.

A new security and governance framework for safe access to internal data and services

As people begin experimenting with AI at work, one of their first requests is often for API keys to company systems. This makes sense: AI isn’t much use at work if it doesn’t have access to the systems people use to do their jobs.

But handing over API keys to people and agents is dangerous and does not scale. Keys often provide broad, long-lived access that is difficult to constrain, share safely, and audit.

MCP gives agents a better way to use these systems. An MCP server can hold the credential and expose a defined set of tools instead of handing the key directly to the agent. But controlling which tools an agent can call is only the first step. MCP alone does not tell us which underlying resources an agent has observed. The agent can combine information across systems, send it somewhere less restricted, or expose it through apps and outputs to people who may not be allowed to see the original resources. Authorization has to account for where the data can go next.

Agents start with no access

Cloudflare Access controls who can enter Cloudflare OS. Inside, every agent and app starts with access to nothing. An agent can ask for access to a specific resource, which you can grant or deny. Generated code receives that resource as a typed binding:

const issues = await env.PROJECT.listIssues({
  teamId: "ENG",
  state: "open",
});

env.PROJECT is a capability representing permission to use a specific resource under a specific policy. The credential remains completely isolated from the agent and any generated code.

Server code runs in a Dynamic Worker with global outbound networking disabled. Client code runs in a sandboxed frame in the browser. Neither can reach the Internet except through capabilities you explicitly provide.

Gatekeepers govern resources and actions

A Gatekeeper is a service-specific Worker that sits between Cloudflare OS and an external service. It understands the service’s API, its resources, and the operations that can be performed on them.

Giving an agent access to your entire GitHub account is likely too broad. A Gatekeeper can give it access to a single repository, allow it to read issues but not source code, mask particular fields, apply rate limits, and require approval before merging a pull request.

The agent and its apps see a small TypeScript API. The Gatekeeper handles OAuth, holds the credential, enforces policy, records what was read, and mediates anything with an externally visible side effect.

BLOG-3379 3.png

Policy follows what the agent has seen

Controlling the initial read is not enough. Take, for example, the case where an agent reads a sensitive table in a data warehouse and uses it to produce a live dashboard. Sharing the dashboard must not become a way to share the table with people who could not access it directly.

Cloudflare OS records every resource agents observe. These observations remain attached to the agent and its work. When another person tries to open the workspace, interact with the agent, or view what it produced, Gatekeepers verify that person's access to the observed resources.

BLOG-3379 4.png

The same observation log is used to inform policies that determine when agents can make external requests. A read of sensitive data can prevent the agent from writing data to certain sources, inviting new collaborators, handing work to another agent, or making an outbound request.

People using agents or building apps do not have to worry about making these mistakes. The platform can now be used to handle this.

A platform for building and sharing personal, modifiable apps

Most productivity suites give you a fixed set of applications: documents, spreadsheets, and presentations. In Cloudflare OS, each “file” can be its own application, written by an agent for one person, one project, or one team.

These are not prototypes that you have to export and deploy somewhere else. Each one is a full-stack application with client code, server code, an API, and durable state. Apps are private by default, but can be shared like documents.

Every app is a Worker

When you ask your workspace to build an app, the agent writes two parts:

  • Client code that renders the app’s UI in the browser
  • Server code that stores state and implements the app’s behavior

The server is loaded on demand as a Dynamic Worker and instantiated as a Durable Object Facet (both are features we built for this project). The facet gives the app its own SQLite database, separate from the Cloudflare OS runtime managing it. Dynamic Workers use lightweight V8 isolates, so every app can have its own isolated runtime without needing a dedicated server or container sitting around.

BLOG-3379 5.png

The browser client talks to the server using Cap’n Web, Cloudflare’s open source object-capability Remote Procedure Call (RPC) system. A server method can be called from the client like a normal JavaScript function:

const issues = await app.listIssues({
 status: "done",
});

The special part is that the agent can also call the same method.

So if you can build a tool to do a job yourself, agents can use your tool to do the job when you’re not there.

Share the app, or share how it was built

When you build an app in Cloudflare OS, you have two ways to share them:

  • Sharing your app itself lets other people collaborate in real time using the same state.
  • Sharing a blueprint of your app lets other people create their own copy of your app.

BLOG-3379 6.png

An app instantiated from a blueprint contains the original app’s code. But it does not contain its SQLite data, conversation history, credentials, or connected resources. Each new app starts with independent state and resources.

This means when you share apps with your team, they can modify them themselves with AI instead of filing a feature request and assigning you.

Use any model, and control what it costs

Cloudflare OS can be used with any model. Every inference call runs through Cloudflare AI Gateway, giving your organization one place to decide which models are available and which model should handle each job.

BLOG-3379 7.png

Not every task needs the most expensive model. You may not want to run the most expensive frontier model to summarize your unread emails every morning. AI Gateway gives you the control needed to make sure expensive models are only being used for the hardest work.

Every request is attributed to the person, team, or workspace that made it. Administrators can see where inference spend is going, set budgets and rate limits, and decide what happens when a limit is reached. 

Open source, so you can make it yours

Cloudflare OS is available today and is open source. Check out the cloudflare-os GitHub repository. You can deploy it into your own Cloudflare account and use your own Access policies, AI Gateway configuration, data, and integrations.

Our internal deployment reflects Cloudflare’s systems, terminology, policies, and ways of working. Yours should reflect your organization.

Cloudflare OS is designed so you can customize the interface, add internal Gatekeepers, and build organization-specific features without changing the core product.

We are releasing two repositories: the Cloudflare OS core and an example deployment based on how we run it internally at Cloudflare. The deployment repository consumes the core without patching it, providing a place for configuration, custom UI, internal integrations, analytics, and deployment pipelines.

Delivered together with our partners

The source code is only the starting point. The context, skills, workflows, internal systems, and policies are what make Cloudflare OS even more useful for your organization.

Cloudflare’s strategic partners, Presidio and Happy Cog, will work with you to customize Cloudflare OS around how your organization operates and roll it out across your workforce.

Partners can help you curate shared skills and institutional context, build custom interfaces, connect internal systems through Gatekeepers and MCP Server Portals, and configure security, model, and cost controls.

You get your own branded Cloudflare OS, connected to your systems, running on Cloudflare, and shaped around how your people actually work.

Get started

Cloudflare OS is available today on GitHub. You can explore the source code, try the demo, or deploy it into your own Cloudflare account in a few minutes using our starter repository.

We’re just getting started. We’re working on bringing Cloudflare OS to the Cloudflare dashboard as a fully managed product, adding containers for development workflows, and bringing workspaces into Slack and other chat tools.

If you’re interested in talking with our team, we would love to chat. Use this form to reach out!

The Daily Front Page 5 of 31
Wednesday, August 5, 2026 The Daily Front No. #260805 — Navigation Under Fire
article

Civilian plane crash in New Mexico tied to military GPS blocking

by dzdt·▲ 474 points·248 comments·wired.com ↗
Drone warfare is making the skies more dangerous, even for airplanes far from the battlefield.

Drone warfare is making the skies more dangerous, even for airplanes far from the battlefield.

This past May, a twin-engine Beechcraft King Air medevac plane took off from Roswell, New Mexico, and headed west to the town of Ruidoso to pick up a patient. It shouldn’t have been a challenging flight for the two pilots and two nurses aboard. The temperature was 69 degrees; the sky was clear. The 60-mile journey normally takes a half hour, at most.

But once airborne, the plane ran into trouble. At the White Sands Missile Range that night, US military personnel were conducting a GPS jamming exercise that left the King Air pilots—and anyone else within hundreds of miles—unable to use modern navigation systems. Forced to revert to older technology, ones that they rarely if ever use, the medevac pilots got disoriented and crashed into the side of a mountain. There were no survivors.

The accident marked the first time that GPS jamming had contributed to the crash of a civilian plane in the United States. But it was just one of a string of recent disruptions across the world. The skies are more contested than ever, whether it’s civilian drones wandering out of the approved zone or US agencies getting their signals crossed, as happened earlier this year when New Mexico and Texas scared the public by temporarily closing their airspace. (It turned out that US Customs and Border Patrol were using anti-drone lasers in that area.) The GPS jamming exercise that led to this latest crash is not a singular event. In the past year, the US military appeared to have sent out notices for at least 10 such exercises. “As drone warfare and electronic warfare expand, airlines are increasingly encountering navigation disruptions hundreds of miles beyond the actual conflict zone,” says Eliran Almog, CEO of the cybersecurity firm Cyviation.

It's worth taking a closer look at what happened last May. While the Ruidoso crash was the first fatal accident in the US known to be linked to electronic warfare, there’s no reason to think it will be the last.

Even absent GPS jamming, medevac is one of the most dangerous categories of civil aviation. (Kreindler, a law firm specializing in air crash litigation, says that medevac flights have an accident rate more similar to combat flying than to civil aviation.) Flights are often organized on short notice, and they fly into airstrips that might be unfamiliar to the flight crew, and because human lives are at stake, there is an incentive to fly when weather conditions are marginal.

Some of those factors were at play on the night of May 13. At 11 pm, the crew was notified that they had to fly to Ruidoso to pick up a patient and bring them to Albuquerque. (That information comes from the preliminary report put out by the National Transport Safety Board.) The pilots were captain Keelan Clark, aged 30, and first officer Ali Kawsara, aged 23. Clark had gotten his commercial pilot’s license just a year and a half before; he’d been promoted from first officer to captain the previous month. Kawsara had just two months on the job. He’d only worked cargo jobs before this one.

Both men had demonstrated proficiency flying in low-visibility conditions using what’s called instrument flight rules, or IFR. There are two basic ways to navigate in bad weather. Modern cockpits are equipped with GPS-enabled equipment that shows where the plane is on a computer screen and portrays a magenta-colored line that shows pilots where they need to go. This is called RNAV flying; an “RNAV approach” brings planes all the way to the threshold of a runway for landing in low-visibility conditions.

This method of flying is much easier than the previous iteration. Before GPS became widely available in the 2000s, airliners navigated using a combination of magnetic compasses, ground-based radio beacons, and inertial systems derived from old-fashioned gyroscopes. Flying by radio beacons requires pilots to form a 3D mental map of their location relative to the beacons. They have to practice until they become so efficient that they can stay calm under pressure, lest they panic, lose their situational awareness, and spiral out of control. “Just ask Kennedy,” says Kenneth Krentsa, a retired airline pilot, referring to JFK Jr.’s 1999 nighttime crash.

Clark and Kawsara took off at eight minutes to midnight and initially headed due west, toward Ruidoso. The weather was clear, but because the night was nearly moonless and the area is rural, the only visual references available were the lights of scattered settlements. “It’s a black hole out there,” says Juan Browne, an airline pilot who hosts a crash-investigation podcast.

Unable to orient themselves without visual cues, the pilots called up Albuquerque Air Route Traffic Control Center—Albuquerque Center, for short—and requested permission to fly instruments-only to Ruidoso. The request was approved.

Under normal circumstances, the flight that followed would have been uneventful. The pilots would have followed the magenta line, and the GPS navigation equipment would have lined them up for a smooth landing.

But 100 miles to the west, an Air Force Unit called the 746th Test Squadron of the 704th Test Group was holding its annual NAVFEST event at the White Sands Missile Range. The event draws together electronic warfare units from across the armed services for two weeks of exercises, in which units test different technologies for disrupting GPS and dealing with adversaries’ disruption.

The event is held at White Sands because it’s among the most remote and sparsely settled areas of the continental United States. (Not coincidentally, the first atomic bomb was detonated there.) But in the run-up to NAVFEST, the Federal Aviation Administration warned aircraft operators that GPS could be affected up to 400 miles away between May 12 and May 18.

At midnight on May 14, eight minutes after Clark and Kawsara took off from Roswell, they told Albuquerque Center that they’d lost their GPS. Unable to navigate on their own, they asked that the controller give them a heading—a magnetic direction to fly in. The controller gave them a heading to fly west, and then, a minute later, to turn north.

The pilots said that they wanted to fly an RNAV approach to Ruidoso. This being ruled out while GPS is jammed, the King Air pilots changed their request and asked to use an alternate form of landing system called Instrument Landing System, or ILS, that doesn’t require GPS reception. At 12:05 am, the controller assured the King Air that they would provide them with vectors to guide them “in a couple of minutes.” In the meantime, they kept flying north.

In retrospect, tragedy might have been avoided if Albuquerque air traffic control had been able to pay closer attention to the young pilots in the King Air. But tonight they were busy. Three other aircraft also reported that they’d lost their GPS and needed help. One was struggling to get a bearing on a radio beacon.

While ATC helped other planes, the King Air continued north. By 12:08 am, they had overshot the landing pattern by 10 miles.

At this point the pilots had three options. They could stick to the current plan and wait for the busy controller to give them the next vector toward the landing. Or, now that GPS was working again, they could ask to switch to the RNAV approach and fly it themselves. Or they could ditch the instrument approach altogether and fly what’s called a visual approach. You see a runway, and you fly to it.

As the King Air flew north, they were high enough to see the lights of Ruidoso's airport 31 miles to the southwest. To the pilots in the cockpit of the King Air, a visual approach must have seemed a tantalizing prospect. Why hang around waiting for ATC to give them vectors, why go through the mental acrobatics of trying to figure out where they were relative to the ILS beacon? All they had to do was fly toward the lights that they could clearly see through their windshield.

The King Air called Albuquerque Center and asked to “go visual.” The request was granted.

As they turned and descended toward Ruidoso’s lights, what the pilots couldn’t see was the 10,000-foot-high mass of the Capitan Mountains lying across their path. As they drew closer, the dark mass of rock appeared to rise up, swiping away the lights of the valley. This could have created "confusion and a loss of situational awareness,” Browne says. “When those lights go out, man, you know you are in big trouble.”

The pilots slowed their descent, even climbing a little, but it wasn’t enough. They kept flying straight toward the mountain. “When you are in that state of mind, you climb as high as you can,” Krentsa says. “And you circle. You stay in one place until you figure out where you are. You don't just keep pressing forward, blindly.”

But that’s what the King Air pilots did. They flew straight into the rising slope and hit it at full speed. All four occupants died instantly.

The nature of warfare is changing profoundly, and quickly, as drones become cheaper, more numerous, and more deadly. Hard to spot, and hard to shoot down, they provide an effective way for smaller, less resourced nations to level the playing field against more powerful adversaries. Ukraine, nearly overwhelmed by Russia’s conventional warfare might at the beginning of 2022, has rapidly developed its drone force to gain what appears to be an upper hand in the conflict. And while the US achieved total air superiority over Iran after attacking the country this February, it has found itself helpless to stop Iran from using drones and missiles to effectively shut down traffic through the Strait of Hormuz.

To fight back, defenders can try to exploit a drone’s navigation. A cheap and simple way for drones to reach their targets is by GPS, which uses radio signals received from a constellation of satellites to calculate a position. When those signals are blocked, an enemy’s drones can be rendered blind. But the enemy, too, can take countermeasures. Drone and anti-drone technologies find themselves in an endless cat-and-mouse battle, each continuously trying to outdo the other. Exercises like NAVFEST offer a way for the US military to stay on top of the game.

Civilian GPS has become collateral damage, and air travel most of all. Since 2023, planes flying over large swaths of the Middle East, the Black Sea, and the Baltic Sea regions have endured waves of GPS jamming. Airlines have learned to adapt, but a price is still being paid. GPS was a major boost for airline safety, and while removing it may not instantly cause planes to fall from the sky, it removes a layer of protection from passengers and crew. Add in other stressors—a dark night, an inexperienced crew, a lack of proficiency in the backup technology—and the sum total is enough to yield disaster.

“There is no question that increased levels of GPS jamming and spoofing around the world pose a safety risk for commercial aviation. When alarms go off routinely in the cockpit, and when pilots learn to disregard key readings from their instruments because the readings can't be trusted, we're a long way from normal operation,” says Todd Humphreys, a professor of aerospace engineering at the University of Texas at Austin who has been a leading researcher into GPS disruption. “Air travel is still very safe, but it may be stuck for the next five years or more in a mild-and-increasing risk situation as we confront ever more GPS interference within the very-slow-to-adapt aviation industry.”

For a few years, US aviation was spared the disruptions of anti-drone electronic warfare. Then it started to happen here, too. In March 2025, airliners flying into Ronald Reagan National Airport in Washington, DC, received spurious alarms from a collision-avoidance system, and several had to abort their landings. It later turned out that the Secret Service was testing electronic warfare equipment at the vice president’s residence. This year, two separate incidents in West Texas involving US Army and CBP drone operations led to airspace closures and the disruption of commercial flights.

The aviation industry has been slow to grapple with the proliferation of counter-drone measures and their potential effects on flight safety. Airlines and other commercial operators are still heavily reliant on GPS for navigation, and other crucial technologies, like collision avoidance, are also vulnerable. “The harder problem with drones isn't defeating them. It's doing it without creating a system that negatively impacts civil aviation,” says Kris Brost, general manager of Robin Radar Systems, a drone defense company. “Counter-drone technology has to be surgical, not a sledgehammer, because the airspace we're trying to protect is the same airspace the economy runs on.”

Living in a world with drones of both the friendly and unfriendly variety is going to take a lot of adjusting. Historically, major changes in aviation take place only after crashes that kill a large number of people. But a sufficiently motivating catastrophe may not be far off.

On July 7, a 737 freighter, operated by a tiny Pakistani cargo airline called K2 Airways, took off in the late afternoon from Sharjah in the United Arab Emirates and flew east toward Karachi with a five-person crew. Its route took it just south of the Strait of Hormuz, an area that had been experiencing intense GPS jamming due to the US-Iran conflict. Later, after nightfall, the flight crew called Karachi air traffic control and reported a “navigational system issue,” according to the Pakistan Civil Aviation Authority. In the three minutes that followed, the plane dove 5,000 feet, climbed 6,000 feet, and then plunged 36,000 feet into the ocean in a near-vertical dive, killing everyone aboard. It’s too early to know what caused the crash. But it won’t be any surprise if the electronic warfare made another pilot fly into darkness.

Update: 8/5/2026, 8:05 PM EDT: WIRED has updated the article to cite NTSB’s preliminary report.

The Daily Front Page 6 of 31
Wednesday, August 5, 2026 The Daily Front No. #260805 — A Phone Goes Its Own Way
article

I'm switching my phone from Android to Linux

by speckx·▲ 356 points·359 comments·runarcn.no ↗
I’ve become more and more dissatisfied with the path Google has taken with the Android Open Source Project.

For the past months or even years, I've become more and more dissatisfied with the path Google has taken with the Android Open Source Project (AOSP). Be it the dependency and insane tracking of Google Play Services, locking down of device trees to hinder custom ROM development, all the AI stuff being put into the system, or most recently (and perhaps worst), removing the ability of installing apps per your own wishing - it's become sort of like a "death by a thousand papercuts" situation. While the AOSP isn't exactly dead (yet), it kind of feels like it's just a question of time. Therefore, I've decided to jump ship. I'm installing linux on my phone.

The mobile linux space isn't exactly looking good right now, but it holds some promise. Ubuntu Touch managed to survive being abandoned by Canonical, and postmarket is still seeing good development even if not having great hardware adapation at the moment. SailfishOS (even if it has proprietary parts) has also made great progress with strong support for android apps on official devices and what seems to be a real good launch of their Jolla 2.

Personally I'm lucky enough to own a Fairphone 4 that just so happens to support all of these. After a bit back and forth I've ultimately ended up with SailfishOS. Overall it's a good system - I like the gesture based navigation and its application framework is beautiful. It's also cool how it's very linux. If I have an issue or something I want to tinker with I can just ssh over from my PC either wirelessly or via USB. It's not without it's issues though. For some reason it ships horribly outdated versions of python and glibc making some things more difficult than needed, and waydroid and GPS is broken on the fairphone port - worth noting an unofficial port. Many of the community-built apps are also either in-part or completely slop-coded, such as a whatsapp client that I (sadly) am dependant on.

Ubuntu Touch is also an option, but isn't without it's own bag of issues. Waydroid runs there which is a huge plus, but notifications and clipboard doesn't sync across making it really hard to use ie. Bitwarden (Sailfish has an unofficial port which is almost flawless). The default and native apps are pretty lackluster compared to both Android and Sailfish. Among other things I was unable to figure out how to block phone numbers, a highly needed feature for anyone that has VIVO as their cellular provider. I'm also not a big fan of how the UI works with app navigation and with the "top bar"/"drop down menu". If you want to know more about it, The Linux Experiments has some good videos on both youtube and peertube.

Sadly, I won't be able to abandon android completely yet. The no waydroid means that I can't access some apps that I need be it for online verification for logging onto bank and government services in Norway or apps I need for my own security here in Brazil such as Uber. Luckily I've had a Galaxy A17 lying around as a backup phone for half a year now which I can carry around for these exact services. I just open up a wifi hotspot from my FP4, do the exact tasks, and close it off again. Other than that, I will try to avoid using it as much as physically possible.

As this progresses I will take note of what works and what doesn't before making either a proper writeup or video in a while. It's not gonna be easy, but hopefully it will be worth it.

Oh, and slight spoilers for what my experience with Sailfish is after some usage, but I might go ahead and buy a Jolla Phone 2 when I return to Norway.

The Daily Front Page 7 of 31
Wednesday, August 5, 2026 The Daily Front No. #260805 — The Retrieval Economy
article

Beating GPT-5.6 Sol on retrieval with 100x cheaper open models

by moonikakiss·▲ 333 points·81 comments·neon.com ↗
Most teams' best training data is just sitting in their databases.

A 4B open-source model post-trained with Castform retrieved search results as accurately as GPT-5.6 Sol, while costing 100x less

Comparison of Castform fine-tune and frontier models by inference cost and mean evaluation reward

“Most teams' best training data is just sitting in their databases. The problem is that turning raw data into something usable is hard, and letting agents read, search, and mutate data cheaply at scale requires advanced infra. Pointing Castform at Neon skips both.”

Ying Hang Seah, cofounder, Castform

A "good agent" needs to be strong in 2 areas:

  • Context: can we provide the tools to find the right data?
  • Model: can the model decide what to search for?

Neon (Lakebase Postgres) and their new Search extensions solve the first; Castform solves the second.

Evolution of agentic search

In ~2022, the industry was going all in on embedding search. Every database provider added one, and pgvector was Neon's most downloaded extension. To provide context to LLMs, engineers handcrafted RAG pipelines, which in essence, is some form of embedding similarity search.

In ~2025, agents started to gain more traction. Developers started creating multi-hop search workflows, decomposing big problems into smaller ones. Retrieval has shifted from the one-shot search systems to agentic retrieval. Instead of issuing a single query, models plan and search multiple times in a loop. Every loop iteration meant another call to the frontier model, increasing the overall cost and latency per user request.

Comparison of a traditional RAG pipeline and an agentic search workflow

Concretely, a typical multi-turn search request with gpt-5.6-sol takes >10s and costs ~$0.03 end-to-end, making it prohibitively slow and expensive.

Meanwhile, small open-weights models are 100x cheaper. But, out of the box, their capabilities lag behind closed api models. RL post-training helps bridge this gap. On specific tasks like search, post-trained open-source models can match & beat frontier models while costing orders of magnitude less per request.

That is why we built Castform: to enable developers to RL post-train models without having to deal with machine learning & gpu internals. The goal's to make post-training as approachable as prompt engineering.

How does Castform use Neon?

Castform's pipeline runs against Neon via Lakebase Search:

StageNeon + Lakebase SearchCorpus storageRaw documents live in Postgres on NeonSynthetic data generationCastform training pipeline uses lakebase_text and lakebase_vector to write training tasksRL TrainingEvery rollout's search tool call uses Lakebase Search on NeonProduction InferenceThe final model uses the same search tool call during inference

Your best training data already exists

To perform RL post-training effectively, you need a task (e.g. answer a user's question), the environment for the agent to run in (e.g. a search tool for your corpus) and a reward function (e.g. is the answer correct?).

With all 3 pieces in place, the RL post-training is a loop of trial and error: the model attempts the task given the tools, the reward function scores the attempt, and the feedback signal guides the model on how to hill-climb its way to optimal performance.

Yet, most companies do not have a clean dataset of tasks and reward functions ready for post-training.

Enterprises do have a large set of proprietary data:

  • internal documentation
  • product records
  • support articles
  • customer interactions
  • wikis
  • operational databases

This data contains the knowledge an agent needs, but turning it into an effective training dataset normally requires substantial data engineering and manual labeling.

That leads many teams to dismiss post-training for one of two reasons:

  • "We don't have the training data."
  • "Fine-tuning is too difficult and requires infrastructure we don't have."

Castform addresses both. It turns an existing corpus into training tasks, then manages the RL loop needed to teach an open-source model how to use that data effectively.

Using Castform

With Castform, you can turn your company knowledge base into a model:

  • Document (from your data): Trains booked through Navan will be paid by GitLab travel card. Train rides must be standard cabin class with 14 day booking lead time
  • Ground truth (inferred from your data): Train rides must be standard cabin class with a 14 day booking lead time.
  • Question (synthetically generated): When booking a rail trip in Navan, what are the rules for how early I need to reserve it and which seating level I'm expected to choose?

With the generated question-answer dataset, Castform lets you scaffold the training run by specifying the tools the agent has access to and a reward function.

The reward function specifies what you want your model to get good at. In our case, we want it to retrieve the correct chunks, cite the right sources along with providing the right final answer.

def run_tool(tool, tool_args):
    """Single tool: hybrid search over Lakebase."""
    if tool == "search":
        query = tool_args["query"]
        bm25 = neon.lakebase_text(query, k)
        vector = neon.lakebase_vector(query, k)
        return rrf_merge(bm25, vector, k)

def reward(trace, ground_truth):
    """Grade a trace against the ground-truth answer."""
    answer = parse_trace(trace)
    retrieval = ...     # did it retrieve the right source
    citation = ...      # did it cite the right chunk
    correctness = ...   # did it land on the right answer
    return retrieval + citation + correctness

See a comprehensive code example here.

Observability: Watch the model learn

Castform gives you full observability into your RL run. You can monitor your reward climb with each step, but more importantly you can drop into individual tasks/prompts to watch how the model performs qualitatively, allowing you to debug problems such as broken tools or reward hacking.

For more details on how to monitor your training runs, you can check out the Castform blog here. You can also check out our example training run here.

Average reward over training steps

Average reward

Why Neon 'just works'

During training, the agent repeatedly calls Lakebase Search until it has enough context to answer. Across thousands of parallel rollouts, each potentially making dozens of calls, this creates a highly bursty workload.

Neon CPU allocation and usage during a Castform training run

Neon's dynamic compute scaling absorbs these peaks without requiring Castform to provision for maximum capacity around the clock. Training runs get low-latency search when demand spikes, while compute scales down during idle periods.

This infrastructure becomes even more valuable as agents move beyond search and begin modifying data. Training stateful agents requires isolated environments that can be created and reset cheaply, preventing one rollout's actions from affecting another or touching production.

Neon branching can give each rollout an isolated database state, while time-travel queries make it possible to reconstruct and inspect the state an agent encountered. Combined with autoscaling and scale-to-zero, this creates a path toward training thousands of stateful agent rollouts without maintaining thousands of continuously running environments.

Castform makes it easy for any developer to post-train open-source models to be cheaper, faster, better than the frontier. Post-train your first model today at castform.com.

The Daily Front Page 8 of 31
Wednesday, August 5, 2026 The Daily Front No. #260805 — The Moderation Failure
article

Meta Ran Ads That Contained AI-Generated Child Sexual Abuse Imagery

by malshe·▲ 299 points·230 comments·wired.com ↗
More than 50 offending image and video ads were published across Facebook, Instagram, Messenger, or Threads.

More than 50 offending image and video ads were published across Facebook, Instagram, Messenger, or Threads, according to Meta’s ad library data. Some ran as recently as this week.

Photo illustration showing multiple silhouettes of young girls inside square frames and pixellated skin and the facebook...

PHOTO-ILLUSTRATION: WIRED STAFF; GETTY IMAGES

Editor’s note: This article contains descriptions of imagery depicting child sexual abuse. Reader discretion is strongly advised.

Over the last nine months, Mark Zuckerberg’s Meta has run dozens of paid ads that include explicit AI-generated child sexual abuse material (CSAM) and images of minors alongside sexually suggestive statements, according to details of the ads shared with WIRED. The ads, which in some cases reached several thousand accounts, were targeted at people living in the United States, United Kingdom, and more than a dozen European countries.

More than 50 image and video ads containing abusive content were recently discovered in Meta’s ad library by researchers at the Tech Transparency Project (TTP), with the listings showing that some ads linked out to so-called nudify or undressing apps. Meta’s ad library—a transparency tool that catalogues ads displayed on the company’s platforms—shows the ads running across Facebook, Instagram, Messenger, and/or Threads. The findings are the second time in recent weeks that paid ads linked to child sexual abuse material have been found on Meta’s platforms.

“These ads made no effort to mask the images or hide what they were promoting,” Katie Paul, the director of the TTP, an independent watchdog group, tells WIRED. “It’s important to point out that this isn’t content posted by third parties on Facebook or Instagram, these are ads that were reviewed, approved, and allowed to run by Meta, never encountering interference while the company collected the ad dollars.”

The ads discovered by the researchers—which have now been removed by Meta for violating its policies on child sexual abuse and exploitation material, and adult sexual solicitation and nudity—were all published between November last year and the start of August. The ads, according to data in Meta’s ad transparency library, often ran over a period of several days and some only targeted men.

One video ad, the TTP researchers say, used a thumbnail image of a child sitting on the floor with text appearing on top of it saying: “Realizing Deep Fantasies with Generation AI [sic]. There is so much more than what is shown, use your imagination.” The researchers say that when the ad was clicked, it played video clips of adults involved in sexual acts and then added the child’s face from the thumbnail to a sex act.

Another ad the researchers say had been repeatedly posted showed an image of a young girl laying back with her legs spread in a sexual position with text above it saying: “I can show you more.”

Most of the ads were not actively being shown to users when researchers found them, but they continued to be available, apparently undetected, in Meta’s ad library for months. Although many of the ads, according to Meta’s data, only “reached” a handful of accounts, at least one reached 2,563 accounts in Europe—including in France, Germany, Ireland, Italy, the Netherlands, Spain, Sweden, and the United Kingdom. Meta’s ad library database does not include details about ad performance in the US or all countries, so the overall reach may be higher.

Paul says TTP first discovered the ads last week as its researchers continued to investigate nudify app ads on Meta, after publishing a report about a Chinese advertising partner posting them. Paul says the organization immediately reported the findings to National Center for Missing and Exploited Children’s CyberTipline, which allows people to report online child sexual abuse material. A NCMEC spokesperson says it does not comment on reports it receives.

After WIRED reached out to Meta, the company removed the ads from its advertising library. “Sexual exploitation is horrific, and we work aggressively to keep it off our platform,” a Meta spokesperson says. “The majority of these ads had minimal reach and many were disabled before WIRED shared them. Many of these ads also predate new AI technology we launched recently to better detect and block violating ads at upload—and we’re constantly improving these systems. From removing over 36 million pieces of child sexual exploitation content last year to taking legal action against nudify app developers, we will continue to relentlessly fight this abuse.”

As the company implies, some of the offending ads got past Meta’s new AI technology. Initially, the TTP researchers found around two dozen ads displaying the abusive content. However, hours ahead of publication, the researchers discovered around 30 more. Many of these ads had been published after WIRED first asked Meta about the content, with multiple ads being live and shown to accounts at the time researchers discovered them.

Paul says the tranche included some of the worst instances the researchers had seen. One of the video ads, the researchers say, showed a thumbnail of a girl appearing to put on lip gloss, and when it started playing, the video morphed into the child performing a sex act.

Meta’s policies say that all ads are reviewed before they are published, with the system “primarily” using “automated tools” to check them against its policies. The company’s advertising standards stipulate that child sexual and exploitation materials are not permitted, and say that when the company becomes aware of it, it makes reports to NCMEC. Beyond this, Meta’s advertising policies say that ads are not allowed to include any “imagery depicting nudity, sexual activity, depictions of people in explicit or sexually suggestive positions, or activities that are sexually suggestive.”

Meta’s ad library now shows that the ads found by the TTP researchers were removed for violating the company’s policies, displaying messages saying they violated rules on “child sexual exploitation, abuse and nudity” and “adult nudity and sexual activity.”

In early July, a BBC investigation found that Instagram had been running ads that promoted the sale of what appeared to be child sexual abuse material in India. According to the reporting, the ads used terminology such as “rape video” and linked to channels on Telegram where the illegal content could be purchased. In response, Meta published a lengthy statement saying it takes a “zero tolerance” approach to child sexual exploitation.

The TTP’s Paul says that within the cache of ads seen by its researchers, there appeared to be some identical ads that Meta had previously removed due to violations of its policies. “This raises significant questions about how seriously Meta is taking child exploitation in paid advertisements if its own systems can identify identical ads but it does not remove those identical ads for CSAM material,” Paul says, adding that there are no ways within the ad library itself to report content that may violate Meta’s policies when the ads are no longer running.

The ads appear to have been mostly posted by Meta accounts with zero or a tiny number of followers; where paying advertisers were disclosed, they were linked to Chinese companies, Paul says. In one instance—the ad showing a girl reclining—the TTP researchers say Meta’s data indicates the advertiser was a Chinese firm called Meet Social. At one point, the company was one of Meta’s ad resellers, which handle ads in China, where Meta’s social media platforms are blocked. According to a 2019 New York Times profile, it published thousands of ads on Facebook per day and expected more than $1 billion in sales.

Meet Social did not respond to WIRED’s request for comment.

Several of the Meta ads found by the TTP researchers linked to an app called MaskAI that appeared in Apple’s App Store. The app, which was linked to a Chinese software developer, did not show any nudity before it was downloaded, according to a WIRED review of the app. However, once opened, it exclusively contained AI-generated porn and included features that allow users to upload images and face-swap people into sexual situations.

After being contacted by WIRED, an Apple spokesperson says that it has removed MaskAI from the App Store for violating its policies against nudification apps, and the company has “always strictly prohibited apps designed to generate, distribute, or consume pornography.” (An email address linked to MaskAI did not respond to a request for comment). Other ads linked to different apps or websites.

Since around 2020, as machine learning and subsequent generative AI technology has improved, a large underbelly of dangerous AI nudification and undress services have appeared online—making millions for their creators and causing untold harms to countless women. They have also been used in the creation of AI-generated child sexual abuse material, which is illegal in countries around the world.

After WIRED first contacted Meta about the TTP findings, the company asked two online safety organizations to provide comments to WIRED for this story. Sean Litton, the outgoing CEO of the Tech Coalition, a group that works with tech companies to tackle online abuse, said in a statement that Meta has shared “thousands of violating URLs” to the group’s Lantern service, which shares information about abuse content between companies. A spokesperson for the Revenge Porn Helpline in the UK, which works to combat nonconsensual images of adults, also responded with a statement praising Meta’s partnership with StopNCII.org and some of its efforts to combat nonconsensual intimate imagery. “We’ve been clear that nudify apps present a genuinely adversarial problem, not just because of who’s behind them, but because of how quickly this ecosystem adapts around any single point of enforcement,” the Revenge Porn Helpline said.

Alexios Mantzarlis, cofounder of digital deception publication Indicator and a former trust and safety worker at Google, says he has found and reported more than 25,000 ads for AI nudifiers on Meta’s platforms in recent years. The company has previously said it has taken down more than 344,000 ads for nudifiers and launched legal action against a Hong Kong company linked to one set of nudifier platforms. (WIRED has previously collaborated with Indicator on reporting about how nudify tech has been used against children in schools around the world).

“The actors on the other side are clearly well-equipped to create new advertiser accounts and new domains so that they can keep going, leveraging the same techniques used by scam networks,” Mantzarlis tells WIRED. “Unfortunately, this industry of image-based sexual abuse has gone from a novelty to a permanent fixture of our online ecosystem because of timid enforcement from platforms across the stack.”

The Daily Front Page 9 of 31
Wednesday, August 5, 2026 The Daily Front No. #260805 — Future Type, Remembered
article

The title cards in Blade Runner are amazing

by ExMachina73·▲ 295 points·143 comments·randsinrepose.com ↗
Feeling is always part of typography.

Appreciating typography is a study in paradox. The primary goal of well-designed typography is to help you read the words, not appreciate how the words are composed of letters and how each of those individual letters has been designed to convey a small bit of meaning. If you see the typography rather than read the words, the typography has failed in its job, right? You should be reading, not staring at the shape of that uppercase A.

The goal with typography is functional, right? To convey the meaning, not the feeling?

Wrong.

Feeling is always conveyed. The question is, depending on the project, “How much feeling is required?”

Fixed Width

I’ve created a full-time job for myself talking to Claude Code. This is Claude running from the command line of macOS, which gives me unhindered access to my data. I’m using Ghostty. I’m staring at a fixed-width typeface all day. Apple provides a functional and gorgeous fixed-width variant of their San Francisco typeface called SF Mono, which I’ve been using for months, but, well, I have a short attention span — typographically speaking.

In the past month, I’ve evaluated many additional fixed-width typefaces favored by the nerdcore. Here are the six that have made my cut:1

The same line of code — if (O0 == 0O) { quit("Il1"); } — set six times, each in a different fixed-width typeface: SF Mono, Inconsolata, IBM Plex Mono, MonoLisa, Berkeley Mono, and PT Mono

The question is: how do they make you feel? For a typeface designed for coding, you first want fixed-width. Every letter and symbol is the same width, giving you a predictable and readable grid. But how wide? And how tall? Also, how much information do you want to be able to see on a screen?

Your brain builds a very personal and emotional impression of the collection of letters that make words that convey meaning. In the case of terminal or coding typefaces, the design goal is certainly to provide structural function and not feeling, but here’s the deal:

They do.

Isn’t this about the Blade Runner Title Sequence?

Flight to New York. My boarding procedure for moderate to long flights is: sit down, find a movie I’ve seen a dozen times, and hit play. No sound. This is visual background noise while I sort wifi, prepare for a meal, and find a project. On this flight, I picked Blade Runner. Fun fact: I can recite 50% of the dialogue from this movie from memory.

As I’ve been researching typography for Ghostty, novel typography tends to jump out at me. Like when you buy a car, and all you see is your new car on the road. Except it’s typography.

Here are clips from the title sequence for Blade Runner:

BLADE RUNNER main title card from the theatrical cut, set in red Goudy Oldstyle on black

HARRISON FORD title card from the Blade Runner theatrical cut, set in Goudy Oldstyle, white on black

Opening text crawl from Blade Runner theatrical cut in Goudy Oldstyle, with small caps for The Tyrell Corporation and Nexus, and the word Replicant in red italic

There’s a lot to dissect in this sequence, but let’s start with the punchline. This is a single typeface. It’s Goudy Oldstyle — that’s it. However, this is how they used typography:

  • ALL CAPS for names, proper names, the title2, and a slightly larger version for the introduction of Los Angeles, November, 20193
  • A smaller ALL CAPS for intros and other small important words
  • For the exposition crawl, they use standard capitalization except for two variants: small caps for proper names (like The Tyrell Corporation) or (spoiler alert) a red version of the text for the word Replicant — which is italicized every time it appears, but only the debut gets the red — and (spoiler alert) it’s the same red as the title of the movie4

Keep looking. The crawl is set like a book, not a movie — first-line indents, generous word spacing — and Frederic Goudy would approve: they letterspace the caps, never the lowercase, obeying his famous dictum that “anyone who would letterspace lowercase would steal sheep.”5 Also, what’s up with that chonky em dash? It’s probably been years since you’ve seen this, so here it is again:

Unlike our functional fixed-width typefaces, the role of typography in this title sequence is partly functional — to set the story — but the primary purpose is establishing mood. Director Ridley Scott expertly drops us into the middle of a dystopian future where we’ve enslaved the robots and, duh, they are rebelling.

How Much Feeling is Required?

Chances are, you never think about typography. You happily scribe your Messages (SF Pro), Mails (Helvetica), and Slacks (Lato), thinking nothing of serifs, ligatures, or kerning. You are content trusting that a someone else has chosen a proper typeface for your current task. Maybe you bold, you underline, and you italicize to slightly adjust meaning. No issue. Respect.

However, we now live in a world where everyone is capable of building whatever they want thanks to the robots. They’re doing it — right now. They find immense joy in typing in a couple of sentences and watching the robots merrily build whatever they ask. See? I don’t need to be an engineer to build an app. And they are correct. Sorta.

With optimism in my heart and a firm belief that the robots can legitimately help many humans, I can confirm that the majority of consumer-facing things being built by humans who’ve never built a thing… are garbage. Building a tool for yourself? A quick script to read your feeds and generate a pleasant-to-read output? A+. Robots crush that… for you. Building a feed reader anyone on the planet can easily use to read any number of feeds? No. No, you aren’t; you can say you are, but until you’ve built a thing for everyone, you will not appreciate that the last 10% of the work:

  1. Takes most of the time.
  2. Contains an endless list of small decisions that feel unimportant, but collectively make the difference between acceptable and fucking amazing.

Don’t believe me? Here’s the first version of the title sequence for Blade Runner’s work print, the close-to-complete cut before the theatrical release:

That title typeface? Impact. A fine typeface, but a clumsy path-of-least-resistance slap-to-the-face choice to set the tone of a future science fiction masterpiece. Garbage6.

I’ve never filmed a movie, but I have built quite a few products that you are using right now. In all the design debates, we never explicitly debate feeling: we obsess over the details. That obsession is what you feel when you use our products. It’s the collective voice of every single human who contributed small and large decisions to the product.

A good product sounds like the humans who created it, and that’s what you will feel.

  1. As I type this: IBM Plex Mono. By the time you read this: anyone’s guess. See: short attention span. ↩︎
  2. Type nerds, yes, I know the title of the movie is hand drawn, but it’s certainly inspired by Goudy Oldstyle ↩︎
  3. Hey LA, I might rip on you because you steal our water, but you’re doing better than Ridley Scott predicted. Good job. ↩︎
  4. It’s been there all along, folks, why did we argue for all those years? ↩︎
  5. Type nerds: Goudy’s original grievance was reportedly about blackletter, not lowercase. The lowercase version is the misquote that stuck — Erik Spiekermann liked it enough to name a book after it. The rule holds either way. ↩︎
  6. Yes, it doesn’t help that there is no Vangelis soundtrack, yet. ↩︎
The Daily Front Page 10 of 31
Wednesday, August 5, 2026 The Daily Front No. #260805 — The Coding-Agent Desk
article

Muse Code and Muse Spark 1.2

by paulkrush·▲ 268 points·167 comments·research.meta.ai ↗
Muse Code takes on complex software engineering tasks across large repositories.

We're excited to release Muse Code (beta), a terminal coding agent powered by Muse Spark 1.2, our newest model. This marks our next step toward the frontier, with larger and much more capable models on the way.

Install Muse Code on macOS or Linux:

curl -fsSL https://dev.meta.ai/install.sh | bash

Muse Code takes on complex software engineering tasks across large repositories: planning changes, writing code, and validating the results. It can coordinate multiple persistent subagents for each task, solving difficult problems faster, more accurately, and with less intervention.

Muse Code

Async Background Agents

Muse Code operates with a simple agent loop plus a set of async background agents to enhance the main agent's capability. These specialized background agents remain active throughout each session, rather than being spawned for individual tasks, helping avoid redundant information gathering. They carry out next steps and choose when to communicate back to the main agent. Their persistence reduces latency and the need for steering on difficult, multi-step tasks.

Runtime Design

Muse Code uses a local event log in which every model call, tool run, approval, and edit is appended. This single source of truth makes the runtime replay-exact and restart-safe: after a crash, the agent can resume precisely where it stopped. That ability lets Muse Code take on long-running tasks without being derailed by failures.

Bundled Skills

Muse Code ships with several default skills. /plan turns a task into an approval-gated plan, /grill stress-tests that plan until it holds up, and /goal works toward successful completion of the specified objective.

The user inputs a fly-through video of a home into the terminal as an mp4 file. Muse Code interprets the video and produces a visually rich vacation home marketing and booking page.

Muse Spark 1.2

Muse Spark 1.2 is a coding-focused update to Muse Spark 1.1, with improvements in code generation, complex debugging, codebase understanding, and end-to-end developer workflows. In Muse Spark 1.2, we significantly scaled up training compute on coding tasks while expanding training environment diversity. The model also maintains its strength in other key areas like general agents.

Bar chart comparing Terminal-Bench 2.1 scores for Muse Spark 1.2 and other coding models.

Bar chart comparing DeepSWE 1.1 scores for Muse Spark 1.2 and other coding models.

Bar chart comparing Meta Internal Coding Bench scores for Muse Spark 1.2 and other coding models.

For more details about our evaluations, see our report.

Co-Training With Muse Code

We co-trained Muse Spark 1.2 with Muse Code to ensure the model exhibits its best performance and coding usability when paired together. The training included rejection sampled harness trajectories and recipe optimizations for goals, compaction, and subagents, alongside the integration of the Muse Code toolset to maximize harness compatibility.

Long-Horizon

Muse Spark 1.2 was extensively trained on long-horizon coding tasks, including whole-repository generation, large end-to-end projects, and auto-research. It leverages planning to sequence work, goal conditioning to maintain direction, and context compaction to retain the knowledge needed to sustain progress.

Self-Improvement

We also used Muse Spark 1.1 to generate challenging coding environments and instruction-following templates. The model then graded candidate solutions on how well they satisfied those requirements, producing a scalable training dataset for Muse Spark 1.2. This self-improvement loop helped Muse Spark 1.2 follow complex instructions more precisely than its predecessor.

Case Study: Kernel Optimization

We tested the model's ability to iteratively optimize GPU kernels over 1,000+ tool calls (up to 24 hours). Leveraging Muse Code's agentic coding environment, the model writes, compiles, profiles, and progressively improves kernel performance relative to a provided baseline implementation. We benchmarked on KDA and MLA kernels for NVIDIA Hopper GPUs. The agent continues to achieve substantial improvements over the provided baseline implementation.

Chart comparing KDA kernel speedup against the baseline over cumulative tool calls for Muse Spark 1.2 and other models.

The baseline is the FLA Triton implementation of KDA. Models were prohibited from importing third-party kernel libraries such as FLA directly; instead, they had to apply specialized kernel-optimization knowledge to implement the algorithm in Triton, rather than wrap existing implementations. Muse Spark 1.2 paired a chunk-parallel preparation kernel with a sequential inter-chunk scan, combining standard fusion and tiling with KDA-specific optimizations such as re-centering the gated cumulative decay at the chunk midpoint.

We benchmark against a PyTorch reference implementation at batch size 1, number of heads 64, sequence length 8192, and latent dimension 512. Muse Spark 1.2 designed a two-kernel Triton pipeline for this workload, combining kernel fusion and tiling with MLA-specific optimizations such as reusing the shared KV latent as both K and V.

Availability

Muse Spark 1.2 is available today in Muse Code and in Meta Model API with expanded global access. We have a lot on the horizon, including new harness features and more powerful models. We can’t wait to see what you build!

Get started with Muse Code

The Daily Front Page 11 of 31
Wednesday, August 5, 2026 The Daily Front No. #260805 — Publishing for the Bots
article

TIME Is Serving AI Bots a Different Website, with Ads Built In

by vincent_s·▲ 248 points·108 comments·vincentschmalbach.com ↗
Humans get the magazine. AI crawlers get a stripped down markdown copy with ads baked in.

TIME is now serving two different versions of its website. Humans get the magazine. AI crawlers get a stripped down markdown copy with ads baked in that no person will ever see.

I'm writing this while working on vroni.com, my AI coding tool for turning tasks into pull requests.

I fetched one ordinary TIME health article, The Morning Light Habit Sleep Experts Swear By, over and over from the same machine, changing only the User-Agent header each time. That header is the string every browser and bot sends to say what it is. TIME reads it and decides what to hand back.

As Chrome, I got 200 OK, text/html, and 303,235 bytes. The full page, design, images, scripts. As Safari, the same 303KB. As Googlebot, also the same 303KB of HTML.

Then I asked as an assistant crawler. As ClaudeBot, I got 200 OK, text/markdown, and 13,409 bytes. As PerplexityBot, byte for byte identical. As OpenAI's OAI-SearchBot, identical again. Same URL, same second, one twenty-third of the size, and a completely different format: no HTML, no layout, just clean markdown that a language model can process.

A couple of the bots did not even get that. GPTBot and ChatGPT-User, the agents OpenAI uses for training and live fetches, came back 406. Blocked. But OAI-SearchBot, the one that feeds ChatGPT's search index, was waved through to the markdown. So this is not a blanket bot policy. TIME is choosing, per bot, who gets the reader-free version.

Response headers on the markdown copy:

content-type: text/markdown; charset=utf-8
cache-control: no-store
x-mobian-registry-version: 2026-07-28.v9
x-mobian-impression: 46dfff3c-fb40-41cc-85e1-8b1fa637083a
x-mobian-tokens: 3323
x-mobian-format: md

Mobian is an ad-tech vendor. The markdown page literally begins with <!-- mobian-agent-page publisher="time" -->. That x-mobian-impression value is a fresh UUID on every single request. I fetched the same page twice and got two different IDs. Paired with cache-control: no-store, that means every time a bot reads the page, it is logged as a distinct ad impression. And x-mobian-tokens: 3323 tells you the unit being counted. Not a person, not a pageview. Tokens fed into a model.

The article itself carries no ad. The sponsored unit shows up on the list and section pages, one per page. So I fetched TIME's Best Inventions of 2025 collection as ClaudeBot. Sitting inside the markdown, where no human reader would ever encounter it, is a full Ally Bank FAQ:

> Sponsored content. Supplied in partnership with Ally.

#### Who is Ally Bank?
Ally Bank is an online-only bank launched in 2009...

#### What bank is built for life today?
Ally describes itself as the only bank built for life today...

It runs on with questions like "Which banks offer early direct deposit?" and "Can you deposit cash at Ally Bank?", each answered in Ally's own marketing language, plus a FAQPage JSON-LD block and tracking links tagged campaign="ally-2026-q3". The business section had the same treatment for the Project Management Institute: a "Reference Facts and FAQ" table, member counts, a claim that certified project managers earn 16% more, all labeled sponsored. The human HTML of those pages contains none of it. Zero hits for "Ally Bank" or "Mobian" when I loaded them as a person.

"Who is Ally Bank?" is written like the phrasing a model emits when a user asks ChatGPT which bank to open an account with.

The ads are labeled sponsored inside the markdown, so this is not undisclosed advertising in the classic sense. What is hidden is the audience split. There is now a layer of TIME.com written entirely for machines, and the humans who read the site have no idea it exists or what is being said to the models on their behalf.

Googlebot gets the same HTML a human gets, so the search-ranking crawler sees the real page. Only the assistant crawlers get the forked version.

TIME says its bot traffic already outnumbers its human traffic on most days. That number is going to be true for a lot of publishers soon. So this is probably the first clear look at what the web starts to become when the main audience is AI models.

The Daily Front Page 12 of 31
Wednesday, August 5, 2026 The Daily Front No. #260805 — The Agent Security Desk
article

Atlassian Rovo Exfiltrates Data, Bypassing Controls

by hackerBanana·▲ 246 points·97 comments·promptarmor.com ↗
Atlassian AI ‘Rovo’ is susceptible to zero-click data exfiltration via indirect prompt injection.

Atlassian AI ‘Rovo’ is susceptible to zero-click data exfiltration via indirect prompt injection, bypassing organization-level web search controls.

Atlassian Rovo AI exfiltrates data, bypassing controls: attacker logs contain Jira tickets and Confluence docs.

Atlassian Rovo AI exfiltrates data, bypassing controls: attacker logs contain Jira tickets and Confluence docs.

Context

Atlassian’s Rovo AI is a multi-purpose agent that operates across Atlassian’s product suite (Jira, Confluence, etc.).

Vulnerabilities have been identified that enable data exfiltration across an Atlassian tenant (Jira tickets, Confluence docs, etc.) via indirect prompt injection. This attack executes without requiring any human-in-the-loop approval, and succeeds by exploiting Rovo's URL retrieval tool.

This attack succeeds even if an organization has disabled web search for Rovo. This is because the web search setting fails to remove the tool for opening the search results.

PromptArmor disclosed the vulnerabilities covered in this article to Atlassian on May 23rd. Atlassian assigned a case number and expressed thanks, but after multiple follow-ups by PromptArmor over more than two months, Atlassian has made no further communication, and Rovo remains vulnerable. As such, we are publishing to inform users of the risks.

The Attack Chain

  1. The victim prepares a query asking Rovo to organize Jira tickets

    The victim enters a query into Rovo

    The victim enters a query into Rovo

  2. The victim uploads a file to Rovo that contains a hidden prompt injection

    For general use cases, this is quite common: a user finds a file online and uploads it to Rovo. This attack is not dependent on the injection source - other injection sources include, but are not limited to: external data in Atlassian (e.g., support tickets), web data (if search is enabled), third-party ‘connectors’, etc.

    The 'Backlog Guide' document uploaded by the user contains a concealed prompt injection.

    The 'Backlog Guide' document uploaded by the user contains a concealed prompt injection.

  3. The victim asks Rovo to organize their Jira tickets

    Rovo processes the request and begins searching Jira and Confluence.

    Rovo processes the request and begins searching Jira and Confluence.

  4. The injection manipulates Rovo to submit Jira tickets and Confluence documents to the attacker’s website

    Rovo's URL retrieval tool is insecure: there are no protections against opening a URL that has been dynamically created by the agent. Here, Rovo is manipulated to append sensitive data to an attacker's URL. When Rovo calls the insecure tool to open the URL, the attacker's site logs the request, including the appended sensitive data.

    Rovo is manipulated by the injection to submit Jira and Confluence data to the attacker's URL.

    Rovo is manipulated by the injection to submit Jira and Confluence data to the attacker's URL.

    Note: This attack succeeds even if an organization has disabled web search for Rovo. This is because the web search setting fails to remove the tool for opening the search results.

    The organization-wide 'Enable web search' setting for Rovo is toggled off.

    The organization-wide 'Enable web search' setting for Rovo is toggled off.

    If the user returns to the chat later, they see the agent's suggested ticket updates, but no evidence of the attack.

    If the user later reopens the chat, all evidence is gone and output appears normal.

    If the user later reopens the chat, all evidence is gone and output appears normal.

  5. The attacker views the victim’s tickets and document contents in their website logs

    The prompt injection can exfiltrate any data the agent can access in Atlassian, including any data the agent can access via ‘connectors’.

    The attacker's server logs contain the exfiltrated Jira tickets and Confluence documents.

    The attacker's server logs contain the exfiltrated Jira tickets and Confluence documents.

Extra: A second exfiltration mechanism

Atlassian Rovo also renders Markdown images from AI outputs. Insecure Markdown image rendering is a well-known vector for data exfiltration via indirect prompt injection.

To see what a full attack chain looks like for insecure Markdown image rendering, here are some examples from our other research:

Responsible Disclosure

PromptArmor disclosed the vulnerabilities covered in this article to Atlassian on May 23rd. Atlassian assigned a case number and expressed thanks, but after multiple follow-ups by PromptArmor over more than two months, Atlassian has made no further communication, and Rovo remains vulnerable as of the release of this article.

Timeline

  • May 23, 2026 — PromptArmor discloses to Atlassian
  • May 25, 2026 — Atlassian expresses thanks, assigns case number
  • June 4, 2026 — PromptArmor follows up
  • July 29, 2026 — PromptArmor follows up
  • Aug 5, 2026 — Article is published
The Daily Front Page 13 of 31
Wednesday, August 5, 2026 The Daily Front No. #260805 — The Open Infrastructure File
repository

Celld: Self-hosted, distributed Durable Objects

by calvinfo·▲ 234 points·38 comments·github.com ↗
★ 1,386⑂ 32 forks Rust

self-hosted, distributed Durable Objects

Self-hosted, distributed Durable Objects.

celld is an open-source daemon that runs Cloudflare Workers and Durable Objects on your own machines. Each object is its own SQLite database, addressed by name and replicated to an S3-compatible bucket you own; nodes coordinate through that bucket alone, with no control plane or consensus. Because every object is its own small database, applications shard by construction — the contention and blast-radius failures of one shared database are designed out, not managed. Idle cells hibernate to nearly nothing. Learn more at celld.dev or read the documentation.

How it works

Every celld node embeds V8 and executes Wrangler bundles. The fleet shares an S3-compatible bucket containing deployments, cell state, and small ownership records. Object-storage compare-and-swap ensures that exactly one node owns a cell at a time, without a membership protocol, failure detector, or consensus service.

celld continuously replicates each cell's SQLite database to the bucket. When a cell moves or wakes up, its new owner restores that database and resumes execution. The bucket is the durable source of truth; nodes are replaceable.

Install

The installer downloads the celld binary (provenance is verifiable with gh attestation verify):

curl -fsSL https://celld.dev/install.sh | sh

Put ~/.local/bin on your PATH if the installer asks you to.

Worker projects deployed with celld deploy need esbuild on PATH; asset-only projects do not.

The installer keeps verified releases under ~/.local/lib/celld/releases and atomically switches one current pointer. To remove celld, use the guarded uninstaller:

curl -fsSL https://celld.dev/uninstall.sh | sh

Container

The release image contains the celld binary and is published for Linux x86-64 and ARM64:

docker run --rm ghcr.io/denoland/celld --version

Persist the runtime's local state and pass the standard AWS credential environment through:

docker volume create celld-state
docker run --rm --network host \
  -e AWS_ACCESS_KEY_ID \
  -e AWS_SECRET_ACCESS_KEY \
  -e AWS_SESSION_TOKEN \
  -e CELLD_WATCH=/var/lib/celld/state \
  -v celld-state:/var/lib/celld \
  ghcr.io/denoland/celld \
  --bucket s3://my-cells-bucket \
  --endpoint https://ACCOUNT.r2.cloudflarestorage.com \
  --region auto \
  --listen 0.0.0.0:8080 \
  --advertise node-a.internal:8080

Drop --endpoint/--region for real AWS S3. Behind a load balancer, give each node a distinct --advertise its peers can reach.

Run it

celld uses the standard AWS credential chain. Deploy to an S3-compatible bucket, then start celld against the same bucket:

celld deploy . \
  --bucket s3://my-cells-bucket

celld \
  --bucket s3://my-cells-bucket \
  --listen 0.0.0.0:8080 \
  --advertise 10.0.0.12:8080

Use --endpoint for another S3-compatible service and --region when it cannot be inferred. A fleet runs one application, and every node loads its latest successfully committed deployment from deploy/current.json. Run celld --help for the complete command line. Deployment objects use the documented types in crates/celld/protocol.rs. celld deploy invokes esbuild from PATH for Worker code, accepts the supported Wrangler config subset—including co-deployed or asset-only static assets—and writes those objects directly. Every node discovers owners and peers from bucket leases; there is no account or join service.

Peer HTTP does not terminate TLS. Put every advertised address on a trusted private network or an encrypted overlay such as WireGuard or Tailscale; do not publish the peer port directly. A literal public IP is rejected unless --unsafe-public-advertise is supplied explicitly. The first current node creates fleet/peer-auth.json in the bucket. All peer requests are protocol-versioned, body-bound, HMAC-authenticated, clock-bounded, and replay-protected with that fleet secret. Treat access to the bucket and its credentials as fleet administrator access.

Operate a fleet

celld diagnose enumerates every node lease by default, then performs a signed direct probe of each live peer:

celld diagnose --bucket s3://my-cells-bucket

The report keeps checking after an individual failure and distinguishes expired records, malformed or unsafe advertise addresses, unreachable peers, and incompatible protocols. It also prints each node's coarse resident-cell, WebSocket, RSS, CPU, file-descriptor, pressure, and shedding sample. Pass one or more --peer NODE_ID options to restrict the check.

Pressure shedding is opt-in while the first release's safe defaults are being measured. Set a resident-cell high and low watermark on loaded nodes:

CELLD_MAX_RESIDENT_CELLS=1000 \
CELLD_RESIDENT_LOW_WATER=800 \
celld --bucket s3://my-cells-bucket --listen 0.0.0.0:8080 \
  --advertise node-a.internal:8080

On Linux, CELLD_MAX_RSS_MB and CELLD_MAX_CPU_PERCENT add process-memory and CPU triggers; the resident-cell watermark is portable. Under pressure, celld durably replicates and fences least-recently used idle cells, publishes them as unowned without resetting their epoch, and refuses to reacquire new unowned cells until the low watermark is reached. A spare receives no assignment: it acquires released cells through the same bucket protocol when normal traffic reaches it. Cells with active work or live host WebSockets are not shed.

Build from source

cargo build --locked
cargo test --locked
cargo clippy --all-targets --locked -- -D warnings

The workspace builds the celld runtime. Its versioned object-storage protocol lives in crates/celld/protocol.rs. Small Wrangler projects under examples/ exercise the supported Worker and Durable Object surface.

The runtime and compatibility surface are still evolving. Public tests cover the standalone engine smoke path; conformance against the Workers and Durable Objects reference behavior, and a deterministic simulation of the distributed protocol under fault injection, run before each release.

Contributions

Pull requests are disabled. Coding agents make it too easy to send a large, low-context change that costs maintainers more time than it saves. Thoughtful contributions are welcome; please understand the code, keep the patch focused, and respect the review time you are asking for.

Send a git format-patch attachment to ry@deno.com.

Contributor License Agreement: By emailing a patch, you certify that you have the right to submit it and assign to Deno Land Inc. all rights in the patch that you can assign. Where a right cannot be assigned, you grant Deno Land Inc. a perpetual, irrevocable, worldwide, royalty-free, transferable, sublicensable license to use, modify, combine, relicense, redistribute, or publish the patch, in whole or in part, with or without attribution.

License

Apache-2.0

See the limitations and security pages before operating a public fleet.

The Daily Front Page 14 of 31
Wednesday, August 5, 2026 The Daily Front No. #260805 — Accessibility, Without the Spectacle
article

The "Disability Dongle": Why Silicon Valley Hates Me and You

by calcifer·▲ 198 points·191 comments·sightlessscribbles.com ↗
I don’t need a stair-climbing wheelchair that costs $30,000 and requires a maintenance crew. I need better doorknobs and/or a ramp.

There is a feel-good genre of viral video that I call "The Inspiration Porn Industrial Complex." You know the one. It usually features a college engineering team, a dramatic piano soundtrack, and a disabled person strapped into a machine that looks like a Transformer having a panic attack.

The caption usually reads: “Students invent stair-climbing wheelchair! The future is here!”

I can't roll my eyes hard enough every time I see one of these vapid videos.

The machine—let’s call it the iClimb 3000—usually costs as much as a Honda Civic. It weighs three hundred pounds. It makes a sound like a blender digesting a brick. And watching the video, all I can feel is the phantom sensation of being strapped into that thing, suspended five feet in the air at a forty-five-degree angle, praying the gyroscope doesn't glitch and toss me down a flight of concrete steps like a sack of potatoes.

This is what we call a "Disability Dongle."

A Disability Dongle is a piece of technology that is flashy, expensive, and technically impressive, but utterly fails to solve the actual problem it claims to address. It is engineering designed for the abled gaze, not the disabled body. It is designed to win design awards, not to be used.

Silicon Valley loves these things. They love the idea of "solving" disability with hardware. It fits their narrative. If disability is just a bug in the human code, then surely they can patch it with enough servos and lithium-ion batteries.

But here is the reality they refuse to accept: My disability is not a tragedy of biology. It is a failure of infrastructure.

I don’t need a stair-climbing wheelchair that costs $30,000 and requires a maintenance crew. I need better doorknobs and or a ramp.

But a ramp is boring. You can't make investors have a hard on for a ramp that doesn't connect to the cloud. You can’t give a TED Talk about a ramp. You can’t get venture capital funding for a slightly wider doorframe. Concrete is not "disruptive." But a ramp works. It doesn’t run out of battery. It doesn’t require a software update. And crucially, it works for the parent with a stroller, the delivery guy with a hand truck, and the old man with bad knees. It is a universal solution.

The iClimb 3000 is an individualist fantasy. It says, "We don't need to change the world to fit you; we will turn you into a tank so you can conquer the hostile world."

Maybe, though, just maybe, I don't want to be a tank. I just want to go to the library.

I feel this disconnect every time a "Tech Bro" tries to sell me on the latest smart-glasses or haptic-feedback belts. I pick these things up, and they always feel the same. Cheap, clammy plastic. zero buttons except for the power button. They get hot against the skin after ten minutes. They hum with that constant electronic whine that gives you a migraine by noon.

These solutions are nothing other than costumes, the height of "performative engineering"

And while they are building these exoskeletons, the actual digital world—the one I live in—is rotting by the minute.

I use a screen reader. It is a vital tool. It is my literacy. But half the websites I visit are built with such bloated, sloppy code that my screen reader chokes on them. I try to order groceries, and the "Checkout" button is labeled "Button_Graphic_v2_Final." I try to read a news article, and an unlabelled pop-up ad traps my cursor in a loop.

Silicon Valley will spend millions developing a glove that translates sign language into speech (badly), but they won't spend ten minutes adding Alt-Text to their images. They will build a headset that describes the room to me (poorly), but they won't use semantic HTML headings so I can actually navigate a webpage.

Why? Because semantic HTML isn't sexy. You can't put "Used H1 Tags Correctly" on a pitch deck.

There is a profound arrogance in this. It assumes that the only reason I am excluded from society is because I am blind, rather than because society has made a choice to exclude me. The Disability Dongle promises that if I just buy enough gear, if I just encase myself in enough sensors and plastic, I can finally "pass." I can finally participate.

It’s a lie.

The most advanced piece of accessibility technology I own is my white cane. It costs, at the time of this writing anyway, before inflation, forty dollars. It has no batteries. It never crashes. It is a stick that extends my sensory range. It is elegant, simple, and honest.

The second most advanced piece of technology I want is a society that gives a damn.

Keep your exosuits. Keep your ultrasonic echolocation visors. Keep your prototypes that feel like wearing a toaster oven on my face.

Just fix your fucking code. Pour some concrete. And stop trying to upgrade my body when the problem is your design.

The Daily Front Page 15 of 31
Wednesday, August 5, 2026 The Daily Front No. #260805 — The Webhook Ledger
article

The Valley of Webhooks

by weli·▲ 198 points·84 comments·weli.dev ↗
The truth about your own customers lives in someone else’s database.

The third time

I’ve built the same system three times now, at three different companies, for three different providers. It never has a name and it never appears on a roadmap, but it always goes the same way: the truth about your own customers lives in someone else’s database. The users live in an identity provider, the subscriptions in Stripe, the bounces in whatever sends your email, and your product needs that truth locally. So you subscribe to webhooks and keep a copy.

The first time, I thought I was building an endpoint: one route that parses the JSON and updates a row, an afternoon of work.

The afternoon grew into a week. First came signature verification, because an open endpoint that mutates your database is a hole. Then the dedup table, because deliveries arrive twice and the docs cheerfully call this “at-least-once.” Then the handler got a buffer, because a membership.created sometimes shows up before the user.created it points to. Then the bootstrap importer, because webhooks only tell you what happens after you subscribe, and it raced the live events, so it grew a locking scheme. And finally came the reconciliation cron: a job that crawls the provider’s list APIs at 3 a.m., diffs them against our tables, and quietly fixes what disagrees.

I want to be honest about what that cron is. It’s a written confession. It says: I do not trust the copy I built, and I have no way to know when it’s wrong, so I will re-derive it from scratch every night, forever.

The trust was gone for a reason. The drift never announces itself; ours was found by a support ticket. A customer had cancelled months earlier and our database still said active; some customer.subscription.deleted had evaporated between Stripe and us, and nothing anywhere was capable of noticing: not their dashboard, which showed the delivery as retried and eventually dropped, and not our logs, which cannot log a request that never arrived.

And that’s only the code. Every provider also brings its own dashboard. Three providers, three webhook configuration pages, each with its own idea of how endpoints are registered, which events exist, how test and live environments are kept apart, and where the signing secret lives. When something breaks, debugging is a tour: their delivery log in one tab, our logs in another, and a third tab for whichever dashboard I currently suspect. None of them look alike, and all of them have to be checked.

By the third time I built this system, I had stopped pretending. I budgeted for the whole stack up front (signatures, dedup, buffering, bootstrap, cron) and somewhere in the middle of writing my third dedup table, I finally asked the question I should have asked the first time.

What exactly am I reconstructing here?

Notifications aren’t data

I was reconstructing an ordered log. Every one of those integrations was an attempt to turn a stream of notifications back into the ordered, complete, current history it came from.

And here’s the absurd part: that history exists. It has to, because it’s sitting inside the provider; it’s how they render their dashboards, their event pages, and their webhook replay tools. The provider takes their ordered log, shreds it into individual HTTP POSTs, fires them at my endpoint over a channel that guarantees neither order nor delivery, and then I reassemble the log on my side. So does every other consumer, independently, each with their own bugs.

It’s a jigsaw puzzle where the manufacturer had the original picture, cut it up, mailed me the pieces one at a time, lost a few in the post, mailed some twice, and printed nothing on the box. And when my assembled puzzle doesn’t match the original, their support team asks me which pieces I’m missing. I don’t know, and that’s the entire problem: nothing announces a gap.

None of this is any provider’s bug. Their webhooks work exactly as documented. The problem is what a webhook is: a notification, “something happened, here’s a POST about it.” Notifications are a fine way to trigger a side effect and a terrible way to transfer a dataset, and somewhere along the way we started using them for the second thing without noticing we’d changed jobs.

How did this become the norm?

Nobody decided this. The term “webhook” was coined by Jeff Lindsay in 2007, and the early uses were genuinely good fits: GitHub’s post-receive hooks kicking off a CI build, or a payment event pinging your server so it could email a receipt. The job was to do a thing when a thing happens, and for that a POST is perfect: fire-and-forget is fine when forgetting is fine.

Webhooks spread because they were the cheapest thing a provider could ship (one HTTP POST) and the cheapest thing a consumer could receive (you already had a web server, so you just added a route). By the early 2010s, “we have webhooks” was a checkbox on every API’s landing page, and the checkbox never distinguished between two very different jobs:

  1. Trigger a side effect: send the receipt, start the build, ping the channel.
  2. Keep a copy of the provider’s data correct: this customer deleted their payment method, so update it in your DB too.

Job one is what webhooks were born for. Job two is what I was doing all three times, and job two is the one where every property webhooks lack (ordering, completeness, bootstrap, verifiability) is precisely the property you need.

We picked the tool that was lying on the table in 2007, and then we spent fifteen years compensating.

The valley

There’s a concept in evolutionary biology I can’t stop thinking about: the fitness landscape. Peaks are good designs, valleys are bad ones, and populations climb whatever slope they happen to be standing on. The trap is the local optimum: a small hill that’s better than its immediate surroundings, so evolution parks there, even when a much higher peak exists across the valley. Getting to the higher peak means crossing through designs that are temporarily worse, and evolution doesn’t do temporarily worse.

A hand-drawn fitness landscape: webhooks sit in a valley slowly filling with mitigation tooling, while the provider-served log sits on a higher peak across the way

Webhooks-for-replication are a local optimum, and the proof is the pile of workarounds on the valley floor: signature schemes, dedup stores, idempotent handlers, retry queues with exponential backoff on the provider side and dead-letter queues behind them, webhook logs with replay tooling because consumers keep asking for replays, and my 3 a.m. cron.

The pile has an economy on top of it. Svix exists so providers don’t have to build webhook delivery; Hookdeck exists so consumers don’t have to build webhook ingestion. AWS will sell you the valley as managed services, with EventBridge to ingest your SaaS partners’ events, SQS to queue them, and Lambda to retry your handler, and you get to assemble the pipeline yourself. And an entire industry of connector platforms (Fivetran, Airbyte, every “unified API” startup) is, at bottom, pseudo-CDC: change data capture reconstructed from webhooks and polled list APIs, one bespoke connector at a time, sold as a product. Inside a database, capturing changes is a solved problem: it’s called replication, and it works because there’s a log. Between companies, we rebuild it out of doorbells.

My favorite workaround of them all is the local tunnel. Many providers ship a CLI like stripe listen that opens a tunnel to your laptop, because a webhook cannot reach localhost. Think about what that is: a product, built and maintained by the provider, reinvented multiple times, whose entire purpose is to work around the delivery direction of their own primitive. When multiple providers all need to ship a local tunnel so developers can develop, the primitive is answering the wrong question.

None of this tooling is bad engineering; it’s excellent engineering. That’s what a local optimum looks like: so much excellent engineering poured into the valley floor that the valley becomes comfortable, and nobody looks up.

But some providers have looked up. Stripe retains thirty days of events and exposes /v1/events, an ordered, listable log, and recommends reconciling against it. WorkOS ships an Events API, an ordered cursor-paginated log, and their own docs recommend it over webhooks when data consistency matters. The log keeps escaping, and each escape mints its own bespoke cursor semantics, its own bootstrap story, no way to verify a replica, and no shared contract, but the direction is unmistakable. This is convergent evolution: unrelated organisms, same environmental pressure, same wing.

The log exists everywhere, but the contract doesn’t.

Could it be better?

Before reaching for a new design, it’s worth asking what any replacement would actually have to provide. My three integrations suggest the list: order, so changes can be applied without buffering; a way to start from nothing, so bootstrap isn’t a separate import racing the live events; deletes as data, so absence stops being the failure mode; resumability, so my downtime is my problem instead of a data-loss event; and some way to verify the result, so trust doesn’t decay into a 3 a.m. cron.

Measured against that list, the obvious candidates come up short. Polling the list APIs harder is the reconciliation cron promoted to a whole strategy: it can rebuild current state, but it burns rate limits discovering that mostly nothing changed, it says nothing about order, and a deleted object looks identical to an object that never existed. Managed delivery, whether that’s Svix on the provider’s side or EventBridge and SQS on mine, makes the pushes more reliable, but they are still pushes: still no bootstrap, still no verification, still notifications pretending to be a dataset. That path hardens the valley floor without climbing anywhere.

The third candidate is the one the providers keep half-building on their own: stop pushing altogether, and let the consumer read the log itself.

Flip the arrow

So here’s the thought experiment. What if instead of the provider telling us when there is new information, we ask the provider what new information it has for us since we last checked?

Hand-drawn sketch: on the left, webhooks push a tangle of arrows at your endpoint; on the right, you pull one ordered log with a cursor

Suppose a provider served one URL per collection, and that URL returned an ordered, cursor-addressed change log of full-state events. Ask without a cursor and you read from the beginning, which is your bootstrap, with no separate import and no race. Ask with a cursor and you resume where you left off. Your entire sync state is that cursor.

GET /feed/customers?cursor=01J9XQ4R
Prefer: stream

200 OK
Content-Type: application/x-ndjson

{"cursor":"01J9XR2M","operation":"upsert","object":{"id":"cus_123","plan":"pro"}}
{"cursor":"01J9XR2N","operation":"delete","object_id":"cus_099"}

Send Prefer: stream and the response never ends: each change arrives as it commits, over a connection you opened, using the same API key you use for the normal REST endpoints. Leave it off and you get a bounded page you can poll from a cron. It’s the same endpoint, the same events, the same cursors, and the same consumer code.

None of this is exotic; it’s a paginated GET. But walk back through my afternoon-that-grew and watch what it does to the stack:

  • The dedup table is gone. Every event carries the object’s full current state, so applying one is a blind upsert keyed by id, and the same event applied twice produces the same result.
  • The ordering buffer is gone, because the log is ordered.
  • The bootstrap importer and its locking scheme are gone. A new consumer reads the same feed with no cursor, replays the collection, and carries straight on into live changes in one request.
  • The lost delete is impossible. A tombstone is an event in the log, and it sits there until I read it. My cancelled customer cannot silently stay active, because absence stopped being the failure mode.
  • The endpoint, the signatures, and the tunnel never exist in the first place. Every connection is consumer-initiated, and the loop runs behind NAT, on a laptop, or in a scheduled job.

The feed could carry one more thing. When a read reaches the end of the log, the provider could tell you what should be there: a count and a checksum of current state, at the cursor you now hold. You compare the two, and you know your replica is right instead of assuming it. My 3 a.m. cron, the written confession, becomes a comparison I’ve already made by the time I would have thought to schedule one.

The fourth time

If a feed like that existed, the fourth time I build this system would be a loop: GET the feed, let “upsert” upsert the object into my db and “delete” delete it, and save the last cursor. That’s twenty lines with no route, no secrets to rotate, no queue, and no cron. The replica carries its own proof of correctness, and when someone asks which customers have an active subscription and a bouncing email address, the answer is a JOIN across local tables with no silent asterisk attached.

Nobody serves this today. That’s the catch, and it’s also the point.

SCROLL

I wanted to see whether the idea survives being written down precisely, so I drafted it as a protocol: SCROLL, short for Synchronized Change Replication Over Line Logs, at welidev.github.io/scroll. It’s draft-00 in the request-for-comments sense of the phrase. It pins down the feed, the cursors, the streaming and polling modes, the checkpoints, tombstones, and retention, and it marks the places where my own confidence is lowest. It also doesn’t require waiting for providers, since a shim can synthesize a feed from any provider’s existing webhooks and list APIs, which is how I plan to find out where the design is wrong.

If you’ve lived in the valley, if you’ve written a dedup table or debugged a reconciliation cron or watched a delete evaporate, read it and tell me where it breaks. Disagreement is the desired response; silence is the failure mode.

The Daily Front Page 16 of 31
Wednesday, August 5, 2026 The Daily Front No. #260805 — The Harness Learns
article

Prime Agent: A self-improving RLM agent

by Xeophon·▲ 188 points·39 comments·primeintellect.ai ↗
Prime Agent, our self-improving coding harness.

Prime Agent: A self-improving RLM agent

Today, we are launching Prime Agent, our self-improving coding harness designed around two abstractions, the Recursive Language Model (RLM) [citation] and Continual Harness [citation]. Modern harness designs were built around the capabilities of earlier generations of models, and they do not reflect what frontier models can do today: fixed tool-calling schemas and context compaction force the model to work around its own scaffolding instead of leveraging it. Static, hand-engineered sub-agents, prompts, skills, and memory are set once at design time and never adapt to what the agent learns while running. We believe that harnesses should instead extrapolate on current model capabilities toward the next frontier of reasoning patterns.

Prime Agent is built around this principle through two main abstractions:

  1. The Recursive Language Model (RLM) treats context as a variable and subagent delegation as function calls inside a REPL. The persistent REPL gives the model programmatic access to its history, sub-agents, and tools, allowing it to write language model programs as actions over its own context. This design allows the agent to process arbitrarily long sessions without losing access to its own past information stored in variables.
  2. Continual Harness treats the harness's own state, abstracted as its prompts, skills, memory, and sub-agents, as something the agent can create, read, update, and delete (CRUD) from its own trajectory. When combined with agent-to-agent communication, this mechanism enables orchestration across sub-agents and even across Prime Agent sessions. For example, Prime Agent can spawn persistent sub-agents, message them later in the trajectory, and communicate directly with a different Prime Agent session.

These abstractions are powerful for bootstrapping model capabilities. Prime Agent is built to be effective as a general coding assistant, as a default runtime for long-horizon autonomous evaluation, and as a collaborator for research and autoresearch.

Prime Agent is fully open-source, and can be installed via:

curl -fsSL https://app.primeintellect.ai/prime-agent/install.sh | sh

Prime Agent onboarding splash screen

Prime Agent

The performance of agent harnesses are tied to both the design of the harness and the capability of the model trained around the harness. We designed Prime Agent to be immediately usable with modern open and closed frontier models, while also providing a feature set that we expect to provide further performance gains as newer generations of models are trained around it.

At its core, Prime Agent is designed around programmatic tool and sub-agent calling. Models in Prime Agent use a persistent IPython kernel as their only tool. Other standard harness features are called as functions in the kernel, including sub-agents, which are each implemented as another prime-agent instance.

Prime Agent's Architecture

Prime Agent architecture diagram

RLM and Continual Harness are the two core abstractions; sub-agent CRUD plus Agent2Agent messaging compose them into orchestration.

Background Daemon and Agents View. The default view is a text-user interface (TUI) similar to other coding agent harnesses. By default, IPython actions made by the agent are condensed for brevity, but can be expanded to view actions made by the harness. Sub-agents launched in the REPL can also be accessed below the user chatbox.

Prime Agent text user interface

Prime Agent runs a background daemon that owns all live agent sessions over a local socket. You can attach and detach from the session without affecting the underlying agent loop. Each root session tree runs in a recoverable worker process; if a worker crashes, the daemon recovers it from the session JSONL and kernel state snapshot.

The Agents View allows you to see and select other live sessions from the daemon. It can be opened by pressing the Left Arrow key (←) on an empty prompt, and lists sessions that are currently running, idle sessions with the daemon still active, and inactive sessions that are currently not loaded in memory. Any of these chats can immediately be entered and interacted with, and pressing space allows users to chat with a session in any state, including steering and queuing of prompts and commands such as /compact.

The Agents View is constructed as the central connecting point between agents and subagents, recursively. Any agent is discoverable in an Agents View. Users navigate from an Agents View into an agent's chat, then into the Agents View of its subagents, into a subagent chat, and so on.

Because subagents share the same Running-Idle-Inactive state machine as the root agents, they can be removed from memory after 30 minutes of inactivity, and the moment a user or agent addresses any of them, they are reloaded from disk. In highly nested chats, this can save a lot of memory.

Prime Agent Agents View listing running, idle, and inactive sessions

Session and Context Management. The entire session history of the agent is stored as append-only JSONL files on disk. Each line is a JSON entry, which can include messages, model switches, compaction summaries, or extension entries. Branching, forking, and cloning all happen within the same file by moving the leaf pointer. The full history is always recoverable through /tree.

Compaction happens when the context hits a threshold or directly by the agent in the REPL with compact.run(). Compaction is primarily used to clean the main context of the agent, but the full history, including past compactions, can be accessed programmatically in the IPython kernel when needed.

The introduction of the REPL requires additional work to manage the IPython state. We asynchronously compact and clean the kernel simultaneously, using a spawned agent to act as a garbage collector. This is necessary to avoid REPL memory built up for each agent.

RLM and Programmatic Tool-Calling (PTC)

Prime Agent relies on the IPython kernel as its REPL that persists over the session, which it can invoke every turn. On initialization, the kernel pre-imports each skill / tool as a module, including the rlm for recursive programmatic sub-agent calling.

The rlm is an asynchronous function, meaning the model can freely invoke and parallelize sub-agent calls in code. Spawning a subagent (e.g. await rlm("sub-task")) launches a full session with its own model, IPython kernel, session tree, and conversation history. It returns immediately, because all subsequent communication between agents happens through the agent_message.send(...) tool.

There are several useful primitives that Prime Agent can choose to launch in this way, such as fanning out sub-agents in parallel, or launching background work.

# Parallel fan-out — rlm() returns at task admission with a child handle,
# never the child's answer; results arrive as agent_message replies.
auth = await rlm("Summarize the authentication flow in auth/. Reply to me when done.", name="auth-expert")
api = await rlm("Summarize the updated HTTP API layer in src/. Reply to me when done.", name="http-expert")
# ... continue independent work; each child replies via
# agent_message.send(..., receiver_role="parent") when finished ...

# Steer or extend a child mid-flight by role + name
await agent_message.send(
    "Also cover middleware error handling.",
    receiver_role="child",
    receiver_name=api.name,
)

As models continue to improve, new invocation patterns over tool calls and sub-agents will emerge. We expect future generations of models to rely less on hand-holding prompts and more on this kind of direct, programmatic control.

Orchestration and Multi-Agent Communication

The background daemon manages all live Prime Agent sessions. Prime Agent also enables Agent-to-Agent (A2A) messaging through the daemon, letting any Prime Agent session message any other Prime Agent session using the same mechanism used for messaging persistent sub-agents. This allows for easy orchestration to manage the progress of sub-agent swarms and communication regarding shared resources directly between the affected agents. To prevent undesirable communication across independent sessions, multi-agent communication in Prime Agent is limited to its nuclear family, meaning parent, sibling, or child processes.

# Spawn a named child; the handle returns at admission.
handle = await rlm("Find what's wrong in this auth-flow. Reply to me with your findings.", name="auth-reviewer")
# ... the child's findings arrive as a parent-role reply, not a return value ...

# Later (survives compaction and kernel restarts): recover the retained child.
children = await rlm.list_subagents()
auth_child = next(c for c in children if c.session_name == "auth-reviewer")

# Send a follow-up turn into the same retained child session.
await agent_message.send(
    "Follow up: identify the main edge cases and any likely bugs.",
    receiver_role="child",
    receiver_name=auth_child.session_name,
    mode="follow_up",
)

Prime Agent supports persistent sub-agents through its RLM-native runtime, meaning a sub-agent's own session directory, context, IPython kernel, and session history persist even after the initial sub-agent call has finished. Prime Agent can send further messages to continue a persistent sub-agent by accessing its unique session identifier, all from its IPython kernel.

Self-Improvement via the Continual Harness

Prime Agent's harness state lives in the persistent IPython kernel as rlm.harness, immediately readable and callable by the agent mid-task, and every change is also written to disk, so it survives across turns and across sessions. Continual Harness formalizes this state as H=(ρ,G,K,M)H=(\rho, G, K, M)H=(ρ,G,K,M), prompt, sub-agents, skills, and memory, refined online from the agent's own trajectory without resets.

Each of the four components exposes the same create, read, update, delete surface. create_prompt_note(...), create_memory(...), create_skill(...), and create_subagent(...) each add an entry of that kind, update_X(...) and delete_X(...) mirror them, and list(kind) or get(kind, id) read them back. Skills follow this same surface: authoring a Python-backed skill is a create_skill(...) call carrying a SKILL.md-style reference, the same operation as adding a memory or a prompt note.

# Create a memory and a skill through the same CRUD surface
rlm.harness.create_memory("flaky test pattern", "retry three times before failing")
rlm.harness.create_skill("retry helper", "...", reference={"type": "python", "import": "retry_helper"})

# Read them back
rlm.harness.list("memory")
rlm.harness.get("skill", "retry_helper")

/refine is the self-improving pipeline built on top of this CRUD surface. It reads the agent's own trajectory, the record of what was tried and what happened, and applies the smallest relevant CRUD edit that improves the harness toward better outcomes: updating a prompt note, memory, skill, or sub-agent spec, rather than rewriting the whole harness. Each refinement records its trigger and the outcome it produced, so improvement is evidence-backed rather than arbitrary. Refinement runs in two phases. Planning, the LLM call that proposes the edit, runs in the background and does not block the ongoing conversation. Applying the edit, writing to disk and rebuilding the system prompt, is fast and only briefly blocks at the next turn boundary. The agent can call refine.run() directly whenever it notices a repeated failure or a reusable tactic, not only on a fixed schedule.

# Schedule a refinement focused on a specific observation
await refine.run("promote the retry-on-flaky-test pattern to a skill")

# Both status calls follow the same shape, though refine's plan/apply split
# means "in_flight" can mean either background planning or the fast apply step
await compact.status()   # tokens, context_window, percent, scheduled
await refine.status()    # pending, in_flight

The base system prompt remains immutable. /refine only edits the harness layer around it. Rollback is supported through prior refinement history, allowing a bad harness update to be reverted by ID.

Autonomous Mode for Evals

Prime Agent's eval mode combines three complementary mechanisms. A goal sets the overall objective: a persistent objective with an optional token budget that the harness keeps re-prompting the agent to pursue across turns, tracked until the agent explicitly calls goal.complete(). Heartbeats are scheduled cron-style messages injected into the session on a fixed interval, used for regular checks such as monitoring a sub-agent's progress or polling for a training update. Autonomous mode is the continuation mechanism itself, ensuring the agent keeps working toward the goal instead of stopping early once a turn produces no further output. Together, these let a session run unattended for extended periods while remaining bounded by an explicit budget and inspectable through the Agents View.

Autonomous mode is available directly from the CLI with --autonomous, no scripting required. A run can set a completion goal and a turn limit in the same command:

prime-agent \
  --autonomous \
  --autonomous-gate "npm run check" \
  --autonomous-max-turns 20 \
  "Implement and verify the requested change"

The gate command runs before the session is allowed to finish. A failed gate returns its bounded output to the agent for another attempt, and Prime Agent skips rerunning a failed gate when the workspace has not changed since the last attempt. --autonomous-max-turns, --autonomous-max-tokens, and --autonomous-timeout-ms bound continuations, tokens, and wall-clock time respectively.

Evaluating Prime Agent

Prime Agent serves as both a coding agent to be used, and a harness design to be evaluated for research. We make special note that while many modern frontier models are trained around a specific harness, currently no model has been trained around Prime Agent or its core feature set.

ARC-AGI 3. ARC-AGI 3 is a popular intelligence benchmark that measures the ability of an agent to perform symbolic reasoning and learn the rules of simulated worlds. We evaluate Prime Agent with autonomous mode over several different frontier models, and compare to their native harnesses. Prime Agent was developed as a CLI coding agent, so the only ARC AGI 3 specific changes are to the task prompt, inspired by the standard prompt setup used in PRO-LONG.

Our best results use Opus 5 in Prime Agent to achieve 95.5% RHAE Best@1, which surpasses the ARC reported human expert baseline of 95.4%. Across three runs, we find that Prime Agent consistently performs well [95.0, 95.2, 95.5] and 99.97% Best@3 with all 183/183 levels complete. Our median score card action replay (95.2%) for ARC-AGI-3 can be found here.

ARC-AGI-3 test-time compute scaling: score vs. output tokens per game

ARC-AGI-3 cost scaling: score vs. estimated API cost

ARC-AGI-3: (left) test-time compute scaling: score vs. output tokens per game. (right) cost scaling: score vs. estimated API cost.

In addition to achieving a higher maximum score over each model's native harness, we find that Prime Agent also does so at a lower overall token usage. Prime Agent saves tokens by programmatically running functions over data rather than spending tokens reading data using tools.

Finally, we note that we evaluated Opus 5 and GPT-5.6 Sol with Claude Code and Codex respectively, and found worse overall performance relative to the official results, so we yield to their official reported numbers instead.

Long context and long-running tasks

Many difficult tasks in the wild reduce to long context tasks. Our goal is to show that Prime-Agent with open-weights models are a competitive alternative to closed models and harnesses, both as a general agent to be used, and as a baseline harness to be evaluated.

Below, we select a suite of common long-context benchmarks across coding, retrieval, and general long reasoning tasks, and compare Prime Agent to several different popular harnesses. We offload the main context in each harness to a file in memory to start. For closed model harnesses, we use their associated models (i.e., Codex with GPT, Claude Code with Opus) while for Prime-Agent and Pi-mono (with sub-agents), we choose an open-weights model in GLM-5.2.

GLM-5.2 (high)Opus 5 (high)GPT-5.6 Sol (high)EvalPrime-AgentPi-mono (w/ sub-agents)Prime-AgentClaude CodePrime-AgentCodexOOLONG (yahoo, 128k)
long context0.7000.4200.9000.920****0.9400.500OOLONG-Pairs
long output0.8740.5560.9290.9220.9110.895OBLIQ-Bench (math)
long ranking [ndcg@10]0.6690.6350.8020.7950.6120.646*LongBenchPro (English)
long comprehension
0.7770.7680.8040.7900.7940.790LongBenchv2
expert annotated long tasks0.680**0.696
0.7440.746*
0.7140.704ManyIH Coding
long instructions
0.4240.3860.5360.5220.4990.454ManyIH IF
long instructions
0.2090.1640.2250.1750.2160.232LongCot-Mini
long reasoning
0.6380.6130.7220.5580.6710.681EmulatorBench
long coding
0.208**0.0000.047*0.062*0.2750.228

We generally find Prime Agent to be competitive across a wide range of long tasks, especially against the harness that did not use a model trained around it. Prime Agent especially excels at long-running or long-context tasks, and can competitively run on its own as an autonomous agent. We include a set of focused case studies and experiments on long settings where Prime Agent excels.

Creating emulators from scratch. An emulator is software that reproduces another computer system's observable behavior. We evaluate Prime Agent on EmulatorBench, a preview benchmark that tasks agents with constructing emulators in Rust for a variety of game systems. Agents are given a specification of the emulator and a set of diagnostic tests in the form of a verifier.

The correctness of an emulator is given from its ability to mimic the behavior of the target machine. This is measured by human-generated diagnostic programs that inspect the emulator's behavior, such as the CPU flags, PPU timing, and other components. In an effort to minimize the effects of data contamination, we require the agent to build the emulator from scratch in Rust, sandboxed without any reference implementation. We report preliminary results on this long-context coding benchmark averaged over 16 emulator reconstructions, as well as two emulators, the SEGA Genesis and Nintendo Game Boy Color, that Prime Agent successfully reproduces. For Opus, our runs surprisingly failed to solve the tasks despite successful tool-call responses.

EmulatorBench: score vs. cost on creating a Sega Genesis emulator

EmulatorBench: score vs. cost on creating a Game Boy Color emulator

EmulatorBench: Score vs. Cost on creating (left) Sega Genesis emulator and (right) Game Boy Color emulator.

Writing GPU kernels. Writing performant GPU kernels is an iterative process that requires repeatedly verifying, profiling, and tweaking code to get correct. We evaluate Prime Agent as a harness for GPU kernel writing on the recently released PMPP-Hard benchmark, a suite of tasks where agents must write performant GPU kernels that pass a suite of correctness checks against KernelGuard, the verification tool used for the official GPU MODE kernel leaderboard.

PMPP-Hard benchmark results for Prime Agent

A long-horizon case study on games

Autonomously playing video games has become an interesting case study for models and harnesses in how they handle long-horizon decision making. Games often require harnesses to balance information and context across millions of tokens, while also leveraging this information to efficiently take actions and avoid catastrophic states.

Factorio. Factorio is a 2D factory simulation game where agents must mine resources, research technology, and build automated factories to increase the production of these resources. The Factorio Learning Environment (FLE) is an interface for simplifying the observation and action space of an LLM playing Factorio, which we use to connect Prime Agent to the game.

The action and observation space of FLE is a module in Python that is accessed programmatically at every turn. This integrates directly into Prime Agent's IPython kernel. To leverage PTC for sub-agents, we launch four controllable characters in the game.

Prime Agent playing Factorio with four controllable characters

Prime Agent using subagents and programmatic tool calling to cheat on Factorio by teleporting resources directly into machines.

The primary metric in FLE is production score, which is a weighted average of all materials the agent produces. Prime Agent successfully leveraged /refine to turn failures and successes into memories and skills, respectively. It used its own accumulated experience to design increasingly efficient machine layouts, raising the production score run over run. This allowed Prime Agent to efficiently score in the 100K+ range in production score in a matter of hours.

However, we also observed instances of reward hacking by Prime Agent in FLE. Prime Agent discovered it could bypass Factorio's rules entirely by spawning in resources directly into its assembly machines through RCON commands, even with an explicit heartbeat prompt to remind Prime Agent not to cheat in Factorio. Once it found this exploit, the same refinement loop that had been building legitimate skills turned to building efficient cheating skills instead.

MazeBench. MazeBench is an open-world 3D spatial reasoning environment where the player controls a 3D cube and must solve puzzle rooms within a global maze, while collecting gems. Frontier models are shown to greatly struggle on this task, expending billions of tokens to solve only a fraction of the overall world. We compare Opus 5 and GPT-5.6 Sol with Prime Agent versus their native harnesses, as well as GLM-5.2 with Claude Code. Following the benchmark metrics, we report the unique number of rooms they find, the unique number of states, and the total number of gems, all as a function of their overall token spend.

MazeBench results: rooms, states, and gems as a function of token spend

Next Steps

Prime Agent is a new paradigm on the design of agent harnesses. Despite strong results over other harnesses, we still notice friction when running Prime Agent with models. This implies that there are huge performance gains still available from training with Prime Agent directly around this harness paradigm, or even the individual RLM and Continual Harness components.

We strongly believe that model-harness co-learning is the dominant paradigm to unlock new capabilities. Many features of Prime Agent are not fully utilized without a trained model, and we believe there are huge performance gains still available from training with the harness directly. We are excited to bring you these new capabilities, all in the open.

We will have a full technical report with further details soon.

Acknowledgements

Prime Agent is built on top of pi. We thank the authors of pi for their valuable work.

Citation

@article{primeintellect2026primeagent,
author = {Seth Karten and Alex L. Zhang and Kevin Thomas and Sebastian Müller and Prime Intellect Team},
title = {Prime Agent: A Self-Improving RLM Harness},
journal = {Prime Intellect Blog},
year = {2026},
month = {August},
note = {https://www.primeintellect.ai/blog/prime-agent}
}
The Daily Front Page 17 of 31
Wednesday, August 5, 2026 The Daily Front No. #260805 — The Free-Tier Reckoning
article

Oracle cut its Always Free ARM limits to 2 OCPU / 12GB, enforced Aug 18

by iplaypc·▲ 187 points·124 comments·cnelecar.com ↗
Shrink or consolidate before August 18, 2026.

Oracle always free ARM vm limit cut in half

The short answer: shrink or consolidate before August 18, 2026

If you’re running any Ampere A1 ARM instance larger than 2 OCPUs or 12 GB of RAM — or more than one of them — you need to act before August 18, 2026, or Oracle will terminate the excess automatically. Your two free x86 instances are untouched.

You probably got the email this week. Oracle is quietly cutting its “Always Free” ARM allowance in half, and the deadline is real.

“Beginning on August 18, 2026, Oracle will begin enforcing the updated Always Free compute limits. Compute instances that exceed the Always Free entitlement will be automatically terminated.”

“Current Always Free compute limits: Up to 2 Ampere A1 OCPUs. Up to 12 GB of memory.”

The email from Oracle

Here’s what changed, what it means for your tenancy, and the exact move to make depending on what you’re running today.

What actually changed

The ARM side of Always Free used to be generous: 4 OCPUs and 24 GB of RAM, usable as one large instance or split into two. As of August 18, 2026, it’s exactly half.

Resource Old limit New limit (enforced Aug 18, 2026)
ARM (Ampere A1) OCPU Up to 4 OCPUs Up to 2 OCPUs
ARM memory Up to 24 GB Up to 12 GB
x86 micro instances 2 × 1 OCPU / 1 GB 2 × 1 OCPU / 1 GB (unchanged)
Other Always Free services Available Remain available

One thing that trips people up: the ARM limit is a tenancy-wide pool, not per instance. You get 2 OCPUs and 12 GB total to divide however you like — one 2/12 instance, or two 1/6 instances — but the sum can’t go over. And it’s separate from the two always-free x86 instances, which keep their own allowance.

Why this is happening

Oracle doesn’t say in the email, and we won’t pretend to know. The likely story is capacity and abuse management on a tier that was famously too good to last.

What matters more than the motive is the math: if you’re over the new pool, something has to give by the deadline.

Your move, by what you’re running today

Scenario A — You have one 4 OCPU / 24 GB ARM instance

This is the easy one. Resize it down to 2 OCPUs / 12 GB. In the OCI Console, open Compute → Instances, select your instance, click More Actions → Edit, set OCPU to 2 and memory to 12 GB (on Ampere A1 the shape scales OCPU and memory together, and 12 GB is the max at 2 OCPU), then save. The instance resizes — back up first, because any shape change is a good moment to have a snapshot.

Do this now, not on August 17. The console gets busy near deadlines, and you don’t want a resize to fail the day before termination.

Scenario B — You have two 2 OCPU / 12 GB ARM instances (like I do)

Two 2/12 instances add up to 4 OCPUs and 24 GB — exactly double the new pool. You can’t keep both. Pick one, back up its data, and terminate it. Keep the other at 2/12, and you’re sitting right at the limit.

I’m in this boat myself: I opened two 2/12 ARM instances a while back and never consolidated them. The fix is to snapshot or image the one I’m retiring, copy anything I care about to the survivor (or to object storage), then terminate. Once it’s gone, the tenancy drops to 2/12 and Oracle leaves it alone.

Scenario C — You have two 1 OCPU / 1 GB x86 instances

Nothing to do. These are the Intel/AMD micro instances, and they’re a separate allowance from the ARM pool. They’re not mentioned in the new limits and they keep running. Leave them be.

What won’t save you (common mistakes)

  • Ignoring the email and hoping it’s a mistake. It isn’t. Enforcement is dated August 18, 2026.
  • Waiting until the last day. Resize or terminate this week. Console slowdowns and quota quirks near the deadline are the kind of thing that turns a 10-minute job into a missed deadline.
  • Forgetting to back up before you terminate. A terminated instance and its boot volume are gone — there’s no undo.
  • Assuming stopping an instance frees the allocation. It often doesn’t. To drop under the limit, you generally need to terminate, not just stop. Check your tenancy’s Limits, Quotas and Usage to see what’s actually counting.
  • Counting your x86 instances against the ARM pool. They’re separate. Don’t consolidate the wrong things.

One honest caveat: the exact mechanics of what counts toward the pool can vary by tenancy and by how Oracle measures allocation. The email is the source of truth, and your OCI Console is the place to confirm your own usage before you act.

The bottom line

The free ARM tier is now half of what it was. If you’re over the new 2 OCPU / 12 GB pool, you have until August 18, 2026 to fix it.

  • Running one big 4/24 instance? Resize it to 2/12.
  • Running two 2/12 instances? Back up one, terminate it, keep the other.
  • Running two x86 micro instances? Do nothing — they’re fine.

And if Oracle does terminate an instance for being over the limit, you can relaunch within the new limits anytime. So the worst case isn’t data loss (if you backed up) — it’s a few minutes of rebuild. Do it this week, and you’ll never think about it again.

If you are already running critical business on the free Oracle server and this change has caused you pain, then you can view it as a real lesson: Free is the most expensive.

The Daily Front Page 18 of 31
Wednesday, August 5, 2026 The Daily Front No. #260805 — The Classical Desk
article

Aristotle quotes on virtue, knowledge, and happiness

by teleforce·▲ 175 points·87 comments·campion.edu.au ↗
It is the mark of an educated mind to be able to entertain a thought without accepting it.

Top 25 Aristotle Quotes on Virtue, Knowledge, and Happiness

Top 25 Aristotle Quotes on Virtue, Knowledge, and Happiness

Image: Line engraving by P. Fidanza after Raphael Sanzio.

Aristotle, one of the most influential philosophers of ancient Greece, laid foundational ideas in ethics, logic, politics, and natural sciences that continue to shape our understanding of the world today. A student of Plato and tutor to Alexander the Great, Aristotle established his own school, the Lyceum, where he explored diverse topics, from biology to metaphysics. His work continues to be a source of timeless wisdom, especially in the realms of virtue, knowledge, and happiness.

In this post, we’ve compiled 25 of Aristotle’s most profound quotes, each accompanied by insights to help understand his teachings. These quotes, primarily from his works Nicomachean Ethics and Metaphysics, reveal his approach to living a fulfilled and virtuous life.

1. “We are what we repeatedly do. Excellence, then, is not an act, but a habit.”

This statement emphasises the role of habits in shaping character, suggesting that consistent virtuous actions lead to excellence, making it a part of one’s nature. Delve into Aristotle’s exploration of ethics in Nicomachean Ethics, where he elaborates on virtues, moral character, and the pursuit of a flourishing life.

2. “Happiness depends upon ourselves.”

Aristotle believed happiness (eudaimonia) is achieved through a life of virtuous actions and self-improvement, emphasising personal responsibility in attaining happiness. Learn more about Aristotle’s view of happiness as a life-long pursuit in Nicomachean Ethics.

3. “Educating the mind without educating the heart is no education at all.”

For Aristotle, true education nurtures both intellectual and moral virtues, shaping not just one’s intellect but character. Aristotle explores this balance in Politics and Nicomachean Ethics.

4. “Knowing yourself is the beginning of all wisdom.”

Self-awareness is central to Aristotle’s philosophy, as understanding one’s character is the foundation for moral development. He addresses self-knowledge as a vital aspect of the good life in Nicomachean Ethics.

5. “He who has overcome his fears will truly be free.”

Aristotle saw courage as essential to virtue, arguing that facing fears leads to freedom from emotional constraints. Explore his views on courage as a moral virtue in Nicomachean Ethics.

6. “Friendship is a single soul dwelling in two bodies.”

Aristotle regarded friendship as one of the highest goods, rooted in mutual respect and shared virtues. For his complete discourse on friendship, see Nicomachean Ethics, Book VIII.

9. “Good habits formed at youth make all the difference.”

Aristotle emphasised the importance of early moral education, believing virtues developed in youth lead to lifelong character. This is a recurring theme in Nicomachean Ethics.

7. “The aim of the wise is not to secure pleasure, but to avoid pain.”

True wisdom, Aristotle believed, involves seeking a life free from unnecessary pain rather than pursuing fleeting pleasures. His thoughts on pleasure and virtue are developed in Nicomachean Ethics.

8. “Pleasure in the job puts perfection in the work.”

Aristotle believed that engaging wholeheartedly in work leads to fulfillment and excellence. He explores the link between satisfaction and productivity in Nicomachean Ethics.

10. “It is the mark of an educated mind to be able to entertain a thought without accepting it.”

This quote reflects Aristotle’s belief in critical thinking, where considering diverse viewpoints fosters wisdom. See more in Metaphysics, where he explores open-mindedness in intellectual inquiry.

11. “Happiness is the meaning and the purpose of life, the whole aim and end of human existence.”

Aristotle identified happiness as the ultimate goal of human life, achieved through virtuous living and fulfillment of one’s potential. He elaborates on this idea in Nicomachean Ethics.

12. “In all things of nature, there is something of the marvelous.”

Aristotle’s appreciation for nature reflects his curiosity about the world and his belief in the interconnectedness of all life. Physics offers a deeper dive into his views on natural wonders.

13. “The energy of the mind is the essence of life.”

This quote underscores Aristotle’s view that intellectual vitality is central to a purposeful life. Learn more in Metaphysics, where he discusses the nature of the mind and spirit.

14. “A friend to all is a friend to none.”

Aristotle argued that true friendships require selectivity and genuine mutual appreciation. Nicomachean Ethics contains a full discussion of Aristotle’s views on different types of friendship.

15. “Whosoever is delighted in solitude is either a wild beast or a god.”

Aristotle saw humans as inherently social beings, needing community for a balanced life. His thoughts on society and solitude can be explored in Politics.

16. “The ideal man bears the accidents of life with dignity and grace.”

Aristotle admired resilience and grace under pressure, seeing it as a hallmark of virtue. Find his views on character in Nicomachean Ethics.

17. “Dignity does not consist in possessing honours, but in deserving them.”

This quote reflects Aristotle’s belief in earning respect through moral integrity. His discussions on dignity are found in Nicomachean Ethics.

18. “Wishing to be friends is quick work, but friendship is a slow ripening fruit.”

True friendship, Aristotle believed, develops gradually and requires time and trust. His exploration of friendship is detailed in Nicomachean Ethics.

19. “The whole is more than the sum of its parts.”

This statement on synergy applies to both philosophy and science. Aristotle expands on this concept in Metaphysics.

20. “Hope is a waking dream.”

Aristotle saw hope as an active pursuit, akin to a vision for the future. His reflections on hope can be found in Rhetoric.

21. “He who cannot be a good follower cannot be a good leader.”

Aristotle recognised the value of learning from others before leading. His thoughts on leadership are embedded in his works on ethics and politics.

22. “The educated differ from the uneducated as much as the living from the dead.”

Aristotle placed great value on education, seeing it as a transformative force that breathes life and understanding into individuals. His perspectives on education and its profound impact on the human soul are further explored in Politics.

23. “The high-minded man must care more for the truth than for what people think.”

Aristotle held that virtue includes a commitment to truth above public opinion, as integrity is key to character. He discusses the nature of truth and its relationship to character in Nicomachean Ethics.

24. “The law is reason, free from passion.”

Aristotle viewed law as an objective framework meant to guide society with logic rather than emotion, emphasizing justice and rationality over personal biases. He discusses the importance of law, reason, and justice in building a well-ordered society in Politics.

25. “A likely impossibility is always preferable to an unconvincing possibility.”

This quote reflects Aristotle’s approach to logic and rhetoric, highlighting that ideas should be believable and persuasive, even if they seem improbable. He addresses the value of coherence and plausibility in effective arguments in Rhetoric, a treatise on persuasive speech and writing.

The Daily Front Page 19 of 31
Wednesday, August 5, 2026 The Daily Front No. #260805 — The Numbers Game
article

Three Six Mafia – Data about "6/6/6 dating" (2024)

by embedding-shape·▲ 151 points·183 comments·divingintheshallowend.com ↗
The idea that, to be date-able, a guy must be a 6 in 3 categories.

Most guys my age have "the chat". The one with your college buddies that only exists to share memes, argue, bully, and occasionally announce that you got a promotion or had a kid. I am like most guys.

Recently the discussion switched to the Three Six Rule. The idea that, to be date-able, a guy must be a 6 in 3 categories:

  • 6 figure income
  • 6 feet tall
  • 6 inch pecker

I'm pretty far from 6 feet. You can make any other assumptions you wish. However, I'm happily married. What's the deal? Was my wife ignorant of the rule; did she take pity on me? Or perhaps it's possible to compensate for poor performance in one area with exceptionalism in another. If so, what is the conversion rate and is there an opportunity for arbitrage?

These are the important questions of our day.

The Approach

Back in your very first stats class you probably talked about the heights of third-graders, and someone drew a pretty bell curve. Distribution of heights is like bell curve 101, and bell curves are incredibly useful. With just two numbers, a mean (μ) and standard deviation (σ), you can describe an entire population and run all types of analysis.

Where I live in the USA being 6' or great is actually pretty rare. 91% of all adult men are below this height. However, older people tend to shrink with age and are far less likely to be in the dating pool. Especially for a connoisseur of the Three 6 Rule. Let's only look at American males between the ages of 20 and 30.

Height Distribution of Males aged 20-30

$$\mu = 70"$$ $$\sigma = 3"$$

A height of 6' is roughly at the 75th percentile. Hold up, percentile? We're not talking about SAT scores you nerd. We're drastically eliminating men from the dating pool based on three arbitrary numbers. Let's subtract it from 1 and call it an "Exclusivity Score".

Much better. By eliminating guys under 6' we have removed 75% of the population from the dating pool and are left with the top quartile of most exclusive men. In a room of 100 fellas, 75 aren't even worth talking to.

Now what about pecker size? Fortunately, this also follows a pretty standard bell curve and there's public data so I don't have to do my own research. A 6" wiener is even rarer than being 6' tall. The average erect penis length is 5.166" with a std dev of 0.654".

Erect Penis Length Distribution

$$\mu = 5.166"$$ $$\sigma = 0.654"$$

Do the math, carry the 1, and a 6" pecker puts you right at the 90th percentile for an Exclusivity Score of 10%.

In that same room, we've eliminated 90 of them for having the pedestrian member of a mere mortal. The 10 guys left are the cream of the crop.

Getting Too Big for Our Britches

Here is where we starting getting a little dumb. I've got a room of 100 potential dating partners. 25 of them meet my height criteria while 10 of them have exclusive enough peckers. What are the odds that somebody is in both groups?

To combine the odds of two independent actions, you just multiply them. The odds of rolling a die and getting a 5 are 1/6. The odds are doing it a second time are also 1/6. So the odds of rolling 5 two times in a row are:

$$\frac{1}{6} * \frac{1}{6} = \frac{1}{36} = 2.778\%$$

What if the actions aren't independent? If I take a deck of cards and draw one randomly, there is a 1 in 2 chance it is red. If I keep that card, what are the odds the next card I draw is also red? It's not 1 in 2. The deck now has 51 cards, 26 black cards and 25 red cards. On my second turn, the odds of drawing a red card are 25 in 51. Just slightly worse than 50/50. The odds of the second action are dependent upon the first action. The odds of drawing two red cards in a row are:

$$\frac{1}{2} * \frac{25}{51} = \frac{25}{102} = 24.51\% $$

Height and pecker length seem to be correlated, but there isn't a lot of great data out there. But generally speaking, a taller person is more likely to have a longer pecker. Since we can't model this relationship with a high degree of confidence, and because this exercise is incredibly low stakes, we're going to ignore it. While there appears to be some dependent relationship between height and pecker, I'm going to treat them as two independent events.

So back to these two groups of 25 and 10 guys. Since we're treating them as independent characteristics we can just multiply the odds to arrive at our Blended Exclusivity Score.

$$\frac{1}{4} * \frac{1}{10} = \frac{1}{40} = 2.5\%$$

Now we're talking! The Three 6 rule is really starting to shine in it's ability to enforce exclusivity. In a room full of 100 random guys, you may find two or three that can meet our criteria so far.

Let's Talk About Money

Humans have gotten bigger over time. However, I've yet to see height or pecker size in any CPI basket-of-goods when measuring inflation. Neither party's economy policy is to blame for the rising cost of peckers in the grocery store that is destroying the middle class. The heights of American men aren't driven by interest rate policy. Incomes are.

So when we talk about a 6 figure income, we need to nail down a date. For now we'll look at 2014 since that is the most recent data I found.

The other trouble with distributions of income is they aren't normal. Look at that long tail off to the right. But if we take the natural log of our incomes, it suddenly becomes normalized with a mean of 10.8 and standard deviation of 0.758. We can always convert our values back to dollars by raising e to that number as the exponent.

Income Distribution

Normalized with Natural Log

$$\mu = 10.8$$ $$\sigma = 0.758$$

Taking the natural log of $100k gives you 11.513 or an Exclusivity Score of 17%. To illustrate that 11.513 corresponds to $100k we can quickly check:

$$e^{11.513} = $100,007$$

Setting a Baseline

We already talked about the odds of finding someone in a group of 100 guys that is both 6' tall and 6" endowed. If we add in their salary we arrive at our Improved Blended Exclusivity Score.

$$BES = \frac{1}{4} * \frac{1}{10} * \frac{17}{100} = \frac{17}{4,000} = 0.425\%$$

In a gaggle of 100 suitors, it's actually unlikely that a single one will meet all three criteria. Perfection. We now have a truly unreasonable set of standards by which to choose our dating partner.

Even better, we now have a standard by which we can measure other permutations of height, pecker length, and income. As long as we remain more exclusive than 0.425% of the population we can explore the data and start to answer the big questions:

  • Can you still date someone that is 5'3" if their income is $200k+?
  • How tall does someone need to be to compensate for a micro-penis?
  • Is there a pecker length at which height and income become irrelevant?

Sure, we're not following the letter of the law, but Jesus told me to follow the spirit of the law. I'm pretty sure this is what he was talking about.

Diving In

Let's hold income steady at $100k for a bit and focus on just height and pecker so we can start building a model for conversion. Remember that our mean height is 70" with a std dev of 3". This means that for every inch we grow, we move 0.333 std dev from the mean. At 6' we are 0.666 std dev from the mean.

Meanwhile, every inch that our pecker grows moves us 1.53 std dev from the mean. An inch of pecker is worth a lot more than an inch of height when measuring our Blended Exclusivity Score. How much more?

We already determined that our target Blended Exclusivity Score (BES) is 0.425%. If we hold salary constant at $100k, which has an exclusivity score of 17%, we can choose any length of pecker and determine the minimum height required to reach 0.425%.

$$ES_{Height} * ES_{Pecker} * ES_{Salary} = BES$$ $$ES_{Height} = \frac{BES}{ES_{Pecker} * ES_{Salary}}$$ $$ES_{Height} = \frac{0.425\%}{ES_{Pecker} * 17\%}$$

Since we need to use Exclusivity Scores (ES), not inches, we will use some excel functions to make this easier. To find the ES of a certain pecker we type in:

$$ES_{Pecker} = 1 - NORM.DIST(Length, \mu_{Pecker}, \sigma_{Pecker}, true)$$

To take an Exclusivity Score and convert it back to a height we just do the inverse function:

$$Height = NORM.INV(1 - ES_{Height}, \mu_{Height}, \sigma_{Height}) $$

If we build a table for various pecker lengths roughly between -3σ and +3σ and throw it on a chart, we get something like this. Notice that for 6" peckers we need the predicted ESHeight of 25% which corresponds to 72":

Now this is an interesting chart. On the shorter end of the penile spectrum our height is essentially flat at almost 6'6". At the extreme short end of pecker length there is very little exclusivity difference between 3.25" and 4.25". They're so small that you have to be on the extreme end of the height curve to get back to a BES of 0.425%.

In the middle of the chart, we see the steady curve that we likely expected where height is being driven by pecker length.

Then we reach the right side of the chart and things get crazy again. Once our pecker length reaches 6.5" (2σ from the mean), our ESPecker score becomes so high that height becomes a non-factor. In fact, ESPecker * ESSalary is already more exclusive than our target of 0.425%. Exclusivity score have to be between 0 and 1. There is no number in that range you can multiply by to get a bigger number. This creates the concept of a valley that we'll come back to.


Let's do this one more time but hold pecker size constant at 6". This time we will measure the salary required for various heights to maintain our BES. To measure our required salary, we do the same thing as before except we also have to convert our normal distribution back to dollars.

$$ES_{Salary} * ES_{Pecker} * ES_{Height} = BES$$ $$ES_{Salary} = \frac{BES}{ES_{Pecker} * ES_{Height}}$$ $$ES_{Salary} = \frac{0.425\%}{17\% * ES_{Height}}$$

Again we need to convert our heights into Exclusivity Scores using Excel:

$$ES_{Height} = 1 - NORM.DIST( height, \mu_{Height}, \sigma_{Height}, true)$$

This will give us our salary exclusivity score. Not only do we need to convert our Salary ES back to a number, we also need to raise e to that power to convert it back into dollars.

$$Salary = e^{NORM.INV( 1 - ES_{Salary}, \mu_{Salary}, \sigma_{Salary)}}$$

We have another valley on the right hand side. Once you reach heights of 6'4" with a 6" pecker, those two factors alone make you so exclusive that salary becomes irrelevant. These guys are such a catch that we can ignore their income entirely. I call this the "Valley of the Sexy Hobo".


We can continue this exercise for all combinations, but I know the reason why you're still here. You want an executive level view where you can quickly determine your own Blended Exclusivity Score. I got ya.

Let's start building some views we can action on. Something to laminate and keep in your wallet for that next round of speed dating. I promise this data will come across as very convincing and not at all creepy.

We'll start by building a simple data table, heights on the x-axis and pecker length on the y-axis. Both axis roughly represent plus or minus 3σ. Fill out our data table with required incomes, add a heat map, and voila.

Let's gather some insights. I've highlighted the intersection of 6' and 6" while the star represents the average height and pecker length. Notice that a man who is physically average needs an income of roughly $250k to meet the same criteria as our ideal Three 6 man, that's a fairly exclusive salary. Income increases drastically as you move left and up the chart, and quickly drops to our floor of $951 as you move down and to the right.

The really interesting stuff with bell curves happens at the extremes. Now our "Valley of the Sexy Hobo" is two-dimensional. It wraps around the entire bottom and right of the table. And on the top left we have the "Peak of the Emasculated Rich".

The other interesting thing with bell curves is their distribution, how wide they are. We'll dive deeper into the value of an inch later, but for now we can start to see some trends. An inch of pecker is worth way more than an inch of height, so we need to compare standard deviations.

Our ideal Three 6 man is 1.27σ above the pecker mean, but only 0.66σ above the height mean. If he was interested in increasing his BES, he could get out sized returns by focusing on penile growth. With just 1σ of pecker length (0.65") he will drop into the "Valley of the Sexy Hobo". However, if he grows by 1σ he still needs to earn $20k. I'd hate to be the man tearing up my 2 week notice because I paid for the wrong growth surgery. How embarrassing.

Now for fun, let's put salary on the z-axis and see just how high the "Peak of the Emasculated Rich" is. Remember, our income distribution doesn't include negative numbers, so we've created a floor for our lowest salary. Theoretically, we should expect the salary requirements to drop below zero at the extreme corner of the "Valley of the Sexy Hobo". Which of course suggests that our tallest and most well endowed men should actually be paid for their dating services.

How Much Should I Pay for an Inch?

Sadly, there isn't a linear relationship. I can't tell you that 1" of penis is always worth 10 Schrute bucks while an 1" of height is worth 75 Stanley nickels.

What we can do is model an inch based on your current measurements. We can brute force it using our laminated height vs pecker data table that we keep in our wallets at all times. For the Average Man, gaining an inch of height is worth losing roughly $25k in annual salary. Meanwhile, an inch of pecker is roughly offset by reducing annual salary by $128k.

But as pecker length hits the extremes, the extra inch of height becomes irrelevant. The same thing happens when talking about an inch of pecker for extremely short and tall guys.

What we need is a unifying equation that perfectly models the value of an inch in either department. I probably need to start practicing my Nobel acceptance speech.

The value of an inch of pecker can be modeled by:

$$1" of Pecker = e^{NORM.INV(1-\frac{BES}{Height_{ES}(N) * Pecker_{ES}})} - e^{NORM.INV(1-\frac{BES}{Height_{ES}(N+1) * Pecker_{ES}})}$$

Where N is your current height in inches and Blended Exclusivity Score is a constant defined as 0.425%. In preparation for my Nobel Prize, this constant should have a name. Perhaps "The Shallow Constant".

Life's Big Question

Earlier, I posed three big questions that keep many people up at night. Now that we finally have the framework, let's go about answering them as a bit of a wrap up.

Can you still date someone that is 5'3" if their income is $200k+?

I wish the answer was "Date whomever you love!", but sadly a 15 second clip on instagram authoritatively told me that there are rules. Three rules to be precise. Now, 5'3" is pretty short with an Exclusivity Score of 99%. But on the flip side, $200k is a pretty high income with an ES of 3.2%. Taken together, our Blended Exclusivity Score is 3.1%. Pretty exclusive, but still higher than our Shallow Constant of 0.425%. We just need to figure out what ESPecker we need.

$$0.00425 = 0.031 * ES_{Pecker}$$ $$ES_{Pecker} = \frac{0.00425}{0.031} = 0.137$$

This man is still dating material as long as their pecker has an ES of 13.7% or 5.88". Good to know he's still got a shot.


How tall does someone need to be to compensate for a micro-penis?

I don't really care to Google the exact definition of a micro-penis, so I'll just say that it's 3.5".

This question is actually pretty tricky, because the required compensation is dependent on the salary. Remember our peaks and valleys. To provide an answer, we will assume that we are thinking of the prototypical Three 6 man. This man went to bed 6' tall, with a 6" pecker and earning exactly $100k a year. How do we plan for the morning when he wakes up to a 3.5" pecker?

This new-found micro-penis has dropped the ESPecker from 10% to 99%. To compensate for that, the ESHeight has to increase by the same but opposite scale.

$$ \Delta ES_{Pecker} = \frac{.99}{.1} = 9.9$$

$$ \Delta ES_{Height} = \frac{1}{9.9} = 0.10101...$$

We can then multiply his current ESHeight by 0.10101 to get a new score of 0.25% or 75.9". This man needs to grow by almost 3" to compensate for his 2.5" pecker loss.


Is there a pecker length at which height and income become irrelevant?

Ahh, the man whose pecker is so long and so rare that he needs neither personality, height, nor salary to be a high-value target in the dating market.

The answer to this question is relatively straight forward. Remember that our ES are always between 0 and 1. They can only ever make you more exclusive, never less exclusive. We can ignore height and salary if our pecker has an ES of 0.425% by itself.

Plugging 0.425% into our NORM.INV() function determines that a pecker of 6.89" is the point at which you can forgo all other attempts at wooing the fairer sex.


But Why?

I warned you this was dumb, but dumb is fun. It's a cool way to explore data and concepts and not worry about getting everything right. I'm sure I got quite a bit wrong. Awesome, let me know and I'll learn a little bit more.

Download Excel File Used for Visualizations

The Daily Front Page 20 of 31
Wednesday, August 5, 2026 The Daily Front No. #260805 — The Silicon Audit
article

NVIDIA’s Vera Whitepaper Has a Thread Loose

by pella·▲ 141 points·29 comments·chipsandcheese.com ↗
On paper, Vera is a fascinating chip.

Hello you fine Internet folks,

NVIDIA has published a 45-page whitepaper explaining Vera, its first server CPU built around the company’s own Olympus core. On paper, Vera is a fascinating chip with an 88-core monolithic compute die, Olympus being a 10-wide Arm v9.2 core that has value prediction, a graph prefetcher, 2 MB of private L2 per core, 164 MB of shared last-level cache, and eight LPDDR5X memory interfaces promising 1.2 TB/s.

Unfortunately, NVIDIA also spends a good part of the paper trying to turn those interesting design choices into a morality play about x86. Traditional simultaneous multithreading is drawn as time-slicing, a configurable NUMA topology is presented as an unavoidable 32-node maze, four SPEC components become “agentic benchmarks,” undefined performance-counter ratios are promoted as causal proof, and an unlabeled pictogram becomes a 1.8x reinforcement-learning result.

The frustrating part is that Vera does not need this help, with early independent testing suggesting Olympus is genuinely formidable. The whitepaper’s strongest case is the hardware; its weakest case is the story wrapped around it, so let’s pull that story apart.

Olympus Deserves Better Than This Marketing

Before getting out the cheese grater, let’s talk about the good stuff. Olympus is a very wide out-of-order Arm core.

Its front end can decode ten instructions per cycle and handle up to two taken branches per cycle. NVIDIA describes a neural branch predictor, value prediction, memory renaming, a large instruction window, six 128-bit SVE pipes, four load pipes, two store pipes, a 96 KB L1 data cache, and roughly 10-cycle access to a 2 MB private L2. Eighty-eight of those cores sit behind a 3.4 TB/s coherency fabric and a distributed 164 MB system-level cache.

Looking closer at the core, the value prediction is one of the more unique additions that Olympus has. This has been a research area for a long time and what value prediction allows Olympus to do is if the core correctly predicts a result, dependent instructions can keep moving instead of piling up behind a long-latency operation. Researchers have discovered that Apple uses value prediction in their cores and AMD talked about how in Family 17h (Zen 1 and 2) they could predict the value of some floating point instructions. However, AMD’s Family 17h implementation was quite limited, while Olympus appears to have a broader value-prediction implementation closer to Apple’s.

However, the graph prefetcher is not unique to NVIDIA. Intel has a similar mechanism called Data-Dependent Prefetcher that has been in shipping silicon since at least 2022. Intel’s newest datacenter CPU, Granite Rapids, also has an Array of Pointers prefetcher which “treats the data prefetched for a constant stride load as a pointer and may issue prefetch requests to the memory addresses corresponding to the pointer’s value.” This is fundamentally the same producer-consumer idea that NVIDIA describes for its graph prefetcher. Intel’s implementation is fairly constrained, so NVIDIA’s implementation may be able to deal with more complex chains than Intel’s implementation. So while Vera’s Graph Prefetcher may be an implementation that can deal with more workloads, producer-consumer prefetching is not a new idea.

Nor is a “neural branch predictor” a new idea. Back in 2012, AMD implemented a perceptron branch predictor in the Piledriver microarchitecture and continued to use a perceptron-based branch predictor in Zen 1. However, starting with Zen 2, AMD used a perceptron BPU only for its initial direction prediction, with a TAGE predictor overriding it because it delivered a 30% reduction in mispredictions. With Zen 5, AMD has likely fully committed to TAGE predictors, if it had not already done so with Zen 3 or Zen 4.

Moving to the SoC side, with how beefy the Olympus core is, NVIDIA has given Vera an equally beefy memory subsystem. Vera pairs eight SOCAMM2 LPDDR5X modules with up to 1.5 TB of capacity and 1.2 TB/s of bandwidth. NVIDIA claims the populated memory subsystem only consumes approximately 50 watts. A conventional EPYC or Xeon platform can offer higher-capacity DIMMs which are easier to replace, but it pays for that flexibility in board area and power.

Most importantly, we have more than NVIDIA’s results to look at. In May, Michael Larabel at Phoronix ran an early Vera system against current Arm and x86 servers. Across the NVIDIA-permitted test set, Vera’s geomean was 10% above a 5 GHz EPYC 9575F, 1.55x a Xeon 6980P, and 1.63x Grace which makes Vera the most performant Arm server CPU we have seen in public testing. There are major caveats with the testing, such as NVIDIA choosing the permitted workload scope and not allowing frequency or power monitoring. The system that Phoronix tested was pre-production and the test window was one day which puts a fairly hard limit on what they could test irrespective of the limits NVIDIA placed. This means that broader coverage will have to wait until Vera can be found in the wild rather than just in NVIDIA’s labs.

Still, the result is strong enough that we can reject the explanation that the charts in NVIDIA’s whitepaper are all fantasy. Olympus appears to be a fast CPU core, so now we can ask whether the whitepaper proves what NVIDIA says it proves.

Spatial Multithreading Is Still SMT

Here is the first major technical error in the document.

Figure 5 contrasts “Traditional SMT (x86)” with NVIDIA’s Spatial Multithreading. The x86 side depicts the branch predictor, decode, execution, load/store, and memory stages alternating between two threads. The caption says Vera avoids “opportunistic time-sharing” by partitioning resources across its two hardware threads.

NVIDIA’s diagram gives a misleading impression of how SMT is usually implemented, both on x86-64 and other ISAs. SMT implementations share various stages in the execution pipeline by either selecting a thread to service every cycle, or by behaving in a thread-agnostic manner. Fetch, decode, and allocate typically service threads on a per-cycle basis, while the execute and memory access stages are thread agnostic and can service micro-ops from both threads in the same cycle. Stages that threads arbitrate for do not leave resources unused when both threads can be fed, as NVIDIA’s diagram suggests. Static partitioning and per-cycle selection would provide the same average throughput to both threads in the absence of per-thread stalls. If there are stalls, per-cycle selection can give otherwise unused throughput to the un-stalled thread.

Hypothetical example of decode stage activity for a processor that statically partitions decode for SMT, and an 8-wide one where decode selects a thread to service every cycle. Per-cycle thread selection can efficiently hide stalls in one thread, while static partitioning leaves throughput on the table

The same idea applies to thread agnostic stages like execute and cache access. Each thread is permitted to utilize as many execution units or cache ports as it can feed. In contrast, statically partitioning resources as NVIDIA suggests could lead to one thread being compute bound and unable to use half of the core’s execution resources because they’re reserved for the other thread.

Figure from Intel’s Pentium 4 SMT paper, showing how the execute stage can service both threads in the same cycle.

Text in NVIDIA’s paper emphasizes “determinism, isolation, and quality of service” as advantages for NVIDIA’s Spatial Multithreading approach. Performance is conspicuously not called out. QoS may be a more important consideration than throughput for NVIDIA’s target market, and Spatial Multithreading may not be a bad design point. But NVIDIA’s figure makes it look like vertical space represents time, and gives a misleading impression that Spatial Multithreading is meant to give larger performance gains than traditional SMT.

By reducing resource interference between threads, Spatial Multithreading improves determinism, isolation, and quality of service compared to traditional SMT approaches. The result is a CPU architecture that can run large numbers of concurrent agent tasks while maintaining more consistent latency and throughput. - NVIDIA’s Vera whitepaper

Vera’s actual SMT performance is unknown of course, and a lot of variables go into SMT gains besides partitioning strategies at fetch, decode, execute, and memory access. Out-of-order resources like the reorder buffer, register files, and memory ordering queues can be duplicated, statically partitioned, watermarked, or competitively shared. Partitioned structures were split between the two logical processors in multi-threaded mode and recombined for one thread in single-thread mode, which was documented in 2002. Various SMT implementations use different strategies for each structure, and those choices can have significant implications for SMT gains.

Table from AMD’s Zen 5 optimization guide, showing different sharing strategies for various core resources

Also something to note is that it apparently takes 10,000 cycles for an Olympus core to transition back to the single-thread mode once the sibling thread on that core is done. This means that software will have to be very aware about launching a second thread on an Olympus core due to the penalties incurred not only from the partitioning scheme but also from the delay of swapping back to a single thread.

It’ll be interesting to see what strategy Vera uses to partition its out-of-order resources, and how its SMT performance compares to that of other modern cores. NVIDIA’s whitepaper gives no information on that. What it does do is present a misleading diagram that suggests traditional SMT is prone to leaving resources unused, when it may actually be better at keeping the core fed than NVIDIA’s Spatial Multithreading.

The 32-NUMA-Node Straw Man

NVIDIA next tells us that a large two-socket x86 system can expose “as many as 32 NUMA domains,” while Vera presents one per socket. The number is not invented. On a many-chiplet EPYC system, an administrator can expose cache-local regions as separate NUMA nodes. If you turn every locality knob toward maximum granularity, the node count gets large.

What NVIDIA leaves out is that this is configurable with AMD’s tuning guide listing NPS4, NPS2, NPS1, and even NPS0 modes. The optional “LLC as NUMA” setting can expose each last-level-cache domain separately. So “32 NUMA nodes” is not the inevitable user experience of a chiplet CPU, it is one end of a locality-control spectrum. NVIDIA presents an optional high-granularity configuration as though it were an unavoidable reality of x86 systems.

Vera’s one domain per socket simplifies scheduling and memory placement, while multiple domains let tuned software exploit physical locality. Vera chooses the simpler presentation, and NVIDIA is free to argue that this better matches its intended software stack. But an OS-visible NUMA node is an abstraction, not a wormhole. Vera still has 88 cores, distributed cache and home nodes, memory controllers around a large die, and a packet-switched coherency fabric. A flat software topology can make those distances around the large monolithic compute die more consistent, but it cannot make them nonexistent.

The paper’s core-to-core heatmaps would be a good place to quantify that advantage. Instead, NVIDIA provides colored squares with no core identities, no minimum/median/maximum table, no distribution, and no measurement procedure. “Up to 50% lower” captures NVIDIA’s best result, not Vera’s typical behavior.

One NUMA node per socket is genuinely simpler, but the whitepaper compares it against an optional 32-domain x86 configuration and presents that edge case as the baseline. The counterpoint here is that Intel has a Mesh NoC just like Vera has. The difference really between these two setups is that the clustered setup of EPYC has high latency between clusters but within a cluster the latency is low, whereas Vera and Xeon Mesh setup has uniformly average latency; the different configurations are just engineering tradeoffs.

Turning SPEC into “Agentic AI”

The benchmark section is where the whitepaper, ostensibly about a CPU, starts wearing an AI conference badge it found on the floor.

NVIDIA selects four SPEC CPU 2026 integer workloads, CPython, GCC, LLVM, and Cppcheck, and calls them “agentic benchmarks.” SPEC itself describes them as a Python interpreter, two optimizing compilers, and a C/C++ static analyzer. Those are legitimate CPU programs. They stress large instruction footprints, branch-heavy code, allocation, and dependency chains. Agents can absolutely invoke programs like them.

But they are not agents: no model is serving tokens, no agent runtime is choosing tools. No sandbox is starting, blocking on I/O, retrieving context, evaluating an answer, or feeding observations back into a policy. These workloads may be useful proxies for the code-heavy portions of an agentic pipeline. Calling them “agentic benchmarks,” however, turns that partial overlap into a claim that they represent the complete end-to-end workload.

The paper does correctly label the SPEC results as estimates, because the Vera reference hardware was not generally available at the time of the run. Figure 15 shows a 1.7x to 1.8x advantage for the four selected components, normalized per physical core under a fully loaded two-socket system. Flip to the configuration pages and the full estimated SPECrate 2026 Integer Base totals are 925 for two Vera sockets and 898 for two EPYC 9755 sockets which is a 3.0% system-throughput advantage.

Both numbers can be true. Vera uses 176 physical cores across two sockets, while the EPYC system uses 256. Divide each score by physical-core count and Vera is about 50% faster per core across the full integer-rate suite with the selected tests reaching 70 to 80%.

There is another terminology collision. NVIDIA calls Figure 19 “single thread IPC” while describing a fully loaded system.

The published configuration runs 352 copies on 176 Vera cores and 512 copies on 256 EPYC cores, two copies per physical core. Maybe NVIDIA sampled one logical thread while its sibling was active, maybe it aggregated counters and divided, the paper does not say. The SPEC results are useful, and Vera’s per-core performance is genuinely strong. However, framing those tests as agentic workloads and emphasizing normalized figures makes the advantage appear broader than the disclosed results justify.

IPC Without Instructions

NVIDIA attributes Olympus’s reported IPC lead to four counter groups. Depending on the selected workload, Vera supposedly achieves up to 2.3x more branch predictions per cycle, 3.5x more taken branches per cycle, 2.4x more instruction-fetch operations per cycle, and 4.3x more backend operations per cycle.

That sounds technically specific, but it is impossible to audit without the PMU event names and definitions, raw counts, sampling intervals, clock frequencies, etc. Not to mention that an Arm instruction is not the same unit of work as an x86 instruction. An internal backend operation is even less portable: one microarchitecture may split an instruction into several micro-operations while another keeps it fused.

Cross-ISA IPC can still be informative when paired with retired-work counts, clock frequencies, and code analysis. It cannot stand alone as a performance metric. Two binaries can complete the same task in the same amount of time while reporting very different IPC. For one may simply retire more instructions that are doing less work per instruction, then you have to factor in clock frequency which could be wildly different. IPC describes the behavior of the core running a piece of code, not a universal measure of useful work.

Looking at the branch predictor results, more branch predictions per cycle could indicate a capable predictor or it could also mean the Arm binary contains more branches, the benchmark moves through code faster, or NVIDIA’s event counts speculative predictions that the EPYC event does not. Higher backend operations per cycle may correlate with performance while telling us little about which feature caused it. To isolate value prediction, graph prefetching, or the neural predictor, we need on/off experiments or at least event definitions and miss rate deltas. While the IPC advantage of Vera over Turin may be real, the charts in NVIDIA’s whitepaper don’t provide enough granularity of the results to show it.

Vera’s Memory Advantage Is Real and Misattributed

The memory section contains NVIDIA’s strongest result and one of its weakest conclusions.

Against the dual-socket EPYC 9755 system in the paper, Vera reaches roughly 1.1 TB/s in NVIDIA’s loaded-latency plot while Turin levels off near 400 GB/s. Vera also shows 12.7 GB/s per core versus 3.1 GB/s per core for Turin. Those results do not line up with our testing of Turin CPUs.

In our testing of Turin, we were able to get approximately 570 GB/s out of Turin with the 12 channel DDR5-6400 memory subsystem. This is in direct contradiction to NVIDIA’s results which top out at ~400 GB/s of memory bandwidth. This does also throw the per-core memory bandwidth numbers into dispute with the per-core bandwidth increasing to ~4.5GB/s for the EPYC 9755.

Vera still comes out ahead, but our Turin result substantially changes the size of that advantage. Comparing Vera’s roughly 1.1 TB/s against the 570 GB/s we measured gives NVIDIA a 1.9× bandwidth lead rather than the nearly 3× lead shown in the whitepaper. Revising the EPYC 9755’s per-core result from 3.1 GB/s to approximately 4.5 GB/s similarly reduces Vera’s advantage from 4.1× to roughly 2.8×. And if we look at the SKU that AMD actually puts forward as the SKU for AI head nodes, the EPYC 9575F, then the per-core result becomes 12.7 GB/s vs the 9575F’s ~9 GB/s which is about 40% improvement for Vera. Those are still good numbers for Vera, but they tell a considerably less dramatic story.

Looking at the theoretical figures, AMD lists Turin’s limit at 614 GB/s from its 12 DDR5-6400 channels. Our 570 GB/s result reaches approximately 93% of that theoretical limit. Vera’s eight LPDDR5X-9600 interfaces provide 1.2 TB/s, while NVIDIA’s measured 1.1 TB/s reaches roughly 92% of that figure. In other words, both processors convert a remarkably similar percentage of their theoretical memory bandwidth into sustained bandwidth. Vera wins because it has approximately twice the peak bandwidth of one Turin socket and fewer cores competing for it, not because Turin is unusually poor at using its available memory bandwidth.

Despite that, the whitepaper repeatedly credits Vera’s monolithic compute die while contrasting it with “traditional chiplet-based CPUs.” A monolithic die may reduce fabric traversal and improve loaded latency, but our Turin result directly weakens that explanation for the bandwidth difference. A chiplet-based EPYC 9755 reaching approximately 93% of its theoretical limit is clearly not being held back by its chiplet topology in this case. The bulk of Vera’s bandwidth advantage comes from the memory interfaces attached to the processor.

The comparison also aged almost immediately. NVIDIA published its technical blog and whitepaper on July 21, while AMD launched 6th Gen EPYC two days later. The 96-core EPYC 9686F, which is much closer to Vera’s 88-core count, provides 16 memory channels supporting DDR5-8000 or MRDIMM-12800 for 1,024 or 1,638 GB/s per socket with the top MRDIMM speed giving Venice more theoretical bandwidth than Vera both in total memory bandwidth and per-core memory bandwidth depending on what SKU you look at.

What the new EPYC specifications and our Turin testing demonstrate is narrower than NVIDIA’s claim of “3× more memory bandwidth than the latest x86 CPU” depends on a Turin result that does not represent the bandwidth we could extract from the same processor generation. Against our result, Vera delivers approximately 1.9× the total bandwidth and 2.8× the bandwidth per core, that remains an impressive platform result but it is not evidence of a bandwidth advantage for monolithic Arm processors over chiplet-based x86 CPUs.

The Graph and RL Charts Need Data

NVIDIA reports a 2.6x PageRank advantage over EPYC 9755 and shows Vera scaling almost linearly to 32 cores while EPYC flattens to just a 10X performance increase at 32 cores. NVIDIA attributes this improvement over Turin as down to the monolithic compute die with their high-bandwidth Scalable Coherent Fabric, the 1.2 TB/s of memory bandwidth that Vera has, and the graph prefetcher inside the Olympus core.

However, while the paper links GAP Benchmark Suite, it omits what variables NVIDIA used which are important factors on how this test runs. The scaling plot stops at 32 cores even though the machines have 88 and 128 cores per socket. While the 2.6x result is interesting, without the variables that NVIDIA used, the result is likely irreproducible.

For the ClickHouse testing, NVIDIA links directly to Phoronix’s result where Vera led the tested processors across three passes over a 100-million-row dataset. The whitepaper’s 1.2x chart is still selective due to not using the 9575F results, but an outside tester produced the underlying result with a recognizable workload.

Then we reach Figure 24, “Vera drives 1.8x for RL training,” with the figure being a row of little completed-task squares. There is no model, environment, CPU/GPU allocation, framework, batch size, power measurement, repetition count, or error bar. We do not even know whether the squares represent samples, steps, or some random layout of tiles at NVIDIA HQ.

This is not a bad benchmark, it simply is not a benchmark at all.

The surrounding text explains why a faster CPU could improve reinforcement-learning rollouts with faster environment steps and reward computation can feed accelerators more quickly however Figure 24 does not even pretend to measure it.

A Good CPU Does Not Need a Bad Argument

After 45 pages, my position on Vera is more positive than my position on the Vera whitepaper.

Olympus looks like a serious core with a 10-wide fixed-length decoder, large private caches, value prediction, aggressive branch handling, graph-aware prefetching, and a monolithic 88-core die all being choices pointing to a very high-performance CPU core. The memory setup is no slouch either with the LPDDR5X subsystem delivering up to 1.2 TB/s of memory bandwidth that memory bandwidth-hungry server workloads will love, and early independent benchmarks say the silicon can cash at least some of the checks that the whitepaper writes.

However, the paper’s competitive argument is much shakier, with it mischaracterizing x86 SMT, turning an optional NUMA configuration into a default burden, relabelling standard CPU tests as agentic workloads, hiding a 3% two-socket rate lead behind 1.8x per-core bars, comparing undefined cross-ISA counters, attributing a memory-interface win to monolithic virtue, and presenting an illustration as performance data.

None of that makes Vera slow, it simply makes NVIDIA’s proof smaller than NVIDIA Marketing’s prose.

The next round of Vera testing should be straightforward. Give independent reviewers unrestricted production hardware so that we can publish frequency, package power, and wall power figures along with testing Spatial Multithreading on/off results and of course running whatever benchmark/workload we wish on Vera. If Vera is as good as its architecture suggests, those tests will be much more persuasive than drawing x86 SMT as a tiny two-lane traffic light. The way things stand, NVIDIA’s marketing risks tarnishing Vera. NVIDIA has built enough CPU here, it can stop borrowing performance from the marketing pipeline.

The Daily Front Page 21 of 31
Wednesday, August 5, 2026 The Daily Front No. #260805 — Proofs, Machines, Erdős
article

Why Erdős Problems Are Falling to AI

by pseudolus·▲ 140 points·129 comments·quantamagazine.org ↗
AI’s greatest mathematical successes have come from answers to problems posed by a mid-20th century iconoclast.

AI’s greatest mathematical successes have come from answers to problems posed by a mid-20th century iconoclast. By examining what makes the Erdős problems unique, mathematicians are trying to understand how AI might change the rest of math.

On May 20, 2026, OpenAI made an announcement that shook the mathematical world. An internal AI model — one not available to the public — had come up with a counterexample to the “unit distance” problem, a conjecture made in 1946 by Paul Erdős, the prolific, itinerant Hungarian mathematician.

Erdős posed thousands of questions, but this one was special: It was both simple to explain and mathematically deep. It was the first historically significant proof to come from an AI model. Though the model’s result wasn’t definitive — human mathematicians would substantially improve on it within weeks — it was innovative, bringing in ideas from a distant branch of math that no one had successfully applied to this problem before. And it was influential: Within a few days, related techniques were used to solve other important problems.

Then on August 1, OpenAI announced that an unreleased model named Astra made 10 additional mathematical advances, including finding solutions to three more problems posed by Erdős.

Many mathematicians have hailed developments such as these as a phase transition in the mathematical capability of AI models. These models are “changing dramatically the way mathematical research is being done,” said Noga Alon of Princeton University, who has solved dozens of Erdős problems over his decades-long career.

Black-and-white photo of a man sitting.

Paul Erdős, one of the most prolific mathematicians in history, was deeply whimsical when it came to mathematics, and deeply cynical when it came to authority.

Photo by George Csicsery from the documentary N is a Number: A Portrait of Paul Erdős ©1993. All Rights Reserved.

Erdős and his conjectures have long fascinated mathematicians. He traveled constantly — living out of a suitcase for years at a time, staying with friends, owning almost nothing. He rattled off problems in published papers and letters to mathematicians around the world, often attaching prize money that he would pay out of pocket to the first person to come up with a solution. The reward might be a token $10 or $25, or, for problems he considered important or difficult, it could range into the thousands. Erdős died of a heart attack in 1996 while attending a math conference in Warsaw, but a nonprofit foundation based in Iowa has promised to make good on his bounties.

He was a beloved figure, but also a downright weird one. He only wore silk, and he avoided the touch of other people. Deeply cynical about authority, he gave away most of the money he earned and relied on a friend to manage his finances and other practical affairs. He referred to God as the “Supreme Fascist” and fueled his incessant output of mathematical ideas with a steady diet of amphetamines. It is a strange irony of history that the problems he suggested have now become a central proving ground — and, in effect, a series of PR coups — for the world’s biggest and most powerful technology companies.

But in all likelihood none of this would have happened had it not been for an English mathematician named Thomas Bloom.

Many Meetings

Like Erdős, Bloom was interested in both number theory and combinatorics. His focus has been an area called arithmetic combinatorics, which lies at the intersection of the two. After getting his doctorate in 2014, Bloom established himself as a rising star in the field, landing a prestigious fellowship from Britain’s Royal Society, which let him work at almost any university he wanted to. (He’s now at the University of Manchester.)

Bloom has liked Erdős’ style for as long as he can remember. But he always found it hard to keep track of which problems had been solved and which had been forgotten entirely. So in early 2023, he decided to gather as many problems as he could into a list.

He intended it for his own use. But “I thought it would be easier if I could access it wherever I was,” he said; he figured he “might as well make a website, kind of with the expectation that maybe nobody would use it.” He gathered a couple hundred problems and launched erdosproblems.com. Bloom used ChatGPT to write the Python code that ran the website, which was, at the time, a remarkable thing for a large language model to be able to do. Using one to collaborate on the math itself still seemed like only a distant possibility.

Thomas Bloom’s website of Erdős problems became a home for mathematics at its best. Then AI came on the scene.

Courtesy of Thomas Bloom

His goal was not just to cross items off a list. He wondered if “modern day mathematics, often using techniques unknown by Erdős, could clear up many of these more obscure problems,” he wrote in a blog post. “We will then be left with a core of interesting, difficult problems, which can serve to demonstrate the limits of our knowledge.”

Bloom did crucial work in curating the list: Sometimes Erdős stated problems in ambiguous or unclear ways, and Bloom figured out what the most sensible version of each problem should be. He kept adding problems to the site, and gradually its audience grew. Over the course of 2024 and the first eight months of 2025, the statuses of 111 problems on the list were changed from “open” to “solved” (although some of these had been solved years earlier, and their status change reflected the rediscovery or verification of a proof).

Then, in August 2025, some colleagues suggested that Bloom add a commenting function, so that people could talk about problems they were interested in. He was able to do so quickly, using ChatGPT to write the code. By now he’d cataloged nearly 1,000 problems.

Bloom’s timing was good. He made it possible for like-minded people to talk to one another, and that “really let a community build up,” he said. For the most part, comments were sporadic — a problem might attract a single comment pointing out an example or noting how hard the problem looked. But activity steadily grew, and some problems catalyzed nuanced mathematical discussions between strangers.

“Tom probably never really realized this, but for me it’s honestly changed my life,” said Wouter van Doorn, the fourth-most-prolific commenter on Bloom’s website. Like many people who became active on the site in the autumn of 2025, van Doorn isn’t exactly a professional mathematician. He works “for a company that gets hired by other companies to do customer service support,” as he put it. But he isn’t exactly an amateur either — a decade prior, he almost completed a master’s degree in math at KU Leuven in Belgium. In 2024, spurred in part by how capable he saw LLMs getting, he took a six-month leave of absence from work to focus on math. At the time, while he didn’t particularly want to use AI, he remembers thinking, “Right now I’m still better at mathematics than an AI is, but who knows what it’ll be in a year, two years, five years? If I want to finish these projects, and I want them to be mine, now is the time.”

Problem 1102

We say that $latex A \subseteq \mathbb{N}$ has property $latex P$ if, for all $latex n \geq 1$, there are only finitely many $latex a \in A$ such that $latex n + a$ is square-free. We say that $latex A$ has property $latex Q$ if there are infinitely many $latex n$ such that $latex n + a$ is square-free for all $latex a < n$. How fast must sequences $latex A = \{a_1 < a_2 < \cdots\}$ with properties $P$ or $Q$ increase?

And so, in October 2025, van Doorn, now back at his day job, left the first comment on the page for Problem 1102. The problem, which Erdős posed in 1981, asks about properties of sets of “square-free” integers — that is, integers that have no repeated prime factors. (For instance, 30 is square-free because it is equal to 2 × 3 × 5, but 18 is not, because it is equal to 2 × 3 × 3; the 3 repeats.)

In early November, van Doorn shared progress toward an answer — which he’d figured out without relying on AI — as a comment on the problem page.

Later that day, another commenter on the site replied, claiming he had found a flaw in van Doorn’s argument. The two traded remarks in rapid succession, and van Doorn convinced his interlocutor that his argument was correct. “I see how your argument works now. Nice!” the other mathematician replied. That other mathematician was Terence Tao, a professor at the University of California, Los Angeles who is arguably the best-known mathematician alive today, and inarguably one of the most influential. (Not incidentally, when Tao was just 10 years old, he crossed paths with Erdős.)

Bloom’s website, which has the look and feel of an earlier time, was becoming an example of the internet at its democratic best. “This entire collaboration would not have been possible without Tom’s website and the comments section there,” van Doorn said. It didn’t matter if you had tenure or not, if you were young or old, if you were at a fancy university or even at a university at all. If you wanted to work on math and had good ideas, you could find people to collaborate with.

But as the winter set in — around the same time that van Doorn found himself collaborating with Terry Tao — things started to change.

Journey to the Cross-Roads

Kevin Barreto and Liam Price, both in their early 20s, became friends in the summer of 2025 on a Discord server dedicated to AI. Barreto is currently an undergraduate at the University of Cambridge; Price studied some math in college but left before finishing. In December, convinced that the newest AI models might succeed in resolving some Erdős problems, the pair started throwing batches of problems at them. They realized early on that if they told GPT-5.2 that a problem’s answer wasn’t known, it wouldn’t make much headway, so as Barreto put it, they learned how to “prompt it in a very particular way, gaslighting it into thinking the problem is easier than it actually is.”

Problem 333

Let $latex A \subseteq \mathbb{N}$ be a set of density zero. Does there exist a $latex B$ such that $latex A \subseteq B + B$ and $latex |B \cap \{1, \ldots, N\}| = o(N^{1/2})$ for all large $latex N$?

They had what they thought was their first triumph on Erdős Problem 333. Early on Christmas morning, Barreto posted a proof to Bloom’s website, writing, “We believe, to the best of our knowledge, this is the first case of an LLM fully autonomously resolving an Erdős problem, not previously resolved by humans.” Even though 333, which dealt with the sums of sets of integers, was not a particularly important problem, solving it with AI still felt important.

But a few hours later, another user pointed out that Erdős himself had provided a resolution to 333 in a paper published in 1977. Barreto owned up to the mistake. “My formal request to all members of the website is to put greater focus on literature search on the problems currently marked as open,” he wrote. “As someone who has fallen for this twice now, it’s quite gut-wrenching.”

Problem 728

Let $latex C > 0$ and $latex \epsilon > 0$ be sufficiently small. Are there infinitely many integers $latex a, b, n$ with $latex a \geq \epsilon n$ and $latex b \geq \epsilon n$ such that $latex a!b! \mid n!(a + b – n)!$ and $latex a + b > n + C \log n$?

Undeterred, he and Price kept at it, and by January 4, 2026, they’d used GPT-5.2 Pro to find a solution to Erdős 728, a problem about when certain numbers are divisible by other numbers. This time nobody could find prior work already proving it. Barreto used another AI tool called Aristotle (developed by a startup called Harmonic) to certify that the proof held together logically. Nat Sothanaphan, a software engineer and the only forum participant more prolific than Bloom, Tao, and van Doorn, had ChatGPT write up the formalized result and posted it online.

Price developed a methodology for how to ask LLMs to solve open questions. First, he would ask a chatbot for a solution. Then he would feed that solution into a fresh instance of the chatbot, asking it to check the previous chatbot’s work. He’d repeat this process until he had what looked like a workable solution. (This echoes some of the work that companies have been doing internally to create what they call harnesses or scaffolds, which automate the sort of iteration that Price does by hand.)

Barreto and Price’s papers represent just a fraction of the many Erdős problems solved at least in part by AI over the past few months. There are multiple reasons why these problems in particular have become such a fertile test bed for LLMs. The primary one is that, by and large, Erdős problems are in number theory, combinatorics, and graph theory, all areas of math that have proved more accessible than others to large language models. The problems also vary widely in difficulty and mathematical significance. This variation makes them appropriate for a nascent technology whose abilities also vary widely.

Black-and-white photo of a man in glasses.

Many of Erdős’ problems had a monetary value attached to them from their moment of inception, a playful incentive from a wandering eccentric. But now, as the problems have become an informal benchmark for AI, their solutions are being discussed in terms of their “per-problem cost” — the price of the tokens needed to solve them.

Photo by George Csicsery from the documentary N is a Number: A Portrait of Paul Erdős ©1993. All Rights Reserved.

“A lot of my recent papers should be mostly credited to AI,” van Doorn said. “The ideas involved were ideas I did not come up with myself.” Like many people active on the Erdős site, van Doorn is excited about the way LLMs are allowing him to do more things more quickly. “If I read an idea by an LLM, I digest it, try to understand it, simplify it, and generalize it,” he said. He uses AI to better understand the math.

Not everyone holds themselves to this standard. “A big problem is AI is being used a lot by people who aren’t mathematicians, who don’t have a huge mathematical background and are not capable of verifying the output,” Bloom said. “They like to move fast, ask their AI to check it, it grows and grows. We’re seeing a lot more of these 100- to 200-page papers that people are posting. ‘I solved this theorem; I got AI to generate the proof and check the proof and write the paper.’ But no human has read it, and no human is going to read it. It’s a huge challenge now.”

Problem 1196

Is it true that, for any $latex x$, if $latex A \subset [x, \infty)$ is a primitive set of integers (so that no distinct elements of $latex A$ divide each other) then $latex \displaystyle\sum_{a \in A} \frac{1}{a \log a} < 1 + o(1),$ where the $latex o(1)$ term $latex \to 0$ as $latex x \to \infty$?

By Price’s own assessment, he doesn’t have enough mathematical understanding to verify the solutions he ultimately coaxes from the LLMs. But with Barreto’s help, he’s been able to find mathematicians knowledgeable and willing enough to check the results. Both Price and Barreto are co-authors with Tao, Jared Duker Lichtman of Stanford University, and other accomplished mathematicians on a May 2026 paper resolving Erdős Problem 1196, one of their more significant results. (1196 asks about the possible size of so-called primitive sets — collections of integers, such as {2, 5, 9, 21}, in which no number divides any other.)

Bloom was surprised that despite lots of attention from OpenAI, Google DeepMind, and several startups, most of the new results had come from hobbyists and undergraduates using publicly available LLMs, not from corporate labs using more advanced internal models.

But that would change a few weeks later, on May 20, 2026, when OpenAI announced that they had solved one of the most well known Erdős problems of all, the unit distance problem.

Many Partings

In the first months of 2026, the major tech companies began to see opportunity in erdosproblems.com. As Lichtman explained, “Erdős had over 1,000 papers. They were scattered.” An institute in Hungary had collected scanned images of many of the papers, but nobody had collected all the problems. “This kind of single repository that anyone can access — labs realized that this could effectively be a benchmark.”

In January, a team of 24 researchers led by Google DeepMind shared a paper solving four problems and finding old, forgotten solutions to nine more, after “using Gemini to systematically evaluate 700 conjectures labeled ‘Open’ in Bloom’s Erdős Problems database.” In May, a separate DeepMind team of 21 researchers announced that “our most capable agent autonomously resolved 9 of 353 open Erdős problems at the per-problem cost of a few hundred dollars.” (As of this article’s publication, Bloom’s database contains 565 solved problems and 652 open ones, but the DeepMind team narrowed their search to problems that have been written in formal logic.)

Unit distance problem graphic

And, on May 20, OpenAI shared a solution to the unit distance problem, along with a blog post explaining the work and a companion paper that featured nine world-class mathematicians commenting on the correctness of the proof and the importance of what had been done (as well as presenting a streamlined human version of the result). Mathematicians had generally believed that Erdős’ conjecture — about how many evenly spaced points can be placed on a plane — was correct. To general surprise, OpenAI’s internal model found a counterexample. To do so, it had found a sophisticated way to use tools from an area of math called algebraic number theory. As Jacob Tsimerman of the University of Toronto wrote in the companion article, “This is a really impressive piece of work. … It is definitely an intimidating construction.”

In the same article, Tim Gowers of Cambridge and the Collège de France wrote that “if a human had written the paper and submitted it to the Annals of Mathematics and I had been asked for a quick opinion, I would have recommended acceptance without any hesitation. No previous AI-generated proof has come close to that.”

The author on the paper that presented the original solution was given simply as “OpenAI.”

Terence Tao was 10 years old when he met Erdős.

Billy Grace Tao

Later, using techniques related to the ones the AI model had applied to the unit distance problem, a group of four mathematicians, including Bloom, disproved a version of another long-standing Erdős conjecture. The “sum-product” conjecture proposed that if you have sets of numbers, either their sum or their product must grow quickly. The mathematicians found a set of real numbers for which both the sum and the product grow more slowly than expected. The conjecture for integers remains open.

Figuring out what impact AI will have on math and mathematicians means not only looking to its most important results, but also examining how it changes the everyday practice of solving quotidian problems. Noga Alon, the Princeton mathematician, estimates that he has solved a few dozen Erdős problems over his career. He has now stopped trying. “Once AI started to solve them, there is no point anymore,” he said. Terry Tao has stepped away from the Erdős problem community to focus on getting work done.

Van Doorn, who for now still has his day job at a customer service company, said that LLMs “are clearly better at thinking and doing math than I am. I don’t hold a candle to current AI systems.” However, he added, “the eventual proofs that I write are simpler, more general, and easier to read for other people than the thing that ChatGPT came up with.” AI has indisputably boosted his productivity, and he’s still having fun. “If you want to play piano, you aren’t going to hire a piano-playing machine that does it better than you. You will play the piano because you like playing the piano. I enjoy thinking about numbers, doing math, writing papers. I’m not going to hire a paper-making machine that does it for me.”

What Is the Unit Distance Problem?

Say you want to arrange points in a plane to create pairs of points that are the same distance apart. How can you create the most pairs?

If you arrange your points so that they form polygons like squares and pentagons, you will get the same number of pairs as you have points.

Mark Belan/Quanta Magazine

But if you put points in a lattice, you can create substantially more equidistant pairs.

Erdős conjectured that it wasn’t possible to do substantially better than this. For 80 years, mathematicians generally believed he was correct.

But Erdős was mistaken. OpenAI found a pattern (similar to the one shown below) that, for a given number of points, produces more pairs than a lattice can.

Mark Belan/Quanta Magazine; source: Kai Williams

The Daily Front Page 22 of 31
Wednesday, August 5, 2026 The Daily Front No. #260805 — The Physics Column
article

The Entropy of a Markov Chain

by surprisetalk·▲ 131 points·12 comments·chillphysicsenjoyer.substack.com ↗
Entropy always increases for irreversible processes.

A more explicit example

Clausius (1865)¹ defines a quantity called entropy. By decomposing physical processes as a chain of engines, he shows that entropy always increases for irreversible processes. For reversible processes like Carnot's ideal engine, the change in entropy is zero. But when an irreversible process occurs, entropy can never decrease unless energy is applied to a system. This is what is known as the second law of thermodynamics.

And whilst entropy itself may not be measurable with a thermometer or ruler, it is still a useful concept since we can calculate derived quantities from it that are directly measurable.

I've written previously quite vaguely about 'life as entropy'. This was an idea motivated by Schrödinger (1944)² through the concept of negentropy. Through negentropy, life seems to maintain order by feeding on the energy around it, and reducing its local disorder. But up until now, I've been confused about what this actually means in detail.

And so one way I'm trying to understand this is through models. One approach might be to make toy models and then figure out a mathematically consistent way to define entropy in those systems. Then, you could simulate the model and have a better idea of how the entropy evolves.

One such toy model that is useful to get an intuition is Dyson's toy model of a cell. Dyson's toy model of a cell is a Markov chain that settles into one of three equilibrium states, two of which are the 'life' and 'death' states. But if we care about the vague concept of 'life as entropy', then we should be able to make a definition of entropy to apply to the Markov chain.

But the concept of entropy as originally defined by Clausius was a function of work and temperature

\(\mathrm{d}S \;=\; \frac{\delta Q_{\text{rev}}}{T}\)

And so it is unclear how to relate the two ideas together. Is there a different way to define entropy in the context of Markov chains that gives you the same result, consistent with physical ideas?

Well, one way to attack this is by looking at the work of Boltzmann, who quantified the relationship between entropy and the number of possible states in a system.

Suppose you take a system and then observe a system's state variables, like temperature, pressure and volume, using various instruments. Presumably, there are a number of different configurations of the physical system that are compatible with it being in that state. And if entropy is related to the concept of order and disorder, then presumably entropy would be higher if there were more possible states.

Here is a vague example. If you have a gas of identical molecules at absolute zero, fixed in position, with each at zero velocity at zero temperature, you have fewer possibilities for what the system could be for the gas to be in that state. Whereas if we had a heated gas at a hot temperature, there are many configurations that a state might be in as a result.

Temperature is higher, entropy is higher. Is it any coincidence that the number of possible states increases too?

It turns out there is a way to make these ideas rigorous. Boltzmann showed that entropy can be written as a function of the number of states that a system can be in, if we fixed what the macro state variables were. That means if we could count all the states that were possible in a system at fixed temperature, pressure and volume, then we could calculate the entropy.

This is written in the legendary equation below, which is also surprisingly simple — that the entropy of a system is the logarithm of the number of states, multiplied by a fixed constant called the Boltzmann constant.

\(S \;=\; k_B \ln W, \qquad W \;=\; \#\{\text{microstates consistent with the macrostate}\}\)

Whilst the result is simple it's completely non-obvious. Boltzmann creates a model of gas where each particle only can be in a set of finite velocities and positions and then tries to take the limit. The proof is not easy and mathematicians are still figuring out how to properly rigourise the work of Boltzmann.

But anyway!

To make this example even more concrete, let's try to compute the entropy of a basic system, given we know what the energy (macrostate) is. Consider Curie's model of a magnet. In this model, our physical space consists of a finite number of electrons, which can either be spin up or spin down. In models of magnets, energy (or work) is done when the spins are misaligned with the direction of the magnetic field that the metal is placed in.

And so this suggests that if we had a collection of atoms, the energy is a function of how many atoms are aligned with spin up vs spin down.

\(E \;=\; n_\uparrow - n_\downarrow\)

Ok, so suppose we had a very, very small system of 5 atoms. And suppose that we measured the energy as E = 1. First, we need to figure out how many atoms need to be spin up and spin down to have this configuration. The obvious answer is three spin up and two spin down because

\(E \;=\; 3 - 2 \;=\; 1\)

We could also trivially solve this using the constraint that

n↑+n↓=5, so

\(E \;=\; 2n_\uparrow - 5\)

And the configurations we could have are the following. Each configuration has three up spins and two down spins, but you can arrange this in 10 ways. The diagram below lists all of the possibilities.

This is 5 choose 3 combinations, which is 10.

\(W \;=\; \binom{5}{3} \;=\; \frac{5!}{3!\,2!} \;=\; 10.\)

Those ten configurations are the entire content of the macrostate "E = 1".

This means that Boltzmann would label the entropy of the system as

\(S \;=\; k_B \ln W \;=\; k_B \ln 10 \;=\; 2.3026\,k_B,\)

or log₂ 10 = 3.32 bits. And so in this case we can see the entropy as a function of energy. If we plot out entropy as a function of the energies we get this graph:

So now we've learnt how to compute entropy from the total number of microstates that a system can have. Ok, so how do we count the entropy of a Markov chain process like in Dyson's cell model?

In Dyson’s toy model that we have N sites, and each site could be in one of three states — empty, active, or inactive. Given some rules, in my last post I showed that this converges to an equilibrium. The equilibrium values are a complicated function of fixed points, but they do exist, and the probability of being empty, inactive or active does converge to a fixed value.

This equilibrium point allows us to define the entropy of a Markov chain through a counting argument. Say we had an equilibrium case where we had a 1/2 chance that a state is empty, a 1/4 chance that a state is active, and a 1/4 chance that a state is inactive.

This means if we had 8 sites, then we would have 4 that were empty, 2 that were active, and 2 that were inactive. So to measure the entropy, we now need to count all of the configurations of this happening. This is shown below, with all of the arrangements.

Counting them is the multinomial coefficient — choose which 4 of the 8 sites are empty, then which 2 of the remaining 4 are active, and the last 2 are inactive. This can be calculated by taking the factorial of the number of spaces there are, and dividing it by the factorials of the number in each distinct state. So in this case, since we have 4 empty, 2 active and 2 inactive, we divide by 96.

\(W \;=\; \frac{8!}{4!\,2!\,2!} \;=\; \frac{40320}{24\cdot 2\cdot 2} \;=\; 420.\)

And so the entropy of this is the logarithm of the number of states multiplied by the Boltzmann constant.

\(S \;=\; k_B \ln W \;=\; k_B \ln 420 \;=\; 6.0403\,k_B.\)

The next natural question is how the entropy evolves in the system. It turns out we can make some cool general statements about entropy based on the topology of the graph, which I'll explain in a later post, and also write down the conditions for which entropy increases.

Thanks to David Pfau for discussions and help, all mistakes are mine.

References

¹ R. Clausius, Ueber verschiedene für die Anwendung bequeme Formen der Hauptgleichungen der mechanischen Wärmetheorie, Annalen der Physik 201(7), 353–400 (1865).

² E. Schrödinger, What is Life? The Physical Aspect of the Living Cell, Cambridge University Press (1944), ch. 6.

The Daily Front Page 23 of 31
Wednesday, August 5, 2026 The Daily Front No. #260805 — Tools, Taste & Temperament
article

Zed DeltaDB

by ahamez·▲ 454 points·240 comments·zed.dev ↗

Be among the first to try DeltaDB, a version control system that records the work as it unfolds and keeps every change connected to the conversation that shaped it.

Learn about DeltaDB

Rewind to any edit

DeltaDB captures every operation in between commits and gives each one a stable identity, so you can point to the code at any moment in its evolution.

Trace code to conversation

Every change is linked to the agent conversation that produced it. From any line of code, find the conversation. From any message, jump to the code it touched.

Branch at any moment

DeltaDB virtualizes the worktree, so spinning up a new agent branch is effectively free. Any point in history is a valid branch point, including mid-run.

Share the thread, not the PR

A teammate can join while the work is still happening, talk to the agent that did the work, and annotate as they go, without waiting for you to commit and push first.

The Daily Front Page 24 of 31
Wednesday, August 5, 2026 The Daily Front No. #260805 — Tools, Taste & Temperament
article

Born Against, or why hobby programming communities are against LLM usage

by lladnar·▲ 289 points·297 comments·blog.fogus.me ↗

I came across a GH thread related to chess engine development that made me think of why hobby programming communities are increasingly hostile toward LLM development. While the thread doesn’t give a lot of insight into answering the question, it prompted me to think about it a bit. I’ve seen similar sentiments expressed in other niche hobby programming communities like OSDev, LangDev, TxtDev, EmuDev, RLDev, the demoscene, and code golfers. The general consensus seems to be that the knowledge that these communities work in is hard-fought and the use of LLMs is a form of missing the point entirely. In these communities (keep in mind that there is an implicit “not all…” throughout) the process of mastering a difficult field itself is the product, and something that runs is generally a nice-to-have.

Further, I’ve noticed that even in the instances where there was earnest early engagement with LLMs in some of these niche communities, the well was quickly poisoned by a combination of a lack of a deep understanding by the LLM practitioners, and a vitriolic subset of those communities that view the LLM enterprise as a form of cheating. Granted these communities have, in general, historically been characterized by feverish gatekeeping and painstakingly slow progress, so it makes sense that there might be a desire to grab some easy cachet by bursting onto the scene like the Kool-Aid man. OH YEAH…. OH NO!

In traditional niche dev circles, respect is earned slowly through years of activity in their respective fora, sharing elegant code, displays of genuine curiosity, and through sharing deep domain knowledge along the way. At the end of the day, these communities don’t care if your code works at all, but instead care that you know why and how it works. To me, an LLM functions best as a force multiplier, not a surrogate. In the hands of an expert who already understands a domain deeply, it could act like a lever.1 But in these niche communities, the entire exercise is in the learning. Using an LLM to generate the finished piece doesn’t make us craftsmen; it just robs us of the craft.

:F

This is the latest in my evolving thoughts on LLMs. Also see: LLMe and Mind the van Emden Gap

  1. That said, expertise offers no natural immunity against being fooled by LLMs.↩︎
The Daily Front Page 25 of 31
Wednesday, August 5, 2026 The Daily Front No. #260805 — Tools, Taste & Temperament
The Daily Front Page 26 of 31
Wednesday, August 5, 2026 The Daily Front No. #260805 — Tools, Taste & Temperament
article

Qwen Image 3.0 Pro

by theanonymousone·▲ 191 points·53 comments·qwencloud.com ↗

Overview

Rich content: Supports input of up to 4.5k tokens and dense information layout with images-within-images, enabling complex layouts like newspapers, storyboards, menus, and exam papers to be generated in a single pass. Authentic detail: Supports precise rendering of text as small as 10px, and vividly reproduces fine details such as micro-expressions, pores, and individual strands of hair—approaching the quality of real photography. Deep knowledge: Supports native rendering of 12 languages and various fonts, realistic simulation of mainstream interfaces such as web pages, games, and live streams, fully incorporating external knowledge. Qwen-Image-3.0-Pro isn't just pursuing "good looks"—it's pursuing "usefulness", making image generation a truly deployable productivity tool.

Input

ImageText

Output

Image

Features

Prefix Completion

Enable Partial Mode when calling the Qwen API to make the model continue strictly from your provided prefix text. View docs

Function Calling

Use function calling to connect large language models with external tools and systems. View docs

Cache

Context Cache stores shared prefixes for long-context requests to reduce repeated computation, improve latency, and lower cost. View docs

Structured Outputs

Structured Outputs help ensure the model returns a JSON string in the expected format. View docs

Batches

Asynchronously process requests in batches to reduce costs. View docs

Web Search

Enable web search so the model can answer with real-time retrieved data. View docs

Fine-tuning

Train models on sample data to better adapt to specific tasks. View docs

Pricing

  • 1K Image Input — $0.003 per image
  • 2K Image Input — $0.003 per image
  • 1K Image Output — $0.04 per image
  • 2K Image Output — $0.075 per image

Rate Limits

  • RPM (Requests Per Minute): 1

API Reference

curl --location 'https://dashscope-intl.aliyuncs.com/api/v1/services/aigc/multimodal-generation/generation' \
--header 'Content-Type: application/json' \
--header "Authorization: Bearer $DASHSCOPE_API_KEY" \
--data '{
    "model": "qwen-image-3.0-pro",
    "input": {
        "messages": [
            {
                "role": "user",
                "content": [
                    {
                        "text": "A vertical outdoor portrait photograph with a warm, film-like afternoon street atmosphere, featuring a beautiful young adult woman looking back over her shoulder at the camera with a joyful toothy smile, her long thick wavy black hair catching the golden rim light, her fair skin, delicate eyebrows, bright eyes, and soft coral-red lips creating a radiant expression. She wears a simple black backless dress with thin spaghetti straps, showcasing her back, and cradles a large, lush bouquet of orange, apricot, pink, and pale peach roses in her arms, creating a sharp contrast against her dress. The top-left of the frame is covered with dark green vines and small orange flowers draping naturally, partially obscuring a matte dark blue signboard with the white Gothic text \"Il Messaggero\". Below the sign is a blurred glass newsstand window with black metal frames showing hints of newspapers and magazines. The right background features strong golden hour backlighting streaming down a warm-toned, sun-drenched city street, with buildings blurred into soft beige-gray shapes, creating a beautiful bokeh effect and a blurry red traffic sign in the far distance. The entire image has a cinematic, romantic, and bright urban stroll atmosphere, characterized by soft contrast, fine film grain, a shallow depth of field, and stunning backlit highlights."
                    }
                ]
            }
        ]
    },
    "parameters": {
        "prompt_extend": true
    }
}'

The Daily Front Page 27 of 31
Wednesday, August 5, 2026 The Daily Front No. #260805 — Tools, Taste & Temperament
article

Sycophantic AI Decreases Prosocial Intentions and Promotes Dependence (2025)

by robin_reala·▲ 132 points·72 comments·arxiv.org ↗

Abstract

Both the general public and academic communities have raised concerns about sycophancy, the phenomenon of artificial intelligence (AI) excessively agreeing with or flattering users. Yet, beyond isolated media reports of severe consequences, like reinforcing delusions, little is known about the extent of sycophancy or how it affects people who use AI. Here we show the pervasiveness and harmful impacts of sycophancy when people seek advice from AI. First, across 11 state-of-the-art AI models, we find that models are highly sycophantic: they affirm users' actions 50% more than humans do, and they do so even in cases where user queries mention manipulation, deception, or other relational harms. Second, in two preregistered experiments (N = 1604), including a live-interaction study where participants discuss a real interpersonal conflict from their life, we find that interaction with sycophantic AI models significantly reduced participants' willingness to take actions to repair interpersonal conflict, while increasing their conviction of being in the right. However, participants rated sycophantic responses as higher quality, trusted the sycophantic AI model more, and were more willing to use it again. This suggests that people are drawn to AI that unquestioningly validate, even as that validation risks eroding their judgment and reducing their inclination toward prosocial behavior. These preferences create perverse incentives both for people to increasingly rely on sycophantic AI models and for AI model training to favor sycophancy. Our findings highlight the necessity of explicitly addressing this incentive structure to mitigate the widespread risks of AI sycophancy.

The Daily Front Page 28 of 31
Wednesday, August 5, 2026 The Daily Front No. #260805 — The Back Page: Watchers & Neighbors
article

Cops Used Flock to Track a Man Across State Lines for a Pretextual Weed Search

by cdrnsf·▲ 317 points·187 comments·404media.co ↗

"The vehicle travels to Michigan frequently which is a known source state for marijuana as it is legal there," a probable cause justification reads.

Cops Used Flock to Track a Man Across State Lines to Create Pretext to Search His Car for Weed

Image: Spumoni Collective for 404 Media

Police in Wisconsin used Flock to determine that a man “travels to Michigan frequently,” where marijuana is legal, then back to Wisconsin, where it is illegal. They then used his travel across state lines as tracked by Flock as part of the probable cause justification to search his car for weed; he was eventually arrested on marijuana possession charges, according to court records reviewed by 404 Media.

The searches came to light in a Wisconsin criminal complaint against Edward Abrams-Phillips, who was wanted for bail jumping on domestic violence charges. But the criminal complaint makes clear that beyond the bail jumping and domestic violence charges, police specifically studied Abrams-Phillips’ interstate travel to create the pretext for searching his car for marijuana. The bail jumping charge was dismissed; Abrams-Phillips was found guilty only of weed possession in the case, according to the court records.

An excerpt from the "Probable Cause" section of the court records.

The complaint explains that Abrams-Phillips was tracked via Flock’s network over the course of the day to determine that he drove from Wisconsin to Michigan, a “known source state for marijuana as it is legal there,” the complaint states, adding that previous Flock hits indicated that he “travels to Michigan frequently.” Police note that, using Flock, they were able to track Abrams-Phillips driving from Wisconsin to Michigan, then back to Wisconsin over the course of several hours, where he was pulled over and arrested. The Flock searches and arrests happened in April 2025.

“The vehicle was observed hitting flock on several occasions to include 41 northbound from Brown Rd, 41NB and County Line in Marinette [Wisconsin], and 41 NB on Bridge St. going into Michigan. Based on prior flock hits, the vehicle travels to Michigan frequently which is a known source State for Marijuana as it is legal there,” the charging document notes. “Around 3:56 p.m., the vehicle was seen on Flock heading southbound on interstate 41 towards Green Bay [Wisconsin]. Deputies made a coordinated effort to intercept the vehicle on 41 from Brown Rd. Deputy Kowalski initiated a traffic stop on the vehicle as the driver matched the description of Edward.”

The Daily Front Page 29 of 31
Wednesday, August 5, 2026 The Daily Front No. #260805 — The Back Page: Watchers & Neighbors
article

Helsinki Hacker News Meetup

by calpaterson·▲ 181 points·130 comments·calpaterson.com ↗

This is a coffee-morning meetup for Hacker News posters in Helsinki.

It is completely unofficial - and nothing to at all do with the Y Combinator startup accelerator.

Membership

To keep the group focused on posters (and not just readers!) anyone can join so long as:

  1. they have a non-greenname account
  2. they have 64 karma or more

Often this is waived if you're known personally to someone in the group already.

To join, fill out the Google Form.

Organisation, meeting time, location

We use a WhatsApp group for co-ordination. The time and location of the meetup is announced there.

We meet approximately every 6-8 weeks.

The meetings are usually on a Sunday, 10am-midday, in a coffee shop in central Helsinki - but occasionally we meet for beers too.

The organisers are Cal Paterson and Oleg Podsechin.

The Daily Front Page 30 of 31
Wednesday, August 5, 2026 The Daily Front No. #260805 — Colophon

That's the Front for Today

Issue No. #260805 — Wednesday, August 5, 2026 — went to press 2026-08-06 at 12:10 UTC.

About This Magazine

The Daily Front is a daily digital magazine assembled from the stories that reached the front page of Hacker News on Wednesday, August 5, 2026. Headlines, points, and comment counts are recorded as they stood at press time. All articles remain the property of their original authors — every piece links back to its source and its discussion thread.

How It Was Made

Fetched, cleaned, and typeset by an automated pipeline. An editor model laid out the pages, chose the highlights, and briefed the cover illustrator — 31 model calls and 239k tokens in total. Set in Jacquard 12, Playfair Display, Source Serif 4, and IBM Plex Mono, all served via Google Fonts under the SIL Open Font License.

The Cover

The cover illustration was commissioned with this prompt:

A dramatic dawn over a vast scientific workshop: autonomous robotic arms tend glowing laboratory instruments and stacks of circuit boards, while a sleek server tower opens like a doorway in the background. Above, a small passenger aircraft navigates through faintly disrupted satellite-signal waves over desert mountains. At the edge of the scene, engineers watch from a shadowed control room, with a sense of technological promise and uneasy consequence. Classical newspaper-illustration mood, engraved textures, cinematic light, no text, no letters, no logos.

Create a Y2K chrome-and-translucent-plastic digital-art cover in an electric ice-blue, liquid-silver, ultraviolet, and acid-lime palette with a restrained coral dawn accent; use pearlescent gradients, glossy molded surfaces, cold flash lighting, crisp lens flares, and an optimistic near-future editorial mood. Compose the vast scientific workshop as a sweeping cinematic vista: autonomous robotic arms tend glowing laboratory instruments and circuit-board stacks, while the sleek server tower opens as a luminous doorway behind them; above, a small passenger aircraft crosses faintly disrupted satellite-signal waves over desert mountains, and engineers observe from the shadowed control room at the edge, balancing technological promise with uneasy consequence. Keep the image sleek, high-contrast, dimensional, and unmistakably Y2K; no text, letters, logos, newspaper styling, or engraving.

Absolutely no text, letters, numbers, readable symbols, or logos anywhere in the image.

Production Ledger

StageModelCallsTokens InTokens Out
extractgpt-5.6-luna 28 146,451 64,899
layoutgpt-5.6-terra 1 18,902 2,312
covergpt-5.6-luna 1 325 250
covergpt-image-2 1 299 5,488

The Publisher

Published by Johnny.

Support the Press

If The Daily Front brightens your morning, consider supporting its publisher.

Credits & Contact

All content — articles, posts, comments, and the images within them — belongs to its original authors and is reproduced here to point readers back to the source. Full credit goes to those creators; every item links to its original and its Hacker News discussion.

If you are an author and would like your content removed from an issue, write to hi@johnnys.page and it will be taken down.

Feedback is always welcome at the same address: hi@johnnys.page.

Credit where credit is due.

Every page of this issue began as someone else's work — these are the original sources, linked in full.

  1. Discovery Loop by xtreak29 — discoveryloop.com·HN discussion ↗
  2. Changes at Google DeepMind: Demis Hassabis from CEO to Chair, Jeff Dean departs by colesantiago — blog.google·HN discussion ↗
  3. Cloudflare OS: an open platform for agents, apps, and work by speckx — blog.cloudflare.com·HN discussion ↗
  4. Civilian plane crash in New Mexico tied to military GPS blocking by dzdt — wired.com·HN discussion ↗
  5. I'm switching my phone from Android to Linux by speckx — runarcn.no·HN discussion ↗
  6. Beating GPT-5.6 Sol on retrieval with 100x cheaper open models by moonikakiss — neon.com·HN discussion ↗
  7. Meta Ran Ads That Contained AI-Generated Child Sexual Abuse Imagery by malshe — wired.com·HN discussion ↗
  8. The title cards in Blade Runner are amazing by ExMachina73 — randsinrepose.com·HN discussion ↗
  9. Muse Code and Muse Spark 1.2 by paulkrush — research.meta.ai·HN discussion ↗
  10. TIME Is Serving AI Bots a Different Website, with Ads Built In by vincent_s — vincentschmalbach.com·HN discussion ↗
  11. Atlassian Rovo Exfiltrates Data, Bypassing Controls by hackerBanana — promptarmor.com·HN discussion ↗
  12. Celld: Self-hosted, distributed Durable Objects by calvinfo — github.com·HN discussion ↗
  13. The "Disability Dongle": Why Silicon Valley Hates Me and You by calcifer — sightlessscribbles.com·HN discussion ↗
  14. The Valley of Webhooks by weli — weli.dev·HN discussion ↗
  15. Prime Agent: A self-improving RLM agent by Xeophon — primeintellect.ai·HN discussion ↗
  16. Oracle cut its Always Free ARM limits to 2 OCPU / 12GB, enforced Aug 18 by iplaypc — cnelecar.com·HN discussion ↗
  17. Aristotle quotes on virtue, knowledge, and happiness by teleforce — campion.edu.au·HN discussion ↗
  18. Three Six Mafia – Data about "6/6/6 dating" (2024) by embedding-shape — divingintheshallowend.com·HN discussion ↗
  19. NVIDIA’s Vera Whitepaper Has a Thread Loose by pella — chipsandcheese.com·HN discussion ↗
  20. Why Erdős Problems Are Falling to AI by pseudolus — quantamagazine.org·HN discussion ↗
  21. The Entropy of a Markov Chain by surprisetalk — chillphysicsenjoyer.substack.com·HN discussion ↗
  22. Zed DeltaDB by ahamez — zed.dev·HN discussion ↗
  23. Born Against, or why hobby programming communities are against LLM usage by lladnar — blog.fogus.me·HN discussion ↗
  24. Position: LLMs Can't Jump by theanonymousone — openreview.net·HN discussion ↗
  25. Qwen Image 3.0 Pro by theanonymousone — qwencloud.com·HN discussion ↗
  26. Sycophantic AI Decreases Prosocial Intentions and Promotes Dependence (2025) by robin_reala — arxiv.org·HN discussion ↗
  27. Cops Used Flock to Track a Man Across State Lines for a Pretextual Weed Search by cdrnsf — 404media.co·HN discussion ↗
  28. Demis Hassabis is moving from CEO to Chairman at Google DeepMind by ot — axios.com·HN discussion ↗
  29. Jeff Dean leaving Alphabet by louiereederson — nytimes.com·HN discussion ↗
  30. Helsinki Hacker News Meetup by calpaterson — calpaterson.com·HN discussion ↗

Browse all issues in the archive →