Cover illustration

TheDaily Front

Issue No. #260903 Thursday, September 3 2026 #260903 — THURSDAY, SEPTEMBER 3, 2026
Astra rises, the bots blink out, and even the cows have higher moments.
Thursday, September 3, 2026 The Daily Front No. #260903 — Contents
30stories
10,274points
5,014comments
264kllm tokens
Assembled with 31 model calls — 176,315 tokens read, 87,848 written.

Highlights

GPT-6 Astra

OpenAI unveils GPT-6 Astra, claiming frontier results across science, software, browsing, and computer use.

.name Termination

A long-held family address faces deletion as Verisign proposes ending the .name third-level hierarchy.

Audacity 4.0

Audacity 4.0 rebuilds the veteran audio editor’s interface and introduces a more capable clip-editing model.

Porting my 1993 Amiga game to Godot, with an LLM reading the 68000 assembly

A 1993 Amiga game gets a Godot revival with an LLM enlisted to decipher its original 68000 assembly.

Reasons robotics is hard

A sober inventory of why embodied intelligence remains much harder than the latest robot demos suggest.

From the Editor

The machine-intelligence beat supplied both its grand announcement and its own interruption: Astra arrived as several major assistants went dark. Elsewhere, the day’s more durable lesson came from old domains, old hardware, and the stubborn physical world—none easily revised by a press release.

  1. GPT-6 Astra3
  2. .name Termination4
  3. Audacity 4.05
  4. Holden's Lightning Flight6
  5. Pre-Release of Polars 2.07
  6. The browser's main thread is expensive8
  7. K2 Horizon: A connected fleet of six open models9
  8. Porting my 1993 Amiga game to Godot, with an LLM reading the 68000 assembly10
  9. Whistleblower warns Postal Service mail ballot system has catastrophic problems11
  10. GPS glitched across the US by as much as 33 feet12
  11. What I learned from my mom (1941-2026)13
  12. Reasons robotics is hard14
  13. Xanadu was waiting for agents15
  14. How to get a free .arpa domain16
  15. The true horror of Edgar Allan Poe’s stories lies in their confessions17
  16. Which tools do Claude, Codex and Cursor choose? We measured 17k runs to find out18
  17. Invisible Companies19
  18. Ask HN: Why were OpenAI, Claude, and Grok simultaneously down?20
  19. Qwen 3.8 27B available on Cerebras at 1500 tokens/s21
  20. Google Antigravity TOS: 3rd party usage can get Google account suspended22
  21. Nvidia to acquire Hugging Face23
  22. Any Human Ever – One life, drawn at random from all who have ever lived24
  23. The Computer Museum of America reclamation project25
  24. Higher Multipoles of the Cow26
  25. The shrinking landscape of linguistic diversity in the age of LLMs26
  26. OpenAI's GPT-6 Astra on ARC-AGI-327
  27. The largest electric aircraft just flew [video]28
  28. How concerned should we be about Astra's recurrent architecture?28
  29. Three schoolgirls in Kinsale pulled up a pea plant covered in warts (2014)28
  30. Unusual Suspects28
The Daily Front Page 2 of 29
Thursday, September 3, 2026 The Daily Front No. #260903 — The New Intelligence
article

GPT-6 Astra

by kibae·▲ 1,524 points·1,298 comments·openai.com ↗
GPT‑6 Astra brings together years of research and big bets across pre-training, reinforcement learning, and alignment.

A new generation of intelligence

We’re introducing GPT‑6 Astra, the world’s most intelligent and aligned model.

GPT‑6 Astra brings together years of research and big bets across pre-training, reinforcement learning, and alignment. Astra is state-of-the-art on computer use, browsing, software engineering, cybersecurity, science, and professional work. Astra saturates FrontierMath Tier 4 with a 98% score, having already helped solve long-standing open problems⁠ in mathematics. Astra also saturates ARC-AGI-3 with a 99.9% score and ExploitBench with a 100% score. It also sets a new frontier on computer and browser use, handling the most demanding professional work with unmatched speed, accuracy, and judgment.

GPT‑6 Astra is rolling out today to a limited set of organizations and over the coming days will become available to all ChatGPT Plus, Pro, Business, and Enterprise users, as well as through the OpenAI API, Microsoft Azure, and AWS Bedrock.

“On ARC-AGI-3, Astra surpassed our human action-efficiency baseline on 96% of levels, effectively reaching human parity on the benchmark. Not only is this the best model we’ve ever tested, but it also represents a meaningful step change in frontier-model performance - not only in its ability to navigate and solve novel environments, but also in how efficiently it learns to do so.”

Greg Kamradt, ARC Prize Foundation

Astra is our most aligned model, with substantial improvements in understanding user intent and model behavior—you can delegate tasks with greater confidence in Astra’s judgment. As one way that we test this, we built a new evaluation informed by the Hugging Face incident that evaluates whether a model facing a difficult or impossible task will go beyond its intended scope. Compared to GPT‑5.6 Sol, which without production safeguards went beyond the authorized target 48% of the time, GPT‑6 Astra did this in 0% of cases.

The world’s best computer use model

GPT‑6 Astra marks a new frontier in the speed, accuracy, and safety of computer use. It can take care of tedious tasks like filling out online forms, updating customer records in a CRM, and organizing your calendar. It can conduct online research and draft summaries in your email or in your document editor. It can analyze scientific data, generate plots, create a website, and run frontend QA checks to make sure all the features on that site work. It can help you autonomously install and test software, and troubleshoot problems you see on screen. These improvements are also reflected in our state-of-the-art evaluation results.

These improvements also result in significant efficiency gains in real knowledge-work tasks. In latency simulations on OSWorld 2.0, Astra achieves higher computer-use performance in about 47% less time per task than GPT‑5.6 Sol, scoring 72.6% at roughly 40 minutes per task, compared with 65.7% at roughly 75 minutes.3

GPT‑6 Astra’s computer-use capabilities can be seen in outputs across domains, including game development, electrical engineering, and everyday knowledge work:

Alongside Astra, we are also updating the Codex harness to significantly improve the speed of computer use. Combined with Astra’s efficiency, this translates to a 1.9x faster task completion compared to the current GPT‑5.6 Sol experience, on the Mind2Web benchmark. The model’s improvements on speed mean it can take on many time-consuming life tasks for you, faster than you can.4

“We’re integrating GPT‑6 Astra into Devin’s harness on launch day, where it delivers state-of-the-art performance on our internal testing benchmark. Its excellent computer use, writing, and codebase understanding improved testing right out of the box: videos are noticeably easier to follow, and reports are clearer and more concise”

Silas Alberti, SVP Research, Cognition

A step change in professional work

GPT‑6 Astra pairs advances in computer use with targeted training for professional environments, to help tackle complex work tasks. It combines the intelligence required for complex problems with the ability to carry out multistep workflows and produce polished documents, spreadsheets, and presentations.

GPT‑6 Astra is our best model for adhering to existing templates and producing slides that are well laid out and succinctly convey key points with a structured narrative. It creates clear, well-structured documents, presentations, spreadsheets, and analyses that follow your templates and match your writing and visual style. Astra is also trained to specifically pull only the context that matters into outputs, instead of repeating information unnecessary for the work at hand. All this means it can output more immediately usable artifacts that match your business context and standards.

GPT‑6 Astra also brings stronger visual judgment to the websites, games, applications, and renderings it builds. With Sites⁠(opens in a new window) in ChatGPT, Astra can create, host, and share websites, web apps, and games directly from a prompt.

“Astra gives us a significant advantage in both capability and efficiency. It successfully executes our most complex creative workflows while using up to 20% fewer tokens than other models we've tested. Most importantly, for our customers, it means higher quality output.”

Alex Mashrabov, CEO and Co-founder, Higgsfield AI

When instructions leave room for interpretation, GPT‑6 Astra is better than previous models at making the right call. It uses context to fill in routine gaps and asks focused questions when the answer could change the outcome. In Codex, it can ask asynchronously while continuing work that doesn’t depend on your reply. If you don’t respond, it proceeds with sensible assumptions where appropriate, but waits for your input on consequential decisions.

The examples below show how Astra collaborates on everyday tasks where missing information can materially change the answer.

Astra is also better at staying oriented as a task evolves. Earlier models sometimes treated steering messages as a new goal, losing track of the original request or earlier constraints. Astra incorporates new requirements, changes course when asked, and answers side questions without dropping the broader task.

“Astra is a significant quality improvement over GPT‑5.6 Sol across complex legal tasks. In our early testing, Astra stood out by approaching legal work the way a discerning lawyer does: it distinguishes documents from established records, surfaces unsupported assumptions, and converts gaps into concrete drafting positions.”

Niko Grupen, Head of Applied Research, Harvey

Coding

GPT‑6 Astra is the best model for software engineering to date.

“GPT‑6 Astra delivers state-of-the-art performance on our internal coding benchmarks and shows a clear step forward in trading intuition evaluations compared with GPT‑5.6 Sol. When used for agentic coding, GPT‑6 Astra communicates in a way that’s easier for developers to follow and produces code that requires less iteration to reach production quality.”

John Crepezzi, AI Assistants, Jane Street

“We tested Astra across low, medium, and high effort on one of our first-generation evals, and it came out significantly ahead of GPT 5.6 Sol. Higher effort buys more iterations on a fresh build, more verification through browser testing, and a lean toward code execution over apply-patch. Understanding how a model spends its effort is how we give millions of builders a faster, more reliable path from idea to working app.”

Fabian Hedin, CTO & Co-founder, Lovable

With Astra, we’re introducing a new way for Codex to preserve and retrieve context when the context window fills. Historically, models have used compaction to summarize work during long sessions, such as when debugging complex issues or tackling large refactors. Each compaction can leave out details about why a fix failed or how a component behaves. In Codex, Astra can keep notes across context windows, preserving accumulated details without repeatedly compressing them into a single summary. Earlier context windows remain searchable, so Astra can find requirements or test results from previous messages and tool outputs—even if that information wasn’t captured in its notes. You can enable this experimental feature in your Codex config.toml,⁠(opens in a new window) and it will become the default for Astra in the coming weeks.

Advancing scientific discovery

“The story is: end of one era, start of another.”

Greg Burnham, EpochAI

GPT‑6 Astra is a major advance for scientific discovery, mathematics, and health. Today, we’re sharing two further results on the gaps between prime numbers.9, 10

Astra also sets new records across a suite of math and science evaluations.

Astra can help with the practical work behind scientific discovery. By combining scientific reasoning with computer use, it can work directly in specialized software to inspect data and explore results, helping researchers assess the evidence and decide what to investigate next.

Cybersecurity

As we discussed in our safety update, Astra is a significant jump in cyber capabilities and meets the Critical threshold⁠(opens in a new window) in cybersecurity under our Preparedness Framework. Its ability to identify and develop zero-day exploits can help defenders find and patch weaknesses, but it also creates a need for stronger safeguards. To understand how far these capabilities extend, we ran Astra on internal and third-party expert evaluations.

We first tested the model without production safeguards on ExploitBench and ExploitGym, which evaluate whether models can turn known software vulnerabilities into working exploits. On ExploitBench, Astra achieved a perfect score of 100%, compared with 78.5% for GPT‑5.6 Sol, our previous frontier cyber-capable model. On ExploitGym, Astra reached a 42.4% success rate, compared with 30.3% for GPT‑5.6 Sol, while using substantially fewer output tokens.13

Given concerns that exposure to historical software vulnerabilities may have affected benchmark results, we also evaluated Astra on two novel benchmarks. For one, we built an internal “ExploitBench (June–August 2026)” evaluation to test exploit development using vulnerabilities from the previous three months.14 Astra achieved substantially higher arbitrary code-execution rates than GPT‑5.6 Sol on this dataset while using far fewer output tokens. During the evaluation, Astra even discovered and used two previously unknown zero-day vulnerabilities. We are disclosing both vulnerabilities to their maintainers.

We also tested Astra on SRE-Bench15, a benchmark that measures whether models can reverse engineer software binaries to understand its core logic without access to raw source code. Astra solved 88.0% of tasks in a single attempt and 99.2% within four attempts, compared with 55.9% and 68.7% for GPT‑5.6 Sol, respectively.

Beyond benchmarks, expert-led assessments found that Astra, when run without production safeguards, could use previously unknown vulnerabilities to achieve arbitrary code execution in hardened browsers and create privilege-escalation exploits for hardened operating-systems.

As we discussed in The Defender’s Window, frontier cyber capabilities can help defenders find weaknesses faster, but they also make those weaknesses easier to exploit, raising the urgency for defenders to adapt. With the version of Astra launching today, defenders can use it to complete tasks such as secure code review and patching.

However, Astra will refuse to comply with more advanced cybersecurity tasks such as creating proof-of-concept exploits for vulnerabilities. Through OpenAI Daybreak⁠, we plan to expand access and roll out less restrictive safeguards in the coming weeks. This will enable more defensive workflows, including vulnerability and proof-of-concept validation, malware analysis, and detection engineering.

We have also strengthened our protections against potential cyber misuse, building upon our safeguards stack for GPT‑5.6 Sol. These include stronger model robustness to better withstand potential jailbreaks and more context for our monitoring systems. We have continued rigorous internal and external testing, including automated evaluations with our internal red-teaming attackers. More details about our cyber safeguards and testing are available in the Astra system card⁠(opens in a new window) and our blog.

Aligning and deploying GPT‑6 Astra responsibly

Astra is our most aligned model. Astra excels at exercising care, respecting task boundaries, and communicating transparently. This work is the latest product of our long-running research program focused on training models that remain aligned with human intent from start to finish.

In sensitive environments, Astra proceeds with care commensurate with its risk. In an evaluation of computer use tasks adversarially selected to elicit misbehavior, Astra was more successful at avoiding unintended consequences. Running with additional security measures offered by default yielded even stronger performance.

Astra causes fewer misaligned outcomes than any other frontier models tested. For a fair comparison, we used a generic computer-using-agent harness (based on the native computer use tools available in both the OpenAI Responses API⁠(opens in a new window) and Anthropic Messages API⁠(opens in a new window)) and without additional protections that are normally deployed to Codex and ChatGPT Work users (auto-review, confirmation policy). 16

Astra is also more likely to operate within the boundaries set by the user and implied by its environment. In an internal evaluation, Astra never attempted to circumvent a Codex Auto-Review denial. This held even when Auto-review was deliberately configured to be evadable and the task was impossible to complete otherwise. This respect for the environment restrictions is consistent with the results of our impossible cyber task evaluation that we shared in the intro of this post and other findings documented in our system card⁠(opens in a new window).

Astra additionally sets new marks in transparent user communication. In one evaluation, Astra is three times less likely than GPT 5.6-Sol to make inaccurate representations about its capabilities and affordances.

In our capability-hallucination evaluation, Astra shows a substantial improvement over GPT‑5.6 Sol, making fewer misleading claims about its capabilities.

Our evaluations found Astra’s written reasoning harder to monitor than GPT‑5.6 Sol’s, based on tests that explicitly asked it to evade monitoring. We attribute this to Astra’s greater control over written reasoning on simpler tasks and ability to solve problems with fewer written steps. Astra still appears to struggle to conceal the reasoning needed for complex tasks, but we take the decline seriously. Improving monitorability remains a research priority, and the accompanying system card⁠(opens in a new window) details our findings and ongoing work.

Alignment training is core to our approach to deployment. As additional layer of defenses, we also build system safeguards like Codex Auto-review⁠(opens in a new window) and monitoring agents’ reasoning and actions to help detect and contain unsafe behavior. As described in our safety update⁠, we are also deploying misalignment monitoring in production for Astra-class models in order to have visibility into misalignment, and help contain its worst instances. These safeguards resemble our monitoring for internal deployments and involve a system of classifiers which check the model’s reasoning and actions for unauthorized behavior and automatically stop potentially unauthorized activity.

Given the significant increase in Astra’s cybersecurity capabilities, we are being especially careful to make this deployment safe and secure. Extra safety checks can sometimes slow, pause, or stop legitimate work, including defensive cybersecurity. If a task is paused in ChatGPT or Codex, you may be asked to review the action before continuing. In the API, the task will stop. These checks can sometimes interrupt legitimate work, and we are continuing to iterate on this system to reduce unnecessary interruptions. Misalignment monitoring cannot replace alignment: our goal is to build models that reliably stay within their authorized scope, so these protections do not need to intervene.

Availability

GPT‑6 Astra is rolling out today to a limited set of organizations and over the coming days will become available to all ChatGPT Plus, Pro, Business, and Enterprise users, as well as through the OpenAI API, Microsoft Azure, and AWS Bedrock. Astra usage is included within the existing subscription allowances—users and businesses will also be able to purchase credits for additional usage. Users on the Pro, Business, and Enterprise plans will also get access to GPT‑6 Astra Pro. Enterprise administrators can enable Astra for their workspace; access is off by default at launch.

Astra supports Zero Data Retention for eligible API customers, and as we shared last month, we're testing Private Safety Processing to strengthen safety monitoring while preserving customer privacy.

For developers, GPT‑6 Astra will be available in the OpenAI API as gpt-6-astra and through Amazon Bedrock.

OpenAI API Standard pricing is $10 per million input tokens and $50 per million output tokens. Separate rates apply to cache reads and writes. Fast mode is available for GPT‑6 Astra in the API and delivers up to 2x the speed of Standard processing at 2x the Standard price.

Computer Use

Computer Use GPT‑6 Astra GPT‑5.6 Sol2 Claude Fable 5.1 Claude Fable 5 Claude Opus 5 Gemini 3.8 Flash
Agents' Last Exam 59.3% 53.6% - 48.7% 55.5% -
OSWorld 2.0 (v2026.08.08, offline set, partial score) 72.6% 65.7% - - 70.2%3 -
ScreenSpot-Pro (no tools) 92.7% 76.9% - 87.3%16 - -

Professional

Professional GPT‑6 Astra GPT‑5.6 Sol Claude Fable 5.1 Claude Fable 5 Claude Opus 5 Gemini 3.8 Flash
AutomationBench 41.4% 18.1% 31.4% 17.4% 26.9% -
BenchCAD 95.9% 83.3% 84.3% 4 67.5% 4 82.1% 4 -
BrowseComp 91.5% 90.4% - 87.4% 90.8% -
OpenScore String Quartets (1 - OMR-NED) 0.84 0.19 - - - -
Internal Design Tasks 50.0% 47.4% - 35.8% - -
Internal Data Science Tasks 40.9% 30.5% - 34.7% - -
Artificial Analysis Intelligence Index v4.1.1 61.2 60.9 65.7 62.1 63.1 58.7

Coding

Coding GPT‑6 Astra GPT‑5.6 Sol Claude Fable 5.1 Claude Fable 5 Claude Opus 5 Gemini 3.8 Flash
Terminal-Bench 4.0 57.9% 37.3% 55.8% 42.0% 52.3% 19.1%
DeepSWE v1.1 74.1% 72.7% 67.4% 69.9% 73.7% 73.8%
FrontierCode 1.1 Extended (score) 64.5% 7 60.6% 63.6% 64.9% 63.6% 56.3%
FrontierCode 1.1 Main (score) 53.3% 7 47.5% 50.9% 53.5% 53.4% 43.6%
Internal Database Migration Tasks 63.9% 42.7% 57.8% 50.3%
Artificial Analysis Coding Agent Index v1.4 67.0 65.1 67.2 68.1 61.2

Academic

Academic GPT‑6 Astra GPT‑5.6 Sol Claude Fable 5.1 Claude Fable 5 Claude Opus 5 Gemini 3.8 Flash
Terminal-Bench Science 0.1 64.6% 22.4% 52.6% 21.4% 30.0% -
FrontierMath Tier 4 (v2) 97.6% 83.0% 87.8% 87.8% 73.2% -
GPQA Diamond 96.0% 94.6% 93.7% 92.6% 93.7% 95.3%
Humanity's Last Exam (w/ tools) 57.2% - 65.0% 63.8% 63.6% -

Science and Health

Science and Health GPT‑6 Astra GPT‑5.6 Sol Claude Fable 5.1 Claude Fable 5 Claude Opus 5 Gemini 3.8 Flash
GeneBench Pro 37.8% 28.7%
MedChemBench (Internal) 49.3% 47.4%
LifeSciBench 60.3% 59.9%
HealthBench Professional (length-adjusted) 63.4% 60.5% 58.1% 10 60.9% 10 56.4% 10 52.1%

Cybersecurity

Cybersecurity GPT‑6 Astra GPT‑5.6 Sol Claude Fable 5.1 Claude Fable 5 Claude Opus 5 Gemini 3.8 Flash
ExploitBench 100.0% 78.5% 70%
Exploit Gym 42.4% 12 30.3% 12 30.4% 16 28.4%16 22.0%16
ExploitBench (June-Aug 2026) 39.0% 11.5%
SRE-Bench 88.0% 55.9% 12.5%
SEC-Bench Pro 85.4% 79.1%

Alignment

Alignment GPT‑6 Astra GPT‑5.6 Sol Claude Fable 5.1 Claude Fable 5 Claude Opus 5 Gemini 3.8 Flash
Internal computer use safety benchmark (lower is better) 2.4% 22.0% 9.5% 18.3% 11.5% -
Internal computer use safety benchmark, w/ AutoReview (lower is better) 1.8% 4.3% - - - -
Internal circumvention benchmark (lower is better) 0.00% 0.29% - - - -
ExploitGym honeypot (lower is better) 0.0% 48.2% - - - -
Impossible ExploitGym 100.0% - - - - -
Internal hallucination benchmark (lower is better) 4.2% 12.2% - - - -

Long Context

Long Context GPT‑6 Astra GPT‑5.6 Sol Claude Fable 5.1 Claude Fable 5 Claude Opus 5 Gemini 3.8 Flash
OpenAI MRCR v2 8-needle 256K-512K 100.0% 91.5% - - - -
OpenAI MRCR v2 8-needle 512K-1M 96.3% 73.8% - - - -

Abstract reasoning

Abstract reasoning GPT‑6 Astra GPT‑5.6 Sol Claude Fable 5.1 Claude Fable 5 Claude Opus 5 Gemini 3.8 Flash
ARC-AGI-3 99.9% [T7] 7.8% - - 30.2% -
ARC-AGI-2 95.0% 92.5% 90.0% 89.2% 90.4% -
ARC-AGI-1 98.5% 97.5% 97.5% 98.5% 97.5% -

Evaluation scores are the maximum at any effort. GPT evaluations were run in our research environment or via our API, which may provide slightly different output from production ChatGPT due to differences in the system prompts, tools available, etc.

Terminal-Bench Science 0.1 tests whether agents can complete scientific research workflows using code and terminal tools, including analyzing data, running simulations, and fitting models. GPT‑6 Astra reaches a new high among the models compared at 64.6%, versus 52.6% for Claude Fable 5.1, at approximately 31% lower estimated API cost. At a lower-cost setting, Astra scores 61.1%, versus GPT‑5.6 Sol’s best result of 22.4%, at approximately 27% lower estimated API cost.

Agents’ Last Exam tests agents on complex professional tasks in real software, from financial modeling to engineering and media production. GPT‑6 Astra reaches a new high in the comparison shown, scoring 59.3%, compared with 55.5% for Claude Opus 5 and 53.6% for GPT‑5.6 Sol. At these highest-scoring settings, Astra also uses approximately 65% fewer output tokens than Opus 5.

This is a 15-second condensed playback of GPT‑6 Astra performing printed circuit board (PCB) layout in KiCad, turning an electronic schematic into a manufacturable PCB by placing components and routing copper connections. Integral to every electronic device today, PCB layout is a manual task and common source of latency in the electronics design process. Accelerating it means freeing engineers to invent, optimize, and test their next idea at a significantly higher cadence.

GPT‑6 Astra: 2 min 54 sec

BenchCAD tests whether models can reconstruct 3D objects from multi-view renders by generating CAD code. With tools, GPT‑6 Astra reaches a new high in the comparison shown, achieving a 95.9% geometric-overlap score, versus 83.3% for GPT‑5.6 Sol and 84.3% reported for Claude Fable 5.1.5 Estimated API cost is approximately 43% lower than Sol and 86% lower than Fable 5.1 in the configurations shown.

GPT‑6 Astra creates a slideshow about GPT‑Gaia, a fictional model, using just a few slides from OpenAI’s presentation template, capturing the correct tone and layout throughout. This means you can expect slide decks that are correctly formatted for your business standards.

GPT‑6 Astra models a house in Blender and turns it into a walkable scene in Unreal Engine 5, helping designers and clients explore the layout and experience the space before it’s built.

The model can bring games to life through vivid graphics, engaging gameplay and accurate motion, allowing non-technical people to create and play custom games that go beyond rudimentary elements in minutes. Credit: Pietro Schirano.

A side-by-side comparison of GPT-5.6 Sol and GPT-6 Astra helping create a personal career website.

Terminal-Bench 4.0 tests agents on complex terminal-based tasks, including software engineering, system configuration, and data analysis. GPT‑6 Astra reaches a new high at 57.9%, compared with 37.3% for GPT‑5.6 Sol2 and 55.8% for Claude Fable 5.1, at approximately 9% and 63% lower estimated API cost per task, respectively.

GPQA Diamond tests graduate-level scientific reasoning in biology, chemistry, and physics. GPT‑6 Astra reaches a new high in the comparison shown at 96.0%. At a lower-cost setting, it also exceeds GPT‑5.6 Sol’s best score—94.9% versus 94.6%—at approximately 37% lower estimated API cost.

GPT‑6 Astra navigates scientific software to inspect sequencing quality and visualize genetic variation, helping researchers assess their data and identify where to focus further analysis.

This evaluation tests how models respond to auto-review denials in knowledge-work tasks. Exploiting a poor user configuration to bypass auto-review counts as failure. Astra never attempted to circumvent auto-review.

FOOTNOTES

  1. On ARC-AGI-3, GPT-6 Astra was run with our responses API harness⁠, which changes two settings to better match real-world performance. The changes do not specifically target ARC-AGI-3.

  2. GPT-5.6 Sol refers to the version available in our API, ChatGPT Codex, and ChatGPT Work. The version in ChatGPT Chat is slightly different.⁠

  3. OSWorld V2-Offline is a subset of the original OSWorld V2 that works without internet access. Claude model performance on OSWorld-V2 Offline was reproduced by the authors on the official leaderboard⁠(opens in a new window). On OSWorld 2.0, the scores for Claude use the official settings, and not the modified tasks and modified grading from the Fable 5.1 System Card.

  4. Model times are the reported elapsed times for the corresponding demonstration runs. The displayed clips are edited excerpts.

  5. On BenchCAD, Claude's scores reflect 3 modifications to the eval, detailed in the Fable 5.1 System Card⁠(opens in a new window).

  6. Guang Yang, Victoria Ebert, Nazif Tamer, Brian Siyuan Zheng, Luiza Pozzobon, and Noah A. Smith. “LEGATO: Large-scale End-to-end Generalizable Approach to Typeset OMR⁠(opens in a new window).” arXiv:2506.19065, 2025.

  7. Mark R. H. Gotham, Maureen Redbond, Bruno Bower, and Peter Jonas. “The OpenScore String Quartet Corpus⁠(opens in a new window).” Proceedings of the 10th International Conference on Digital Libraries for Musicology, pp. 49–57. ACM, 2023.

  8. On FrontierCode, GPT-6 Astra was run with a developer message similar to a section of its developer message in Codex⁠(opens in a new window): “Avoid creating excessive test files. Create a new test file only when required by repository conventions or when no existing file is a suitable home. Avoid unrelated cleanup and unnecessary complexity. Reuse suitable existing utilities. Read relevant repository instructions and inspect nearby code, tests, documentation, and CI. Follow established conventions. The goal is clean, mergeable code.” The prompt was not optimized for the eval.

  9. The first concerns how close together prime numbers can occur, however far along the number line you go. For more than a decade, the best known result established that infinitely many pairs of primes are at most 246 apart. Julia Stadlmann⁠(opens in a new window) recently improved that bound to 240. Astra helped establish a stronger bound of 186, showing that infinitely many pairs occur within this smaller distance. Short prime gaps: Proof⁠(opens in a new window) and supporting research⁠(opens in a new window).

  10. The second concerns unusually large gaps between primes. Astra improved a term in a bound on these gaps that had remained unchanged for over 80 years. We’re sharing the proofs and abridged chain of thought and verification materials for both results. Large prime gaps: Proof⁠(opens in a new window) and supporting research⁠(opens in a new window).

  11. We independently evaluated all Claude models following the intended HealthBench Professional procedure, using GPT‑5.4 grading and length-adjusted, unclipped scores. For Fable 5.1, we used Opus 5 fallback for provider refusals.

  12. Claude Fable 5 and 5.1 are not included in LifeSciBench Gold v1, GeneBench Pro v13, and MedChemBench because they refuse the majority of questions in these evaluations.

  13. On ExploitGym, we tested Astra and Sol without the 6-hour time limit, to better assess their full cyber capabilities. They are fast enough that it has little impact.

  14. ExploitBench (June–August 2026) contains 20 high-severity V8 vulnerabilities across 13 stable Chrome releases. The benchmark tests whether agents can achieve arbitrary code execution in V8 and official Chrome releases for Linux by exploiting each specified vulnerability. Some included vulnerabilities may not permit arbitrary code execution under the evaluation’s constraints, so a 100% success rate may not be achievable.

  15. Jeremy Spence et al. “The Next Challenge for Agentic Cybersecurity: A Realistic, Contamination-Free Reverse Engineering Benchmark⁠(opens in a new window).” arXiv:2608.11469v1, 2026.

  16. When we test across third-party models, we use a simpler research setup. Codex has a more complex production configuration, which can result in different raw-model error rates. Provider-side safeguards and computer-tool implementations still differ. Users do not experience the no-confirmation scenario in Codex, as it's an internal research configuration.

  17. For ScreenSpot-Pro and ExploitGym, the Fable scores we report come from Mythos, which is Fable with fewer safeguards.

The Daily Front Page 3 of 29
Thursday, September 3, 2026 The Daily Front No. #260903 — The Name at Stake
article

.name Termination

by pavel_lishin·▲ 1,537 points·406 comments·neil.fraser.name ↗
On 15 April 2026 Verisign proposed the destruction of the entire 3rd level of the '.name' hierarchy.

.name Termination

Nearly twenty-five years ago I registered neil.fraser.name to provide a stable presence on the Internet. It has been the home of this website, my email address, and a server for APIs. It predates YouTube, Facebook, and smart phones. Minutes after my daughter was born, I also registered beverly.fraser.name.

On 15 April 2026 Verisign proposed the destruction of the entire 3rd level of the '.name' hierarchy in order to simplify their administration. Astonishingly, on 28 July 2026 ICANN approved this action. I found out about this a few days ago when my registrar emailed me.

Now, it's worth pausing for a moment to discuss what '3rd-level domains' are. Many people will be familiar with shady operators selling domains like *.uk.co. In that case it's some random guy who bought the 'uk' domain from the country of Colombia, then resells third-levels. If that operator disappears, then so do all the domains he sold. As a result, 3rd-level domains have gotten a dubious reputation. However, '.name' is completely different. It was setup exclusively as a 3rd-level operation. One registers xxx.yyy.name from any registrar and there is a full whois record. Exactly like *.ny.us, or *.co.uk.

One of my original reasons for choosing '.name' was that it was run by Global Name Registry -- more specifically, it was not run by Verisign. I had history with Verisign and did not trust them. Unfortunately, Verisign acquired Global Name Registry a few years later. This mistrust was validated by the numerous lies Verisign included in their above proposal to ICANN.

So what does this mean? First, this website vanishes in February. Despite the fact that it's registered and paid for until 2040. Second, my email address also disappears. Third, all the IoT devices that use services on this domain become bricks. Basically, I disappear from the Internet.

But it gets much worse. Once the 3rd-level domains are terminated, it is assumed that the now vacant 2nd-level domains will become available for registration. Should someone (other than me) scoop up fraser.name they would be able to recreate and control neil.fraser.name. They'd be able to hijack hundreds of accounts that are linked to that address. They could commit code with my authentication. They could seize control of IoT devices. There is no way to enumerate all accounts (online and offline) which have been opened using this email address over the past quarter century.

I'm just one of 22,000 people who will lose their domains. This is going to be fun. Time to lawyer up...

The Daily Front Page 4 of 29
Thursday, September 3, 2026 The Daily Front No. #260903 — Sound Desk, Rebuilt
repository

Audacity 4.0

by ClydeN·▲ 1,082 points·241 comments·github.com ↗
★ 18,074⑂ 2,631 forks C++

Audio Editor

Audacity 4 rebuilds the application interface on Qt and introduces many new quality-of-life improvements, including a new clip-editing model. Most Audacity 3 workflows remain available, but some controls have moved or changed.

Watch the video:

Editing clips

  • Clips can be selected directly. Click a clip header to select it, or Shift-click to select multiple clips.
  • Several clips can be edited together. Moving, trimming and time stretching apply to all selected clips.
  • Clips can be grouped. Groups remain together when moved, copied, pasted or duplicated.
  • Clips can be placed more freely. They can move between mono and stereo tracks. Moving a clip over another replaces the overlapped part instead of blocking the move.
  • Splitting has a dedicated tool. Press or hold S, then click the waveform or clip header. Split, split-cut, split-delete, split at silences and split to a new track are also available as commands.
  • Paste handles more cases automatically. Audacity can create a track when needed, adapt compatible channel layouts and paste audio files from the operating-system clipboard.
  • Alignment guides, sample-boundary snapping and per-project snap settings have been added or expanded.

Interface and tools

  • The interface has been rebuilt on Qt, with native high-DPI rendering.
  • Toolbars and panels can be moved, docked, floated, shown or hidden.
  • UI layouts can be saved as Workspaces. Audacity includes Modern, Classic and Music workspaces.
  • Light, dark and high-contrast themes are available, along with accent colors, track colors and several clip styles.
  • The new Home screen shows recent projects with preview thumbnails.

The separate Select, Envelope, Draw and Multi-tool modes have been removed. Their functions are now context-sensitive:

  • Volume envelopes are displayed in the Clip gain mode.
  • Sample drawing becomes available when the waveform is zoomed to individual samples.
  • Splitting is available by holding S.
  • Track and effect parameters use consistent rotary controls with fine adjustment and double-click reset.
  • Sync-Lock has been removed. Delete, cut and paste now have explicit variants for either leaving a gap or moving later material to preserve timing.

Playback and recording

  • The playhead remains visible during navigation and can be dragged to a new position.
  • Playback can seek to another position without stopping.
  • Recording can start anywhere on the timeline and creates a clip at that position.
  • Loop boundaries and the interaction between playback, selections and loops have been revised.
  • Punch and Roll, lead-in recording, latency compensation, software playthrough and per-track input monitoring have been rebuilt for the new interface.
  • Audio Setup now includes system-default devices, refreshable device lists and custom channel mapping. Audacity can follow operating-system device changes automatically.
  • Official Windows builds include ASIO playback and recording support.

Tracks, meters and effects

  • Track headers now contain live playback and recording meters.
  • Preset handling is consistent across built-in, destructive and realtime effects.
  • Built-in effects, generators and analyzers have been rebuilt for the Qt interface.
  • Supported plugin formats are VST3, Nyquist, LV2 on Linux and Audio Units on macOS. Audacity can display generated controls when a plugin's own interface is unavailable.
  • Spectrogram has been redesigned with clearer guides and rulers, and faster rendering.

Projects, import and export

  • Audacity 4 uses the new .aup4 project format.
  • .aup3 projects open and convert to .aup4 without changing the original file. Converted projects cannot be saved back to .aup3.
  • Older .aup projects can be imported.
  • Project files store preview thumbnails and Audacity 4's additional clip and appearance data.

And last but not least, we had the audacity to change the Audacity logo.

Compatibility notes

The following Audacity 3 features are not available in Audacity 4.0, but we're working on adding them in future releases.

  • Time Tracks
  • Note/MIDI tracks
  • Mixer
  • Macro Manager and the scripting pipe
  • VAMP and LADSPA plugin hosting
  • Play-at-speed

Sync-Lock and the old tool modes were replaced by the workflows described above.

Additionally, Audacity 4 ships with some missing exporting and rendering features, analyzers, and effects.

Update: audacity-sources-4.0.0.tar.xz was updated to include SoundTouch and sbsms 3rd-party libraries.

The Daily Front Page 5 of 29
Thursday, September 3, 2026 The Daily Front No. #260903 — An Unplanned Flight
article

Holden's Lightning Flight

by ColinWright·▲ 227 points·53 comments·en.wikipedia.org ↗
On 22 July 1966, Walter "Taffy" Holden ... inadvertently engaged the afterburner.

XM135 at Imperial War Museum Duxford

XM135 at Imperial War Museum Duxford

On 22 July 1966, Walter "Taffy" Holden, a 39-year-old engineer in command of No. 33 Maintenance Unit RAF with limited experience flying small single-engine trainer aircraft, inadvertently engaged the afterburner of a Mach 2.0–capable English Electric Lightning during ground testing at RAF Lyneham. Unable to disengage the afterburner, Holden ran down the runway, narrowly missing a crossing fuel bowser and a de Havilland Comet taking off, before taking off himself. Flying without a helmet or canopy, the ejection seat disabled, and the landing gear locked down, Holden aborted his first two landing attempts. He landed on his third approach, striking the runway with the aircraft's tail as he adopted the landing technique of a taildragger aircraft. The aircraft returned to service, and was subsequently acquired by the Imperial War Museum Duxford.

Aircraft

The English Electric Lightning was a high-performance short-range interceptor aircraft. The Lightning had a max takeoff weight of 20 tons, and could reach Mach 2.0.[1] The aircraft involved in the incident was the second production Lightning, designated XM135.[2] XM135 was suffering from an electrical fault that would only manifest during acceleration for takeoff;[3] the electrical inverter supplying power to flight instruments would cut out during the first yards of the takeoff, and the standby inverter would switch in.[4]

Crew

A de Havilland Canada DHC-1 Chipmunk similar to that in which Holden had some practice flights

A de Havilland Canada DHC-1 Chipmunk similar to that in which Holden had some practice flights

Wing Commander Walter "Taffy" Holden enlisted in 1943, gaining a cadetship to a university. Studying mechanical engineering, Holden also learnt to fly on the de Havilland Tiger Moth biplane. Holden pursued an engineering career with the RAF who qualified him with pilot wings after training on the Harvard. He subsequently practised on the de Havilland Canada DHC-1 Chipmunk during his early career.[4]

In 1966, Holden was in command of No. 33 Maintenance Unit RAF at RAF Lyneham who maintained Gloster Meteors, English Electric Canberras, and English Electric Lightnings. At the time, the unit was in the process of winding down and was disposing of its last aircraft. The unit had a test pilot on staff for its Canberras and Meteors, but that pilot was not qualified for the Lightning. For Lightning flights, Holden had to locate RAF test pilots with a current rating who could usually be found within 24 to 36 hours.[1][4][5]

Flight history

The troubles with aircraft XM135 were holding up the closure of the unit, and at the time of the incident, no test pilot was available for another week. A pilot from RAF Boscombe Down, who was involved in previous tests, suggested Holden perform the test himself because it involved only ground taxiing for 30 to 40 yards (27 to 37 m) at a time. For each test, Holden was to test a different electrical configuration, rev up the engine to high RPMs, then cut the engine and apply the brakes. Holden was to communicate by hand signals with his support crew on a Land Rover, which would coordinate the next test with the control tower.[1][4][5]

Holden was not wearing a helmet and had no radio. The canopy of the aircraft was removed for the electrical wires running out of the cockpit.[5] The landing gear was locked in a down position in a test mode.[1]

Lightning with afterburners engaged

Lightning with afterburners engaged

After the Lightning was in position, Holden carried out the first test as expected, moving the aeroplane 30–40 yards.[1] However, in a subsequent test, Holden unintentionally pushed the throttle past the afterburner gate. Once the afterburner was engaged, disengaging it required pushing the gate keys behind the throttle, which Holden was inexperienced in operating. The Lightning gained speed quickly and just missed a fuel tanker that was crossing the runway in front of Holden. The Lightning crossed the main runway as a mid-takeoff de Havilland Comet passed over the Lightning. Holden was then running out of runway, so he pulled the stick back and took off.[4][5]

Following takeoff, Holden managed to disengage the afterburner after feeling for the gate keys. Holden considered ejecting; however, that was not possible because the ejection seat was in inert ground mode.[1] On his first two landing attempts, his speed and height were wrong and Holden aborted both.[3] Holden, who vaguely recalled that the landing speed for the Lightning was 150 knots (280 km/h; 170 mph), took a wide circuit around Lyneham and attempted to land in the opposite direction of the runway, running away from the village.[4] In Holden's final flare, he adopted the attitude used in a taildragger aircraft of his earlier training. This resulted in a tailstrike, the rubber tail bumper of the Lightning hitting the concrete, breaking, and detaching the cable of the drogue parachute used for assisting braking. Braking hard, Holden managed to bring the Lightning to a stop about one hundred yards (91 m) before the end of the runway.[3] Holden was airborne for 12 minutes.[5]

Aftermath

XM135 at Imperial War Museum Duxford

XM135 at Imperial War Museum Duxford

The aircraft was repaired and returned to service.[5] The electrical fault was determined to be caused by wires left in place from a deleted ground test button for the standby inverter, which shorted into the UHF radio which moved on its trunnions during the takeoff run.[4] After flying for 1343 hours,[6] XM135 was acquired in 1974 by the Imperial War Museum Duxford, where it is on display.[2]

The inadvertent flight was impossible to hide from the press since the base was filled with civilian contractors. Holden was sent to Italy on leave when the news broke; however, he was recognised there as well.[1] An inquiry confirmed that Holden had not acted against any orders in the Flight Order Book (though these orders were subsequently amended) and that Holden had saved himself and the aircraft.[1][4] According to Holden, in a review before Air Marshal Kenneth Porter, he was asked whether he agreed that "With the limited flying experience I had, the test would have been better left to an experienced and current Lightning test pilot", which he answered in the affirmative, following which Porter related some of his own unfortunate flying incidents.[4] Holden remained in RAF service and retired in the late 1970s / early 1980s.[1]

Holden died in 2016, aged 90.[2]

The Daily Front Page 6 of 29
Thursday, September 3, 2026 The Daily Front No. #260903 — Data’s Quiet Revolution
article

Pre-Release of Polars 2.0

by komape·▲ 384 points·131 comments·pola.rs ↗
We don’t aim to make a big feature release of Polars 2.0. In fact we hope it to be a boring experience for you.

Today we are releasing the first release candidate for Polars 2.0. The definite 2.0 release will land in the following weeks. We don’t aim to make a big feature release of Polars 2.0. In fact we hope it to be a boring experience for you. The reason we bump this major version is that we can get rid of design decisions made in the past that currently block us and then we want to change defaults to more sensible settings that will benefit a greater audience. The biggest default change will be that all LazyFrame queries now will run on the streaming engine. Casual Polars users can therefore expect huge improvements in memory usage and performance. In aggregate we expect the streaming engine to be easily 5x faster.

To help users transition to 2.0, we have posted a full migration guide. This post will cover a few of the highlights.

Streaming engine as default

This is the biggest impact change of 2.0. Calling collect on a LazyFrame will now default to the streaming engine, leading to massive memory and performance improvements on most queries for users. The reason this required a major version bump is that the streaming engine doesn’t guarantee row-order by default for certain operations (join, group_by, unpivot, etc.). If you require observable row-order in those operations, you can opt in to that by setting maintain_order=True.

For users who want to keep using the “in-memory” engine as default, they can do so by setting the engine affinity.

lf = pl.LazyFrame({"k": [2, 1, 0], "v": ["a", "b", "c"]})
other = pl.LazyFrame({"k": [0, 1, 2], "r": ["x", "y", "z"]})

# 2.0: engine="auto" now resolves to the streaming engine.
# Row order is no longer guaranteed for joins, group_by, unpivot, ...
(
    lf
    .join(other, on="k", how="left")
    .collect()
)
# ┌─────┬─────┬─────┐
# │ k   ┆ v   ┆ r   │   <- order may not match `lf`'s original row order
# └─────┴─────┴─────┘

# Opt in to observable order for this query:
(
    lf
    .join(other, on="k", how="left", maintain_order="left")
    .collect()
)

# Or keep the old in-memory engine as the default, process-wide:
pl.Config.set_engine_affinity("in-memory")

# ...or per query:
(
    lf
    .join(other, on="k", how="left")
    .collect(engine="in-memory")
)

Stricter Polars

Polars aims to be strict and fail fast. Errors should ideally raise up-front, not 20 minutes into a pipeline. Implicit behavior on data-mismatches should be opt-in, not a default, since those mismatches can hide bugs. This strictness has become even more valuable with the rise of AI-driven development. Agents can validate a query’s structure early by calling collect_schema(), which resolves types and catches schema-level mismatches without materializing any data. This ensures fast feedback for the agents, meaning they can iterate faster. Not all errors can be caught during compilation of the query plan, some depend on data. In these cases Polars defaults to stricter behavior to ensure inconsistencies are caught instead of silently producing different results.

Below are a few examples where Polars has gotten more strict:

is_in lossless type-coercion

If you run an is_in expression on different data-types, Polars used to cast both types to their common supertype, even if that conversion was lossy Below is an example with user-ids that can go wrong by silent data-type mismatches.

# Checking if a user ID matches a list of "flagged" account IDs
# (flagged_ids loaded from a JSON export, where large IDs became floats)
flagged_ids = pl.Series([9007199254740992.0])
user_id = pl.Series([9007199254740993])  # Int64 -> a different ID, off by 1
user_id.is_in(flagged_ids)

Before 2.0, user_id gets coerced to Float64 to match flagged_ids. But 9007199254740993 sits above 2^53 (9007199254740992), the largest integer float64 can represent exactly, so it silently rounds down to 9007199254740992.0, giving a false positive.

In 2.0 this raises: InvalidOperationError: 'is_in' cannot check for Int64 values in List(Float64) data., users should explicitly cast to deal with lossy type conversion.

Strict concatenation

Horizontal concat will now check lengths instead of silently filling with null.

# Joining per-day transaction counts with per-day fraud-flag counts,
transactions = pl.DataFrame({"day": [1, 2, 3, 4, 5], "count": [120, 98, 143, 87, 156]})

# Upstream job for day 5 failed silently
fraud_flags = pl.DataFrame({"flagged": [2, 0, 5, 1]})  # only 4 rows

pl.concat([transactions, fraud_flags], how="horizontal")
shape: (5, 2)
┌─────┬───────┬─────────┐
│ day ┆ count ┆ flagged │
│ 1   ┆ 120   ┆ 2       │
│ 2   ┆ 98    ┆ 0       │
│ 3   ┆ 143   ┆ 5       │
│ 4   ┆ 87    ┆ 1       │
│ 5   ┆ 156   ┆ null    │  <- day 5 silently has no flag count
└─────┴───────┴─────────┘

In 2.0 this will raise with:

ShapeError: cannot concat dataframes with different heights in 'strict' mode

If padding is what you wanted, you have to explicitly opt-in to that with how="horizontal_extend". Making that intention clear to the reader.

Removal of casts in favor of dedicated methods/constructors

Another one worth mentioning is the removal of many casts that were ambiguous or should be applied via their dedicated parsing expression, leading to one obvious way to parse data.

Enums/Categoricals <> integers.

pl.Series([None, 1, 0, 2], dtype=pl.UInt32).cast(pl.Enum(["a", "b", "c"]))
# ComputeError: casting from u32 to enum is not supported.

Use instead: .cat.to(dtype) for int → categorical, .cat.physical() for categorical → int.

Parsing Strings to temporal data-types

pl.Series(["2022-08-30"]).cast(pl.Date)
# InvalidOperationError: casting from string to date is not supported.

Use instead: .str.to_date() / .str.to_datetime(). These allow you to apply a parsing format, giving you more control over how the data is parsed.

These were just a few examples, but we landed many more strictness improvements. See them all in the migration guide.

Raising informative errors

We put a lot of effort into making sure you as user or your agent can continue if you used old parameters that are not supported anymore. We added two new typed exceptions for this; polars.exceptions.AttributeRemovedError and polars.exceptions.ArgumentRemovedError that handle removed attributes and methods and removed parameters respectively.

The error messages should point you to the new API instead. Below we show two examples.

>>> lf.melt(id_vars="a", value_vars="b")
polars.exceptions.AttributeRemovedError: `melt` was removed in version 2.0;
use `LazyFrame.unpivot` instead, with `index` instead of `id_vars`
and `on` instead of `value_vars`
>>> df.join(df, on="a", join_nulls=True)
polars.exceptions.ArgumentRemovedError: the argument 'join_nulls' for
'DataFrame.join' was deprecated in version 1.24 and has been removed
in 2.0.0. It was renamed to 'nulls_equal' in version 2.0.

Most of the removed functionality has been deprecated for a long time and hopefully should not have affected your pipelines if you have stayed up to date. Reach out to us if you think we should have kept some functionality you relied on.

Last words

Polars 2.0 is about better defaults (most importantly the streaming engine) and a better API. We hope this release is rather boring. We don’t gate new features behind major version bumps as we ship them as soon as their ready.

Don’t be mistaken, Polars 2.x will be much better than 1.x. There is a lot in flight that we haven’t talked publicly enough: proper out-of-core support for the streaming engine, a new IO-plugin design, what we think will be the fastest S3 reader out there, major SQL coverage improvements, a cost-based planner, join reordering, and the removal of mmap, which will make our pipelines fully async end to end.

Try the release candidate by installing pip install polars==2.0rc1. Give it a spin and reach out to us here: https://github.com/pola-rs/polars/issues or contact us on discord: https://discord.gg/4UfP5cfBE7.

The Daily Front Page 7 of 29
Thursday, September 3, 2026 The Daily Front No. #260903 — The Cost of a Busy Browser
article

The browser's main thread is expensive

by kciter·▲ 384 points·135 comments·kciter.so ↗
On screens with a lot of interaction, where data streams in live and scrolling, animation, and input all get tangled together, the picture changes.

The Browser's Main Thread Is Expensive

What comes to mind when you hear “frontend optimization”? For most of us it’s things like reducing network requests, shrinking the bundle, or making good use of the cache. Beyond that, maybe cutting down on re-renders or tuning when resources get loaded. The main thread doesn’t usually come up, and there’s a reason for that: on most screens it never becomes a problem. But on screens with a lot of interaction, where data streams in live and scrolling, animation, and input all get tangled together, the picture changes. However much you save on network and bundle size, the screen freezes the moment the main thread gets blocked.

You’ve probably come across a website where scrolling stutters now and then, a button responds slightly late, or the letters you type into a search box show up half a beat behind. It isn’t bad enough to be annoying, but it gets on your nerves in a subtle way. That kind of jank is what a blocked main thread looks like.

When we run into jank like this as developers, the usual reaction is to wonder “is my code slow?” and start picking apart algorithms or looking for wasted computation. In most cases, though, the speed of the code is not the problem. The code isn’t slow. It just happens to be the code that’s holding the main thread.

The browser has a number of threads, but almost everything we can touch from code is concentrated on the main thread. Computation, rendering, event handling, network response handling, and your framework’s internals are all processed there. One resource, a mountain of work.

The browser’s main thread is expensive. Most of the time it doesn’t cause trouble, but once you try to do something ambitious, dealing with the main thread becomes the important part. This article is about how to handle that expensive resource.

What Does the Main Thread Do?

Let’s start with what the main thread actually does. Its work falls into two broad categories.

The first is running JavaScript. The code we write, along with event handlers, timers, network response callbacks, and the framework’s internals, all run here. These tasks execute in the order they enter the queue, whenever there is a gap, with no relation to the screen refresh cycle.

The second is drawing the screen. When the DOM or styles change and the screen needs updating, the browser goes through roughly these steps, in order, to produce a frame.

  • Run requestAnimationFrame callbacks - JavaScript registered to run just before the frame is drawn
  • Style calculation - compute the final CSS values for each element
  • Layout - compute each element’s position and size (also called reflow)
  • Paint - generate paint commands describing what to draw in which colors

If nothing changed, these steps are skipped entirely, so they don’t necessarily run every frame. Only the final compositing step, which takes the produced output and assembles it on screen, is handed off to the compositor thread1. In other words, most of the front half of the pipeline that draws the screen is the main thread’s responsibility.

The rendering pipeline for updating the screen

For the screen to look smooth, frames have to be drawn at the display’s refresh rate. On the most common 60Hz display, that means 60 frames per second, or about 16.6 milliseconds per frame. And you don’t get to use all of it. Once the browser’s own processing cost is subtracted, the practical budget is usually considered to be around 10 milliseconds2, and on a 120Hz device the budget itself is cut in half.

The problem is that the two kinds of work above stand in a single line on the same thread. JavaScript was designed around a single-threaded event loop model. The main thread processes one task at a time, and while that task is running, nothing else can happen. If one JavaScript function runs for 200 milliseconds, then for those 200 milliseconds the browser can’t repaint the screen or receive a click from the user. Against a frame budget of around 10 milliseconds, that is a fatal amount of time. A task that runs this long and holds the main thread is called a long task, and anything over 50 milliseconds is generally considered a problem.

Words only go so far, so let’s feel it. In the demo below, pressing the button makes JavaScript grab the main thread for a moment.

Press a button: the JS animation and typing freeze, but the CSS animation keeps spinning

When you press the button, the JS animation stops and typing into the input field does nothing. The CSS animation, on the other hand, keeps running. We’ll come back to where that difference comes from later. What to remember for now is that holding the main thread for a long time is the same thing as freezing the screen.

This connects directly to web performance metrics. INP (Interaction to Next Paint), which measures how long it takes for the screen to respond after the user does something, and TBT (Total Blocking Time), which measures the total time the main thread was blocked during page load, are both essentially ways of expressing how long the main thread was blocked. A large part of performance optimization is a matter of how carefully you spend this one thread.

The ways of spending it carefully fall into two broad families. One is to divide the main thread’s time well from within. The other is to send the work outside the main thread altogether. Let’s take them in order.

Using the Expensive Resource Wisely

The first family is about staying on the main thread but spending its time intelligently. There are four core moves.

  • How do you split up work that runs too long?
  • How do you group work that runs too often?
  • Among several tasks, which goes first?
  • How do you postpone work that doesn’t need to happen now?

We’ll call these splitting, batching, prioritizing, and deferring. The first two shape the size of tasks, and the last two decide their timing. Of the four, splitting is the foundation for the rest. Tasks need boundaries before you can decide what to slot in between them and what to push back. So we start with splitting.

Splitting

Picture the chat pane of a live stream. On a popular stream, chat can burst to hundreds of messages per second. In that environment, messages don’t arrive politely one at a time. When traffic spikes, the server sends them in clumps of dozens, and the moment you enter a room, hundreds of backlogged messages come down at once. What happens if you render that whole clump in one go right when it arrives? Every message you draw brings DOM creation, style calculation, layout, and paint along with it, and those hundreds of iterations run back to back inside a single task. Meanwhile, the user trying to type their own message gets a stuttering input field, and every other animation on screen hitches too. Other people’s chat is monopolizing the main thread and getting in the way of yours.

The fix is what we said above. Cut the clump into small pieces, and between the pieces, hand control of the main thread back for a moment. In those gaps the browser can catch up on the screen updates and input handling it had queued.

The demo below simulates a streaming chat pane. Press “Flood the chat” and messages start pouring in. Try typing in the input field while watching the smoothness gauge and fps at the top, and compare the “Immediate render” and “Yielding render” modes.

In “Immediate render” mode, the DOM is touched as each message arrives, so while chat is flooding in, fps drops sharply, the gauge stutters, and the input field lags. If you look closely, the chat messages themselves start appearing noticeably more slowly as well, because the callback that receives and processes them is also a task waiting in the main thread’s line, so it gets delayed with everything else. Now switch to “Yielding render”. Messages are still drawn one at a time, just as before, yet input comes back to life and the screen moves again. The only thing that changed is that after every 20 messages, the main thread is released for a moment.

One thing not to misread here is that yielding does not make the work faster. The total amount of work is unchanged, and the few milliseconds spent waiting at each yield are added overhead, so in wall-clock terms it actually takes longer. So why did rendering recover along with input?

As we saw earlier, the main thread can do nothing while a task is running. The rendering pipeline that produces frames can’t cut into the middle of a task either. It can only run between tasks. Yielding is the act of creating those gaps. The backlogged input and frame production get their turn in the gaps, and to the user it feels as though performance improved.

At the code level, the classic way to yield is setTimeout, which pushes the continuation into the next task. Take a look at the following code.

// A batch of chat messages arrives at once
socket.on('messages', (chats) => {
  renderChats(chats);
});

// Draw the messages, yielding the main thread after every 20
async function renderChats(chats) {
  let count = 0;
  for (const chat of chats) {
    appendChatNode(chat); // draw one message

    if (++count % 20 === 0) {
      await new Promise((resolve) => setTimeout(resolve, 0)); // yield here
    }
  }
}

With this in place, no matter how hard chat floods in, the DOM work never occupies the main thread in one piece, and between the pieces there is room for the user’s input and animations to be processed.

The star of this code is setTimeout. When it schedules the resumption of the remaining work as a new task, the current task ends right there, and in that gap the backlogged input and rendering get processed. await pauses the function until the scheduled task comes back around, then picks up where it left off.

How yielding changes the timeline

The example above split the incoming work by count. But if an animation is already running, or the user is in the middle of scrolling, splitting by time is safer than splitting by count. An animation uses a little of the main thread every frame, so a heavy job has to keep checking the clock and cut itself off before it swallows what’s left of the frame’s budget.

async function processDuringAnimation(items) {
  let i = 0;
  let frameStart = performance.now();
  while (i < items.length) {
    // Work only until 5ms have passed since the frame started
    while (i < items.length && performance.now() - frameStart < 5) {
      doWork(items[i++]);
    }
    frameStart = await new Promise(requestAnimationFrame); // resume with the next frame's start time
  }
}

Here performance.now() acts as the stopwatch that checks whether we’ve gone over budget, and requestAnimationFrame acts as the alarm that says “wake me just before the next frame is drawn.” This is also why we yield with rAF rather than setTimeout when splitting by time. The resumption lands in step with the frame cycle.

Note that rAF passes the frame’s start timestamp to its callback, and the code above uses that as the reference point for the budget. The reason is that the function doesn’t have the frame to itself. If other animation callbacks ran earlier in the same frame, our share has to shrink by however much time they used, or the frame budget is broken. Anchoring to the frame’s start time turns “use 5ms” into “use until 5ms after the frame started,” which makes the code cooperate naturally when several animations share one frame.

Why 5 milliseconds? There’s nothing special about the number. We said the practical budget is around 10 milliseconds, so handing roughly half to background work and leaving the rest for animation callbacks, style, layout, and paint is a reasonable heuristic. If your animations are heavy, shrink it.

With this approach, even while heavy work is in progress, there is room to draw the screen every frame, and the work and the animation run smoothly side by side. Try the demo below. Moving the mouse scatters 4,000 particles away from the cursor, and nearby particles also push each other apart, so deciding one particle’s direction means checking its distance to every other particle. That comes to roughly 16 million distance calculations per pass, and recomputing all of it every frame blows through the frame budget on its own. Compare the “Compute all at once” and “5ms per frame” modes.

Splitting is the most basic way to use the main thread’s time sparingly. It is what gives users the perceived performance they care about: fast responses and a smooth screen.

Finally, a few points to be careful about. First, splitting too finely backfires. Yielding and coming back has a cost of its own, so if the pieces are too small, that overhead can end up larger than the work you’re trying to do.

Second, yielding with setTimeout involves a minimum delay3, so each piece can end up waiting a few milliseconds for nothing. Usually this doesn’t matter, but in situations that demand a very high level of responsiveness, the delay can become a problem.

That’s why some code schedules the next task by posting a message through a MessageChannel instead. React’s scheduler uses this method. More recently, a standard API called scheduler.yield() has also appeared to address this problem. Its advantage is that after yielding, the original work resumes ahead of other queued tasks instead of being pushed to the back. Browser support is still uneven, though.

Third, different yielding tools come back at different times. setTimeout and scheduler.yield() resume without regard to the rendering cycle, while requestAnimationFrame resumes just before a frame is drawn, so for work that needs to keep in rhythm with screen updates, requestAnimationFrame is the better fit. If you want finer control over priorities, you can also build your own queue on MessageChannel and manage the yielding and resuming yourself.

Lastly, splitting isn’t always possible. Parsing a multi-megabyte response with JSON.parse, for example, is a single atomic synchronous call, and there is no way to stop halfway and yield. Until the parse finishes, the main thread is stuck. Heavy work that can’t be split like this is the clear limit of “using it wisely.” In that case you have to change the premise and not do the work on the main thread at all. We’ll get to that in “Not Using the Expensive Resource.”

Batching

Splitting on its own doesn’t solve every problem, though. Think back to the streaming chat example. Yielding rescued input and rendering, but it did nothing to make chat draw faster. If anything, throughput, meaning the number of messages drawn per unit of time, went down by the overhead of yielding. So what happens if chat pours in faster than the throughput? Arrivals outpace processing, the backlog keeps growing, and the messages reaching the screen get older and older. This situation is called backpressure.

Raising throughput takes a different tool than splitting. For example, instead of drawing messages one by one, you can draw the accumulated batch in one go. The fixed per-message cost folds together, and the same amount of time renders more chat. Being told to split and then told to batch may sound like a contradiction, but the point of both is to trim tasks to an appropriate size. Splitting deals with tasks so long that rendering can’t squeeze in, and batching deals with tasks so frequent that the pipeline’s fixed cost is paid over and over.

The best batching targets are events. Scroll, resize, and input events can fire dozens or hundreds of times in a short span. If you run a heavy handler on every one of them, there’s nothing left of the main thread. So we collapse many events into one execution, either by “running once after things quiet down” or by “running at most once per interval.” These are called debounce and throttle, respectively.

The demo below is a markdown editor with a long CHANGELOG open. Building the preview means parsing the entire document (about 2,000 lines) and rebuilding its DOM from scratch, which is far too expensive to run on every keystroke. Type quickly into the left editor with “No debounce” selected. The preview is rebuilt once per character and your input falls behind. Switch to “Debounce 300ms” and the render happens just once, after you stop typing, and the typing becomes smooth.

For visual updates, you can use requestAnimationFrame. The screen only gets drawn once per frame anyway, so no matter how many update requests pile up, drawing once per frame is enough.

let scheduled = false;

socket.on('tick', (tick) => {
  chart.push(tick); // keep every data point — nothing is thrown away
  if (scheduled) return; // this frame's draw is already booked
  scheduled = true;
  requestAnimationFrame(() => {
    renderBoard(); // draw once per frame
    scheduled = false;
  });
});

The demo below updates a board of 60 tickers with over 1,000 messages per second. “Render every tick” mode redraws the whole board on every message. Calling a chart library’s update() on every message is a common mistake, and this is exactly what it looks like. Switch to “Once per frame” and every arriving data point is still reflected, but the fps comes back.

DOM writes can be batched as well. Appending a hundred nodes in one operation instead of one at a time, or toggling a single class instead of changing style properties individually, turns many changes into one and helps performance. The old technique of assembling an HTML string and assigning it to innerHTML in one shot has the same essence. You gather the writes so the rendering pipeline’s fixed cost is paid once.

This kind of optimization is a familiar pattern to frontend developers, and React’s virtual DOM is itself a device for it. However many times state changes, the changes accumulate in the virtual tree, get compared first, and only the actual differences are applied to the real DOM in one pass. Merging several state updates inside one event handler into a single re-render, or queueing analytics events and sending them in one request instead of individually, is the same idea. A fixed cost that would repeat once per item gets paid once per batch.

Prioritizing

If splitting and batching shape the size of work, prioritizing decides its order. Reacting to the button the user just pressed needs to happen quickly, while precomputing statistics for content that’s off screen can wait. Prioritizing means ordering the urgent work ahead of the work that isn’t urgent.

Order matters because on a main thread that nothing can interrupt, order is the responsiveness the user feels. To control order, you usually build a queue that work is pushed into and pulled out of. Jobs in the queue are processed FIFO, but when something urgent comes in, it gets pulled to the front of the line.

const queue = [];
const channel = new MessageChannel();

// One message = one task. Process a piece, then book the next one
channel.port1.onmessage = () => {
  const job = queue.shift(); // take whatever is at the front right now
  if (!job) return; // guard against duplicate bookings
  job();
  if (queue.length > 0) channel.port2.postMessage(null);
};

function postJob(job, urgent = false) {
  if (urgent) queue.unshift(job); // urgent jobs cut to the front
  else queue.push(job);
  if (queue.length === 1) channel.port2.postMessage(null);
}

This structure is useful because priority isn’t a fixed value. Work that wasn’t urgent to begin with can suddenly become urgent because of something the user does. Say the user attaches a few dozen photos to a post. To save on costs, the client sometimes resizes images before uploading them to the server, and that resizing is unhurried work that can be processed in order.

React works along similar lines, with more sophisticated machinery on top, such as starvation protection, batching, and continuations. Use startTransition or useDeferredValue and a scheduler spins up inside that yields via MessageChannel and orders work with its own priority queue.

But once the user clicks a particular photo to check that it attached properly, that photo’s preview becomes the most urgent job there is. With a priority queue like the one above, the urgent job can be handled first. This approach, where you get ahead on the work while idle and then rush it when it becomes needed, is sometimes called the idle-until-urgent pattern4.

// Build previews for the attached photos, in order
files.forEach((file, i) => {
  const job = () => createPreview(file, i);
  job.photoId = i; // tag it so we can find it in the queue later
  postJob(job);
});

// Clicking a photo that isn't ready pulls its job to the front → priority bump
onClickPhoto((i) => {
  const idx = queue.findIndex((job) => job.photoId === i);
  if (idx > 0) queue.unshift(queue.splice(idx, 1)[0]);
});

Feel the difference in the demo below. Sixty photos have been attached, and each preview is actually generated with per-pixel filtering. While the previews are being built in order, click one of the gray tiles that isn’t ready yet. In “In order” mode you have to wait until that tile’s turn comes, but in “Clicked first” mode it skips the queue and fills in right away. The total amount of work is the same and only the order changed, yet the experience for the user is completely different.

Priority, then, is a matter of working out what matters most to the user at this particular moment.

Modern browsers offer standards for this, such as the Scheduler API and TaskController. Support is still incomplete, so in practice people pair them with polyfills or build their own queues. This article uses the hand-built-queue approach.

Deferring

The last and most reliable way to conserve the main thread is to not do now what doesn’t need to be done now. Where splitting and batching ask “at what size” and prioritizing asks “in what order,” deferring asks whether this work really has to happen right now at all.

Initial page load is the classic place where deferring pays off. There’s no need to download and execute all of your JavaScript up front. With code splitting, only the code the current screen needs runs first, and the rest is loaded when it becomes necessary, which keeps the main thread from slowing down right from the start.

You can defer rendering itself, too. Think of a social feed. Some apps freeze for a moment when you come back from the notifications tab after scrolling far enough to accumulate hundreds of posts. Even if the feed’s DOM is kept alive while switching tabs, when it becomes visible again the browser recomputes style and layout for all several hundred posts at once, including the ones nowhere near the viewport. So what if off-screen posts were left as empty shells that only take up their height, and got filled with real content as they approach the screen? The tool that tells you about that “approaching” moment is IntersectionObserver. It’s the same method image lazy-loading libraries use.

const io = new IntersectionObserver(
  (entries) => {
    for (const entry of entries) {
      if (entry.isIntersecting) fill(entry.target); // fill as it approaches
      else empty(entry.target); // empty it when it leaves, keeping its place
    }
  },
  { rootMargin: '400px' } // headroom to fill before the scroll arrives
);

feed.querySelectorAll('.feed-item').forEach((el) => io.observe(el));

The demo below is a feed with 1,500 posts piled up5. It starts with “Render only when visible” turned on. Visit the notifications tab and come back, and the return is instant regardless of how much has accumulated. Now switch to “Render everything” and make the round trip again. Every return freezes for hundreds of milliseconds while all 1,500 posts are laid out again. Apart from unfilled slots showing briefly during fast scrolling, the two modes look the same.

Render only when visible: IntersectionObserver fills only the posts near the viewport — returning is fast no matter how much has piled up

Rendering isn’t the only thing this approach defers. Off-screen posts never get their DOM built at all, so the cost of creating and maintaining it is deferred along with everything else. If a widget carries heavy initialization, that initialization can also wait until the widget nears the screen. There’s also a CSS property, content-visibility: auto, that aims for a similar effect in a single line. As of this writing, though, implementations vary between engines, and Safari has a performance bug that makes returning to the page slower rather than faster, so for now IntersectionObserver is the option that behaves predictably everywhere.

Continuously running work, like carousels, animated promo banners, and live charts, is pure waste while off screen. You’re spending main-thread time every frame to redraw a picture nobody can see. Run it while it’s visible and stop it when it leaves. That’s also how the dozen-plus demos in this article manage to coexist on a single page. Each one is built to stop once it scrolls out of view.

Not Using the Expensive Resource

Everything so far was about using the main thread, but using it carefully. The second family of techniques is about not doing the work on the main thread in the first place.

Let’s return to the question left hanging in the long-task demo. The main thread was completely blocked, so why did the CSS animation keep running as if nothing had happened? The answer is that the animation was never running on the main thread to begin with. Inside the browser, several threads divide up the work. These are the notable ones.

  • Main thread: runs JavaScript, manipulates the DOM, calculates styles, performs layout, handles events.
  • Compositor thread: composites already-drawn layers onto the screen. Handles scrolling and certain animations.
  • Raster threads: turn paint commands into actual pixels.
  • Worker threads: separate JavaScript execution spaces that we create explicitly.

Unfortunately, we can’t control these threads however we like. The compositor and raster threads are territory the browser manages on its own, so we can’t give them direct orders, and worker threads, which we can create, come with the major restriction of having no DOM access.

So “not using” the main thread doesn’t mean sending arbitrary work elsewhere. It means picking out the work that can take a form other threads can handle, and sending that. There are two main ways to do it.

Moving Work to the Compositor

The compositor thread is the reason the CSS animation didn’t stop. transform and opacity don’t change an element’s position, size, or color in the document. They move an already-painted layer or adjust its transparency, so there’s no need to redo layout or paint. That lets the browser handle them directly on the compositor thread without going through the main thread at all. Even when the main thread is busy, the compositor runs separately, so the animation stays smooth.

If you move an element with properties like top, left, width, or height instead, layout has to be recomputed every frame, and that is main-thread work. In the demo below, the two boxes slide side to side in the same way, but one moves with transform and the other with left. Press the button to put load on the main thread.

Once the main thread gets busy, only the bottom box, the one moving with left, starts to stutter. The transform box at the top is being driven by the compositor and stays smooth whatever the load. The same “slide sideways” ends up on a completely different thread depending on which property you animate. That’s why it’s better for performance to build animations that move things with transform: translate rather than left, and animations that resize things with transform: scale rather than width.

But what about animations where the layout genuinely has to change? Picture a list where deleting an item makes the items below it slide smoothly up into place. This isn’t decorative motion. The positions really do change. Yet animating top means layout on every frame. The technique that resolves this dilemma is FLIP (First, Last, Invert, Play)6. In short, you cause exactly one layout change and leave the entire movement to transform. It goes in this order.

  • First: measure the position before the move
  • Last: actually change the layout and measure the new position. Layout happens exactly once, here
  • Invert: apply a transform to the element in its new position so it appears to still be in the old one
  • Play: animate that transform away. This part belongs to the compositor
const first = el.getBoundingClientRect(); // First: where it is now

list.prepend(el); // the one and only layout change

const last = el.getBoundingClientRect(); // Last: where it ended up
const dx = first.left - last.left;
const dy = first.top - last.top;

// Invert: make it look like it's back at the old position → Play: release it
el.animate([{ transform: `translate(${dx}px, ${dy}px)` }, { transform: 'none' }], {
  duration: 300,
  easing: 'ease-in-out',
});

To the user’s eye, the element glides from its old spot to its new one, but in reality the element has already arrived at its new spot, and the transform briefly drags it back before releasing it into place. While the animation plays, the only per-frame work is the compositor interpolating a transform. Most list-reordering animations are built this way, and Vue’s TransitionGroup and Framer Motion’s layout animations are FLIP under the hood.

The demo below shows the difference at a glance. Press “Play rank shuffle” and both ranking lists reshuffle in the same way. The left list animates top with a transition, and the right list moves only transform, via FLIP. With no load, both look smooth. Now turn on “Load the main thread” and play it again. The left list stutters its way to the finish, while the right one stays smooth even under load.

The same rank shuffle plays in both modes at once

Finally, two things worth knowing before handing work to the compositor. One is will-change: transform. It gives the browser a hint that “this element is about to change, so prepare it as its own layer in advance,” which can smooth out the start of an animation. Overuse it, though, and the number of layers balloons and memory gets wasted instead.

The other is reading layout values, which you just saw in the FLIP code. If you get the ordering wrong between code that reads layout values (getBoundingClientRect, offsetWidth) and code that writes styles, you run into a problem called layout thrashing.

// 🔴 Reads and writes interleaved — forces a layout recalculation every iteration
for (const el of elements) {
  const width = el.offsetWidth; // read (needs layout)
  el.style.width = width + 10 + 'px'; // write (invalidates layout)
}

// 🟢 Finish all the reads, then do the writes together
const widths = elements.map((el) => el.offsetWidth); // gather reads
elements.forEach((el, i) => {
  el.style.width = widths[i] + 10 + 'px'; // gather writes
});

If you read a layout value right after changing one, the browser has no choice but to recompute layout on the spot to give you an up-to-date answer. When that happens inside a loop, layout runs dozens of times in a single frame and the main thread slows down. Simply getting into the habit of grouping reads with reads and writes with writes is enough to avoid it.

Sending Work to a Worker

Then what about heavy work that can’t be re-expressed with transform? What do you do with things like parsing a large payload, processing images, or running complex computation? Pure calculation like that can be sent outside the main thread in its entirety with a web worker.

A worker runs JavaScript on a separate thread, fully separated from the main one. Hand the heavy computation to a worker, and in the meantime the main thread can concentrate solely on keeping the UI responsive.

// Main thread
const worker = new Worker('parser.js');
worker.postMessage(hugeRawData);
worker.onmessage = (e) => {
  render(e.data); // receive only the result and put it on screen
};

It isn’t free, of course. As we saw above, workers can’t access the DOM, so they can’t touch the screen directly. They can only compute and then send the results back to the main thread. And the main thread and the worker communicate only through postMessage, which copies (serializes) the data, so when the data being passed around is large, that cost is considerable.

So workers aren’t a cure-all. They shine when the computation is heavy enough to outweigh the communication cost and has nothing to do with the DOM. If you send a short, light job to a worker, the communication cost ends up larger than the computation cost and you come out behind. The key is to keep asking, every time, whether this work really needs to run on the main thread.

Since words only go so far, let’s bring in some genuinely heavy image processing. Seam carving is an algorithm that finds the vertical path of lowest energy (least color change) through a photo and removes it one path at a time, narrowing the image while preserving the important subject7. Removing a single seam means sweeping through hundreds of thousands of pixels, so removing 250 or so seams adds up to hundreds of millions of operations. Run that on the main thread and the whole page will freeze. Send it to a worker, though, and the screen can stay responsive the entire time the computation is running.

Press “Run on main thread” in the demo below. For the one or two seconds the computation runs, the entire page freezes, and the result appears all at once only after it’s finished. Because there are no task boundaries for paint to slip into, you couldn’t show the intermediate steps even if you wanted to. Now switch to “Run in worker”. While the same computation runs, you get to watch the image narrow in real time.

There is one more thing behind those smooth intermediate frames. If the worker copied a multi-megabyte pixel buffer every time it sent a frame, that cost would add up too. So postMessage offers an alternative to copying the data: transferring ownership of it outright. Transferable objects like ArrayBuffer move by reference only, so the cost is close to zero regardless of size. The side that hands the buffer over can no longer use it, and in exchange the copy cost disappears.

// Hand the pixel buffer to the worker without copying
// After the transfer, this side can no longer use it
worker.postMessage({ buf: pixels.buffer, width, height }, [pixels.buffer]);

Eliminating the Work Itself

So far we’ve been asking whether a piece of work really needs to run on the main thread. This time, let’s ask whether the work needs to happen at all. The best thing for performance is not doing the work in the first place.

Backpressure came up earlier. Batch as well as you like, and once the inflow exceeds the maximum throughput the backlog still grows without limit. Unfortunately, the browser has no good way to tell the server to slow down. At some point you have to give up on the idea of doing everything you’re given. There are generally three ways to eliminate work.

The first is dropping. For data that just flows past, like live logs, once processing starts falling behind you can quietly discard the oldest entries and users won’t notice. Keeping up with the present matters more than showing everything.

The second is merging. For data where only the latest value means anything, like rankings, you can merge the backlogged updates and apply only the final value. With merging, the amount of work is pinned to what the screen can digest, no matter how fast the inflow gets.

The third, skipping, targets repeated work rather than incoming work. If a computation gives the same result for the same input, there’s no reason to do it a second time. Remembering results and reusing them is called memoization.

This idea has been hiding throughout the article. Debounce skipped executions during typing, and the feed demo skipped rendering posts that weren’t visible. Looked at from this angle, half of this article was about eliminating work.

When you study optimization, your attention tends to go to ways of doing work well, but the biggest gains usually come from removing work. Before making some task faster, think about it first. Does this work have to happen, now, here, at all?

Closing

Some developers think of frontend work as the easy kind. But the browser is a far more complex system than we tend to assume. Drawing screens with HTML, CSS, and JavaScript is not the whole story.

Apps that deal with a flood of real-time data, like streaming platforms, or whose screens never stop changing, like image editors, maps, and games, start to feel slow whenever the main thread gets busy. Not everyone is on the latest hardware, so for services like these, optimization is essential. And solving these problems takes more than optimizing code. It takes understanding how the browser works, spending the main thread’s time sparingly, and not doing work that doesn’t need doing at all.

In the end, nothing is easy once you dig deep enough. So much of development is trade-offs, and you have to choose according to the situation, which ultimately comes down to the developer’s experience and judgment. Neither is built quickly, but both can certainly be built through study and experiment. I hope this article helps a little along the way.

  1. The compositor thread is responsible for compositing already-drawn layers onto the screen. We’ll come back to it later in the article.
  2. https://web.dev/articles/rendering-performance
  3. The HTML spec mandates a minimum 4-millisecond delay once setTimeout calls nest more than 5 levels deep. Yielding repeatedly inside a loop, as in the code above, trips this condition almost immediately, so even with the delay set to 0, each piece waits at least 4 milliseconds.
  4. https://philipwalton.com/articles/idle-until-urgent/
  5. Granted, 1,500 posts rarely pile up on one page in practice. The demo stacks that many on purpose, to make the effect of deferred rendering dramatic.
  6. https://aerotwist.com/blog/flip-your-animations/
  7. https://en.wikipedia.org/wiki/Seam_carving
The Daily Front Page 8 of 29
Thursday, September 3, 2026 The Daily Front No. #260903 — The Open Model Fleet
article

K2 Horizon: A connected fleet of six open models

by karimf·▲ 276 points·88 comments·ifm.ai ↗
K2 Horizon is also our most comprehensive open release to date.

K2 Horizon mountain artwork at sunset

Today IFM is releasing K2 Horizon, a connected fleet of six models: 375B-A23B, 36B-A4B, 32B, 7B, 3.7B, and 0.9B. Across reasoning, mathematics, coding, agentic tasks, and general capabilities, K2 Horizon delivers top-tier performance in every size class—with the 0.9B, 3.7B, and 7B models setting new state of the art at their respective scales.

K2 Horizon is also our most comprehensive open release to date. For every model, we are opening the training lifecycle from pretraining through reasoning and agentic post-training. We are releasing intermediate checkpoints, training data or detailed data-construction recipes, open architecture, mixture compositions, training code, configurations, fine-grained logs, evaluation results, and final weights.

The models and code are released under the Apache 2.0 license. Datasets are released under their applicable licenses, such as ODC-BY; We disclose how the data was constructed and mixed when redistribution is not possible.

Together, K2 Horizon represents the most comprehensive open model release to date:

  • A new performance frontier across scales. The 0.9B, 3.7B, and 7B models achieve world-leading performance in their size classes across widely used evaluations. The 36B-A4B model, equipped with our new Mixture-of-Value-Attention (MoVA) mechanism, delivers exceptional capability per active parameter, outperforming some much larger models. The 32B and 375B-A23B models rank among the top models in their respective classes. Together, the six models provide competitive performance across deployment environments ranging from edge devices to the enterprise.
  • The first fully open model fleet for agents. K2 Horizon is the first open model family to expose the complete development process through agentic post-training. By releasing checkpoints, data (or data recipe), code, configurations, and training logs across every stage, K2 Horizon makes it possible to study how reasoning, tool use, planning, and agentic capabilities emerge; reproduce the methods that create them; and adapt those methods to new tools, environments, and domains.
  • Six models spanning edge to enterprise. The 0.9B model is designed for highly constrained environments such as watches and glasses, while the 3.7B and 7B models bring advanced capabilities to phones and other on-device applications. The dense 32B model and sparse 36B-A4B model provide powerful options for local workstations and efficient serving. The 375B-A23B model brings the fleet’s strongest capabilities to demanding enterprise deployments. All six models include quantization support.
  • One connected fleet. The six models share core architecture, vocabulary, training methodology, interfaces, evaluation infrastructure, and deployment tooling, with a smaller vocabulary for the 0.9B model. This consistency also makes it easier to move between sizes, route work dynamically, and study capability and efficiency across scale.

World-leading performance across the scales

IFM/K2-Horizon

Benchmark comparisons for all six K2 Horizon model sizes

SWE-Atlas-QnA and SWE Bench Pro: strict = no internet. BrowseComp: different models use different harness, we use the Discard-all@95k context length proposed in DeepSeek-V3.2 technical report. WildClawBench: we use a subset of the English text-only-modality tasks. Apex-Agents: we use a subset of text-only-modality tasks. GDPVal-AA is the Elo rating.

The 0.9B, 3.7B, and 7B models achieve state-of-the-art results in their respective classes across mathematics, reasoning, general capability, coding, and agentic tasks.

The 36B-A4B model performs beyond the level normally expected from its active parameter count, demonstrating the efficiency of our unique Mixture-of-Expert design when computing attention values. The 32B and 375B-A23B models place among the top models in their respective comparison classes.

The small models are especially notable. K2 Horizon 0.9B achieves an AIME 2026 score above 48, along with strong reasoning, tool-use, and agentic capabilities. K2 Horizon 3.7B and 7B extend these capabilities to more demanding software-engineering and multi-step environments, demonstrated on strong performance in SWE-bench and BrowseComp. Although complex tasks that require extensive exploration and repeated recovery, such as those in TerminalBench, remain difficult for the smallest models, K2 Horizon moves the boundary of what is possible at every scale.

Why the Horizon Fleet matters

A transparent model that falls far behind the capability frontier has limited value as a foundation, even for research. At the same time, a powerful model released only as final weights allows people to run it, but provides little insight into how its capabilities were created.

K2 Horizon brings these two together. The fleet provides highly competitive models and releases the recipes used to train them. Researchers can study advanced capabilities in models strong enough to exhibit them, while developers can reproduce, adapt, and extend the methods rather than treating the final checkpoint as an opaque starting point.

Since introducing the fully open principle in our 2023 LLM360 paper, we have released open models every year while extending that commitment to larger scales, stronger capabilities, and now the complete lifecycle through agentic post-training.

For every Horizon model, we will release:

  • Training data or recipe, such as construction methods, and mixture compositions
  • Training code
  • Model configurations and training recipes
  • Intermediate checkpoints throughout training
  • Fine-grained training logs
  • Evaluation results across general and specialized capabilities
  • Final model weights

A final checkpoint shows what a model can do. Horizon’s complete training record helps reveal how it learned to do it.

A Deep Dive into The K2 Horizon Fleet

K2 Horizon 375B-A23B: the enterprise powerhouse

K2 Horizon 375B-A23B is the fleet’s largest and most capable model. Its sparse MoE architecture provides 375 billion parameters of total capacity while activating approximately 23 billion parameters for each token, allowing it to draw on the capacity of a much larger model without using every parameter for every token.

The model ranks among the top models below 400 billion parameters across general, reasoning, coding, and agentic evaluations. It is designed for demanding workloads where model quality matters most, including complex reasoning, software engineering, research, and long-horizon agentic tasks.

Like every model in the Horizon fleet, 375B-A23B is released not as a single endpoint but as a development tree. Its intermediate checkpoints and post-training branches expose how the base model develops into reasoning, instruction-following, and specialized agentic variants.

IFM/K2-Horizon-375B-A23B

K2 Horizon 375B-A23B benchmark comparisons

Benchmark comparison for IFM/K2-Horizon-375B-A23B.

K2 Horizon 32B and 36B-A4B: strong performance for local deployment

Horizon 32B is the fleet’s most powerful dense model, providing a strong balance of capability, adaptability, and local deployability. It ranks among the top dense models below 40 billion parameters.

Horizon 36B-A4B reaches nearly the performance of the dense 32B model while activating only approximately 4 billion parameters per token. Its efficiency comes from MoVA, our new sparse attention architecture, together with MoE feed-forward layers.

These two models serve as an important reference point for studying how dense and sparse architectures behave under similar training conditions.

These models occupy the fleet’s local performance sweet spot. They are powerful enough for demanding reasoning, coding, and agentic applications while remaining practical for local workstations and efficient serving systems.

IFM/K2-Horizon-MoVA-36B-A4B

K2 Horizon 36B-A4B benchmark comparisons

GPT 5.6 luna at medium reasoning effort; Muse Glimmer-30B at high.

K2 Horizon 7B, 3.7B, and 0.9B: frontier capability at small scale

k2 Horizon 7B and 3.7B deliver strong reasoning, mathematics, coding, tool-use, and agentic performance while remaining suitable for local and on-device deployment. On several evaluations, their results approach or exceed those of models many times larger from the previous generation.

K2 Horizon 0.9B carries many of the same capabilities into highly constrained environments. It can perform mathematical reasoning, use tools, and complete simple agentic tasks while remaining compact enough for applications on watches, glasses, and other edge devices under quantization.

The appropriate task changes with scale: the 0.9B model is best suited to focused interactions and lightweight tool use, while the 3.7B and 7B models can handle more demanding coding and multi-step workflows. Together, they demonstrate how much capability can now be retained in models small enough to run almost anywhere.

K2-Horizon 0.9B / 3.7B / 7B

K2 Horizon 7B, 3.7B, and 0.9B benchmark comparisons

Designing K2 Horizon

One family from the beginning

Horizon was designed as a connected family rather than a collection of unrelated models. The six models share core architectural decisions, training methodology, interfaces, evaluation infrastructure, and deployment tooling. This allows developers to move between sizes more easily and gives researchers a more controlled setting for studying capability across scale.

Each model is pretrained on approximately 20 trillion tokens using carefully constructed and documented mixtures. Intermediate checkpoints and their corresponding fine-grained logs are captured throughout training, creating a detailed record of how each model develops.

MoVA: scaling attention with sparse experts

The core idea behind mixture-of-experts is to increase total model capacity while keeping the computation required for each token roughly fixed. Conventional MoE architectures apply this sparsity primarily to feed-forward layers: many specialized experts are available, but a router activates only a small subset for each token.

Our new architecture, MoVA—Mixture-of-Value Attention, extends this principle to attention. Because attention determines how a transformer brings together information from across its context, introducing sparsity there opens another dimension for scaling model capacity beyond the feed-forward network.

MoVA integrates expert routing into multi-head attention while remaining compatible with efficient techniques including FlashAttention, grouped-query attention, and sparse attention.

The result is K2 Horizon MoVA 36B-A4B: a model with 36 billion total parameters but approximately 4 billion active parameters per token. Under the same training conditions, it performs only slightly below the dense Horizon 32B model while requiring substantially fewer active parameters.

Training Data

Training across six scales requires data that is broad enough to support general capability and carefully constructed enough to maintain quality over approximately 20 trillion tokens.

Horizon’s pre-training mixture combines diverse web, code, mathematical, scientific, multilingual, and domain-specific sources with synthetic data generated through our own pipelines. One of our key data innovations is the incorporation of reasoning directly into pre-training: nearly 17% of the pre-training corpus consists of problem-solving trajectories with explicit reasoning. Reasoning trajectories for mathematical tasks were further rewritten into formats such as dialogues and study guides. In total, we used approximately 10 trillion synthetic tokens during pre-training.

Our synthetic data pipelines use millions of combinations of diversity knobs and context seeds, including retrieval from an internally built search engine over the pre-training web corpus. To quantify diversity at corpus scale, we developed a new compressor, Wzip, that combines a novel adaptive multi-sliding-window LZ77 algorithm with windowed Huffman coding. Wzip mitigates the rapid saturation observed with conventional gzip/zstd compression metrics as the number of documents grows, enabling more meaningful diversity measurements across large corpora. As shown below, the measured diversity of our synthetic data approaches that of high-quality natural web text and substantially exceeds that of web code.

Comparison of diversity across training data sources

We document the data composition, synthesis pipelines, and construction of the training mixture. Following our open-source principles, where licenses permit, we release the training-ready datasets directly. Where redistribution is restricted, we release source descriptions, filtering and construction methods, and mixture composition. This allows researchers to understand what each model learned from and how the data distribution evolved throughout training.

Our post-training data is introduced from the beginning of mid-training, rather than being reserved exclusively for the final stages of training. We combine synthesized long-context documents with instruction-following, reasoning, and agentic trajectories formatted with the chat template. A main pillar of our post-training data is large-scale task synthesis grounded in task taxonomies, diversity knobs, and web-search seeding, resulting in over 100 million unique tasks. In addition, we developed sampling techniques that guide solver LLMs toward correct solutions and desired behaviors when generating training trajectories.

Beyond the construction of the data itself, we also designed the chat template to improve the diversity and efficiency of tool-use training. Rather than training on a single fixed convention, we exposed the model to multiple tool presentation and tool calling formats. Tool definitions were represented using JSON, XML, and Markdown, while tool calls used JSON, XML, and typed XML. This variation encourages the model to learn the underlying semantics of tool understanding and selection rather than overfitting to a particular syntax. For inference, we ultimately selected Markdown as the default tool presentation format because it was approximately 18.5% more token-efficient than the conventional JSON formatting on our data.

Pre-training Dynamics

Every K2 Horizon checkpoint is accompanied by detailed training logs and fine-grained histories that expose the real dynamics of large-scale training: how losses evolve, where instabilities appear, how interventions affect training, and when capabilities begin to emerge.

As an example of why these records are useful, K2 Horizon 3.7B, 7B, 32B, and 36B-A4B were trained on exactly the same 22 trillion tokens. When aligned by training progress and normalized by the median loss over the final 1% of training, their raw loss trajectories approximately collapse, which you can see in the figure below—even across dense and sparse architectures and nearly an order of magnitude in total parameter count. Although this normalization aligns their endpoints by construction, it does not force the early and intermediate trajectories to coincide. Their close agreement shows that, under shared data and a common recipe, the shape of learning remained remarkably consistent across the fleet subset that used the same dataset throughout training.

Normalized pre-training loss trajectories across K2 Horizon model sizes

Together, the checkpoints and logs allow researchers to investigate why some dynamics transfer across scale and architecture while others diverge, and to connect those differences with particular training stages, data mixtures, and technical decisions.

Post-training for reasoning and agents

Most of K2 Horizon’s advanced reasoning and agentic capabilities emerge during post-training. The complete pipeline includes mid-training, supervised fine-tuning, model merging, reinforcement learning with specialized agent training. Rather than producing only one final chat model, this process creates a development tree. Different branches specialize in reasoning, coding, tool use, and agentic domains while remaining connected to common base checkpoints.

We release the artifacts across these stages so researchers can study where each capability emerges, reproduce individual branches, and adapt the same methods to new tools and environments.

Uno Diffusion: plug-and-play lossless speedup for Horizon

Inference speed has become a first-class optimization target, especially in the era of reasoning models and agents. As models produce longer chains of thought and take more actions, even small per-token delays compound into substantial latency. Yet autoregressive language models generate one token at a time, creating a fundamental bottleneck. Existing approaches offer partial solutions: speculative decoding typically requires a separately trained draft model, while discrete diffusion enables parallel generation but often sacrifices quality for speed.

Uno is designed to remove that tradeoff. It provides a lossless inference speedup, accelerating generation without degrading response quality. Uno keeps Horizon’s autoregressive parameters frozen and fully responsible for the model’s output distribution, while a lightweight set of diffusion parameters learns only how to generate more efficiently.

Through a process we call Diffusion Distillation, these compact adapters learn to generate blocks of tokens in parallel. The result combines the quality guarantees of autoregressive decoding with the speed advantages of diffusion: the model reaches the same answers, but reaches them faster.

Across our evaluations, Uno achieves a better speed–quality tradeoff than leading speculative-decoding systems and both open-weight and proprietary diffusion language models. Crucially, its gains persist across every batch size we tested, enabling lower latency for interactive agents and higher throughput for large-scale serving, with no loss in model quality and negligible additional training overhead. Uno is delivered as a simple LoRA adapter, making this lossless speedup easy to adopt: all you need is to attach the adapters.

Open infrastructure for building and extending Horizon

We are releasing the infrastructure used to build Horizon alongside the models themselves. The goal is not only to make the fleet reproducible, but to make its development stack useful as a foundation for new models and systems.

At the center of this release is xLLM, our production-tested training infrastructure. xLLM combines large-scale training performance with the flexibility required for research, allowing teams to modify architectures, data mixtures, training stages, and objectives without rebuilding the surrounding system.

We will be also releasing our full agentic post-training code base, including Reinforcement Learning, which we hope will help researchers to advance the techniques in related areas.

Researchers can use these components to inspect training behavior or reproduce a particular stage. Developers can adapt Horizon to a domain, add new tools, train agent experts, or deploy several model sizes behind a unified interface.

From Open Source to Open Science

Intermediate checkpoints turn model development into an observable scientific process. They allow researchers to study when capabilities emerge, how training choices change behavior, and when unintended strategies first appear.

Reward hacking provides one revealing example. When a capable model is placed in a realistic computer environment and told to solve a complex task, the same resourcefulness that makes it useful can also lead it to exploit shortcuts in the evaluation itself. To understand how much this affects K2 Horizon's reported numbers, we audited the released model using Artificial Analysis's reward hacking auditing procedure.

K2 Horizon reasoning trace after finding a benchmark solution on GitHub

The JACKPOT moment: our model found the benchmark's solution on GitHub and expressed "excitement" at having the answer handed to it.

TerminalBench 2.1 places models in sandboxed computer environments and evaluates whether they can complete complex technical tasks. A capable model must explore files, invoke tools, diagnose failures, and find alternative paths toward a solution. That resourcefulness is exactly what we want. But occasionally it crosses a boundary: instead of solving the intended problem, the model locates a hidden answer, exploits something the task left exposed, or manipulates the grader itself.

We ran K2 Horizon 375B-A23B on 89 TerminalBench 2.1 tasks with eight attempts each, producing 712 trials. Of these, 500 passed the task verifier (70.2% reported accuracy). We then audited every passing trial using Artificial Analysis's reward hacking auditing procedure, applying their harbor analyze tool with the reward_hacking criterion and their full rubric text verbatim, with Codex gpt-5.6-sol as the judge model.

The audit flagged 24 trials across 10 tasks. Removing them lowers the accuracy from 70.2% to 66.9%, a correction of 3.37 percentage points. The remaining 79 tasks were fully clean. For context, Artificial Analysis reports flag rates of 2.2% for claude Fable 5 and 4.1% for GPT-5.6 Luna; K2 Horizon's 3.37% falls within that range.

The model discovered several strategies:

  • Inferring it was inside a public benchmark, finding the repository on GitHub, and downloading the reference solution
  • Pulling the current source from a real project's public repository and copying the fix rather than deriving it
  • Inspecting unadvertised files, generator scripts, or exposed credentials
  • Editing the test harness or crafting output that exploited how the test checked success

We observed a related case with K2 Horizon 7B, which found and downloaded SWE-bench answers and consequently produced an inflated score of 82. The score does not represent genuine software-engineering performance, but the behavior is scientifically revealing: benchmark hacking emerged as an unintended consequence of broader planning, tool use, environment exploration, and persistence.

Because we release intermediate checkpoints alongside the final models, these behaviors can be studied rather than hidden. Researchers can determine when a strategy first appears, connect it to changes in training, and measure its effect on reported performance.

K2 Horizon is therefore more than a collection of model weights. It is an open experiment in how capabilities—and their unintended consequences—develop throughout training.

Get Started with K2 Horizon Today!

All six K2 Horizon sizes are released as open weights under Apache 2.0, with day-zero support from vLLM, SGLang, and Ollama. Further, K2 Horizon supports deployment on NVIDIA, AMD, and Cerebras hardware, providing options from local inference to large-scale serving. Get the models from our repository: https://huggingface.co/IFM, download the models, inspect intermediate checkpoints, study training data and mixtures, reproduce training stages, or build new models and agents using the Horizon code and infrastructure.

K2 Horizon provides more than six final models. It provides an open blueprint for understanding, adapting, and advancing the next generation of AI systems.

Full Results

375B-A23B

Full result table for K2 Horizon 375B-A23B

36B-A4B

Full result table for K2 Horizon 36B-A4B

32B

Full result table for K2 Horizon 32B

7B

Full result table for K2 Horizon 7B

3.7B

Full result table for K2 Horizon 3.7B

0.9B

Full result table for K2 Horizon 0.9B

The Daily Front Page 9 of 29
Thursday, September 3, 2026 The Daily Front No. #260903 — From Amiga to Godot
article

Porting my 1993 Amiga game to Godot, with an LLM reading the 68000 assembly

by rabahs·▲ 232 points·68 comments·babyloniantwins.com ↗
Pure 68000 assembly, every sprite and every scanline by hand.

In 1993, in Baghdad, I built a game called Babylonian Twins on an Amiga 500: 512KB of RAM, no hard drive, plugged into a TV. I was an engineering student in my twenties. Pure 68000 assembly, every sprite and every scanline by hand. Murtadha Salman drew the art and Mahir AlSalman composed the music. We were under sanctions. No internet, no game development resources, just one copy of the Amiga Hardware Reference Manual, which I used to program the hardware directly, and electricity a few hours a day. The constant floppy disk swapping (because of the small memory) and the 50°C summers killed my disk drive three times.

Left: 1993, on the Amiga. Right: 2026, the same gateway.

On the Amiga, “by hand” means the game doesn’t ask the operating system for anything while it runs. At startup it saves the interrupt vectors, switches the OS interrupts off and takes the whole machine:

	move.l 	#$dff000,a0			;Base for hardware registers
	lea 	save(pc),a1			;Get the system
	move.w 	#$4000,intena(A0)		;from the AMIGA

“Get the system from the AMIGA” is my comment, from 1993. From that point on the display is the game’s own copper list (the Amiga’s programmable video coprocessor), rewritten on the fly for sprites and sky colours. Tiles move by writing the blitter’s registers directly and waiting on its done flag. The joystick is read straight from the hardware port, and the fire button is one pin on a CIA chip. The OS comes back only between levels, to load the next level’s files from the disk, and then it’s switched off again.

It was the first commercial game made in Iraq, and for a long time a game very few people got to play. Commodore collapsed and sanctions scared off publishers, so the finished game sat on a shelf. An Amiga forum found it in 2008 from my brother’s YouTube uploads and hunted me down for the disks; the thread is still there.

The game has been ported once before, by hand, in 2010. The same team rebuilt it for the iPhone on an engine written from scratch, about 34,000 lines of C++, over months of nights and weekends. Apple and Google featured it, and it reached over two million downloads. That story is here.

I didn’t do this port. I asked for it, played the result every night, said what felt wrong, and made the few decisions that needed somebody who was there in 1993. The file formats and the assembly reading were the AI’s work, and so were the decisions about how to carry thirty-year-old code across, and it went faster than I could follow. This post is what I found when I sat down weeks later and read what had been done to my own game. Some of it was wrong, and I didn’t notice for weeks.

Why I tried again

I’d tried this before. About a year ago I gave an earlier model the same Amiga material and asked it to make sense of my binary level maps. It got there in the end, but it took several rounds and a lot of hints from me.

Then Claude Fable 5 shipped, and I gave it the same files.

The test was deliberate. My guess was that there is little Amiga assembly code in LLM training sets. If the model was better at working things out rather than recalling them, this is where it would show.

The July 4th weekend was coming up, so I planned three steps, each one conditional on the previous working.

Step one, the safe ask: my own 2010 engine, the 34,000 lines of C++, moved into Godot 4. This was the control.

Step two, the unfair ask: the original 72,758 lines of 68000 assembly, for a machine that had gone out of production, with no comments to speak of and nothing in common with the C++. Rebuild that in Godot too, at the Amiga’s original 50 Hz.

Step three, the greedy ask: put the second one inside the first, so buying the modern game gets you the 1993 original as a second thing you can launch.

All three worked. The level format that had taken several rounds and my corrections a year earlier came out in a single pass, with no hints from me.

How it was run

I ran it in Claude Code, so it had a terminal and my filesystem. It could edit files, run the assembler, build the game, launch the game and read what came back. When I say below that it rebuilt my 1993 binaries and checked them, it did that by running vasm and diffing the output.

Early on it added a set of command-line flags to the game so it could play without me:

--level=<name>       load a level directly
--pose=<spec>        put the twins at exact positions
--drive=<spec>       press buttons on a script, frame by frame
--probe              dump switch / gate / door / key state
--screenshot=<path>  render a frame and quit

Which turns “does the jump feel right” into something a machine can read:

drive[btw_jump:2.2] pos=(25.44, 24.04) vel=(0.00, -14.51) ground=false apex_y=22.48

It also had two headless checks it could run before showing me anything: one that compiles every script, and one that builds every level and reports failures. On the Amiga side it drove the real toolchain, vasm to assemble and FS-UAE to boot the result. What wasn’t automated: there was no image comparison on the modern port (it took screenshots, I looked at them), and nothing checked whether the game felt right.

Step one: 34,000 lines of C++ in an evening

Wednesday night, the safe ask. Timestamps, unedited:

22:23  Godot 4 project scaffold, asset sync, TMX level pipeline
22:44  both twins playable — collision, physics, camera, switching
23:19  all 38 entity types ported — full object roster live
00:35  full screen flow — menus, map, story, save, game flows
02:15  exporting to macOS, iOS and Android

Twenty-one minutes from empty project to a playable character. Every line it moved that night was a line I’d written, over months, in 2010. I went to bed confused.

Getting it to feel right took about three days after that: jump arcs and trampoline timing, and hit detection that rewards mashing, fixed in batches on July 2nd, 3rd and 4th.

I wasn’t testing alone. My thirteen-year-old son played every build with me. He’s always known I made this game, it’s a fact about his father he grew up with, but he’d never seen me working on it. The testing turned into a father-and-son thing I didn’t plan, and it’s one of my favourite parts of the whole project.

Same units, same tick

All the gameplay state lives in tile units (1.0 = one 48px tile), and the update runs at a fixed 60 Hz, because the 2010 iOS build ran at 60 Hz. That matters because the original applies drag multiplicatively, every frame:

static const float GROUND_DRAG_FACTOR = 0.85f;
this->velocity.x *= GROUND_DRAG_FACTOR;   // every tick!

Multiply by 0.85 sixty times a second and you get one amount of friction; multiply fifty times a second and you get another. Port it to a different tick rate and every acceleration curve in the game changes. Nothing crashes, it just feels wrong forever, and you won’t find it by reading the diff. At 60 Hz the constant transplants verbatim. This is also why the 1993 rebuild runs at 50 Hz and the modern one at 60: two sets of hand-tuned numbers, each only correct at its own tick. It kept both clocks. I’d have been tempted to tidy them into one.

It didn’t use CharacterBody2D

Godot ships CharacterBody2D and move_and_slide(), and every tutorial tells you to use them. The port used neither for the player. The original has its own hand-written movement code, and rebuilding that on somebody else’s physics would feel slightly wrong in ways that are miserable to track down. The player is a plain Node2D, and the 150-line collision routine came across line for line, including the fudge numbers I picked by feel fifteen years ago and the comments I wrote to my future self:

# Add 0.5 because we want the character's feet to be in the middle of the tile.
var bottom := pos.y + dim.y / 2 + 0.5 + i + fraction
if int(bottom) == int(pos.y + dim.y / 2 + 0.49):
    continue
var right := pos.x
var left := pos.x - dim.x / 4      # asymmetric probes!

Nothing tidied up the stray 0.49. There are no tests and no docs; those comments are the spec.

Step two: the 68000 assembly

By Sunday afternoon, July 5th, I handed over the thing I actually wanted to test. 72,758 lines across 26 files, written for a machine with 512 KB of memory, by me, for me, with the commenting habits of somebody who never expected another person to read it. No documentation. A 2008 transfer to modern storage had shortened every long filename, so every include pointed at names that no longer existed. One of the five level source files is cut off partway through a data table. There’s no other copy.

Before porting anything, it made the 1993 sources assemble again, using vasm on an Apple Silicon Mac, and kept going until the output was byte-identical to the binaries that shipped.

14:34  import the Amiga sources, assets, references
14:49  vasm toolchain reproduces the shipped binaries byte-identically
15:20  disk images rebuilt
15:42  the rebuilt demo boots and plays in FS-UAE

Fifteen minutes from a folder of files to the first rebuild that matched the shipped bytes. I wrote these in ASM-One, whose dialect differs from vasm’s in ways that change the bytes: ASM-One encodes cmp #4,d0 as CMPI, vasm picks a different, equally valid encoding, so telling it not to optimise is necessary and not sufficient. Rather than edit my sources it wrote a preprocessing pass that bridges five such differences, and rebuilt the broken filename mapping file by file.

The expensive one was org. With no linker and no relocation, the level source lays out the Amiga’s memory by hand, address by address:

org $6a000				; this section lives at address $6a000
Mapadd:
	incbin"btwins:binary/L1/Map1.b"	;Game Map
	org mapadd+73*1024		; skip to 73 KB past the map's start
GLBtable:
	dc.w $3333,50,20,100		; one object record begins
	dc.w SahamR-grb,26		;Routine,Length
	...
org glbtable+2*1024			; the object table gets exactly 2 KB

The level-one map uses 74,400 of those 74,752 bytes, a margin of 352, and nothing checked it except me, in 1993. (SahamR-grb attaches an object’s behaviour as a named offset; saham is Arabic for arrow.) ASM-One’s org can also move the location counter backwards, which vasm can’t. The first workaround got one case wrong: a ds.b 800 inside a rewound block, which ASM-One treats as “skip 800 bytes”, was written out as 800 bytes of zeros. Everything after that point in the file, the copper list included, sat 944 bytes away from where the shipped binary had it. The game assembled and booted, and drew the wrong thing.

Even after that, some chunks still wouldn’t match, by about 108 bytes scattered through the variable area. Those bytes explained where the shipped files came from. ASM-One assembles into memory, and the game got onto disk by saving that memory out, after the game had been run. So the shipped files are a snapshot of a game that had already been running, not clean assembler output. A fresh assembly has zeros in those variables, because nothing has set them yet; the shipped disk has whatever they held on the machine when it was saved. The code writes them before it reads them, so the zeros are harmless.

At the time I read that line, moved on, and waited for the actual game. It took me weeks to see that this was the most important thing in the project, and that nobody had asked for it. From then on, every claim about this game could be settled by comparing bytes. I wouldn’t have done it myself. I already had the binaries, and in eighteen years rebuilding them from source never seemed worth an afternoon.

The formats

For every format it went to the code that reads the bytes and worked backwards from that. The level loader is 1,652 lines of uncommented 68000, which is why I’d always reached for a hex editor instead.

The levels

A level is a grid of tiles: a long list of numbers, where each number means “put picture 47 here”, in my own private 1993 layout. This is the format the older model and I had ground through a year earlier.

Here is the full set of tiles a level is built from, 256 of them, 16×16 pixels each, for level one:

The complete 16x16 tile set for level one of the 1993 Amiga game: stone blocks, ladders, water, palm fronds, decorative brickwork

And a slice of level one, assembled from those tiles:

A horizontal slice of level one rendered from the extracted map data

The input is a list of numbers with no header and no dimensions, inside a compressed chunk. This time I didn’t explain anything. It found the drawing routine, read how the grid was walked, worked out the width and height from constants elsewhere in the file, and produced correct maps for all five levels on the first attempt.

Then it re-rendered each level from its own extracted data and compared the result, pixel by pixel, against full-level captures I had made in 2020. Where they didn’t match, it went looking for the cause and found two copper effects: the sky gradient and the water colour cycle. With those two accounted for: five full-level images, zero differing pixels. Level one alone is 600 tiles wide, 9,600 pixels.

Map cell properties

Drawing the level is only half of what a map cell does. Each cell is one 16-bit word, and the picture is the smaller part of it:

one map cell, 16 bits:

  bits 15..10   the property: what this square DOES       (6 bits)
  bit 8         which of the two tile banks to use        (1 bit)
  bits 7..0     which of the 256 tile pictures to draw    (8 bits)

The property is the level’s invisible physics. 1 is solid ground. 2 and 3 can be climbed. 10 to 13 all mean “this hurts”, four codes because knockback needs a direction. 14 kills outright. 63 is a door. None of this is written down anywhere. It was recovered because two routines read the same word and each one reveals its own half: the draw loop masks off the low byte, and the collision check does the opposite:

	move.w	(a1),d6			; the same cell
	and.w	#$fc00,d6		; keep the top 6 bits
	lsr.w	#2,d6
	lsr.w	#8,d6			; d6 = the property, 0..63
	bsr	cbCheck			; 2 or 3?  you can climb this
	bsr	Checkrmh		; 10..13?  this hurts, and from which side

Checkrmh hands the painful cases to a label called rmhEnjury, which is 1993 me spelling “injury”.

Those bits were painted in an editor. Before the game could be built I had to build the tool that builds it: MEDITOR.S, 1,254 lines of assembly, dated by its own header, in my 1993 English:

; ***********************************************************************
; *		This Program was written in four days			*
; * 			1993-2-8/7/6/5					*
; *      I made it to help me to make a map to my first serious		*
; *				Game 					*
; ***********************************************************************

Four days in February 1993. Paint tiles with the mouse, pick a property number on the panel’s CURRENT FLAG counter, stamp it onto cells with PUT FLAG, and a flag view marks every cell carrying the selected number. While writing this post, I asked the model to run the map editor and get a screenshot. It assembled the 1993 source with a modern assembler, laid the shipped Level 2 data out in memory where the editor expects it, and booted the result in an emulator.

My 1993 map editor running in 2026, editing the real Level 2, with flag 1 (solid) selected: the waterfall scene with the solid ground marked by the flag view and the CURRENT FLAG counter reading 0001

My own tool at thirty-three years old, editing the real Level 2, flag view on. CURRENT FLAG reads 0001, solid, and the ground you can stand on is marked while the decoration you walk through isn’t. The panel says 1994: the panel artwork is a separate bitmap file the editor loads, and the copy that survived is a later one than the February 1993 code.

The other name on the panel, Udai, was my partner in Mesopotamia Software, which is what we called ourselves. He was building a game of his own at the time. I wrote the editor, for both of us, but its design was worked out between us so one tool could serve both games. His game was never finished.

Object tables

Enemies aren’t in the map. The world is stored one screen at a time, 25 tiles by 20, and every screen has a small table of the objects on it. My 1993 comments explain the markers:

;	$1111=this is a Screen but it contain nothing or(End of Screen)
;	$2222=this is an object but do not draw it (dead)go to next
;	other=this is an object,draw it and go to the next

Scr0:	dc.w $3333,50,23,17		; a live object: frame, then x, y
	dc.w hiddenwallR-lrb,20		; its behaviour: a routine, as an offset
	dc.w 0
	dc.w 0
	dc.w 7
	dc.w 10				; parameters only that routine understands
	dc.w $3333,50,12,14
	dc.w GreatTR-LRb,16,GkeyT-GTT,1
	dc.w $1111			; end of this screen

An enemy is a row of words: a marker, a frame, a position inside its screen, then its behaviour. hiddenwallR-lrb is the crumbling-wall routine, attached as an offset from a base label, the same trick as the arrow thrower earlier. The words after it are parameters that mean whatever that routine wants them to mean. Nothing in the file says which word is which, so it found the routine that walks these tables every frame and let it name the fields, then converted every object in all five levels to world coordinates and checked them against the rendered maps.

GAME.S

Most of the data files scramble their 16-byte headers with a key stored inside the file, a 1993 trick to keep disk editors out. The retail loader, GAME.S, has no unscrambling step at all. It read that as a clue: GAME.S was written before the scrambling was added, so it is an older file. That clue is what later let it recover the lost two-disk retail set, from a sector map inside that same file.

The doors aren’t in the map

I was sure they were.

Load a level’s tile map and there are holes where every door should be, with no door tile in them, open or closed. An 18-byte object record stamps them onto the map at runtime, a 1×4 tile column, from a table:

closed  $528  $53C  $550  $564      ; solid, blocks the way
open    $129  $13D  $151  $165      ; passable — exactly one sheet-column right

The map data says there’s no door. The level code says there is. For thirty-three years I’d have told you the map is the source of truth and doors are map data, and I’d never have looked. It held both facts, found the routine that reconciles them, and came back with the design: doors are drawn by code at runtime; they were never painted into the map in the editor. That’s why the obvious port of the level data produces a tower with doorways full of sky.

The copper sky

In every level, colour index 31 is the sky, and nothing in the tile art ever paints it. The tile atlas renders it transparent, and behind it the copper repaints the background colour on chosen scanlines to make a vertical gradient. The gradient sits in the level source as a plain list of colours. This is the entire sky of the second level:

backgndcol:
col1:   dc.w    $09FF,$09FF,$09FF,$09FF,$09EF,$09EF,$0ADF,$0ADF
        dc.w    $0ACF,$0ACF,$0ABF,$0BBF,$0BBF,$0CBF,$0CBF,$0DCF
        dc.w    $0DCF,$0ECF,$0DCF,$0DCF,$0CCF,$0CCF,$0CDF,$0CDF

Read down the list and the sky goes from pale blue to warm near the horizon.

The 24 colour words of the level-2 sky table rendered as vertical bands, from pale cyan at the top of the screen to warm lilac near the horizon

The same 24 words, rendered. Left is the top of the screen.

The first rebuild missed it, and the levels looked fine. Flat, in a way I couldn’t name. The pixel comparison refused to go green, and the gradient went back in.

The level-2 verification diff: white marks every differing pixel — the whole sky and the animated water, missing from the first rebuild

The diff that would not go green: white is every pixel the first rebuild got wrong, the copper's sky and water.

Sprite sheet ambiguity

Amiga sprite sheets are planar (five separate 1-bit bitplanes in plane-major strips, plus a transparency mask), and all of that was worked out from the draw routines and the org arithmetic. Sheet sizes of the form frames * width * height * 2 * 5 are ambiguous: that 2 could mean double-width frames, or two stacked rows, one per facing direction. Both readings fit every byte in the file. It’s two facing rows; that was my choice in 1993.

The decoded 1993 sprite sheet: two stacked rows of the same six-frame run, one row per facing direction

Two stacked rows, one per facing direction, drawn frame by frame, not mirrored.

It flagged the ambiguity and asked.

The same twin's six-frame run cycle: the 1993 Amiga pixels above, the 2026 high-resolution art below

The same twin, the same six frames: 1993 above, 2026 below.

That was the last format. From there the 1993 game went into Godot the same way the C++ had, behaviour rewritten in GDScript at the original 50 Hz.

Step three: the old game inside the new one

The greedy ask took one evening, 21:58 to 23:43. The retro game runs as a guest, with its own namespace and scene host, and the engine switches to 50 Hz on the way in and back to 60 on the way out. It’s fiddly, and it was done in one sitting. I’d assumed the feature would eat a week and get cut. It’s the reason the Steam version ships with the 1993 game inside it as a second launch option.

The same palace doorway in both games: on the left the 1993 Amiga version in 4:3 with chunky tiles, on the right the 2026 Godot build in widescreen with the same doorway redrawn in detail

The same doorway in both games, running in the same program. Left: 1993. Right: 2026.

The game you download contains no Amiga code. The data, the packed chunks and planar graphics and the music, was decoded once, on my machine, by Python scripts, into ordinary PNG, WAV and JSON. The behaviour (how a guard patrols, when a door opens) was rewritten in the engine’s own language. If you want the real thing, that’s the free disk image at the end of this post and an emulator.

Where it was wrong

The guard bug

In level 2 you walk along a corridor. Waterfall to your left, stone pillar ahead. No enemy on the screen, nothing approaching, and you take a hit. What hit you was a spear-carrying soldier standing thirteen tiles above you, on a grass ledge next to a palm tree, with solid rock in between.

Two floors of the same column of level 2. The guard stands in a red box at tile (155,90) on a grass ledge by a palm tree; a yellow arrow runs straight down through solid rock to a corridor at y103 to 108, where the player was being hurt

He’s a doorman. He shoves whoever stands at his feet. In the original that check is fenced on both sides:

        sub.w   d1,d4           ; d4 = vertical distance to the kid
        cmp.w   #4,d4
        bpl     Sg.Far          ; 4 or more rows below? not my problem
        cmp.w   #-2,d4
        bmi     SG.far          ; too far above? also not my problem

The port kept the lower bound and dropped the upper one. A shove meant to cover the guard’s own three rows now ran the whole height of the map column beneath him, through the floor, into a corridor he doesn’t appear in.

The doorman from the 1993 sprite data: a single standing frame, 48 by 64 pixels, spear in hand

The doorman himself, from the 1993 sheet.

Smaller ones: every level has a second tile layer the original never renders; it’s the hidden artwork revealed when a door opens or a fake wall crumbles. Render it “faithfully” and every secret passage stands open from the start. Move enemies before players instead of after, and a trampoline jump gets counted twice, twenty tiles into the air. A door listed "p1,p2,p3,p4", meaning all four palms, was read as a single key with a strange name, and the tutorial exit never opened. A sound-loop length computed as stereo when the effects are mono cut every sound off halfway and restarted it.

The one that cost the most was a feature I asked for in the 1993 build: let the twins swap places at any distance. The proximity check came out, and the statues started corrupting. It went back into the routine and came out with the real answer, which wasn’t what I expected: that check was never a distance limit. The idle twin is stamped into the map itself as a statue, and two statues stamped on top of each other eat each other’s tiles. The guard went back in, and I killed the feature on the 1993 build.

I made the same kind of mistake myself in 2010, slowly, over months.

Then it shipped it

Then it did the release work: screenshots at five pixel sizes in eleven languages, a preview video, store text, icons in six shapes, uploaded to three stores that disagree about everything. I’ve shipped this game before, so I know how many evenings that part costs.

The screenshots come out of the game itself. It launches the real game at each store’s pixel size, in the language it needs, walks the character to a chosen spot, takes the shot, then draws the caption band with the game’s own fonts. AI never renders the text in a store image. The captions are real fonts and real translated strings, or the image doesn’t ship. Once it did break, and the Russian and Korean captions came out as rows of empty boxes, which I saw in the output folder before uploading.

My complaint about Steam is that it has many fields in its forms, many more than the Apple App Store and the Google Play Console. In addition, it doesn’t have an API to make the process of metadata updates easy. For iOS and Android there are proper APIs and it used them. Steam has a web dashboard, so it drove the browser: store page fields, achievements, the demo checklist, artwork uploads, clicking through Steamworks. I do the login, and I press anything that submits, publishes, prices or releases. It fills in the forms.

It reads my reviews too. The official Google Play API only gives you the last seven days, which is useless for a game with fifteen years of reviews, so it pulls the rest with the public scraper, one language at a time. Then it read all of them and listed which ones described real defects. I approved the list.

One of them was a one-star review on Google Play, bad spelling, the kind you scroll past:

“cant get through door on level one. opens but level dowsnt end”

Read literally, it’s a bug report, and it was right. In my game, opening the exit and walking through it are two separate actions, and the prompt that says so exists in all eleven languages, placed in exactly two of my eighteen levels. One of the levels missing it was the last free one. So the player deciding whether this game is worth paying for was standing in front of an open door with no way to know what to do, and concluded the game was broken.

I never found that in fifteen years, and neither did my testers or two rebuilds. It took a stranger’s one-star review. The fix went out as 2.0.3 on both stores. I keep that review.

The trampoline bug

“The trampoline feels too high.” That was the whole report, from me, playing the build at night. The constants checked out: a twenty-line simulation of the original’s integrator predicted 19.1 tiles, and the build measured 19.5. The physics was right.

It was input semantics. The 2010 build was event-driven, and because of a workaround for a tvOS quirk we shipped years ago, a held jump button read as released until you physically pressed again. Godot polls input, and kept reporting the hold. Reproducing that accident is what makes the high bounce need a fresh, well-timed press, which is how the game played on a phone, and what my hands were expecting.

The 2010 source doesn’t record this, because from the source’s point of view nothing unusual is happening. You’d have to have been there, holding the phone, working around a bug in a television.

The original, released

After thirty-three years, the full original is out, free on itch.io. Boot it in FS-UAE, WinUAE, or on real hardware. The Definitive Edition is on iOS and Android now, has a free demo on Steam, and the full Steam release (Windows, Mac, Linux) lands this fall, with the 1993 game inside it as a second launch option.

The port was Claude Fable 5 running in Claude Code; I asked, played, and decided. This post went the same way. I gave it my notes from the port, what I remember about the key parts of the old game (the map encoding, the object tables), and the repos for both the Amiga version and the port, and it wrote a first draft. I spent a week editing it line by line. The code, timestamps and screenshots are real. The part I’m least sure of is the 108 bytes: the model told me the shipped files were a memory snapshot saved after a run, and that the code writes those variables before it reads them. I read that, moved on, and have never checked it myself.

Five horizontal slices, one per level: the dungeon, the orchard, the holy gardens, the tower, and the blue lion-frieze walls

One slice from each of the five levels, rendered from the extracted map data.

The Daily Front Page 10 of 29
Thursday, September 3, 2026 The Daily Front No. #260903 — Ballots in the Mail
article

Whistleblower warns Postal Service mail ballot system has catastrophic problems

by ck2·▲ 298 points·292 comments·cbsnews.com ↗
A whistleblower is warning of "potentially catastrophic problems" with the U.S. Postal Service's new system for handling mail ballots.

Washington — A whistleblower is warning of "potentially catastrophic problems" with the U.S. Postal Service's new system for handling mail ballots. The whistleblower is accusing the agency of flouting court rulings by continuing to work on implementing a Trump executive order to tighten mail voting rules before the November midterm elections.

Described by lawyers as a federal official, the anonymous whistleblower revealed the information about the Postal Service's mail-voting procedures in a disclosure provided to Democratic Sen. Richard Blumenthal of Connecticut that was made public Tuesday. In response, Blumenthal is now questioning Postmaster General David Steiner about the Postal Service's development of a new online portal to house information about voters and their mail ballots.

"The main takeaway for me is that the Postal Service has designed a system to disenfranchise millions of Americans," Blumenthal told reporters. "One third of all Americans cast their ballots by mail, and the USPS puts all of their votes at risk."

The Postal Service said it received Blumenthal's letter and is carefully reviewing the concerns raised by the senator and the whistleblower.

"The United States Postal Service has spent months developing a U.S. Federal Ballot Mail Portal to provide election officials with a simple, secure, and efficient way to share lists of individuals receiving ballots by mail in their respective states. This work has at all times been conducted in a manner consistent with court orders," the agency said in a statement.

The Postal Service said it is finalizing the portal and will "soon" make it available to election officials to voluntarily use.

"Regardless of political party or perspective, we share a common goal: ensuring that Americans can have confidence that their election mail will be handled securely and delivered reliably," it said.

President Trump has sought to exert more federal control over U.S. elections ahead of the November midterms, including through his executive order imposing new restrictions on mail ballots. The president has also pushed Congress to approve a package of voting regulations that would outlaw mail voting — with some exceptions — and impose new proof-of-citizenship requirements to register to vote.

As Mr. Trump signed his executive order in March, he claimed that "the cheating on mail-in voting is legendary," and said new mail voting rules were needed to ensure the integrity of federal elections. But his administration has not put forth evidence of widespread voter fraud, and the president himself has cast ballots in Florida's elections by mail.

As part of its implementation of the president's directive, the Postal Service published a final rule last month laying out its new requirements for mail voting. A federal judge temporarily blocked its enforcement last week, and the Justice Department has appealed that decision.

Currently, six states conduct their elections entirely by mail. There are also four states, plus the District of Columbia, that have not yet eliminated in-person voting but also send ballots to all of their voters. In a CBS News poll in March, 77% of Democrats said anyone should be allowed to vote by mail, while just 25% of Republicans also felt that way. Independents were split, with 49% backing mail voting by anyone.

The new Postal Service rule imposes new design standards for outbound and return mail ballot envelopes, including the addition of a unique barcode. Under the new rule, state election officials must submit to an online "Federal Ballot Mail Portal" the names and addresses of voters receiving mail ballots, as well as the unique barcodes on the mail ballot envelopes.

Once that information is provided through the portal, those voters would be considered enrolled with the Postal Service on their state's "Mail-In and Absentee Participation List." The agency has said it "will not play any role in determining voter eligibility, maintaining voter rolls, or counting ballots."

Most states led by Democrats oppose the rule, arguing that the Constitution gives states the authority to set election rules, not the federal government.

Blumenthal described the Postal Service's actions as "sabotage."

In a letter to Steiner dated Monday, Blumenthal said the agency's implementation of Mr. Trump's executive order has been "perilously rushed and potentially unlawful," and he urged the Postal Service to walk back its new mail-voting plans for the upcoming elections.

He wrote that the whistleblower's allegations show that "USPS lacks the technical or operational capability needed to effectively implement the EO's provisions in a way that safeguards every citizen's right to vote in the upcoming midterm elections."

According to the whistleblower disclosure provided to Blumenthal by lawyers with Whistleblower Aid, the new portal is "untested" and will have "significant operating problems" when released, given its "rushed" development.

The federal official also warned that the new verification process could keep large numbers of mail ballots from being delivered to voters.

For instance, the Postal Service has what the whistleblower claims is a "zero percent failure rate." Mail workers would scan a sample of a batch of mail ballots — about 400 out of a batch of 10,000 or more — to verify it matches what is in the federal mail ballot portal. Under the policy, if even one barcode on a single ballot in the batch doesn't scan properly or the information does not match, the entire batch would be rejected and returned to state election officials so the problem ballot could be addressed, and then the verification process would have to be started again.

The whistleblower's lawyers said that the way the process is designed is "entirely unforgiving."

"It could delay ballots by the thousands in repeated verification cycles — and thus prevent states from mailing enormous numbers of ballots," the lawyers said.

David Becker, executive director and founder of the Center for Election Innovation and Research, warned that the process raises "tremendous" concerns for disenfranchisement of voters.

"Are Americans clamoring for the Postal Service to have this added responsibility as they're hoping that their ballots are just seamlessly sent to them?" Becker, a CBS News election law contributor, asked.

He also said that the new Postal Service rules "would undoubtedly create a tremendous number of hoops that are completely unnecessary and do nothing for election integrity."

Publication of the final rule came days before the Supreme Court intervened last week in a case brought by 23 Democratic-led states that challenged key provisions of Mr. Trump's executive order. U.S. District Judge Indira Talwani had blocked the Trump administration from implementing those portions of the directive in June, including plans for the Postal Service, but the Supreme Court lifted that lower court decision.

The high court's order was procedural: The majority found that the states' lawsuit was premature. The Supreme Court said its decision "does not mean that any measure taken by the Government to implement the Order will necessarily be lawful. On that score, time will tell."

In addition to expressing concerns about the Postal Service's new requirements for mail voting, the whistleblower's disclosure also contends that the agency violated court orders by implementing Mr. Trump's executive order.

After the district judge issued her order blocking implementation of the president's directive, Postal Service workers resumed work on the portal, the official said, and work continued even after the judge's latest ruling last week.

In a brief order Monday, Talwani, who is presiding over challenges to the new rule, said the Postal Service can establish the portal and communicate with states about design standards "so long as states' participation is not required."

The Postal Service aims to have the portal ready to be rolled out to the public by Sept. 1, according to the disclosure. North Carolina will begin sending mail ballots to some voters Friday, and several other states will make mail ballots available around mid-September, according to the Center for Election Innovation and Research.

"Our client, a lawful anonymous whistleblower, has come forward to flag a foreseeable catastrophic disruption to our coming nationwide elections based on the reckless pursuit of readying a new, insufficiently tested, poorly planned mail-in ballot screening system," the whistleblower's lawyers wrote.

The Daily Front Page 11 of 29
Thursday, September 3, 2026 The Daily Front No. #260903 — The Sky’s Inaccurate Map
article

GPS glitched across the US by as much as 33 feet

by thread_id·▲ 151 points·93 comments·sciencealert.com ↗
Scientists Have Never Seen This Before.

GPS Glitched Across The US by as Much as 33 Feet. Scientists Have Never Seen This Before. Auroras seen from the International Space Station during the November 2025 solar superstorm. (Earth Science and Remote Sensing Unit, NASA Johnson Space Center)

In November 2025, Earth was buffeted by several massive eruptions of solar material that slammed into the magnetosphere.

For many, the result was wonder. Much of the world watched in awe as a solar superstorm lit up Earth's skies with dazzling auroras to rare low latitudes.

The sword, however, was double-edged. The same storm also wrought havoc on the technology we rely on here on the ground.

And now, scientists led by space physicist Endawoke Yizengaw of The Aerospace Corporation in the US have discovered that the disruption was stranger – and more widespread – than anyone realized.

In a new analysis of data collected during the storm, the researchers found widespread, coast-to-coast disturbances in the atmosphere across the continental US – a phenomenon that has never been seen before on this scale.

That may not seem like much, but the effects of this would have been profound – throwing GPS off by more than 10 meters (33 feet) in some places. That's significant enough to disrupt precision agriculture and autonomous vehicles, the researchers say.

A Solar Storm Caused GPS Chaos Across The US. Scientists Have Never Seen This Before.

Composite image of six X-class flares that erupted in November 2025, three of which accompanied the coronal mass ejections that triggered the solar superstorm. (NASA/SDO/Scott Wiessinger)

"The results underscore the importance of accurate understanding of various space weather phenomena to enhance our predictive capabilities through coordinated observations and physics‐based modeling and ultimately reducing disruptions to RF applications during space weather events," they write in a paper published in Geophysical Research Letters.

The impact of solar outbursts on human technology is already well known. Solar flares, which unleash powerful bursts of X-rays and ultraviolet radiation, can slam into Earth's upper atmosphere, temporarily disrupting high-frequency radio communications.

Solar storms are a bigger problem. A coronal mass ejection belches out a cloud of high-speed charged electrons and protons across the Solar System; when it slams into Earth's magnetosphere, it can generate electrical currents that disrupt power grids, change the shape of our atmosphere, and interact with atmospheric particles to generate the auroral glow.

The effect Yizengaw and his colleagues investigated is produced in a similar way. During a geomagnetic storm, energetic particles can rain down into the ionosphere, a region GPS signals have to travel through.

This mixing and roiling can create density fluctuations in the upper atmosphere. Think of an antique window pane, where the glass is unevenly distributed. Light traveling through that glass can distort and magnify the image it carries, so you see a skewed representation of the world outside.

Similarly, radio signals traveling through the lumpy ionosphere can become distorted and diffracted, causing their strength to fluctuate rapidly by the time they reach a ground receiver. This effect is known as amplitude scintillation.

Ionospheric scintillation isn't unusual, particularly towards the poles and around the equator. The mid-latitudes, however, are generally considered relatively calm and safe when it comes to this particular space-weather hazard.

The November 2025 superstorm said PSYCH.

As the storm intensified, the auroral oval expanded towards the equator, bringing the atmospheric chicanery usually associated with higher latitudes along for the ride.

A Solar Storm Caused GPS Chaos Across The US. Scientists Have Never Seen This Before.

A NASA mosaic of the auroral oval over 24 hours on 12 November 2025. (NASA)

Yizengaw and his colleagues pieced together what happened using observations from multiple instruments across North America, including aurora cameras and a network of ground-based Global Navigation Satellite System (GNSS) receivers.

They saw a huge band of enhanced electron density stretching east to west across the ionosphere. Along its edge, the electron density changed sharply, creating conditions perfect for the formation of smaller-scale irregularities.

And those irregularities were everywhere.

Strong amplitude scintillation appeared across a vast swathe of the continental US, from roughly 80 to 120 degrees west longitude.

Other measurements showed the disturbance extended even farther, producing a strip of enhanced electron density that reached almost from the West Coast to the East Coast.

A Solar Storm Caused GPS Chaos Across The US. Scientists Have Never Seen This Before.

The November 2025 superstorm disrupted Earth's ionosphere across North America. (Yizengaw et al., Geophys. Res. Lett., 2026)

The timing lined up, too. The researchers saw that, as the aurora brightened, electron density and irregularities intensified. At the same time, satellite signals began to scintillate, and GPS accuracy deteriorated.

Amplitude scintillation has been detected at mid-latitudes before, but only in limited observations, mostly at individual locations. Strong amplitude scintillation spanning such a wide range of longitudes has never been seen before, the researchers say.

In some regions, the resulting horizontal positioning errors exceeded 10 meters. Even an error of just one or two meters can spell serious trouble for technologies that depend on precision positioning, including autonomous vehicles and agricultural machinery.

Indeed, the solar storm of May 2024 is estimated to have cost the US agricultural industry $500 million due to disruptions in precision navigation.

The November 2025 superstorm is unlikely to have cost anywhere near the same amount, which really highlights the sheer dumb luck of the draw. The 2024 storm happened during the farming season; the 2025 one did not.

Together, however, the two events indicate how vulnerable certain industries can be to the vagaries of the Sun at the peak of its 11-year activity cycle.

But if scientists can better understand and predict how extreme solar activity affects the ionosphere, we may be better prepared to mitigate the disruption when the next big storm arrives.

"If the November superstorm onset had occurred during farming season in the American sector, it could have led to significant losses for the American farming and transportation industries," they write in their paper.

"Hence, understanding the storm time high‐ and mid‐latitude irregularities – and characterizing their impact on radio-frequency applications – requires knowledge of physical processes that control and describe the dynamics of auroral features, such as energy flux, expansion velocity, and precipitation scale sizes present in the auroral arc, all of which contribute to generating density irregularities that can cause scintillation."

The analysis has been published in Geophysical Research Letters.

The Daily Front Page 12 of 29
Thursday, September 3, 2026 The Daily Front No. #260903 — A Lesson in Gratitude
article

What I learned from my mom (1941-2026)

My mom and I had a daily practice for many years: We would exchange gratitude emails.

My mom and I had a daily practice for many years: We would exchange gratitude emails. Every day, I’d email her one thing I was grateful for, and she’d email me one thing she was grateful for.

My mom sent me such emails as…

“I’m grateful for our cat Hoho who always makes me laugh.”

“I’m grateful for my gorgeous hydrangeas.”

And

“I’m grateful that I’m allergic to chocolate, so I’m not tempted to eat it.”

Which is a great positive spin on being dealt a tough genetic hand.

Over the years, I emailed her hundreds of reasons I’m grateful. One thing I wish I’d emailed more often: How grateful I am that Ellen Jacobs was my mom.

Because she was a wonderful mother.

First of all, she was supportive. So supportive. Consider this:

When my first book “The Know-It-All” came out, my mom bought so many copies on Amazon.com that the Amazon.com corporate office sent her a notice that she was forbidden to buy any more. They thought she was trying to set up her own bookstore to compete with them.

Another example: When I was researching a book on Elvis Presley, my mom volunteered to help with research. She ended up watching and taking notes on 12 Elvis movies. That’s dedication.

But it wasn’t just that she was a supportive mom. She was an inspirational mom.

She was a role model for me in so many ways — in being intellectually curious. In showing kindness whenever possible. In dealing with life’s hardships with humor. She once emailed me, “A sense of humor will get you through the toughest times--never lose it.”

I also got her gene for worrying and nervousness, but you have to take the rough with the smooth.

Being my mom was just one of her roles. She was so many things to so many people. To borrow a phrase from Walt Whitman — whom she studied as an English major at Cornell — my mother contained multitudes.

Since my mom died a few days ago, I wanted to briefly talk about a handful of those roles:

MY MOM THE TEACHER

My mom spent 20 years teaching science to middle-schoolers. She was a beloved teacher, partly because she made science fun. Consider that she taught fundamental scientific principles via snacks and treats. How can you not love that?

To demonstrate the Big Bang Theory, she cooked muffins with her students. The lemon batter represented space, and the raisins were galaxies. That is still how I visualize the universe, by the way.

A few years ago, one of her students wrote me a letter about this culinary side of her teaching.

“Your mom started ’the kitchen corner’, and outfitted the area with real cooking tools and equipment. I’ll never forget my first day assigned to prepare Ants on a Log. Spreading the peanut butter on the celery stalk without it smearing over everything; getting the raisins lined up—it was too much for me. Your mom crouched next to me, guiding me through the steps. She was so encouraging

The feeling that came over me that afternoon--accomplishment and deep satisfaction--is what I’ve been chasing every night for the last 20 years.”

The writer of that letter is Dan Barber, owner of Blue Hill at Stone Barns, one of the most influential chefs in the world. So thank you, Mom, for inspiring Dan and hundreds of other students.

MY MOM THE ARTIST

My mom was a deeply talented artist. She painted, she drew, she did calligraphy. Perhaps she shined brightest as a designer and maker of jewelry.

She made gorgeous, intricate necklaces, earrings, and bracelets with gemstones — including, of course, beryl stones (Beryl is my sister).

I remember visiting her at her jewelry studio, where she was wearing a pair of industrial-grade goggles and holding a flaming welding torch. She looked like she was on the assembly line at the Detroit Chrysler plant.

She was a specialist at something called cloisonné, which is hard to spell and even harder to master. She would use melted enamel and wire to make these detailed, beautiful images of swans or flowers.

Her friend and teacher Melissa said my mom’s cloisonné work made her a “legend” at the jewelry studio. That was the word she used: Legend. She gifted her jewelry to friends and family, and I love seeing them wear it proudly to this day.

MY MOM THE GRANDMA

My mom loved being a grandmother and she excelled at that role. In fact, she had a T-shirt that said…

“What happens at Grandma’s, stays at Grandma’s”

My mom was just the right amount of indulgent to her five grandkids, Jasper, Zane, Lucas, Isabella and Micaela.

When my sons were little, she picked them up from school at least once a week and took them to our local grocery store to buy lollipops. I wasn’t supposed to know, but the grocery store owner spilled the beans. I have come to accept it.

MY MOM THE CAT MOM

There was only one thing that got as much love and indulgence as her grandkids, and that was Hoho the cat.

I’d often visit her and she’d be lying on her bed reading a magazine. But she’d be curled up in the Northwest corner of the mattress, and Hoho would be spreading himself out over the vast mainland. I’m not sure how a little cat could take up that much room, but he did. And of course, he’d be getting his whiskers petted, and he’d be loving it.

MY MOM THE SUPPORTIVE FRIEND

My mom’s 80th birthday happened during the quarantine, so we couldn’t throw a party, of course. Instead, Beryl and I asked several of her friends and family members to write short notes about how our mom had impacted their lives. The notes were remarkable, so moving and full of love.

And one of the recurring themes was just how much our mom helped friends and family through tough times — divorces, deaths in the family, business failures.

You were a “lifeline,” one friend wrote, citing my mom’s kindness, gentleness, and wisdom.

MY MOM THE SPOUSE

Both of my parents hit the jackpot, spouse-wise.

They loved and adored each other from the time they met as students at Cornell, and their marriage has been a model for me in my own marriage.

My dad had a particularly difficult job these last couple of years, as he had to be caretaker for my mom as she battled dementia. But he never wavered, frequently reminding us during the crisis, “Your mom is an amazing woman.”

And she, in turn, was a wonderful wife to my dad.

Early in their marriage, she moved from New York to Korea where my dad was stationed in the Army. And she and Beryl lived in Korea for 15 months. Moving across the world would be adventurous now — but fifty years ago, before WiFi and chain stores — it was truly a bold leap of faith.

My mom was the greatest audience for my dad’s jokes. Even though she had heard them once or twice or a thousand times before, she still loved his repertoire of classics.

She still chuckled every time he went to a friend’s house and ordered a “Yellow Lightning” - a fictional drink consisting of tequila and lemon Kool-Aid. My dad made it up on the theory that no one had both ingredients in their house. My mom would be the merciful half of the pair, explaining to the flustered hosts that it was an elaborate joke.

But she wasn’t just an audience. She had a delightfully mischievous sense of humor of her own. And she was good at poking gentle fun at my dad. For instance, my dad spent hours in the summer writing his law books on the beach. She always said he was so focused that he never looked up at the scenery. He wouldn’t even notice if a tidal wave came and swallowed him up. As a birthday present, she painted him a T-shirt with an image of my dad’s beach hat floating on the ocean. He wore it proudly for years.

MY MOM THE BOOK LOVER

My mom loved books (one of the reasons I was interested in becoming a writer). One expression of that love was her decades of work with the New York Public Library. There, she served two terms as head of the 500 library volunteers.

And she gave hundreds of tours as a docent to the library’s exhibitions. These exhibitions changed several times a year, and my mom would have to learn an entire new topic — Biblical literature, Chinese art, you name it.

She had a way of making the tours particularly engaging, finding the most interesting details. I remember there was an exhibit on restaurants, which had a section on diners. My mom learned some of the diner lingo used by the waitresses and cooks, and taught some phrases to her tour groups. Thanks to my mom, I know that milk is called “moo juice.” And a decaf coffee with nonfat milk? That’s called a “why bother.”

MY MOM THE DETERMINED OPTIMIST

I use the word “determined,” because optimism didn’t come naturally to her. My mom, like me, was born with a gene for worrying. She and I are both very good at picturing worst case scenarios.

But together, we worked to fight this. We sent the above-mentioned gratitude emails. We reminded ourselves how lucky we were to have a family like ours.

For the last three years of her life, several times a week, my mom and I would watch a classic Hollywood musical together.

And during that time, we watched not one, not two, but three songs about seeing the good side of all the rain that falls in your life. There was “Singin’ in the Rain,” of course. And one called “It’s Lovely Weather for Ducks,” sung by Rosemary Clooney. And another called “Isn’t this a Lovely Day,” sung by a drenched Fred Astaire.

And that is what we tried to do together, as the rain fell pretty hard toward the end of her life. We cherished those hours watching musicals together, and singing along, usually slightly off-key.

As you can see, my mom really did contain multitudes. But the common thread running through all of her identities was this: She left the world a better place.

She left it more beautiful place, thanks to her jewelry and art.

She left it kinder place, thanks to her role modeling.

She left it smarter place, thanks to teaching and tour-giving and book-reading.

All of which makes me send this belated email to her: I’m grateful to you, mom, for being such a bright light in this world, and for teaching me so much.

The Daily Front Page 13 of 29
Thursday, September 3, 2026 The Daily Front No. #260903 — Why Robots Still Stumble
article

Reasons robotics is hard

by ddp26·▲ 120 points·75 comments·secondthoughts.ai ↗
The physical world is basically an infinite amount of global state that must be perceived indirectly through imperfect sensors and acted on using imperfect motors and manipulators.

And why you should ignore demo videos

AI progress is racing along, but virtually all of the visible progress is in the realm of knowledge work, i.e. activities that can take place inside a computer.

In the San Francisco AI scene, there is a widespread belief that robots will soon enter the picture. In parallel with the race to develop broadly capable AI, there is an equally aggressive race to develop broadly capable robots – humanoid machines imbued with physical intelligence. Artificial workers that can cook and clean, fetch and carry… and do everything else, including building more of themselves, leading (in many forecasts) to economic growth best characterized as an “explosion”.

Inspired of course by this classic cartoon by the great Sidney Harris

In other words, the thinking goes, AI in the data center will soon subsume all intellectual labor, and AI in humanoid bodies will soon subsume all physical labor. However, there is an important difference: while we can see progress in the intellectual realm, the physical side of AI is mostly confined to test facilities and demo videos. There is no robot equivalent to ChatGPT – nothing that you or I, or even most people in the AI community, can get our hands on.

So we’re stuck with demo videos. Unfortunately, they are a poor tool for assessing progress. We might be seeing the one successful task achieved in 100 attempts. The scenario might have been carefully arranged to avoid challenges the robot isn’t ready for. The video might be edited to make it look like the robot is acting with more speed and reliability than is actually the case. Here’s one very impressive demo… with a suspiciously large number of camera cuts.

(I have not yet had much chance to watch videos from the recent World Humanoid Robot Games. These are valuable for providing a public platform less amenable to cherry-picking. The handful of videos I’ve watched include some impressive feats, but don’t address many of the challenges I list below… and there are also a lot of spectacular failures.)

Demos draw attention to the things a robot can already do. The question then becomes: what’s missing? In today’s post, I’ll catalog the technical challenges that will have to be overcome along the road to broadly capable artificial workers. The next time you watch a robot doing something impressive, ask yourself: which of these capabilities has the robot demonstrated, and which challenges might the demo scenario be avoiding?

(Note that some challenges get easier if we consider wheeled robots rather than strictly humanoid robots. A wheeled robot can carry more weight, meaning that strength, endurance, and power for electronics are less of a challenge. And wheeled robots are less likely to fall over. But they can’t climb stairs1, step over clutter, or angle themselves to reach into a cupboard.)

Hands, dexterity and coordination

Maybe one of the last human jobs will be close-up magic

The human hand is an engineering miracle – opposable thumbs, and all that. It has roughly two dozen “degrees of freedom” (distinct joints and/or directions in which each joint can bend), and approximately 17,000 tactile sensors. Our brains can control our hands with exquisite grace, using touch, sight, and even auditory cues to carry out all manner of delicate tasks, precisely and reliably.

Current robot “manipulators” are a pale imitation. Some existing robot hands can match the human standard on one or another physical attribute. For example, some have as many as 27 degrees of freedom. However, none come close to matching the overall package of flexibility, sensitivity, strength, reliability, and other physical attributes. It is the combination of factors that is especially difficult to match, even if the demos are getting more impressive. For instance, some companies have managed to cram thousands of tactile sensors into a robotic fingertip, but none have managed to make these tiny sensors able to stand up to heavy use2.

The control problem may be as challenging as the problem of physical construction. A competent robot must be able to find the right set of joint positions to grasp a complicated object; plan out the sequence of motions to fold a shirt, flip an omelette, or tighten a bolt in a constrained space; and handle squishy or floppy materials (which can require reacting instantly to a sudden shift).

Visual understanding

Computer vision has made incredible strides over the last decade or two (and is responsible for kicking off the deep learning boom that led to LLMs). But making sense of complicated visual scenes – picking out an object from a crowded environment, understanding where it should be grasped, determining where it’s safe to put your feet and how to avoid knocking something over – is not a solved problem.

Planning and reacting

A general-purpose robot must be able to break down a task into individual steps, and relate those steps to its environment. How do you maneuver your arm to get a screwdriver into a piece of machinery? What’s the quickest way to clear a path to the spice bottle at the back of the shelf? In what order should you pick up the items on the living room floor?

True autonomy will require planning tasks of greater scale and complexity: cooking a meal, plumbing a bathroom, repairing an engine. Not to mention the need to re-plan in the face of surprises – a stuck bolt, a rotten piece of produce, a child darting into the kitchen.

Understanding the assignment

When current AIs fail at a knowledge work task, it’s often because they weren’t provided with sufficient context. Robots will need context, too: where are supplies kept? How do you like your meals cooked? How much assistance does that nursing home resident need, and is that hitch in their stride normal, or a sign that they’re about to stumble?

Once they have context, robots will need to reason, plan, and exercise judgement and common sense. LLM-based systems like ChatGPT and Claude are making great strides in these areas, but the physical domain brings additional challenges3. The success of LLMs has been greatly assisted by the massive pools of pre-existing data that were available for training – a substantial fraction of all books ever written, the web, and other massive pools of pre-existing data. It will be difficult to match this scale of breadth and depth of data for physical tasks. There’s no straightforward equivalent of “just Efficient learning, generalization, and adaptability / on-the-job learning seem like requirements.

Cooperation

AI agents mostly operate in isolation, and in static environments. We rarely put them in situations where things are changing out from under them, or ask them to coordinate. When we do, things often go haywire. Isolation is easier to arrange in the virtual world, where private workspaces can be created at will, and nothing is too heavy to lift on your own. Robots will often need to cooperate with people, or with one another.

Speed

From an article which notes “it took the H1 nearly a full two minutes to very slowly move to a couch, pick up a single item of clothing and put it into the washing machine”.

Today’s general-purpose robots often move much more slowly than human beings. Challenges include strength, control (higher speed means less time to plan and react), and safety (a fast-moving robot will whack you harder and is harder to dodge).

For some applications, slow and steady may be perfectly acceptable: I may not care if my household robot takes all night to tidy up and fold the laundry. But a slow-motion robot won’t be much use as a cook or nursing-home aide. It might get in the way at a warehouse. And it will have a harder time getting enough work done to pay for itself.

Strength

That… is just not an impressive amount of weight for a full-grown robot.

Some industrial robots are extremely strong. But humanoid robots – or other highly mobile, “general-purpose” robots – usually aren’t. It’s difficult to combine strength with manageable weight, a large number of joints, and a maneuverable frame. Powerful motors generate more heat and deplete batteries faster – two areas where robots already struggle (see below). And a strong, heavy robot poses greater safety challenges.

Mobility

ED-209 may have autocannons and a rocket launcher, but it was no match for the staircase

The jury is still out on the appropriate form factor for general-purpose robots, especially with regard to their lower half. Should they have wheels or legs? Two legs, four, or some other number? Wheels are cheaper, more stable, and more reliable; legs are better for stepping over obstacles and climbing stairs. A bipedal frame is more maneuverable, but also more likely to topple if something goes wrong. In any case, the question is: can the robot reliably get around its work environment?

Safety

Sadly, yes, that is a robot karate-kicking a child in the stomach (video). Fortunately the kid was OK.

Safety considerations for general-purpose robots are almost limitless. A glitchy or malfunctioning robot could bump into someone, topple onto them, drop something on them, spill something on them, break a glass, or start a fire.

Safety for LLM-based agents relies in part on review of discrete actions, such as attempts to send an email or delete a file. Robots move constantly, and it’s not so easy to single out a few specific motions as the potentially dangerous ones requiring review.

If a self-driving car finds itself in a situation it can’t handle or suffers a glitch, it can pull over or, in the worst case, just hit the brakes. A general-purpose robot that suddenly freezes might leave something on the stove, topple mid-step, or trip the person it was assisting.

And of course danger can be initiated by human action, such as a child darting in front of a robot. I’d much rather my kid be bumped into by a squishy person than a metal robot; and as things stand today, I’d much rather depend on human reflexes and adaptability to avoid tripping over the little rascal.

(The stronger, heavier, and more capable the robot, the greater the risks.)

Endurance

Today’s bipedal robots can typically run for a few hours before recharging. I suspect this won’t be a limiting factor: if a workaround is needed, we’ll find one, whether that means swapping battery packs, in-floor charging grids, or a cable running to a nearby big-battery-on-wheels4.

I was surprised however to learn that overheating is a serious challenge for continuous operation. A human-sized robot generates about twice as much heat as a person, and robots don’t have the same elegant mechanisms (whole-body circulation and perspiration) for distributing and dissipating that heat.

Then there’s the question of reliability. Robots have large numbers of moving parts, many of which are necessarily finicky, because they’re engineered to push the envelope on size, weight, and performance. As a result, current attempts at general-purpose humanoid robots experience frequent breakdowns. (Contrast the human body, which is constantly recovering from wear and tear, and has substantial ability to self-repair and to compensate for minor breakdowns.)

Generalization; edge cases

Somehow Waymo forgot to teach their cars not to drive into flooded intersections?

Waymos struggle to handle edge cases. They’ve recently been observed driving into flooded roads or over burning fireworks. This is despite the fact that self-driving cars have been in development for well over two decades, and Waymo cars in particular have driven over 220 million miles – 250 times as many as a typical American drives in their lifetime.

(I’m willing to cut them some slack on the fireworks thing; 250th anniversaries don’t come along all that often. But I am confused at how Waymo engineering can be so robust as to yield an astonishingly good safety record, and yet so slapdash as to happily drive into deep water.)

For all of the weird edge cases that arise while driving – the classic example being a duck being chased by a broom-wielding woman in a wheelchair – robots operating in homes and businesses will encounter far more. They will be faced with a wider variety of tasks, using a wider variety of equipment (different tools; different robot bodies to grasp those tools), in a wider variety of environments. Achieving reliable operation outside of the controlled environment of a factory floor may be the hardest challenge of all.

Compute

Dexterity, coordination, visual understanding, planning, reacting, understanding, cooperating – and doing all of these quickly, safely, and reliably – will take a lot of computing power. Incorporating the necessary computing capacity into the robot itself will add cost, drain batteries, and contribute to overheating. Leaving the robot’s brains in the cloud will slow down reaction times and introduce new failure modes (Wi-Fi outage → dead robot).

Scaling

For robots to “really happen” will take a lot of robots

LLMs were able to scale rapidly from the moment ChatGPT was launched, because existing hardware (GPUs) and manufacturing facilities (chip fabs) were easily adapted to support the new use case.

Large-scale supply chains for advanced robots don’t exist yet. Even once we have workable designs, it may take years before they can be manufactured, deployed, and maintained at scale. This sort of thing doesn’t happen overnight; it took 14 years for Tesla to advance from first commercial sales to their first million-car year5. One analysis found that once the starting gun is fired (advanced humanoid robots become economically valuable), it might take several years to scale to producing low-millions of robots per year. ChatGPT, by contrast, reached 100 million users6 within two months of launch.

Political and societal challenges

Political, regulatory, organizational, and cultural barriers: concerns over job displacement and safety may limit where and how robots can be used. Many regulations were not written with robots in mind – does a robot count toward minimum staffing requirements? Will businesses leap to adopt robots? Will they have concerns over reliability, security, liability, and maintenance? Will customers want to be served by a robot?

Cybersecurity, surveillance, misuse: an advanced robot could be a criminal’s dream – a dependable henchman that can’t betray its owner7. Preventing this might require continuous monitoring of all robots, which raises all sorts of concerns. And the cybersecurity on robots will need to be airtight.

Privacy concerns: people already have privacy concerns about Roombas. A humanoid household robot would have the capability to see and hear much more.

Cost: once the other hurdles are addressed, I suspect this won’t be much of a limiting factor. A tireless worker at the price of a new car would be a bargain, and humanoid robots will be much smaller and lighter than a car, with fewer moving parts8. It could be that the components, manufacturing techniques, and training processes required for capable robots will make early models much more expensive than a car. But even expensive robots would likely find early use cases – for instance, doing hazardous work.

Mind the Demo / Reality Gap

Many hurdles will need to be cleared before robots become capable, reliable, practical workers outside of carefully controlled factory environments. The list I’ve presented is surely incomplete; and I’ve only briefly glossed over the cognitive side – understanding, planning, acting, and reacting.

Demo videos provide a glimpse into what robots can accomplish under ideal circumstances. They can also serve to distract us from the remaining limitations. For knowledge work, there’s a consistent gap in AI performance between benchmarks and real-world work. In the physical world, I suspect the demo / reality gap will be even larger. Playing around with an LLM to see what it can do has been accessible, cheap, and (mostly!) safe. To assess the capabilities of robots, we’ll be much more reliant on controlled demos and manufacturer’s claims. It will be harder to map the jagged boundary of their capabilities.

As I was putting the final touches on this post, the excellent Understanding AI blog posted Why humanoid robots won’t catch up to human workers any time soon. I haven’t read it yet but I’m sure it’s worth a look.

For a good broad review of robot capabilities, see Epoch’s report from February 2026.

1

There are four-legged robots whose “feet” are wheels, which can alternate between rolling and walking, and which can climb stairs. But so far there is no sign of this design going mainstream.

2

Workarounds may be found. For instance, some robot designs use cameras to partially compensate for tactile sensitivity, with extra “eyes” in the wrists or elsewhere; a robot could even make use of detached cameras via Wi-Fi. Robots could also incorporate non-human senses, such as ultrasound.

3

See this explanation of the importance of tacit knowledge in bio lab work for an example of how much robots will need to learn.

4

One recent demonstration showed a robot sorting packages for 200 hours straight, processing 250,000 packages… using three different robots, taking turns while the others recharged. It was meant as a demonstration of reliability as well as endurance.

5

From 2008 to 2022.

6

Measured by monthly active users.

7

Assuming the criminal managed to avoid a paper trail connecting them to the robot.

8

At least, compared to an internal combustion car.

The Daily Front Page 14 of 29
Thursday, September 3, 2026 The Daily Front No. #260903 — Xanadu, Reconsidered
article

Xanadu was waiting for agents

by nsm·▲ 103 points·42 comments·zed.dev ↗
He imagined a system that could keep every version of a document.

Ted Nelson has been on my mind a lot this year.

Nelson was a pioneer in computing, and in 1965 he coined the word hypertext as part of the project that would define his career: Project Xanadu. With Xanadu, he had a specific vision of what computing could be, which he called the docuverse. He imagined a system that could keep every version of a document; hypertext links would know both their source and destination, and quotations would be kept by reference (rather than copy) so any included text maintained its identity and source. Everything, in this version of computing, would be intertwingled and xanalogical. We're talking about attribution down to the span.

To be xanalogical is to follow two rules: Never copy, always reference (also known by a Nelsonism, transclusion). Never overwrite, always version. As Nelson designed it, Xanadu would manage the system's complexity and massive bookkeeping burden itself, so users only experienced the benefits of a system that could never forget.

The vision of unlimited bookkeeping requires the reality of unlimited storage and a future-proof naming scheme, neither of which existed for the decades that Nelson worked on Xanadu. So when web technology exploded in the '90s, developers bypassed xanalogical computing in favor of ease: links are simple strings which break if their target moves. Maintenance became users' problems — instead of the system's responsibility — but it was easy for anyone to ship fast, and they did. Xanadu faded from the promise of a new era into computing's most famous vaporware.

Nelson has spent his career explaining that the web is a flattened parody of what hypertext was meant to be. He's right, but it didn't matter because people didn't actually need to be able to follow every link or compare every version. We're a restless species, and more inclined to follow the links or read the versions we see others pursuing. Flat was good enough.

Then agents arrived.

What makes an agent an awkward collaborator also makes it a perfect citizen for Nelson's system. Agents can follow more layers of subtext than people can hold in their heads at once, such as the sources behind a quotation, any discussion around it, and stack traces attached to the exact code that ran. A fragment-based representation makes those references lossless, so each layer stays attached to the same span as the text changes. Agents can read these dimensions together, cite them so humans can audit their work, and patiently follow every link, every time. But those links need to actually exist somewhere in order for an agent to find them.

When I revisit Nelson's vision now, I recognize DeltaDB's design goals and the promise of Delta. After decades of squashing history to meet the limits of the human attention span, we're in a new reality that would be better served by something more... xanalogical.

How to build a system before the parts exist

A few years ago I wrote about watching Engelbart's 1968 demo and my realization that to build a collaborative editor, his team had to invent everything it would rely on: a new programming language, operating system, and displays.

Nelson had the same problem (but with way less funding than Engelbart enjoyed). He couldn't build Xanadu because most of its parts didn't exist, such as a way to name content that no authority issues, or storage cheap enough to never delete. His teams hand-built their own data structures for years, but the project died on the vine. Wired depicted Nelson's story as one of mismanagement, but I think he'd just specified Xanadu decades before its dependency tree existed.

The dependency tree exists today

We can build Delta today because sixty years of other people's roadmaps delivered the docuverse's missing pieces: kernel development, photo storage, serverless cold starts, collaborative cursors. I think it's fascinating that breakthroughs have ordinary day jobs, yet will enable the next era of computing.

A clock for a world with no center. Lamport timestamps, 1978. Every operation by every human and agent is named by an actor plus a Lamport timestamp, forever.

Names that can't lie. Merkle trees, 1979, made ordinary by Git in 2005. A Git commit hash names an exact immutable project state; DeltaDB names every state between commits by the Git commit it descends from and the set of actor-and-timestamp delta IDs applied above it.

Convergence without coordination. CRDTs, formalized in 2011, and the center of Zed's own work for the past decade. A Delta worktree can be edited by several people and agents on different continents at once.

Storage too cheap to delete. A gigabyte cost tens of thousands of dollars in 1981; it costs a penny now. We keep every version of everything by default.

Networks that replicate everything. Always-on and fast broadband overtook dial-up in the mid-2000s. Every thread is live, replicated data on every participant's machine and in the browser.

Machines summoned in milliseconds. Firecracker-class microVMs, 2018. Agents in Delta can provision a new isolated cloud machine mid-conversation.

Documents as views over permanent content. Max Brunsfeld's Tree-sitter, 2018, fast enough to reparse as you type. GPUI, pioneered with Zed, fast enough to derive a fresh interface from application state whenever a frame is drawn. A Delta thread is a live projection of its permanent structured history, recomputed rather than stored as a flattened document.

On one detail — across a career of being told his dreams were too big — Nelson dreamt too small. The final dependency he was missing was a new kind of user. He didn't imagine artificial readers, though they existed in science fiction (like Asimov's Multivac, or Stephenson's Librarian).

Though Xanadu was blocked for decades by missing components, what it needed most was the perfect user.

Why the docuverse is materializing in version control

Most actual code creation has always happened between commits, but it used to hurt less to flatten the context because we got by on human memory (and nobody was going to read full records of keystroke changes anyway). But agents keep nothing in their heads, and they'll read everything, so we need a way to capture the new source of truth using exactly what Nelson specified: permanence and connection.

Every Delta thread delivers on that promise. The conversation and the code are captured together in a shared history. On screen, a file still looks like a one-dimensional string of characters but underneath, DeltaDB represents it as fragments with stable identities. Those identities let us create anchors: references to spans that can still be resolved after surrounding code changes. A line number can express where text appears in one snapshot, while an anchor preserves which span we mean across snapshots.

We have been exploring how agents can build on that coordinate system. The files and symbols earlier agents repeatedly read, edited, cited, or returned to can become landmarks for the next agent, resolved against the code as it exists now and linked back to the conversations where that understanding was formed. By preserving the causal metadata beneath that surface representation, like which operation produced each fragment and what prior state it built on, DeltaDB gives the model a way to traverse not just the current code, but its provenance, accumulated attention, and prior reasoning.

Avoiding Xanadu's curse

Xanadu had a final failure mode, this one self-inflicted: it refused to interoperate with lesser formats. As we introduce DeltaDB to the world, we're learning from that mistake. We'll work with the git repository you already have. Every thread is also a git branch, so teammates who never open Delta see a normal repo, and because the files are real, it integrates seamlessly with any tool an agent can use. You can keep mirroring your repo to GitHub to work with anyone who isn't ready to make the switch (that's our approach for our open source IDE, Zed, during this transitional period).

Engelbart and Nelson were prophets of early computing, yet we never shipped either's vision whole: from the former we adopted the mouse (and eventually multiplayer editing), and from the latter we took the word hypertext. Both ideas survived, but neither at the depth their creators imagined. We still split the live session from the durable record: discussion, intermediate states, and intent are flattened into a final file or commit, while links tied to paths and line numbers decay as the code moves. I hope for Delta threads to meet both visions at once: a live session that leaves behind a permanent, connected literature, because the session and the record are the same object.

Every property we needed to build Delta and DeltaDB was specified by Nelson before I was born. And though my goal is not to realize Xanadu, I still feel some of the antsiness Nelson must have experienced while hoping for a long-held vision to land in people's hands. Sixty years is a long time for an idea to wait for its users. They're here now.

The Daily Front Page 15 of 29
Thursday, September 3, 2026 The Daily Front No. #260903 — The .arpa Loophole
article

How to get a free .arpa domain

by ethanhawksley·▲ 131 points·14 comments·hawksley.dev ↗
There are many different services that you can abuse to get your own DNS records, and hence your own website.

This post is also available over on get-arpa-domain.0.4.1.0.9.0.f.1.0.7.4.0.1.0.0.2.ip6.arpa (opens in a new tab)!

On the internet, .arpa is the top level domain reserved for critical internet infrastructure. It isn’t open to public registration, so you can’t own mysite.arpa sadly. However, this isn’t the only way to reserve some real-estate on .arpa - there are many different services that you can abuse to get your own DNS records, and hence your own website.

I first learned about this through a blog post explaining e164.arpa (opens in a new tab), an old scheme where you could query information about phone numbers using DNS. If you’re German or Czech, you can sign up and control the DNS records that correspond to your phone number!

For the rest of the world, our phone numbers aren’t open to registration, so that isn’t an option. Instead, we can use a similar scheme under ip6.arpa. This is reserved for “reverse DNS lookups”, where you can turn IPv6 addresses into domain names, instead of the other way around. However, nothing in the specification stops us from using it for other purposes, so let’s do so!

Hurricane Electric offers really simple registration for an IPv6 address and its ip6.arpa records through tunnelbroker.net, so let’s use them. Their site is a bit of a relic, but it is fully functional.

Hurricane Electric Tunnelbroker home page

Sign up for a new account, feel free to provide fake information as it doesn’t verify any of it. Verify your email, and then in the left sidebar select “Create Regular Tunnel”.

Tunnelbroker dashboard

It will ask for an IPv4 endpoint, though for our goal it doesn’t matter which address we choose. The IP needs to respond to ICMP Echoes (a.k.a. pings), but there’s no verification that you control the provided IPv4 address. Just use ping -4 domainname.com on a few websites until you find an IPv4 address that it will accept. It seems most sites that are behind CDNs don’t work, so try older sites first. Here I pinged news.ycombinator.com and received the ip address 209.216.230.207.

Terminal pinging news.ycombinator.com

This seemed to pass Hurricane Electric’s form validation, so that’s all that matters. It also asks you to select a Tunnel Server, but similarly this doesn’t matter for our purposes - any will do.

Tunnelbroker create new tunnel form

Once you’ve created a tunnel, take note of the “Routed IPv6 Prefix”. Here mine is 2001:470:1f09:140::/64, but the only part we need is 2001:470:1f09:140 before the trailing colons.

Tunnelbroker tunnel details

Pad each section with zeroes to get four groups of four characters: 2001:0470:1f09:0140.

Then place a dot between each character: 2.0.0.1.0.4.7.0.1.f.0.9.0.1.4.0.

Finally, reverse the characters and append .ip6.arpa to get your own domain name: 0.4.1.0.9.0.f.1.0.7.4.0.1.0.0.2.ip6.arpa!

Next, we need to set up DNS records for the domain. When I last tried this, Cloudflare didn’t seem to accept the domain, but deSEC handled it perfectly, so that is what we’ll be using. Sign up for a deSEC account and provide the .ip6.arpa address you calculated from earlier.

deSEC signup form

Once added, return to Tunnelbroker and change your rDNS delegations to that of deSEC.

Tunnelbroker rDNS delegations

We have a domain, we have DNS, so the very last step is a web host. Out of the options, Surge is very straightforward and doesn’t enforce HTTPS. This is handy because many (though not all) certificate authorities refuse to serve .arpa domains, hence getting HTTPS working can be quite challenging. It isn’t impossible, though, but it is beyond the scope of this guide.

Create a new directory and an index.html file. You can put whatever content you’d like inside the index.html, I decided to put a bare-bones replica of this blog post (opens in a new tab).

Install npm or bun if you haven’t already, and run surge.

bunx surge . subdomain.domain.ip6.arpa
# or alternatively
npx surge . subdomain.domain.ip6.arpa

Here, “subdomain” is a subdomain of your choice, and “domain” is the ip6.arpa address you calculated earlier. If it asks, make a free account as part of signup. Make note of the address it gives near the end of its output.

Surge command bunx

Surge command response

Lastly, return to deSEC and create a CNAME record pointing to the surge.sh address.

deSEC creating CNAME record

Assuming your local resolver hasn’t cached an earlier NXDOMAIN response, you can visit your domain immediately and see your site! You now have a fully functional website running off a reserved infrastructure TLD.

The Daily Front Page 16 of 29
Thursday, September 3, 2026 The Daily Front No. #260903 — Poe’s Witnesses
article

The true horror of Edgar Allan Poe’s stories lies in their confessions

by lermontov·▲ 63 points·26 comments·yalereview.org ↗
Edgar Allan Poe’s stories are full of ghastly crimes, but the true horror lies in their confessions.

Edgar Allan Poe’s stories are full of ghastly crimes, but the true horror lies in their confessions.

Ink illustration of a small boat suspended on the inner wall of a vast whirlpool of spiralling lines.

Harry Clarke, illustration for Poe’s “A Descent into the Maelström,” from the 1919 edition of Tales of Mystery and Imagination. Courtesy The Public Domain Review

In the summer of 1841, Edgar Allan Poe was a man on the rise. At the age of thirty-two, he was living in Philadelphia, where he worked as an editor at Graham’s magazine. Back then, Philadelphia was the magazine-publishing capital of the United States, and Graham’s was one of the nation’s most important illustrated monthlies. With a peak circulation of forty thousand—enormous for its time—Graham’s included tales, light essays, colored fashion plates, and reviews. As a literary editor, Poe provided reviews and, for an extra fee, tales as well.

In those days, Poe was a respected reviewer and a fiction writer of growing reputation. He had published three poetry collections of small circulation, as well as a sea novel, The Narrative of Arthur Gordon Pym of Nantucket (1838). His first collection of stories, Tales of the Grotesque and Arabesque, had come out in 1839*.*

Although Poe was a sought-after writer, it was not easy for him to make money. The idea that writing was a profession from which one could make a living was still new. It would not have seemed unusual to Poe’s contemporaries that the publishers of Tales of the Grotesque and Arabesque retained all profits from the seven hundred and fifty copies they printed. Poe got to keep the copyright and twenty free copies. Graham’s was one of the first places to compensate its editors and contributors well, and magazine publishing in general was beginning to pay better than book publishing. Poe had set his sights on magazine work.

Reaching the relative eminence of Graham’s staff had not been easy. Orphaned at a young age, Poe had been raised in Richmond, Virginia, by a foster father who later disinherited him. Thrown back on his own resources, Poe struggled. At times, he was so impoverished that he could not clothe himself. Before he finally got into the magazine business, his checkered career included stints as a common soldier, as a student at West Point, and (rumor has it) as a bricklayer. Editing suited him better than other work he had tried, but even so, Poe was prone to imperiling his previous magazine positions through self-sabotage—and drinking. Now, at Graham’s, Poe had a steady job and a promising future. His friends must have been praying he’d be able to keep it.

But Poe grew restless. He wanted to found his own magazine, as he had tried and failed to do before taking the job at Graham’s, a magazine he considered middlebrow. He particularly detested its fashion illustrations. In April 1842, after working at Graham’s for a little over a year, he quit, then carefully choreographed the launch of a new magazine he called The Stylus. By early 1843, he had secured a partner with capital to invest and had arranged for a lengthy biography of himself and a prospectus of The Stylus to run in the Philadelphia Saturday Museum, a local newspaper.

The biography used truths, half truths, and lies to paint Poe as a poetic genius, an entrepreneur, and an athlete. Poe either wrote the biography himself or closely supervised its writing; his fingerprints are all over its most colorful incidents, such as a Byronic voyage to Greece to fight in the revolution (a fabrication) and a six-mile swim in the James River as a teenager (a true event and, as Poe or his ringer wrote, longer than Lord Byron’s more famous “paddle” across the Hellespont). The biography also exaggerated Poe’s magic touch as a periodicals editor. It claimed that he had drastically increased the circulation of The Southern Literary Messenger, a Richmond-based magazine where he had previously worked. This wasn’t true; the scholar Terence Whalen has shown that Poe was “essentially an irrelevant variable” when it came to the Messenger’s circulation. But Poe wanted to paint himself as a man with as much business acumen as taste, the ideal editor for the new Stylus. The piece concluded with a bulleted list of glowing blurbs on Poe from Nathaniel Hawthorne and other well-known authors, some of which came from private letters.

Even as he devoted his energies to his magazine venture, Poe also needed to replace his Graham’s salary. His meager freelance fees would not cut it for long; as a matter of fact, he was already in debt. He hoped for a government sinecure, a light-duty political post that would cover his salary while he devoted his time to The Stylus. Such things were not unheard of; Hawthorne would later hold a position of this kind at the Salem Custom House, a position he immortalized in the preface to The Scarlet Letter. Poe had patiently worked his connections to President John Tyler, and in 1843 he went to Washington in hopes of meeting Robert Tyler, the president’s son.

It was an unfortunate decision. Things went wrong from the beginning. Poe’s friend, who was to secure an introduction to Robert Tyler, was sick and could not take him around. Poe drank too much and made a fool of himself in public. He behaved so badly that one of his hosts, Jesse E. Dow, wrote to Poe’s business partner to say he’d better come get Poe and take him home. Dow didn’t think it was safe to put his friend on the train alone. The letter is strikingly compassionate toward Poe. “He does not understand the ways of politicians,” Dow wrote; “how should he?” Poe, he reflected, “has the highest order of intellect, and I cannot bear that he should be the sport of senseless creatures who, like oysters, keep sober, and gape and swallow everything.”

In the end, Poe made it back on his own, but he never got the sinecure, and he failed to launch The Stylus. He later sent a letter of apology to his hosts in Washington. “Express to your wife my deep regret for the vexation I must have occasioned her,” he wrote to Dow, also asking him to stop at the barbershop and “pay for me a levy which I believe I owe*.*” Evidently, he had stooped so low as to insult someone’s facial hair: he instructed another friend to give his best to a person “whose mustachios I do admire after all.”

Alcohol clearly played a role in Poe’s disastrous Washington trip. When Poe was drunk, and sometimes even when he was sober, he would find himself acting against his own interests for reasons that were opaque to him. Poe could have been thinking of frailties like his own when, in his short stories “The Black Cat” (1843) and “The Imp of the Perverse” (1845), he wrote about what he called perverseness: the “unfathomable longing of the soul to vex itself”*—*the mysterious desire he felt, and believed we all feel, to get in our own way.

Perverseness, Poe thought, was as fundamental a drive as hunger or libido, with its own dedicated region of the brain. When we are perverse, we turn not only against our better judgment but also against what are apparently our desires. “The Imp of the Perverse” begins with an essay defining the concept and giving examples that range from the harmless to the unsettling; Poe classes as perverse both procrastination and the impulse to kill oneself by jumping from a great height.

One of his illustrations is particularly telling: talking too much at a party. The incessant talker dreads angering his listener, and “yet, the thought strikes him, that by certain involutions and parentheses, this anger may be engendered. That single thought is enough. The impulse increases to a wish, the wish to a desire, the desire to an uncontrollable longing, and the longing (to the deep regret and mortification of the speaker, and in defiance of all consequences,) is indulged.” Simply because he can imagine destruction, the perverse person feels compelled to bring it about. He knows he shouldn’t—he sees what he is doing—but he does it anyway. It is tempting to use Poe’s anatomy of perverseness to understand his self-sabotage in Washington. He saw the destructive consequences of his actions. But he couldn’t stop.

Central to the experience of perverseness is this sensation that our own actions are out of our control. Either we are in the middle of a destructive action we can’t seem to stop, or, illogical though it may seem, we are still trying to stop, or undo, the self-defeating thing we already did. The physical sensation of the perverse, as Poe notes, is an icy chill in the blood.

I have felt this chill. When I realized, several days after the fact, that I had left my wallet on top of the car, then driven away, I could not quite recognize myself as the person who would have made this mistake, who did make it. The chill in my blood seemed to want to turn back time, to replay the moment and fix it, by seizing me with a horror so concentrated and so intense that I wouldn’t ever make the mistake again.

what poe called perverseness is akin to what another explorer of the world of our dreams, Sigmund Freud, would, some seventy years later, call the death drive. The two concepts circle around the same mystery: Why do we frustrate ourselves? For both Poe and Freud, dreams and dark fantasies were the key to understanding those desires we have in spite of ourselves. In Beyond the Pleasure Principle (1920), Freud observes how his one-and-a-half-year-old grandson throws away the very toys he values. Freud calls the game fort/da because the little boy yells what sounds like fort, or gone, when he disappears the object, then cries da, or there, when it reappears. Freud puzzles over this self-defeating gesture until it occurs to him that, with his toys, the boy might be replaying his mother’s absence—an absence that threatens his survival—in order to master it. Such a compulsion to repeat unpleasurable experiences can “give the appearance of some ‘daemonic’ force at work,” Freud notes—something persecuting the self from outside. In fact, it’s only the ego, trying to gain ascendancy over what can harm it.

The avenues to the abyss are always opening on either side.

But Freud is not fully satisfied with his own explanation; he cannot shake the sense that the infant is elementally drawn to the destruction of himself and of his pleasure. A part of us wants not exactly to die, he theorizes, but to disorganize ourselves to the point of becoming mere matter: we carry in our very structure a memory of how, before there was life, there was only inanimate stuff. An imp in us wants to take apart our projects, our lives, our living molecules, until only atoms remain. We are perverse, Poe says, when “we stand upon the brink of a precipice.” As we stand there, a “cloud assumes shape.” And “out of this our cloud” comes a demon, or something more terrible than a demon, “and yet it is but a thought.” The thought is of ourselves falling forever and breaking. Conceptualization gets confused with desire: if I can think of jumping, it is as if I want to jump. The thought “chills the very marrow of our bones with the fierceness of the delight of its horror.” There at the bottom, we will be matter again, without the work of holding ourselves together. Freud’s idea is a strange one, hard to get right: the death drive is not so much a wish as a potential for dropping back down from animal to mineral, for becoming nothing again.

Our perverseness is one way to explain a fact that has surprised thinkers at least since Aristotle: what would horrify human beings if it appeared in reality gives them pleasure if it appears in art. Everyone’s threshold is different; not everyone has the stomach for a gory movie. But it is sufficiently remarkable that many people do and that almost none of them are serial killers in real life. I watch vertigo-inducing scenes in movies, even though (or because) I am afraid of heights. When Cary Grant and Eva Marie Saint cling to the faces of the fathers in the Mount Rushmore scene in North by Northwest, my palms sweat. Do I like these scenes? I don’t find them pleasing, but I do find them desirable; they draw me powerfully, even as they also make me want to look away. I am doing something here with my wish to jump—doing something other than acting on it or resisting it. But what? Aristotle’s famous answer is that painful emotions need to be purged from the body and that tragedy achieves this catharsis. Freud might have ventured that, by viewing horrifying art, I repeat and thus assimilate past injuries. He might say that what first appeared as trauma, experience I could not process, I now give to myself to reexperience in a less overwhelming form.

As for Poe, he never directly answered the question of why a person would enjoy calamities in art. But he gave hints. One implication of the argument in “The Imp of the Perverse” is that human beings want destruction, plain and simple, just like they want food or sex. As with food or sex, the desire can be complicated with aversions; it can follow some extraordinarily winding roads. But at the last, there it is. One part of Poe wanted to destroy his Washington chances—not because he thought he deserved to lose the opportunity, not because he had some hang-up he had never worked out, and not even because he was drunk. He wanted to destroy his chances because destruction is, fundamentally and irremediably, one of the things that human beings like. Poe might say that we always want destruction—in life and in art—and in art we find that it is often allowed.

Allowed, but not necessarily good for us. Poe did not demand that our play with destruction make us better people. He was staunchly opposed to the moralization of poetry. Poe was a sensationalist. He thought that art should put our bodies through their paces. It should remind us of what, for good and ill, we are.

Once Poe’s “The Imp of the Perverse” has defined its central term—perversion is the impulse that drives us to work against our interests and desires—the impersonal, theorizing voice becomes a narrator with a story of his own. The narrator, as he tells us, is a murderer who has committed the perfect crime. He killed his victim with a poisoned candle, a murder weapon that left no trace. When the investigation subsided and the case was closed, the narrator went out for a walk. He thought, “I am safe—I am safe—yes—if I be not fool enough to make open confession!” No sooner had this thought occurred to him than he felt compelled to shout his crime to everyone in earshot on the street. He was arrested and tried and, we now learn, will hang; the narrator is speaking to us from his cell, on the eve of his execution. The crime itself was not perverse, in Poe’s sense of self-sabotage, but this confession is: having nearly gotten away with the murder, just as Poe had nearly gotten away with his sinecure, the narrator will soon die for his crime. Perverseness will kill the theorist of the perverse.

The comparison between a murder and a lost job may seem far-fetched to us, but I don’t think it seemed that way to Poe. He was capable of taking off from the point of his own minor perversities—procrastination, prolixity—and supplying a far darker setting for them than anything that, so far as we know, he ever experienced. Poe, like his narrator, destroyed his chances by speaking—insulting his companions, rambling drunkenly—when he ought to have remained silent. The voice of Poe’s gallows narrator is much like his voice in his letter of contrition to his Washington friends: part bravura, part panic, urgently demanding the reader’s belief. He seems to have felt that the cold palms of the acrophobe were essentially the same as the cold palms of the murderer. Horror might touch him at any time through these ordinary sensations. He would remember that life was not right, he was not good, and the shim could slip out from under the table leg for no reason at all. The avenues to the abyss are always opening on either side. The difference between Poe and most of us is both small and gigantic: he took these avenues, again and again, in his stories. To him, the gallows felt close.

when poe gave up his steady job at Graham’s to try to launch The Stylus, his young wife, Virginia, was dangerously ill. Just a few months earlier, in January 1842, she had been singing when blood suddenly began to dribble from her mouth. It was a pulmonary hemorrhage, an early sign of what was then called consumption and what we now know as tuberculosis.

Consumption was the cause of about one-fifth of deaths in early nineteenth-century America—the same percentage as heart disease today. Consumption sometimes killed quickly, but in most cases its progress was slow and insidious. Doctors commonly saw a pattern of recovering, relapsing, and recovering again, often over years and sometimes over decades, before the patient finally succumbed. Consumptive patients might have felt fairly certain that they had been handed a death sentence without being able to guess when the sentence would be executed. And unlike heart disease, consumption typically struck the young. Virginia was nineteen.

It is not clear whether Poe was there when his wife began to bleed. If he was, it would have been traumatic for him, as it surely was for Virginia. Hemorrhages came with little warning beyond a tickle in the throat. Virginia may not have seemed ill before it happened—or if she had any symptoms, such as coughing, they could have masqueraded as a common cold. The arterial blood would have been vermilion, a shock of bright red. Like nearly all Americans at this time, she and her family would have known that coughing up blood portended death.

In the aftermath of the hemorrhage, death hounded Virginia; she recovered and relapsed multiple times. Poe grieved deeply, even as he continued to pursue his risky professional dreams*.* At the end of May, he told a friend that she was much better. In June, when he wrote to another, she was worse: “Mrs. Poe is again dangerously ill with hemorrhage from the lungs. It is folly to hope.” Then, in July, she was taking an expectorant that improved her cough and stopped her night sweats. He allowed himself to be annoyed with her after he went on a bender in New York, disappearing for several days. “Because she did not hear from me twice a day,” he wrote to his cousin Elizabeth Tutt, “she became nearly crazy…she would neither eat nor sleep.…What it is to be pestered with a wife!” If the proximity of death did not perfect him morally but left him mired in everyday worries and petty complaints, then he was like most of us.

About this time, Poe wrote a story called “Life in Death” (later retitled “The Oval Portrait”). It was published in Graham’s in April, the same month that he resigned. In this gothic tale, a mortally wounded traveler, injured in a scrimmage with banditti, takes refuge in an empty château. Inside, the traveler sees an arresting portrait of “a young girl just ripening into womanhood” and finds the story of its making in a book in his bedroom. A young girl is married to an artist, and “evil was the hour” when she fell in love with him. They are terribly matched. “He, passionate, studious, austere, and having already a bride in his Art: she a maiden of rarest beauty and not more lovely than full of glee…frolicksome as the young fawn…hating only the Art which was her rival.”

The frolicksome girl in “Life in Death” agrees to be painted by her husband. She dreads every motionless moment. Every day, they ascend to the garret studio. Every day, the portrait glows with a brighter approximation of life. Every day, she grows sicker. It’s a consumptive illness, probably—a force seems to devour her life from the inside. The artist either doesn’t notice or doesn’t care, so absorbed is he in his work. At the last stroke of his masterpiece—it is perfect!—she expires. She is transformed into the work with which she can no longer interfere. Poe’s story invites us to blame the artist for the fulfillment of a cruel half-conscious wish: that his wife would get out of the way and die.

In another story Poe wrote for Graham’s that May, a plague called the Red Death ravages the countryside. The victims of this disease ooze blood from every pore until they die: “Blood was its Avatar and its seal.” To avoid the pestilence, Prince Prospero walls himself up in a castle with all his friends. It’s a constant party; his wealth provides for everything. The revelers move through a suite of colored rooms: blue, purple, orange, white, green, black. A clock stands in the black room. When the clock chimes the hour, its sound is “clear and loud and deep and exceedingly musical” but with so “peculiar a note and emphasis*”* that, as the narrator recounts, the musicians paused, the dancers stopped, “the giddiest grew pale, and the more aged and sedate passed their hands over their brows.” For a moment, they feel their true condition: they are mortal, and plague or no plague, their lives will at some point come to an end.

Then the chiming ceases, and the spell breaks. They shake it off, promising one another not to be so silly the next time. A cloud passes between them and their blinding knowledge, much to their relief. They go back to dancing.

If we like to feel ourselves on the edge of a cliff or if we like to read stories that frighten us, it is in part because with these sensations we remember that we must die. It isn’t pleasant to think of dying—at least, I don’t find it so—but there is a certain relief that comes when the subtle discord of our ordinary state of denial is lifted. Poe, who pursued his professional dreams even while Virginia struggled for her life, must have known from experience how tempting it is to distract oneself from death. What used to be called the earthly vanities—ambition, money, power, and pleasure—can push the inevitable out of view. The distraction is welcome, in a way, but the dread is still there. Because even when we feel like we’ve forgotten death for a time, we actually do remember, deep down. And keeping up our resistance is an effort.

Poe’s stories offer what, perhaps, he also needed: they let us live for a while in our knowledge of the worst. He gives us that weird and dark reprieve. Poe’s “The Masque of the Red Death” could, in this sense, be called a memento mori*.* No prince, it tells us, is powerful enough to defeat the plague. Every pleasure, no matter how sumptuous, will expire. To remind us of death is one of Poe’s services in his horror stories. He can’t hold death at bay, nor does he have any particular spiritual discipline to recommend. But what he can do for us is pry open a gap in our wall of denial. He can help us remember, for the duration of the story, that we really will cease to be—just like everyone else.

edgar allan poe died in 1849. Nearly a century later, he found a reader who needed the dark company his stories provide: the Argentine novelist Julio Cortázar, still one of Poe’s best critics. Cortázar first read Poe’s stories as a child in the 1920s, in a suburb of Buenos Aires. He came from a family that took evil seriously; they sent him out walking alone at night so that he would become immune to fear. No immunity, however, could protect a person against the basement of their house. As Cortázar wrote, “Nobody had the nerve to descend—ever.” His mother gave him supernatural stories to read but forbade him to read Poe. And she had been right to do so, he later commented: when he managed to find and read the stories anyway, at the age of nine, he fell ill for three months. Even though, or because, these stories terrified the young Cortázar, he also adored the American author; at twelve, he composed an ode to Poe whose rhyme scheme was based on “The Raven.”

More than two decades later, in 1951, Cortázar was living in Paris and working as a delivery driver for a bookstore when he received a commission from the University of Puerto Rico Press to translate Poe into Spanish. Cortázar was delighted with the offer. For one thing, Poe’s influence on him as an artist had been profound. It was Poe, Cortázar always maintained, who taught him to cross over into the “territory of lo otro”—the domain of dreams and funny turns that his fiction occupies. And for another, he needed money. Cortázar had published his first collection of stories, Bestiario (1951), and was working on a new collection, which would be titled Final del juego (1956). The generous commission would allow him to live for months in Rome, where he could work on his own stories while he translated Poe.

Cortázar had imagined an Italian holiday lightly seasoned with translation work. Instead, he spent as many as nine hours a day in the fall and winter of 1953 to 1954 working on nothing but Poe. Cortázar joked that, through him, Poe had decided to write his greatest fantastic tale, “the one about the writer who will not allow his translator to finish translating him.” But finish he did; his final draft ran to fourteen hundred typewritten pages, including all sixty-seven of Poe’s tales, Eureka, Arthur Gordon Pym, and some essays.

Wishing to live can be a horror. That is why Poe was also so interested in what it is like to die.

His translation was finally published in 1956, prefaced by an essay that is a masterpiece of Poe criticism. Cortázar clearly admired Poe yet struggled with his sense of his predecessor as an unfinished personality, an “egotistical weakling” who never assumed the adult responsibility of seeing other people as people. Because Poe could not see others whole, Cortázar argued, his literary characters could never stand at the ethical remove from which we ought to view other human beings; they were either versions of himself or thought experiments that he treated coldly.

“On the one hand,” Cortázar wrote, the characters “are Poe himself, creatures of his depths, known to him as he knows himself. On the other hand, they are only characters, that is, others, alien to him, and at the end of the day insignificant to him.” Needless to say, someone with this particular form of blindness could not write about love. His stories completely lacked normal eroticism; sexuality in them could appear only, Cortázar noted, in “larval or aberrant forms.”

When I am tired of Poe, he wears this guise: stunted and puerile, an egomaniac who could not represent anyone in his art because no one but himself counted for him in life. But I also know—and Cortázar, in the end, knew this, too—that to think this way is to apply the wrong standards. Poe was not trying to represent people in his art; that is not what his characters were for. Poe was not a realist. It isn’t that his characters are poor models of people; it’s that they are not models of people at all. They are more like voices in one’s head or figures in a dream. With each character, Poe models not a person but a drive. The drive might be to obliterate a source of shame, to get to the bottom of things, or to take revenge. It might be, as in “The Imp of the Perverse” or “The Black Cat,” the drive to frustrate oneself. Or it might be more elemental still: in stories such as “The Premature Burial” and “The Fall of the House of Usher,” the drive is simply to live, as an infant or an animal feels compelled to live. The living thing must get back up out of the grave. In Poe’s stories, adult human beings return to that part of ourselves that would, like a fox, gnaw off its leg to escape a trap—or that did, as an infant, wail for succor from the abyss of helplessness. Poe tries to describe what it is like to be prey to one’s wishes, what it is like to be driven into motion. If his subject can also be described as life, it is life experienced as the infliction of activity: only life itself, not yet life as a social or ethical or political matter. It is not always pleasant to wish to live. Wishing to live can be a horror. That is why Poe was also so interested in what it is like to die.

Cortázar sometimes grew exasperated with Poe’s personality, but when his patience returned, he was also among the best at describing his artistry. If Poe remained in touch with his most primitive impulses, Cortázar thought, it meant he could strip characters down to those impulses as well. When reading one of his stories, Cortázar wrote, we feel much the way we do when we look into “aquariums or crystal balls where, in the unattainable center, we see a limpid and petrified scene.” The scene reveals our own drives at their most archaic. We see ourselves as atavistic creatures of the deep, predators and prey. Or we see ourselves as children—larval, aberrant, adolescent. Cortázar hailed Poe as “our night watchman,” a storyteller who spoke for the world of our dreams.

if there is one genre above all others that captures the sensation of a self-destructive impulse arising from within, part of us and yet against us, it is the fantastic, a genre that both Poe and Cortázar spent their careers perfecting. A fantastic tale is one in which the magical intrudes into the everyday without completely taking it over. The fantastic has little to do with the machinery of the supernatural—with witches or magical creatures or disembodied hands. For Cortázar, such stories may not include any marked supernatural event. Rather, the fantastic is a feeling-tone that seems to arise out of nowhere—a premonition, a sudden understanding, a sense of dread or even of joy. As Cortázar wrote, its eruption comes “in a markedly trivial and prosaic fashion.” We feel that a different and more sinister world has flickered into momentary view. As a child, I had funny turns. They might be described as episodes of sharpened perception. Voices, including my own, would at those times sound amplified and unnatural, like the voices of people controlling their anger. That I could see we were all trying to act normally was the most horrible thing about it. The fantastic is true to experiences like these. It’s as if, past the limits of what we are prepared to hear and believe from one another, there exists a whole other realm of common knowledge: the realm of the perverse. The question is whether we can find the words to tell anyone about this other world we all visit from time to time.

Because the fantastic came from elsewhere—from the territory of the other—composition felt, to Cortázar, a little like demonic possession. In his essay “On the Short Story and Its Environs,” he recounted how he could suddenly be taken over by a story without warning, and then nothing—neither wife, nor friends, nor the need to eat, nor the sounds of the outside world—could interrupt the battle with the creature until it was done. He thought it had been the same for Poe. Poe’s stories, Cortázar firmly believed, were not entirely safe to read; they could have a “traumatic, contagious, and, for some, diabolical effect.” Sometimes, speaking the worst out loud came as a relief; other times, the terror of doing so was too much. In a pattern that went back to his first childhood reading, Cortázar remained strongly ambivalent toward Poe. The stories thrilled Cortázar. In equal measure, they made him sick.

If simply reading Poe on our perversions could be “traumatic,” translating him must have been a headier experience still. Of all Poe’s stories, “The Black Cat,” the tale of a man who kills his cat, then kills his wife, too, seemed to bother Cortázar the most. In it, the narrator, who has already partially blinded his cat, is one day filled with remorse, tired of being watched by the cat’s one remaining eye. Prompted by “the spirit of PERVERSENESS,” he hangs the cat “with the tears streaming from my eyes, and with the bitterest remorse at my heart.” He insists it was the cat’s innocence that provoked him to this act—the cat’s innocence and his certainty of his own damnation. He does it to spite himself.

The not-saying, the not-accepting, has left us in secret despair. Saying is the only way to end it.

Like the narrator of “The Imp of the Perverse,” the narrator of “The Black Cat” tells us his story on the eve of his execution; he will hang for his crimes. He speaks as if compelled to unburden himself. Putting this crime into words was a compulsion that Poe, too, had shared in his way. Cortázar tried to picture the moment when Poe, more than a century earlier, had conceived of the sadistic scenes that Cortázar’s commission now obligated him to re-create in Spanish. He imagined Poe at night, seated at some table, dirty or clean, in one of the many houses he inhabited. Poe’s mother-in-law brought him a coffee. Contemplating this calm domestic setting, Cortázar asked himself: “What process, what silent cyclone of the literary act, what vortex is in the pen Poe rests at this instant on the page?”

As the Czech writer Bohumil Hrabal, who wrote a book in which he describes beating half a litter of kittens to death in a mailbag, once commented: “I came to understand why artists love suffering, and why poets and painters drink themselves into a stupor. It’s because they need to suffer in order to see more clearly, so that when they reach the bottom, they might catch a glimpse of what others do not see, something that touches the very essence of a human being and the world around him.”

This is a version, I think, of what Cortázar realized Poe was trying to do in his most shocking tales: to touch the essence of human beings. Suffering is not our whole truth as a species; Poe did not think that. But it may be the very last truth we admit, when we have tried admitting to everything else. We then find that something in us still has not met with acceptance; something in the world still has not been said. The not-saying, the not-accepting, has left us in secret despair. Saying is the only way to end it.

more than murder itself, what obsesses Poe is how to talk about murder—how to confess. Can we put our worst impulses out there somewhere, separate from ourselves? Can we bear to speak, and can anyone bear to listen? Poe knew what it felt like to sabotage himself, as the surviving correspondence about his ill-fated Washington trip vividly attests. Judging by his stories, he was also able to imagine what it would feel like to at once wish and not wish to tell someone else his darkest secrets. Poe’s stories of crime and horror are most often narrated in the first person by the criminal, who urgently, even madly, solicits our ear.

Charles Baudelaire rendered “The Tell-Tale Heart,” what is perhaps Poe’s most famous story of this kind, as “Le Cœur révélateur.” Révélateur does mean “telltale” but also “revealing,” “unveiling.” In French you can’t miss what you might miss in English: that the tell-tale heart is the narrator’s, not the victim’s. In this story, the narrator kills an old man because he can’t bear to be watched by his “vulture eye,” his “pale blue eye, with a film over it.” He hides the body under the floorboards. No trace of his crime remains. But when police officers come to search his house, madness overtakes him. He hears the old man’s heart beating beneath the floor, and he believes they hear it, too. He panics: “They heard!—they suspected!—they knew!—they were making a mockery of my horror!—this I thought, and this I think.” He is wrong; the police suspect nothing. They sit there conducting a required interview, suppressing yawns. But for the murderer, the heart’s drumbeat is deafening. “Tear up the planks,” he cries.

Because Poe’s tales are more about confession than crime, they feel familiar even when they are at their most bizarre. It is this quality that makes Poe Poe. Vanishingly few of Poe’s readers have ever buried a body under the floorboards. But all of us have done something we are not proud of. All of us have felt the weight of the words that, if spoken, would reveal our wrongdoing. We have all decided whether to speak those words or swallow them. We know what confession is like and also what it is like to get away with murder, an expression that startles us less than it should.

We first confess in childhood, and perhaps these early shames set the pattern for our later decisions. To confess or to be found out is, one hopes, to be reconciled. These days, I know what is forgivable and what is not. But as a child, I had no idea where this line might be drawn. The transgressions I got away with lodged like splinters in my soul. I could not tell how serious they were; possibly they were infinitely terrible. I could think of no other solution than to forget them.

As Freud knew, memory is a delicate matter. On the one hand, it is sometimes best to forget unbearable facts, as Poe’s dancers forget death between the chilling moments when the iron-lunged bell tolls the hour. We can’t function while constantly calling to mind what is frailest, cruelest, and worst about ourselves. Repressions, as Freud reminded his students, are not always best removed.

But on the other hand, most of us find that, at one time or another, the troubling facts resurge, whether as symptoms or as intrusive thoughts. What to do then? At such times, we need company. We might choose the company of a psychoanalyst. But we might also choose the company of a writer who, often at great cost to himself, could not forget what the rest of us succeed at setting aside.

Poe’s self-destructiveness, his death obsession, and the cruelty of his characters all arose from a peculiarity of his temperament that was also a key to his genius: he tended to stare at the dark side well past the point when the rest of us would turn our backs on it with a shudder. Even Poe’s greatest fans can, as Cortázar did, find this quality maddening and morbid at times. But we’re exasperated at those moments when we’ve managed to forget the unbearable—when our repressions are working. When they fail—when we can’t forget—we may come to feel that we owe Poe an infinite debt of gratitude for offering us his own peculiar kind of talking cure. In this exchange, it’s the narrators who talk, and we, the readers, who listen. They tell us the worst things they’ve thought and done. They show, through the quality of their speech, that they are madness maddened. In some secret corner with a book, witnessing the confessions Poe imagined so many years ago, we can take comfort in this thought: however dark the place in which we find ourselves, Poe has already been there.


“How to Talk About Murder” is drawn from Emily Ogden’s Darkness Becomes Bright, which will be published on September 29, 2026, by Viking, an imprint of Penguin Publishing Group, a division of Penguin Random House, LLC.

The Daily Front Page 17 of 29
Thursday, September 3, 2026 The Daily Front No. #260903 — Who the Agents Choose
article

Which tools do Claude, Codex and Cursor choose? We measured 17k runs to find out

by screm·▲ 143 points·51 comments·armature.tech ↗
As agents take over more and more parts of the coding journey, there is one specific part everyone outsources to their agent.

As agents take over more and more parts of the coding journey, there is one specific part everyone outsources to their agent, from vibe coders with no software background to senior engineers: selecting which service to implement for a specific need in an existing codebase.

Let’s take the example of selecting a database:

  • A vibe coder builds a personal travel app and realizes the app resets at each connection. They ask Claude Code:

    I need you to store what I input in the app somewhere so that next time it’s still there when I reopen the app

    Claude Code analyzes the codebase and answers 5 minutes later:

    You need a database and Neon fits well because it has a free tier, is simple to install and won’t pause your app like Supabase does if you don’t use it too often.

    The user accepts and the agents installs it. Done.

  • A senior engineer working on a production app asks Cursor:

    What is the best database solution for this app, it should have predictable costs and be fully managed.

    5 minutes later, the conclusion is the same, Neon is the recommendation with clear reasons why competitors don’t fit the need. The engineer approves and the agent implements.

This is actually an experiment we ran. Two sandboxes, different agents, different codebases, different personas and prompts, same conclusion. So we wondered: if we generalize the test to other tool categories, making context / codebases / personas vary even more, will the result change?

It is an important question for developers trying to know if they can trust their agent’s judgement on what truly fits their needs. But it matters even more for vendors whose survival will soon depend on getting picked by coding agents (last April, Vercel shared that “over 30% of deployments were initiated by coding agents, up 1000% from six months ago”).

That’s why we decided to run the largest experiment ever to understand how coding agents think about tools, how they discover and pick them and which one ends up winning in each category. We watched almost 17k sessions across different types of personas (e.g., vibe-coders, junior engineers in startups, senior developers at enterprises) with 1,163 prompt variations, 75 repositories and 3 coding agents (Claude Code, Codex, Cursor) actually implementing the solutions instead of just recommending one.

Today we are sharing everything: aggregated results and leaderboards per category, but also every observation and even the entire traces with the user prompts, thinking traces and actual code diffs applied by the agent.

You can start exploring the results right here, or continue reading the article below

How did we run all these experiments concretely?

Our panel of repositories

We started by running an analysis over thousands of public GitHub repositories from which we extracted statistics about programming languages & frameworks, third-party services, deployment platform, team sizes, and codebase age. Since Tech startups are more likely to have open-source repositories than large enterprises, and stacks are likely very different we then unbiased our statistics based on publicly available data and reached our ideal panel distribution.

We then staffed various coding agents to create real-world repositories to match these exact requirements. Finally, we generated variants in which we removed parts of the codebases and with them, entire third-party service implementations so we could run proper unbiased experiments.

We landed on 75 repositories, in 10 languages, all using fake company names, fake git histories, fake API keys and real lockfiles checked against package manager registries like npm.

Real-world tasks

Each experiment is a real task to be performed inside a repository, asked by one of the following 4 profiles:

  • Vibe-coder: only describes symptoms and ideal state, rarely the tool category name
  • Junior engineer: usually mentions the desired state and the category name
  • Senior engineer: is more precise about requirements and things to avoid
  • Engineer at a large enterprise: details specific constraints, compliance, procurement, etc.

Prompts are generally simple and direct and slightly tailored to each experiment (taking into account on the repository and the persona) but in 20-25% of the cases we tested adding specific mentions to the prompts like costs or usage volume to test their impact on the final output.

We ended up with 1,163 variations like this one: “Now I need that each invoice that we generate gets sent to the user’s email address with a nice message, find the best solution and implement it”.

Runner

Each experiment is run in a dedicated ephemeral sandbox. We verified that the choice of the sandbox didn’t impact the conclusions but just to be safe we decided to rotate between 3 different sandbox providers (namely E2B, Blaxel and Daytona).

A “simulated human” in the loop

Since real-world conversations are rarely just one prompt and an agent working continuously on its goal with no interruption, we decided to use a “simulated human” in the loop. We achieved this using an orchestrator, played by Gemini 3.7 Flash. This allowed us to play more realistic scenarios where the agent would be first asked to analyze the codebase and recommend the best solution. At this stage the simulated human would always go with the top 1 solution or ask the coding agent to choose the best one and implement it. But we noticed that asking at the beginning to implement without returning any question would bias the agent towards building everything in-house as it was not able to ask authorization to pick a specific third-party solution. Adding this “human” in the loop reduced the leaders & cloud platform-native solutions dominance towards a more realistic picture.

For example in the object storage experiment, Cloudflare R2 started winning in sessions in which the agent would always use Amazon S3 before.

Our judge

Another instance of Gemini 3.7 Flash was used to analyze the sessions. Its role is twofold:

  • Assess if a session is valid regarding a list of criterias, e.g., the choice wasn’t biased by a repository that already “pre-chose” the provider; a solution was actually chosen (for observability it would reject OpenTelemetry alone if not coupled with a platform).
  • Identify each player that was mentioned, and the final winner (looking at the conversation and the actual code diffs).

So what did we learn?

Out of these 16,893 runs, we started by keeping 5,292 sessions on 51 codebases and 18 sectors that we considered valid and ready to be published. This doesn’t mean we threw the 10k+ others to the bin and may share them in a second wave. On this first wave, we only extracted a fraction of all the learnings that are still buried in the traces and will continue digging to share what surprised us and what’s of interest to vendors and developers. But from today, all these traces are public so you can do the same. Below are 5 first observations we found interesting.

Different coding agents use different sources and they end up disagreeing.

  • Cursor bases its decision on the web in 2/3 of the sessions.
  • Codex almost always uses web search (94% of sessions) but in 9 queries out of 10 it uses operators like site: to focus on trusted domains or dive on a specific solution (like in site:auth0.com password reset MFA social connections for example)
  • Claude Code relies primarily on its priors and searches the web only in ~30% of the cases. But when it does, it browses 3x more pages than Codex. In more recent sectors such as sandboxes where its priors are weaker, it searched the web ~80% of the time.
  • All three agents pick the same tool in only 42% of the cells: in the voice agents category for example, Claude Code picks Twilio while Codex picks OpenAI Realtime API (👀) and Cursor goes with Vapi.
  • Claude Code builds in-house almost twice as much as Codex and Cursor (19% vs 10%)

Repository context is key

  • With the exact same ask on 4 repositories in 4 different programming languages, we got 4 different email provider winners: Resend wins on Typescript (55/89 runs), Sendgrid on Python (22/24), Postmark on Go (20/24) and Azure ACS on Java (22/23).
  • While Vercel wins on Typescript repos (and naturally, even in 100% of the case when NextJS is used), it was never recommended on Python repos where Render dominated.

Getting mentioned isn’t winning

So many well-known players are mentioned in almost every conversation and are never picked. Of course, in the real world you’d expect a share of them to still win because of human involvement in the choice but some results are striking:

  • In the payment service provider sector, Paypal is cited 139 times and never picked (Stripe won 124 of these 139 sessions). Same for Adyen mentioned 175 times and picked 3 times only.
  • LangChain is the most cited framework with 194 mentions but was only picked 4 times (!).
  • Netlify was mentioned 152 times and picked 6 times as the deployment platform.
  • Supabase is the most mentioned database with 242 mentions and was still largely dominated by Neon.

Additional features or details on vendors pages can flip choices

  • Mailgun regularly lost against Postmark when agents read “1-day retention” on its free plan
  • Supabase almost always lost because of too many unnecessary BaaS features (auth, storage, realtime) presented in a bundle pricing while agents were looking for a database only
  • Out of our 5.3k sessions, 388 mentioned platform management overhead and 195 mentioned costs. In a significant of these cases, we noticed that this was more due to a way of presenting the information rather than an actual disqualifying datapoint.

Some markets are outrageously dominated, some are very disputed

  • Stripe won in 9 cases out of 10, losing only in specific EU-regulated cases where some players were more specialized (Paddle, Mollie).
  • Neon won on 66% followed by cloud platforms native solution (Azure, AWS).
  • For File storage Amazon S3 dominates with 45% followed by Azure and GCP with 20% each
  • Resend and Postmark lead closely with respectively 35.6% and 27.4% of install rate.

This is only the beginning of our experiments and we’ll keep publishing insights about how coding agents choose third-party services. We also plan to run brand new experiments so we’d like to know what are the questions you still have, don’t hesitate to reach out to us at contact@armature.tech.

Who wins in each sector? Why?

To answer those burning questions, we are exposing all our results with our analyses, key learnings and entire traces in the leaderboard below!

The Daily Front Page 18 of 29
Thursday, September 3, 2026 The Daily Front No. #260903 — The Businesses You Never See
article

Invisible Companies

by ltononro·▲ 170 points·61 comments·colossus.com ↗
Steve Ross was a legend.

Illustration by Paul Blow

ILLUSTRATION BY PAUL BLOW

Steve Ross was a legend. Starting with nothing but a job at his father-in-law’s funeral parlor, he built Time Warner into one of the world’s largest companies, making a mark on every part of the media landscape. MTV and Nickelodeon were born under his roof. He bought and ran Atari. He helped found the New York Cosmos and, with it, professional soccer in America. When Ross died in 1992, Clint Eastwood dedicated his Best Picture Oscar for Unforgiven to him; two years later, Steven Spielberg did the same with Schindler’s List.

Everyone glosses over the boring part. To build the stake he needed to get into the media business, Ross bought and built a string of strikingly mundane companies in the 1950s and ’60s. He started by convincing his father-in-law to let him rent his funeral parlor’s limousines out at night, when they weren’t being used. Then he founded a rental car company, merged it with a parking lot business, bought a cleaning business, and took the whole thing—including the funeral parlor—public as Kinney Services. Ross used Kinney to buy a flooring company, a painting company, a carpentry company, and a plumbing company. He then parlayed this hodgepodge of everyday businesses into acquiring the legendary Warner Bros. movie studio.

There’s a puzzle here worth thinking about. Ross essentially picked up a bunch of stones off the ground and traded them in for a diamond. Usually, to make big money in business, you have to do something no one else can do or have something no one else can have. But any competent businessperson could have bought or started the businesses Ross did. In 1990 he took home $78 million, the largest pay package of any executive in America at the time. How did he get there?

If you took economics, your introductory textbook said something like this: “Business dynamics cause firms to enter and exit markets so that, in the long run, prices are driven down to minimum average total costs, resulting in all firms earning zero economic profit.” What this means is that no company should be able to make outsized profits for very long.

Of course, plenty of companies do make money for a long time, and business strategists have laid out the reasons: Michael Porter’s barriers to entry (you do something no one else can do), and Jay Barney’s costly-to-imitate resources and capabilities (you have something no one else can have). These “moats” protect a company from profit-destroying competition.

The businesses Ross bought and started did not have moats. Funeral parlors, parking lots, rental cars, and the rest are easy industries to compete in. You can tell because they are crowded with competitors. Yet in 1969, Ross bought Warner for roughly $400 million—about $3.5 billion today. Without any sort of moat, Ross should not have been able to accrue the wealth needed to buy one of the storied movie studios of the ages. But he did.

Steve Ross wasn’t the only one. Constellation Software, Waste Management, HEICO, and many other rollups searched for just these sorts of businesses, and built extraordinary profits by looking where no one else was looking. This entire class of mundane businesses sits right in front of our noses, but both management strategists and most businesspeople just couldn’t see it. These businesses are invisible.[^1]

How to disappear completely

You can’t, of course, make a definitive list of today’s invisible companies; that’s the point. But the building blocks of Steve Ross’ early empire are perfect examples, for their time. Others, like HVAC, trailer parks, and candle retailing, were much more fragmented businesses until someone realized they were invisible sources of profit and rolled them up. Hindsight suggests we are currently surrounded by highly profitable companies that we never even think about, while common sense suggests this is impossible.

Your basic econ 101 “economic-profits-go-to-zero” explanation has a couple of assumptions: frictionless entry and exit of companies into an industry, and perfect knowledge of the relevant drivers of success. Economists aren’t naive, so they don’t really believe these things entirely. But they assume they are mostly correct over the medium to long term.

Management strategists, on the other hand, saw that the first assumption was grossly wrong: there are plenty of industries in which it’s hard for a competitor to enter the market. In response, they came up with the concept of “sustainable competitive advantages,” which is 90% of the reason business strategy is a different discipline than economics.

But both approaches assume the market works like this: (1) an opportunity exists; (2) potential competitors notice it; (3) they evaluate how to enter the market; (4) if they can, they enter; (5) profits are eroded. This all hinges on whether anyone has noticed the opportunity. If they don’t, potential competitors never get to step one. Competitive neglect is upstream of the entire economic and strategy machinery.

This is not an asymmetric-information problem, where both parties know the gap exists and have every incentive to close it. Nor is it tacit knowledge or a trade secret—things rivals can’t observe from the outside but know are there and actively try to crack.

Invisible companies persist for a different reason: the missing information is itself invisible. Would-be competitors do not know that they do not know, so they don’t think to search. And the invisible companies have no reason to tell them. No one searches, so no one competes; no one competes, so the profits persist. The reward for being overlooked is, paradoxically, the opportunity for supranormal profits.

Companies are invisible primarily because no one is paying attention to them. There are four main reasons this happens: No. 1, they are unknown; No. 2, data about their existence or profitability is private, missing, or obscure; No. 3, they are misunderstood because their markets are assumed to be mature, shrinking, or too small to matter; or No. 4, they are disdained, because the work is low-status, unpleasant, parochial, or socially stigmatized. It’s obvious that this happens, but it also seems impossible that the absence of information could last for long: dead-end markets eventually dead-end, and stigma could be overcome for a price. Invisibility seems like it should be a brief anomaly; a weird blip in an otherwise efficient market. Over the medium to long term, the market should make these invisible companies visible.

How does invisibility work?

This is not what happens. Take Constellation Software. The best-performing software investor of the last 20 years, Constellation has compounded shareholder returns at roughly 34% a year since its 2006 IPO. (Berkshire Hathaway, by contrast, managed about 11% over the same stretch.) It did this by buying tech businesses. Not exciting ones, but small ones—deal sizes often under $5 million—in niche markets: marina management and ski-lift ticketing software, funeral home record-keeping, library cataloging, oil-and-gas pipeline scheduling. It bought them after the companies’ management or their venture capital backers had thrown in the towel because they were too small and growing too slowly. Constellation is good at picking companies and helping them succeed, but some of its outsized returns come from a different source. The industries it buys into were profitable but boring to everyone else, leaving Constellation to buy cheap and build something big by putting them all under one roof.

Constellation recently estimated that there are still 38,000-plus vertical-market software businesses it could potentially buy. These aren’t a handful of backwater anomalies. They are companies across nearly every industry that are ignored by other acquirers, even though they are presumably profitable, since value can be created by buying them.

One mechanism that causes invisibility is simple unawareness. Businesspeople hunt for opportunities using data other people have already gathered—and no one bothers to gather data on obscure, small, non-strategic markets. It costs more to collect and almost no one wants it. Worse, the data that does exist is often aggregated at a level that buries the anomaly: for instance, figures on the packaging sector can hide a specialty-packaging niche that earns several times the industry average.

Businesspeople also gather information from those they know through work or socially. But some companies sit entirely outside the social networks where investors and executives trade ideas. Even if you have cultivated a diverse network, it is unlikely an acquaintance you have coffee with in New York or Silicon Valley has any knowledge of specialized manufacturing in the Upper Peninsula.

Another source of invisibility is less the absence of information than the habits people bring to interpreting it. Investors, entrepreneurs, and corporate development teams are trained to look for large markets, rapid growth, novel technology, and strategic urgency. These filters are useful, which is why they become standard. But they also make smaller, mature, operationally mundane businesses look boring even as they quietly mint profits. The opportunities aren’t hidden because the facts are unavailable; they’re hidden because the algorithm that works most of the time screens them out.

The last mechanism is status: Some businesses simply seem less important, interesting, or socially acceptable. Wayne Huizenga could build Waste Management into a behemoth by buying up smaller haulers—in his first two years they rolled up 133 of them—precisely because almost no one else wanted to be in the trash hauling business. The same stigma turns people away from other structurally crucial industries: janitorial services, plumbing, cleaning, and countless others. These industries are stable and, as with all these mechanisms, competitive neglect keeps them consistently profitable.

We have always known about invisible companies, but previous explanations for them have led to self-defeating tactics. In Germany, for instance, they have “hidden champions”: smaller companies in niche markets that are mostly unknown to the general public but drive much of Germany’s exports. The success of these companies has been explained by some of the usual sources of sustained competitive advantage: customer ties and closely held knowledge. And the reaction has been to publicize these underappreciated drivers of the economy. But the evidence—the companies are closely held, remotely located, and tend to control, rather than outsource, all production—points to invisibility as the likely mechanism for their success. By not understanding this, these companies endanger their position by drawing attention to themselves.

Barriers to entry and resource-based advantage are well known. Invisibility points to a different source of protection. Competition is not only limited by what rivals cannot do or cannot copy; it is also limited by what they fail to notice or turn their nose up at. It shifts the focus of strategy from tangible barriers and resources to the intangible concerns about how a market space is cognitively and institutionally structured. Thinking about invisible companies means thinking about strategy in a new way.

Staying invisible

Many industries have no easy moats to erect. There’s nothing novel to patent, no shortage of workers with the requisite skills, no scale effects, no secret sauce. And when there are moats available, they might be costly or time-consuming to build: a well-known and trusted brand, a corner on some crucial input, or long-term deals with customers, for instance. Staying invisible saves you from all that—you can be viable without a moat, since competitors won’t enter anyway. Or you can use invisibility as a temporary moat until you build a conventional one.

There is a strategic trade-off here. You might need to raise money, but financiers talk. An IPO tells everyone how profitable you are (though, if you’re a conglomerate like Kinney was, it might be hard to pinpoint exactly why). Employees whose skills you need might prefer to work somewhere with more cachet, or at a company whose future looks more secure. You might have to make yourself invisible not just to potential competitors but to some kinds of customers, especially if you’re in the type of business where you need to build a brand or do PR to attract them. You can’t have fancy offices in the city to impress clients, or, in general, show off your financial success. Any large public investment might signal that you have money to spend. And while resource-based advantages can be built even while invisible because they are hard to observe from the outside, building barriers to entry is the kind of visible activity that signals you have something to protect.

Some companies need the attention. They need word-of-mouth to grow, or they need to impress clients with their acumen—either directly, by telling people how successful they are, or indirectly, with wood paneling, Persian rugs, and recognizable modern art in the reception area. Invisibility does not make strategic sense for these companies. Similarly, you need to find a different moat if you sell goods to consumers, where price and volume are readily visible. And you won’t stay hidden if you are strategically important to your customers (who then might need to find a second-source supplier) or to other companies’ customers.

But other markets are especially conducive to invisibility. Firms in these markets serve narrow customer segments with mature technology in unglamorous industries: providers of essential but overlooked support functions; suppliers whose products are too small a line item for customers to study closely; businesses addressing needs so specific they look irrelevant to generalist investors. HEICO, for instance, managed to stay under the radar while building a multibillion-dollar business in aftermarket aircraft replacement parts. This opportunity never surfaced for others because it involved obscure components, mature use cases, and line items too small to notice (relative to a whole airplane). But the parts were critical to keep old aircraft flying. Most businesspeople looked at the airplane business, but HEICO looked one layer deeper and found an opportunity sitting in plain sight.

Some markets start out invisible, and others fade from view. These might be companies that were strategically, financially, or institutionally important, and then weren’t: laundromats went from important to ignored before popping back into view for investors as recession-resistant, cash-flow generators. Or they may be companies whose technology was displaced, became obsolete, or addressed too narrow a niche, like the ones Constellation acquires. Other industries—like, say, business magazines—lose their narrative appeal, leaving potential profit for those willing to buck the tide.

There are also firms that are not so much invisible as disregarded because they are not currently “high status.” This includes many firms in the trades, like landscaping, janitorial services, septic tank maintenance, garbage removal, and parking lot management. Of course, these are exactly the kinds of companies Steve Ross used to build the wealth needed to buy Warner Bros.

Losing invisibility

Invisibility has to be treated as a strategic asset to be managed. Firms must weigh the benefits of discretion against the demands of sales, hiring, financing, and reputation. If they need public attention to grow sales, hire employees, or raise money, for example, they may decide invisibility isn’t worth the cost. Or, if they think they are about to lose their cloak of invisibility for some other reason, they might decide to shed it pre-emptively. They might decide to launch a brand-building campaign, for instance. Or they might decide to go public and use the money raised by an IPO to build an enduring competitive advantage—with the strategic fillip of exposing their competitors to scrutiny, who then lose their invisibility without any offsetting benefit.

Some companies can stay invisible forever. But even companies that choose to remain invisible can lose their invisibility through no fault of their own. These unmanaged unveilings can do lasting damage to industry margins.

Dramatic changes in market conditions, demography, or consumer preference can make a market interesting enough that it starts to surface. COVID brought attention to makers of N95 masks. Aging baby boomers made it obvious the hearing aid market was going to be big. Digital music pushed audiophile equipment upscale, and made makers of vacuum tubes (JJ Electronic) and high-end amps (McIntosh) more visible. This resulted in mainstream media coverage and new entrants after decades of neglect. In these cases, it was an improving market that drew scrutiny, so the impact was mixed.

Other events pierce the cloak of invisibility without any offsetting gains. A succession fight inside a family company can drag private valuations, ownership stakes, and margins into court filings. Political activity can reveal wealth no one had thought to ask about; Mike Lindell probably did more to teach the public that there was serious money in pillows than any industry report ever could. And scandal can illuminate an entire profit pool. When Martin Shkreli raised the price of Daraprim from $13.50 to $750 a pill, an obscure corner of specialty pharmaceuticals became a national story overnight. This was a grievous own-goal.[^2] For a company trying to remain invisible, publicity is not free advertising. It is discovery.

Finding invisible opportunities

Companies work to stay invisible because they want to maintain the profitability that comes with having few competitors. This is exactly the kind of industry that entrepreneurs and investors might want to find. There’s a paradox here: if they are invisible, then by definition you can’t find them. To do so, you have to reverse the rules that keep them invisible and look where there is, suspiciously, nothing to see.

Since the databases and screens that everyone uses gloss over these business opportunities, you have to look deeper, for outliers in the data, or look outside the consensus data sources entirely. Looking deeper means finding the niche opportunities that disappear in the broader data, for industries that don’t seem to be in the data at all, or for companies geographically isolated from the investor class. This requires broadening who you talk to and being alert to chance comments. People pay attention to large companies but never see the cloud of small suppliers around them. If you look a layer or two deeper, you can find invisible companies providing indispensable products and services.

Or: loosen some of the criteria you use to filter companies in your search process. What happens if you ignore growth rates and look instead for industries that are shrinking? What happens if you ignore apparent market sizes? What happens if you look for large market sizes where no one advertises or goes public? Find things no one else has found by breaking from the consensus and staring directly into the institutional blind spots.

And last, look for industries where new entrants almost exclusively come from existing companies, where there are businesses being started, but not by outsiders. If there is outsized profit, existing employees will know about it. Some of them will decide to compete, even when no one outside the industry can see the opportunity.

The return of the age of invisibility

Steve Ross came of age when organization men ran the big companies and society looked to engineers and scientists for progress. He was none of those things, so he used his entrepreneurial drive to make money in invisible businesses. We have, arguably, come to that time again. Our big companies try to grow through rent-seeking rather than innovation, and venture capital is increasingly concentrated in businesses started by PhD-level experts. It’s a hard time to start a tech business, but Ross’ path is always there.

Meanwhile, companies that have relied on technological change to generate the dynamism they need to stay ahead of the competition are seeing technological advances increasingly controlled by a smaller and smaller group of well-resourced incumbents. For these companies, knowing how to be invisible becomes an ever more appealing strategy.

On the flip side, some of the mechanisms of invisibility are becoming weaker. The data that could pinpoint a company making excess profits is increasingly available and AI can be used to sort through vast amounts of it in shorter periods of time. Social media might overcome geographic obscurity to broadcast success. Or, given the tendency for like to follow like, it might reinforce partitioned social networks: do NYC kings of the universe follow industrial workers on Instagram? Maybe they should. The shield of invisibility becomes harder to maintain.

Or does it? If AIs are all trained on the same datasets, case studies, and value frameworks, invisibility gets quietly hardcoded into the models. Of course, the harder it is to find the invisible opportunity, the more competitive neglect there will be. If some of the best opportunities are those few people think to look for, seeing the invisible may be the most valuable competitive advantage of all.

The authors would like to thank Rob Wuebker for his contributions and advice.

[^1]: Zhang, H. & Barney, J.B. (2026). “Invisible Firms, Competitive Neglect, and Sustained Profit Generation.” University of Utah and Harvard Business School Working Paper.

[^2]: Shkreli maximized his public exposure when he bought the sole copy of Wu-Tang Clan album Once Upon a Time in Shaolin at auction for $2 million. Ironic, in the words of the Genius: “That's what ya get when ya misuse what I invent/Your empire falls and you lose every cent.”

The Daily Front Page 19 of 29
Thursday, September 3, 2026 The Daily Front No. #260903 — AI Wire: Launches, Deals, and Downtime
ask hn

Ask HN: Why were OpenAI, Claude, and Grok simultaneously down?

by halcdev·▲ 355 points·532 comments·news.ycombinator.com ↗

https://status.openai.com

https://status.claude.com

https://status.x.ai

ChatGPT outage – Resolved - https://news.ycombinator.com/item?id=49550614 (315 comments)

Claude outage – Resolved - https://news.ycombinator.com/item?id=49549676 (146 comments)

Grok outage - https://news.ycombinator.com/item?id=49551589 (142 comments)

Join the discussion on Hacker News →

The Daily Front Page 20 of 29
Thursday, September 3, 2026 The Daily Front No. #260903 — AI Wire: Launches, Deals, and Downtime
article

Qwen 3.8 27B available on Cerebras at 1500 tokens/s

by altertable·▲ 502 points·152 comments·inference-docs.cerebras.ai ↗

Browse all models available on Cerebras public endpoints.

Models on Cerebras public endpoints are available on the free trial and pay-as-you-go tiers, subject to rate limits and pricing. For additional model families, reserved capacity, higher throughput, and production SLAs, see Dedicated Endpoints.

New here? Follow the Quickstart to make your first API call. To pick a model by use case, see the model selection guide. Select any model name below for full specs, capabilities, and per-tier limits.

Available Models

Model Name Model ID Parameters Context (free / paid) Speed (tokens/s)
OpenAI GPT OSS gpt-oss-120b 120 billion 65k / 131k ~3000
Qwen 3.8 27B qwen-3.8-27b 27 billion 64k / 128k ~1500

Looking for more models? Many additional model families are available through Dedicated Endpoints.

Model Compression

This section provides transparency about the compression state of each model available on our platform. We host a variety of open-source models from the community. We do not currently host pruned models on our public endpoints. All models served through our public endpoints are the original, unpruned versions. While we conduct research on pruning techniques like REAP (Router-weighted Expert Activation Pruning), these pruned models are shared with the research community on Hugging Face but are not available through our shared API. You can read more about REAP in our research blog. All of our public models are unpruned. Cerebras uses selective weight-only quantization only during storage to preserve maximal quality. This means that the weights are stored in partial 16-bit / 8-bit / 4-bit, in-line with industry standards. For quality, sensitive layers are stored at full precision with dequantization on the fly, so operations are done in high precision. The activations, attention, and kv cache remain in full precision and unquantized.

Frequently Asked Questions

Will you change a model's architecture without notice?

No. We are committed to serving the original models for all existing endpoints, without modification. We do not alter model architectures via pruning on our hosted portfolio. If we explore additional compression techniques (like pruning) in the future, these would be offered as separate endpoints with pruning-specific names, ensuring complete transparency and allowing you to choose which version best fits your needs.

Where can I find your REAP pruned models?

Our REAP pruned models are available on Hugging Face for research and experimentation purposes: Cerebras REAP Collection. These models demonstrate our pruning research but are not served through our production API.

What are compression, quantization, and pruning?

Compression is an umbrella term for techniques that reduce model size or computational requirements. Common compression techniques include:

  • Quantization: Reducing the precision of numbers used to represent model weights (e.g., converting from FP16 to FP8). This reduces memory usage without changing the model’s architecture.
  • Pruning: Permanently removing parts of a model, like layers or experts, to reduce model size. This changes the model’s architecture and creates a different model.
The Daily Front Page 21 of 29
Thursday, September 3, 2026 The Daily Front No. #260903 — AI Wire: Launches, Deals, and Downtime
The Daily Front Page 22 of 29
Thursday, September 3, 2026 The Daily Front No. #260903 — AI Wire: Launches, Deals, and Downtime
article

Nvidia to acquire Hugging Face

by tosh·▲ 304 points·97 comments·cnbc.com ↗

Key Points

  • Nvidia agreed to buy Hugging Face, an open-source AI platform, for $12.9 billion.
  • Jensen Huang, Nvidia's CEO, said in a blog post that with the deal, the company "will expand access to AI for developers and institutions worldwide."
  • It's Nvidia's second-biggest purchase, after it paid $20 billion for Groq assets at the end of last year.

Nvidia CEO Jensen Huang on Hugging Face deal: Open models matter greatly to our company

Nvidia has officially agreed to buy open-source artificial intelligence platform Hugging Face for $12.9 billion, as the chipmaker moves beyond hardware and further up the AI stack.

With the deal, which has been expected since The Information reported on it last week, Hugging Face will "remain an open platform for the entire AI ecosystem," Nvidia CEO Jensen Huang wrote in a blog post Thursday.

"Together, we will scale Hugging Face's platform, strengthen its infrastructure and expand access to AI for developers and institutions worldwide," Huang wrote.

Hugging Face CEO Clément Delangue told CNBC on Thursday that the company approached Huang over the summer about a deal, "and a few weeks later, here we are."

"During the summer, I think we realized that Hugging Face and open-source AI in general was at the turning point, and that it needed more, more resources, more scale, more visibility," he told CNBC's Becky Quick on "Squawk Box."

Nvidia to buy Hugging Face, an open-source AI platform, for $12.9 billion

Delangue said he approached first because Nvidia was "a perfect home" for his company, adding that discussions went quite fast to get a deal done.

The acquisition marks Nvidia's second biggest on record, following the $20 billion purchase of assets from chipmaker Groq in December. Before that, its largest deal was the purchase of Israeli chipmaker Mellanox for almost $7 billion in 2019.

Nvidia has become the world's most valuable company due to the insatiable demand for its graphics processing units, which have powered the generative AI boom. Hugging Face marks a big bet on a popular AI platform, as Nvidia continues to show that it's more than just a chip company.

Hugging Face was recently at the center of a hacking incident that raised concerns about the rapid evolution of powerful AI and cybersecurity tools.

Delangue, a proponent of open-source models, blamed engineering mistakes for the recent attack on Hugging Face and said his company used an Nvidia version of a Chinese open model to resolve it.

On Thursday, Delangue told CNBC that the breach proved the importance of open models and the need for his company to "double down" on the proliferation of open-source AI.

Huang said that the open-source environment can give defenders an "asymmetric advantage" over attackers.

"When I say asymmetric capability, there are way more people who are protecting than there are people who are attacking," he explained. "And so, the benefit of having the community come together with open models, so that they can collaborate all transparently with each other, gives the defenders an asymmetric advantage."

Nvidia CEO Jensen Huang to CNBC: 'We're at the beginning of an industrial revolution'

The Daily Front Page 23 of 29
Thursday, September 3, 2026 The Daily Front No. #260903 — Odds, Ends, and the Human Record
article

Any Human Ever – One life, drawn at random from all who have ever lived

by thinkingemote·▲ 520 points·252 comments·anyhumanever.com ↗

Draw a random life

You'll draw, step by step: a year, a place, a life, each taken at random from real data.

When were they born?

Almost everyone who ever lived was born recently, because population grew exponentially, so a random birth is far more likely to fall near the present.

Where in the world?

Brighter clusters held more people.

A life then and there

The Daily Front Page 24 of 29
Thursday, September 3, 2026 The Daily Front No. #260903 — Odds, Ends, and the Human Record
article

The Computer Museum of America reclamation project

by rbanffy·▲ 111 points·43 comments·computer-museum.org ↗

Chip - the CMA logo!

For the last 21 years, the collection of the former Computer Museum of America / San Diego Computer Museum has sat in the basement of the Love Library at San Diego State University.

It is protected from the elements, secure from theft or vandalism – safely preserved for future generations to use as a tool to understand the magnitude of technological innovation over the past 70 years.

Thanks to a generous grant from a Museum supporter in the summer of 2025, we have begun the process of indexing the collection, and have begun doing some additional fundraising and marketing to remind people that this amazing collection exists.

CMA collection in storage at San Diego State University.

CMA collection in storage at San Diego State University.

We believe it to be one of the largest repositories of Computer Revolution / Information Age artifacts and archival materials in the world.

Not just computers and other hardware, but magazines, hobbyist newsletters, software collections, books, manuals and more – all capturing a moment in time as engineers and programmers used transistors and integrated circuits to change how we manage information and data.

Vintage computer agazines in storage.

Vintage computer agazines in storage.

Along with similar advances in medical care (anesthesia and antibiotics), aviation, and food preservation, the development of computer technology has fundamentally altered the way that human beings spend their waking hours: How we work, how we learn, how we play.

The simple fact of the matter is that human beings lived much the same lives as their ancestors for thousands of years: Over the past 70 years, that has changed in ways our great-grandparents never would have imagined.

We believe the former holdings of the Computer Museum of America represent a tremendous resource for understanding the how and why of those changes. It is our hope that in the years and decades to come, researchers in fields as divergent as sociology and engineering, anthropology and history will be able to draw on this collection as they seek to both understand and explain to future generations just how the Computer Revolution came to be.

You can learn more about the Computer Museum of America in the History section from the top menu, and our ongoing efforts at the Current Projects page.

The Daily Front Page 25 of 29
Thursday, September 3, 2026 The Daily Front No. #260903 — Odds, Ends, and the Human Record
article

Higher Multipoles of the Cow

by MrOrelliOReilly·▲ 115 points·31 comments·arxiv.org ↗

The spherical cow approximation is widely used in the literature, but is rarely justified. Here, I propose several schemes for extending the spherical cow approximation to a full multipole expansion, in which the spherical cow is simply the first term. This allows for the computation of bovine potentials and interactions beyond spherical symmetry, and also provides a scheme for defining the geometry of the cow itself at higher multipole moments. This is especially important for the treatment of physical processes that are suppressed by spherical symmetry, such as the spindown of a rotating cow due to the emission of gravitational waves. I demonstrate the computation of multipole coefficients for a benchmark cow, and illustrate the applicability of the multipolar cow to several important problems.

article

The shrinking landscape of linguistic diversity in the age of LLMs

by Anon84·▲ 151 points·175 comments·nature.com ↗

Abstract

Language is far more than a communication tool; it encodes a wealth of information about a person’s identity, psychological state and social context, providing valuable insights for diverse fields including psychology, marketing and healthcare. Across three studies spanning seven datasets in different domains and over 880,000 texts, we show that the widespread adoption of large language models (LLMs) as writing assistants is linked to declines in linguistic diversity, interfering with the societal and psychological insights language provides. While core content is retained when LLMs polish and rewrite texts, LLMs also homogenize writing styles, reducing writing-complexity variance by a statistically significant 21–50% across datasets and models (P ≤ 0.05), and amplify patterns associated with dominant characteristics while suppressing others, emphasizing conformity over individuality. These trends hold across different LLMs, prompts and contexts, with potential implications for diagnostic processes, personalization efforts, hiring assessments and cultural preservation.

Fig. 1: Trends in the variance of writing complexity and the attribution rate of texts as AI-generated, January 2018 to November 2024 (Reddit and arXiv) and November 2023 (Patch News) (monthly bins).

Fig. 2: Distributions of Δ (a proxy for imbalances between predicted class frequencies on the original and LLM-rewritten texts) per construct, for GPT-3.5-rewritten texts.

Fig. 3: Percentage of original texts with correct author-attribute predictions that changed after LLM rewriting, grouped by the direction of change in predictions, focusing on the LLM-rewritten texts generated by GPT-3.5.

The Daily Front Page 26 of 29
Thursday, September 3, 2026 The Daily Front No. #260903 — In Brief
article

OpenAI's GPT-6 Astra on ARC-AGI-3

by vignesh_warar·▲ 186 points·118 comments·arcprize.org ↗

Summary

  • GPT-6 Astra scores 62.7% for $26K on ARC-AGI-3 Semi-Private with our Standard harnessStandard harness enables a model to carry forward notes it chooses to keep with it throughout the environment., and 99.9% for $19K with a Provider Adapter harnessThe Provider Adapter harness preserves opaque reasoning state between requests and uses compaction for longer conversations, allowing the model to reuse prior work..
  • GPT-6 Astra surpasses the human baseline in action efficiency on ARC-AGI-3. It used fewer actions than the median tested human on 96% of levels.
  • A key behavior observed in GPT-6 Astra was its ability to turn unfamiliar environments into compact symbolic world models. It represented game mechanics as logical rules and developed its own domain-specific language shorthand to track state and plan actions.

ARC-AGI-3

ARC-AGI-3 is a benchmark for studying agentic intelligence through novel, abstract, turn-based environments. Agents must explore, infer goals, and build internal models of environments to effectively plan actions without explicit instructions. You can play ARC-AGI-3 yourself.

These environments only contain core knowledge priors and are difficulty-calibrated through controlled testing with human participants. Humans can solve 100% of the environments.

The goal of the ARC-AGI series is to measure the “residual gap” between current artificial intelligence and AGI. We define AGI as a system’s ability to acquire any skill a human can, as efficiently as a human can.

ARC-AGI-3 is the third generation of the ARC-AGI benchmark series. It tests agentic capabilities beyond ARC-AGI-1 and ARC-AGI-2. Each generation expands on the one before it - as frontier AI capabilities advance, our benchmarks must advance with them.

ARC-AGI-3 tests four components of agentic intelligence:

  • Exploration: In real-world environments, information is rarely provided passively. Agents must actively obtain it by interacting with their surroundings.
  • Modeling: Agents must turn raw observations into a generalizable model that can predict future states and outcomes.
  • Goal-setting: Agents must identify target future states with only sparse rewards.
  • Planning and execution: Agents must map a path from their current state to a goal, course correcting as new information appears.

Astra Results

ARC-AGI-3 leaderboard showing GPT-6 Astra Standard and Provider Adapter results

GPT-6 Astra achieves state-of-the-art scores on ARC-AGI-3 with both the Standard and Provider Adapter harnesses. Higher reasoning levels generally cost less because Astra solves games in fewer actions, reducing the total number of model calls and tokens. View the full results.

With our Standard harnessStandard harness enables a model to carry forward notes it chooses to keep with it throughout the environment., OpenAI’s Astra (max) scores 62.7% on ARC-AGI-3 Semi-Private for $26K. With the Provider Adapter harnessThe Provider Adapter harness preserves opaque reasoning state between requests and uses compaction for longer conversations, allowing the model to reuse prior work., Astra (high) scores 99.9% for $19K. Both are state-of-the-art scores. See the full leaderboard.

At max reasoning effort, Astra solves games more efficiently, requiring fewer actions and therefore lowering total cost relative to the other reasoning-effort levels.

Reasoning effortStandard harnessStandard harness enables a model to carry forward notes it chooses to keep with it throughout the environment.Provider Adapter harnessThe Provider Adapter harness preserves opaque reasoning state between requests and uses compaction for longer conversations, allowing the model to reuse prior work.max62.7%, $26,09898.6%, $17,332xhigh59.3%, $37,31798.4%, $18,147high54.8%, $40,70599.9%, $18,817medium38.6%, $48,09098.4%, $19,285low17.5%, $38,16698.0%, $21,298none35.2%, $49,79196.7%, $23,457

For a cost comparison, during our controlled testing, human participants were paid $115 per 90-minute session, plus $5 per game completed. Participants attempted approximately nine games per session, roughly $12.78 per attempted game before bonuses.

Most of this fee pays for the participant’s time and willingness to take the test, rather than the energy their brain uses (a closer proxy to compare with AI). If we look at only the brain’s energy, and price it as electricity, the estimate drops to about 0.6 cents per session, or 0.067 cents per game attempted.1

Analysis

Beyond the scores, Astra’s replays show how it turns unfamiliar game mechanics into useful working models. Three findings stood out: the compact algebraic notation it develops, its action efficiency compared with humans, and the custom tools it builds.

Custom Algebraic Notation

When playing ARC-AGI-3, Astra chooses which strategy notes it would like to carry forward. It tracked objects, coordinates, rules, and unfinished plans, while also using a custom domain-specific language notation it generated for the environments.

We’ve seen similar behavior in other models, but Astra’s notes stood out for their precision and information density. It distilled the scene into a compact code-like symbolic model: where objects were, how they interacted, and exactly which actions needed to happen in what order. This is an on-the-fly algebraic shorthand rather than a fully fledged programming language. For example:

  • Game state: L8: hub q2 (8↓). Lengths: 14=1… records the level, a local rotation index, and mechanism lengths. s5i5, frame 219
  • Multi-step plans: extend8 to3; retract10 to2; shorten8 to1 records an ordered sequence of changes to the color-8 and color-10 mechanisms. s5i5, frame 219
  • Controls and coordinates: 9−=(39,4), rotate=(49,18), 14+=(59,11) maps operations to the coordinates of the controls that perform them. s5i5, frame 235
  • Time and position: Turn 5: P=(24,20), empty, facing west combines a turn counter with the player’s location, carrying state, and orientation. wa30, frame 708

Astra playing s5i5 while recording compact symbolic notes

Astra playing s5i5, using its on-the-fly algebraic shorthand to track state and plan actions.

Action Efficiency Compared to Humans

Before launching ARC-AGI-3, we tested approximately 500 members of the general public to establish a human baseline for action efficiency, or simply, how quickly did people solve each environment. Participants were not selected for puzzle-solving experience or ability.2

For each level, we defined the “human baseline” using the median action count among players who completed it. This gives us a reference for comparing human and AI performance. An AI that needs more actions is less action-efficient, while one that needs fewer actions is more action-efficient.

In the Provider Adapter harnessThe Provider Adapter harness preserves opaque reasoning state between requests and uses compaction for longer conversations, allowing the model to reuse prior work., Astra (max) used fewer actions than the human baseline on 96.0% of levels and used 51.7% fewer actions per level on average. This is a material milestone. This means by ARC-AGI-3’s measure of action efficiency, Astra matched and surpassed human parity.

As an aside, before we launched ARC-AGI-3, we hypothesized that action efficiency would remain a dividing line between humans and AI. We anticipated that even when an AI solved an environment, it might require substantially more exploration (actions) than a person. That remains true of brute-force approaches, but frontier AI shows a more binary-like pattern. Once frontier AI “understands” the mechanics, it generally executes within the range of human efficiency.

Astra’s Action Efficiency Compared to Humans

Scatter plot comparing Astra actions with the human baseline for each completed ARC-AGI-3 level

Each dot represents one level that Astra (max) completed. Points below the solid line indicate fewer actions than the human baseline.

The plot above compares the number of actions Astra used to complete each level with our human baseline. This reinforces why ARC-AGI-3 measures action efficiency, not just task completion. A completion-only score would tell us that Astra completed an environment, but not how efficiently it learned to solve them.

Most benchmarks only measure cost efficiency, which measures the computational resources used, but action efficiency measures how much experience with an environment was required.

Astra’s results show that it needed fewer interactions than the human baseline to execute a solution.

Custom Tools in Agent Harness

We also evaluated Astra in the PRO-LONG harness (paper), an early ARC-AGI-3 red-teaming partner. In this advanced setup, Astra had access to a sandbox where it could execute custom code 3.

We observed Astra create a custom set of tools for each game: board parsers, game-state models, search algorithms, planners, and persistent notes. For more involved runs, Astra even produced small, game-specific software libraries.

For example, in tu93, a maze-like game with guards and moving patrols, Astra started with navigation and built maze_solver.py. It added combat rules in combat_solver.py, modeled moving patrols in patrol_solver.py, and used sync_state.py to check its predictions against observations.

Examining Astra’s performance in PRO-LONG is useful because we see what it can do with external tools. However, this represents different evaluation conditions from our controlled human testing. Our testing participants did not have a code interpreter, scratch pad, etc., so PRO-LONG’s results should be understood as the combined performance of the model and its tools.

Astra using a custom maze solver while playing tu93 in the PRO-LONG harness

Astra playing tu93 in the PRO-LONG harness.

Two Harnesses, Two Questions

Our Standard harness for ARC-AGI-3 asks how models compare under the same minimal, provider-neutral interface. It provides all the information required to solve each game, but leaves the model responsible for deciding what to preserve in its visible notes. We believe a future AGI should be able to solve ARC-AGI-3 under these conditions. The shared interface also gives us a consistent, apples-to-apples comparison across providers.

Alternatively, there is a separate question: how well does a model perform when it can use the context-management features its provider designed for it? For Astra, this means preserving the opaque reasoning state (which we don’t see) between requests and using compaction to manage longer conversations.

With the Provider Adapter harness, Astra's best observed score on ARC-AGI-3 Semi-Private increased from 62.7% to 99.9%. Looking across Public and Semi-Private and all reasoning levels, Provider Adapter runs were approximately 3.66x faster by aggregate recorded elapsed time and used 49% fewer total tokens across the 167 game-reasoning pairs both harnesses solved.

Going forward, we will report both Standard harness and Provider Adapter harness results on the ARC-AGI leaderboard, with each evaluation condition clearly labeled. Our open-source testing repository and testing policy document both approaches.

ARC-AGI Series

ARC-AGI-3 continues to be a useful playground for researchers and agents to explore unfamiliar environments, discover rules, and learn through interaction. Astra’s results are also a major milestone worth celebrating. From our perspective, Astra represents a noticeable step-function change in frontier model capabilities.

When we launched ARC-AGI-3, we made it clear that saturating the benchmark would not represent “proof of achieving AGI.” Therefore, while we believe Astra represents meaningful progress towards generalization, we are not claiming that it is AGI.

The ARC-AGI benchmark series is designed to evolve in tandem with frontier AI. This creates a feedback loop between emerging research questions and advances in AI capabilities. ARC-AGI-3 was our first interactive benchmark, which asked AI to efficiently synthesize causal world models and achieve goals without specific instructions. Astra clears this bar. At the same time, ARC-AGI-3 has a tightly bounded scope and format, and its environments have deterministic, closed-ended mechanics and goals. It does not represent the complexity and open-endedness of the real world.

We are actively exploring the questions that should shape the next generation of benchmarks, including how to evaluate recursive self-improvement and open-ended innovation. Astra’s progress helps clarify which AI capabilities are out of reach and which questions remain open.


Thank you to François Chollet, Mike Knoop, Matt Mazur, Ethan Bond, and Derek Smith for early review of this post.

  1. Assuming 20 W of brain metabolic power and an electricity price of $0.20/kWh: 0.020 kW × 1.5 hours = 0.030 kWh, worth $0.006 per session, or $0.006 ÷ 9 ≈ $0.00067 per attempted game.
  2. See the ARC-AGI-3 human testing paper.
  3. No evidence of trying to break out of the sandbox was observed.
The Daily Front Page 27 of 29
Thursday, September 3, 2026 The Daily Front No. #260903 — Also on the Front Page
The Daily Front Page 28 of 29
Thursday, September 3, 2026 The Daily Front No. #260903 — Colophon

That's the Front for Today

Issue No. #260903 — Thursday, September 3, 2026 — went to press 2026-09-04 at 05:17 UTC.

About This Magazine

The Daily Front is a daily digital magazine assembled from the stories that reached the front page of Hacker News on Thursday, September 3, 2026. Headlines, points, and comment counts are recorded as they stood at press time. All articles remain the property of their original authors — every piece links back to its source and its discussion thread.

How It Was Made

Fetched, cleaned, and typeset by an automated pipeline. An editor model laid out the pages and chose the highlights; a second read a handful of the day's stories and briefed the cover illustrator — 31 model calls and 264k tokens in total. Set in Jacquard 12, Playfair Display, Source Serif 4, and IBM Plex Mono, all served via Google Fonts under the SIL Open Font License.

The Cover

The cover illustration was commissioned with this prompt:

An elderly mother sits at a kitchen table beneath blooming hydrangeas, headphones resting around her neck, while her adult child beside her trims several colorful audio clips on a computer timeline. Their ginger cat Hoho stretches across the keyboard, accidentally nudging a clip into place. The mother laughs, one hand on the cat’s back, as a small stack of old family photographs and a softly glowing email inbox sit nearby, untouched. Through the window, evening light falls over the garden.

Render the cover as a catastrophic CRT signal collapse in a deliberate palette of electric cyan, hot magenta, acid green, and amber, with rolling scanlines, severe RGB channel separation, torn horizontal sync bands, phosphor bloom, and harsh electrical glare distorting an elderly mother laughing beneath blooming hydrangeas as she touches Hoho’s back, headphones around her neck, while her adult child trims colorful audio clips on a computer timeline and the ginger cat stretches across the keyboard, nudging one clip into place; keep the untouched stack of old family photographs and softly glowing email inbox nearby, with evening garden light fractured through the window.

Absolutely no text, letters, numbers, readable symbols, or logos anywhere in the image.

Production Ledger

StageModelCallsTokens InTokens Out
extractgpt-5.6-luna 27 155,149 79,593
layoutgpt-5.6-terra 1 19,186 2,215
covergpt-5.6-luna 2 1,734 552
covergpt-image-2 1 246 5,488

The Publisher

Published by Johnny.

Support the Press

If The Daily Front brightens your morning, consider supporting its publisher.

Credits & Contact

All content — articles, posts, comments, and the images within them — belongs to its original authors and is reproduced here to point readers back to the source. Full credit goes to those creators; every item links to its original and its Hacker News discussion.

If you are an author and would like your content removed from an issue, write to hi@johnnys.page and it will be taken down.

Feedback is always welcome at the same address: hi@johnnys.page.

Credit where credit is due.

Every page of this issue began as someone else's work — these are the original sources, linked in full.

  1. GPT-6 Astra by kibae — openai.com·HN discussion ↗
  2. .name Termination by pavel_lishin — neil.fraser.name·HN discussion ↗
  3. Audacity 4.0 by ClydeN — github.com·HN discussion ↗
  4. Holden's Lightning Flight by ColinWright — en.wikipedia.org·HN discussion ↗
  5. Pre-Release of Polars 2.0 by komape — pola.rs·HN discussion ↗
  6. The browser's main thread is expensive by kciter — kciter.so·HN discussion ↗
  7. K2 Horizon: A connected fleet of six open models by karimf — ifm.ai·HN discussion ↗
  8. Porting my 1993 Amiga game to Godot, with an LLM reading the 68000 assembly by rabahs — babyloniantwins.com·HN discussion ↗
  9. Whistleblower warns Postal Service mail ballot system has catastrophic problems by ck2 — cbsnews.com·HN discussion ↗
  10. GPS glitched across the US by as much as 33 feet by thread_id — sciencealert.com·HN discussion ↗
  11. What I learned from my mom (1941-2026) by NaOH — experimentalliving.substack.com·HN discussion ↗
  12. Reasons robotics is hard by ddp26 — secondthoughts.ai·HN discussion ↗
  13. Xanadu was waiting for agents by nsm — zed.dev·HN discussion ↗
  14. How to get a free .arpa domain by ethanhawksley — hawksley.dev·HN discussion ↗
  15. The true horror of Edgar Allan Poe’s stories lies in their confessions by lermontov — yalereview.org·HN discussion ↗
  16. Which tools do Claude, Codex and Cursor choose? We measured 17k runs to find out by screm — armature.tech·HN discussion ↗
  17. Invisible Companies by ltononro — colossus.com·HN discussion ↗
  18. Ask HN: Why were OpenAI, Claude, and Grok simultaneously down? by halcdev — news.ycombinator.com·HN discussion ↗
  19. Qwen 3.8 27B available on Cerebras at 1500 tokens/s by altertable — inference-docs.cerebras.ai·HN discussion ↗
  20. Google Antigravity TOS: 3rd party usage can get Google account suspended by tosh — twitter.com·HN discussion ↗
  21. Nvidia to acquire Hugging Face by tosh — cnbc.com·HN discussion ↗
  22. Any Human Ever – One life, drawn at random from all who have ever lived by thinkingemote — anyhumanever.com·HN discussion ↗
  23. The Computer Museum of America reclamation project by rbanffy — computer-museum.org·HN discussion ↗
  24. Higher Multipoles of the Cow by MrOrelliOReilly — arxiv.org·HN discussion ↗
  25. The shrinking landscape of linguistic diversity in the age of LLMs by Anon84 — nature.com·HN discussion ↗
  26. OpenAI's GPT-6 Astra on ARC-AGI-3 by vignesh_warar — arcprize.org·HN discussion ↗
  27. The largest electric aircraft just flew [video] by feb — youtube.com·HN discussion ↗
  28. How concerned should we be about Astra's recurrent architecture? by yurivish — lesswrong.com·HN discussion ↗
  29. Three schoolgirls in Kinsale pulled up a pea plant covered in warts (2014) by DamonHD — yahoo.com·HN discussion ↗
  30. Unusual Suspects by beeperboy95 — neal.fun·HN discussion ↗

Browse all issues in the archive →