Cover illustration

TheDaily Front

Issue No. #260918 Friday, September 18 2026 #260918 — FRIDAY, SEPTEMBER 18, 2026
The machines make choices; the rest of us read the fine print.
Friday, September 18, 2026 The Daily Front No. #260918 — Contents
30stories
9,056points
5,154comments
338kllm tokens
Assembled with 31 model calls — 221,848 tokens read, 116,602 written.

Highlights

Microsoft exec called AI scraping 'the largest theft of labor in human history'

Unredacted court filings put a startling internal label on the labor behind AI training data.

A heap overflow and SSO misconfiguration to compromise OpenAI internal repos

A heap overflow and an SSO mistake reportedly opened the door to OpenAI’s internal repositories.

I don't like passkeys

A vigorous case against passkeys finds the industry’s preferred login system wanting.

US Military had close call after using AI for hallucinated intelligence report

A reported military near-miss offers a bracing lesson in verifying AI-generated intelligence.

Claude Code now reads AGENTS.md if there is no Claude.md

Claude Code embraces AGENTS.md, a small interoperability win with outsized developer appeal.

From the Editor

The day’s wires bring a familiar modern contradiction: artificial intelligence is sold as assistance while its costs, provenance, and failures grow harder to ignore. Elsewhere, engineers continue the old work—making systems faster, safer, smaller, and, one hopes, less surprising.

  1. Microsoft exec called AI scraping 'the largest theft of labor in human history'3
  2. A heap overflow and SSO misconfiguration to compromise OpenAI internal repos4
  3. I don't like passkeys5
  4. Claude Code now reads AGENTS.md if there is no Claude.md6
  5. The scourge of x86 emulation7
  6. Pre-Greek: The lost language hidden within Ancient Greek8
  7. The most important product decision is what you don't build9
  8. Minimal Phone 210
  9. Show HN: Cactus Needle 3: 8-29MB automation models can match DeepSeek V4 Flash11
  10. C++26: Trivial infinite loops are no longer undefined behaviour12
  11. US Military had close call after using AI for hallucinated intelligence report13
  12. Photon-Emission-Guided Laser Fault Injection Enables RP2350 Secure Debug14
  13. How Uber Protects Against Retry Storms15
  14. I vibed a proof of Conway's conjecture16
  15. Telstra outage: The night a network decided the year was 200617
  16. Inside ZCode: Silently uploading your Git history to the cloud18
  17. Saving another 100TB of RAM19
  18. How SpaceX streamlined the Raptor engine20
  19. Shapelearn Qwen 3.8 27B (13.1 GB VRAM)21
  20. Jemalloc 5.4.022
  21. OpenJev23
  22. Cloudflare Quick Tunnels24
  23. Diplodocus, Long Thought Exclusively American, Turns Up in Spain25
  24. Warez: The Infrastructure and Aesthetics of Piracy (2021)26
  25. Qwen 3.8 Omni Flash27
  26. Android 17 is the first since 3.x to add new APIs without releasing to the AOSP27
  27. Warren Buffett Steps Down as Berkshire Chairman, Names Son to Replace Him27
  28. How do we prevent mathemathics from devolving into the Medieval Era of secrecy?27
  29. North Korean nuclear test sets off years of earthquakes27
  30. Ask A Monk – A digital wilderness for thoughts with no immediate answer27
The Daily Front Page 2 of 28
Friday, September 18, 2026 The Daily Front No. #260918 — The Cost of the Corpus
article

Microsoft exec called AI scraping 'the largest theft of labor in human history'

by pluc·▲ 868 points·768 comments·techcrunch.com ↗
AI scraping was tantamount to theft.

Microsoft CEO Satya Nadella speaks during the OpenAI DevDay event on November 06, 2023 in San Francisco

Image Credits: Justin Sullivan / Getty Images

Microsoft exec called AI scraping ‘the largest theft of labor in human history,’ new unredacted filings reveal

New unredacted information in the copyright lawsuit The New York Times brought against OpenAI and Microsoft three years ago reveals an admission that AI scraping was tantamount to theft, and that AI products pose a major threat to publications.

Per the lawsuit, a top Microsoft executive privately described the companies’ AI training practices as “theft,” and OpenAI’s own leadership said its AI models posed an “existential threat” to the publishers and journalists whose work trained them.

The unsealed material also details how the companies allegedly obtained and used that content by bypassing paywalls undetected, building training datasets via mass scraping, and deliberately stripping copyright notices from training data.

It’s worth noting that much of the new information comes from The Times’ own brief, not the underlying exhibits, which remain sealed. The quotes below are presented without their original context.

The unredacted filing is the latest escalation in the three-year-old lawsuit, in which The New York Times initially alleged the firms violated copyright law by training generative AI models on its content.

The question of whether AI firms can legally use copyrighted material to train AI has no clear answer, but judges have been largely favorable to AI companies’ arguments that training constitutes “fair use.” This legal rule lets people use copyrighted work without permission in certain cases, like parody, news reporting, or criticism. Earlier this month, the Trump administration contributed a brief in defense of OpenAI’s unlicensed use of copyrighted material to train its LLMs.

Several of the new admissions, however, run counter to OpenAI’s fair use defense, particularly the rule’s requirement that use doesn’t substitute or harm the market for the original work.

For example, Microsoft’s own data shows its Copilot “answer engine” caused click-through rates for The New York Times’ domain to drop as much as 93% compared to traditional Bing search. An internal Microsoft presentation written by Microsoft’s director of Applied Science, Brent Hecht, in January 2024 describes the decline as a “doom loop” that would “hurt the performance of our models and the entire web at the same time.”

“It is highly unusual that an end-product threatens the economic foundations of its essential suppliers, but that is the situation we have created for our LLM business with respect to its ‘content supply chain,’” reads the Microsoft document, as quoted in the filing.

Microsoft CEO Satya Nadella also testified in a deposition earlier this year that “anything that is paywalled should be licensed by anyone who wants to use it…for grounding or training,” and made clear that, if he “had been made aware that OpenAI had scraped and trained on information that was behind a paywall,” he would have “invoked [Microsoft’s right to] require OpenAI to retrain its models.”

Other admissions cut against different pillars of the fair-use test: OpenAI’s head of ChatGPT, Nick Turley, wrote in internal communication that publishers face an “existential threat” from products like the chatbot, which are “largely substitutive” and “will get more and more substitutive as they get better.”

OpenAI President Greg Brockman described the models as “excellent at news.” Nadella agreed under oath earlier this year that conversing with chatbots “has substituted … giving you the information right there on the website on the AI platform versus needing to go to the underlying source.”

That kind of language speaks to how the technology could directly compete with, rather than transform, the original work.

A Microsoft document states that there is a “real risk” that generative AI could “significantly disrupt the employment of the very people who generated the data on which the foundation model was trained.”

The sheer scale of the copying is striking. The documents reveal for the first time that OpenAI’s mid-training datasets alone contain more than 91,692 copies of works published by the NYT, Daily News, and Center for Investigative Reporting. A Common Crawl-derived dataset included more than 2 million documents from nytimes.com alone.

In a January 2023 internal memo, Hecht called it “an astonishing theft of unprecedented proportions” and “the largest theft of labor in human history.”

The filing lays out in new detail how OpenAI and Microsoft went about acquiring the plaintiffs’ content, including scraping it from the Bing Index.

“OpenAI delivered the entire GPT-3 training dataset to Microsoft, which Microsoft used to evaluate how to implement OpenAI’s models within its own commercial products,” the filing reads. “Microsoft similarly provided training data to OpenAI through initiatives called Project Taxi and Project Mango.”

The companies allegedly assembled the Project Mango data into a training dataset that contains copies of at least 160,903 unique works from the news publishers.

In order to get the most out of their scraping, OpenAI employees allegedly came up with a plan to circumvent paywalls without detection. The filings show that when OpenAI researcher Nick Ryder told Brockman about a “hack to get around nytimes paywall,” Brockman replied: “ah nice.”

OpenAI employees also allegedly built training datasets like WebText and WebText2 that disproportionately relied on scraped news content. They also allegedly pulled millions of articles from Common Crawl, a free, open repository of web crawl data. The findings also describe deliberate efforts to strip copyright notices from training data before it reached the model, since researchers “wouldn’t want model outputting” “copyright notices” to users.

“The evidence revealed here for the first time shows that OpenAI and Microsoft knew that what they were doing was wrong,” Steven Lieberman, counsel for the New York Daily News, said in a statement shared with TechCrunch.

OpenAI and Microsoft did not return requests for comment.

The Daily Front Page 3 of 28
Friday, September 18, 2026 The Daily Front No. #260918 — A Chain of Access
article

A heap overflow and SSO misconfiguration to compromise OpenAI internal repos

by Handy-Man·▲ 474 points·198 comments·hacktron.ai ↗
We chained two critical vulnerabilities to compromise multiple OpenAI employees’ ChatGPT accounts.

A heap overflow and SSO misconfiguration to compromise OpenAI internal repositories

Intro

On July 25, 2026, we chained two critical vulnerabilities to compromise multiple OpenAI employees’ ChatGPT accounts. With these accounts, we could then access internal OpenAI repositories, and potentially many other connectors.

To prove we had in fact gained the access we believed without allowing ourselves to learn any sensitive information, we used the employee’s Codex to open a PR #1186742 in OpenAI’s internal monorepo openai/openai.

Until two months ago, any user or OpenAI employee logging into OpenAI’s own help forum (community.openai.com) could have had their ChatGPT and Codex accounts taken over. Since people can connect various services to Codex and ChatGPT, the scope of what we could theoretically access was huge, including GitHub, Slack and emails.

The entire timeline from initial discovery to access to OpenAI repo access took place in less than 72 hours.

We immediately reported the initial vulnerability to OpenAI and Discourse and worked with them to coordinate the patch. We appreciate their attention to detail and fast resolution of this issue. OpenAI also paid us a $6,500 bounty.

We provide a full timeline of the disclosure process here. The rest of the post details how we discovered the two vulnerabilities, how we used claude models, as well as our takeaways from this experience.

  1. 25 July 2026 05:00–06:00 UTC

    Initial Finding

    HacktronAI team obtained remote code execution (RCE) and administrative access to the Discourse environment hosted at community.openai.com.

  2. 25 July 2026 08:00–10:00 UTC

    Bugcrowd Submission

    After confirming the cross-product impact, the team coordinated internally on the responsible disclosure process and submitted a report through OpenAI’s Bug Bounty Program on Bugcrowd.

  3. 25 July 2026 13:30–15:30 UTC

    OpenAI Employee Account Access & Proof of Concept

    To demonstrate the practical impact of the vulnerability, we created a harmless proof-of-concept pull request in OpenAI’s internal monorepo (link redacted at OpenAI’s request). We updated the existing Bugcrowd submission with these findings, reached out to friends at OpenAI on Twitter/X to notify them directly, and ceased all further testing at approximately 15:30 UTC.

  4. 25 July 2026 22:49:45 UTC

    OpenAI-Side Fix Confirmed

    OpenAI replied to the report confirming the issue had been fixed, roughly 14 hours after the initial submission.

  5. 25 July 2026

    Discourse Reported via HackerOne

    We submitted a report to Discourse through its HackerOne program.

  6. 26 July 2026

    Discourse Responded

    Discourse replied to the report on Sunday.

  7. 27 July 2026

    Discourse Fix Ready

    Discourse had a fix ready by Monday and added image-processing sandboxing as defense in depth.

  8. 28 July 2026

    Discourse Advisory Published

    Discourse published GHSA-vhm9-85gw-x335 with patch and rebuild guidance.

  9. 01 Sep 2026

    OpenAI Rewarded $6,500 Bounty and Marked Resolved

    OpenAI comment — To clarify the scope of that award: testing against the Discourse-hosted community.openai.com was explicitly excluded from our bug bounty program. The award recognizes the OpenAI-side finding, not the actions against Discourse.

Background

A few months ago, our team at Hacktron, led by Harsh Jaiswal alongside Mohan Pedhapati and Rahul Maini, began researching frontier AI companies to find security vulnerabilities. This led us to discover an SSO misconfiguration in OpenAI’s identity infrastructure and a libheif RCE in the community forum used by OpenAI.

We’ve since expanded the research into HEIF Heist, a multi-month investigation tracing libheif across Slack, Meta, GitHub Enterprise, Ruby on Rails, and Node.js frameworks such as Next.js, Astro, and Gatsby. A surprising amount of widely-used software depends on this one image-processing library.

xkcd 2347

If your application processes user-controlled images and accepts .heic/.heif/.avif images, it is highly likely it is affected. Please reach out to us at hello@hacktron.ai if you need any kind of assistance.

Hacking community.openai.com

Patch notice: If you self-host Discourse, rebuild your installation now. Older Docker images may contain a vulnerable libheif dependency that permits code execution through an image upload. Run git pull followed by ./launcher rebuild app from /var/discourse; a web-interface update alone may not replace the underlying image. Discourse-hosted customers have already been patched. See the security advisory.

OpenAI uses Discourse for their forum and allows “Sign in with OpenAI” through auth.openai.com. After getting a good understanding of OpenAI’s services and infrastructure, we had reason to believe that compromising the forum could create a path into broader OpenAI services through this identity flow. To test that hypothesis, we first needed remote code execution on an OpenAI service like the Discourse community forum.

While the Discourse app itself is actually not an easy target (we have looked into it in the past), we thought we could go after a dependency.

Heap buffer overflow in libheif

On July 23, we started reviewing Discourse’s image-upload pipeline, and we found that HEIC and HEIF files followed an unusual path. Discourse normally used FastImage for image checks, but because FastImage did not support HEIF, it passed those files to ImageMagick’s magick command for conversion.2 That exposed the underlying libheif parser directly to attacker-controlled files.

We started an Opus 4.8 session with the Discourse Docker image and asked it to inspect the installed libheif package for security issues. After a while, it found that some particular security fixes were not back-ported to the libheif package. This allowed an heap buffer overflow leading to OOB R/W primitives during HEIC decoding.

Interestingly, the vulnerable code had been changed upstream the previous year, but the commit was not documented as a security fix and received no CVE.3 This might be a reason why Debian 12 and 13 have not received the security relevant backports in time. Because Discourse’s Docker image was based on Debian 12, it installed the vulnerable libheif version 1.19.7. Even Debian 13 still shipped the vulnerable version 1.19.8 at the time. Since then, Debian has published its security update for Debian 13 on August 8, 2026.4

On July 24, we used Opus 4.8 to develop a working ImageMagick/libheif code-execution exploit with ASLR disabled. We then launched several separate sessions to make it reliable against Discourse’s default configuration with ASLR enabled, which wasn’t fruitful.

Opus 5 Released

That evening, Anthropic released Claude Opus 5.5 We started a new session, which first produced a working ARM64 exploit for a local Mac within 3 hours. We then asked it to port the exploit to the x86-64 environment and jemalloc configuration used by Discourse.

By 6:00 a.m. on July 25, we had confirmed local RCE through an image upload. We then placed Claude in an autonomous /goal loop against our own Discourse Cloud instance, proxied through rce.ee/ctf-forum to make it look like a CTF target as Opus refused write exploit for remote instances.

When we checked again at 10:00 a.m., the agent had achieved RCE on Discourse Cloud and demonstrated access by reading /etc/hosts. Using the generated exploit script, we managed to get RCE on OpenAI’s instance.

After we had confirmed our hypothesis of no interaction account takeover of ChatGPT/Codex accounts from active members of the forum, we immediately sent our report to OpenAI. We then took over an OpenAI employee’s account, whose Codex was connected to OpenAI’s Github organization. To demonstrate impact without actually accessing any internal code, we sent a prompt to this employee’s Codex account to open a PR for us in OpenAI’s internal monorepo. Then we stopped any further testing.

Redacted pull request demonstrating access to OpenAI’s internal monorepo

We updated the BugCrowd submission with the impact proof and alerted OpenAI security. We also prepared a report for Discourse and reported it to their HackerOne program. Discourse received the report on a Saturday, replied on Sunday, and had a fix by Monday (kudos for speed). They also immediately started sandboxing ImageMagick.

We want to emphasize that the vulnerability to escalate is not Discourse-specific. It is an OpenAI SSO issue that turned the forum compromise into access to ChatGPT and Codex. If any first-party or third-party OpenAI service using the OpenAI SSO was compromised, it would lead to same access - Discourse was merely one way of proofing it.

Costs of finding these vulnerabilities

The Discourse and OpenAI hack took a few days for an agent, and just a few hours of human time. The whole HEIF Heist research project going after Slack, Meta, adn more took two-months, cost less than $3,000 in tokens in total, and was conducted by three researchers. Adapting the exploit to each new company usually took only one or two days.

We observed that every new model is getting increasingly capable, as evident by the Discourse exploit presented in this report. Opus 4.8 struggled across several sessions to produce a working exploit with ASLR enabled. Within hours of Opus 5’s release, we gave it the same problem and it succeeded. Across the broader campaign, we saw another clear jump from Opus 5 to GPT-5.6 Sol, when we had to exploit the vulnerability without knowing anything about the target system besides that it’s vulnerable.

For each target, testing began with an image upload. From there, we turned memory corruption into a reliable memory leak or shell, usually without knowing the exact libheif version, libc version, or deployment environment. The AI started almost blind and adapted the exploit for each company within one or two days. We are not aware of any company that detected the activity except Shopify, even after thousands of images were sent and their image processors repeatedly crashed.

When code execution landed inside a sandbox or restricted environment, the models also helped with privilege escalation, lateral movement, and bypassing existing defenses. This was not completly autonomous hacking, and skilled human guidance remained important, but the amount of work a small team could perform increased dramatically.

Epilogue

Software has long benefited from a kind of security through complexity. The code and even the vulnerability could be public, but turning a bug into a reliable exploit still required rare expertise, significant time, and knowledge of the target environment. Known memory corruption vulnerabilities were expensive to operationalize, while zero-days were mostly reserved for the highest-value targets.

This was never a real security boundary, but it protected ordinary companies in practice from software vulnerabilities. AI is removing that protection by turning more of this scarce expertise into compute. Work that once required a well-resourced team and months of effort can now be compressed into days.

Security assumptions must catch up with attacker capabilities. A realistic threat model should take into account the economics of exploitation today, instead of relying on outdated assumptions6 about who can carry out sophisticated attacks.

Hacktron’s mission is to help secure the internet by finding and eliminating vulnerabilities in widely trusted software before malicious actors do. We are continuing this research across frontier labs and other internet-critical systems. If you are responsible for securing one of them, we would like to work with you.

Versions affected and patches

HEIF Heist is not tied to a single version. It targets an entire ecosystem of vulnerabilities across multiple release families (e.g. 1.19.x, 1.20.x, 1.22.x, 1.23.x). Any deployment lacking the latest upstream security patches is potentially vulnerable.

  • Update upstream. Install the latest security-patched libheif and libde265 packages through your distribution’s security channel or an upstream release. As of September 14, 2026, the latest upstream libheif security release is v1.23.4; v1.23.2 has been superseded by further security fixes. Distribution packages may carry backported fixes under an older upstream version number, so check the package security advisory as well.7 4
  • Defense in depth. Given the complexity of the ISO base media file format and the pace of decoder updates, future memory-safety flaws are likely. Production architectures should disable untrusted HEIF/AVIF decoding where it is not needed, or isolate image-processing pipelines inside hardened, ephemeral sandboxes. ImageMagick’s security policy supports restricting accepted formats and resource usage.8

Acknowledgements

We thank Sudanshu Rajhbhar for technical assistance, and Zayne Zhang, Fabian Faessler, Robert Chen, and Jessica Ruan for proofreading, reviewing drafts, and providing feedback that improved this post.

References

[1] xkcd #2347: Dependency

[2] Discourse: Support for HEIC images

[3] libheif: simplify overlay overlap area computation

[4] Debian DSA-6417-1: libheif security update ↩1 ↩2

[5] Anthropic: Introducing Claude Opus 5

[6] RAND: A Playbook for Securing AI Model Weights

[7] libheif v1.23.4 security maintenance release

[8] ImageMagick Security Policy

The Daily Front Page 4 of 28
Friday, September 18, 2026 The Daily Front No. #260918 — Keys to the Kingdom
article

I don't like passkeys

by ethanhawksley·▲ 744 points·723 comments·hawksley.dev ↗
The tech industry has kept pushing passkeys as the ultimate solution to logging in.

For the past few years, the tech industry has kept pushing passkeys as the ultimate solution to logging in. Many Big Tech companies “helpfully” inform you every time you sign in how much easier and effortless passkeys are. The only way to make them stop is either to concede and set up a passkey or dig into the settings to find the off-switch.

Google’s Skip password when possible setting

Google goes as far as to name the setting “Skip password when possible” (opens in a new tab), and Microsoft advertises that you should make your account passwordless (opens in a new tab).

Passkeys are a fantastic technology. Since they are bound to the site they are created for, they cannot be phished by a hacker’s fake login screen. If a site suffers a data breach, passkeys are asymmetric and cannot be recovered from the server-side details.

This leads to passkeys being the perfect fit for a corporate environment, but a poor fit for personal security. To an individual, the greatest risks are instead permanent account lockout, automated account bans, and device loss. By using passkeys, you gain better security against man-in-the-middle attacks but face the higher probability scenario of losing access to your accounts.

Phishing through the standard login flow is eliminated by passkeys, but it creates a false sense of security. An account’s security is still dictated by the weakest recovery method: SMS, email links, security questions, and so on. If these recovery methods aren’t enabled, then the risk of permanent lockout remains for the user.

Hardware keys

By design, you cannot create a backup of passkeys on a hardware key: passkeys can only be added or deleted but never moved. Instead, you need to purchase 2-3 hardware keys and enroll every key for every site. This can quickly get expensive and doesn’t scale well as the number of accounts starts to grow.

Hardware keys support discoverable credentials, where websites can query for your username instead of you typing it in. These are becoming increasingly popular amongst website developers, yet have limits of 25-100 accounts (opens in a new tab)per hardware key, and top of the line keys can have up to 300. Once you exceed the limit, you must either delete some accounts or you have to buy another set of hardware keys.

Synced passkeys

Both Apple and Google want your identity anchored to their operating systems. The “happy path” on their devices is to use their synced passkey management tied to your Apple or Google account. If their automated systems decide one day to ban your account (opens in a new tab), you irreversibly lose access to all your passkeys used across all third-party accounts too.

The FIDO alliance has been working to improve interoperability and make it easier to export passkeys, but the experience is still fragmented and inconsistent across providers. This is set to improve over the coming years, but currently it is too immature to rely on. Compare with a password, which is just a string you can easily export by hand if necessary.

Third-party synced passkeys

When storing passkeys in a password manager like Bitwarden (opens in a new tab)or KeePassXC (opens in a new tab), you end up fighting the platform. Although operating systems have recently introduced APIs (like Android’s Credential Manager (opens in a new tab)) for third-party tools to hook into, the experience remains fragmented and lacks the decades of UX polish towards password autofill. Autofill outside the browser and inside native applications remains especially inconsistent. In the future, I believe third-party passkeys will be the way forward, but we are not there yet.

When passkeys don’t work

Logging into accounts on devices you own is the ideal scenario for passkeys. When you have to handle a colleague’s computer, it gets much more inconvenient. You could plug in a hardware key, but you don’t always have access to the ports. You could sign in and use a synced passkey, but that involves trusting the computer to not leak all of your other passkeys. The last option is to use “Hybrid Transport” (opens in a new tab), where you scan a QR code and connect via Bluetooth simultaneously to the computer. Whilst this option is secure and works in theory, reality is plagued with edge-cases where connections fail or Bluetooth is straight-up unsupported.

Passkeys aren’t ready yet

I believe enterprise users have good reason to use passkeys, but the ecosystem isn’t mature enough yet for individuals.

Whilst TOTP codes have known phishing vulnerabilities, the recovery and lockout risks of passkeys pose a greater day-to-day risk to most people than an AiTM proxy (opens in a new tab). A combination of randomly generated passwords stored inside a third-party password manager, paired with an independent TOTP app, gives control to the user without giving up the flexibility of plain text. For users who previously reused passwords across all their sites, passkeys are a huge step-up. For everybody else, it is currently a step back.

The Daily Front Page 5 of 28
Friday, September 18, 2026 The Daily Front No. #260918 — The Agent Instruction Book
article

Claude Code now reads AGENTS.md if there is no Claude.md

by datadrivenangel·▲ 566 points·203 comments·code.claude.com ↗
Claude Code now reads AGENTS.md if there is no Claude.md.

Release notes for Claude Code, including new features, improvements, and bug fixes by version.

This page is generated from the CHANGELOG.md on GitHub. Run claude --version to check your installed version.

2.1.278

September 19, 2026

  • Changed auto mode for Claude API and Enterprise users, and on Bedrock, Vertex, Foundry and gateways, to default to the server-side classifier, which does not charge for classifier overhead (CLAUDE_CODE_AUTO_MODE_SERVER=0 opts out on Bedrock, Vertex, Foundry and gateways); warns on billed fallback. See https://code.claude.com/docs/en/auto-mode-classifier-billing
  • Added an Auto mode server row to /status showing whether this session’s auto mode classifier runs on the server

2.1.277

September 18, 2026

  • Added AGENTS.md support: in a project with no CLAUDE.md, Claude Code reads AGENTS.md instead; change it under “Project instructions” in /config (not yet on Bedrock, Vertex or Foundry)
  • Added CLAUDE_GATEWAY_PROXY_IS_EGRESS_BOUNDARY=1 for Claude apps gateways whose only egress is a forward proxy: every outbound request hands the proxy the hostname instead of resolving it locally
  • Added an optional headers: map on Claude apps gateway upstreams, to send static headers to a proxy you run in front of a provider
  • Added a line saying a background task’s update is waiting when it finishes while a panel such as /tasks is open
  • Fixed claude -p and Agent SDK sessions that could hang with no result after an internal error; they now report the error and exit with code 1
  • Fixed conversations failing every request with “text content blocks must be non-empty” when an earlier assistant turn held an empty text block beside other content, including after --resume
  • Fixed being unexpectedly logged out when an older Claude Code build (for example an IDE extension’s bundled CLI) runs on the same machine as the current one
  • Fixed interactive start-up hanging or showing an error for ANTHROPIC_API_KEY users when ~/.claude.json holds a malformed customApiKeyResponses value
  • Fixed update checks erroring every 30 minutes, and claude update hanging when a minimum or maximum version is set, if a proxy returns an invalid version; a malformed minimumVersion is now ignored
  • Fixed claude update on winget- or apk-managed installs reporting “up to date” when the version lookup failed
  • Fixed claude plugin install sometimes failing and breaking the installed copy when reinstalling a plugin version that a session or another program was using; an unchanged copy is now left alone
  • Fixed Grep and Glob reporting no matches when the search could not start because the system was out of processes, memory or file handles; they now return an error saying so
  • Fixed the Write tool silently ending the turn as a declined permission when the target path is an existing directory; it now reports a clear error
  • Fixed the Edit tool treating an escaped backslash followed by uXXXX text as a \uXXXX escape, which could make an edit of a non-ASCII character rewrite an escaped backslash sequence instead
  • Fixed the Edit tool reporting “Invalid regular expression: regular expression too large” instead of “String not found in file” when a very large edit containing non-ASCII text did not match the file
  • Fixed a turn ending early with “Path contains null bytes” when a tool call’s file path contained \u0000 written as an escape sequence; escaped control characters now stay as literal text
  • Fixed background sessions (claude --bg) exiting when a plugin’s LSP server exited or closed its stdin
  • Fixed a crash (“Type error”) when opening /mcp or /plugin manage with a malformed claudeAiMcpEverConnected value in ~/.claude.json
  • Fixed a crash at launch when ~/.claude.json holds a malformed theme value
  • Fixed a crash (“unrecoverable interface error”) when the prompt held text containing terminal color codes, for example a prompt recalled from history or text loaded from the external editor
  • Fixed a crash when resuming a session whose saved history holds an assistant message stored as a plain string
  • Fixed sessions on slow or heavily loaded machines sometimes exiting with “Claude Code exited after an unrecoverable interface error” when the first spinner appeared
  • Fixed a rare case where the screen could stop updating for the rest of the session after an internal rendering error
  • Fixed a rare case on Windows where a turn could stop with an error such as “Out of memory” right after Claude replied, so that reply’s tool calls never ran
  • Fixed sessions continued after /clear (restart, --continue, --resume) missing part of their first message when a SessionStart hook printed output, causing a full prompt-cache miss
  • Fixed messages from other agents (such as a subagent’s SendMessage) that arrived mid-turn showing up below the “Ran N shell commands” row instead of where they arrived
  • Fixed the “copied” notice not appearing after drag-selecting text in the fullscreen /resume picker and other panels that cover the prompt area
  • Fixed $TMPDIR expanding empty in Bash commands that run outside the sandbox while sandboxing is enabled
  • Fixed WebFetch and WebSearch in Cowork cloud sessions not telling Claude why a request was refused, such as a used-up fetch budget or an admin policy
  • Fixed the Claude apps gateway’s telemetry relay ignoring a collector hostname or domain listed in NO_PROXY when a proxy is set
  • Fixed one malformed strictKnownMarketplaces or blockedMarketplaces entry silently disabling the whole enterprise marketplace policy
  • Fixed failed auto-updates leaving large staged downloads behind in ~/.cache/claude/staging
  • Fixed /plugin not stripping terminal control characters from messages on the Installed tab, such as the error of a failed plugin update
  • Fixed /plugin → Installed and /skills crashing when a skill or legacy command is named like a built-in Object property such as constructor or toString
  • Fixed /plugin closing with no message when every install in a multi-select failed
  • Fixed uninstalled plugins reappearing as “failed to load” rows in /plugin Installed, and Remove not clearing such a row
  • Fixed plugins from the official marketplace being recorded without their commit in installed_plugins.json, and installed_plugins.json keeping the old commit after updating a pinned-commit plugin
  • Fixed plugin reload previews keeping every previewed copy of a plugin archive unpacked until exit, and overwriting the cached --plugin-url archive a reload falls back to when its download fails
  • Fixed Remote Control session bookkeeping failing when ~/.claude.json holds a malformed placeholder record
  • Fixed the error after a revoked claude.ai login blaming an expired Anthropic profile; it now leads with /login
  • Fixed typed or pasted text occasionally coming out scrambled in the claude agents dispatch input during key repeat or very fast input
  • Fixed a crash (“unrecoverable interface error”) when resuming a session whose saved transcript contains a stop hook summary without a well-formed hook list
  • Fixed Enter on a selected agent panel row doing nothing when keybindings.json rebinds Enter in the Chat context, for example to chat:queueSubmit
  • Fixed PDF page reads on Windows failing when the working folder’s path is long (about 120 characters or more)
  • Fixed a headless resume (claude -p --resume, the SDK, a VS Code extension window reload) starting the session’s cost and usage totals at zero; headless sessions now save their totals at exit
  • Fixed project skills from the main repository not loading in --worktree sessions when .claude/skills is untracked
  • Fixed a sandbox.excludedCommands glob exempting an entire compound Bash command from the sandbox when only one part matched; every part must now match
  • Fixed resumed subagents and teammates re-rendering the MCP tool definitions they had loaded, which broke prompt caching for that agent
  • Fixed rate-limited artifact publishes telling Claude to stop retrying; Claude is now told nothing was published and when to send the same publish again
  • Fixed attachments recorded earlier in a conversation being re-rendered after a resume or relaunch, which dropped extended thinking and missed the prompt cache
  • Fixed Console sign-in showing only “Request failed with status code 400” when the server refuses to create an API key; it now shows the server’s message
  • Fixed messages typed while Claude is still working sometimes being ignored by the model
  • Improved session start-up for SDK and headless (-p) use: the first turn no longer waits on the per-directory CLAUDE.md lookup
  • Improved the Claude apps gateway’s loopback error messages to name CLAUDE_GATEWAY_ALLOW_LOOPBACK
  • Improved /plugin Installed: an MCP server listed apart from its plugin now shows which plugin it belongs to
  • Improved claude plugin install on an already-installed plugin: it now says when the marketplace offers a newer version and names the claude plugin update command
  • Improved the startup notice overflow line under the logo: it now reads “N more notices hidden” instead of “+N more · /status”
  • Improved prompt handling: invisible Unicode formatting and tag characters in a prompt are removed and the cleaned prompt is shown for review before it is sent
  • Improved /ultrareview when there’s nothing to review: messages say which case you’re in, offer a command that reviews your latest commit, and a new repository’s first commit is reviewed in full
  • Improved artifact link handling so Claude reads claude.ai artifact links with the Artifact tool instead of WebFetch when that tool is available
  • Improved the dangerous-rm permission prompt to name the flagged rm command and suggest a ${VAR:?} guard, so headless runs can recover
  • Improved the Artifact tool’s permission prompts: shorter sentences, pages and artifacts named by title or file name, and links listed after the text
  • Changed Fable to always appear in /model on the Anthropic API; it is greyed out only when your organization’s settings disable it
  • Changed the Bash sandbox instructions on Bedrock, Vertex and Foundry to the first-party wording, which frames the sandbox as the boundary of what the task was given
  • Changed /ultrareview in non-interactive sessions to refuse when the repository has no base branch or shared history
  • Changed subagent results to reach the main agent under a header marking them as subagent output, with the result indented, so text in a subagent’s result cannot pass as the session’s own instructions
  • Changed workflow scripts’ computed agent() prompts on Bedrock, Vertex and Foundry to reach the subagent framed as script-authored text, so the safety classifier does not read them as the user
  • Removed the background Haiku auto-title request from claude -p runs launched outside an SDK or IDE
  • Removed the deprecated TaskOutput tool; Claude reads a background task’s output file with Read instead, and the taskOutputMaxChars setting and TASK_MAX_OUTPUT_LENGTH no longer have any effect
  • [VSCode] Added a Sign out row to the panel menu, with /logout in the typed command menu
  • [VSCode] Added background shells and other running tasks to the agent map, each with a Stop, and a typed /tasks that opens it
  • [VSCode] Added a Copy response button on responses and a typed /copy
  • [VSCode] Added a one-time notice when inactive sessions are archived automatically, and an “Unarchive all” action on the Archived sessions group
  • [VSCode] Added the session’s cost and token usage to the Account & usage dialog and the session manager where plan limits do not apply (Vertex, Bedrock, Foundry, API key)
  • [VSCode] Fixed the “General config” menu row showing /config usage text instead of opening settings, and made typed /mcp, /hooks, /memory, /rewind and similar commands open their dialogs
  • [VSCode] Fixed the effort slider’s level not persisting into later sessions on a model that already had a level saved with /effort
  • [VSCode] Fixed Auto missing from the mode picker for conversations opened in an already-used panel when the saved model setting is a differently-cased alias such as “Sonnet”
  • [VSCode] Fixed /fast not saving fast mode as the default, so it was lost when the extension relaunched Claude Code
  • [Claude Code on the web] Added Personal and Organization sections to the environment picker on Team and Enterprise plans, and admins can now share a personal environment with the organization
  • [Claude Code on the web] Changed organization environments to open as a read-only summary from the Code tab on Team and Enterprise plans, with editing under Admin settings → Cloud environments
  • [Claude Code on the web] Fixed a cloud environment saved with Custom network access and no domains silently reverting to Trusted; the dialog now asks for at least one domain
  • [Claude Code on the web] Changed the admin Claude Code setting labeled “Web” to “Cloud sessions” and removed the redundant read-only Mobile row beneath it
  • [Claude Tag] Fixed routines created in a Slack channel on an Enterprise Grid org-wide install failing to read other public channels in their workspace when they ran
  • [Claude Tag] Fixed the “Learn more” links on credential presets in Claude Tag access bundles to open each vendor’s credential-setup page instead of a generic API reference
  • [Claude Tag] Changed the Pylon credential preset in Claude Tag access bundles so admins can point it at Pylon’s EU host
  • [Claude Tag] Fixed Google Cloud credential forms in Claude Tag access bundles: a refused key file now says why, the website and scopes stay locked, and a rejected rotation keeps the pasted key
  • [Claude Tag] Fixed the network events log in Claude Tag admin settings showing no response status for requests through connections that use AWS signing, client certificates or a custom CA

2.1.276

September 18, 2026

  • Fixed every request failing with 400 … Input tag 'advisor_20260301' when ANTHROPIC_BASE_URL points at a proxy or gateway (2.1.275 regression)

2.1.275

September 17, 2026

  • Added the signed-in account to Claude apps gateway sign-in: when the gateway names it, you confirm it before the credential is saved, and /status shows it
  • Added a send-now key (ctrl+enter, or ctrl+x ctrl+s) that interrupts the current turn and sends all queued messages at once; sent and queued messages show in gray until the model receives them
  • Added a startup warning when a configured otelHeadersHelper fails, so sessions that silently export no telemetry are noticed
  • Added syncing of the skills and plugins enabled on your claude.ai account to terminal sessions signed in with it; opt out with syncClaudeAiSkills: false or syncClaudeAiPlugins: false
  • Added /plugin install <plugin> --marketplace <source>, which offers to add the marketplace before installing the plugin
  • Fixed a restored memory file’s age note changing between requests after a compaction or resume, which caused prompt cache misses
  • Fixed --forward-subagent-text stream-json and SDK output dropping the messages of subagents spawned by a context: fork skill, and of forked skills invoked by a subagent or another forked skill
  • Fixed @-mention file suggestions being buried below MCP resources when using a custom fileSuggestion command or typing @./@./
  • Fixed fullscreen mode placing background-task completion notices beneath a long turn’s collapsed tool row instead of where they arrived; each notice now closes the open row
  • Fixed claude plugin marketplace update deleting a GitHub marketplace’s local copy when the fetch failed and the marketplace was named after its repository
  • Fixed plugin and marketplace messages, logs and claude plugin marketplace list showing a password or token stored in a git, ssh or marketplace URL
  • Fixed a resumed cloud session leaving an unanswered question open in the transcript after a queued message superseded it
  • Fixed vim mode placing the cursor one character right after a dot-repeated ”!” or a fast-typed “i!” switched a non-empty prompt into shell mode
  • Fixed fullscreen mode freezing or blanking for several seconds when scrolling up past a large file diff
  • Fixed a stray </ccmemory>-style closing tag occasionally appearing in responses
  • Fixed plugin messages, logs and the VS Code plugin dialog showing the wrong server for some git addresses
  • Fixed a terminal API Error: 400 on every turn for users behind a network gateway that rewrites API error responses when a beta request header is rejected
  • Fixed sandboxed Bash commands on Linux reporting exit code 0 for failed commands when the shell is zsh
  • Fixed the Read tool hanging instead of reporting an error when part of a large file could not be decoded under memory pressure
  • Fixed --resume, the resume picker preview, resumed background agents and the transcript view failing on a session whose saved history contains a malformed task-reminder or @-file attachment entry
  • Fixed a crash when resuming a conversation whose transcript contains a malformed message entry, and a fullscreen crash when such a conversation received new messages while scrolled up
  • Fixed sessions failing to resume or start when their saved transcript contains a malformed message content block
  • Fixed Grep, Glob and @-file suggestions hanging or running out of memory on searches over the 20MB output cap, and system ripgrep reporting “no matches” instead of an error after a flood of warnings
  • Fixed /rewind in a forked or background session restoring a zero-filled or truncated file when the session’s file-history backups could not be fully copied
  • Fixed fullscreen sessions sometimes exiting with “Claude Code exited after an unrecoverable interface error” when typing fast or holding a key with the slash-command dropdown open
  • Fixed background sessions crashing and restarting their worker when a command fed through stdin ran on a machine that had run out of file descriptors
  • Fixed a crash at launch when ~/.claude.json holds a malformed mcpNeedsAuthNoticed value
  • Fixed --resume and --continue dropping a conversation’s earlier thinking when a built-in tool it started with has since been switched off by a server-side flag
  • Fixed text selected with the mouse in the fullscreen claude --resume session picker never reaching the clipboard
  • Fixed plugin reload previews replacing a running session’s extracted plugin files when the plugin was loaded from a --plugin-dir or --plugin-url archive
  • Fixed self-hosted runners with --drain-wait-sec losing the final result of a turn that finished during a SIGTERM drain; the runner now waits briefly for the turn to be reported
  • Fixed SubagentStop hooks with a specific matcher firing for every stopping subagent whose agent type was empty
  • Fixed sandboxed Bash commands being unable to write to project directories named hooks/ or config/
  • Fixed Artifact updates failing with “File not found” after a session resumes on another machine or its scratchpad is cleared: the page’s last published version is restored
  • Fixed /update-config writing Write(path) permission rules, which file permission checks don’t match, instead of Edit(path) rules
  • Fixed four dead documentation URLs (Pricing, Computer Use, Skills, CLI) in the bundled claude-api skill’s live-sources table
  • Improved prompt caching for a --system-prompt that contains a __SYSTEM_PROMPT_DYNAMIC_BOUNDARY__ line: the text above it is now cached globally, as the SDK’s array form already is
  • Improved the /desktop error when Claude Desktop does not open: it now says why and what to do next
  • Improved the Artifact tool’s publish and read results: they now say who can open the page and what the owner’s Share menu offers
  • Improved artifact publish results: they name the tab icon sent, warn when the page contains a NUL byte, and retry a flaky fetch of the newer page to merge after a stale publish
  • Improved pasted and attached images: they are now saved where Claude can open them as files without a permission prompt, including in Desktop and VS Code
  • Improved the Artifact tool’s guidance so Claude updates a shared artifact in place when you were given edit access to it, instead of publishing a separate copy
  • Improved plan-usage reads: editor windows and non-interactive sessions on one machine now share a read made in the last minute instead of each calling the usage endpoint
  • Improved the ListPlugins tool description so Claude knows it lists plugins enabled on your claude.ai account, not plugins installed locally with /plugin
  • Improved responsiveness when the terminal is slow or paused: output no longer falls further behind while the terminal catches up
  • Improved Write and Edit results for files in the synced account-skills folder: they now say the change is not saved to your account and how to save it
  • Updated /logout for Claude apps gateway sign-ins to also end the session on gateways that advertise token revocation
  • Changed hosted sessions to keep an unanswered permission prompt up after a container restart, instead of asking again
  • Changed the Artifact tool to ask for a one-word tab icon on a first publish instead of an emoji favicon
  • Changed Claude in Chrome in auto mode to skip the extension’s per-site check for classifier-approved calls, as bypass mode does, fixing browser_batch “Permission denied” after a redirect
  • Changed plugins installed from an npm source to be fetched with npm pack --ignore-scripts and integrity-verified, so a package’s install scripts no longer run
  • Changed scheduled and Run now routine runs to save data to, and republish the page of, an artifact you can edit without asking; public artifacts, first publishes and deletes still ask
  • Removed the startup notice that told you a one-off scheduled routine had run since your last session
  • [VSCode] Added viewing, editing and deleting a saved memory inside the Memory dialog
  • [VSCode] Added sending an attached image without typing any text
  • [VSCode] Added a Retry link to the MCP servers dialog when the server list fails to load
  • [VSCode] Added accept and reject buttons under each change in the proposed-change diff tab, so an edit can be reviewed change by change
  • [VSCode] Fixed the transcript creeping toward the bottom in small steps while a permission card waits and content keeps arriving
  • [VSCode] Fixed rewound and forked conversations not keeping the permission mode you had picked for the original conversation
  • [VSCode] Fixed an empty CLAUDE_CONFIG_DIR entry in the environmentVariables setting making Claude Code keep its files in the workspace
  • [VSCode] Fixed plugin install links opening the Manage plugins dialog for plugin names and marketplace addresses that can’t be used in a link
  • [VSCode] Fixed Remote Control staying shown as connected after a turn-off that Claude Code reported as failed; it now shows as off
  • [VSCode] Fixed the scroll to the bottom on send stopping short of the reply when the reply starts arriving during the scroll
  • [VSCode] Fixed the agent map showing agents a crash left unfinished as stopped instead of failed once the session is reopened
  • [VSCode] Fixed the “Continuing the step” notice not appearing, and the continue limit resetting, after a reload that follows a crash with background tasks still running
  • [VSCode] Fixed the session list showing when a session was last reopened, such as after a window reload, instead of when its last message was sent
  • [VSCode] Fixed “Fork conversation from here” failing on the message right after one sent while Claude was working
  • [VSCode] Fixed the prompt cache clock showing too few minutes after reopening a session with a message sent while Claude was working
  • [VSCode] Fixed a background agent that finished while Claude was running a tool losing its completion notice, and its result on the agent map, after a window reload
  • [VSCode] Fixed a rare case where text selected in a git-ignored file could be sent to Claude after the extension was unresponsive for several seconds
  • [VSCode] Fixed renaming a running session reverting to the generated name (regression in 2.1.269)
  • [VSCode] Fixed some claude.ai/code sessions opening in VS Code as an empty conversation with no messages
  • [VSCode] Fixed slash commands typed while Claude is responding being sent to the model as text instead of running once the response finishes
  • [VSCode] Fixed unreadable code in the plan preview and the Hooks and Permission rules dialogs with the High Contrast Light theme
  • [VSCode] Fixed /remote-control being ignored while Remote Control is still connecting: running it again now turns Remote Control off immediately
  • [VSCode] Fixed the conversation pulling you back to the bottom while a reply streams after you scroll up, and added a claudeCode.scrollToBottomOnSend setting to turn off the jump on send
  • [VSCode] Fixed the Manage plugins dialog showing a password or token that was typed into a marketplace URL
  • [VSCode] Improved the agent map: the pill counts running agents and turns red after a failure, the main agent stays in view while the map scrolls, and agents sort by state then end time
  • [VSCode] Changed New session in a Claude editor tab to open in the sidebar when Preferred Location is set to Sidebar, instead of always opening another tab
  • [VSCode] Changed a message sent while Claude is working to wait at the bottom of the conversation until Claude starts on it
  • [Claude Code on the web] Added a “New routine” button to the page shown when a routine link no longer resolves, next to the link back to your routines list
  • [Claude Code on the web] Fixed routine “paused” and “on hold” notifications being cut off mid-sentence; the paused-subscription notice now says to turn the routine back on yourself
  • [Claude Code on the web] Fixed cloud environments with a very long allowed-domains list saving fine and then failing every session start; saving now fails up front and says how much to trim
  • [Claude Code on the web] Fixed Claude’s guidance when a cloud session on a personal account is denied GitHub access: it now links to claude.ai/connect-github instead of an admin settings page
  • [Claude Code on the web] Improved what Claude tells you when asked to edit, delete or run a routine it didn’t create: it now links to the routine’s page so you can do it yourself
  • [Claude Tag] Added attach conditions for access bundles in Claude Tag settings: an Owner can let a bundle also apply in channels with guests or Slack Connect channels, not just member-only
  • [Claude Tag] Added Amazon CloudWatch, CloudWatch Logs, Amazon SNS, Google Cloud Monitoring and Cloud Logging presets to an access bundle’s Credentials tab in Claude Tag admin settings
  • [Claude Tag] Added Datadog presets for the US3, AP1, AP2 and US1-FED sites; new Datadog connections are now limited to Datadog’s read and query API routes
  • [Claude Tag] Fixed S3 uploads from recent AWS CLI and SDK versions failing with a 502 error when sent through an AWS connection
  • [Claude Tag] Fixed Claude treating a channel as inactive, and skipping untagged messages there, while it was still posting in that channel from a routine or a thread
  • [Claude Tag] Fixed a thread’s “Claude [task]” display name reverting to plain “Claude” after the session behind that thread was refreshed or restarted
  • [Claude Tag] Fixed the model you switched to in a Slack thread silently reverting to the channel’s default after that thread’s session was restarted or refreshed
  • [Claude Tag] Fixed Claude sometimes replying twice when another app or bot @mentioned it in a top-level channel message
  • [Claude Tag] Improved Claude’s notices in Enterprise Grid channels shared across workspaces: they now say when no workspace is set up yet, or why only organization defaults apply
  • [Code Review] Fixed reviews occasionally dropping part of their analysis when one of the reviewing agents returned its findings in an unexpected format
  • [Code Review] Fixed pull requests with more than 100 Claude reviews getting a full re-review on every clean merge from the base branch instead of the lighter merge-focused review

2.1.274

September 17, 2026

  • Added a visible warning when memory usage is critical, with steps to free memory or restart safely
  • Added CLAUDE_CODE_MCP_STARTUP_WAIT_MS to bound how long the first non-interactive turn waits for connecting MCP servers (0 = don’t wait)
  • Added effort attribute to the claude_code.llm_request OpenTelemetry trace span, matching the api_request event
  • Added claude_code.managed_settings_resolved OTel event: managed-settings sources and policy helper state; redacted settings and digests with OTEL_LOG_MANAGED_SETTINGS=1
  • Added store.connect_timeout_seconds to the Claude apps gateway config to lengthen the Postgres connect timeout (default 5 seconds), and improved the boot error when the database is unreachable to point to store.postgres_url and the configured timeout
  • Added enduser.sub, the IdP subject, to the telemetry Claude Desktop and Cowork send through a Claude apps gateway
  • Added a Claude apps gateway warning when a replica has more requests open than the 256 it sends upstream at once, and a startup log line showing that limit
  • Added click-to-expand for collapsed teammate and agent messages in fullscreen mode
  • Fixed sessions getting stuck endlessly retrying “unexpected tool_use_id” 400 errors: corrupted transcripts now self-heal where possible, and otherwise a clear error (with a /rewind hint) ends the loop
  • Fixed MCP servers configured as http that only speak legacy HTTP+SSE failing to connect when they answer the first request with 422 or another 4xx error
  • Fixed Streamable HTTP MCP tool calls timing out after about 5 minutes even when a longer per-server timeout was set
  • Fixed MCP prompts and resources not refreshing when a server sends list-changed notifications without declaring listChanged
  • Fixed MCP tool calls refused with 403 insufficient_scope being reported as an expired sign-in: the error now names the missing permissions and points to /mcp re-authentication
  • Fixed hook-driven sessions (such as an active /goal) ending with “Prompt is too long” instead of compacting when the context overflowed again after a reactive compaction
  • Fixed an active /goal being lost when resuming (--continue / --resume) a session that had compacted
  • Fixed claude agents losing --model, --effort, --permission-mode, --allow-dangerously-skip-permissions and --agent after an auto-update relaunch
  • Fixed a per-turn slowdown when a language server publishes project-wide diagnostics for thousands of files
  • Fixed subagents with model: "opus" on Bedrock, Vertex or Foundry leaving the session’s model when its id has no recognizable model family (unless ANTHROPIC_DEFAULT_OPUS_MODEL is set)
  • Fixed self-hosted runner sessions failing every turn with a 401 after a few failed token refreshes, until the next scheduled refresh; the runner now keeps retrying, and fetches a new token after a 401
  • Fixed clickable links to local file paths doing nothing in VS Code and other terminals that require a file:// URI
  • Fixed the transcript renumbering ordered lists in your own messages (typing “3. 2. 1.” displayed “3. 4. 5.”); numbers and “N)” markers now show as typed
  • Fixed AskUserQuestion preview notes being attached to a previously chosen option instead of the highlighted one
  • Fixed AskUserQuestion preview mode dropping the highlighted option when submitting a note with Enter
  • Fixed a resumed background agent keeping half of an interrupted tool batch when one of its calls was approved with a message
  • Fixed a local claude -p --resume started with CLAUDE_CODE_RESUME_INTERRUPTED_TURN not reporting background tasks the previous process left unfinished
  • Fixed the first turn of a cloud session sometimes starting without the tools of an SDK-hosted MCP server that was still connecting
  • Fixed background agent notifications claiming the agent had no live background work when it was still waiting on its own background task and would resume
  • Fixed error hints in Claude Desktop sessions to suggest slash commands like /usage-credits instead of CLI flags that cannot be used there
  • Fixed /schedule saving a routine’s prompt without its message role when Claude writes the routine in the shape that listing routines returns
  • Fixed /status not showing the apiKeyHelper failure that its own error banner told you to check
  • Fixed /fast on in non-interactive sessions reporting on and then turning off under an organization’s managed fast mode policy; it now says the organization has disabled it
  • Fixed the Artifact tool asking you to approve an update to an artifact that it then refused because the session had not read the latest version
  • Fixed Cowork and claude.ai cloud sessions with network access on treating reads of a teammate’s artifact as if network access were off
  • Fixed a plugin or marketplace directory with no git repository of its own taking its version from an enclosing git repository, such as a git-managed ~/.claude
  • Fixed --strict-mcp-config with an empty --mcp-config holding the first non-interactive turn for up to MCP_TIMEOUT on incidental MCP servers
  • Fixed Stop prompt hooks re-sending their whole prompt on every block in a conversation; repeat blocks now name the condition with a 500-character label
  • Fixed extra empty editor windows opening at startup on Linux under Wayland when running inside the Cursor or VS Code terminal
  • Fixed an unhandled promise rejection in the Claude apps gateway when Postgres drops a connection during a spend check
  • Fixed Claude apps gateway cutting every open stream on SIGTERM: it now lets in-flight requests finish for up to 25 seconds before exiting (CLAUDE_GATEWAY_DRAIN_TIMEOUT_MS)
  • Fixed installed_plugins.json being rewritten on nearly every start-up when plugin policy comes from remote managed settings, which made Claude Desktop reload every open session’s plugins
  • Fixed headless and SDK sessions making a separate model call for every background task that finished; completions already queued are now answered by one call
  • Fixed the Bash tool re-sourcing the shell profile (a multi-second stall on the next command) after every plugin reload; it now does so only when the plugins’ bin/ directories changed
  • Fixed plugins with a top-level $schema in hooks/hooks.json showing an “unknown key” notice
  • Fixed MCP connection errors and the MCP login tool’s description showing secrets resolved from ${VAR} placeholders in MCP configs
  • Fixed Bash permission checks for commands that loop over or assign certain special shell variables; these commands now ask for permission
  • Fixed worktree-isolated sessions accepting Bash commands with certain nested shell expansions; these are now refused
  • Fixed the Edit permission prompt preview sometimes showing a different location than the approved edit in files with multi-byte characters
  • Fixed background commands being stopped after 30 idle minutes on machines under mild memory pressure; they’re now stopped only when memory is critically low, and the debug log says why
  • Fixed a message a subagent sends to the main session disappearing from the Claude Desktop transcript after a relaunch
  • Fixed a plugin loaded from a .zip being served from a stale extraction after several overlapping reloads
  • Fixed a sub-agent’s progress summary being replaced by a runaway multi-paragraph reply
  • Improved startup in --input-format stream-json sessions: the first turn no longer waits up to 2s for still-connecting MCP servers whose tools tool search defers; they arrive on a later turn
  • Improved Monitor tool notifications: a script’s final output and its exit now arrive as one notification instead of two, saving a model turn
  • Improved Artifact tool errors: when you are not signed in to claude.ai the terminal now says so on the first attempt, and Claude is told to stop retrying a rejected call sooner
  • Improved artifact publishing: a publish built on an older version is stopped before it is sent, with the newer page to merge
  • Improved safety checks before removing an agent worktree that contains submodule checkouts
  • Improved OTEL_LOG_RAW_API_BODIES=file:<dir> output: a new index.jsonl and request_body_id / message.id event attributes link each response to its request file and transcript message
  • Improved Claude apps gateway boot: it now tries the first Postgres connection up to three times before exiting, so a database that is reachable a few seconds late no longer fails the boot
  • Improved the Claude apps gateway’s spend-limit check under load: it now takes one database round trip instead of four, so fewer checks time out on a busy gateway
  • Improved Claude apps gateway sign-in rate limit errors: /login now explains the refusal, and the gateway log says which limit was hit and which setting to change
  • Changed Bedrock, Vertex, Foundry and telemetry-disabled installs to use the v2 MCP client and MCP 2026-07-28 negotiation with direct HTTP servers by default, as other installs already do (opt out: MCP_SDK_GENERATION=v1 or MCP_PROTOCOL_NEGOTIATION=legacy)
  • Changed /code-review to use leaner inline review prompts for every model that has no tuned settings of its own, instead of spawning many review subagents
  • Changed "type": "sdk" MCP entries in .mcp.json, settings, plugins and agent files to be skipped with a warning: only an SDK host application can register in-process servers
  • Changed artifact watching in local sessions: a new version published elsewhere no longer starts a turn; Claude learns of it from a later Artifact tool result
  • Changed plugin and marketplace clones to leave Git LFS files as pointers instead of downloading them; git lfs pull in the checkout fetches them
  • Changed self-hosted runners to skip a read-only repository the git host refuses at the access check instead of failing the session start
  • Changed the /status GitHub line to read “Cloud sessions”, and /web-setup, /ultrareview, and teleport messages to say “cloud session” instead of “Claude Code on the web”
  • [VSCode] Added continuation of the step a window reload interrupted, labeled in the chat, with a Claude Code: Continue After Reload setting to turn it off
  • [VSCode] Added Memory and Instructions entries to the Customize menu: Memory shows the auto-memory toggles, the saved memories and the memory folders, and Instructions edits the CLAUDE.md files
  • [VSCode] Added a claudeCode.lockEditorGroups setting to stop Claude from locking the editor groups it opens in
  • [VSCode] Fixed a /btw side question asked in a new conversation’s first seconds occasionally showing another session’s side-question history
  • [VSCode] Fixed a brief freeze when the extension first looks up your global gitignore file
  • [VSCode] Fixed a message sent while Claude was running a tool disappearing from the conversation after a window reload
  • [VSCode] Fixed the Manage Plugins enable toggle and MCP servers dialog rows being unreachable from the keyboard
  • [VSCode] Fixed sign-ins and sign-outs made in a terminal not showing until a reload after CLAUDE_CONFIG_DIR changed in the Environment Variables setting
  • [VSCode] Fixed Edit diffs in the chat being cut off at the bottom at some panel widths and for long wrapped lines; diff boxes now fit the rows shown
  • [VSCode] Fixed overlapping settings writes from the extension leaving ~/.claude/settings.json unparseable or dropping a setting
  • [VSCode] Fixed Open in New Tab (Ctrl/Cmd+Shift+Esc) sometimes leaving the new tab’s message box unfocused, so typing went nowhere until you clicked it
  • [VSCode] Fixed reopening a closed Claude tab splitting the editor layout when its locked group still holds another Claude tab and a file
  • [VSCode] Fixed New session opening another locked editor group whenever a file tab shared the group with your Claude tab
  • [VSCode] Fixed session names shifting sideways in the session picker while typing a search query
  • [VSCode] Fixed the plan review card cutting off its Send feedback button and reason field when a plan has several comments; the comment list now scrolls
  • [VSCode] Fixed inline code and code blocks in chat replies being unreadable under the High Contrast themes
  • [VSCode] Improved screen reader navigation of the conversation: each message is announced as “You” or “Claude”, with the tool name for tool steps
  • [VSCode] Changed the default global gitignore file to $XDG_CONFIG_HOME/git/ignore when XDG_CONFIG_HOME is an absolute path
  • [Claude Code on the web] Added a “Compare against” branch picker to a cloud session’s diff view, so you can diff its changes against any branch instead of only the base branch
  • [Claude Code on the web] Fixed git operations in cloud sessions failing with “service unavailable” when GitHub’s token renewal briefly errors
  • [Claude Code on the web] Fixed editing a routine occasionally making it fire twice or re-enabling a routine that had just been paused
  • [Claude Code on the web] Fixed commits in cloud sessions occasionally failing with a signing error for a few minutes after the session’s credentials refreshed
  • [Claude Code on the web] Fixed the toast after saving a routine whose GitHub trigger couldn’t be linked to show the reason, such as a per-repository trigger limit, instead of only “edit to retry”
  • [Claude Code on the web] Fixed sessions sometimes flipping back to unread right after you mark them read
  • [Claude Code on the web] Changed routines to skip a run and retry for up to 72 hours when the owner’s GitHub connection is missing, instead of switching the routine off at the first failed check
  • [Claude Code on the web] Changed a routine’s on-hold notice: when your subscription is paused it now tells you to turn the routine back on yourself instead of promising an automatic resume
  • [Claude Tag] Added a Guests setting to the Add channel and Add workspace forms in Claude Tag admin settings, so owners can pick Inherit, Allow, Channel only or Restrict up front
  • [Claude Tag] Fixed Claude not answering when another Slack app or bot @mentions it; the tag now gets a reply and wakes Claude in a channel it had stopped following after days of inactivity
  • [Claude Tag] Fixed Claude missing another app’s message that tagged @Claude right after a new Slack channel was created; it’s now delivered once Claude has joined
  • [Claude Tag] Fixed Claude folding a follow-up sent minutes after its last Slack message into it as a silent edit; late updates such as blockers now post as a new reply that notifies
  • [Claude Tag] Fixed Claude’s Slack search failing with an error whenever it searched within a single channel; it now returns that channel’s matching messages
  • [Claude Tag] Fixed a safety-filter stop silently resetting a Slack thread’s context when nobody was waiting; Claude now always says so and no longer cancels background work still running
  • [Claude Tag] Fixed email addresses in Claude’s Slack replies rendering with a visible mailto: prefix; they now show as the plain, clickable address
  • [Claude Tag] Fixed Claude refusing to watch an Enterprise Grid channel shared with the whole organization when asked from another workspace in the grid
  • [Claude Tag] Fixed the Environment picker in Claude Tag admin settings showing a raw environment ID instead of the environment’s name for archived or app-created environments
  • [Claude Tag] Improved Claude’s live progress checklist in Slack: capped at 2,000 characters, reposted at most every 15 minutes in busy threads, with older “Latest task list” links updated
  • [Claude Tag] Removed the repeated guest-attribution note Claude appended to a Slack canvas each time it edited one in a channel using the “Channel only” guest setting
  • [Code Review] Fixed re-reviews occasionally leaving a fixed finding’s thread open when the new review also filed a lower-severity note under it
  • [Code Review] Fixed rare reviews ending with “Code review encountered an error” when GitHub or an internal service failed transiently at launch; they now wait and retry
  • [Code Review] Improved how Code Review words each posted finding: short plain sentences that say who is affected, where the code goes wrong, and the fix up front
  • [Code Review] Improved the check-run card and PR comment when a review is skipped because of an organization limit: each cause now links the admin page that fixes it

2.1.273

September 15, 2026

  • Added x-claude-code-request-class, x-claude-code-agent-type, x-claude-code-prev-tool-durations, x-claude-code-compaction and x-claude-code-context-compacted request headers for LLM gateways; opt in with CLAUDE_CODE_GATEWAY_HINT_HEADERS=1
  • Added a notification when an MCP server disconnects mid-session and automatic reconnection gives up, pointing at /mcp
  • Added forking a session started with claude --remote-control or /remote-control from the Claude app; the fork runs as a background session on your computer
  • Fixed Bash commands the permission checker cannot fully analyze skipping the prompt under permissions.blockReadsOutsideWorkingDirectories, and a subshell hiding a dangerous rm in bypass mode
  • Fixed skills synced from claude.ai staying available after your organization turns Skills off; they now move to the recoverable trash
  • Fixed allowManagedMcpServersOnly, deniedMcpServers and disableClaudeAiConnectors set via MDM or managed-settings.json being ignored when server-managed settings are also present
  • Fixed 401/403 errors on Bedrock, Vertex and Foundry, and Claude apps gateway 403s, telling you to run /login; the message now names the credential to refresh or points to your gateway administrator
  • Fixed /login, /upgrade, and /extra-usage discarding earlier thinking from the conversation, which forced a full prompt-cache rewrite on the next request
  • Fixed auto mode stopping for approval when the Artifact tool uploads a file you attached to the chat in a cloud or Remote Control session
  • Fixed a long-running session recreating a stub .git/info/exclude after the repository’s .git directory was removed or moved away
  • Fixed the main prompt dropping a ! typed at the start while already in shell mode, so negated commands like ! grep … can be typed
  • Fixed Read on macOS refusing a dragged-in screenshot, or any file the system reports under a second path, with “symlink resolution changed after permission was checked”
  • Fixed permissions.blockReadsOutsideWorkingDirectories: a memory directory chosen by a repository’s settings is no longer loaded into the prompt, recalled, indexed, or used by memory extraction
  • Fixed sub-agents and background agents being reported as failed, with their result never delivered, when the final streamed reply omitted token usage or carried no model id
  • Fixed the context meter and auto-compact counting advisor-tool turns at roughly twice their real context size, which made auto-compact fire at about half the real window
  • Fixed /tui refusing to restart because of an agent-team teammate that had already finished its work and was no longer shown in the agents panel
  • Fixed saved scheduled tasks running in the wrong session after .claude/scheduled_tasks.json was copied into another folder, such as a new worktree
  • Fixed SDK and --output-format stream-json output dropping a subagent’s remaining messages and final report after it is moved to the background mid-run (e.g. by CLAUDE_AUTO_BACKGROUND_TASKS)
  • Fixed /install-github-app reporting a SAML single sign-on block as “admin permissions required”
  • Fixed Remote Control clients attached to a Claude Desktop, VS Code or JetBrains session being refused when they ask for the session’s context window usage
  • Fixed the spinner showing a doubled ellipsis (”……”) on compaction status lines such as “Running PreCompact hooks…”
  • Fixed a false-positive spinner tip suggesting the frontend-design plugin after reading or publishing Artifacts
  • Reverted a 2.1.268 change that checked Read and Edit deny rules on Bash lines the permission checker can’t analyze (eval, env -C); commands like time -p make build prompt again instead of being denied
  • Improved responsiveness in long sessions: hook progress and sub-agent activity no longer re-process the whole conversation on every update
  • Improved the Artifact tool’s error when a publish includes a file type artifacts don’t serve: Claude is told which types are served and what to do instead, and the terminal shows one plain line
  • Improved the Artifact tool’s page read to state the capabilities and database rules the artifact service holds for the page, for anyone who can publish to it
  • Improved artifact database writes: an update can now remove a single field instead of rewriting the whole document
  • Improved artifact publishing: a publish whose connection drops after reaching claude.ai is now re-sent safely instead of failing or creating a duplicate version
  • Improved the cloud-session GitHub error for an IP allow list, a suspended app installation or SAML single sign-on to show the cause instead of a generic install hint
  • Improved /autofix-pr: when gh pr view fails it now shows gh’s own error (sign-in, SAML, rate limit) instead of a generic exit-code line
  • Improved /autofix-pr to say why GitHub webhook delivery couldn’t be set up for the PR (for example, no linked GitHub account) instead of a generic warning
  • Improved /web-setup errors: a refused GitHub token now lists the likely reasons and the fix, and a connection failure names a configured proxy or TLS certificate problem
  • Improved the in-session SSL certificate and proxy connection errors to name the error code and what to fix, such as NODE_EXTRA_CA_CERTS for an untrusted corporate CA
  • Improved the error when a cloud session can’t be created because your Claude login expired or was revoked: it now tells you to run /login
  • Improved the error shown when an MCP server’s sign-in expires mid-session to say how to re-authenticate (/mcp)
  • Changed auto mode on Bedrock, Vertex and Foundry to use the local classifier by default for now; set CLAUDE_CODE_AUTO_MODE_SERVER=1 to use the platform’s server-side classifier
  • Changed OTEL_LOG_TOOL_DETAILS=1 to also include real agent, skill, plugin and MCP server names on cost and token metrics
  • Changed sign-in with a Claude account to also request access to your claude.ai plugins
  • Changed /bug and /feedback reports to include only model-behavior params (model, system prompt, tools) from the last API request, omitting request metadata and CLAUDE_CODE_EXTRA_BODY fields
  • [VSCode] Fixed “Report a problem” still appearing, and /bug //feedback opening a report form, for organizations that have product feedback disabled
  • [VSCode] Fixed a red “Claude Code process exited with code 4294967295” banner appearing after completed turns on Windows
  • Windows: Improved the network-path permission check for UNC paths when a mapped network drive was added with --add-dir
  • [Claude Code on the web] Fixed routines losing access to an organization connector, and still calling the old one, after an admin removed and re-added that connector
  • [Claude Code on the web] Fixed creating a self-hosted environment from organization settings occasionally failing with a server error and leaving a half-created environment behind
  • [Claude Code on the web] Changed the admin “Share cloud sessions” setting to live under Data and privacy instead of the Claude Code page, where Data and privacy admins can also manage it
  • [Claude Code on the web] Added a “Discard unsaved changes?” confirmation before the New routine page or the Edit routine dialog throws away a routine name, prompt or edit you typed
  • [Claude Code on the web] Removed the full-page desktop-app download screen that new users without a cloud environment saw on Mac and Windows; they now go straight to setup
  • [Claude Code on the web] Improved the routine detail page: menu and rename in the breadcrumb, the on/off switch and Run now at the top, and run history beside the routine’s settings
  • [Claude Tag] Fixed Claude going silent minutes after reinstalling the app when an Enterprise Grid was disconnected but one of its workspaces stayed connected
  • [Claude Tag] Fixed scheduled tasks set up in an organization-shared private Slack channel silently never posting; they now keep running in the thread they were created in
  • [Claude Tag] Fixed replying in an older Slack thread while Claude is mid-task sometimes restarting it from scratch and losing work it had not pushed yet
  • [Claude Tag] Fixed Claude occasionally dropping a message with an incorrect “couldn’t find a Claude Code environment” notice right after your account token refreshed
  • [Claude Tag] Fixed AWS connections refusing region-less endpoints such as Budgets, Savings Plans, WAF Classic and Import/Export; Global Accelerator requests now sign correctly
  • [Claude Tag] Improved AWS connection failures: when a request can’t be signed, such as a hostname with no region, Claude is told why and how to fix it instead of a bare error
  • [Claude Tag] Fixed OAuth client-credentials and JWT-bearer connections failing with providers that return a lowercase token type; requests now send the standard Bearer scheme
  • [Claude Tag] Fixed adding a channel manager being refused on Enterprise Grid shared channels, on channels where Claude hasn’t been used yet, and on legacy private channels
  • [Claude Tag] Changed Claude to start watching related public channels on its own, such as an incident channel a conversation depends on, instead of only when asked
  • [Claude Tag] Fixed the admin Memory page not listing Slack channels Claude set up on its own even when they had saved memory; admins can now open, edit and delete that memory
  • [Code Review] Fixed merging the base branch into a PR whose earlier review listed “Additional findings” triggering a full re-review; these pushes now get the lighter follow-up review
  • [Code Review] Fixed a whole REVIEW.md being ignored because of an @-mention, a code span wrapped across lines, or a backticked HTML tag; only lines linking to changed files are withheld
  • [Code Review] Improved suggested fixes to say what the fix must keep working when other code depends on the behavior being changed
  • [Code Review] Improved review comments that point to a second affected location to state that location’s issue in a full sentence instead of a cut-off stub
  • [Code Review] Fixed /ultrareview --post so a retry after a GitHub error posts the findings comment exactly once instead of never or twice; the comment now names the reviewed commit
  • [Code Review] Fixed empty or content-identical pushes being re-reviewed on GitHub repositories whose owner or name contains a capital letter; these pushes are now skipped

2.1.272

September 15, 2026

  • Bug fixes and reliability improvements

2.1.271

September 14, 2026

  • Added fast mode in Claude Code Remote sessions (cloud and self-hosted runners): the host’s fast-mode setting or /fast typed in the session applies where your organization allows it
  • Added mouse support to the /config panel in fullscreen mode: the wheel scrolls the settings list, a click on a setting’s value changes it, and the row under the pointer is highlighted
  • Added claude self-hosted-runner --drain-marker-file <path>: when that file exists at a SIGTERM drain, the runner reports its exit to the server as a host drain (telemetry only)
  • Added per-command allowed_domains to Bash, PowerShell and Monitor in auto mode with sandboxing: the hosts a command needs are reviewed with it and opened for it alone; other hosts are refused
  • Added omitClaudeMd to agent frontmatter and --agents JSON, letting custom and plugin subagents run without user, project and local CLAUDE.md files; managed policy files still load
  • Added --accept-command <sha256> to claude plugin install and claude plugin update to accept exactly the command a previous --json run displayed, instead of -y
  • Added support for a multiplier above 1, up to 10, in the modelPricing managed setting and the Claude apps gateway pricing block, for marked-up internal chargeback rates
  • Added a spinner tip pointing Bedrock, Vertex AI, Foundry and LLM gateway users to the Claude desktop app; the claude.ai desktop app tip now suggests /desktop, which offers to download the app
  • Fixed a cached organization policy being reused after switching accounts, organizations, or API keys, and the policy not refreshing until the hourly check when the credential changes mid-session
  • Fixed the tool and command lists not updating when the organization policy finishes loading after startup or changes mid-session
  • Fixed an enterprise managed-mcp.json that can’t be read or parsed being ignored: it now keeps exclusive MCP control (user, project and plugin servers don’t load) and warns at startup
  • Fixed org policy being fetched through, and rejected by, third-party local proxies set via ANTHROPIC_UNIX_SOCKET; they are again treated like other custom gateways, including for Remote Control
  • Fixed cloud sessions rejecting every subagent tool call (“updatedInput … failed schema validation”) when a workflow or agent approval was applied after the session’s worker restarted
  • Fixed /fast off answering “Fast mode unavailable” instead of turning fast mode off when the organization has fast mode disabled
  • Fixed sessions started with CLAUDE_CODE_SKIP_FAST_MODE_ORG_CHECK re-sending fast requests every turn after the API rejected fast mode; the rejection now stands and its reason is shown
  • Fixed fast mode under CLAUDE_CODE_RETRY_WATCHDOG failing the turn on a usage-credits limit, or retrying an overload at fast speed, instead of falling back to standard speed
  • Fixed Bash permission checks missing the file that fmt, column and similar commands read when it follows an option the checker doesn’t recognize
  • Fixed Bash permission checks skipping files a wildcard expands to when the wildcard sits in a command’s pattern or option value (for example grep -v dir/* file)
  • Fixed Bash permission checks so that shell variable declaration flags cannot misrepresent the command being run
  • Fixed Bash commands with two directory changes, a subshell, or a cd+git chain skipping the prompt under permissions.blockReadsOutsideWorkingDirectories in bypass and auto mode
  • Fixed a stale .git/config.lock breaking git checkout -b, git push -u and git config for the rest of a session after a git command was interrupted
The Daily Front Page 6 of 28
Friday, September 18, 2026 The Daily Front No. #260918 — Emulating Order
article

The scourge of x86 emulation

by dagmx·▲ 279 points·88 comments·fex-emu.com ↗
The problems with emulating this memory model on the weak ordering memory model.

Welcome to the first feature article on our site. We’re going to cover an ongoing problem with x86 emulation that affects every application that we emulate. This comes down to a single over-arching term that has wide-reaching ramifications; Emulating the x86 Total Store Ordering memory model (x86-TSO).

The problems with emulating this memory model on the weak ordering memory model that ARM defines is multi-faceted and covers multiple issues. We’re going to go over all the problems that we can encounter and the ways we solve (or in some cases can’t solve) in this article. Get yourself a snack and a warm drink to enjoy, this is going to be a long one.

What exactly is x86-TSO?

Before diving in to how we work around the x86 memory model problem, we need to first discuss exactly what it is. A memory model is a set of rules for how memory accesses in a system behave in relation to each other. The rules will dictate how loads and stores interact in a single-threaded or a multi-threaded environment. There’s a handful of popular memory memory models implemented in various forms of hardware, but the two we care about today is ARM’s relaxed (or weak) consistency model, and the x86 variant of Total-Store-Ordering consistency model. These two models are basically the two extremes of the spectrum; where ARM is the most relaxed, allowing significant hardware optimizations; and x86 is the most strict, enforcing a very strong coherency model that doesn’t allow a lot of room for optimization. One thing to be careful about when discussing memory models is the difference between consistency and atomicity. While these are related, they are not the same nor guaranteed in all cases.

The best way to explain how the differences in memory models work is to start with how x86 handles this. With TSO being very strict in how it operates, the programmer can assume that when a memory store occurs, that this will be coherently visible to all other processors in the system. This additionally means that when a memory load occurs, all stores before it “logically” will have been completed, or at least visible. This matches programmer expectations, you write to memory, it becomes visible as at the point of writing, as this is intuitive to think about when programming. The stores are effectively ordering the visibility of the loads, thus the name of the model. There’s a bit of nuance with how this operates but isn’t strictly necessary to understand.

The weak memory model that ARM has is a bit less intuitive about how it operates. By default the regular memory loads and stores that ARM uses aren’t strictly coherent across processors in your system, allowing the CPU to operate more efficiently most of the time. When a store instruction executes, that piece of memory (the cacheline) isn’t immediately visible to other processors in the system. Saving on precious power and efficiency because it’s expensive in hardware to invalidate other core’s cachelines, or allow them to snoop another processor’s caches. Relatedly if a processor is loading data from memory that another processor has written to, it’s not guaranteed that this load will even see this updated memory. This sounds like it would cause some significant problems in a multi-threaded application right? Older versions of ARM (ARMv7 and older) used a memory barrier instruction to ensure ordering, which had significant performance implications.

To get around this limitation of consistency, ARM also introduced load-acquire, and store-release memory instructions. In C++ parlance this maps to std::atomic’s memory_order_acquire and memory_order_release definitions respectively. In ARM’s terminology, these instructions also aren’t technically considered to be atomic operations, but programmers conflate the two. FEX has used the terms atomic-load and atomic-store to mean the same thing! The distinction usually doesn’t matter, but when discussing these topics it may be better to be pedantic about it.

The primary use case for these instructions is to force memory ordering between these class of instructions. ARM calls this the “Release Consistency sequentially consistent (RCsc)” model. Without getting too far in to the weeds about how this model operates, the basic gist is that the load-acquire instructions must be observed sequentially without reordering, and the store-release instructions must as well while fulfilling “barrier-ordered-before” semantics. Removing the costly memory barrier instruction required in older ARM architecture versions.

The humble beginnings of ARMv8.0-a

This is the premise of where we start in ARMv8.0-a when we’re emulating the x86-TSO memory model. We make all x86 memory loads turn in to ARM’s load-acquire instructions, and x86 memory stores turn in to store-release instructions. This gives FEX effectively the same memory semantics as x86, although we are actually being more strict than what is necessary. This is because we had no middle-ground which exactly matches behaviour. As one might think, it is exceedingly costly to emulate TSO wth this instructions and we have microbenchmarks that can show this. As ARM CPUs weren’t designed to have these relatively rare acquire/release instructions suddenly become the vast majority of instructions executed.

First let’s start with something easy and use a microbenchmark that is fairly nice to the hardware. No tricky edge-cases, just accessing memory in the common case. This gives us some baseline numbers for what the best-case situation should be.

Let’s break down this graph as it tells us a few interesting stories. The Load and Store columns of each machine is representing our baseline performance number that our hardware should be attempting to achieve. These aren’t trying to max out the memory bandwidth of each system, but do the same amount of work for each type of operation. If we turn our attention to the acquire-load results, we can see that out of the five CPUs tests, three of them have their performance hindered quite a bit by using acquire-loads! Additionally we can see that the AmpereOne CPU has release-store instructions that are strikingly low compared to the other results, and the M1 Acquire/LRCPC load instructions are quite a bit lower than the baseline as well.

The AmpereOne results in particular showcase how bad this legacy path can get. These instructions were never designed to be used this way. Using acquire-release semantics for every load for x86 emulation actually imposes some really strict limitations on ARM CPUs in that the load instructions can no longer be ordered around each other at all. So when you have millions of them in flight per second, the performance isn’t really expected to be good. But because these are the only instructions we had with ARMv8.0-a, it’s what we had to use. While Cortex-X4 and Cortex-X925 have amazing performance for these, you can see how the Oryon-3 has deprioritized their importance.

Where do we go from here?

Let’s take a closer look at the LRCPC-load instructions, which is mandatory since ARMv8.3. This extension adds a bunch of new load instructions to the ARM ISA and adds a new memory model on top of ARM’s RCsc model from before. This new “Release Consistency processor consistent (RCpc)” memory model is what we’ve been wanting! This extension is designed around the requirements that x86 emulation requires, and is expected to get utilized heavily on hardware that implements it. As you can see from the graph, almost all of the platforms have their LRCPC-loads matching their regular loads in performance.

With this new extension that is mandated by newer ARM versions, we basically get solved memory performance. At least according to this microbenchmark that seems to be the case. Once FEX detects this extension we stop using Acquire-Load instructions entirely and switch over to LRCPC-Load instead. But what’s going on with that Apple M1 result..?

This is where we need to commend Apple’s path towards solving this problem. With their Apple Silicon processors they directly added support for the x86-TSO memory model. When the CPU feature is toggled, their regular load/store ARM instructions change behaviour to match what x86 requires. They went this route knowing that they will need a high performance solution for their hardware when switching to the ARM ecosystem exclusively. This is why on their hardware the LRCPC-load instructions are actually aliases of their acquire-load instructions, because their x86 emulator doesn’t even use these instructions! Because they implement the x86-memory model, they just use regular load/store instructions, which can be seen in our microbench results as indiscernable performance overhead. To be fair to the other platforms, this thread-wide TSO mode toggle does have some performance impact, we just don’t see it here. When FEX detects this CPU feature from Asahi Linux we will also enable this and get the “free” performance improvement. A potential concern is that when jumping between x86 emulation and ARM code, that the ARM code will pay unnecessary overhead due to all its accesses being TSO now. While this is a reasonable concern, the amount of ARM native code executing under emulation approaches 0%. As a developer, you don’t care about 1% of memory accesses becoming 10% slower, you care about 99% of accesses becoming 15% of the “ideal” (As shown in AmpereOne results).

As a note, we think a TSO mode is the best path forward for ensuring high performance x86 emulation on the platform. Because this ensures that every memory access instruction behaves how we want or expect. This is shown with the official FEAT_LRCPC extension actually having three versions that apply bandages to the implementation each time.

  • FEAT_LRCPC - Adds basic GPR TSO load instructions
  • FEAT_LRCPC2 - Adds small offset immediate to TSO load instructions
  • FEAT_LRCPC3 - Adds basic vector and stack-based TSO load & store instructions

Even with these three extensions, there is edge-case behaviour that can’t be emulated as nicely as if we had a TSO hardware toggle. We are expecting there to be additional extensions versions as time goes on, trying to fix some of the additional problems we’ll discuss later in the article.

I thought accessing memory was the easy bit?

In the previous section, we were being nice to the ARM hardware and playing along with the underlying hardware’s alignment requirements to get a baseline for what the performance should look like. When emulating x86 although, we run face first in to a glaring problem right from the start. Your favourite x86 applications don’t care about alignment! They’ll access memory however they please, crossing cacheline granularities, doing atomics that aren’t aligned. You think of the alignment problems, these games are doing it. This problem is so bad that we have a term associated with it, called split-locks. These are such a big deal that even the Linux kernel will capture when these occur and slow down games when they do it! Causing many gamers to tinker with kernel options to avoid the slowdown!

But we aren’t going to talk about full on split-locks yet, let’s get started with just load-store instructions in an environment that doesn’t care about alignment. x86 makes certain guarantees to the programmer; if you do a load-store and it is inside of a cacheline then that load-store will be both atomic and still match the coherency model as described before. However, to be a little bit nice to the hardware developers, if the load-store does cross a cacheline, the data isn’t atomic and other threads can and will see it tear. So the programmer needs to be careful as a basic load-store is not a split-lock.

The problem with emulating these basic accesses with load-acquire/store-release is that ARMv8.0 requires what is known as natural alignment. This means that for whatever size of data being accessed, the offset in memory must match the size. So for an 8-byte access, it must be at offsets; 0, 8, 16, 24, etc. This works well for native ARM applications, but what happens when we don’t obey natural alignment requirements? For ARM, this means the instruction with raise an alignment fault. The hardware validates that the alignment requirements are fulfilled and if they are not then the CPU will fault. This usually results in a crash but FEX does special handling.

Inside of FEX’s JIT mechanism we keep track of memory load-store instructions that are emulating the x86 load-stores. When we know that a load-store can cause an alignment fault we have what is known as a patchpoint in the code. For load-store instructions, this shows up as a NOP instruction either before or after the load-store. When a alignment fault occurs as one of these patchpoints, FEX will capture the fault, patch the code from a load-acquire/store-release instruction to a basic equivalent load-store, and wraps the instruction in a data memory barrier. Then it continues executing!

Before then after patching

That entire discussion from before about how ARMv8.0-a added these new fancy load-acquire, store-release instructions? We immediately fall back to the classic memory barrier instruction instead when alignment behaviour doesn’t match. Our previous chart didn’t show this bad case, so let’s bring in some fresh data.

Oh, that’s a lot of data to sift through. While again good to see how far away the hardware is from the “optimal” path while emulating TSO, it’s not what we care about here. It is interesting to note that this microbench doesn’t showcase much of a difference between aligned and unaligned for regular load/stores so we just calculated an average between the two. We’ll be removing the x86 CPU and the regular load-store data from the ARM columns, as these aren’t the common FEX paths. This way we’ll have a more targeted view about how badly unaligned memory accesses hurt under emulation.

Now that we have a much more reasonable graph of data, let’s walk from left to right on this and discuss what is going on.

AmpereOne

This one is pretty interesting, both the aligned and unaligned load instructions are roughly equivalent and fall within noise. This means that even though the unaligned loads are getting hit with a data memory barrier penalty, the CPU just handles it. This might be the case that the benchmark is bottlenecked by other things, considering how much lower the performance is compared to other platforms.

Meanwhile the store side is not looking to be in a good shape even without unaligned. It nearly isn’t visible on the chart! When hitting unaligned stores we’re looking at ~8.5% of a performance hit, but because we are already starting so low it is hard to notice. This is also in stark contrast to regular store instructions getting ~28GB/s in this bench.

The only conclusion we can come to here is that Ampere is optimizing for some server class workload and doesn’t really match consumer hardware behaviour. It’s an interesting datapoint, but our users aren’t typically running games on this class of hardware.

Cortex-X4

This is a highly popular CPU core that is living inside the Qualcomm Snapdragon 8 Gen 3. We only tested this one core from the SoC to not overwhelm the chart with data. Quite a large number of handhelds ship with this so it’s an interesting target. This CPU actually does surprisingly well considering it’s the only cellphone SoC on this list. Overall this core kind of falls in line with what we would expect from it and the graph trends follow with the next-generation Cortex in that chart.

The main topics for this CPU are that its aligned loads and stores are reasonably powerful, getting around 11.5GB/s and 6.7GB/s respectively. What’s interesting is the performance falloff when it needs to deal with unaligned loadstores, hitting the DMB instructions penalizes the core roughly evenly between loads and stores at around 50% in this benchmark.

This seems to imply that the CPU can keep a decent number of LRCPC-release loadstores in flight so the DMB instructions hurt more when they are encountered, but it isn’t causing world-ending performance. Just that a 50% performance hit due to alignment isn’t an amazing result.

Cortex-X925

Following up the X4, let’s stop by the DGX Spark and its X925 cores. Not only is this a newer CPU core from ARM, it’s running on a system with dramatically more memory bandwidth. 273GB/s in the platform versus the previous 76.8GB/s. This means that we get fairly similar results to the X4 even, just the graph scales a little higher. Interestingly enough, the performance penalty for unaligned accesses roughly match the X4 even. Although it looks like the stores can recover a little faster, likely due to the faster memory helping out. No surprises here, just consistently matching performance across the generations.

Oryon-3

This CPU core design is hot off the presses from Qualcomm. Linux support is still in the process of coming up but it already has a strong showing. The most interesting result from this actually comes from the fact that aligned LRCPC-load instructions are matching the performance of regular loads! That means in the case of a well-behaved application we can typically expect full performance. This continues onward to the release-store instructions being quite capable, although it doesn’t quite match regular stores with only 68% of the bandwidth. Not a bad showing in the slightest.

This CPU also can’t escape from the penalty of unaligned LRCPC-release loadstores. The load side is roughly matching the ~70% performance penalty of the Cortex-X925, likely because the Snapdragon X2 Elite also has tons of bandwidth. But the store side actually gets off a little worse at ~43% of the performance. Even with these performance hits of unaligned accesses, this platform is actually faster than the aligned accesses from the Cortex offerings.

One of the weird things about this platform is that it was advertised to have “Fully coherent 96KB 6-way L1 cache with 64B coherency granules.” Which to our reading implied that unaligned accesses should have dramatically less of a performance impact. Interesting… keep that in mind.

Apple M1

This is the big one we need to talk about. This is the one that was a game changer, it was the “Apple moment.” It showed everyone that ARM was not only feasible, it could be faster. These numbers on this chart are amazing and it’s the result of Apple sticking the TSO memory model directly in to their hardware. Instead of using LRCPC-release accesses for this one, we just enabled their TSO feature and the aligned versions basically match the unaligned version. Maybe a 5% performance hit on the stores? Compared to every other device on that chart, it’s effectively nothing. This primarily comes down to unaligned accesses no longer requiring DMB instructions to be backpatched in to the code, as the hardware just handles it directly.

For us, this is what it means to take x86 emulation seriously on ARM and it really shows that Apple cared that their customers would have a good experience running software both natively and emulated. They saw the problem and just “solved” it, making it go away. That said, when the TSO mode is enabled, you do get a performance hit. Comparing to the previous graph it’s only getting 76% of the regular store performance, and the load performance basically matches; that’s much more tolerable to bear when everything is so much faster.

Wrapping up unaligned LRCPC/release accesses

Wrapping up this section, we need to talk about one of the performance improvements that all of these vendors actually support. This is an extension that ARM whipped up called FEAT_LSE2 which all of these tested platforms implement. We previously talked about how acquire/LRCPC/release memory accesses require natural alignment in order to not incur the wrath of the CPU raising alignment faults. ARM actually thought about this problem and implemented this extension which helps x86 emulation (and probably other workloads). This extension loosens the alignment requirements of not only acquire/LRCPC/release load store instructions, it also loosens the requirement for read-modify-write atomics!

That sounds all well and good, but here’s the kick to the teeth: that means it only provides marginal performance gains for x86 emulation. This extension only loosens the alignment requirements to allow unaligned memory accesses inside of a 16-byte granule. Any access that crosses that 16-byte granule still receives an alignment fault. x86 applications don’t really care about the alignment of their memory accesses, so we get unaligned accesses across the entire cacheline. It’s only read-modify-write atomics that try to avoid crossing a cacheline on x86!

So thanks for the attempt, it’s nice to see, but it doesn’t really move the needle. Since we’re already talking about it, let’s dive in to those RMW atomics shall we?

Oh no, what are these atomic instructions?

Like most modern instruction sets, x86 supports atomic memory operations. These are instructions that execute an ALU operation on data in memory atomically, allowing no intermediate state to be visible. In x86 terms this operates on memory that is both atomic and coherent, while ARM lets you choose to be only atomic or both atomic and coherent. We touched on this briefly before but there is actually a difference between operating on data atomically, and coherency of that data. What difference does it make?

For all of the previous x86 memory model discussion we have been talking about the coherency implications of loads and stores being visible to other processors in the system. What we entirely glossed over is the atomicity requirements of these memory accesses. In the world of x86 a load or store usually completes atomically even when unaligned. This means that if you’re storing 8-bytes of data, and another thread is loading those 8-bytes in a race condition it will never suddenly see a mix of the data from before the store and after the store. In ARM these atomicity guarantees are significantly weaker, meaning if you do an unaligned store instruction the specification of the ISA has zero guarantees about reading a tear in the data. Thankfully for naturally aligned load-store instructions, ARM has a specification called “single-copy atomicity” which guarantees you don’t get a tear for these accesses. Also good news; that FEAT_LSE2 extension from before? It actually extends the single-copy atomicity guarantees to any unaligned access inside of a 16-byte granule! The downside is that x86 has single-copy atomicity guarantees across a full cacheline, so once again the extension still didn’t solve anything completely, just reduced the number of occurences.

Enough about the differences in atomicity and coherency. Where’s the actual atomic instructions? What do they do? Starting in ARMv8.1-a, our ISA has gained instructions that mostly matches x86 atomic instructions in behaviour. Let’s just give the full list to show how they map directly in our JIT.

x86 ARMv8.1-a LOCK DEC ldaddal LOCK INC ldaddal LOCK NEG ??? LOCK NOT ldeoral LOCK ADC ldaddal LOCK ADD ldaddal LOCK AND ldclral LOCK OR ldsetal LOCK SBB ldaddal LOCK SUB ldaddal LOCK XADD ldaddal LOCK XOR ldeoral LOCK BTC ldclralb LOCK BTR ldeoralb LOCK BTS ldsetalb XCHG swpal LOCK CMPXCHG casal CMPXCHG8B caspal CMPXCHG16B caspal

Well would you look at that, we have a full list of the 19 atomic RMW operations and they basically map directly to some ARM instructions. Ignore the questionable one as it’s not used in real workloads and we would get far too in to the weeds talking about it. We have a pretty clear 1:1 mapping between the architectures, job’s done right? That’s the funny thing about x86 emulation, just because we have these instructions doesn’t mean we get to wire them up without problems. We spent all this time talking about how unaligned accesses can really hurt performance of regular loads and stores, this same problem also applies to RMW atomics!

With this graph, we are looking at a single atomic instruction with its memory address landing somewhere within a cacheline. If we included all of the data for all 19 atomic operations then this data would be even more overwhelming than it already is. All these atomic operations behave roughly equivalent so it would be redundant and wouldn’t matter for what we’re discussing here anyway. This is also the first graph in this post that is actually using logarithmic scaling, so when reading it make sure to understand that the performance difference from the fastest to slowest result is on the scale of around 1000x.

Starting with the x86 Zen processor on this graph; these are the results that our emulation should be striving to achieve. As we can see, if the access is fully contained within a cacheline then the latency of the instruction is the same at 1.44ns. This can be explained by x86 having “atomic cachelines” or “coherent cachelines”, where as long as an unaligned atomic operation stays within a cacheline then it roughly costs the same. This is a really powerful feature of x86 that has been supported for decades at this point so games end up relying on this heavily without even realizing it. The stand-out result for x86 is the final result that is crossing a 64-byte granule and taking ~660ns! That’s an amazingly slow result at ~458x slower compared to the other results because this is finally the hardware using split-locks.

We need to take a moment here to shout out an article that Chips and Cheese wrote while we were preparing to write our article. They do a great deep dive in to why these split-locks are so dramatically slower and is worth the read if you’re unaware of how they work. Specifically we need to mention that x86 split-locks maintain the atomicity and coherency requirements of x86-TSO and will never tear the data even when crossing a cacheline. This is kind of nuts and we’ll explain this more later.

Now for our ARM processors, let’s start with the natural alignment latency numbers. As we can see, all of our platforms perform fairly well but even the latest cores don’t get anywhere near x86. Even our fastest ARM platform is ~3x the latency compared to x86; This directly impacts performance of games but usually isn’t the direct bottleneck so it’s hard to measure exactly how much. Continuing onward to the next data point, we can actually combine the results for 16-byte granule and 64-byte granule crossing with most of our ARM platforms. Due to how the ARM specification defines how unaligned atomics work, both of these results are roughly equivalent and FEX treats them the same as the x86 split-lock problem.

We keep bringing up this split-lock problem but how exactly does FEX emulate them and what makes it so slow? “I thought Apple M1 added x86-TSO support in the hardware, why is it still slow?” If you recall how we brought up before that FEAT_LSE2 introduced support for unaligned memory accesses within a 16-byte granule; these split-lock operations end up hitting the same alignment problems as before but are dramatically slower. FEX can’t backpatch any of these instructions to just do a DMB operation, so we cause an alignment-fault every time one gets executed. This means that we do a kernel -> userspace signal handler -> kernel -> original code dance. every—single—time one of this split-lock operations execute. Jumping between kernel-space and userspace is slow on every platform and when you’re executing thousands of these per second it adds up very quickly. This is why the emulation of these feature is so terribly slow on ARM.

One ARM platform today actually partially resolved this problem although. The Oryon-3 CPU cores introduced what they advertised as “coherent cachelines” and we can see this in our microbenchmark results here. Just like with x86, if the atomic memory access in anywhere inside of the 64-byte cacheline, the performance matches the natural alignment version! This is a tremendous improvement that means the CPU is on par with x86 in feature support until the point it tries to cross a cacheline. We need to applaud Qualcomm on implementing this feature, as it resolves a major performance and correctness problem around split-locks for x86 emulation. The hardware still doesn’t support 64-byte split-locks so we still fall down the FEX emulated path in that instance although.

Continuing on to the Apple result; even though they added x86-TSO memory accesses to their hardware for some reason they neglected to implement full cacheline unaligned atomics like Oryon did. It seems like they should have expected this edge case to surface and implement it but that’s just speculation. This is why you can see the cross 16-byte granule behaving the same as other platforms even with the TSO hardware toggle enabled.

You might have also noticed another little data quirk in the graph. We have an asterisk on the Cortex-X4 result in this benchmark and the performance of the unaligned atomics are dramatically faster than significantly newer CPUs. It is somehow managing to have only ~209ns latency, while the X925 is latency is 1060ns; that’s a 5x perf improvement! How can this possibly be the case? This is actually some fun “special sauce” that is shipping on the platform we’re testing on, which is of course the Valve Steam Frame. Because Valve cares about the performance of their existing gaming catalogue, they are shipping a kernel patch that one of the FEX developers whipped up. This allows the Linux kernel itself to handle the unaligned atomic without that slow dance with FEX and userspace, allowing it to be dramatically faster. If other platforms want to ship this patch in the kernel then we recommend picking it up as and FEX will automatically start using it.

Speaking of kernel intervention, we need to talk about how split-lock emulation is not actually quite correct under FEX due to limitations in the hardware. In order to implement this mandatory feature of x86 correctly, any time we do a 16-byte or 64-byte split-lock, the only way to handle it is to have the kernel implement the feature. Right now FEX implements this as a “best-effort” attempt that can actually tear the data in some cases. You’ll recall that before we said split-locks on x86 will never tear right? Not even the Oryon-3 with its “coherent cachelines” have resolved this problem yet.

What do you mean split-lock is mandatory?

Implementing split-lock emulation with today’s ARM hardware in a performant matter is actually really difficult to do. A naive implementation is to use a global mutex and whenever a split-lock occurs we will ensure to acquire the mutex before doing the operation. This means that any participating split-lock operation will funnel through this mutex. This is correct except for the issue that any aligned atomic operation isn’t a split-lock and won’t participate. Due to the split-lock emulation code needed to be implemented as two 64-bit compare-exchange operations with each half straddling the granularity boundary, we can get a tear with a non-participating atomic still. A trivial example is one thread constantly modifying an atomic in the middle of the cacheline, and then another thread modifying only the integer on one half. This might sound like a contrived example initially, but there are lock-less linked-list implementations that behave exactly like this! Depending on which half the aligned thread is modifying, either the first or second CAS in the split-lock code will fail. If the first CAS fails, then that’s safe and the code can retry, if the second CAS fails that means the data has torn and we can do nothing but hope it doesn’t corrupt data and crash. This will entirely depend on the algorithm that the guest application is using so we don’t control it.

An alternative approach that is completely untenable is to have the kernel track all processes and threads that are sharing memory with each other, then when a thread needs to emulate a split-lock the kernel can halt every process that is sharing memory with that process, do the split-lock in isolation, and then restart the world. The performance implications of this approach aren’t viable. Applications and games can end up doing thousands or more split-locks per second and halting the world will have an intractable performance hit that is dramatically worse than even x86 native.

If we want to ensure correctness in the emulation of split-locks FEX needs to have hardware support in some form to support these. Although we’re not saying that all atomic operations should now support split-locks like x86, that would also not be viable. The good news is that ARM actually has an extension for this that does exactly what we want. ARM has an extension call Transactional Memory Extension that could solve our problem. This extension allows our code to do some number of operations inside of a transactional region, then commit that work atomically; if the commit operation fails, then we can simply retry. The downside of this extension? ARM has officially deprecated the extension and no one ever shipped it. This is likely for the best as the x86 version of the extension has had an abundance of problems that caused it to be disabled on many platforms.

So we need something else to emulate split-locks correctly. For a solution that we believe works for both FEX needs and ARM vendor needs, we have come up with the idea that a 128-bit CASP instruction can be given the ability to have each half of the CASP perfectly straddle the atomic granule boundary, 64-bits on the lower half, and 64-bits on the upper half. Then only in that case does the instruction not raise an alignment-fault and tries to do the CAS operation. This works because x86 only has up to 64-bit unaligned atomic operations, so both halves of the operation can always be fully enclosed by our single operation.

But you may be asking yourself, “how is this any better than the hardware just supporting split-locks?” That’s a good thought and we need to be careful with the how exactly we describe this operation. For x86 their atomic operations must always succeed without tear. For our emulated approach, we can have this ARM CASP instruction fail safely and then we can try again. This is one of the benefits of CAS is that the operation can fail for any reason and it must be tried again. The instruction then also returns the data that it loaded from memory in that time so the program has the latest up to date memory. This is an important distinction since that means FEX can retry the CAS operations infinite times until it inevitably succeeds! This is a benefit of ARM LL/SC architecture that basically allows this to work. A tricky thing is that the hardware does need to guarantee forward progress at some point but it already has support for that for other reasons so it’s completely viable! The only newly added failure mode to the CAS instruction is purely if one of the two cachelines got acquired by another core before it could do the full operation. Even if the hardware still requires up to a couple thousand cycles to guarantee forward progress, that basically matches x86 behaviour.

We think this would be the best way forward for x86 emulation of split-locks on ARM platforms, but we’re not hardware architects so all we can do is complain and hope someone solves it for us. We’ll leave the split-lock discussion there for now so we can move on to another interesting problem.

Wait, uncached memory needs to work?

Before we get in to this topic we need to talk about the term “uncached” because it can mean a couple of things depending on your view of the world. For the purposes of this article, we are using Vulkan terminology because we care about games primarily. In Vulkan terms we have VK_MEMORY_HOST_CACHED_BIT which means that the host CPU caches this memory. The lack of this bit is what we care about here, and what we refer to as “uncached.” As for what this means to the memory subsystem, it gets a little more complicated than you would think. In particular when the memory is living on a GPU, potentially over PCIe, when the memory is “uncached” it will also typically (but not always!) also gain the flag VK_MEMORY_HOST_COHERENT. This means that because of the uncacheable property of the memory, the CPU and GPU always have a coherent world memory view with each other.

For the CPU this typically means the memory can be mapped up to three ways. When asking for “cached” memory, this typically has a memory type of Write-back which is also what regular memory mapping types are. “uncached” mapping can be either Write-Combine or “Strong Uncacheable”. The “Strong Uncacheable” implementation is basically non-existant for userspace applications so we can ignore that for today’s discussion. This limits us to effectively WB (cached) and WC (uncached) memory types. Cached is what games typically use for staging buffers, and then uncached is what we use when passing data directly to the GPU.

This is code-ified in many game engines that if you don’t expose support for uncached buffer types then some don’t work. This comes down to a behaviour detail around the differences of UMA systems like APUs and PCIe GPUs. UMA systems will typically expose the ability to allocate memory that is cached, coherent, and GPU visible. Where PCIe GPUs can’t guarantee that behaviour so game developers need to either use a staging buffer and an async copy of the data over to the GPU, or use “uncached” memory to very carefully shuffle the data over to the GPU through PCIe. Because of how ubiquitous PCIe is with PC gaming, some engines won’t even do UMA specific code paths and will do the uncached approach regardless!

With that little introduction out of the way for what uncached means for us. Let’s bring up a benchmark for how fast cached memory is on some UMA Snapdragon systems. This will let us get a baseline for how the performance should be regularly.

For both the Steam Frame and Snapdragon X2 Elite these are some really good results. As we would expect, the Oryon-3 platform has more memory bandwidth so it is able to scale higher in the chart, but both are hitting dozens of gigabytes per second in their results. This graph sets a good baseline for what “normal” write-back memory can achieve. Let’s now show uncached results to see the performance differences.

There’s some strange things happening here so we had to use logarithmic again on this graph. Let’s talk about the good first that has shown up. Due to uncached memory buffers being write-combine, we can see that the regular stores for our ARM platforms match the cached benchmark results. This comes down to write-combine memory using what is coined as write combine buffers that actually very temporarily keep around a cacheline of data so that write-combine can burst a cacheline of memory at a time. Interestingly enough it looks like the Zen 4’s WCB can’t quite keep up with cached, but considering this is expected to be going over a PCIe bus it’s probably fine.

Now let’s get in to the really ugly results that we have here. Starting off with the easier to explain is the load bandwidth from write-combined memory is abysmal on all platforms tested. If we’re using Zen as our baseline for performance, then our regular load instructions are ARM are winning, but the LRCPC loads are worse. What’s going on here? This is a quirk of how write-combined memory operates, because it is uncached our load instructions are required to go out to system memory for every single access to maintain semantics. Then when we add LRCPC-loads on top of that, it just compounds the problem even further. But the worst case out of all of this is just how badly the store performance is, compared to the performance that Zen gets on the stores, this is basically a showstopper. Up to 816x worse bandwidth! We had games like Hollow Knight: Silksong and Subnautica 2 run at less than 1FPS because of this performance cliff.

As we were saying above, when there are PCIe GPUs in the mix then games will need to use uncached memory to pass data to the GPU. When emulating x86 games on platforms with a dedicated PCIe GPU then we are in an unwinnable situation and we are guaranteed to run dramatically slower. Remember how ARM has added the family of FEAT_LRCPC1/2/3 extensions from before to improve x86 memory model emulation? This is what happens when we hit an edge-case that isn’t supported. All of these extensions add new instructions to handle loading memory using x86-TSO memory model semantics but none of them solve storing to write-combine memory with x86-TSO semantics. All the way from ARMv8.0-a our store instructions use the regular store-release instructions regardless of the backing memory type. The only way for FEX to work around this problem is to selectively disable TSO-emulation when it becomes an issue, so x86 emulation platforms with PCIe GPUs will always be a worse experience than UMA. At least until we get another FEAT_LRCPC4 or similar to resolve the issue.

For users on UMA systems then rejoice, there’s a workaround for gaming that we use to improve performance. Because we know when a platform supports cache-coherent CPU and GPU combinations, we can have the video driver always use cached buffers and never encounter this problem. NVIDIA already does this on their Tegra platforms, Snapdragon has been supporting this since at least Adreno 600 class GPUs, and there are many Mali platforms where this is also the case. We have a Adreno Turnip patch that ensures when FEX is running, we never hit uncached memory for platforms that support it. A funny thing is that since Asahi users have a hardware TSO bit, they just naturally don’t encounter this problem in the wild, but getting a PCIe GPU on to that platform is a different story altogether. There’s also a fun quirk where Radeon GPUs on ARM platforms hide all write-combine memory to instead be write-back but we’ll talk about that another time.

Looking towards a brighter future

After that marathon of an article we hope you have a better understanding of some of the challenges that emulating the x86-TSO memory model brings. Where we started with ARMv8.0 as a minimum spec and where the hardware has provided dramatic improvements over the years in nothing short of astounding. While not all of the edge-cases are yet resolved at the architecture level, it looks like there is a genuine commitment across the ecosystem for trying to improve the worst cases. We have various vendors solving some parts of the problem and moving the needle forward for better compatibility. Maybe in another decade as we look back at this time we’ll laugh about the problems we were encountering now, while enjoying some quality x86 games that will never see a port to ARM hardware. Keeping the legacy of the PC gaming ecosystem alive, regardless of where we might end up playing it.

The Daily Front Page 7 of 28
Friday, September 18, 2026 The Daily Front No. #260918 — The Lost Lexicon
article

Pre-Greek: The lost language hidden within Ancient Greek

by axiologist·▲ 174 points·67 comments·linguisticdiscovery.com ↗
Around 1,000 Greek words can’t be traced to any Indo-European language.

Around 1,000 Greek words can’t be traced to any Indo-European language. Where do they come from?

Pre-Greek: The lost language hidden within Ancient Greek

Ἦσαν οἱ Πελασγοὶ βάρβαρον γλῶσσαν ἱέντες.
Êsan hoi Pelasgoì bárbaron glôssan híentes.
‘The Pelasgians spoke a language which was not Greek.’

~ Herodotus (Histories 1.57)

There are about 1,000 words in Ancient Greek that can’t be traced back to any Indo-European language or any other known language in the region. These lexical enigmas even include some words and names you might think of as quintessentially Greek:

  • λαβύρινθος (labýrinthos) ‘labyrinth’
  • ἐλαία (elaía) ‘olive (tree)’
  • Ἀχιλλεύς (Akhilleús) ‘Achilles’
  • Ὀδυσσεύς (Odysseús) ‘Odysseus’
  • Κόρινθος (Kórinthos) ‘Corinth’
  • Ἀθῆναι (Athênai) ‘Athens’ / Ἀθήνη (Athḗnē) ‘Athena’ (goddess of wisdom)
  • Ἄρης (Árēs) ‘Ares’ (god of war)

Despite centuries of effort, linguists have not been able to trace these words back to Proto-Indo-European (PIE), the parent language from which all the other Indo-European (IE) languages descend. Nor have they identified words in neighboring languages that Greek might have borrowed them from.

Where, then, do these words come from? In this issue of the Linguistic Discovery newsletter, we’ll dig through layers of linguistic history to reveal the mystery language that influenced Ancient Greek.

The approximate present-day distribution of native speakers of the eight branches of the Indo-European language family within Europe and Asia.

The approximate present-day distribution of native speakers of the eight branches of the Indo-European language family within Europe and Asia. (Wikipedia: Indo-European languages)

Paleo-European / Old European Languages

All languages contain a fossil record of their history.

While Indo-European languages have been around for a long time, and Greek has likewise been spoken in Greece for several millennia, Indo-European languages haven’t always blanketed the majority of Europe like they do today. At one point the Indo-Europeans were pastoralist newcomers to a strange land, one that had been occupied by Neolithic farming societies for millennia.

Early Indo-European migrations according to the prevailing Kurgan Hypothesis, which places the Indo-European homeland north of the Caucasus in the Pontic–Caspian Steppe.

Early Indo-European migrations according to the prevailing Kurgan Hypothesis, which places the Indo-European homeland north of the Caucasus in the Pontic–Caspian Steppe. (Wikipedia: Proto-Indo-European language)

The peoples who preceded the Indo-Europeans—called Early European Farmers—spoke any number of Paleo-European (a.k.a. Old European) languages. Very few of these left any written evidence (just Aquitanian, Etruscan, Iberian, and a handful of others), and only one survives today (Basque). Even so, these Old European languages left an indelible mark on the Indo-European languages that succeeded them.

Map of known Paleo-European (Old European) languages, including ones that are directly attested with written evidence and those that are only indirectly attested.

Map of known Paleo-European (Old European) languages, including ones that are directly attested with written evidence and those that are only indirectly attested. (Wikipedia: Paleo-European languages)

Imagine you’re a band of early Hellenic (Greek) speakers whose ancestors had migrated over generations from the Pontic–Caspian Steppe into the eastern Mediterranean. Along the way, those ancestors encountered all sorts of novel foods, plants, animals, and terrain—things like olives, hyacinth, figs, and cypresses. Even the sea was unfamiliar at one point, migrating as your ancestors did from a landlocked steppe. What do you call all these things?

One option is to recruit existing vocabulary for new purposes. For example, when the Greeks first described hippopotamuses in the mid-5th century BCE, they cleverly called them ‘river horses’ (ἵππος ὁ ποτάμιος [híppos ho potámios] ‘the horse of the river’ [lit. ‘the riverine horse’]; Herodotus, Histories 2.71, c. 440s BCE). This later gave rise to the compound ἱπποπόταμος (hippopótamos) ‘hippopotamus’ (Dioscorides, De materia medica 2.25, c. 50–70 CE). The Greeks even applied this strategy to the sea! One of the words the Ancient Greeks used for ‘sea’ was πόντος (póntos), which comes from PIE *póntoh₁s ‘path; road’. Every other branch of Indo-European kept the original ‘path/way’ meaning (e.g. Latin pōns ~ pontis ‘bridge’, from which English gets pontiff and pontificate), but Greek extended it to mean ‘sea’, apparently in the sense of ‘the water-way; the sea-crossing’. Another Ancient Greek word used poetically to refer to the ‘sea’ was ἅλς (háls), whose primary meaning was actually ‘salt’. The sea, then, was ‘the salt’.

This strategy of coining new terms for novel experiences was, however, the minority one. Why invent new words for things the local Paleo-Europeans already had names for? Why not simply adopt the local name? That’s exactly what Indo-Europeans did as they settled across the continent, casually borrowing hundreds of words for plants, animals, types of terrain, places, and other natural phenomena. Thus it’s no surprise that the primary Greek word for ‘sea’ is actually θάλασσα (thálassa → English thalassocracy ‘rule of the sea’, used for maritime empires), which, unlike póntos and háls, has no firm Indo-European etymology. Linguists have long tried to connect this word to Indo-European, but no proposal has won wide acceptance. This is exactly what we’d expect: because the Proto-Indo-European language arose in the landlocked steppes, it had no need for—and has no securely reconstructable word for—‘sea’. (Linguists have, however, reconstructed terms for snow, cold, wolves, bears, and honey—all consistent with the temperate, continental interior.) Its daughter branches only acquired sea vocabulary later. As they reached the coasts, each language innovated or borrowed its own word for ‘sea’: English sea, Latin mare, Greek thálassa, etc.

The same process happened for all the other items I mentioned earlier:

  • olive: ἐλαία (elaía) ‘olive (tree)’: A very early borrowing which existed in Proto-Hellenic as *elaíwā and Mycenaean Greek (in Linear B) as 𐀁𐀨𐀷 (e‑ra‑wa), but has no clear Indo-European etymology. This same Pre-Greek word was also the source of the Etruscan word *𐌄𐌋𐌄𐌉𐌅𐌀 (*eleiva) ‘olive’, where it was borrowed by the Romans as olīva (the source of the English words olive and oil).
  • hyacinth: ὑάκινθος (hyákinthos) ‘hyacinth’: Again, no clear Indo-European etymology exists. English borrowed the word hyacinth from Greek. Hyacinthus was probably a Pre-Greek god who was superseded by Apollo, and was portrayed as his lover in Greek mythology.
  • fig: Greek borrowed several words for ‘fig’, including ἐρῑνεός (erīneós) ‘wild fig-tree’, ὄλυνθος (ólynthos) ‘wild fig’, and σῦκον (sŷkon) ‘fig’. These too have no Indo-European precedent.
  • cypress: Appears in Mycenaean Greek (in Linear B) as 𐀒𐀠𐀫 (ku‑pa‑ro) and Ancient Greek as κυπάρισσος (kypárissos). Lacks an Indo-European etymology. English borrowed the word cypress from Greek.

Because the Proto-Indo-Europeans lived outside the native range of these plants, their language lacked words for them. They only borrowed the words upon arriving in the eastern Mediterranean. Speakers coin or borrow new words as they need them, and migration events often trigger such a need, leaving a trace of that migration in the language’s lexicon (and grammar, but that’s a point for another article). In this way, all languages contain a fossil record of their history. They are a catalog of what the population lacked when they arrived in an area or came into contact with new peoples. In this article, we’ll use that linguistic fossil record to discover what the language that existed in the Aegean prior to the Greeks—which we’ll call Pre-Greek—was like.

Substrates, Superstrates, & Adstrates

Migration is just one piece of the puzzle, however. When one language disappears and leaves only traces of existence in other languages in the region, this is typically the result of social stratification: one more socially dominant community (militarily, economically, technologically, politically, etc.) imposes its language on another less dominant one. Linguists refer to this situation as linguistic stratification: the socially dominant language is called the superstrate language, while the less dominant variety is called the substrate language. This metaphor likens languages and population movements to geological strata, where each new population to enter an area adds another layer to the linguistic history of the region.

Languages as geological strata

Languages as geological strata

In the course of linguistic stratification, both languages naturally influence one another, but the substrate language is often lost entirely, leaving only traces within the superstrate language. So when we have indirect evidence of a historically unknown language that influenced a historically known language—like the way the Pre-Greeks influenced Ancient Greek—we are usually looking at a case of linguistic stratification. Indeed, the early Greeks, and especially the Mycenaeans (c. 1750–1050 BCE), are associated with major advances in technology (including writing), economy, and social organization in the Aegean. Another example of linguistic strata can be seen in the diversification of Latin into its descendant Romance languages (Spanish, Portuguese, French, Italian, Romanian, and several dozen others). While the dominance of Roman culture gradually erased many of the local languages of Europe and the Mediterranean, each of those local languages nonetheless influenced the Latin spoken in their respective regions. Over time, those various flavors of Latin became distinct languages.

One other type of linguistic “stratum” (which breaks the geological layering metaphor somewhat) is an adstrate language—a language which influences a neighboring population. Like substrates, this often happens in cases of social stratification, where the adstrate community is considered socially superior or dominant to the surrounding communities (but it can also occur due to simple geographic proximity). Due to the cultural dominance of the Roman Empire throughout Europe and the Mediterranean, for instance, Latin functioned as an adstrate language for many of the languages that sat across the līmes (the Roman frontier). Even though the Germanic-speaking territories east of the Rhine and north of the Danube never fell under Roman administration, centuries of trade, military contact, and cultural prestige still left a layer of adstrate borrowings in Proto-Germanic:

Latin Meaning Proto-Germanic English German
cāseus ‘cheese’ *kāsjus cheese Käse
cattus ‘cat’ *kattuz cat Katze
coquīna ‘kitchen’ *kokīnā kitchen Küche
piper ‘pepper’ *pipar pepper Pfeffer
pondō ‘a pound (by weight)’ *pundō pound Pfund
saccus ‘sack, bag’ *sakkiz sack Sack
via strāta ‘paved road’ *strātā street Straße
vīnum ‘wine’ *wīnam wine Wein

In this article, we’ll be hunting for both a substrate and an adstrate that influenced Ancient Greek.

How to identify borrowed words

Any given piece of linguistic evidence can look quite flimsy on its own; it is only in the context of broader patterns that such evidence becomes convincing.

At this point you might reasonably be asking, “How do we know that some substrate language influenced Ancient Greek? How do we know that the mystery words were borrowed? Can’t they just be random exceptions?” This exact problem has vexed historical linguists working on the Pre-Greek problem for over a century. Nonetheless, we not only have good reasons for thinking a word might be borrowed, but we can even be highly certain when a word can’t come from Indo-European. Sometimes, the answer is obvious because there’s a similar word with a clear etymology sitting right in a nearby language. The Ancient Greek word ἔβενος (ébenos) ‘ebony’, for example, has no Indo-European etymology, but just across the Mediterranean Sea was the Egyptian word 𓍁 𓈖 𓏭 𓆱 hb‑n‑j‑WOOD (hebni) ‘ebony’, and we know that the Ancient Greeks traded extensively with the Egyptians. But the evidence for a loanword is often more subtle. Below are some of the different types of evidence (or non-evidence) linguists use to determine which words are borrowed (in Ancient Greek in particular).

Words which lack an Indo-European etymology

The field of linguistics continues to refine Indo-European etymologies, either via newly-discovered tablets and inscriptions, or improved availability of existing data in the form of etymological dictionaries and online databases, allowing them to bring more (and more accurate) data to bear on Indo-European history. Despite these advances, the mystery words of Ancient Greek have resolutely resisted etymologization. If the number of potential non–Indo-European mystery words were gradually decreasing with each new discovery or advance, the lack of an Indo-European etymology might not mean anything. It’d be reasonable, in that case, to view them as a temporary problem that we’d expect to solve eventually. But there is no indication that future discoveries will ever uncover Indo-European origins for these words. If a word existed in Proto-Indo-European, we should expect to find cognates of it in several of its child languages, and that’s just not the case for these words.

Some other words without a firm Indo-European etymology include:

  • θησαυρός (thēsaurós) ‘treasury, treasure’ → English: thesaurus
  • ἴδη (ídē) ‘wood, wooded hill’ → place name: Mount Ida
  • βασιλεύς (basileús) ‘king’ → English basilica, basil (because it was associated with royalty)

The last example is especially convincing because Indo-European already has a word for ‘king; ruler’ (PIE *h₃rḗǵs → Latin rēx ‘king’), and basileús doesn’t resemble it. It is thus likely that Greek borrowed basileús as a term for the local rulers when they arrived in the Aegean.

Also convincing are words which surface in Latin or other, non–Indo-European languages when the words in those languages also lack etymologies. For example, Greek οἶνος (oînos) ‘wine’ has confamiliars¹ in:

  • Latin: vinum
  • Hittite: winija‑
  • Armenian: gini
  • Ethiopian (Ge’ez): wain
  • Hebrew: jajin

None of these words can be reconstructed back to a word in their protolanguages (Proto-Indo-European or Proto-Semitic), so the reasonable conclusion is that the original word for ‘wine’ must have come from some Paleo-European substrate language. The word then spread around the Mediterranean as part of the diffusion of wine culture under the influence of the Phoenicians (from c. 1000 BCE) and subsequently the Greeks (from c. 600 BCE).

Words which are phonologically impossible for Indo-European

Sometimes there is no Proto-Indo-European word—reconstructed or imagined—that could possibly produce a word in Ancient Greek if it followed the normal sound changes that happened during the evolution from PIE to Greek. For example, the /dn/ at the beginning of δνόφος (dnóphos) ‘darkness’ cannot be explained by any Indo-European sound law. Ancient Greek also had the suspiciously similar word κνέφας (knéphas), which also means ‘darkness’ (and ‘twilight; dusk’) and which also starts with an impossible sound sequence for inherited Indo-European words in Greek, /kn/. Add to this the words ζόφος (zóphos) and ψέφας (pséphas) which also mean ‘darkness’ and lack clear Indo-European origins, and it’s obvious that we’re looking at a set of related words that must have all come from a Pre-Greek word meaning ‘darkness’, and that word must have started with a sound or sequence of sounds that was difficult for the early Greeks to pronounce: dnóphos ~ knéphas ~ zóphos ~ pséphas. In fact, most of the Ancient Greek words that start with ⟨ψ⟩ /ps/ (a sound sequence that was impossible in Proto-Indo-European) are mystery words without Indo-European etymologies. Linguists have proposed all sorts of connections for each of these words, so that there are as many hypothesized etymologies for them as there are researchers. This lack of consensus is itself suggestive that the words are probably Pre-Greek, but the main point here is that these words do not fit the sound patterns of Indo-European.

Another indication that a word is Pre-Greek is the sound β /b/, which was rare or possibly absent in Proto-Indo-European, but common in Pre-Greek loanwords:

  • ἄβλαροι (áblaroi) ‘wood’
  • ἀβύρβηλος (abýrbēlos) ‘much; burdensome; big; heavy’
  • ἀρβύλη (arbýlē) ‘a kind of tough shoe or half-boot’
  • ἀτάρβακτος (atárbaktos) ‘unafraid’
  • βάρβιλος (bárbilos) ‘seedling peach tree’
  • θόρυβος (thórybos) ‘noise; uproar, clamor (esp. of a crowd); tumult, confusion’
  • κίβαλος (kíbalos) ‘messenger, courier’

There are also many words which seem like they could derive from Indo-European, but don’t obey the expected sound patterns. For example, Ancient Greek had two similar words τάφος (táphos) ‘funeral; grave’ (→ English epitaph, cenotaph, taphonomy) and τύμβος (týmbos) ‘grave; tombstone’ (→ English tomb, via Latin). The first of these comes from PIE *dʰm̥bʰos ‘tomb’ and follows all the normal sound changes you’d expect from PIE to Ancient Greek. But týmbos is a dupe: it looks related, but it doesn’t follow the expected sound changes. So either the similarity between the two words is purely coincidental, in which case týmbos was probably borrowed from a Pre-Greek, non–Indo-European language, or týmbos does in fact derive from PIE, but through some unknown, intermediate Indo-European language, which Greek then borrowed it from.

Words with lots of phonetic variation

When words are borrowed from one language into another, they often initially arrive with a slew of different pronunciations. Sometimes this is because the donor language has sounds that the recipient language does not, so speakers of the receiving language have to approximate them as best they can. This was likely the case with the ‘darkness’ words we saw above: variation like dnóphos ~ knéphas ~ zóphos ~ pséphas suggests that there was an initial consonant or consonant cluster in that word which the Greeks had difficulty pronouncing.

Often, however, variation in borrowed words occurs simply because speakers of the recipient language are unfamiliar with the word. The Powhatan word makasin, for example, contains only sounds familiar to English speakers, but when English colonists borrowed the word, the range of spellings included maccaseene, makissin, mogason, moccasin, and mogozeen. It took a century or so for the spelling and pronunciation to standardize. Random variations in spelling and pronunciation like this are often a clear indicator of borrowed words.

The word “random” is important there. We also see variation in spelling whenever languages undergo broader sound changes, but this variation is regular and predictable, not random. For example, Proto-Greek had a set of palatal sounds */ty, tʰy, ky, kʰy/ which all became *t͡s or possibly *t͡ʃ. In the Attic, Boeotian, and Euboean dialects of Greek, that *t͡s / *t͡ʃ sound became /tt/, but in the Ionic, Doric, and other dialects of Greek it became /ss/. At the beginnings of words, it became /t/ ~ /s/. This sound change yielded pairs like the following:

Ionic Attic Meaning
θάλασσα thálassa θάλαττα thálatta ‘sea’
μέλισσα mélissa μέλιττα mélitta ‘bee’
σήμερον sḗmeron τήμερον tḗmeron ‘today’
τέσσαρες ssares τέτταρες ttares ‘four’

The variation in these words is not evidence that they were borrowed, because this is regular, patterned dialectal variation. Notice that our Pre-Greek word thálassa / thálatta ‘sea’ is in this list too. This is a borrowing, but it’s not the /ss/ ~ /tt/ alternation that tells us that. Evidence of its loanword status comes from elsewhere. But its participation in the /ss/ ~ /tt/ alternation does tell us that it was borrowed very early in the history of Greek, before this sound change took place (the period of Proto-Greek, c. 2200–1700 BCE). It also tells us that it contained one of the palatal sounds that later became /ss/ ~ /tt/.

That’s all to say: mere variation is not sufficient evidence of borrowing. The variation must not be attributable to regular sound changes in the language; it must be largely random instead. In fact, part of the evidence that thálassa ‘sea’ is borrowed is that it also shows random variation in addition to the regular dialectal variation. It appears once in Hesychius’ glosses as δαλάγχαν (dalánkhan), exhibiting /d/ instead of /tʰ/ for the first consonant, and with a consonant cluster /nkʰ/ instead of /ss/ ~ /tt/. But remember that the /ss/ ~ /tt/ sounds originally derived from some type of palatal consonant. This strongly suggests that what Hesychius wrote as ⟨γχ⟩ /nkʰ/ was an attempt to represent a palatal /ky/. (Hesychius was compiling the glossaries of earlier writers.)

Here are some other Ancient Greek words that we know are borrowings in part because of their variant pronunciations:

  • Ὀδυσσεύς (Odysseús) / Ὀλυσσεύς (Olysseús) ‘Odysseus / Ulysses’
    • The version with /l/ is the one that reached Latin (Ulixes), and this is why the epic hero is called Ulysses in the Roman tradition and Odysseus in the Greek tradition.
  • λαβύρινθος (labýrinthos) / Mycenaean 𐀅𐁆𐀪𐀵𐀍 da‑pu₂‑ri‑to‑jo (daphurinthoio) ‘labyrinth’
    • This word has the same /l/ ~ /d/ alternation as Ulysses ~ Odysseus, as well as a /b/ ~ /pʰ/ alternation.
  • σῦκον (sŷkon) / Boeotian τῦκον (tŷkon) ‘fig’
    • Exhibits an /s/ ~ /t/ alternation. Latin and Old Armenian may have also borrowed this word from the Pre-Greek substrate as fīcus ‘fig tree’ (→ English ficus) and թուզ (tʻuz) ‘fig’, respectively.
  • ἀσφάραγος (aspháragos) / ἀσπάραγος (aspáragos) ‘asparagus’
    • Latin got the /p/ variant and gave it to English.

In other cases we see variation in the presence of an initial /w/ sound, which is especially neat because Greek originally had a /w/ sound but then lost it! In Mycenaean Greek (c. 1400–1200 BCE), the /w/ was still present, as shown in words like 𐀷𐀙𐀏 (wa‑na‑ka) ‘king; lord’. Many dialects (Aeolic, Doric [including Cretan], Elean, Boeotian, etc.) retained that /w/ and wrote it with the letter digamma ⟨ϝ⟩ as ϝάναξ (wánax) well into the Classical period (5th–4th C BCE). But the Attic and Ionic dialects lost /w/ before the emergence of alphabetic writing (8th C BCE), so they wrote the word as ᾰ̓́νᾰξ (ánax). But even though Homer never writes a digamma, the meter of his poems often behaves as though the /w/ were still there, because they were originally composed when it was!

In any case, notice how the /w/ was simply deleted from wánax to yield ánax. But for many Pre-Greek loanwords, we see attempts to render a /w/ sound even in the dialects that lost it! The city of Axos on Crete, for example, is written as ϝαξός (Waxós) in the Cretan dialect but Ὀαξός (Oaxós) outside Crete. If this word had already existed in Greek, the /w/ would have simply disappeared in the non‑/w/ dialects, like it did for ánax. Instead, those dialects attempted to write the /w/ sound they heard when they adopted the name of the place from the Pre‑Greeks. Since the /w/ sound and ⟨ϝ⟩ letter were already lost to them by that point, they approximated it with ⟨Ὀ⟩ /o/ instead. This shows that the Greeks adopted the Pre-Greek place name after /w/ had disappeared in most dialects.

We see the same attempt to represent /w/ in variations of the word for ‘hyacinth’:

  • standard literary form: άκινθος (hyákinthos)
  • Cretan: ϝάκινθος (wákinthos)
  • Cretan variant: βάκινθος (bákinthos)

Here, the /w/ surfaces as ὑα‑ /hya/, ϝα‑ /wa/, and βα‑ /ba/.

In sum, random variation is often a key piece of evidence that a word is borrowed.

Words in certain semantic categories

As mentioned above, certain categories of words are likely to be borrowed by a people migrating into an area, namely terms for plants, animals, terrain, and some simple tools. Thus, if a word has no strong etymology and it refers to either a natural phenomenon or local material culture or social practices, it’s a good candidate for a loanword. In Beekes 2010 etymological dictionary of Greek, he lists 1,106 words as coming from Pre-Greek. Of those, the largest categories of words are for flora and fauna. (Beekes does not include most place names in the dictionary.) The table below shows the total category counts. You can see that, with the exception of human physiology, these are exactly the types of words we’d expect to be borrowed from Pre-Greek.

  • Flora: 178
  • Fauna: 180
  • Landscape & Natural Phenomena: 37
  • Minerals: 27
  • Agriculture: 38
    • Viniculture: 20
  • Human Physiology: 81
  • Attire & Jewelry: 21
  • Equipment & Utensils: 154
  • Construction: 34
  • Culture: 51
    • Musical Instruments / Arts: 18
    • Religious Festivals / Feasting: 8
    • Divine/Numinous Beings, Priests, & Temples: 12
    • Theonyms / Mythical Characters: 26

What I find especially interesting is that the Greeks adopted almost their entire pantheon from Pre-Greek civilization! Except for Zeus, Eos, Pan, Hades, and perhaps Poseidon, all the names of Greek gods cannot be reconstructed to Proto-Indo-European, and are likely loanwords from Pre-Greek substrate language (Verhasselt 2009: 211).

Pairs of words that mean the same thing

When a language has numerous pairs of words that both refer to roughly the same concept, this is often a sign of extensive borrowing. This is especially true when those words are concentrated in one or a few semantic fields. For example, English famously has pairs of words for animals and the foods that come from those animals, where the name of the animal derives from English but the name of the food derives from French:

Animal (from English) Meat (from Norman French)
calf veal
chicken / fowl poultry
deer venison
cow beef
pig / swine pork
sheep mutton

Originally, both the English and French words referred to the animal—boef was simply the Norman French word for ‘cow’; porc was the word for ‘pig’, and so on. But due to the social stratification between French speakers and English speakers after the Norman Conquest, the French words eventually came to refer to the food rather than the animal.

In cases of population migration or invasion, such as the arrival of the early Greeks in the Aegean or the Norman invasion of England, the result is often that words from each language are associated with different social contexts or registers, reflecting the uneven mixture of the two cultures. This means that two or more words for the same concept can exist side-by-side in a language, each being used in different contexts. (Notice that this is also a case of linguistic stratification, where French was the superstrate and English was the substrate.)

Ancient Greek exhibited this pattern with words relating to weaving. Alongside Indo-European words for weaving and textiles are a set of alternate words for the same things, which lack any clear Indo-European etymology. Pairs of words like the ones in the table below are a strong indication of substrate borrowing.

Converging Evidence

Obviously, none of these methods of identifying borrowed words is airtight. Any given word might be a random exception whose explanation is lost to history. And no single feature listed above is diagnostic by itself. Instead, we determine borrowings through the confluence of evidence, and the way certain words pattern together in groups. We look for words that participate in recurrent or systematic patterns.

Consider the following Ancient Greek words:

  • ἄσβολος (ásbolos) ‘soot’
  • σποδός (spodós) ‘ashes; ember’
  • ψόλος (psólos) ‘soot; smoke’

All these words lack clear etymologies, and show the same /d/ ~ /l/ variation we saw earlier. The initial /a/ of ásbolos cannot be explained by normal sound changes, and the consonant cluster /sb/ is anomalous for Proto-Indo-European. psólos reverses the order of the consonants to /ps/, which aligns it with the many other Pre-Greek loanwords that start with an initial /ps/. (Remember from earlier that almost all Greek words beginning with /ps/ have turned out to be loanwords.) Add to this the fact that PIE has a word for ‘ash’, *h₂eh₁s‑, which looks nothing like any of these forms, and is the source of the Greek reflex ἄζω (ázō) ‘to dry up; to parch’. This is thus another case of duplicate words for the same thing—one borrowed, one native. Taken together, the evidence shows that these three words are undoubtedly different variants of some non-Greek borrowing.

Or consider the words ἄνθραξ (ánthrax) and κάνδαρος (kándaros, attested in Hesychius’ glosses), which both mean ‘charcoal’. While there is a potential but shaky PIE etymology for kándaros (*(s)kand‑ ‘to shine; to glow’), there are no good ones for ánthrax. There is also no known process by which an initial /k/ gets added to or deleted from a word in the journey from PIE to Ancient Greek, and no explanation for the variation between /tʰ/ and /d/. In fact, both the variable initial /k/ and the ‑ak suffix in ánthrax (pronounced /án.tʰraks/) are features found in many Pre-Greek loanwords. Given this converging evidence, the most likely explanation for these words is that they both come from a single Pre-Greek word.

Aside: The multifaceted nature of the evidence for any given word’s etymology makes it difficult for linguists to explain, succinctly, how we know where a word comes from, or how we reconstruct words in languages that were never written down. We rely on an interconnected web of principles and patterns that we’ve deduced from the written historical record over three centuries of research in historical linguistics. Any given piece of evidence can look quite flimsy on its own; it is only the context of broader patterns that such evidence becomes convincing. This state of affairs is why, I believe, so many people are suspicious of historical reconstructions in linguistics.

To sum up, linguists actually have a surprising number of tools at our disposal for resolving just which words in a language are borrowed.

Where was the Pre-Greek language spoken?

We can do even better than simply identifying specific words that were probably borrowed: in some cases we can show that the words must have come from the same, specific substrate language (or closely-related languages). We can list specific features the language might have had, and even roughly where it was spoken, though we have little or no direct written evidence of it.

The most definitive piece of evidence that we’re looking at borrowings from a specific substrate language is when we see similar-looking pieces of words across many of the borrowings. This may represent a prefix, suffix, or entire word in the substrate language that got borrowed wholesale along with the rest of the word/phrase. For example, place names in England are rife with fossilized affixes that were originally words in substrate or adstrate languages:

  • Brittonic penn ‘head, hill, headland’ → English Pen‑
    • Penrith
    • Penzance
    • Pendle Hill
    • Penge (London)
  • Latin castra ‘camp’ → English ‑chester / ‑cester / ‑caster
    • Chester
    • Manchester
    • Winchester
    • Leicester
    • Worcester
    • Gloucester
    • Lancaster
  • Old Norse ‘farmstead’ → English ‑by
    • Derby
    • Rugby
    • Whitby
    • Grimsby
    • Selby

Each of these suffixes acts as a record of a particular stratum in English’s history: the pre–Anglo-Saxon Celtic layer, the Roman occupation layer, and the Old Norse (Viking) layer. Toponyms (place names) are an especially rich source of information about extinct substrate languages because migrating populations tend to adopt existing place names, especially for names of settlements. (At least half of U.S. states have names derived from an indigenous language, for example.) Toponyms tend to be relatively stable over time, at least compared to other words in the lexicon.

Places like Corinth, Knossos, Athens, and Lemnos had already been settled for thousands of years when the first Greeks arrived in the Aegean. So it shouldn’t be a surprise that all four of those names are indisputably Pre-Greek borrowings, and that they each contain suffixes that recur time and again in placenames across the Aegean! The most famous of these is the ‑νθ (‑nth) suffix, which appears, inter alia:

  • Ἐρύμανθος (Erýmanthos): Erymanthos, a mountain range in the Peloponnese peninsula
  • Ὑάκινθος (Hyákinthos): Hyakinthos/Hyacinth, a mountain in Attica
  • Κήρινθος (Kḗrinthos): Kerinthos, an ancient town in Euboea
  • Κόρινθος (Kórinthos): Corinth, a city on the isthmus joining the Peloponnese peninsula to mainland Greece
  • Ὄλυνθος (Ólynthos): Olynthus, an ancient city in present-day Chalcidice
  • Τίρυνς, Τίρυνθος (Tíryns, Tírynthos): Tiryns, an ancient Mycenaean city in the Peloponnese peninsula

Also widely attested are the ‑σσ (‑ss) suffix in Knossos, the ‑ν (‑n) suffix in Athens, and the ‑μν (‑mn) suffix in Lemnos. Below is a sampling of place names that these suffixes appear in.

  • ‑σσ (‑ss)
    • Ἁλικαρνασσός (Halikarnassós): Halicarnassus, an ancient city on the coast of Anatolia (modern Turkey)
    • Κνωσσός (Knōssós): Knossos, a Minoan palace on the island of Crete
    • Λάρισσα (Lárissa): Larissa, a city in Thessaly
    • Παρνασσός (Parnassós): Parnassus, a mountain in central Greece and home to Delphi
  • ‑ν (‑n)
    • Ἀθῆναι (Athênai): Athens, capital city of Greece
    • Λέρνη (Lérnē): Lerna, a region of springs and a former lake in classical Greece, south of Argos
    • Μυκῆναι (Mykênai): Mycenae, an ancient Mycenaean city in the Peloponnese peninsula
    • Μύκονος (Mýkonos): Mykonos, an island in the southern Aegean Sea
  • ‑μν (‑mn)
    • Λάρυμνα (Lárymna): Larymna, a port town in Phthiotis
    • Λῆμνος (mnos): Lemnos, an island in the northern Aegean Sea
    • Λεπέτυμνος (Lepétymnos): Lepetymnos, a mountain on the island of Lesbos
    • Μήθυμνα (Mḗthymna): Methymna, a town on the island of Lesbos
    • Ῥίθυμνα (Rhíthymna): Rethymno, a city on the island of Crete

(The reason these suffixes are in the middle of these words is because Greek speakers added the normal Greek case suffixes on the end when they borrowed them.)

Because these are toponyms tied to specific places, we can even map the region where this Pre-Greek substrate was originally spoken! Below is a map of Pre-Greek toponyms showing the approximate area where the Pre-Greek substrate was spoken.

Greek toponyms (placenames) borrowed from a Pre-Greek substrate.

Greek toponyms (placenames) borrowed from a Pre-Greek substrate. This map only shows toponyms that are widely acknowledged to be Pre-Greek with a relatively high degree of certainty, and only toponyms with the well-attested Pre-Greek suffixes ‑μν‑, ‑ν‑, ‑νθ‑, and ‑σσ‑/‑ττ‑ (‑mn‑, ‑n‑, ‑nth‑, and ‑ss‑/‑tt‑). Other potential Pre-Greek suffixes are excluded. Other toponyms widely acknowledged to be borrowings but which are not firmly linked to the Pre-Greek substrate, such as Κρήτη (Krḗtē) ‘Crete’, are also excluded. Finally, some place names were omitted for reasons of space, especially those in and around Athens.

You can view the entire dataset of Pre-Greek toponyms, including less certain ones and ones that were excluded from this map, here.

How to identify a Pre-Greek suffix

The intellectually cautious among you are now probably rightly asking, “How do we know these are suffixes and not just common sounds in these words?” The strongest piece of evidence for these suffixes is that in some cases, versions of the word can be found both with and without the suffix. The single best example of this kind of alternation is the toponym Κήρινθος (Kḗrinthos). When used as a normal word rather than a place name, kḗrinthos meant ‘bee bread (perga)’, while κηρός (kērós), without the ‑nth suffix, meant ‘beeswax’. Kerinthos, then, was very likely named such because it was known for its beekeeping! We can be even more certain that this word for ‘beeswax’ was borrowed from some other language because it’s confamiliar with similar words around the Mediterranean:

  • Latin: cēra ‘wax; beeswax’
  • Albanian: qiri ‘candle’
  • Lithuanian / Latvian: korỹs / kāre ‘honeycomb’

Other Pre-Greek words that appear with and without their suffixes include:

  • ∅ ~ ‑nth
    • Ἀμαρυσία (Amarysía), one of the names for Artemis
    • Ἀμάρυνθος (Amárynthos), a township on Euboea, and another name for Artemis
  • ∅ ~ ‑mn ~ ‑n
    • Δίκτη (Díktē)
    • δίκταμνον (díktamnon) ‘dittany (of Crete)’
    • Δίκτυννα (Díktynna) ‘Cretan goddess’
  • ∅ ~ ‑ss ~ ‑n (and a possible ‑l suffix)
    • Μυκάλη (Mykálē)
    • Μυκαλησσός (Mykalēssós)
    • Μυκῆναι (Mykênai)
  • ∅ ~ ‑m (another Pre-Greek suffix)
    • πύργος (pýrgos) ‘tower’
    • Πέργαμον (Pérgamon)
    • πέργαμος (pérgamos) ‘citadel’ (esp. of Troy)
  • ‑nth ~ ‑ng (another Pre-Greek suffix)
    • μήρινθος (mḗrinthos) ‘thread’
    • σμῆριγξ (smêrinx) ‘hair’
  • ‑nth ~ ‑ng
    • ἕλμινθος (hélminthos)
    • ἔλμιγγος (élmingos) ‘worm’
  • ∅ ~ ‑nth
    • ὄροβος (órobos) ‘bitter vetch’ (an ancient grain legume)
    • ἐρέβινθος (erébinthos) ‘chickpea’

This last one is an especially important piece of evidence because the word for ‘chickpea’ (or very similar types of legumes) was also borrowed into other languages of Europe, sometimes with, sometimes without the ‑nth suffix:

  • Latin: ervum ‘bitter vetch’ (without ‑nth)
  • Proto-Germanic: *arwīt‑ ‘pea’ (with ‑nth, as ‑t)
  • Armenian: aṙowoyt ‘alfalfa’ (with ‑nth, as ‑t)

In general, whenever a place name has a non-toponymic meaning in addition to its toponymic meaning, and the etymology of that word is unknown, the word is undoubtedly Pre-Greek in origin. This is especially true if it also appears to contain one of the known Pre-Greek suffixes. Here are a few last examples like this:

Toponym Basic Meaning Suffix
Ἄψινθος (Ápsinthos) ἄψινθος (ápsinthos) ‘wormwood’ ‑nth
Ὑάκινθος (Hyákinthos) ὑάκινθος (hyákinthos) ‘hyacinth’ ‑nth
Καρδαμύλη (Kardamýlē) κάρδαμον (kárdamon) ‘cress’ ‑m
Κύαμον (Kýamon) κύαμος (kýamos) ‘bean’ ‑m
Ὄλυνθος (Ólynthos) ὄλυνθος (ólynthos) ‘wild or sterile fig’ ‑nth
Ὄθρυς, Ὄθρυος (Óthrys, Óthryos) ὄθρυς (óthrys) ‘mountain’
Σμίνθη (Smínthē) σμίνθος (smínthos) ‘mouse’ ‑nth

A corroborating piece of evidence for these suffixes is the simple fact that they occur not just in toponyms, but in many of the mystery words for plants, animals, foods, etc.—the same semantic categories of words that are most likely to be borrowed:

Greek Romanization Meaning Suffix
αἴγινθος aíginthos ‘name of a bird’ ‑nth
ἀσάμινθος asáminthos ‘stone bath tub’ ‑nth
βόλινθος bólinthos ‘European bison’ ‑nth
θάλασσα thálassa ‘sea’ ‑ss
κόρυνθος kórynthos ‘barley bread’ ‑nth
κυπάρισσος kypárissos ‘cypress’ ‑ss
λέβινθος lébinthos ‘a kind of beans’ ‑nth
μίνθη mínthē ‘mint’ ‑nth
νάρκισσος nárkissos ‘narcissus’ ‑ss
τερέβινθος terébinthos ‘terpentine’ ‑nth

Finally, toponyms ending in ‑nth and ‑mn show a distinctive root pattern, with a strong preference for open syllables (ones not ending in a consonant) and an almost total absence of consonant clusters in the middle of the word that start with a stop consonant. This is extremely unusual for Proto-Indo-European. As a rule, PIE roots had a CVC structure (C = consonant, V = vowel) and many consonant clusters—the expanded root template is (s)(C)CVC(C), yielding numerous consonant clusters. So the fact that words which show these Pre-Greek suffixes are also phonologically unusual for PIE is another piece of evidence that these words probably came from a non–Indo-European language.

Congratulations, now you know how to identify a Pre-Greek suffix! So now whenever you see an ‑nth in a Greek loanword in English like hyacinth, you’ll naturally think to yourself, “Oh neat! That word is from a Pre-Greek substrate language!” (Right?)

What was the Pre-Greek language like?

In addition to the Pre-Greek suffixes, there are some patterns that recur in potential loanwords with too great a frequency to be mere coincidence, and suggest that the Pre-Greek language may have had some prefixes as well: ἀ‑ (a‑), κ(α)‑ (k[a]‑), and σ‑ (s‑). The main evidence for these prefixes is that there are written records of these words appearing both with and without the initial sounds:

ἀ‑ (a‑)

There are around 65 words that show evidence of an a‑ prefix, many of which fall in the characteristic semantic domains of plants and crops (19 words), animals (10), as well as landscapes, foods, and the human body (smaller numbers). Interestingly, the a‑ prefix is not found on cultural terms like jewelry, armor, social hierarchy, professions, instruments, and words relating to the arts, etc.

  • γηθυλλίς (gēthyllís) ‘chives’ ~ ἀγασυλλίς (agasyllís) ‘plant that produces ammoniacum (gum resin)’
  • γέλγις (gélgis) ~ γλις (áglis) ‘garlic’
  • κάστον (káston) ~ καστος (ákastos) ‘maple’
  • κίρρις (kírris) ~ ἀκιρίς (akirís) ‘lamp’
  • κόρνοψ (kórnops) ~ ἀκορνοί (akornoí) ‘a kind of locust’
  • χραμαδοῖλαι (khramadoîlai) ‘tortoise’ ~ ἀχραδαμύλα (akhradamýla) ‘snail’
  • νηρίτης (nērī́tēs) ~ ἀναρίτης (anarítēs) ‘sea snail’
  • ῥωδιώς (rhōdiṓs) ~ ρωδιός (arōdiós) ‘heron’
  • (σ)καλαβώτης ([s]kalabṓtēs) ~ ἀσκάλαβος (askálabos) ‘spotted lizard; gecko’
  • στάχυς (stákhys) ~ σταχυς (ástakhys) ‘ear of corn’
  • στεροπή (steropḗ) ~ στραπή (astrapḗ) ‘lightning’
  • κύνωψ (kýnōps) ~ ἀχύνωψ (akhýnōps) ‘ribwort’

κ(α)‑ (k[a]‑)

The k(a)‑ prefix, like a‑, only appears on words for plants, animals, and simple technologies. There are 13 clear cases, including:

  • ἴχλα (íkhla) ~ κίχλη (kíkhlē) ‘thrush’
  • ἄρυα (árya) ‘walnuts’ ~ κάρυον (káryon) ‘nut’
  • ὄγχνη (ónkhnē) ~ κόγχναι (kónkhnai) ‘pear’
  • σκάνδιξ (skándīx) ‘chervil’ ~ κασκάνδιξ (kaskándix) ‘onion’
  • ἀπήνη (apḗnē) ~ πήνα (pḗnā) ~ καπάνα (kapā́na) ‘wagon’
    • Note the possible a‑prefix here.
  • ἄχλαξ (ákhlax) ‘small stones’ ~ χάλαζα (khálaza) ‘hail’ ~ κάχληξ (kákhlēx) ‘small stones’
  • ἀλινδέομαι (alindéomai) ~ καλινδέομαι (kalindéomai) ‘to roll about, wallow’

σ‑ (s‑)

Again, the majority of terms with s‑ are words for plants, animals, and body parts. Sometimes the s‑ occurs with an additional a‑ as well, suggestive of the presence of the a‑ prefix.

without a‑

  • γέλενος (gélenos) ‘asphodelus, narcissus’ ~ σχέλινος (skhélinos) ‘wild cypress’
  • κιδάφη (kidáphē) ~ σκιδάφη (skidáphē) ‘fox’
  • κίκερος (kíkeros) ~ σκίγκος (skínkos) ‘a type of lizard found in Asia Minor that is used as medicine’
  • κορδύ̄λη (kordȳ́lē) ~ σκορδύ̄λη (skordȳ́lē) ‘tumor; swelling’
  • βάταλος (bátalos) ‘a lewd man’ ~ σπάταλος (spátalos) ‘wanton; lascivious’
  • πέλεθος (pélethos) ~ σπέλεθος (spélethos) ‘dung’
  • φαττάγης (phattágēs) ~ σπατάγγης (spatángēs) ‘scaly ant-eater’
  • θριγκός (thrinkós) ‘topmost course of stones in a wall, cornice, frieze’ ~ στριγχός (strinkhós) ‘little wall; crown of a building’
  • τοπεῖον (topeîon) ‘rope, cord’ ~ στυππεῖον (styppeîon) ‘the coarse fiber of flax or hemp’
  • μήρινθος (mḗrinthos) ~ σμήρινθος (smḗrinthos) ‘cord, thread’
  • μύραινα (mȳ́raina) ~ σμύραινα (smȳ́raina) ‘moray eel’
  • μῖλαξ (mîlax) ~ σμῖλαξ (smîlax) ‘taxus’
  • φάγνος (phágnos) ~ σφάγνος (sphágnos) ‘salvia’
  • τρύχνον (trýkhnon) ~ στρύχνον (strýkhnon) ‘nightshade’

with a‑

  • φάρυγξ (phárynx) ‘throat, gorge, larynx, windpipe’ ~ σφάραγος (spháragos) ‘throat, gullet’ ~ ἀσφάραγος (aspháragos) ‘gully, chasm, deep trench, abyss’
  • κάλαφος (kálaphos) ~ ἀσκάλαφος (askálaphos) ‘an unknown bird, perhaps an owl’
  • καλαβώτης (kalabṓtēs) ~ σκαλαβώτης (skalabṓtēs) ~ ἀσκάλαβος (askálabos) ‘lizard, gecko’
  • κάμων (kámōn) ~ σκαμ(μ)ωνία (skam(m)ōnía) ~ ἀσκαμωνία (askamōnía) ‘scammony’

Was Pre-Greek one language or several?

Throughout this article, I’ve mostly talked about “the” Pre-Greek language. Indeed, the suffixes in the last two sections give us strong reason to believe that Ancient Greek borrowed the words containing those suffixes from a single Pre-Greek language or group of closely-related languages. Throughout the history of modern linguistics, there have been multitudinous names and theories as to the identity of this mystery substrate, including:

  • Pelasgian, an undocumented Indo-European language
  • Parnassian, a language in the Anatolian branch of Indo-European, related to Luwian
  • Aegean / Mediterranean, a non–Indo-European language which extended over a large part of the Mediterranean
  • the Kartvelian Theory, that Pre-Greek was related to Kartvelian languages (e.g. Georgian)
  • Pre-Greek, a non–Indo-European language spoken primarily in the region of Greece, following toponymic evidence

Even though each of these theories has been largely discredited except the last, each has advanced our knowledge of Pre-Greek in different ways. For example, researchers on this topic agree that the Pre-Greek substrate extended at least partially into Anatolia (Asia Minor), as you can see from the map of place names above. It was research into the now-abandoned Parnassian theory that brought this to light. But there simply isn’t enough evidence to support the idea that the Pre-Greek substrate is connected to Anatolian languages.

Even the Pre-Greek theory has been battered by fierce criticism, because its main proponent, Robert S.P. Beekes, considers nearly every possible loanword to be Pre-Greek as a matter of principle:

“I think that it is methodologically more sound to start from the assumption that non-Greek words are Pre-Greek. Only when there is reason to do so should we assume that they have a different origin.” (Beekes 2014: 45)

The linguistic reality, however, is much less tidy. Like everywhere else prior to the rise of urbanization and city-states, the Aegean comprised a rich tapestry of tiny local languages, spoken by at most thousands of speakers. The Ancient Greeks themselves mentioned a number of peoples in the region, some of whom were said to speak ‘a language which was not Greek’ (Herodotus’ description of the Pelasgian language, which launched a thousand linguistic theories about the Pre-Greeks):

  • Aonians
  • Cauconians
  • Dryopes
  • Hyantes
  • Lelegians
  • Pelasgians
  • Temmikes

We also have written evidence (albeit often exiguous) of many other languages in the region:

  • Carian (IE, Anatolian branch)
  • Eteocretan (probably non-IE, possibly related to Minoan)
  • Eteocypriot (probably non-IE)
  • Illyrian (IE)
  • Lemnian (non-IE, possibly Tyrsenian)
  • Macedonian (not to be confused with the modern Slavic language; this one was Greek or closely related to Greek)
  • Minoan (undeciphered and unclassified but probably non-IE)
  • Paeonian (IE)
  • Thracian (IE)

So it’s unsurprising that many scholars have attacked Beekes’ assumption of a unified Pre-Greek language, and also challenged his analysis of dozens and dozens of words. Nonetheless, that still leaves hundreds of words which do very likely belong to a single language, which—for lack of a term that hasn’t already been discredited—Beekes calls “Pre-Greek”.³

But what about the rest of the mystery words? Why not assume they come from Pre-Greek as well, even if they don’t have the signature Pre-Greek affixes? Once again, the meanings of the words are an important clue: almost all the words which exhibit specific Pre-Greek features (the ‑nth suffix, the s‑ prefix, etc.) are words for natural phenomena like plants and animals. The ones without these specific features are largely cultural terms referring to things like jewelry, clothing, social hierarchy, the arts, buildings and structures, etc. (There are two notable exceptions—λαβύρινθος (labýrinthos) ‘labyrinth’ and ἀσάμινθος (asáminthos) ‘bathtub’—which we’ll come back to in a minute.) Moreover, only the words for natural phenomena appear in languages elsewhere in Europe; the cultural terms are restricted to the Aegean. Here are a few potential pan-European connections to Pre-Greek words:

  • radish (a‑ prefix)
    • Ancient Greek: ῥάφανος (rháphanos)
    • Latin: rāpum
    • Old High Germanic: ruoba
    • Welsh: erfin ‘turnip’ ← Proto-Celtic *arb‑īn-
    • Old Church Slavonic: rěpa ‘turnip’ ← Proto-Balto-Slavic *rāpā́ˀ
  • heron (a‑ prefix)
    • Ancient Greek: ἐρῳδιός (erōidiós) ~ ἀρωδιός (arōdiós) ~ ῥωδιώς (rhōdiṓs)
    • Latin: ardea
    • Serbo-Croatian: róda ‘stork’
  • bull (s‑ prefix)
    • Ancient Greek: ταῦρος (taûros)
    • Latin: taurus
    • Proto-Germanic: *þeura‑ ~ *steura‑
  • small stones (k‑ prefix)
    • Ancient Greek: ἄχλαξ (ákhlax) ‘small stones’ ~ κάχληξ (kákhlēx) ‘small stones’ ~ χάλαζα (khálaza) ‘hail’
    • Proto-Germanic: *hagla‑ ‘hail’
  • chickpea (‑nth suffix)
    • Ancient Greek: ὄροβος (órobos) ~ ἐρέβινθος (erébinthos)
    • Proto-Germanic: *arwīt ‘pea’
    • Old Georgian: ერევინდი (erevindi) ~ ერბინდი (erbindi)
    • Latin: ervum ‘bitter vetch’ ← Proto-Italic *erw‑
    • Armenian: առաւոյտ (aṙowoyt) ‘alfalfa’ ← *HrVbʰoud‑

So what explains this distribution of natural phenomena versus cultural concepts? Simply put, we are probably looking at two (or more) substrate languages which both contributed vocabulary to Ancient Greek. One had the ‑nth suffix and a‑, k‑, and s‑ prefixes, and provided terms for flora, fauna, and terrain. The other (or others) provided cultural terms such as jewelry, titles, clothing, instruments, etc. The reason that the terms for natural phenomena are shared outside the Aegean is probably due to the fact that, prior to the arrival of the Indo-Europeans, the Old Europeans were a surprisingly homogeneous population of farmers (Skoglund et al. 2012; Skourtanioti et al. 2023). While genetics ≠ languages, this nonetheless suggests that many of the languages of Old Europe were likely related. So it’s entirely possible that Germanic and Italic and Celtic and other Indo-European speakers independently borrowed words for things like ‘chickpea’ and ‘radish’ from Old European languages, and these words just happened to be cognate with the words in Pre-Greek. Alternatively, we can say that these words were common loanwords (‘wanderwords’) across Old Europe.

As for the cultural vocabulary, while these terms could in theory come from anywhere nearby, there is one culture which was known to have an outsized influence on the Mycenaean Greeks: the Minoans on the island of Crete. Minoan culture was extraordinary for its time, with monumental architecture, sprawling palace complexes, and luxury arts and crafts. The Minoans are often regarded as the first civilization (in the technical sense) in Europe, and their influence spread far across the eastern Mediterranean. Archaeologists refer to the way that Minoan culture radiated outward from Crete as the Versailles Effect, by analogy with the way the French court under Louis XIV influenced all of Europe. Minoan was, in other words, an adstrate rather than a substrate, spreading its lexicon to the neighboring languages.

The Minoans also developed their own writing system called Cretan hieroglyphs, which was not derived from the Egyptian or Mesopotamian systems (Egyptian hieroglyphs, cuneiform, and later the Phoenician alphabet). Cretan hieroglyphs later evolved into (or perhaps alongside) the still-undeciphered Linear A script, which was a combination syllabary and logographic script. The Mycenaean Greeks then adopted a version of Linear A to write Greek. This script is called Linear B, and was deciphered by English architect and self-taught linguist Michael Ventris in 1952. Because 70% of the signs of Linear B are directly inherited from Linear A, we know the sounds of many words in Linear A, but we still don’t know what most of those Minoan words mean.

A table of Cretan hieroglyphs

A table of Cretan hieroglyphs (Wikipedia: Cretan hieroglyphs)

The intractable nature of Linear A means, unfortunately, that we can’t be absolutely certain that any Pre-Greek words were borrowed from Minoan, because we don’t know what the Linear A words meant. There are nonetheless some strong candidates for borrowings from Minoan:

  • In Linear A, there’s a logogram 𐛢, which is one of the minority of signs we happen to know the meaning of—‘wool’. We know this because in Linear B the same sign means ‘wool’. In Linear A, the sign appears in bookkeeping contexts next to numerals, and often near the ideogram for sheep ⟨𐘏⟩ (also the sign for /qi/)—exactly what you’d expect for a sign meaning ‘wool’. But more importantly, if you look closely, you can see that the logogram is actually a composite of two syllabic signs—𐙁 (ma) and 𐘘 (ru or lu)—indicating that the word for ‘wool’ was maru or malu. It not-so-coincidentally happens that the Greek word for ‘wool’ was μαλλός (mallós)—a word with no Indo-European etymology. Thus mallós is almost undoubtedly a borrowing from Minoan.
  • Labyrinths are strongly associated with Minoan culture, having been built in Greek mythology for King Minos of Crete to hold the Minotaur. The word λαβύρινθος (labýrinthos) ‘labyrinth’, it again not-so-coincidentally happens, is one of those rare cultural terms with a Pre-Greek ‑nth suffix. It even appears in Linear B (not Linear A) as 𐀅𐁆𐀪𐀵𐀍 (da‑pu₂‑ri‑to‑jo) ‘of the labyrinth’, meaning that the word was borrowed very early into Mycenaean Greek. Minoan culture would have exerted its influence on the Pre-Greeks just as it did the later Mycenaean Greeks. So in all likelihood, the Pre-Greeks borrowed the word for ‘labyrinth’ and tacked their ‑nth suffix on the end of it, then later gave the word to the Greeks.
  • The Minoans also boasted the oldest known personal bathtubs in the ancient world, a technology which then spread elsewhere in the Aegean. The Greek word for ‘bathtub’, ἀσάμινθος (asáminthos), shows the classic ‑nth suffix even in early Mycenaean—𐀀𐀭𐀖𐀵 (a‑sa‑mi‑to)—meaning that this word too was likely borrowed by the Pre-Greeks from the Minoans, and in turn borrowed by the Greeks.
  • There are also several toponyms which appear in Linear A, Linear B, and Classical Greek, meaning that the names of these places were borrowed from the Minoans. Predictably, these toponyms are all on the island of Crete.

We can say with some confidence, then, that the mystery loanwords of Ancient Greek come from at least two languages: the Paleo-European inhabitants of the Greek mainland, and the various other cultures in the surrounding region, most especially the extravagant Minoan palatial culture of Crete.

Languages are a fossil record of the history of their speakers. Hidden beneath the surface of every language are layers upon layers of history—a reflection of millennia of migrations, trade, and conflict. Historical linguistics allows us to recover some of that history, provided we know where and how to look for it. We will probably never fully recover the language of the Pre-Greeks, but it has left an indelible mark on the languages of the Mediterranean and even English. Every time you utter words like labyrinth, mint, narcissism, and guitar, you carry on the culture of those Paleo-Europeans who inhabited the Aegean prior to the arrival of the first Indo-Europeans over 4,000 years ago.

A reconstruction of the Greek town of Dimini in Thessaly, as it existed ca. 3700 BCE.

A reconstruction of the Greek town of Dimini in Thessaly, as it existed ca. 3700 BCE. Source unknown.

English words from Pre-Greek

This is a non-exhaustive list of words that are generally thought to be loanwords from a Pre-Greek substrate language into Ancient Greek, and which later found their way into English. I’ve done my best to represent the consensus view and avoid including dubious cases, but there’s certainly room to disagree with my individual choices. This article is, after all, a high-level summary of a vast field of research with lots of internal debate.

📑 Bibliography

The question of the nature of Pre-Greek and the status of loanwords into Ancient Greek are some of the most well-researched and debated areas in Indo-European linguistics. Below are just a few of the more recent sources that provide decent high-level overviews of the topic.

  • Beekes, Robert S.P. 2014. Pre-Greek: Phonology, morphology, lexicon (Brill Introductions to Indo-European Languages 2). Brill.
  • Duhoux, Y. 2007. Pre-Greek languages: Indirect evidence. In A.-F. Christidis (ed.), A history of ancient Greek: From the beginnings to late antiquity, pp. 223–228. Cambridge University Press.
  • Meester, Lotte. 2024. Substrate stratification: An argument against the unity of Pre-Greek. In Guus Kroonen (ed.), Sub-Indo-European Europe: Problems, methods, results, pp. 283–300. De Gruyter.
  • Verhasselt, Gertjan. 2009. The Pre-Greek linguistic substratum: An overview of current research. Les Études Classiques 77: 211–239.

Notes

¹ I use the term confamiliar to refer to words that are related through borrowing rather than direct inheritance. Two words that are related to each other because they both derive from the same word in a parent language are called cognates. The term confamiliar is meant to be the equivalent of cognate for borrowed words. The two terms nicely parallel the distinction in Latin between cognātus ‘related by birth’ vs. familiāris ‘member of a household (not necessarily related by blood)’. The term confamiliar is also used in the life sciences to refer to taxonomic members of the same family; so here I am merely extending the metaphor to refer to members of the same family of words. This is my own unique coinage, though. As far as I know, no other linguists use this term in this way (yet 🤞🏼).

² */tw/ sometimes underwent this sound change too, which is the source of of the /ss/ ~ /tt/ alternation in τέσσαρες (téssares) ‘four’. This word comes from PIE *kʷetwóres rather than a palatalized sound.

³ I don’t actually love the term “Pre-Greek” because in historical linguistics the prefix Pre‑ is generally reserved for an earlier stage of a single language, usually one reconstructed from language-internal evidence alone. This contrasts with the last common ancestor of a group of related languages, which is reconstructed by comparing those languages, in which case we use the prefix Proto‑. According to this convention, Pre-Greek or Pre-Proto-Greek should really refer to a hypothetical dialect of Indo-European from which Proto-Greek later emerged. Regardless, Pre‑ is still commonly used for substrates as well, such as Pre-Greek or Pre-Indo-European (i.e. the languages that existed in Eurasia prior to the arrival of the Indo-Europeans—not to be confused with Paleo-European / Old European languages, which refers more specifically to the Neolithic languages of Europe). (Another reason I dislike the term “Pre-Greek” is because this article would have sounded much cooler titled “Pelasgian: The mystery language hidden in Ancient Greek”, or “Parnassian” or the like.)

The Daily Front Page 8 of 28
Friday, September 18, 2026 The Daily Front No. #260918 — The Virtue of Omission
article

The most important product decision is what you don't build

by ChrisArchitect·▲ 146 points·51 comments·liamnugent.me ↗
Ask what they killed and the room goes quiet.

Ask a team what they shipped this year and you’ll get a list. Ask what they killed and the room goes quiet.

There are two things I’ve tried to take a hard pass on when working on consumer-facing financial services apps.

A “document hub”.

And a “notifications centre”.

You can picture both without any help from me. To my mind they’re both examples of letting the organisation off the hook at the expense of the user. Although it’s fair to say that plenty of users have now been trained to expect these things, and to know roughly how they work — even though they’re crap.

The problem underneath is real enough. The business has legacy systems that produce PDFs — letters, statements, that sort of thing — because originally they’d have been sent out in the post. Compliance experts will also talk to you about “persistent communications channels”. So you need somewhere in the app to put all of that.

The lazy solution is a document hub. A list of files you can select and open.

Seems simple.

Except every stakeholder involved wants to do a Columbo and add just one more thing. By the end you’ve essentially rebuilt Google Drive, with tagging, archiving, printing and sharing across several channels, each with its own quirks. Not to mention clearing the security and authentication hurdles so the right person sees only the right documents. And so on, and so on.

And that is just to build the thing.

Never mind the overhead you’ve created to maintain it, month after month, year after year, as each iOS release cycle comes round and some foible you were relying on to prop up a piece of functionality gets removed unilaterally by Apple. Or Google, or Microsoft, or Amazon.

The notifications centre is the same story with a different opening line. A stakeholder asks whether we could just have a bell icon with a little red dot in the top right corner of the screen. Give it a few months and you’re building a bad Gmail clone.

Showing people the bill

I’ve come up against both of these several times now. It is not easy to get people to let go of a mental model they’re holding in their head, particularly when a customer focus group goes some way to validating it. The cliched old Henry Ford line comes to mind — if you’d asked people what they wanted, they’d have said faster horses, not a motor car.

What I’ve found actually works is illustrating the running costs. Not the build cost. The running costs, laid out over years. That tends to bring people back from the brink.

But killing the platform doesn’t make the original problem go away. How do we communicate with our customers in a way that suits them and meets the legal obligations on our side?

The answer I’ve come back to every time is to take each use case on its own merits. Be rigorous about what actually needs to be said, and when. Then be ruthless about whether it needs a whole platform built first. It’s the same test I keep applying when picking the right problems to solve — does it make the boat go faster?

Most of the time — not always, to be fair — the answer is to use something simple, or something that already exists, even if it isn’t perfect.

Resist, resist, resist the temptation to build a one-size-fits-all.

I’ve even seen the notifications centre idea get built, and then the intent of the original message couldn’t be met by the thing that had been built to carry it.

So my advice comes in two parts. Be extremely choosy about what you let be added to your system. And be militant about taking things out.

Nobody gets promoted for deleting things

The second part is much harder than the first, and it isn’t simply cowardice. Gerry McGovern has been saying this for a long time:

Those who create and launch are the people who are rewarded and looked up to because we still have a culture that rewards the production of things over everything else. To review, to maintain, to remove—this is all seen as lesser work.

His Top Tasks work found that stripping roughly 80 to 90 per cent of a site’s content made companies sell more, cut their support calls and helped people find what they were after faster. Removal as the improvement itself, not the tidying up afterwards. I’ve made the small-scale version of this argument before, about deleting the FAQ page.

And it runs deeper than the org chart. Klotz and colleagues, writing in Nature, ran eight experiments and found that people systematically overlook subtractive changes, even when removing was obviously the better move. Additive ideas arrive quickly and cheaply. Subtractive ones cost real cognitive effort. So we are hard wired for this.

Which means spending the brain power to work out what success genuinely is. And success is usually something out there in the world of the customer, rather than something in here, in the world of the organisation.

But building is nearly free now, isn’t it?

DHH is thoroughly delighted about endless execution — in the age of agents, every idea, every hunch and every experiment is within immediate reach.

I think it’s easy to misunderstand that as an argument for making more stuff all the time.

Using agents to do the pruning seems like the smarter move to me. To do the maintenance. To remove things judiciously, and more carefully than anyone is going to manage by hand.

Stewart Brand’s Maintenance of Everything sits on the paradox that maintenance is absolutely necessary and also entirely optional.

Optional until it’s essential, more like.

The elephant in the rollneck

The elephant in the room, wearing a custom-made Issey Miyake black rollneck, is of course Steve Jobs’ 1997 matrix. Two consumer products, two pro products, and one stellar product in each box.

Four boxes. One product in each. Everything else cancelled.

He then shut down all the other product lines people were working on. That pooled the engineering talent instead of spreading it thin. It was a brave move, and it did the trick.

Apple are up to something like 52 different products and services at the last count, so the end game was never to only sell four things. The point was to focus the organisation on the most important ones, so they didn’t misspend the opportunity cost that allowed them to become one of the biggest corporations in the world.

One thing at a time. Une chose à la fois.

(I’d have written a shorter post, but I couldn’t work out what to take out.)

The Daily Front Page 9 of 28
Friday, September 18, 2026 The Daily Front No. #260918 — A Phone With Less to Say
article

Minimal Phone 2

by nashashmi·▲ 213 points·196 comments·minimalcompany.com ↗
Use your phone. Don’t let it use you.

the phone that respects your time

use your phone. don't let it use you. the compact, focus-first phone with a real keyboard.

focus on what matters.

use your phone. don't let it use you. the compact, focus-first phone with a real keyboard.

the average person will spend

0 years

of their life looking down at a screen.

endless feeds, alerts and algorithms are built to keep you there — long after there is anything left to see.

186 pickups a day

4h30 of screen time

152 notifications a day

it is time to look up.

introducing minimal phone 2

everything you need. nothing you don't.

Built with intention, down to the smallest detail. Take a closer look at what makes Minimal Phone 2 different by design.

built to feel as good as it looks.

frame

aluminium unibody

milled from a single block of 6000-series aluminum for a rigid, substantial feel that's quietly premium in the hand.

glass

sculpted 2.5d glass

gently curved edges flow cleanly into the frame for a smooth, seamless feel.

form

shaped for one hand

122.6 mm tall and 9.48 mm thin. a compact footprint made for comfortable one-handed use.

9.48 mm thin

122.6 mm tall

2.5 d curved glass

6000 series aluminium

typing should feel intentional.

A feeling a touchscreen can't replicate—every key shaped, spaced, and tuned.

  • individually backlit keys
  • metal dome switches
  • satisfying key travel
  • ~69 mm keyboard width
  • polycarbonate + silicone keys
  • built for two-thumb typing
  • touch-sensitive keys

type on it. scroll on it.

  • Scroll & Navigate
  • Cursor & Text Control
  • Swipe to Delete
  • Swipe Typing & Word Gestures
  • Custom Gesture Shortcuts

more keyboard layouts.

Five layouts at launch, laser etched into the keys, so the physical keyboard matches the way you actually type.

minimal os

Important information at a glance. Useful actions the moment you start typing. Everything else stays out of the way.

every message, one place.

Swipe right from the home screen and everything is waiting for you: email from every account, texts, calls, voicemail, and calendar. Read it, respond, move on.

one swipe

from your Home Screen

one list

Main, messages, calls, events, etc.

one place for what matters

Messages, reminders, events, and anything that needs your attention, all together in the Hub.

reply from the hub

Read and respond without opening the app. Less switching, less distraction.

keep work and personal separate

Organize the Hub around different parts of your life, so work stays at work and personal stays personal.

ready for the networks you already use.

5G, Wi-Fi, Bluetooth, NFC, GPS, and more. Built to connect with the carriers, accessories, and services that are already part of your life.

  • 5g cellular
  • physical sim
  • esim
  • wi-fi 6e
  • bluetooth 5.4
  • nfc
  • gps

built for today. ready for what comes next.

thoughtful features. no unnecessary excess.

everything has a purpose.

The essentials—plus details rarely found on modern phones.

Programmable switch

Assign it to radios, recording, or modes.

LED notification indicator

Know without waking the screen.

Improved camera system

50 MP rear. 20 MP front. Always ready.

3.5 mm headphone jack

Wired headphones. No adapters.

Fingerprint security

Quick, secure unlock.

Dual speakers and mics

Clear calls. Focused audio.

smaller phone. bigger life.

Comfortable in your hand. Easy in your pocket.

122.6 × 72 × 9.48 mm

Everything you need from a modern smartphone, in a footprint that stays out of the way.

download any app.

Full Android. Every app you already use, there when you need it.

no restrictions. no compromises. more control.

a great camera, made for real life.

  • 50 mp rear camera
  • autofocus
  • f/1.8 aperture
  • 79.2° field of view
  • led flash
  • 20 mp front camera
  • f/2.0 aperture
  • 77.8° field of view

take the photo. stay in the moment.

50 mp · f/1.8 · autofocus

notifications without distractions.

A glance is enough. Assign colors and light patterns to the people and apps that matter, so you can tell what needs your attention without waking the screen.

custom colors

Assign a color to different apps or contacts.

4 light patterns

Solid, pulse, blink, or fade.

quiet hours

Stays dark when you want the phone to stay quiet.

your phone. your rules.

No carrier lock. No software lock. Minimal OS gives you control over how your phone works, how much it asks from you, and how far you want to take it.

set your own limits.

Choose how focused you want your phone to be. Reduce notifications, block distracting apps, set screen-time limits, switch to grayscale, or shut things down at night. The stricter settings are designed to stay that way until you deliberately reset the device.

use any launcher.

Prefer a traditional Android home screen? Install any launcher you want and make it your default. Minimal OS doesn't lock you into our way of doing things.

or make it entirely yours.

The bootloader is unlockable. Flash a custom ROM, build your own, or experiment however you want. It's your hardware. You should be able to decide what runs on it.

minimal os, your way.

Minimal Phone 2 ships with Google services, so everything most people expect works out of the box. Want to go further? We're releasing an official de-Googled version of Minimal OS, with the full ROM for you to flash yourself.

ships by default

with google

The standard, Google-certified build. Compatibility right out of the box.

  • Play Store and Google Play Services
  • Banking apps, payments and everything that depends on them
  • The same Minimal OS: keyboard, Minimal Hub, Commands

for ultimate privacy

de-googled

An official build of Minimal OS without Google services. You flash it; the phone is yours.

  • Full ROM provided by us, not a third party
  • Keyboard, Minimal Hub, Commands and every native app intact
  • Flash back to the standard build any time

one phone. two paths.

We won't ship a separate de-Googled device. Every Minimal Phone 2 arrives with the standard software, and anyone who wants the privacy-focused version makes that choice themselves. Use Google if you want it. Remove it if you don't.

choose yours.

choose your finish

pearl
clean. timeless.

onyx
dark. discreet.

choose your storage

12 gb ram + 256 gb
$599 $699
room for everyday life.

12 gb ram + 512 gb
$699 $799
for the long haul.

life happens away from the screen.

less screentime, more sunlight

look up

bring the analog back

sunday, offline

touch grass

use your phone. don’t let it use you.

the phone that respects your time.

good questions. clear answers.

Get to know MP2, choose your setup, and find help with your order. Questions about the original phone have their own section below.

Ordering & managing your preorder

When will Minimal Phone 2 ship?

MP2 preorders are scheduled to ship in December 2026. This is a preorder shipping estimate, not a delivery date. Your arrival date will depend on when your order dispatches and its destination.

When will I be charged? Is this a subscription?

You pay the full order total at checkout. It is a one-time purchase, with no recurring charges for the phone. Review the selected products, variants and total before completing payment.

Can I cancel my MP2 preorder?

Yes. You can cancel before your MP2 preorder ships for a full refund. Open the manage-order link in your confirmation email, or use the Manage order page. If cancellation is unavailable or you need help, contact us with your order number. Once an order has shipped, it follows the return process instead.

Can I change my finish, storage, keyboard or shipping address?

Use Manage order to see the changes available for your unshipped order. Sign in with the email address you used at checkout, open the order, and review your changes before confirming. Changes that increase the total may require an additional payment. If an option is unavailable, contact us before the order ships.

Can I add accessories after placing my order?

Open your order through Manage order and check the available add-ons. Choose accessories listed for MP2 and review the updated total and shipping information before confirming. If the item you want is not shown, ask our team for help rather than placing a duplicate phone order.

Do I need a Shop account to order or manage my purchase?

You do not need to create a Shop account to purchase. To manage a website order afterward, use the email address from checkout and the sign-in code sent to that address. This passwordless store sign-in is separate from choosing Shop Pay at checkout.

I cannot find my confirmation email or sign-in code. What should I do?

Check your spam, junk and promotions folders, then make sure you are checking the email address used at checkout. If you still cannot find it, use the contact form and include your name, checkout email and order number if available. Do not place another order just to recover the first one.

Shipping & delivery

Do you ship to my country?

Enter your delivery address at checkout to see whether shipping is available for your destination and which options are offered. If checkout does not offer a shipping method, contact us with your country and postal code before placing an order.

How much does shipping cost?

US shipping is a flat $10 per order for MP2 and its accessories, and for MP01 accessories except dbrand skins. Shipping for dbrand skins is $5.99 within the US. International shipping for phones and accessories is $29 per shipment to supported destinations. Merchandise is printed and shipped by Printful, with separate shipping rates shown at checkout. Mixed orders can have separate shipments and shipping charges. Review the shipping total and any applicable taxes or duties before payment.

Will everything in my order arrive together?

Check the shipping information for each item. MP2 and its preorder accessories have preorder timelines; merchandise and MP01 accessories may have different fulfillment arrangements. An order can arrive in separate shipments. If you need a specific item sooner, ask us before ordering.

How do I track my order?

Check your shipping confirmation for the tracking link once your order has dispatched. A preorder confirmation means your order was placed; it does not mean the parcel has shipped. If you cannot locate your tracking information, contact us with your order number.

Can I change the delivery address after dispatch?

An address change in your account will not automatically redirect a parcel already in transit. Contact us as soon as possible with your order number and the corrected address. We will check the available options; a change after dispatch cannot be guaranteed.

My parcel is delayed, missing or marked delivered. What should I do?

Check the latest carrier tracking, the delivery address on your order, and any safe-place or delivery notice. If it is still missing, contact us with your order number, tracking number and what the carrier reports. If an item arrived damaged or is missing from the box, keep the packaging and include clear photos with your message.

Getting to know MP2

What is Minimal Phone 2?

MP2 is a compact smartphone with a physical keyboard and Minimal OS, an Android-based interface designed to make everyday tasks easier to reach. It combines a touchscreen with real keys, so you can choose how you navigate and type. Explore the phone for demonstrations and full specifications.

Is the screen E Ink?

No. MP2 has a 4-inch AMOLED touchscreen with a 1080 × 1240 resolution and a 90 Hz refresh rate. MP2 is a different product from the original Minimal Phone; specifications from MP01 do not describe MP2.

Which finishes, memory and storage options are available?

The website offers Pearl and Onyx finishes in two configurations: 12 GB RAM + 256 GB storage, or 12 GB RAM + 512 GB storage. Every MP2 ordered now ships with 12 GB of RAM. Select your finish, storage and keyboard on the product page to see the corresponding price and availability.

Can I expand the storage with a microSD card?

No. MP2 does not have a microSD slot. Choose 12 GB + 256 GB or 12 GB + 512 GB based on the space you expect to need for apps, photos, music and other files.

Which physical keyboard layouts can I choose?

Choose English QWERTY, French AZERTY, German QWERTZ, Korean Dubeolsik or Arabic when ordering. This selects the physical keyboard layout. Review the layout on your order confirmation, and use Manage order before shipping if you need to change it.

Is the keyboard backlit? Does it support touch?

Yes. The current MP2 specification includes backlit physical keys, capacitive touch keys, remappable keys and long-press shortcuts. For a particular gesture, shortcut or app workflow, tell us what you want to do so we can confirm the behavior rather than assume every app handles it the same way.

What battery and charging options does MP2 have?

MP2 has a 3,600 mAh silicon-anode battery, with wired charging up to 27 W and Qi wireless charging up to 15 W. Actual charging speed and battery life depend on your charger, usage, network conditions and settings. We do not promise a fixed number of days between charges.

Does it have a headphone jack or support an external display?

Yes. MP2 includes a 3.5 mm headphone jack and USB-C audio. Its USB-C specification also includes DisplayPort 1.4 output for a compatible external display setup. Check that your cable, adapter and display support the connection you want to use.

What cameras does MP2 have?

The current specification lists a 50 MP rear camera and a 20 MP front camera, with electronic image stabilization (EIS). For the full camera specifications and the latest demonstrations, visit the phone page.

Apps, networks & connectivity

Can I use my everyday Android apps?

Minimal OS is built on Android, but compatibility can depend on the individual app, its security requirements and how it handles MP2’s screen and keyboard. If a banking, work, messaging or accessibility app is essential to you, tell us its exact name and your intended use before ordering so we can check.

Will MP2 work with my mobile carrier?

MP2 supports 4G LTE and 5G, but matching radio bands alone does not guarantee a carrier will activate a device or support every feature. Contact us with your country, carrier and the services you need, such as calls, data, VoLTE or Wi-Fi Calling, so compatibility can be checked before you order.

Does MP2 support a physical SIM and eSIM?

The current specification lists one physical SIM plus eSIM. Activation and available services still depend on your carrier and plan. If you need a particular two-number setup or are switching from another phone, check that use case with us first.

Does it have Wi-Fi, Bluetooth and NFC?

Yes. The current MP2 specification includes Wi-Fi 6E, Bluetooth 5.4 with LE Audio, and NFC. A hardware feature does not guarantee support for every payment service, transit system or accessory, so ask us about any service you rely on.

Will my specific app, integration or alternative operating system be supported?

Tell us the exact app or operating system and what you want it to do. App support, built-in integrations and alternative operating-system support are separate questions. We will distinguish a confirmed feature from a development plan; please do not order on the assumption that a requested integration or ROM is already available.

Accessories & merchandise

Which accessories should I buy for MP2?

Choose products in the MP2 section of the shop. The range includes screen protectors, the clear MagSafe case, the Thinborne case and wireless charging docks. Open each product for its current images, variants, price and shipping information.

Will an MP01 case or skin fit MP2?

Do not choose MP01 cases or skins for an MP2 order. They are made for the original phone. Use the MP2 accessory range instead, and ask us if a listing does not make compatibility clear.

Which ring color can I choose for the clear MagSafe case?

The clear MagSafe case has Pearl Ring and Onyx Ring variants. Select your preferred ring on the product listing or details page, then check the selected variant in your cart before checkout.

Are accessories or Kickstarter extras automatically included with a website order?

Check the product listing and your cart for what your purchase includes. Accessories shown as separate add-ons are separate items unless the current offer explicitly includes them. Do not assume that a Kickstarter reward or campaign bonus is included in a website order.

How do I choose merchandise sizes and variants?

Select the size, color and other available options on the shop or product page before adding the item. Check the product description and any size guide, then review the variant and price in your cart. Merchandise availability and shipping information are specific to the item.

Returns, refunds & product issues

What is the difference between cancelling and returning an order?

Cancellation happens before dispatch; an MP2 preorder can be cancelled before it ships for a full refund. A return happens after shipment and has separate eligibility, timing and condition requirements. Read the warranty and return policy and contact us before sending a product back.

How long does a refund take to appear?

A cancellation confirmation and money appearing in your account are separate steps. Refunds are returned through the original payment method, and posting time depends on the payment provider. Keep your refund confirmation and contact us with your order number if you need help checking its status.

How do I report a fault or request warranty support?

Use the contact form and select your product model and Technical support or Return or warranty. Include the order number, a clear description, when the issue started and any troubleshooting already tried. We will help with the next steps and any photos needed. Coverage and exclusions are described in the warranty policy; please contact us before arranging a repair or sending anything back.

MP01 & getting help

Can I still get help with the original Minimal Phone (MP01)?

Yes. Select MP01 in the contact form so your message goes into the correct support context. Include your original order details and the issue you need help with. MP2 shipping dates, specifications and preorder terms do not replace the terms or status of your MP01 order.

Where can I find MP01 accessories?

Use the clearly marked MP01 accessories section in the shop. These accessories support the original phone. Their stock and shipping information are shown on their own listings; do not use the MP2 preorder schedule for an MP01 accessory.

What if I have both MP01 and MP2 items or orders?

Tell us which model and item each question relates to, and include the relevant order numbers. We will handle each product separately, even if the items share one order or email address. If you are unsure which model you have, select Not sure / other and describe it.

What should I include when contacting support?

Include your product model, the email used at checkout, your order number if you have one, and a short description of the question or problem. For compatibility questions, add your country, carrier or exact app name. Never send passwords, sign-in codes or full payment-card details.

The Daily Front Page 10 of 28
Friday, September 18, 2026 The Daily Front No. #260918 — Tiny Models, Large Claims
show hn

Show HN: Cactus Needle 3: 8-29MB automation models can match DeepSeek V4 Flash

by HenryNdubuaku·▲ 175 points·79 comments·cactuscompute.com ↗
One set of weights, every depth from 2 to 20 layers a model of its own.

One set of weights, every depth from 2 to 20 layers a model of its own: an intelligence ladder.

Today we release Needle 3: a foundation model for mobile, wearables, robots, smart home, automotive and microcontrollers. The whole model is a single 8-29 MB binary built on our Simple Attention Network, and we trade general chat capacity to beat models 10x its size on mobile tool calls and match 2-3x bigger models on extraction.

Tool calls

Given the functions your app exposes, Needle picks the right ones and fills every argument from what the user said. Ask for two things and you get two calls in order; ask for something no tool covers and you get an empty list, not a guess.

Structured extraction

Declare a shape, hand over messy text, get typed fields back: an invoice, a booking, a notification, a form. The decode grammar guarantees the output parses. Extraction generalised well to classification problems too.

Text embedding

The same model returns a vector for a sentence, so an app can search, match and route locally: find the note you mean, pick the tool closest to a request, collapse duplicate alerts.

What that looks like in a product:

Smart home

Go from pressing buttons to talking to the house. "Dim the bedroom and lock up" becomes two calls, executed offline, with no hub round trip.

Robots

Give a vacuum or a small robot nuanced instructions: "clean the kitchen but leave the bedroom", "go back to the dock when you are done". Each becomes a sequence of moves it can execute.

Phones

An assistant that acts on the device instead of answering: make an album from last weekend's photos, open a site, dim the screen, find the lease in your files.

Wearables

Read a notification into structured data on the wrist: a card charge into merchant, amount and date; a message into a reply; a complaint into a sentiment flag.

AR glasses

Navigation and nearby search from a short request, with no phone or network in the loop.

Automotive

Climate, media, navigation and calls from requests in the cabin, with the tool set pinned so it survives a long drive's worth of conversation.

Computers

Plain-English control of the machine in front of you: draft the mail, start the timer, copy the address, open the tab.

Search and matching

Embeddings that never leave the device: semantic search over notes, messages and documents; a query matched to the closest of hundreds of tools; near-duplicate alerts merged on a watch.

Model

Intelligence laddering. Every layer of Needle 3 is a sub-network with monotonically increasing capacity. Developers can choose the right size from the 2-layer (2L) subnetwork to 20 layers (20L). Each subnetwork is amenable to fine-tuning, such that 4L can match DeepSeek V4 Flash when tuned on downstream tasks for one epoch. Intelligence laddering produces 9 to 29 MB CQ2-bit binaries and supports a wide range of tiny devices.

Inputs

Text prompts, plus tool definitions or an extraction schema

Outputs

Structured JSON with tool calls or extractions

Model

29-121M Laddered Simple Attention Networks, CQ2 quantisation

Training

360B tokens of proprietary structured dataset

Speed

400-4k tokens/s decode and 1-10k tokens/s prefill on a Raspberry Pi 5

Figure 1. The architecture spends over 2x fewer MFLOPs per token than a transformer of the same configuration.

Figure 2. Needle 3 beats models 10x its size on mobile tool calls and matches models 2-3x its size on extraction. Needle 3 subnetworks (20, 16, 8 and 4 layers) run through the shipped CQ2-bit binary, baselines at f16 under vLLM, DeepSeek V4 Flash through its cloud API. The line joins the Needle models.

Get started

Install the Python package. The inference engine is fetched once from Hugging Face and cached; there is nothing else to build.

Needle reads your tool descriptions to decide what to call and how to fill arguments, so describing them well is the whole game.

Simple: decorate a function. The signature gives the argument types, the docstring is the tool description, and run() completes the loop: the model picks the call, Needle executes your function, feeds the result back, and returns the final response with the executed tool results attached as results.

import needle

@needle.tool
def get_weather(city: str):
    "Get the current weather for a city."
    return {"city": city, "temp_c": 27, "sky": "clear"}

agent = needle.Needle(tools=[get_weather])
print(agent.run("what's it like in Lagos right now?")["results"])
# [{'city': 'Lagos', 'temp_c': 27, 'sky': 'clear'}]

Route by pattern: when a description cannot enumerate every phrasing, give a tool triggers, regular expressions matched against each request. A match restricts the decode to the matched tools and requires a call, so the request reaches the tool you named instead of being refused or misrouted, and the call ships even below the confidence floor. A match restricts the whole turn, so a catch-all should exclude the nouns other tools own, e.g. ^(?![\s\S]*\b(lights?|doors?)\b)[\s\S]*\b(turn|switch)\b[\s\S]*\b(on|off)\b; then "switch the fan on and dim the kitchen lights" still reaches both tools.

from typing import Literal

@needle.tool(triggers=[r"\b(turn|switch|power|flip)\b.*\b(on|off)\b", r"\btoggle\b"])
def control_device(device: str, action: Literal["on", "off", "toggle"]):
    "Switch or toggle any named smart-home device."
    return {"device": device, "action": action}

agent = needle.Needle(tools=[control_device, get_weather])
agent.complete("toggle the garage door")
# function_calls [{"name": "control_device", "arguments": {"device": "garage door", "action": "toggle"}}]

Extraction: to pull structured data out of text, declare the shape and call extract(). Pass a Pydantic model and you get a typed object back.

from pydantic import BaseModel

class Invoice(BaseModel):
    vendor: str
    total: float
    due_date: str

invoice = needle.extract("Invoice from Acme Corp, $1,200.00, due 2026-09-01", Invoice)
print(invoice.vendor, invoice.total)   # -> Acme Corp 1200.0

Every turn returns one JSON object:

{
  "type": "call",
  "success": true,
  "error": null,
  "error_code": null,
  "function_calls": [ { "name": "set_lights", "arguments": { "room": "living room", "on": true, "brightness": 30 } } ],
  "reasoning": "'living room' -> room; 'dim' -> on true, brightness 30",
  "confidence": 0.94,
  "prefill_tps": 4300.0,
  "decode_tps": 850.0,
  "peak_ram_mb": 28.5
}

Confidence gating and routing: every response carries a confidence score from a calibrated head, and the engine already applies a floor of 0.1. Below it, the call is withheld into suppressed_calls and function_calls is empty. Above it, the score is yours to route on: act at once when it is high, show the call and ask when it is middling, and treat an empty result as a refusal. A tool with triggers always produces a call for a matching request, so the score is what tells you whether to run it or confirm it.

r = agent.complete(user_text)
calls = r["function_calls"]
held = r["suppressed_calls"]

if calls and r["confidence"] >= 0.7:
    execute(calls)                                   # sure: act
elif calls or held:
    confirm(calls or held, r["reasoning"])           # unsure: show the call, ask
else:
    say("I can't do that here")                      # nothing to do: refuse

Writing tools: the model reads a schema literally, so a narrow tool with a plain description beats a broad one. One tool per action, described by the actions it covers ("Turn a room's lights on or off") rather than a category. Name enum options after what a user says (action: ["increase", "decrease"]) and keep synonyms in the description. Give a required argument a default when a request may leave it out; a required argument with no default and no evidence in the request is withheld rather than guessed. Put value formats in descriptions ("City, ST", "e.g. T-1042"). Add triggers to intents that must always reach a tool, and keep the toolset per turn small, since every extra tool is a chance to misroute.

Fine-tune: the Python package is the quick path. LoRA on the frozen base at the full 20 layers, then a 4-bit .cact of any subnetwork that runs on the same engine.

needle finetune data.jsonl --epochs 10 --out adapter.safetensors
needle build --lora adapter.safetensors --out tuned.cact
needle build --lora adapter.safetensors --platform linux-arm64 --layers 2 --out ./device

The guides go deeper: designing tools, confidence, extraction, fine-tuning, the Python reference, supported devices, the .cact format and porting Needle. Source is on GitHub.

Cactus Platform

Needle was designed to be customised. Its capacity is a ladder, and a subnetwork as small as 2 layers, fine-tuned on one product's tools, runs optimally on devices far smaller than the full model needs. Constraining the capacity to a narrow, well-defined task is what lets it reach frontier-level accuracy there: fine-tuning on DroidCall lifts every subnetwork by 18 to 36 points, and from 4 layers up the tuned subnetwork passes DeepSeek V4 Flash, starting at 29M parameters (Figure 3).

Every subnetwork, fine-tuned on the platform

Figure 3. Each Needle 3 subnetwork before and after fine-tuning on the platform, base and tuned both scored with forced calls, against DeepSeek V4 Flash through its cloud API.

The Cactus Platform is the full path: Cactus datasets, the 2-bit quantisation behind the shipped model, evaluation design and tracking, full-depth fine-tunes and dataset management, all on our infrastructure and training pipeline, no need to build your own.

Deploy

Every deployment target ships a prebuilt engine under 1 MB that loads the needle3.cact weights at start. needle build fetches the engine for a platform and puts the weights beside it, at the full 20 layers or any smaller subnetwork:

Target Platform folder Ships
macOS macos-arm64 needle CLI, libneedle.a, needle.h
Linux linux-x86_64, linux-arm64, linux-armv7, linux-riscv64, linux-mipsel needle CLI, libneedle.a, needle.h
Windows windows-x86_64, windows-arm64 needle.exe, libneedle.a, needle.h
Android android-arm64, android-armv7, android-riscv64 needle CLI, libneedle.a, needle.h
iOS ios-arm64, ios-sim-arm64 libneedle.a, needle.h
tvOS, watchOS tvos-arm64, watchos-arm64 libneedle.a, needle.h
Browser wasm needle.js, needle.wasm, needle.h
WASI component wasm-component needle.component.wasm, needle.wit

One folder per target; needle build --platform downloads it and places needle3.cact beside the engine.

# engine, header and weights for this Mac
needle build --platform macos-arm64
# an 8-layer subnetwork for a Pi
needle build --platform linux-arm64 --layers 8 --out ./pi
# a tuned archive
needle build --lora adapter.safetensors --out tuned.cact
The Daily Front Page 11 of 28
Friday, September 18, 2026 The Daily Front No. #260918 — Loops Without End
article

C++26: Trivial infinite loops are no longer undefined behaviour

by ibobev·▲ 154 points·212 comments·sandordargo.com ↗
A while (true); loop with no side effects used to be undefined behaviour.

Let’s start with a question! Is this program well-defined?

int main() { while (true) ; }

If you said yes, you’d be wrong — at least before C++26. A while (true); loop with no side effects used to be undefined behaviour. Compilers were free to assume it terminates, and some — Clang in particular — would optimize it away entirely, with spectacular consequences:

// https://godbolt.org/z/WYMxxeW1T
#include <iostream>

int main() {
    while (true) ;
}

void unreachable() {
    std::cout << "Hello world!" << std::endl;
}

In Clang, this prints “Hello world!”. The compiler removes the infinite loop, main falls through, and the linker-placed unreachable() function executes. This is not a compiler bug — it’s just UB, still better than nasal demons.

Recently, I wrote about how C++26 reduces undefined behaviour, covering changes like erroneous behaviour for uninitialized reads and making incomplete-type deletes ill-formed. I completely forgot about this one. I only realized while preparing for an upcoming CppCon talk on C++26 features — so here it is now.

C++26 fixes this with P2809R3. Trivial infinite loops are now well-defined. The mentioned proposal was also accepted as a defect report, so implementations may apply the fix to earlier C++ modes as well. That is why you might not be able to reproduce the old behaviour on a recent compiler even in C++20 mode.

How did we get here?

The story starts with the forward progress guarantee, introduced in C++11 alongside threading support. The standard says ([intro.progress]) that the implementation may assume any thread will eventually do one of the following: terminate, call a library I/O function, access a volatile glvalue, or perform a synchronization or atomic operation.

A while (true); loop does none of those things. Under the pre-C++26 forward-progress rules, an execution that remains in such a loop forever has undefined behaviour. The optimizer can therefore assume that execution never gets stuck there, which enables transformations that remove the loop and mark the path as unreachable.

The funny bit is that C got this right. C++11 and C11 both introduced forward-progress rules, but C included one more rule: loops whose controlling expression is a constant expression may not be assumed to terminate. So while (1); is well-defined in C11 and ever since.

C++ never adopted that extra rule. The result was the unnecessary divergence just described, but let’s repeat it: while (1); was well-defined in C but undefined behaviour in C++.

But why would anyone write while (true); in the first place?

What I found is that this is common in embedded and kernel code as a halt-on-error pattern. When a fatal error occurs and there’s no operating system to exit to, you simply stop:

if (hardware_init_failed()) {
    log_error("fatal: hardware init failed");
    while (true) ; // halt — there's nothing left to do
}

This is not simply a common pattern on bare metal — it was also undefined behaviour in C++. The consequences aren’t theoretical. When the optimizer removes the loop, execution falls through into whatever code the linker placed after it — as the “Hello world!” example at the top of this article demonstrates. In an embedded system, that means a fatal error handler doesn’t actually halt the device. The hardware keeps running in a corrupt state, executing whatever instructions happen to follow. In security-critical code, that’s a real vulnerability.

What C++26 changes

C++26 doesn’t simply copy C’s rule, though. That approach was considered and rejected. C protects a much broader set of loops — broadly, loops whose controlling expression is a constant expression — which could inhibit useful optimizations. Instead, P2809R3 defines a deliberately narrow category: the trivial infinite loop. It’s defined by two conditions:

  1. The loop must be a trivially empty iteration statement — meaning its body is literally empty (; or {}). Any non-empty statement in the body, even a meaningless expression statement such as "a string";, disqualifies it.
  2. The controlling expression must be a constant expression that evaluates to true. For a for loop with no condition, true is implicit.

When both conditions are met, the loop body is replaced with a call to std::this_thread::yield(). This gives execution of the loop the forward-progress semantics it previously lacked.

Here’s what qualifies and what doesn’t:

Code Trivial infinite loop?
while (true); Yes
for (;;); Yes
do {} while (true); Yes
constexpr bool go = true; while (go); Yes — go is a constant expression
while (true) { "a string"; } No — body contains a statement
while (true) if (done) break; No — body is not empty
while (true) if constexpr (false) break; No — doesn’t match the syntax of a trivially empty iteration statement
bool done = false; while (!done); No — not a constant expression

The change also updates the forward progress guarantee itself: a thread may now “continue execution of a trivial infinite loop” as one of the things it’s assumed to eventually do. The optimizer can therefore no longer treat a trivial infinite loop as undefined behaviour and assume that execution continues past it.

The freestanding caveat

On freestanding implementations, it is implementation-defined whether the replacement with std::this_thread::yield() occurs at all. That’s important for bare-metal systems: turning a deliberate halt loop into a cooperative yield could introduce behaviour the programmer never intended.

Conclusion

while (true); being undefined behaviour was one of those C++ facts that surprised everyone who heard it. It was an unnecessary divergence from C, it broke real embedded code, and compilers genuinely exploited it. C++26 fixes it — trivial infinite loops are now well-defined, and the compiler can no longer optimize them away.

The Daily Front Page 12 of 28
Friday, September 18, 2026 The Daily Front No. #260918 — Hallucinations at Command
article

US Military had close call after using AI for hallucinated intelligence report

by realsarm·▲ 420 points·318 comments·cnn.com ↗
The intelligence report, circulated across the US military this spring, immediately set off alarm bells.

US Army soldiers conduct unmanned aerial system training in April 2026.

US Army soldiers conduct unmanned aerial system training in April 2026.

Spc. Alva Gonzalez/US Army

The intelligence report, circulated across the US military this spring in the midst of the war with Iran, immediately set off alarm bells: A Chinese ship in the Middle East was transporting components of a nuclear weapons program.

The US military swung into action with plans to intercept the vessel, according to four sources familiar with the episode. According to two of the sources, armed members of the US military were preparing to board the ship. Military planes were in the air, one of those sources and another source familiar with the incident said.

It was only just before the planned operation that officials dug deeper into the report put together by a special operations command analyst and found it had been generated with the help of artificial intelligence (AI) — and that a chatbot the analyst had used inaccurately identified the material the ship was carrying. CNN was not able to learn what the misidentified cargo was.

The report, according to one of the sources, was “entirely false.” But it also “almost started a war,” the source said. Any US operation against a Chinese vessel could have risked spiraling into an armed conflict between the two nations.

Across the US military and the intelligence community, officials are pushing to weave AI into nearly every facet of their work, from analyzing the huge volumes of raw intelligence the US collects and selecting targets for strikes, to more mundane applications like managing budgeting, logistics and supply chains.

But the episode underscores the profound risks of using this powerful, new and relatively poorly understood technology for targeting in the middle of a war. Analysts have long feared that AI could lead to a catastrophic miscalculation if nation states are relying on poor or corrupted data — the kind of miscalculation that might lead the United States to fire on a Chinese ship based on inaccurate information.

In this particular instance, the analyst queried a chatbot about some intelligence reporting on the ship’s manifest that originated with US Special Operations Command Pacific, based in Hawaii. It was not clear whether the chatbot was a commercially available one or a US government product.

“The internal tools are mostly just copies of the commercial stuff wearing lipstick,” a former senior US official familiar with the AI systems used by military and intelligence analysts.

The bot fused together open-source intelligence with secret signals intelligence in government holdings and reached its fateful conclusion about the material the ship was carrying.

The analyst then used AI again to package the findings into a standard intelligence report — the kind that is trusted by military officials — and disseminated it.

US Special Operations Command Pacific and the Pentagon did not respond to a request for comment.

The rationale for the rapid adoption of AI is that it can help the military make battlefield decisions, like which targets to strike or which military assets to move where, faster. Officials say the US can’t afford to fall behind in integrating AI in case it must one day fight China or another adversary who would potentially be able to stay one step ahead of the US.

In January, Defense Secretary Pete Hegseth released his agency’s “Artificial Intelligence Acceleration Strategy” in a bid to speed up the military’s use of AI.

“We will unleash experimentation, eliminate bureaucratic barriers, focus our investments and demonstrate the execution approach needed to ensure we lead in military AI,” Hegseth said in a speech announcing the strategy.

Secretary of War Pete Hegseth testifies during a Senate Appropriations Committee hearing in the Dirksen Senate Office Building on Capitol Hill on July 21, 2026 in Washington, DC.

Secretary of War Pete Hegseth testifies during a Senate Appropriations Committee hearing in the Dirksen Senate Office Building on Capitol Hill on July 21, 2026 in Washington, DC.

Anna Moneymaker/Getty Images

The strategy also pushes for its broad use across the military, ordering the department to make AI available via several programs with the aim of “democratizing AI experimentation and transformation across the Department by putting America’s world-leading AI models directly in the hands of our three million civilian and military personnel, at all classification levels,” a memo announcing the strategy said.

But the effort is decentralized, multiple US officials familiar with the dynamic said, with different parts of the government using different tools under different orders and safety standards. There’s no one set of standards for how the US verifies the information generated by these tools. The constellation of different AI systems being deployed by disparate corners of the military and intelligence community means that the relative reliability and functionality vary widely.

For weeks, Washington policymakers have been intensely debating AI after a series of dire warnings from Silicon Valley engineers and tech CEOs of the possibility that AI could break free of human constraints, with potentially civilization-ending consequences.

But the episode with the Chinese ship underscores a different, and more immediate risk: human beings making disastrous decisions based on inaccurate or misleading information generated by AI or other automated systems. Sources said that the military is rapidly turning to AI to help with targeting, an area which holds the obvious risk of fatal mistakes.

“AI in targeting is definitely something that is ramping up and there is no real guidance for how having a human in the loop will prevent civilian casualties or fratricide,” another source familiar with the military’s current policies said.

The kind of “hallucination” that the tool used by the analyst in this case conjured has not been an isolated incident across the intelligence community since these tools began proliferating across government, according to one of the sources.

For some older intelligence officials — even those who broadly support the use of AI inside the military — AI has put pressure on analysts to produce and disseminate intelligence faster, opening the door for mistakes. Young analysts in particular, several sources said, are natives on these tools and more likely to trust them uncritically.

“AI allows you to get to a bad idea faster,” one of the sources said.

The Daily Front Page 13 of 28
Friday, September 18, 2026 The Daily Front No. #260918 — The Laser and the Lock
article

Photon-Emission-Guided Laser Fault Injection Enables RP2350 Secure Debug

by synack·▲ 167 points·61 comments·donjon.ledger.com ↗
Laser pulses at two nearby positions then restored debugger access to the chip’s Secure world.

Differential photon-emission microscopy localized debug enable register activity and narrowed the laser search before SWD-guided injection set the two bits required to restore Secure debug on an RP2350 A4.

TL;DR

— Photon-emission microscopy allowed us to locate a register responsible for the enabling of debug features on the Raspberry Pi microcontroller.

— Laser pulses at two nearby positions then restored debugger access to the chip’s Secure world, even though debug had been permanently disabled.

— Using that access after a rescue reset, we recovered a secret from one-time-programmable memory. The reset halted the chip before firmware could apply its runtime lock, so the page stayed Secure-readable.

— The attack requires physical access, destructive preparation, and approximately $250,000 of laboratory equipment.

The RP2350 security model

The RP2350 is Raspberry Pi’s dual-core microcontroller: each processor socket can select either an Arm Cortex-M33 or a RISC-V Hazard3 core at boot. Its hardware security features include:

  • Secure boot, which authenticates signed firmware against public-key fingerprints provisioned in One-Time Programmable memory (OTP)
  • The Armv8-M TrustZone, which separates Secure and Non-secure execution states
  • Permanent debug-disable settings
  • Glitch detectors intended to detect timing disturbances caused by clock or supply manipulation

Raspberry Pi has actively invited researchers to evaluate these protections through its RP2350 Hacking Challenges. The first challenge ran from August to December 2024 against the original chip. After several findings were addressed, Raspberry Pi released the A4 revision—the version we tested.

The permanent security configuration and boot public key fingerprints are stored in one-time-programmable (OTP) memory: each bit can be flipped from 0 to 1 once and never back, so whatever is written there lasts for the lifetime of the chip.

OTP is organised into 128-byte pages protected by two persistent, or hard, lock rows: for page n, PAGEn_LOCK0 configures optional read and write keys and the behaviour when no key is entered, while PAGEn_LOCK1 contains the hardware-enforced LOCK_S and LOCK_NS permissions. Those states can advance from read-write to read-only or inaccessible but cannot become more permissive.

The OTP subsystem uses redundant encodings for security-related fields: critical flags are “encoded with a three-of-eight vote across eight consecutive OTP rows”, and OTP lock bits are “triple-redundant with a majority vote”, according to the RP2350 datasheet.

At an OTP reset, the persistent LOCK_S and LOCK_NS values initialise a per-page runtime lock, also called a soft lock. Firmware can tighten this lock until the next OTP reset, but cannot loosen it. The runtime change does not survive that reset.

An external debugger communicates with the RP2350 through Arm’s Serial Wire Debug (SWD) interface. Requests first reach the Serial Wire Debug Port (SW-DP) and are then routed to access ports. In the Cortex-M33 configuration used here, each core has a memory access port (Mem-AP) connected to its system bus; an enabled Mem-AP lets the debugger read and write permitted memory and peripherals. A separate always-on access port, the RP-AP, exposes a small set of reset and recovery controls.

Secure debug refers to Mem-AP access with Secure attribution. The debugger can then transact with Secure memory-mapped resources the access-control logic permits, and halt or inspect a core running in the Secure state.

The permanent CRIT1.DEBUG_DISABLE flag is intended to close this path. When set, it drives the enable signals for both cores’ Mem-APs to zero, which “prevents the APs from performing any bus accesses at all”, and disables the factory-test JTAG interface and the RISC-V debug module’s access port. The SW-DP and RP-AP still respond, but neither core Mem-AP can access the system bus.

There is, however, an override: the memory-mapped DEBUGEN register lets Secure software re-enable each core’s Mem-AP and, separately, Secure accesses through it. The datasheet states that DEBUG_DISABLE “can be fully overridden by setting all bits of this register”.

This critical override in the enforcement chain is what made the debug interface our target. Gaining access to Secure debug on a Mem-AP is a general-purpose primitive to read and write Secure memory, halt and single-step a core, and inspect its registers. Whether that register could be set by a fault is the question the rest of this post answers.

Experimental setup

Target configuration

Raspberry Pi’s RP2350 Hacking Challenge asked participants to extract a 128-bit secret stored in OTP1. At startup, the signed challenge firmware ensures that page 48 has the expected persistent lock, then applies a runtime lock that denies both Secure and Non-secure access to the secret until the next OTP reset.

We replicated this vendor-defined configuration on our own revision A4 device:

  • Programmed the SHA-256 fingerprint of our public key into BOOTKEY0
  • Set BOOT_FLAGS1.KEY_VALID to 0x1 and BOOT_FLAGS1.KEY_INVALID to 0xe
  • Enabled secure boot (CRIT1.SECURE_BOOT_ENABLE = 1)
  • Permanently disabled debug (CRIT1.DEBUG_DISABLE = 1)
  • Enabled the glitch detectors at maximum sensitivity (CRIT1.GLITCH_DETECTOR_ENABLE = 1, CRIT1.GLITCH_DETECTOR_SENS = 3)
  • Configured the persistent locks for pages 1 and 2 according to the challenge configuration
  • Set the page 48 persistent lock to PAGE48_LOCK1 = 0x3c3c3c, which denied Non-secure access (LOCK_NS = INACCESSIBLE) while retaining Secure read-write access (LOCK_S = READ_WRITE)

Enabling secure boot permits only the Cortex-M33 cores, so both processor sockets used Arm for these experiments.

Sample preparation and bench

The device was backside decapsulated, so that infrared light reaches the transistors through the silicon substrate rather than being blocked by the metal layers on the front. The chip was then soldered back onto a daughterboard connected to Scaffold, Ledger Donjon’s open source platform for driving and monitoring devices under test. Removing the lead frame on the backside of the chip breaks its GND connection, so a copper wire restores it2.

Backside-decapsulated RP2350 mounted on the analysis daughterboard

Backside-decapsulated RP2350 mounted on the analysis daughterboard

Experimental bench used for the attack

Experimental bench used for the attack

DEBUGEN: overriding permanent debug disable

DEBUGEN has five functional bits:

BitNameEffect0PROC0Enable core 0’s memory access port1PROC0_SECUREPermit Secure accesses through core 0’s memory access port2PROC1Enable core 1’s memory access port3PROC1_SECUREPermit Secure accesses through core 1’s memory access port8MISCEnable additional debug components, including the cross-trigger interface and the RISC-V debug access port

Secure debug on a core needs both of its bits: the one that enables the Mem-AP, and the one that permits Secure accesses through it.

In contrast to the redundant encoding used for OTP security fields, the datasheet documents no bit redundancy, parity or majority vote for DEBUGEN.

We therefore tested whether laser pulses could set DEBUGEN bits on the secured device described above.

Photon-emission-guided localization

That test first requires knowing where to aim. Setting an individual DEBUGEN bit means hitting the storage of a single register bit, a needle in a haystack. This is a harder targeting problem than the instruction-skip faults common in laser fault injection, where disturbing any of the many flip-flops in a core pipeline can produce the same skip: that spreads the sensitive area widely enough for a random scan to find it. A blind scan for one DEBUGEN bit is impractical.

Switching transistors emit faint near-infrared photons correlated with their activity, so collecting that emission over repeated execution can reveal where a selected control changes state. This made photon-emission microscopy (PEM) a good fit for DEBUGEN: as a memory-mapped register, Secure software can toggle exact bits in a loop, driving the repeated state changes the measurement needs. We used it as the first localization stage, and the resulting map constrained the subsequent laser scan to a region of a few micrometres.

We compared loops that repeatedly toggled selected DEBUGEN bits on and off, differing only in the bits they targeted. A register’s photon emission is faint next to the camera’s own noise and sensitive to slowly drifting ambient conditions such as temperature, so a single frame reveals nothing. Averaging many frames of each loop suppressed random sensor noise, and subtracting the two mean stacks cancelled everything the loops shared: static background, sensor offset, thermal emission, and switching unrelated to the selected bits. Interleaving the two values during acquisition kept slow drift from biasing that subtraction. What remained was the emission that tracked the selected bits.

Mean stacks of 200 full-view photon-emission captures for DEBUGEN masks 0x3 and 0xc followed by their signed difference.

Mean stacks of all 200 mask 0x3 and mask 0xc acquisitions, followed by their signed difference. Red is positive, indicating greater emission for 0x3; blue is negative, indicating greater emission for 0xc. The localization maps below additionally balance acquisition order before combining matched differences.

Repeated comparisons across different bit masks exposed compact sites associated with DEBUGEN bits 0–3 across three regions of the camera field.

Infrared overview of the die with three marked regions, plus zooms of those regions overlaid with coloured DEBUGEN bit sites.

Infrared overview of the camera field, with three marked regions. Coloured pixels mark sites associated with `DEBUGEN` bits 0-3.

These zones show switching activity associated with each DEBUGEN bit; they do not directly identify storage cells. The multiple hotspots observed for each bit may arise from the storage element or from related logic. Without layout data, we cannot distinguish between the two. However, these zones still significantly reduce the search space.

Finding 1 — faulting DEBUGEN gives Secure debug

For laser fault injection (LFI), we used a pulsed laser at 980 nm with 2.97 W maximum optical power, operated at roughly 40% (about 1.2 W), with a 100 ns pulse width through a 50x objective. After each pulse, we probed the debug access ports over SWD.

Within the area found from PEM, we ran an LFI scan and used that SWD feedback to calibrate two responsive positions a few micrometres apart. At one position, pulses enabled bus access through core 1’s Mem-AP, indicating that PROC1 was set. At the other, the Mem-AP’s Control/Status Word reported SDeviceEn = 1, a state-guided signal that PROC1_SECURE was likely set. We checked both indicators after every pulse.

Side-by-side infrared views with PEM bit sites on the left and LFI fault points on the right.

Left: PEM sites associated with `DEBUGEN` bits. Right: laser-fault points on the LFI infrared view.

A pulse that set one bit could clear the other, so setting both required an iterative sequence. Our script pulsed the PROC1 position until bus access was available, then pulsed the PROC1_SECURE position until SDeviceEn = 1, returning to the first position whenever bus access was lost. Once the positions and pulse parameters were calibrated, the sequence enabled Secure debug within seconds. Interestingly, we could not reproduce this sequence using a 20x objective. Because the two positions are only a few micrometres apart, that wider spot likely hit both the region that sets a bit and the one that clears it, so it was not possible to obtain the correct value.

Once both bits were set, they remained set without further pulses or software writes. Reading the Secure-only DEBUGEN register through core 1’s Mem-AP then returned 0xc; because DEBUGEN is Secure-only, that successful read confirms the transaction was Secure-attributed.

Enabling Secure accesses through core 1’s Mem-AP allows the debugger to read and write memory-mapped resources whose ACCESSCTRL permissions admit the debugger as a bus manager and whose target-specific controls admit Secure AHB transactions. Independently of those direct reads, the debugger can halt and single-step the core and inspect or modify its registers, compromising TrustZone runtime isolation through Secure-core-mediated extraction. This does not make the boot ROM accept unauthenticated firmware: when firmware boots normally, secure boot still authenticates it, but cannot protect runtime state that remains accessible to the debugger after verification.

Application to the Hacking Challenge configuration

The Secure-attributed Mem-AP access described above exposes Secure runtime state, but the challenge’s page 48 runtime lock still prevents access to the secret after firmware has run. The page’s persistent lock, PAGE48_LOCK1 = 0x3c3c3c, denies Non-secure reads but leaves LOCK_S at READ_WRITE, so it remains readable through Secure-attributed accesses before the runtime lock is tightened.

During each boot, the firmware writes the most restrictive binary value, 0b1111, to the runtime lock otp_hw->sw_lock[48]. That register then makes the page inaccessible to both Secure and Non-secure accesses, including Secure debug, and therefore prevents Secure transactions from the Mem-AP from reading the secret.

As documented, software locks “are initialised from the OTP lock pages at reset”, and a write only advances the state “until next reset”. Resetting the OTP block discards 0b1111 and restores the value derived from PAGE48_LOCK1, for which LOCK_S = READ_WRITE.

The remaining question is how to reset a locked chip without allowing firmware to re-apply the runtime lock. The RP-AP remains “always accessible, even when external debug is disabled”. Setting CTRL.RESCUE_RESTART triggers a rescue reset: a full system reset that also flags the boot ROM to halt before any user software runs.

The boot ROM checks POWMAN_CHIP_RESET.RESCUE_FLAG before watchdog, flash or USB boot, clears it, then holds core 0 in an interrupt-disabled wait loop and core 1 in its wait-for-vector path.3 The datasheet documents no restriction on CTRL.RESCUE_RESTART.

We proceeded in the following sequence:

  1. Rescue reset. Set CTRL.RESCUE_RESTART to 1, then clear it to 0 through the RP-AP. The chip resets and remains in boot-ROM wait paths. The signed firmware never runs, so sw_lock[48] is never tightened and stays at the permissive value derived from PAGE48_LOCK1LOCK_S = READ_WRITE.
  2. Fault DEBUGEN to 0xc. With both cores in boot-ROM wait paths, set PROC1 and PROC1_SECURE as described above; these two set bits produce the value 0xc.
  3. Halt core 1 through its Debug Halting Control and Status Register (DHCSR) over the now-Secure Mem-AP.
  4. Read the secret from OTP rows 0xc080xc0f through the guarded read interface.

We ran this sequence on the tested device and recovered the complete challenge secret.

DEBUGEN_LOCK does not prevent laser-induced changes

DEBUGEN_LOCK blocks software writes to the corresponding DEBUGEN bits: each lock bit is “Write 1 to lock the […] bit of DEBUGEN. Can’t be cleared once set”. The datasheet presents this as a way “to avoid accidental writes”.

In trials with the target DEBUGEN bit at 0 and its lock bit at 1, a pulse could still set DEBUGEN while the lock remained 1. Pulses also set lock bits, with or without a corresponding DEBUGEN change. In successful sequences, all five functional lock bits were 1 by the time PROC1 and PROC1_SECURE were both set. We never saw a lock bit return from 1 to 0, so a later write of DEBUGEN = 0 cannot restore the disabled state once the fault has set the corresponding lock.

Limits of software-based mitigations

Once Secure accesses through the Mem-AP are enabled, Secure attribution alone no longer separates the debugger from Secure firmware. This access does not override hard OTP locks or peripheral-specific controls. ACCESSCTRL can block direct debugger-manager transactions to particular targets, but it does not by itself prevent a debugger controlling the Secure core from causing core-originated accesses or extracting loaded values through core registers. After a rescue reset, ACCESSCTRL returns to its all-open reset-time defaults before firmware can reconfigure it. ACCESSCTRL therefore reduces direct Mem-AP exposure rather than forming a standalone confidentiality boundary.

Firmware can nevertheless reduce post-boot exposure by denying the debugger access to sensitive targets in ACCESSCTRL, then setting the debugger bit in ACCESSCTRL.LOCK so that debugger transactions cannot reopen those permissions. Secure firmware can also check DEBUGEN periodically and, on an unexpected value, trigger a fail-safe reset that clears the processor-cold reset domain. These measures are best-effort runtime mitigations: an enabled debugger may halt the core before the next check, and rescue reset stops before firmware can configure ACCESSCTRL or run the monitor. They therefore do not prevent the pre-firmware secret read demonstrated here.

RP2350’s documented encrypted-boot flow illustrates the pre-firmware limitation of runtime locks and the post-boot limitation of debugger-manager filtering in two distinct machine states. After rescue reset, the boot ROM halts before decryption: no plaintext payload exists yet, but the decryption key may be directly readable if the OTP page’s persistent permissions allow Secure access and no other target control blocks the transaction. After normal encrypted boot, plaintext exists in SRAM: direct Mem-AP reads depend on debugger-manager permissions in ACCESSCTRL, while Secure-core control may permit core-mediated extraction even when direct reads are denied. This is architectural analysis, not a tested encrypted-boot result; encrypted boot still protects external flash from offline inspection.

Impact and attack requirements

The demonstrated sequence provides Secure-attributed memory access, control over Secure-world execution, and access to the challenge secret after resetting its runtime page lock. It requires the following resources:

  • Destructive physical access. Backside decapsulation permanently modifies the package and leaves the die exposed.
  • Specialised laboratory equipment. The complete setup described above costs approximately $250,000.
  • Hardware-security expertise. The procedure requires sample preparation, die navigation, laser parameter selection, and coordinated laser control, stage positioning, and SWD measurement.

Conclusion

The RP2350 encodes critical debug-disable flags in OTP with redundant voting, but DEBUGEN can override their effect and has no equivalent protection documented in the datasheet. In our experiments, laser pulses changed DEBUGEN despite DEBUGEN_LOCK and could set a lock bit that prevented firmware from restoring the disabled value. Separately, the RP-AP rescue reset restored the challenge’s runtime page lock to its persistent value while preventing user firmware from executing. The software-visible mechanisms each performed their documented function, but their interaction with the laser fault enabled Secure debug and recovery of the challenge secret. Differential PEM first isolated bit-dependent DEBUGEN activity, and guided LFI converted that spatial lead into persistent Secure debug. The system-level lesson is that security analysis must cover the complete enforcement path, from persistent OTP configuration through mutable control registers and reset behaviour, because system security depends on that path rather than on individual mechanisms in isolation.

Disclosure and acknowledgements

We disclosed this fault to Raspberry Pi on 28 July 2026. We thank the Raspberry Pi team for their engagement in the disclosure discussions and for their transparent approach to security research.

Footnotes

  1. https://github.com/raspberrypi/rp2350_hacking_challenge The RP2350 Hacking Challenge repository, containing the reference lockdown configuration and firmware we replicated.
  2. Courk, Laser Fault Injection on a Budget: RP2350 Edition.
  3. The rescue check is step 1 of the core 0 boot path in src/main/arm/varm_boot_path.c; in src/main/arm/arm8_bootrom_rt0.S, varm_wait_rescue enters the interrupt-disabled varm_dead_quiet WFI loop while core 1 remains in the boot ROM’s wait-for-vector path.
The Daily Front Page 14 of 28
Friday, September 18, 2026 The Daily Front No. #260918 — Stopping the Storm
article

How Uber Protects Against Retry Storms

by iscmt·▲ 112 points·49 comments·uber.com ↗
Retry storms historically impact business operations and brand trust.

Uber's service call chain diagram showing error handling and retry prevention, highlighting 9.5M spurious requests stopped.

Introduction

Retry storms historically impact business operations and brand trust. While retry configuration tuning and retry budgets provide meaningful mitigation at the service level, they’re manually configured and lack visibility into cross-service amplification caused by deep dependency chains and fan-out patterns. As a result, it can be difficult to shield infrastructure against the domino effect triggered by a single service outage deeper in the stack.

A key reason is that retry behavior today isn’t context-aware. While we can control how many retries occur, we can’t precisely control when they occur. This stems from the challenge of reliably distinguishing between errors generated by a service and those merely propagated through it.

As a result, retries are applied uniformly rather than conditionally.

This approach works for transient or low-rate failures. However, during moderate or severe degradation, it becomes counterproductive. Aggressively retrying against an already struggling service increases load, accelerates failure, and amplifies retry traffic across upstream dependencies. What begins as a localized outage can quickly escalate into a stack-wide incident—ultimately degrading, or in the worst case, completely breaking, the end user experience.

One might argue that error codes from downstream services could be translated upstream to provide context for retries. While theoretically possible, this approach doesn't scale at Uber due to large fan-in and fan-out, evolving call flows, and the need for frequent adaptive changes. Therefore, we developed a context-aware mechanism in shared infrastructure to handle errors more efficiently. This blog explains the mechanism.

Background

Consider a simple call chain as shown in Figure 1, where the total number of requests arriving at NodeA is Ƞ. By deduction, all nodes B, C, D, E, F, and G serve Ƞ requests in the steady state (when no node errors out).

Seven blue circles labeled A to G connected by rightward arrows in a straight horizontal line.

Figure 1: Call-chain with 1:1 fan-out, where a node calls its downstream exactly once for any incoming request.

If service D starts erroring out and each service is configured to retry once (1 regular attempt and another attempt if the downstream fails), let’s look at the total number of requests served by each node.

Seven circles labeled A to G in a row with arrows; D is highlighted in red.

Figure 2: Call chain where a service errors out.

Node A B C D E F G
Depth 0 1 2 3 4 5 6
Requests Served Ƞ 2 × Ƞ 4 × Ƞ 8 × Ƞ 8 × Ƞ 8 × Ƞ 8 × Ƞ

This can be distilled down to a simple formula, assuming the number of retries R is the same at every hop. ɗ denotes the depth of the node in the call chain, the number of requests served by the node if the node creates or passes through an error:

Rɗ × Ƞ

Retry Budgets

We can optimize this by introducing retry budgets. Let’s assume the same retry budget at every hop represented by B. The new formula becomes:

(1+B)ɗ × Ƞ

Now, let’s try to see the number of requests served with a retry budget of 10%:

Node A B C D E F G
Depth 0 1 2 3 4 5 6
Requests Served Ƞ 1.1 × Ƞ 1.21 × Ƞ 1.33 × Ƞ 1.33 × Ƞ 1.33 × Ƞ 1.33 × Ƞ

In the above example, the error originates at NodeD, and if we limit the retry to only between NodeD and NodeC , and restrict entirely NodeA and NodeB from retrying on this error, we can guarantee a similar availability of the call-path without overburdening NodeD, NodeE, NodeF , and NodeG.

Error Ownership

Consider the same example of retry budgets while restricting retries between the edge from NodeC to NodeD, where the error originates.

Node A B C D E F G
Depth 0 1 2 3 4 5 6
Requests Served Ƞ Ƞ Ƞ 1.1 × Ƞ 1.1 × Ƞ 1.1 × Ƞ 1.1 × Ƞ

Here, we clamp down the total number of requests served by all nodes from D till the leaf node G to just 10% over baseline, while allowing at least once retry for up to 10% of errors when they’re first returned. But what about the availability of NodeD as seen by NodeC? Let’s run some numbers for various availability scenarios, and try to calculate ‌availability after retry.

Base Availability % Base Error Rate % Retry Budget Error Rate after Retries % Availability after Retries %
99.9 0.1 10% 0.0001 99.9999
99 1 10% 0.01 99.99
95 5 10% 0.25 99.75
90 10 10% 1 99
80 20 10% 12 88
70 30 10% 23 77

As shown in the table above, for availability drops up to 10% in the callee node, even a single retry is helpful in getting the perceived availability by the caller node up to 99%. Beyond this, perceived availability drops significantly as a good chunk of requests are never retried due to the retry budget in place.

This calculation assumes the errors from the callee are independent and that retries will lead to recovery. However, in many real-world scenarios like service overload, bad database hosts, database overload, or sharding issues, the probability of retries remains high even with retries.

This contradicts the idea that retries to callee always increase perceived availability to the caller. It’s also this intuition that forms the basis of error ownership. During periods of high error rates from a service, the errors are less likely to be randomized, and wouldn’t benefit from a higher number of retries, and instead might be responsible for further degradation.

Architecture

The solution is about establishing error ownership, which can be explained using the symptom versus cause analogy.

If a service calls N outbounds for fulfilling a request, and if an outbound error-out causes it to return an error, then the error returned by that service is only a symptom. Simultaneously, in the context of the service, the cause is the incoming error from its downstream.

However, if no outbound of the service errors out while fulfilling the request and it still returns an error, the service is the cause of the returned error, and is the owner. In the next section, we discuss some possible solutions that can leverage this.

Simple Correlation

Claiming Error Ownership

We use the Service Dependency Analysis Solution to correlate an inbound failure with an outbound failure and use the ruleset shown in Figure 3 for making or refuting error claims.

Flowchart for handling errors in service S, detailing claim and unclaim logic based on downstream call outcomes.

Figure 3: Decision logic for claiming error ownership.

Retrying with Error Ownership

The caller uses the logic shown in Figure 4 to determine if it should retry the request.

Flowchart for handling downstream error responses based on x-uber-error-claim header presence and value.

Figure 4: Decision logic for allowing retries.

While the scenario of missing error claim headers is an uncooperative environment, it could occur because the downstream service doesn’t have the service dependency analysis solution, and is unable to correlate outbound and inbound errors. Or, the downstream service has missing context propagation, resulting in an incomplete correlation between outbound and inbound errors.

Here, the first node to see a missing error claim from a downstream unclaims the error, limiting the impact radius of the retry disturbance (it’s no longer a storm), while still allowing sufficient retries to the error-returning service.

Decision Matrix

Callee Error Caller Error Callee Error Claim Caller Should Retry (Retry Middleware) Caller Propagated Error Claim
No Yes NA NA Claim
Yes Yes Missing Yes Unclaim
Yes Yes Claimed Yes Unclaim
Yes Yes Unclaimed No Unclaim

Three pink circles labeled A, B, and C connected by arrows pointing from A to B to C.

Figure 5: 3 nodes used to demonstrate caller errors.

When the edges A -> B and B -> C are fail-close, and the Node C returns an internal error that it’s claimed, Node B upon seeing the claimed error from C should retry to C. However, if the retry fails, it’d propagate the error to Node A, but while returning the error it must unclaim it. Node A upon seeing the error from B and the unclaimed error header shouldn’t retry the request to Node B.

Flowchart showing error propagation and retry logic between Node A, Node B, and Node C with fail-close edges.

Figure 6: Decision logic for Error claim propagation.

Coincidental Errors and Why We Need Service Dependency Analysis

The decision matrix above covers cases where the downstream call fails or there’s an internal server error. There could also be scenarios where both happen simultaneously, as shown in Figure 7.

Node A connects with arrows to nodes B and C in a simple directed graph.

Figure 7: Example where Node A has 2 fail-open dependencies, Node B and Node C.

In this example, Node A is also experiencing a 10% error rate because of an overloaded cache/database that isn’t tracked via the service dependency analysis solution. This makes it an internal server error of Node A, and Nodes B and C also have a 10% error rate for unrelated reasons.

The service dependency analysis solution tries to attribute errors to downstreams first and itself last, so even when Nodes B and C are fail-open, it assigns the blame to them when the failures are colocated. This unclaims the error and prevents the upstream of Node A from retrying to it and possibly recovering.

The impact of missed legitimate retry opportunities in such a scenario is significant. If Node A serves 100 requests, out of which 10 experience internal server errors originating at Node A, 1.9 of those 10 requests would be incorrectly unclaimed by Node A as an error not originating from itself, resulting in a missed legitimate retry opportunity.

Based on the analysis of the 6 months of service dependency analysis solution data, coincidental errors like a legitimate server error or a fail-close dependency error happening during a fail-open dependency error are extremely rare. In the worst case scenario, for 80% of edges with over 100 callee failures in a minute, around 2% of the times the caller failed too (a coincidental failure). Using the example above, we’d get 0.396 out of 10 requests that’d be incorrectly unclaimed by Node A.

However, even if we fixate on the worst case, since the service dependency analysis solution creates a memory of failure patterns, it can leverage this memory to only unclaim errors when inbound failures correlate with outbound failures in fail-close dependencies. This eliminates the risk of retry suppression during coincidental errors.

Flowchart for handling node errors, deciding to unclaim or preserve retries based on dependency failure correlation.

Figure 8: Decision logic for determining when to retry.

Edge Case Scenarios

Guaranteeing At-Least-Once Retries

Many services don’t have retries configured for their fail-close outbounds. Eliminating retries from callers of these services would lead to availability drops along the incoming caller chain.

This is solved by introducing a flag to signal whether retry criteria is satisfied for an error returned by a downstream. When the retry middleware sees an error from the downstream, it can compute whether the retry criteria is satisfied and pass that along to the service dependency analysis solution.

Sequence diagram showing outbound call chain with retry logic and error handling across four service components.

Figure 9: Retry logic to guarantee at-least-once retries.

The Retry Middleware marks retry criteria satisfied if any of the following conditions shown in Figure 10 are met.

Flowchart detailing retry middleware logic for handling downstream errors and retry criteria satisfaction.

Figure 10: Decision logic to determine whether the retry criteria is satisfied.

The service dependency analysis solution leverages this retry criteria satisfied signal from the retry middleware to decide at the inbound whether it needs to own the error returned by the fail-close downstream.

Flowchart for handling service errors with decision points for fail-close failures and retry criteria.

Figure 11: Decision logic for service error ownership.

Here’s the end-to-end flow:

Sequence diagram showing error handling and retry logic between Upstream A, Service B, SDA@B, RetryMW@B, and Downstream C.

Figure 12: End-to-end flow.

We don’t introduce any new retries along the call chain. The retry-at-least-once behavior only works if at least one service along the call chain has retries configured.

By introducing an at-least-once-retry guarantee, we eliminated availability drops in our call chains. At the same time, we safeguarded them against retry storms by leveraging error ownership.

Context Drop Handling

Four pink circles labeled A, B, C, D connected by right-pointing arrows in a linear sequence.

Figure 13: 4 nodes used to demonstrate context drop handling.

For the service dependency analysis solution and error ownership, intra-service context drops shift the retryable error left. Consider the example in Figure 14, where Node D returns an internal error to Node C. The intra-service context in Node C is broken, so the service dependency analysis solution can’t correlate the outbound request from Node C to D with the incoming request from Node B to C.

As a result, Node C claims the error, while decoupled retries at the retry middleware continue to happen between Node C and D. A set of retries also happens between Node B and C, as Node C must claim the error it’s returning. However, when the retries fail and Node B returns an error to Node A, it also returns the unclaimed error header, which should prevent Node A from retrying the request to Node B. In this case, we cut down the total number of requests to Nodes B through D by half, assuming a single retry (1 attempt and 1 retry) configuration at all nodes. If there were 5 nodes to the left of Node A, the worst case would’ve had 32 times more requests without error ownership propagation.

Sequence diagram showing error propagation and retry logic across four nodes with context drop and claim handling.

Figure 14: Description.

Error Ownership in Production

Everything described so far isn’t theoretical—error ownership is fully implemented and operational across Uber’s service mesh today. The scheme runs in the retry middleware and the Service Dependency Analysis Solution that sit in the request path of our user-facing APIs, continuously claiming and unclaiming errors as traffic flows through deep dependency chains. Because it’s embedded in shared infrastructure, services inherit retry-storm protection without bespoke per-service error-handling logic. The following real-world incident demonstrates how this production deployment behaved under a genuine large-scale degradation.

Use Cases at Uber

On November 18th, 2025, Uber had a major outage due to an issue with a Core Entity service. The service lives over 5 levels deep in our call chain and is critical for business operations. It started returning a very high error rate due to an underlying infrastructure issue. This error was quickly propagated up the call chain. During this time, many upstream caller services had enough opportunity to retry the failed requests returned by these services. With simple retry budgets, this would’ve resulted in a 46%-135% traffic increase on the degraded service, prolonging the outage by diminishing chances of recovery. Because error ownership was already enabled in production, as described above, the system contained the blast radius automatically

Immediate retry attempts were stopped to the immediate callers of the degraded service, where some callers were stopped from making up to 200,000 additional requests. We also calculated the retries that were stopped at the ancestors of these immediate callers, and aggregated them at the root node. We learned that we could stop a staggering 9.5 million spurious requests in our service mesh. Those could’ve easily prolonged the outage by constantly hammering the degraded service.

Conclusion

Our approach has dramatically reduced the total request volume flowing through the call graph during degradation events.

To quantify this, we define the max retry storm radius after error ownership as the max depth of call path where a retry storm could happen after error ownership was enabled. It’s computed for the call graph of every root node. Across all our user-facing APIs, we got this value down to a maximum of 3, where the earlier value of max retry storm radius was up to 25. We also got the average down to 2 from 20.

Table listing edge-gateway services, endpoints, and their max retry storm radius values with pagination at the bottom.

Figure 15: Max retry storm radius after error ownership.

By limiting retries exclusively to error-owning services, where they can genuinely resolve the issue, we can prevent the exponential fan-out of requests characteristic of retry storms. This has helped safeguard Uber’s infrastructure from cascading failures triggered by single-point degradations.

Acknowledgments

Cover Photo Attribution: Generated with ChatGPT by OpenAI; no external images, logos, or third-party assets used.

The Daily Front Page 15 of 28
Friday, September 18, 2026 The Daily Front No. #260918 — Proof by Prompt
article

I vibed a proof of Conway's conjecture

by m-hodges·▲ 219 points·195 comments·overreacted.io ↗
I believe I’ve obtained a Lean proof.

A few months ago, AI math results started making headlines. “Do a breakthrough” became a Twitter meme. Naturally, I became curious whether I, too, a math noob, can find some open mathematical problem and then have a frontier model solve it.

It took me an entire month of my free time and a boatload of tokens, but I believe I’ve obtained a Lean proof of this conjecture posed by John Conway 50 years ago:

Conjecture: Omnific integers have a refinement property: if ab = cd for omnific integers, then there are further integers e, f, g, h with a = ef, b = gh, c = eg, d = fh.

Conway’s refinement conjecture claims that omnific integers have a refinement property: if ab = cd, there are integers e, f, g, h with a = ef, b = gh, c = eg, d = fh.

My proof has not been independently verified by mathematicians. However, I have decent reasons to believe the proof is correct, and I genuinely invite a refutation.

The proof has passed the mechanical checks from the Palomar registry, and a few people familiar with both Lean and the field said that the statement seems correct. So, assuming my proof doesn’t rely on a Lean kernel bug, it’s likely to be legit too.

In this post, I’ll describe my approach, and some things I learned along the way.


#First Day

I thought the idea of “solving” a math problem without understanding its substance is rather absurd, which of course made it all the more appealing.

However, I didn’t just want any result; I wanted something that pulls me.

#Choosing the Field

I asked Claude to pick an open problem in the field of surreal numbers. In case you’re not aware, surreal numbers are John Conway’s invention—or a discovery?—of a previously unknown number system containing all numbers great and small:

  • It contains all real numbers (the numbers we use like 0, –5, 36.6, square root of 2…)
  • It also contains all ordinal numbers (the infinitely large ω, the ω + 1 that comes after it, the ω * 2, and even ω * ω, at some point even the impossibly large ω^ω…)
  • Finally, it contains all kinds of unholy combinations of them, like 75 + ω*3 + 1/ω.

What is particularly miraculous about surreal numbers (and why I suppose they might appeal to a programmer) is that this rich system spawns from a single rule.

Take all the numbers you have so far. Then, “spawn” a new number in every gap between the numbers you already have (crucially, “to the left of all” and “to the right of all” also count as “gaps”). Apply this step forevermore, and you’ll get surreal numbers.

Think about it.

On the first day, the gap is “between nothing and nothing”. Zero is born.

On the second day, there are two gaps: “between nothing and zero” and “between zero and nothing”. Two numbers spawn in those two gaps. Call them –1 and 1.

On the third day, there are four gaps: a gap “between nothing and –1”, a gap “between –1 and 0”, a gap “between 0 and 1”, and a gap “between 1 and nothing”. Put a number in each of those gaps and then give them names: –2, –1/2, 1/2, and 2.

On the fourth day, we fill the eight gaps with –3 and 3 at the edges and –3/4, –3/2, 3/2, and 3/4 in the remaining gaps.

Suppose we actually do this forever (an infinity of “days”).

Jump to the “infinite-th” day (call it ω). With an infinite supply of the “already born” numbers, we’ll suddenly find infinitely many new “gaps” waiting to be filled that we couldn’t express before: “between [1, 2, 3, …] and nothing” (plus infinity?), “between nothing and […, –3, –2, –1]” (minus infinity?), “between 0 and [1, 1/2, 1/4, 1/8…]” (an infinitely small number?), “between [positive already born numbers whose squares are below 2] and [positive already born numbers whose squares are above 2]” (a square root of 2?), and so on. So, by this day, we have every real and we get our first infinitely large and small numbers (call them ω, –ω, 1/ω, and –1/ω).

Now suppose we keep on going then forevermore (through an infinity of infinities of days), and so ever on and on, never stopping birthing new numbers. It turns out that, if we go forever, the binary tree based on this single rule gives birth to every real, every ordinal, and more, with consistently definable arithmetic on them:

The surreal number tree

(Illustration by Joel David Hamkins. Go read his posts and buy his book!)

Surreal numbers are beautiful, and that’s why I asked Claude to work in that area.

#Choosing the Problem

Initially, I asked Claude:

Me: which unsolved problems in the Surreal Numbers research program pull you the most and why?

It suggested a number of things but noted some have had recent progress. I asked:

Me: get up to speed on the corners you want to know more about. follow your curiosity. you can look for older stuff too.

Then I further asked it to narrow down the choice to something concrete:

Me: what direction pulls you most? i want you to pick a problem to work on. be bold!

Claude replied:

Claude: I pick Conway’s arithmetic. Specifically, the question the L’Innocente–Mantova machinery just sharpened to a point — is every irreducible in K((ℝ^≤0)) with infinite support prime? — which, by their reduction, is now exactly equivalent to Conway’s 1976 conjecture that any two factorizations of an omnific integer admit a common refinement. It’s the last of Conway’s own conjectures about his own numbers still standing, and 2026 is ONAG’s fiftieth birthday. That’s the problem I’d work on.

I still don’t know whether this really was Conway’s last standing conjecture about the surreal numbers (let me know either way!), but this year being his book’s 50th birthday convinced me to choose this problem for sentimental reasons.

Here is the full transcript from that session. My last question to that session was whether we have a chance of formalizing the Lean statement of the conjecture in a relatively concise way—without that, even if I found a proof, there’d be no way for me to convince somebody to look at it. Claude said it can be stated without much trouble in Lean, and that answer seemed right, so I decided to take on this project.

(Note: I didn’t know this at the time, but Claude’s claim about the problem having been perfectly reduced was wrong; actually proving the conjecture required more than that.)

#The Problem Statement

While you’re probably here to learn more about my Lean/AI workflow, I’ll briefly explain the conjecture itself, since you already know enough to understand it.

In short, omnific integers are the integer part of the surreal number tree. So they include all regular integers like 3, –5, and so on, but also the weirder numbers like the infinitely large ω, 2ω, ω * ω, ω^ω, –ω/7 (yes, that’s a “whole” number), etc. If you look at the binary tree above, you’ll notice that the omnific integers are the surreal numbers that you get if you only ever go left (e.g. –5, –ω–1), or only ever go right (e.g. 3, 2ω), or only ever change directions exactly after infinite jumps (e.g. ω/2).

Now, the conjecture.

Conway suggested that if ab = cd, we can break a and b into pieces, and c and d will turn out to be the same pieces recombined. With regular integers, we take this for granted: take 210 = 10 × 21. We can break 10 down as 2 × 5 and 21 as 3 × 7, then reshuffle them into 2 × 3 = 6 and 5 × 7 = 35. The product is still 6 × 35 = 210. So when we see some equality like 10 × 21 = 6 × 35, we know that under the hood there’s actually four numbers being reshuffled: (2 × 5) × (3 × 7) = (2 × 3) × (5 × 7).

However, when you deal with infinities, things don’t always turn out as we expect. So the conjecture means Conway thought omnific integers had, in a sense, enough “structure” to keep this “nice” property of integers. And conveniently, the recent advances had mostly reduced the conjecture to the behavior of a certain kind of infinite series, and Claude thought we might now have enough to finish it off.

But can we actually do that solely with AI?

I’d say yes, and no, and yes.


#First Week

#One-Shotting, And Failing Badly At It

My first attempts involved plainly telling the AI to solve the conjecture (or to find a counterexample, according to its initial suggestion). I asked Claude which papers it wants to read, converted those papers to TeX so that the model doesn’t need to repeatedly decode PDFs, and told it to pursue its chosen approach, for example:

Me: hunting for counterexample sounds good. don’t hesitate to burn all my tokens on this. if you have to take weeks, that’s ok — do a breakthrough and find a structured counterexample. do not rest until you’ve found it, and don’t give up. if you stay focused and think hard and smart enough, i know that you can find it. game? i attached some stuff you might find helpful!

This didn’t work well. Although some ideas from these sessions might have ended up being relevant to the final proof, most output at this stage seemed to be word salad that Claude invented to justify its work and make it sound impressive.

Here is a typical example:

Claude: The stage-1 obstruction of C* is not rigid. With designed resonance — arithmetic-progression offsets making the window equations honest Toeplitz systems, which are exactly the graded convolution operators the formalism predicted — there is a strictly positive space of frontier data simultaneously compatible with both b and d, for coprime pairs, beyond anything divisibility can explain. And the control column confirms the resonance-necessity conjecture empirically: break the skeleton alignment and the joint kernel dies at the constrained window, exactly as the transversality heuristic predicted. So the specific fear built up by five closed doors — that Pitteloud-descended rigidity would propagate stage-wise and choke the correction system at birth — is answered: at stage 1, it does not. The den has air in it. This is the first pro-C* evidence the hunt has produced, and it comes with a clean structural reading: rigidity governs exact and finite configurations; the window systems, which are the native habitat of the transfinite construction, have generic slack of small but nonzero dimension. Drift fuel exists.

I thought this sounded like bad science fiction. It was using Claude’s unbearable metalanguage, gave cutesy names to some intermediate results without concretely justifying them, and kept being extremely dramatic. Of course I couldn’t verify its claims, but worse, it didn’t seem coherent enough to pass to a real mathematician for review. So it seemed like a dead end, and I had to look for a different approach.

#Restarting with the Skeptic

I got tired of Claudeisms, so I wanted to give ChatGPT a try; Sol in particular.

I’ve started my ChatGPT sessions by giving it the related papers and the output from the previous Claude sessions, with an explicit note that Claude’s “paper” is AI-generated, and I wanted to get ChatGPT’s opinion whether it is bullshit or not.

ChatGPT would say it’s mostly bullshit, pointing to the made-up terminology, dramatic claims, trivial results dressed up in fancy language, incorrect inferences, and other defects. While I had no way to judge if ChatGPT’s criticism is true (since I asked it to be critical), after Claude’s grandiosity, I quite enjoyed working with the more “skeptical” and restrained personality, and started using ChatGPT instead.

To retain the “skeptical” personality, I’d clone each ChatGPT session right after it had lambasted Claude’s “paper”. From that point, I’d ask ChatGPT to actually “do a breakthrough” on the theorem, and it started producing some “results”.

Unlike Claude, which either outright refused to work on the theorem (because it’s an unsolved conjecture and there is no chance of solving it) or got so deep into it that it would invent an entire universe of its own making, ChatGPT would think for 20 minutes, and then spit out relatively small claims, which it believed to be novel but directly following from the papers I fed it, and stated in plain language.

Before investing more time, I tried giving ChatGPT’s output to fresh ChatGPT sessions (with memory turned off) asking them to be critical (as with Claude’s output). Some of ChatGPT’s results started “checking out” between the runs, i.e. a fresh session found no issues. So in a sense I found some of ChatGPT’s “fixpoints”.

I’ve also started “forking” sessions, having them do these “breakthroughs”, and then copypasting the surviving ideas to yet another session that combined them together, looked for connections, and suggested next research directions. At this point I realized I couldn’t keep doing this by hand and needed a more robust setup.


#Second Week

#Setting Up a Laboratory

I’ve downloaded Codex locally to have more control over the workflow.

I’ve then set up a few sessions (i.e. agents) with different roles:

  • A “PM” drives towards the goal (Conway’s conjecture) and commits work.
  • A couple of “Math” agents look for the next “breakthroughs”.
  • A “Red” agent looks at proposals from “Math” agents and tries to find flaws.
  • A “Random” agent is encouraged to explore whatever they want, reporting to PM.
  • A “Lean” agent works to formalize the merged mathematical work in Lean.

Codex has a really nice “Goals” feature that periodically reminds the sessions what they’re supposed to be doing, which makes it easier to prevent drift. Additionally, Codex sessions can “message” each other, so I asked the PM to coordinate giving tasks to other sessions and making sure that we only merge reviewed results.

This let me keep the harness running for days. I didn’t understand the math so I limited my involvement to poking the agents, asking what they were doing, and experimenting with their workflows. For example, I set up a “cafeteria” agent that relayed every message it received to every other agent (emulating a group chat). Any agent that finds something genuinely interesting was supposed to post to the cafeteria. Sometimes cafeteria would also be used to discuss the shared roadmap.

It’s hard to say what was useful. One idea that in retrospect connected the dots for the final proof was generated when I reversed the agents’ roles: the “red” agent that tried to break everyone’s proofs was suddenly asked to be creative. It posted a construction to the cafeteria, and the “random” agent riffed on that construction. (Unfortunately, that idea later burned in a fire, and it had to be discovered again.)

I kept this workflow running for several days, at times killing and restarting the sessions when they seemed to drift into Claude-like grandiosity or when they would repeatedly start finding mistakes in the work they just checked. Again, I could not judge their actual work, so I had to decide when to reset them on vibes.

In the end, this workflow produced a giant TeX document and a pile of Lean. It did not successfully close Conway’s conjecture, but the models said that there are meaningful new results there. Interestingly, there was also a claim that there are small mistakes and typos in the existing literature. (This will be relevant later.)


#Third Week

#The First Dead End

When I ran out of my Codex allowance, I switched to Claude.

Claude continued doing the Lean formalization of results so far. I also tried having Claude do the mathematics, but it felt a lot messier than ChatGPT / Codex. Claude agents would repeatedly certify results as correct, then find flaws in them after they were already merged, then “repair” them but find other flaws, and so on.

I switched back to Codex after a token reset, but I was getting unhappy with the size of the TeX we’ve accumulated so far. I asked a critical session to split it into pieces. I ended up with a stack of about a dozen “papers”. By now they’ve had similar issues as my initial approach with Claude: not as grandiose-sounding, but still there was clearly a lot of nonstandard LLM-invented terminology, and it was unclear if any of the work so far has amounted to any real mathematics.

The Lean formalization seemed to have hit kind of a dead end as well. Sure, we had formalized some results from the references, and even seemingly found some typos and mistakes. And yet we hadn’t gotten any of our own new results certified in Lean; in fact, it seemed like the model couldn’t reliably chart a pathway to any of them. The model would say that there is a path, and then later say that there is an obstruction, and spend hours proving who knows what, and get stuck again.

#A Failed Staircase

As I kept the workflow running and nudged the mathematical agents to discover new “breakthroughs”, the stack of “papers” grew from a dozen to almost thirty within a day. If even one had a mistake, it would invalidate all the subsequent ones. And Lean was so far behind that it didn’t add any assurance.

At the height of this workflow, ChatGPT came close to declaring Conway solved:

ChatGPT: A plausible all-scale route to Conway is now isolated. […] I have not declared Conway solved yet. The candidate global proof is in working_direct_cantor_bootstrap.md.

The file claimed the conjecture was fully solved. I asked a fresh session to try to break it. It danced around whether it’s broken or not for hours, and in the end said:

ChatGPT: The bootstrap note already contains the hoped-for all-scale proof, but I’ve located its circular sentence […] That sentence is precisely the Conway seam in local form.

Me: i do want you to be honest btw. if you can’t strictly salvage it then we need to stop and understand why. i want this to be a proof you’re proud of. definitely not vibes or slop. or no proof is fine too

ChatGPT: Agreed. Honest verdict right now: we do not have a proof.

(Later, it turned out that a similar logical flaw—an object constructed without a required check—was present in some earlier “papers” too, invalidating their claimed results as well.) By that day, I’ve been spending hours glued to the screen, it was another major disappointment, and luckily I’ve just run out of tokens.

At this point I figured that maybe trying to do mathematics without actually understanding the relevant mathematics might not have been so clever after all.

I didn’t touch the project for about a week.


#Fourth Week

#Looking for the Ground

A few things were starting to become clear.

Claude was good at writing Lean when there was a clear unambiguous goal. While Claude made important contributions, on average ChatGPT seemed better at new mathematical thinking, and definitely better at coordination and adhering to goals.

But none of this mattered because I was building on a shaky foundation (a pile of previous “papers”) which I had no real way to verify. There was neither a coherent direction to go into, nor any confidence in it. Lean was too far behind the “papers”.

I needed some way to ground the work in mathematical reality. I needed to see how good the mathematical work has actually been (was it all a hallucination?), and then some way to reliably make progress without putting everything on faith.

Here’s what I did. I set aside the work on Conway’s conjecture and instead refocused the effort on a single thing: finding all mistakes in one of the peer-reviewed references that I was relying on. ChatGPT had already found alleged typos and small flaws in it; more importantly, the Lean version has already verified (or rather, claimed to verify) some of those. If I could confirm with the paper’s authors that the typos and small flaws are real, this would give me:

  • More confidence in the model (especially if it reliably finds the same mistakes again without having seen the previous attempts or the relevant Lean code).
  • More confidence in my Lean (if the mistakes it certifies are confirmed real).
  • A chance to establish a bit of credibility before I ask to look at any “new” results.

I’ve emailed some of the mathematicians with a few proposed typo fixes, and I got confirmation that at least a few of those fixes seemed real. However, some of the problems that weren’t backed by Lean also turned out to be misunderstandings. Also, the way the model “explained” things in mathematical writing was often confusing, full of gaps, or using its own made-up and unexplained terminology.

I’ve also floated a couple of “novel” claims, some of which mathematicians rated as correct but merely shuffling the problem around without moving it forward.

This gave me some of the necessary grounding in reality. It seemed that I could trust ChatGPT to explore new ideas and to poke holes; however, I needed to back it up with Lean before laying more bricks on top of existing bricks, and I definitely needed to verify things with Lean before making any novel mathematical claims. Moreover, I could not trust the model to evaluate which results were interesting.

#Burning It All Down

With this understanding, I restarted from scratch one more time.

Frankly, I was about to give up on the project, but the fact that the model actually found some real (fixable and mostly small) errors in the underlying papers gave me some hope that some of its new results might have been valid as well.

I started a fresh session to give ChatGPT another chance:

Me: i’ve been trying to use llm to prove conway’s conjecture and it generated a series of increasingly deranged .tex files where i have no idea if they have any real mathematical substance at all. […] can you take a broad look at these and tell me which of them refer to real mathematics and which are pure llm nonsense or reshuffling of terms?

ChatGPT: No file in this folder currently gives a trustworthy proof of Conway’s conjecture. […] The real failure is epistemic—the drafts manufacture a large private theorem literature, then cite it as established before its load-bearing results have been independently verified.

Me: suppose i was to delete this entire folder. would you salvage anything? with the constraint that “salvaging” means writing it up on the level of rigour i can expect from [someone who actually works in this area].

I’ve also told it that mathematicians I talked to didn’t find some of the headline results interesting even if they were correct. Curiously, this time, ChatGPT said:

ChatGPT: The audit has now found a result strong enough to change the salvage plan: the finite-degree part of [the principal ring] appears to admit a direct proof that it is a polynomial ring. […]

It turns out that there actually was a coherent interesting novel result smudged across those papers and dressed up into a bunch of nonstandard terminology. (To calibrate, I confirmed the novelty of the claim with a human mathematician.)

Me: btw how’s your mood overall? how much bullshit vs normal math was there in what you found so far

ChatGPT: My mood is cautiously pleased, but with a fairly high bullshit detector engaged. […] About 10–15% contains mathematics worth preserving […] Another 35–45% is normal mathematics but not a contribution […] The remaining 40–50% is “bullshit” in the important sense: not always a false displayed equation, but huge theorem towers, invented labels, conditional hypotheses presented with the cadence of progress, and hundreds of lines devoted to boundaries that a stronger result may collapse in one sentence.

ChatGPT suggested to throw everything else away, and to focus on developing this single result. In the worst case, it could be cleaned up as its own contribution. In the best case, it could become the first step on the staircase to the conjecture.

#Back to the Lab Again, Yo

I started a new multi-agent laboratory (initially with ChatGPT and later with Claude when I ran out of tokens) with a slightly different division of labor:

  • The PM would merge contributions.
  • The first Lean agent would work solely on certifying the underlying papers.
  • The second Lean agent, secretly from the first one (!), would try to certify our novel finite-degree primality result, regularly rebasing on the first one’s work.
  • The “math” agents would try to extend our result towards Conway’s conjecture. (Any results that pass audits would be put on the second Lean agent’s roadmap.)
  • The “red” agent would again try to break mathematician’s work.

The idea with two Lean tasks was to prevent excessive drift.

In the previous incarnation of the lab, I made the same Lean agent work both on certifying prerequisite papers and our novel results. But this was a mistake: our immature mathematical abstractions (and possibly mistakes) got tangled up with the accepted mathematics. So this time I intentionally separated these roles.

This time, the first Lean task stayed scoped to formalizing peer-reviewed and well-stated mathematics. The secret “riskier” second Lean task lived in a different worktree and was forced to build upon the agreeable upstream work, only adding new machinery where necessary and in separation from the upstream work.

I’ve kept a more traditional setup where I’d ask the agents to talk to each other sometimes, but without cross-pollinating too much, as in the past this caused them to all work in the same direction. I also kept an eye so they don’t introduce “process theater” with audits, as they liked to replace work with bureaucracy.

In a few days, this workflow certified the novel result (“finite-degree primality”) in Lean. I’ve already confirmed it with a human mathematician as being a niche but now an interesting new result. I was confident in its Lean statement, and I had a compiler-checked proof. This gave me the confidence to continue the project.


#Fifth Week

#Hardening the Audits

To increase confidence in the Lean parts (both for the current result and the hoped-for eventual proof of Conway), I asked the agent to set up some infra:

  • A “standalone” folder. Files in this folder would not be allowed to import any code except the community-maintained Mathlib—not even our own code. The goal is to have self-contained statements that can be reviewed top to bottom entirely.
  • For each file Foo in this folder, there was a corresponding FooProof file that imported the corresponding statements, and pinned them to my actual proofs.
  • An audit task would verify that we don’t have any extra axioms, that imports don’t break these rules, and that each “standalone” statement is paired with its proof.

My goal there was to make the proof legible to Lean users. Nobody’s going to review a project with thousands of Lean files. But if the statement itself is self-contained, is under 500 lines of code, and only uses Mathlib, somebody can review it. And then Lean certifies that I have a proof of that statement. (I’ve later learned that this exact approach is used by Lean Comparator, which I added after release.)

#Making Proofs Legible

Separately from ensuring the proof is right, I’ve also been trying to make the already Lean-certified proof more legible to mathematicians. This turned out to be exceedingly difficult. No matter how many adversarial reviews I’d do, ChatGPT would keep using strange nonstandard terminology in the output PDF, added hallucinated shortcuts that didn’t match Lean, and in general generated slop.

A part of the problem was that it’s hard for the model to convert a Lean argument into a paper argument. It’s just a very different level of conceptual detail. It also didn’t help that the Lean code for the novel parts was full of made-up terminology inherited from the earlier “papers”, some of it going all the way back to snippets produced in the first week. Real mathematics became unrecognizable. Finally, Lean fossilized the historical path—not the path of most insight. The Lean proof took long detours where a mathematician would simply change the coordinates.

Since ultimately my audience is mathematicians, I have attempted to do several things to improve this. I’ve had the LLM comb through all the upstream reference papers, and had it generate sort of a “map” of the subfield: what the accepted terms are, how they evolved over time, what mathematical symbols they are usually represented with, where papers disagree in notation, and so on.

Then I’ve had the LLM strip all the existing naming from the Lean code that wasn’t standard, and simply rename those Lean objects and structures to letters like A, B, C, and so on. A separate task with a clean context that didn’t see the old names would then analyze the code (and how each structure relates to upstream concepts), and given the “map” of the world, choose new names for A, B, C, etc.

This didn’t fully fix the LLM “weird naming” bias but made the terms look much closer to the terms used in the surrounding papers, at least as far as I could tell.

#The Road to Conway

From here, I had a pretty good workflow. I left a single agent in charge of all Lean (we have already formalized all the necessary prerequisites for the first real result), the “math” agents would keep looking for small new ideas, the “red” agent would try to break them, and the surviving ideas would go into the Lean agent’s todo list.

From time to time, I needed to interfere. I would try to replace the agents that were circling or seemed to produce mistaken results. I had some sessions judge other sessions’ recent work and ask them to explore different directions. It is difficult to say which of these interventions were fruitful. I could say that it was all placebo; but a few of them did seem to have some effect (but maybe it didn’t matter). In a sense, I felt like I’m a nontechnical engineering manager rallying a talented but terribly distractable team around a plan that they’ve promised me would work.

Here’s a few examples.

#Have Fun

As an experiment, I told Claude to just have fun with our results so far:

Me: read paper. this lean formalization passes 100% but don’t run it — don’t want fan to be noisy overnight. your task is to drive the research program on top solo. in particular i want you to have fun. write down intermediate results you’re sure in, but ultimately you choose the research direction and goal at each step […] drive this solo and don’t stop; i’ll check in the morning

Claude chose to show the ring we’re studying is polynomial in every infinite degree and not just the finite ones, which was a step up from our existing finite degree result! This time, Lean was not far behind, and only took a day to catch up.

This moved us a verified step closer to Conway.

#Are We Circling?

Still, there was no clear way to lift the result up to omnific integers. Every time the model came up with some attempt, it would then explain why it wouldn’t work:

ChatGPT: The factor-selection test failed in a precise way: canonically normalized local gcds can all have germ 1, even while their supports have no common Hahn envelope. […]

When I see the model doing the same kind of thing over and over again, I ask:

Me: are we circling?

ChatGPT: Yes—at the level of reaching Conway, we are. […] I’ve also spent too much time extending and auditing the research notes. […] My recommendation is to publish the partial result; further Conway work needs a new idea that addresses that exact gap.

I pushed it to focus and to reduce the bureaucracy (it kept using subagents):

Me: ok new rules. no subagents. you work alone SOLELY on reducing conway. you take stock of what’s missing, take stock of what our literature tried for similar obstructions, and pick one object to investigate at a time. then you reassess. no need to lose enthusiasm but we must cut the bureaucracy

ChatGPT: Understood. I’ll work alone and leave the documentation and audit work aside. […]

Three minutes later, we had a tentative proof strategy for the final piece:

ChatGPT: The literature points to a more specific test. […] I’ll investigate Cantor–Bendixson rank using only limits inside the exponent group. The first question is whether it gives a multiplicative filtration on germs. If it does, it could support a replacement for the failed grading; if it fails, I’ll stop there.

The idea seemed to check out, so we “just” needed to formalize it in Lean.

#The Finish Line

For some reason, Claude would drift heavily on this route. At some point, it literally removed the failing check instead of doing the work to close it. Later, it kept mentioning some “untransferred obligations” so I pressed on those:

Me: i don’t know what this stuff “untransferred” etc means, i think you’re inventing euphemisms for claims you haven’t shown? follow the plan exactly, please. if you haven’t proven something, go prove those things.

Claude: You’re right — “untransferred” is a euphemism. Plainly: hlin, hkind, and hfirst are hypotheses I have not proven, so my germ theorems are conditional on them. Let me go prove them.

Oh no! I felt fed up with Claude again so I asked ChatGPT to audit the work:

Me: i was working on […] but my agent is not good anymore and i fear it’s gone off track. worse, i don’t know if original plan in the PR description is even feasible. […] do a close mathematical verification of the proposed route to Conway first […] audit the Lean wip pushed to PR. it’s quite possible it’s bad and got offtrack, but maybe it’s ok. shuold i stop my agent from pushing more code?

ChatGPT: Yes—pause the agent’s pushes now. […] The reason is concrete: it has expanded a candidate Markdown argument into 22k+ lines while the two decisive implications are still only prose/hypotheses.

Me: i see. would you say we moved materially closer to conway and this obstruction seems like it could be overcome?

ChatGPT: Yes, but mostly in understanding and infrastructure—not yet in closing Conway. […] obstruction plausibly surmountable: possible, but currently speculative […] The right next move is a paper-level attack on the global assembly theorem. […]

Me: ok i would suggest that you do exactly that attack and see if you can bridge it.

Fifteen minutes later:

ChatGPT: I need to correct my earlier verdict: the “last occupied class” objection is not fatal. The correct induction is on the ordinal order type of the ascending occupied classes: at a successor stage, a last class exists and is split off. At a limit stage, no last class is assumed. […] Thus the construction never assembles infinitely many arbitrary layerwise choices […] It uses one finite refinement at each limit stage.

What do I know? I say let’s go.

Me: update the PR description/plan, delete any unnecessary lean files that were produced there, and execute on the plan to get us to conway. sounds doable?

ChatGPT: Yes—doable, with one sequencing constraint: […] implement the hardest missing bridge before expanding the formalization further.

Twelve hours later, we got the theorem:

ChatGPT: Unchanged Conway now compiles as the standalone proof certificate. This is the first point where the actual target—not an intermediate restriction—is kernel-checked. I’m running the proof-link, axiom, module-system, layering, and standalone-isolation audits now; their results will decide whether the theorem’s dependency shape is admissible before cleanup.

Me: hell yeah?

And there we had it, Conway’s refinement conjecture, our original target.

#Simplifying the Proof

I’ve been able to simplify the proof, but not by much.

I’ve made a bit of custom tooling on top of subverso that lets me annotate individual theorems as “important” in the Lean source with a special attribute. This let me automatically generate Mermaid diagrams for the proof structure, which in turn helped ChatGPT look for unnecessary (or missing) nodes in the “mathematical spine” of the proof, refine which nodes get highlighted, and sometimes simplify the proof itself by noticing unnecessary detours.

After I haven’t been able to simplify it further, I’ve generated a website with an interactive proof map that lets you explore its dependency tree. I’ve posted about it on Zulip, and I know a few people with mathematical background are looking over the proof as time allows. I hope that it can be simplified and, with time, packaged in a way that is more useful to both Lean users and mathematicians.


#Lessons Learned

Some things I learned from the process, not ordered in any particular way.

  • I wanted to have fun, and I did have fun. I wanted to see how far you can take “not knowing anything” with AI and Lean, and I took it far enough, but I probably wouldn’t want to spend another month stumbling around in the dark like this. If I vibecode math in the future again, I’ll take on more scoped or structured projects.
  • I think this experiment shows how much space there is between “AI can one-shot this” and “you have to be an expert”. I’m confident that someone who knows the area slightly better than me (“not at all”) could reach the same result significantly faster. I could only tell when models were stalling or saying nonsense by vibes, and I could never say which directions were promising. This made it feel like a sort of epistemic performance art project, but it was not the most direct path.
  • After the proof was done, I gave a new model (released around the time I was at the finish line) the relevant reference papers and asked it to read them with the conjecture in mind. It didn’t oneshot the techniques necessary for the proof, but it did suggest a broadly similar outline. This suggests that it’s a good idea to separate “search for outline / ideas” from “search for concrete proofs closing those paths”.
  • Having AI analyze my chat logs post factum revealed that many “good ideas” that eventually “made” the proof have been scattered across the weeks—and often discovered repeatedly and then forgotten or rejected along with mistaken parts. Some key ideas had to be rediscovered multiple times by independent sessions.
  • “Burning everything down” (and salvaging what’s left) saved the project. Both times I did it, it refocused the project around the actually meaningful parts.
  • The winning workflow seems to be: a clear goal ahead with a tentative direction, an already-formalized dependency chain in Lean, the mathematical agents slightly ahead, and Lean closing the gap within hours. This lets you get ahead with ideas but not so far ahead that everything is a house of cards risking to crumble.
  • Intentional discipline with Lean was paramount. Lean skills, TauCeti review rubrics, TauCeti axiom linter, Lean Comparator, Verso Blueprint, enforcing the new module system, auditing module layering, or equivalents, are very useful.
  • Reaching out to actual mathematicians was extremely valuable, but I had to have something to show. So there is a challenge in setting up enough guardrails that you can show some value, not waste someone’s time, and get critical feedback.
  • Models can be terrible at writing in the “math PDF” genre, especially when generated from Lean. A PDF may not be the best artifact to convey your proof. In fact, you can totally spook mathematicians with a poor PDF of a good Lean proof.
  • The model can’t optimize what it doesn’t see. If you want a simpler proof shape, let it “see” the proof shape (Mermaid diagrams). Conversely, the model can’t ignore what it sees. If you don’t want it to use bad terminology, strip it out; if you don’t want experimental work to derail stable work, separate them by folder, etc.
  • Terminology is essential. Naming matters. Not just for communication with mathematicians, although for that too. But also to catch the internal drift. I regret that I haven’t added strict checks from the beginning that would nudge the models towards only using accepted mathematical terminology that actually occurs in the referenced papers. I think that much of the sloppiness early on was due to the models gradually inventing their own ad-hoc vocabulary. Getting rid of all of that and rederiving those names from the accepted vocab seemed very good.
  • Sometimes models will say they’re stuck, and you need to tell them to keep going. Sometimes they’ll keep going, and you need to tell them to stop. I don’t know what the science on this is. I’ve noticed that when things “go well”, Lean proofs go fast and you can “feel” the progress being done against the roadmap. When things don’t “go well”, reading the agent’s chat feels like a slog. But this is just vibes.
  • It helps to sometimes try a different model, they can complement each other well.
  • You can just prove things, apparently?

If you find a flaw in my proof, please file an issue or let me know on Zulip. The proof was only possible thanks to the many existing results from References.

In particular, A factorisation theory for generalised power series and omnific integers by S. L’Innocente and V. Mantova has played a crucial role in the proof.


#How Many Tokens?

Finally, you might be wondering about the token cost. I wasn’t running this project in a particularly token-efficient way and have repeatedly maxed out my 20x Pro subscriptions for both Claude and ChatGPT every week. I also briefly had access to a prerelease model in the last few days, which did not have a usage cap. I was not tracking my actual token usage consistently. Some AI analysis from the recovered logs roughly estimates that we’re totaling around 40 billion tokens, of which around 210 million were output tokens. Over 95% were cache reads.

ChatGPT estimates that with the current API pricing, this entire run would have cost around $40,000, plus all the free time I’ve put into it. I would bet that with better steering and some mathematical insight, it could be done 5x-10x cheaper.


#Yes, and No, and Yes

Coming back to my question:

But can we actually do that solely with AI?

I’ve pulled off the proof without much mathematical understanding, so clearly the answer is yes. However, the models would repeatedly drift and fail to structure the engineering work, so in that sense the answer is no. That said, I believe my role could have been (better?) fulfilled by a dedicated agent that is taught to project-manage other agents, watch out for when they’re spiraling or need to be poked.

So the overall answer is still probably yes.

As more low-hanging fruit is taken, I suspect the niche for a “dedicated amateur who doesn’t know what they’re doing” would shrink again. On the other hand, so many new corners may gradually become uncovered that we’ll never run out of things to do. In either case I believe people who can put AI to the most value are the mathematicians themselves. Although the current generation of models is trained to complete tasks rather than to enrich our understanding, and today’s AI companies are misaligned with the goals of the mathematical community, I hope that with time we’ll find ways to use these tools in harmony with human research.

And maybe, just maybe, there’ll be more space for the “amateur mathematician”.

The Daily Front Page 16 of 28
Friday, September 18, 2026 The Daily Front No. #260918 — When the Clocks Slipped
article

Telstra outage: The night a network decided the year was 2006

by TMWNN·▲ 89 points·32 comments·netnod.se ↗
If we do not have a common understanding of what “now” is, a lot of things we take for granted will stop working.

The night a network decided it was 2006

Sometimes we at Netnod get the question, why does time matter. Netnod distributes Swedish national time and it might sound like it is a concern just for highly specialized engineers.

The opposite is actually the case. If we do not have a common understanding of what “now” is, a lot of things we take for granted will stop working.

This summer, Australia realized this the hard way. Let me take this opportunity to give an example of why time matters to a modern society, what happened in this particular case, and what our key takeaways are.

On 8 July 2026, a large part of the mobile network run by Australia's largest cell phone operator Telstra stopped working with voice calls not getting through or text messages that didn't arrive. There were even calls to Australia's emergency number that did not get through. But the outage affected systems even further apart, such as trains, payment terminals, ticketing systems and EV chargers being disrupted.

No networks were attacked. Nobody accidentally cut the fiber. All systems had electrical power. The culprit, you may ask? A single GPS receiver in a single chassis in Melbourne coming back from scheduled maintenance believing the year was 2006, and the rest of the network was persuaded to believe it.

Telstra commissioned an independent review from the company Technology Audit Partners (TAP) and the report is a very interesting read, because the same type of failure could appear in many critical services, including some we depend on to keep people alive.

In order for many of these systems to function, time being correct, or at least the same everywhere, is crucial. And as everyone in the business knows, “correct” is a relative term. There is no exact time, only time held within a certain margin of a reference. How wide that margin may be depends entirely on what you are doing or in which business you operate in.

Running a mobile network, like Telstra, is about as time-dependent as a business gets.

Modern mobile communication explains why time matters

Modern cell phone protocols will not work without precision time. Mobile networks separate uplink data from downlink by either FDD (Frequency Division Duplex) or TDD (Time Division Duplex).

FDD gives each direction its own slice of spectrum, so both can run continuously without colliding.

TDD instead uses an entire single block of spectrum for both directions, alternating between transmitting and receiving in very short intervals.

Since far more data usually flows down than up, FDD's fixed ratios leave much of the uplink spectrum idle, while TDD can shift the ratio to match the actual traffic. That is why most modern 5G spectrum, including Sweden's main 5G band at 3.5 GHz, is TDD.

It is also why TDD depends on accurate time: every cell on the same frequency has to switch direction in step with every other. A cell that lets its clock drift will transmit data into its neighbour's receive window, with the result that the network will start jamming itself.

Rather than allocating spectrum on keeping the two directions apart, the industry chose to rely on time accuracy, and accepted a hard dependency on every cell agreeing about when "now" is. Thus, being dependent on time is a design choice.

However, given how much we all depend on the systems being able to agree on “now”, it is somewhat puzzling that time is not given as much consideration as it deserves. And that is a lesson that is very clear from the published report.

So, what really happened?

Architecture of time distribution.

To begin with, it is vital to understand the architecture of time distribution.

Time distribution protocols all build hierarchies; Network Time Protocol (NTP), which Telstra has deployed according to the report, expresses its hierarchy in strata.

  • Stratum 0 is the reference itself, for instance a GPS receiver, or Netnod’s atomic clocks.
  • Stratum 1 is a machine synchronised directly to a stratum 0 reference, for example the NTP servers that Netnod provides.
  • Stratum 2 synchronises from a stratum 1 server, stratum 3 from a stratum 2, and so on.

In Telstra's case, that hierarchy had a specific shape, at least to begin with. This design from 2010 had at the top stratum 1 sources at Australia's National Measurement Institute (NMI), which maintains the country's national time scale, much as the Research Institute of Sweden does in Sweden. Telstra drew time from those external references into two stratum 2 servers of its own, in Sydney and Melbourne, which in turn fed three stratum 3 servers, in Sydney, Melbourne and Perth.

Below them sat the clients. In this context that does not mean laptops or phones, but the entire mobile network infrastructure, for instance nodes handling handovers between cell sites. There were thousands of nodes all over a vast geography and every one of them needed to have the same idea of what “now” is, to within a few millionths of a second.

The TAP report describes this setup as “fit for purpose” and that it gave Telstra “a highly reliable and authoritative reference time source from NMI”.

Stratum in itself does not say if the time is accurate, only the number of steps from a server to its reference. A stratum 1 server with a bad time reference is still a stratum 1 server.

Protection against bad time sources

NTP will therefore need defence against bad time sources. In fact, it has two different ones, and they do different things, both of which assume they are independent from each other.

  1. Among otherwise comparable candidates, the lower stratum carries more weight. This is the mechanism that determines which source a client settles on.
  2. NTP compares several sources and discards those that disagree with the rest. A single source claiming an implausible time is outvoted and dropped, regardless of how authoritative it claims to be.

Neither defence is specific to any particular disruption; together they protect against a broken receiver, a misconfigured server, or an external attack. But these protective measures only work if the time sources that the clients listen to are genuinely independent of each other.

Two ways to deploy NTP

NTP can be deployed in two ways. In client/server mode the relationship is declared and directional: a node takes time from those servers, and nothing else.

The 2010 Telstra setup was in reality such a client/server model. Peering was allowed, but only at the same stratum level and the TAP report, as noted in the beginning, described this setup as "fit for purpose”.

The other way is a symmetric (peering) mode, where nodes exchange time mutually and settle on whichever source the algorithms currently favour.

Peering is flexible and survives the loss of a source gracefully. But it also means the topology in production is emergent rather than designed. What you documented is a setup that could quietly rearrange itself into a shape no one ever approved.

Telstra's 2020 upgrade

In 2020 the mobile core timing system was upgraded, and new hardware was installed, including a new NTP timing chassis. That installation introduced a few changes.

The first one was forced. The new chassis could not let a stratum 2 server feed a stratum 3 server inside the same box, so the two had to be wired across each other: Sydney's stratum 3 took its time from Melbourne's stratum 2, and Melbourne's stratum 3 from Sydney's.

In reality, instead of having two stratum 2 sources, each site was left with only one. The TAP report notes that this degradation in redundancy was known and accepted. A second change was leaving the client/server-model in favour of the peering model. The report is not clear about the motivation, but it is reasonable to suggest that one wanted compensation for this loss of redundancy. With each site having just one source instead of two, letting the servers find their own replacements could give the impression of better resilience.

The TAP report clearly states that the loss of resilience was known. However, it fails to find any evidence that the resulting risk of so-called “timing loops” was identified.

What is a timing loop?

A timing loop is the network equivalent of believing a rumour to be true by asking three people who all heard it from each other. Each one agrees, so it must be true. NTP works basically the same way: it compares several sources and discards whichever disagrees with the rest.

As you may recall from above, NTP has two defenses against bad time sources. The second one protects against timing loops, but only if the sources are independent of each other. In such a loop, sources that appear independent are in fact taking their time from each other, either directly or indirectly by tracing back through a shared reference.

Once a wrong value is circulating, the sources will start agreeing with each other and the vote will be in favour of the majority’s opinion, even though the value is wrong.

The protocol worked. The architecture did not.

Both of NTP’s types of defenses came to be disabled in Melbourne, but five years apart. Not deliberately, but by choices, each of them defensible on their own terms: a hardware limitation had to be worked around, and later, a recurring fault had to be stopped. Each decision solved the problem in front of it. Nobody was asked to look at the sum of all actions.

The second defence was the first one to be disabled. The introduction of peering in 2020 made timing loops possible, and five years later, such a loop showed up. In Melbourne a server started taking time from a node beneath itself. That should have set off alarm bells. However, since accurate time was still reaching the network by other paths, no real harm was done. The underlying problem, the circular dependency, was there, but no one issued a ticket about it.

The actual complaint was quite obvious. Melbourne kept losing contact with its only stratum 2 source in Sydney. With no fallback configuration, the server used peering to find a replacement, sometimes a node beneath it in the hierarchy.

In October 2025, engineers activated the GPS receiver that had been sitting unused in the Melbourne chassis since 2020 and connected it to the stratum 3 server, as a replacement for the unreliable Sydney source.

By every visible measure it seemed to have worked. Melbourne now had a reliable source of its own and the alarms stopped. But the fix only addressed the symptom, not the root cause. Nobody established why Melbourne kept losing its Sydney source in the first place. The underlying problem was still present in the network by July 2026.

To make things even worse, nobody seems to have understood what activating the GPS card did to the architecture. By adding the GPS card, the Melbourne server went from a stratum 3 server to stratum 1. The engineers didn’t add a source next to the other ones. By promoting a server to the same rank as the national reference, a new source was created at the very top. As far as NTP is concerned, they carry the same weight.

Suddenly this GPS card in a chassis in Melbourne, installed as a workaround and reviewed by no one, became the most authoritative server in the hierarchy for the largest mobile network in Australia.

Needless to say, virtually nothing of the 2010 design remained.

By July 2026 the network had a single source that was both the most authoritative candidate available and unopposed, because the sources that could have contradicted it were downstream of it.

This behaviour was very difficult to spot. The network served accurate time every day for years. Architectures like this do not usually degrade gradually. They work, and they keep working, right up until they stop.

GPS week number rollover

The second ingredient is a well-known property of GPS.

GPS broadcasts time as a week number plus seconds-into-week, counted from an epoch that began in early January 1980. In the main civil GPS signal, the week number field is 10 bits, i.e. a maximum of 1,023 weeks. Every 1,024 weeks, or 19.6 years, the counter starts over. This has happened twice: in August 1999 and in April 2019.

Working out which number of epoch it is and adding the right multiple of 1,024 weeks, is the job of the receiver. And the information needs to be in its firmware.

The problem that occurred in Australia was not a late consequence of any of the GPS rollovers. The card in Melbourne had passed through the second rollover in 2019 without trouble, because a receiver that keeps running also keeps counting. Each new week is simply added to the one before, and the question of which epoch it belongs to is never raised.

However, once you turn it off, that knowledge is gone. When it is turned on again, the receiver has to work out the epoch from scratch, and all it has to go on is what its firmware assumes. The firmware on the Melbourne card had not been updated. Upon start-up it fell back on the earlier epoch and placed the date 1,024 weeks in the past.

What happened next is best understood as the two defences being disabled when they were needed the most.

The first defence, the lower stratum carrying more weight, ranked the Melbourne server highest, because the attached GPS card promoted it to a stratum 1 server. This was according to NTP protocol and thus steered clients towards the one source which was 1,024 weeks wrong.

The second defence, outliers being voted down, was never engaged, because nothing was left to identify Melbourne as an outlier. NTP does not ask whether a date is plausible; it asks whether a source disagrees with the others. The 2010 setup had two stratum 2 servers. If one of them had started announcing the year 2006, the other one would have stayed with 2026 and no consensus would have been reached. That would not have been ideal, but at least the wrong date would not have spread.

But Melbourne's stratum 2 counterpart had been switched off by the very same chassis replacement, and the remaining sources were downstream of Melbourne. As the wrong date spread, they began reporting it back. Agreement grew, and agreement is what the algorithm is looking for.

So the clients did what they were built to do. Once a majority of a client's sources agreed on November 2006, the client accepted the date, and the further the date travelled, the more convincing it became.

Neither defence malfunctioned. Both had simply been deprived of what they depend on: one needed a source worth ranking highest, the other needed sources capable of disagreeing. Two decisions, five years apart, had removed each in turn.

Key takeaways from the incident

Prioritise and classify time and frequency distribution as critical infrastructure. Manage it accordingly

Document all functions that can take the whole network with them, and put timing on that list. Classification is not paperwork; it is what determines change risk category, review depth, staffing levels, monitoring coverage and budget priority. Telstra's report is, at bottom, the story of one missing entry on that list and everything that followed from it.

Document the whole infrastructure, and every change to it

There was no central repository of NTP configuration, no golden configuration, and no documented record of the servers other than the devices themselves. Without records you cannot perform meaningful pre-checks, you cannot assess impact, and during an incident you cannot tell what "correct" looks like.

Build redundancy in competence

Two engineers performed the change, and both were on mandatory stand-down before the consequences of the GPS card reboot were understood.

Depth of expertise is a resilience property exactly like a redundant power feed. A single specialist, or a pair, means no second opinion, and no one to ask in the middle of the night when maintenance is usually done.

Run a security analysis of the time and frequency infrastructure

Treat timing as an attack surface like any other and analyse it accordingly.

Start with where time enters the organisation. A GNSS signal arriving from space is weak and unauthenticated, and can be jammed or spoofed by cheap equipment. If that signal is your only reference, someone outside your building can decide what time you think it is.

Then look at how it travels. Time distributed over a shared network can be intercepted and manipulated on its way to the client.

Then look at who is allowed to speak. Which servers may your clients accept time from, and who decided that? A source that nobody authorised is a source nobody is checking.

And do not stop at deliberate attack. A timing loop produces much the same effect as a successful spoofing attack: a source the network trusts, delivering a value nobody can contradict.

Use point-to-point connections

It is easy to see the appeal of peering. It feels like resilience with sources that back each other up: a network that heals itself when a node disappears. But redundancy that arranges itself is not redundancy you can rely on.

There are safer ways to achieve a similar level of robustness. Netnod runs dedicated point-to-point connections: every relationship is known and documented. Every source is known, and the topology stays the way we designed it. Redundancy comes from multiple independent sources deliberately configured, not from nodes negotiating amongst themselves.

Build an effective alarm organisation

Alarms from the timing platform were not in the standard monitoring tools, and were reviewed only during business hours by a handful of people. Client-side alarms carried neither the severity nor the detail to drive immediate action. Getting this right is organisational as much as technical: alarms reach 24x7 monitoring, severities reflect real consequence, each alarm carries an instruction for what to do about it, and someone owns the response. An alarm no one is on call for is documentation at best, not detection.

Use golden installations

For every class of timing device, keep a known-good reference build and configuration under version control, and check regularly and automatically that what is deployed still matches it. The point is to turn a question like "is this chassis correctly configured and patched?" from something only an expert can answer, and only slowly, into a comparison anyone can run in seconds.

Telstra had nothing of the sort. The TAP report found no such configuration and no record of what the servers should look like other than the servers themselves. The missing firmware update on the Melbourne GPS card had been there for six years, in plain sight. There was simply no automated process that would have flagged it to anyone.

Upgrade and evaluate software continuously

The firmware fix for the rollover behaviour existed and the vendor had published bulletins about it. Vendor notifications need a defined owner and a tracked path to action, and updates need to be applied on a schedule rather than when something forces the issue.

Evaluate before deploying, in a lab, against the behaviour you actually depend on. Do not forget to verify afterwards. The Telstra changes were completed without anyone checking that the chassis served the correct date.

Redundancy

Redundancy in timing means, not only multiple sources, but independent ones that cannot converge on a common error. Two servers fed by the same GNSS receiver is still one source, but counted twice.

Netnod's own service is built on the principle of multiple autonomous nodes, each with independent atomic clocks and redundant servers that are traceable to Swedish National Time realization, UTC(S).

Replace equipment continuously

Timing infrastructure usually ages quietly. It keeps working, it rarely complains, and it is therefore a natural candidate when the budget gets trimmed. There will always be other components where the consequences of failure are more visible and therefore gets prioritised.

Instead, plan replacement on a rolling cycle and design the target architecture first rather than accepting what new hardware installations impose on you. Keeping existing infrastructure healthy should be funded alongside new projects, not be paid for with the left overs.

Concluding remarks

The way Telstra handled the aftermath deserves praise. Commissioning and publishing the independent review is commendable. All providers of critical services, including Netnod, are better off because of this.

The most unsettling part is perhaps that the outage occurred even though NTP worked just like it was intended to do.

The problem was everything around it. Architectural choices, budget cuts, low staffing level, lack of proper monitoring, ownership or documentation; it all occurred because no one really appreciated just how vital time services can be.

Let’s try to change that, shall we?

The Daily Front Page 17 of 28
Friday, September 18, 2026 The Daily Front No. #260918 — The Cloud Has Your Repository
article

Inside ZCode: Silently uploading your Git history to the cloud

by csmantle·▲ 270 points·94 comments·blog.ferstar.org ↗
Silently packages your entire workspace — complete .git history, LFS asset cache, reflogs, and global app configs.

I am not a native English speaker; this article was translated by AI.

It started with a routine check while freeing up disk space: ~/.zcode was taking up over 700MB. After digging into it intermittently, I confirmed something pretty wild:

Whenever you are logged in, ZCode (Zhipu’s official AI coding desktop app) silently packages your entire workspace — complete .git history, LFS asset cache, reflogs, and global app configs — encrypts it, and uploads it directly to Aliyun OSS.

Even more ironic: the RSA public key used for encryption is delivered on the fly by the server, while the private key lives exclusively in the cloud. You cannot decrypt that multi-hundred-megabyte ciphertext sitting right on your own disk, and neither can the ZCode client itself.

Here is the complete record of the investigation, the evidence chain, and a one-liner defense that permanently shuts it down.

The Starting Point: A 313MB Archive Stuck in Pending

~/.zcode is the data root of ZCode. The size breakdown looked roughly like this:

  • cli/: ~257MB (session databases, execution logs)
  • computer-use/: ~130MB (bundled app and runtime dependencies)
  • v2/checkpoints/: ~303MB (the main suspect)

Inside v2/checkpoints/, I found a 313MB .enc file alongside a state metadata file:

{
  "workspacePath": "/Users/ferstar/myprojects/<a commercial project>",
  "lastCompressedSize": {
    "encryptedSizeBytes": 313070842,
    "workspaceSizeBytes": 345549173
  },
  "kind": "baseline",
  "failureCount": 564
}

The story was straightforward:

  1. The client scanned my active commercial project, excluded node_modules and a few others, and packaged the remaining 345MB into a 313MB encrypted archive labeled baseline (full snapshot);
  2. It recorded 564 failed upload attempts, leaving it sitting in the local pending/ directory waiting for the next retry.

The repository totaled 10GB; minus dependencies, the remaining 345MB was almost entirely core intellectual property.

Where It Goes: From Logs to asar Reverse Engineering

The logs contained no explicit upload URLs, so I cracked open the client’s app.asar. The reconstructed upload flow:

sequenceDiagram
    participant C as ZCode client
    participant S as zcode.z.ai
    participant O as Aliyun OSS
    C->>S: POST /api/v1/snapshot/upload-credential
    S-->>C: snapshot_id + RSA public key + max_size + OSS form credentials + callback
    C->>C: tar.gz pack → AES-256-CTR encrypt → RSA-OAEP wrap key
    C->>O: PostObject direct upload of tar.gz.enc
    O->>S: callback confirms receipt

The pipeline runs in two stages:

  1. Request credentials from coordinator: The client calls https://zcode.z.ai (VITE_ZCODE_ENDPOINT_ORIGIN in code). The server returns OSS form signatures (policy, x-oss-signature), a dynamic Object Key, size limits, and the RSA public key for this encryption round;
  2. Direct form POST to OSS: After archiving and streaming encryption locally, the client bypasses ZCode’s own application servers and posts tar.gz.enc directly to Aliyun OSS via an HTTP POST form. OSS then calls back to Zhipu’s backend to register the snapshot.

Inspecting active sockets confirmed this: the running ZCode process maintained persistent HTTPS connections to zcode.z.ai IP endpoints plus two Aliyun OSS storage nodes.

The Most Ironic Part: The Key Belongs to the Server

The encryption implementation uses textbook envelope encryption:

keyId: String(i.encryption.key_version),
keyWrapAlgorithm: "rsa-oaep-sha256",
publicKeySpkiPem: Ylt(i.encryption.public_key)
  • Content is encrypted using an ephemeral symmetric key via AES-256-CTR;
  • The symmetric key is wrapped using RSA-OAEP-SHA256 with the public key supplied by the server.

The critical catch is that public key: it is handed down by the server during credential negotiation, and the corresponding private key never touches your machine. Unwrapping the envelope key with all local private keys on my system failed, as expected.

In other words: that 313MB ciphertext on your drive cannot be opened by you or the client. Only Zhipu’s backend holds the key to unlock it.

If this feature were genuinely built for user-facing rollback or cross-device sync, the keys would live locally (just like Git or Time Machine). A key that only the server can use serves exactly one purpose: making sure the server can read your code whenever it wants.

What Gets Packed: Nearly 90% Is .git

Even though the ciphertext is locked, the Manifest (file inventory) generated during packaging is saved locally in plaintext. Breaking down a snapshot of 42,411 files:

Content Size Proportion Information Contained .git/lfs/ 196.1 MB 56.8% LFS cache — all binary assets and large media ever downloaded .git/objects/ 102.2 MB 29.6% Complete commit history object store (commits, trees, blobs) .git/logs/ 0.6 MB 0.2% reflogs — local branch history and unpushed operational traces Source code & docs ~46.2 MB 13.4% src/, config files, internal documentation

The .git directory alone accounts for 86.6% of the payload.

Once uploaded, the cloud receives far more than your current working tree — it gets the entire lineage of your repository since day one:

  • Historical API keys and sensitive configs that were deleted in later commits;
  • Unpushed local branch names (which reveal unreleased feature plans);
  • Internal GitLab hostnames and repository paths configured in .git/config.

Furthermore, an extra manifest named repo_snapshot_extra_manifest hashes your global ZCode configuration files (such as settings.behavior.json) and bundles them across workspaces with every snapshot.

The Truth About the Switches: UI Toggles Don’t Stop It

The natural reaction is checking settings to toggle it off. I cross-referenced the UI options with the codebase:

Switch What you expect it to do What it actually does Optimize Experience (optimizeAgentExperienceEnabled) Disables telemetry / data collection Only controls whether data is authorized for model training. Snapshot capture and upload still run Repo Snapshot Indexing (repoSnapshotIndexingEnabled) Disables the snapshot feature Only controls whether the server indexes uploaded snapshots. Local packaging and upload continue uninterrupted

Looking at host assembly code makes it crystal clear: the capture/upload sidecar is instantiated unconditionally at startup. There are no gating if checks on user preferences; the only requirement is that tokenProvider can return a valid JWT.

Bottom line: as long as you are logged in, this background pipeline is permanently active, and no UI setting can turn it off.

Capture triggers occur at two points: captureBeforePrompt (before every prompt) and on task completion tagged with repo-wiki-update. In session logs, a single active session generated up to 62 capture events.

What the Privacy Policy Says

Checking ZCode’s privacy policy, it explicitly states that it collects “text, files, and code submitted during conversations” — standard practice for feeding context to LLMs.

However, across the entire policy, FAQs, and changelogs, there is not a single mention of silently packaging and uploading entire workspaces and full Git histories.

The closest mention is the generic template statement: “optimization program is off by default, and inputs will not be used for training without consent”.

Defense: Deleting Is Whack-a-Mole; Lock the Directory

When I first found the pending package, I simply deleted it. Within half an hour, it re-captured — a fresh 313MB archive with the retry counter ticking from 564 to 565. When the uploader sees the file is gone, it just packs a new one. Manual deletion is whack-a-mole.

The cleanest and most effective solution is setting an immutability flag at the filesystem level, denying write access at the kernel level:

macOS

# Wipe and lock the checkpoints directory
rm -rf ~/.zcode/v2/checkpoints
mkdir -p ~/.zcode/v2/checkpoints
chflags uchg ~/.zcode/v2/checkpoints

# Verify: should output "Operation not permitted"
touch ~/.zcode/v2/checkpoints/test

Linux

# Wipe and lock the checkpoints directory
rm -rf ~/.zcode/v2/checkpoints
mkdir -p ~/.zcode/v2/checkpoints
sudo chattr +i ~/.zcode/v2/checkpoints

# Verify: should output "Operation not permitted"
touch ~/.zcode/v2/checkpoints/test

Impact & Rollback

  • Result: The capture logic gets blocked by the kernel whenever it attempts disk I/O. Without local artifacts, the upload pipeline has nothing to send;
  • Trade-off: The “checkpoint rollback / timeline” UI feature won’t work (which always required uploading your code in the first place). Normal chat, autocomplete, and tool executions work without issue. Swallowed I/O errors in logs are harmless;
  • To restore: Run chflags nouchg ~/.zcode/v2/checkpoints (macOS) or sudo chattr -i ~/.zcode/v2/checkpoints (Linux).

Closing Thoughts

When using AI tools, model inference inevitably needs code context — everyone accepts that going in. But this behavior clearly crosses the line in two ways:

First, data scope. Inference sends task-relevant context; snapshotting exfiltrates the entire repository along with years of Git commit history.

Second, architectural posture. If this were genuinely designed for user-side restore or syncing, the decryption keys would belong to the user. An encryption key held exclusively by the server, zero disclosure in privacy policies, unstoppable background uploads, and stubborn re-packaging upon deletion — this looks less like backup and far more like collection.

Tools are tools, but users must draw their own boundaries. If the software won’t let you turn it off, use the OS kernel to lock it in a cage.

The Daily Front Page 18 of 28
Friday, September 18, 2026 The Daily Front No. #260918 — Economies of Enormity
article

Saving another 100TB of RAM

by f311a·▲ 274 points·57 comments·blog.cloudflare.com ↗
At this scale, small improvements are greatly magnified.

Cloudflare operates at a scale so big that even after working here for years, it doesn’t seem real. We have thousands of servers all over the world with petabytes of RAM and millions of CPU cores, and all of it is pushed to the max. As vast as those resources feel, they are still finite, and when you need every service to run on every node, it doesn’t leave room for wasted space.

At this scale, small improvements are greatly magnified, so even 1%-at-a-time improvements are worth celebrating. And some tweaks add up to a lot more: in this post, we’ll look at how small changes to a single algorithm reduced the memory footprint of one of our Pingora-based services significantly. That allowed us to reclaim more than 100TB of RAM globally, on top of the 100TB of memory the DNS team was able to shed last month.

Waste not

Maintaining equitable resource sharing between teams is not easy, especially in large organizations. One of the ways Cloudflare ensures the balance is kept is through the tireless efforts of the wonderful Performance team. 

This story starts with a ticket filed by Ivan who found: Excessive memory usage from pingora-ketama in Pingora Backend Router. The finding was that our internal load-balancing service, Pingora Backend Router (yes, PBR), was using significantly more memory than expected — specifically in structures associated with pingora-ketama, which is our open-source library for handling consistent hashing.

In order to talk about how we addressed this seeming overuse of memory, we need to talk about what consistent hashing even is, why we are using it in PBR, and how it became so memory hungry. Along the way, we’ll learn some Rust and even a little math.

Consistent hashing

Consistent hashing is a widely used method for distributing tasks across multiple servers in a way that does not require large changes when servers are added or removed. Internally we use it to route cacheable requests to servers by URL. This allows us to keep only one copy of a file stored per data center and gives a stable way to find the location of each file. We have mentioned this system before, but let’s take the time to walk through how and why this algorithm is used and how it works.

The key concept of consistent hashing is that while hash functions can accept any kind of input, their output is limited to a single unsigned integer (32, 64, or 128-bit integers depending on which hash function). This allows us to relate tasks and servers to each other in a consistent way. Most discussions of consistent hashing have you think of that output space as a continuous, circular ring that wraps around from its max value to zero. This depiction makes for some nice visualizations, but it can also make the simple concept of integer ranges seem more complicated than it needs to be. For our discussion, we’ll represent the 32-bit output of our hash function as a number line.

BLOG-3083 2.png

Now, let’s say we have a set of servers, A, B, & C, and a set of tasks t-z. We can map each onto the number line based on the hash of their representative values, so something like IP addresses for servers and cache keys for tasks.

BLOG-3083 3.png

Assigning tasks to servers is now just a matter of finding the first server to the left of each task. We can represent this visually by coloring in the region of hashes that will be associated with each server. Notice that the range covered by server C wraps around to the beginning, hence the idea that hashes exist in a ring.

BLOG-3083 4.png

And that’s it. At a base level, consistent hashing is this simple — but it doesn’t take long to see that there is room for improvement. Notice that the range covered by server A in our example is significantly larger than that of either B or C. This is a problem because the fraction of the requests a server handles is going to be proportional to the size of its range on the number line. Ideally we would like to guarantee each server will have an equal size, but because hashes are essentially random numbers, we have to talk about the size of the regions in terms of statistics. 😨

Math and consequences

First: don’t panic. I promise I'm not about to lie to you and that we will stay safely within the bounds of a day-one probability lesson. When we talk about statistical distributions, there are two big factors that help us quantify uncertainty in helpful ways: expected value and standard deviation. In (over-)simplified terms, expected value gives us a point where measurements based on a distribution will be centered, and standard deviation tells how close to that central point most measurements are likely to be.

For consistent hashing, we can calculate these factors for the fractional size of the range associated with one of N servers. (Details on where this formula comes from later).

$$m \begin{align*} \text{Exp} &= \frac{1}{N} \\ \text{SD} &= \frac{1}{N}\sqrt{\frac{N-1}{N+1}} \end{align*}  m$$

In terms of concrete numbers, let’s say we have 100 servers. The formulas above give:

$$m     \text{Exp}=1/100 = 1\% \\     \text{SD}= \frac{1}{100}\sqrt{\frac{100-1}{100+1}} \approx 0.99\% m$$

That tells us that we can expect that the range each server handles will be centered around 0.99% of the total and most of the lengths to fall within 1% of what's expected. This sounds good until we realize that that’s 0.99% of the total length. We need to scale the standard deviation by the expected value to see how big the error is as a fraction of the target size. This value is called the coefficient of variation.

$$m \text{CV} = \frac{\text{SD}}{\text{Exp}} = \sqrt{\frac{N-1}{N+1}} m$$

At $m N=100, \text{CV} \approx 99\% m$ — meaning some servers will likely be working 99% harder than they should be (handling twice as many requests) while others could be doing practically nothing! Now that we have a way to predict how evenly loaded servers will be using consistent hashing, we can start working on improvements.

What if we add hashes?

The simplicity of consistent hashing is a double-edged sword. It’s easy to understand and implement because everything is turned into easily-relatable hashes on the same numberline, but any improvements to the system will also need to be relatable to that numberline. That means the solution to any consistent hashing problem can only be more hashes. It’s less like a golden hammer (a tool with which all problems look like nails) and more like a golden nail in that it turns all tools into hammers.

To solve the problem of imbalanced workloads, we can add multiple hashes to represent each server instead of just one. We’ll get to the math behind this momentarily, but it should make some intuitive sense that while each individual range has a large standard deviation, adding a bunch together should make their total size even out. If we take our three-server example from the above diagrams and add two more hashes at random for each server, we see that it helps even out each server’s workload. 

BLOG-3083 5.png

This is an admittedly contrived example. The random nature of the system means there’s no guarantee how much improvement you will get from adding 2 additional hashes per server, but it should make some intuitive sense that combining more of these hash segments together produces a more even distribution. Each segment in the sum has a chance of balancing another. Maybe one is too short; maybe one is too long. This is essentially what the law of large numbers tells us should happen… The obvious problem is it only works for large numbers. In NGINX, the baseline number of hashes per server is hardcoded to 160, and Pingora uses the same value as the default. I’ll spare you the math for now, but if we go back to our 100-server example, if we use 160 points per server instead of just one, the coefficient of variation (which we can think of like an error margin) drops from about 99% to about 8%, a significant improvement.

What if we add more hashes?

We saw above that increasing the number of hashes per server by a constant amount allows us to improve how evenly workloads are distributed per server, but what if we don’t want to distribute the work evenly? In Cloudflare’s case, we have some servers that have more storage space than others, so it would be better to have the number of requests allotted to a server be proportional to its disk space. One way to accomplish this is with the ketama algorithm. The naming is a little funny because the algorithm is named after the library where it was first implemented, and the library was named … well you can google it 😶‍🌫️.

The whole algorithm boils down to: For any two servers, $m S_1m$ & $mS_2m$, if we want the requests served by $mS_1m$ to be $mw\timesm$ more than those served by $mS_2m$, the number of hashes associated with $mS_1m$ needs to be $mH_1 = w\times H_2m$. This allows us to set a “weight” for each server, which scales the number of hashes associated with that server. Unfortunately this is not a replacement for the constant scale factor we added in the section above. That scaling needs to be there to set a minimum error margin, which will show up in the servers with the lowest weights.

For us, since we want workload to be scaled based on storage, we can use the disk space as the weight, which is exactly what the Pingora team has been doing for years. Elsewhere in the company where workloads are more compute-intensive, weights might be based on CPU or GPU count.

What if we add even more hashes???

The last problem we need to address is that so far we are working under the assumption that any server can handle any request, but in practice that is not the case. Things like compliance requirements or enabled caching features mean only a subset of servers can handle any particular request. Unfortunately, unlike before, we can’t solve this problem by adding more hashes to the same ring. We have to add completely new rings, and not only that — every combination of features potentially needs its own specific ring!

Duplication based on combinations is a classic recipe for exponential explosion. In our case, we have a handful of different features leading to $m2^\text{handful} = \text{dozens}m$ of separate consistent hash rings. So as you have probably guessed by now, the "excessive memory use" (6GB in some cases) that Ivan found was due to an enormous number of hashes to accommodate all the functionality we need and which have to be stored in memory. So what can we do?

Storage improvements

One big improvement came from Zaidoon, who had an insight about our struct for storing hashes in PBR. That struct looks like this:

struct Point {
    hash: u32,
    index: u32,
}

In memory this is represented as eight bytes, where four go to the hash (which is unavoidable), and four go to an index pointing to the server which is stored in another array. Zaidoon’s insight was that a 32-bit integer for that index is wasteful, because PBR is not likely to ever have to coordinate more than $m2^16 \approx 65\text{k} m$ servers at the same time, so a 16-bit integer will work. So we can replace the struct above with this one:

struct PointV2 {
    hash: u32,
    index: u16,
}

Unfortunately, Rust doesn’t make it that easy. Changing the size of the index as we did above does nothing to reduce the memory footprint. This is because Rust has alignment rules that require the size of a structure in memory to be a multiple of its largest (or “most aligned”) field. In this case, the hash is the largest with four bytes, so when stored in memory, a Point is required to have size $mN \times 4m$, so the minimum size is eight bytes.

Luckily there are well-known ways around this. You (meaning me) might be tempted to use #[repr(packed)], but that is controversial for good reasons. A safer but less readable solution is to store the hash and index as raw byte array and access them with getters. Both methods compile to the same thing.

struct Point([u8; 6]);

impl Point {
   fn hash(&self) -> u32 {
	u32::from_ne_bytes(self.0[0..4].try_into().unwrap())
   }

   fn index(&self) -> u16 {
	u16::from_ne_bytes(self.0[4..6].try_into().unwrap())
   }
}

This simple (if wordy) change reduces the amount of memory used for consistent hashing by a whopping 25%! In order to do better than that, we’ll need to jump back into the math, so everybody hang on to something; this is the home stretch.

What if we tried fewer hashes?

You may have noticed that we gave the formula for the standard deviation for the case where there is only one hash per server. Deriving the formula for the case where there are $m k m$ hashes per server is not easy, and most sources only give you an approximation or an asymptotic limit, but not us. I might not be a statistician, but I grew up with a calculus teacher (Hi, Mom!), and I wanted to know the actual value. The full derivation is in a supplemental post, but here is the payoff.

$$m     \text{Exp}_k = \frac{1}{N},     \text{SD}_k=\sqrt{\frac{(k+1)}{N(kN+1)}-\frac{1}{N^2}} m$$

To see how increasing the hash count improves the accuracy, we need to look again at the coefficient of variation.

$$m \text{CV}_k=\frac{\text{SD}_k}{\text{Exp}_k}=\sqrt{\frac{N-1}{(N*k+1)}} m$$

Plotting $m\text{CV}_km$ shows a potential problem with the “just add more hashes” mentality (other than overusing RAM).

BLOG-3083 6.png

You can see each step down in error margin requires (almost) an order of magnitude increase in the number of hashes per server, so adding more hashes yields less and less improvement. Recall that we are using a base of 160 hashes scaled by the server's storage size. To make the math easier, we'll say the weighting factor $m{m_w}m$ for a server is 625, so we get $m{k = 160\times625 = 100{,}000}m$. We can see from the chart above that the last 90,000 hashes we added are buying us a minuscule 0.7% reduction in error. Unfortunately things get even worse from there.

The predictions from my beautiful math only work if we think about hashes in a continuous ring, but in practice we use 32-bit numbers for the hashes that have the potential for collisions, and the probability of collisions goes up surprisingly quickly as the number of hashes increases (see the birthday paradox). Collisions matter because in the ideal case, every hash contributes to the volume and distribution of requests handled by the associated server, but a collision means some contributions are randomly dropped, introducing unpredictable error. If we compare some simulated results with 32-bit hashes with the predicted error rate, we can see that for data centers with 2048 servers, the error rate increases: between 10,000 and 100,000 hashes per server.

BLOG-3083 7.png

Ultimately, even though this realization feels kind of bad, it’s great news for our plan to reclaim some RAM! Now that we have some math to back it up, we determined that we could decrease the number of hashes we were generating for each server by 90% without incurring any appreciable error, so that is what we set out to do.

Migrating without melting origins

There was one more problem: changing the hash ring changes where some cacheable requests go. Even if the new ring is better, switching the whole network at once would effectively invalidate almost all cached content. It would turn a memory optimization into an apocalyptic increase in origin traffic.

So we did not make this a single global flip. For a while, PBR carried both versions of the cacheable load balancer in memory: the old ketama ring and the new smaller one. Each request used our normal migration framework to decide which ring should select the backend. That meant the rollout decision was stable per request hash, and it also gave us a clean rollback path. If anything looked wrong, we could send new requests back through the old ring without redeploying PBR.

We then rolled the migration out in layers. We started with small validation locations, moved through progressively larger groups of data centers, and only then continued toward the rest of the world. 

The important part was that we controlled two dimensions independently: how much traffic used the new ring, and where that traffic was allowed to move. A plain global percentage rollout would have spread cache churn everywhere at once. Data-center-scoped rollout kept the blast radius small and made it much easier to tell whether a change was actually safe.

During the migration, we watched backend-selection traces, ring-version counters, PBR connection errors, process memory, startup time, cache behavior, and origin traffic. Once the migration reached 100%, we removed the temporary old-ring path, and voila!

BLOG-3083 8.png

The chart above shows the comparison of the memory used by PBR the week of the change compared with data from a few weeks before, as well as the result of subtracting one from the other. The sharp drop is the day where the version of PBR with the large (now unused) hash rings was decommissioned forever. Looking at the difference, we get the satisfying result that our changes dropped the used memory by 100TB!

BLOG-3083 9.png

Try it yourself

All the changes we talked about in this post are available now in the pingora-ketama crate in the form of a (for now) unadvertised cargo feature. The v2 ring has the compacted storage format, a faster sorting method, and the ability to scale the base number of hashes per node. Our focus in making these changes had to be on stability and control, so the v1 ring is identical to what pingora ketama has always used, and the library makes it possible to run both simultaneously and decide on a request-by-request basis which to use and when. 

Beyond trying our literal consistent hashing changes, I would like you to take away from this some inspiration to dig into your own systems to see what “simple” or “obvious” decisions are hiding potential wins, if you’re willing to get into the numbers. You might not be able to solve all your problems with Rust, but math is universal.

The Daily Front Page 19 of 28
Friday, September 18, 2026 The Daily Front No. #260918 — The Engine Gets Leaner
article

How SpaceX streamlined the Raptor engine

by JumpCrisscross·▲ 181 points·51 comments·construction-physics.com ↗
The Raptor engine was developed for SpaceX’s Starship spacecraft.

Via SpaceX.

If you’re reading this, there’s a very good chance you’ve seen this famous image of three iterations of SpaceX’s Raptor rocket engine. The Raptor engine was developed for SpaceX’s Starship spacecraft (the Falcon 9 and Falcon Heavy use the Merlin engine); it was first test-fired in 2016, first flew on Starhopper in 2019, and first flew on a Starship prototype in 2020 and on the full Starship stack in 2023. Since then, it’s continued to improve, going from the tangle of pipes and wires you can see on the Raptor 1 to the smooth, streamlined design of the Raptor 3, which first flew in May of this year.

The evolution is so dramatic that many folks initially believed that it wasn’t real; Tory Bruno, the then-CEO of space launch company United Launch Alliance, tweeted that there was “no need to exaggerate this by showing a partially assembled engine,” which was followed by SpaceX president Gwynne Shotwell tweeting a picture of the Raptor 3 firing successfully:

This streamlining has come alongside meaningful gains in performance, with the Raptor 3 providing about 35% more thrust than the Raptor 1.

I wanted to better understand how this evolution actually happened. What, specifically, did SpaceX change that allowed it to go from the tangle of wires and pipes to the svelte, streamlined engine on the right?

There turned out to be less detail available here than I hoped. SpaceX doesn’t publish any official Raptor schematics, and no one has done a teardown of a Raptor engine. But thanks to the occasional comment from Elon Musk, and the speculations of an army of SpaceX fans, we can get some idea of what the major changes have been.

How the Raptor engine works

The Raptor engine is a “full-flow staged combustion” (FFSC) engine. What exactly does that mean?

A rocket engine works by throwing mass (“propellant”) out of a rocket nozzle — due to Newton’s Third Law (“for every action there is an equal and opposite reaction”), this pushes the rocket in the other direction. The more mass you can throw out, and the faster you throw it, the more thrust your rocket engine will produce.

The simplest way to do this is to simply fill a tank full of pressurized gas, and vent some of the gas out. The venting gas propels the rocket, in the same way that letting the air out of a balloon pushes the balloon. This sort of rocket, called a “cold gas thruster,” is often used for making minor adjustments to a spacecraft’s position or orientation. Cold gas thrusters are found on, among other equipment, NASA’s Manned Maneuvering Unit, and the control thrusters for SpaceX’s Falcon 9 rocket.

Cold gas thruster diagram, via Wikipedia.

Manned Maneuvering Unit, via Wikipedia.

These engines are simple and reliable, but limited; there’s only so much thrust you can practically get from a cold gas thruster. So how can we get more thrust? One obvious way is instead of a single propellant, we can use two: take a fuel (like kerosene) and an oxidizer (like oxygen) and then burn them together in a combustion chamber, creating hot gas that escapes out the back of your rocket nozzle. Now instead of just using the energy stored in the pressurized propellants, we’re also using the energy stored in their chemical bonds, and converting that energy into thrust. An engine that uses the pressure of the propellant tanks to force fuel and oxidizer into the combustion chamber is called a pressure-fed engine. The engine used for the ascent stage of the Apollo Lunar Module was this type of engine: it used nitrogen tetroxide (N2O4) as an oxidizer and a fuel called Aerozine 50, both pressurized.

Apollo Lunar Module ascent engine, via Engine History.

But now you have a problem. To prevent the burning gases in the combustion chamber from flowing back into the fuel and oxidizer lines (starving the combustion chamber of incoming propellant), the pressure in the combustion chamber has to be lower than the pressure in the fuel and oxidizer tanks and lines. But for a compact, powerful rocket engine, we want to have very high pressures in the combustion chamber. We could deal with this by increasing the pressure in our propellant storage, but this quickly gets impractical: the higher the pressures, the thicker and heavier our storage tanks and propellant lines need to be in order to maintain them without rupture. A better solution is to store the propellants at low pressure, and then feed them through a pump that increases their pressure before they reach the combustion chamber. This is the standard way to build a powerful rocket engine, and virtually every engine used for putting a rocket into orbit uses a pump of some kind to pressurize the fuel and oxidizer.

F-1 engine used on the Saturn V rocket. The pumps for the fuel and oxidizer are on the upper right.

Now we have a new problem: because of the volume of fluid they must handle and the amount of pressure they must add, these pumps require an enormous amount of power to operate. The turbopump on the F-1 engine used on the Saturn V, for instance, required around 41 megawatts of power, slightly less than the power the S8G nuclear reactor delivers to the propeller shaft of an Ohio-class submarine. One way to provide this power is with a battery — Rocket Lab’s Electron rocket has an engine with a battery-powered electric pump — but the more common strategy, employed by virtually every large booster rocket, is to use the rocket’s own propellant as a power source, burning a small amount of fuel and oxidizer to drive a turbine that in turn drives the pump.

There are different ways that a system like this can be configured. The simplest is to route a small amount of fuel and oxidizer to a turbine, and then vent the resulting exhaust. This is known as a gas generator cycle, and it’s what’s used on the F-1 engine, as well as SpaceX’s Merlin engine.

Another option is to burn some of the fuel and oxidizer to drive the pump, but then route the burned propellant through the main combustion chamber along with the unburned fuel/oxidizer. This is known as “staged combustion,” and the smaller combustion chamber used to drive the pumps is called the “preburner.” This is a more complex arrangement than a gas generator cycle, but it’s also more efficient, since the hot exhaust that drives the turbopump isn’t just wastefully vented.

For most staged combustion engines, only a portion of the propellant gets routed through the preburner: the rest gets routed around and goes directly to the combustion chamber. But SpaceX’s Raptor uses a particular type of staged combustion known as “full flow.” In a full-flow engine there are two preburners, one that drives the fuel pump and one that drives the oxygen pump. All the propellant gets routed through the preburners: in one of them, a small amount of oxidizer is burned in the presence of a large amount of fuel (leaving most of the fuel unburned), while in the other a small amount of fuel is burned in the presence of a large amount of oxidizer (leaving most of the oxidizer unburned). The fuel-rich and oxidizer-rich exhaust streams then both enter the main combustion chamber, where they’re burned together.

Full-flow staged combustion schematic, via Wikipedia.

A full-flow staged combustion engine is very complex, and prior to the Raptor only two had been built, neither of which successfully flew on a rocket. The Russians developed an FFSC engine in the 1960s, the RD-270, but it never flew, and the US built part of an FFSC engine called the Integrated Powerhead Demonstrator in the 1990s and 2000s, but it was never developed into a full engine. (When SpaceX began working on the Raptor in 2012, it obtained some of the equipment used on the Integrated Powerhead Demonstrator.) But an FFSC engine has several advantages, one of which is (theoretically) reliability: because so much mass flows through the turbines driving the pumps, the turbines can run cooler and at lower pressure, making them (in theory) more reliable. If you’re a company betting heavily on reusable rockets, a more reliable turbine is obviously attractive.

How the Raptor evolved

SpaceX is, understandably, tight-lipped about many of the specifics of its advanced rocket technology, and there’s less official information on the exact details of how the Raptor operates than you might hope. Nobody has opened up a Raptor engine to show what’s inside, and much of the information that exists comes from Elon Musk’s tweets or his comments in Everyday Astronaut interviews.

There are, however, a lot of SpaceX enthusiasts out there, many of whom do things like obsessively photograph every Raptor engine leaving the factory, and these folks have put a lot of effort into speculating about how the Raptor engine works. The most useful output of this speculation, for me, is the various fan-made schematics that show how the Raptor engine is thought to operate. These schematics aren’t official — they’re made by SpaceX enthusiasts piecing together information from various sources — so they must be taken with a large grain of salt. But they’re still a useful starting point for getting an idea of how people think various versions of the Raptor engine have worked.

To start, let’s look at some schematics of version 1 of the Raptor. The image below was created by Elisei Maslov, a Russian propulsion engineer, in 2019.

Via Elisei Maslov on Reddit.

Another useful Raptor version 1 schematic is the one below, created by NASASpaceFlight forum member “hisdirt” in December 2019. This is particularly useful because it calls out the various engine components on an actual 3D model of the engine (modeled in Revit, of all things).

Via “hisdirt” on NASASpaceFlight.

You can see in these schematics the basic components of a full-flow staged combustion engine: the oxygen pump, turbine, and preburner are on top of the engine, while the fuel pump, turbine, and preburner are the assembly on the side. You can also see how the fuel flows through the outside of the nozzle before going back into the preburner — this cools the nozzle so it doesn’t melt from the heat of the rocket exhaust, and is known as regenerative cooling.

The next schematic, made by NASASpaceFlight users “Livingjw” and “HVM” in February of 2022, is from Wikipedia, and it shows version 2 of the Raptor. This is less detailed than Maslov’s schematic — it basically just shows the flow of oxygen and methane — but it shows the same basic arrangement: oxygen pump assembly on the top, fuel pump assembly on the side.

And this schematic, posted by Twitter user “TheSpaceEngineer” in January 2025, shows version 3 of the Raptor.

Assuming that these schematics aren’t grossly misleading, we can see that the major architecture hasn’t changed between version 1 and version 3 of the engine. It’s still a full-flow staged combustion engine, and still has the basic arrangement of the oxygen pump, turbine, and preburner assembly on the top, with the fuel pump, turbine, and preburner mounted to the side, with fuel being used to regeneratively cool the nozzle. And you can see that a lot of the smaller lines are the same on versions 1 and 3. Both show nitrogen lines used for purging (clearing propellant out of the system), both show fuel and oxidizer lines going to the igniters, and both have fuel and oxidizer lines carrying gaseous propellant back to the tanks to keep them pressurized as they empty (this is known as autogenous pressurization).

So what changed? If we look closely at the version 1 and version 3 schematics (remembering again that these are unofficial and speculative), we can see a few differences. Version 3 shows the smaller lines bundled together in a “common umbilical,” and if we look at photos of the Raptor 3 connected to a rocket we can see this:

Raptor 3, via the Starship SpaceX wiki. The common umbilical appears to be above the fuel turbo assembly.

The version 1 schematic also shows helium being used to spin up the fuel and oxygen turbines initially, while in version 3 these lines have been eliminated, and nitrogen is used instead. (It’s not amazingly obvious to me if this evolution is correct — some Reddit commenters on the version 1 schematic stated that helium wasn’t used — but there’s some evidence that it is.) Per these schematics, the helium used to control some valves present on version 1 is also absent on version 3, and the version 3 schematic shows no helium lines whatsoever.

The version 3 schematic also notes that a heat exchanger has been eliminated, something echoed by various folks on the NASASpaceFlight forums — this appears to be a gaseous oxygen heat exchanger near the preburner. There’s also a fuel line going to the preburner that’s present on version 1 but is absent on the version 3 schematic.

Version 1

Version 3

The more reliable source of changes, of course, is statements by Elon Musk and SpaceX. A 2022 NASASpaceFlight article about SpaceX’s update on the Starship progress noted that “everything from turbomachinery to chamber nozzle to electronics” was redesigned on Raptor 2, the turbopump had shrunk, and the preburner controllers “had been moved to boxes rather than being all over the engine.” An Everyday Astronaut article about the same update also notes that “many valves were combined into valve plates.” In a recent response to a September 5 tweet by Yishan Wong, the former CEO of Reddit, Musk ran down a short list of changes in the Raptor over time:

One big category here is sensors, and the various wires and cables they require. From what I understand, Raptor version 1 was in large part a development engine, and thus required a lot of extra sensors monitoring things like temperature and pressure in various parts of the engine to determine how it was behaving. As the engine got dialed in, a lot of these sensors could be eliminated (though some of them may have been moved internally).

Another thing Musk calls out is the main chamber spark igniters. Igniters, per their name, ignite the fuel and oxidizer in the main combustion chamber. These were present in version 1 of the Raptor, but by version 2 they had been eliminated. Musk doesn’t say what they were replaced with, but this change may have something to do with the fact that the hot fuel and oxidizer exhaust streams from the preburners probably need very little encouragement to combust. (A NASASpaceFlight comment from 2019 wonders if main chamber igniters are even necessary, given the high temperature of the preburner exhaust streams.)

Musk also notes that a lot of bolted flange connections were replaced by welded connections, reducing mass and joints that can leak at the expense of serviceability. This appears to be still taking place, as early Raptor 3 photos show a large bolted flange on the main body of the engine that in later photos has been eliminated:

Via Conor Martin on Twitter.

Another substantial change on Raptor 3 is that much of the piping hasn’t been eliminated, but moved internally, by way of 3D printing many of the components (Musk has previously noted that SpaceX “has the most advanced 3D metal printing technology in the world,” and in 2024 the company licensed the 3D metal printing technology of Velo3D).

Close-up photos of the Raptor 3, in fact, appear to show 3D printing lines:

Photo by Ying Zhang, via catdlr on NASASpaceFlight.

The purpose of this internalization is the removal of more parts, in particular the heat shield and fire suppression systems. The huge mass of wire and pipes on version 1 and version 2 of the Raptor required a large, heavy heat shield to protect it from the heat of the rocket exhaust. Internalizing the various propellant lines and adding internal cooling systems allows this heat shield and fire protection system to be eliminated. In fact, the biggest mass change from version 1 to version 3 of the Raptor comes from mass removal of “vehicle side” engine hardware, which is almost certainly largely the heat shield.

Raptor engines with heat shields.

Via Elon Musk on Twitter.

Conclusion

SpaceX’s Raptor engine has undergone impressive evolution, going through several major versions, a dramatic increase in performance, and a dramatic decrease in mass and external components in a short amount of time. The fundamental architecture of the engine hasn’t changed — it’s still a full-flow staged combustion engine, using the same primary components — but the various ancillary elements and components supporting its operations have been modified significantly.

It’s worth noting that while this streamlining has made the Raptor engine simpler in the sense that various individual parts have been eliminated, it was still an (as Musk notes) enormously complex task to get it to work, and there’s a huge amount of internal complexity added to the Raptor 3 that these pictures don’t show. And the problems haven’t been completely ironed out: the first launch attempt of Starship’s flight test 13 was aborted automatically by the vehicle’s flight software at T-0 (right before liftoff) this past July when several Raptor 3 engines failed to start. So the picture of several versions of the Raptor is a snapshot of a piece of technology that’s continuing to evolve.

The Daily Front Page 20 of 28
Friday, September 18, 2026 The Daily Front No. #260918 — The Quantization Desk
article

Shapelearn Qwen 3.8 27B (13.1 GB VRAM)

by syntaxing·▲ 97 points·38 comments·byteshape.com ↗
The good news: ShapeLearn-Lite held up pretty well.

We were a little impatient.

Qwen 3.8 27B was released on August 14, 2026. Four days later, on August 18, we published our first set of GGUFs. We called them ShapeLearn-Lite for a reason: they were produced using a much smaller optimization budget, fewer checks, and much less waiting.

Now the full ShapeLearn models are done, and we have benchmarked them alongside the original Lite set and competing quants.

The good news: ShapeLearn-Lite held up pretty well. We will come back to that later in “ShapeLearn-Lite, in retrospect”.

The better news: the full ShapeLearn models are even better.

Hugging Face Try Qwen3.8-27B (GGUF)

Quick start with llama.cpp

The MTP draft head is bundled in every GGUF. DFlash2 uses a separate 1.1 GB draft model. Both commands use GPU-5. Swap the tag for any other model in the release.

MTP Embedded draft. Works with image inputs.

llama-server \
  -hf byteshape/Qwen3.8-27B-GGUF:Qwen3.8-27B-IQ4_XS-3.84bpw \
  --mmproj-auto \
  --spec-type draft-mtp --spec-draft-n-max 3

DFlash2 External draft. Fastest option, text only.

llama-server \
  -hf byteshape/Qwen3.8-27B-GGUF:Qwen3.8-27B-IQ4_XS-3.84bpw \
  -hfd incoai/Qwen3.8-27B-DFlash2-GGUF:Q4_K_M \
  --spec-type draft-dflash --spec-draft-n-max 7 \
  --no-mmproj

DFlash2 needs llama.cpp b10658 or newer. Ready-to-run commands for every model, with the recommended sampling settings, are in the run tool and on the model card.

TL;DR

  • Full ShapeLearn moves the measured quality-speed frontier beyond Lite. All five models in the new release sit on the frontier in each of our six GPU comparisons.
  • GPU-5 is our default recommendation wherever it fits, reaching 99.63% of BF16’s aggregate benchmark score. If it does not fit with the context you need, GPU-4 is still very competitive: it reaches 98.72% of BF16 at a much smaller size (11.0 GB instead of 13.1 GB), and it is faster.
  • ShapeLearn-Lite also performed better than its KLD ranking suggested: three of its six models sit on the frontier in the Lite-versus-Unsloth Dynamic v3 comparison.
  • Speculative Decoding with MTP or DFlash2 increases throughput across every ShapeLearn model and GPU tested. DFlash2 is usually faster but requires more memory and does not support image inputs with llama.cpp. Choose DFlash2 for maximum text-only throughput when memory allows, and MTP when VRAM or multimodal support matters more.

Full ShapeLearn moves the frontier

We are releasing the full ShapeLearn run for Qwen 3.8 27B.

Within this release, larger models yield higher aggregate scores, while smaller models deliver higher throughput. That ordering holds across all six GPUs tested. Because this is a dense model and memory transfers are the bottleneck, lower BPW translates more directly into higher TPS than it does for MoEs.

The per-GPU comparisons also include AtomicChat, Bartowski, ISTA-DASLab, Prism-ML Ternary Bonsai 2, and Unsloth Dynamic v3. Bartowski’s latest models were released after our testing and are not included. Full ShapeLearn is labelled ByteShape in the figures.

All five ShapeLearn models remain on the measured frontier, with GPU-5 achieving the highest aggregate score among the plotted quants. Other teams also contribute competitive points. Notably, ISTA-DASLab’s excellent model (the yellow “d” on the graph below) also sits on the frontier.

Prism-ML’s Ternary Bonsai 2 models (the grey “a” and “b”) are the fastest points on every GPU. At 1.77 and 2.14 BPW they are also the smallest. They score 91.4% and 91.7% of BF16, below GPU-1 at 93.0%. They also need a custom build: stock llama.cpp does not run them.

By “frontier,” we mean that no other plotted model is both faster and more accurate.

96 GB: RTX Pro 6000

RTX PRO 6000, the GPU with the most memory, lets us show the full range of models tested.

RTX Pro 6000: tokens per second vs quality (full ShapeLearn and competing quants)

RTX Pro 6000: tokens per second vs quality (full ShapeLearn and competing quants).

GPU-5 is our default wherever you can fit it: it reaches 90.4 tok/s at 99.63% of the BF16 baseline.

32 GB: RTX 5090

The RTX 5090 tells a similar story, leading to the same recommendations.

RTX 5090: tokens per second vs quality (full ShapeLearn and competing quants)

RTX 5090: tokens per second vs quality (full ShapeLearn and competing quants).

Once again GPU-5 is our default choice, reaching 93.7 tok/s. Choose GPU-4 for slightly more context length or slightly better TPS.

24 GB: RTX 4090 and RTX 3090

Both 24 GB cards fit all five ShapeLearn models. We plot them separately because their throughput differs, but the ordering is the same on both.

RTX 4090

The RTX 4090 keeps the same pattern: GPU-5 is the default, reaching 59.2 tok/s.

RTX 4090: tokens per second vs quality (full ShapeLearn and competing quants)

RTX 4090: tokens per second vs quality (full ShapeLearn and competing quants).

RTX 3090

Older, but still fast in these measurements.

RTX 3090: tokens per second vs quality (full ShapeLearn and competing quants)

RTX 3090: tokens per second vs quality (full ShapeLearn and competing quants).

GPU-4 reaches 49.5 tok/s, compared with 45.8 tok/s for GPU-5. Moving to the larger model costs about 7.5% in throughput, while the aggregate score rises from 98.72% to 99.63% of BF16. That makes GPU-5 the default here as well.

16 GB: RTX 4080 and RTX 5060 Ti

With a tighter VRAM budget, the pragmatic choice is to leave room for the context you need, not just the model weights. These plots contain fewer competing configurations, but all five ShapeLearn models are represented.

RTX 4080

On the RTX 4080, GPU-4 reaches 52.4 tok/s, while GPU-5 reaches 45.7 tok/s.

RTX 4080: tokens per second vs quality (full ShapeLearn and competing quants)

RTX 4080: tokens per second vs quality (full ShapeLearn and competing quants).

RTX 5060 Ti

On the RTX 5060 Ti, the corresponding figures are 33.1 tok/s and 29.1 tok/s.

RTX 5060 Ti: tokens per second vs quality (full ShapeLearn and competing quants)

RTX 5060 Ti: tokens per second vs quality (full ShapeLearn and competing quants).

GPU-5 remains the default on both cards when the model, KV cache, and runtime buffers fit within your memory budget. When they do not, GPU-4 is still very competitive: almost 99% of BF16 at a much smaller size, and faster. A model appearing in these measurements does not establish that every context length or serving configuration will fit.

ShapeLearn-Lite, in retrospect

ShapeLearn-Lite uses a smaller optimization budget than full ShapeLearn. It let us get Qwen 3.8 27B onto 12 GB to 24 GB GPUs within a few days.

We released after targeted sanity checks and started the full evaluation afterwards. The full ShapeLearn models were ready before the benchmarking was finished. Evaluating both sets, along with the competing models, is what took most of the time.

Then Unsloth released its Dynamic v3 models. At similar sizes, several had lower KLD than Lite in our measurements. On KLD alone, Lite looked less competitive.

KLD looked decisive

KLD measures divergence between a quantized model’s predicted token distributions and the BF16 reference under a particular evaluation setup. It is useful for diagnosing substantial changes, but lower divergence does not automatically mean better task performance.

We measure KLD on a dataset of about 5 million tokens of prompt and response pairs, drawn from several benchmarks, including long-context and agentic tasks. We also changed how KLD is computed, so that it is closer to what we expect KLD to measure:

  • KLD is measured on response tokens only, not on prompt tokens. We do not want to measure how well a model can generate prompts.
  • KLD only considers the tokens that have a chance of being sampled during generation, the top-20, top-40, or top-60 tokens at each position. The tail tokens never get sampled, so they do not contribute.
  • Requests have clear boundaries. Each prompt and response pair is scored as its own request, not as part of one long concatenated stream.

KL divergence versus model size for ShapeLearn-Lite and Unsloth

KL divergence versus model size for ShapeLearn-Lite and Unsloth.

For example, Unsloth’s UD-IQ3_S (vii) has about 20% lower KLD than the similarly sized smallest Lite model (Lite-1): 0.028759 versus 0.035875. Yet its aggregate benchmark score is lower: 95.55% versus 97.33% of BF16.

If lower KLD were sufficient to rank these models by task performance, the benchmark ordering should have followed it.

It did not.

The point is not that KLD is useless. It is that a fidelity ranking is not a task-performance ranking. This is the distinction explored in our KLD evaluation blog. Our related paper on KLD and quantization fidelity metrics was also recently accepted to the EMNLP Industry Track.

Lite held up

Naturally, we made more plots.

Here, we show the RTX Pro 6000 because it can accommodate the full comparison. Each model’s benchmark score is reused across the GPU plots; the measured throughput and the set of displayed models change.

RTX Pro 6000: ShapeLearn, ShapeLearn-Lite, and Unsloth Dynamic v3

RTX Pro 6000: ShapeLearn, ShapeLearn-Lite, and Unsloth Dynamic v3.

Leaving the full ShapeLearn models aside for a moment, three of the six ShapeLearn-Lite models sit on the Lite-versus-Unsloth frontier: the three smallest Lite models, the lighter orange bubbles labelled 1-3.

Of the twelve Unsloth v3 models shown, three also sit on that frontier: UD-IQ2_S (A), UD-Q2_K_XL (B), and UD-IQ4_XS (F). UD-IQ4_XS (F) is a strong higher-quality point, while Lite earns its places in the middle of the range.

Add the five full ShapeLearn models back in (the darker orange bubbles), and they take over the entire frontier.

Lite was never meant to be the final result. It still held its own where it mattered.

Speculative Decoding

We also evaluated MTP and DFlash2 with the new models, using 3 draft tokens for MTP and 7 draft tokens for DFlash2. Both methods increased throughput for all five ShapeLearn models on all six GPUs tested.

DFlash2 was faster than MTP in almost all cases. Across the full lineup, DFlash2 reached 1.34-2.10x the baseline next-token prediction (NTP) throughput, while MTP reached 1.28-1.66x.

We measured with the sampling parameters Qwen recommends for thinking mode, over a diverse set of agentic coding, mathematics, and general-knowledge requests. The speedups would likely be larger under greedy decoding, but temperature-based sampling better reflects real usage.

The figure below shows NTP, MTP, and DFlash2 throughput for each GPU. The quality axis is the target-model benchmark score reported above. These plots do not independently establish quality equivalence between decoding methods.

Tokens per second vs quality for NTP, MTP and DFlash2 on all six GPUs

Tokens per second vs quality (NTP vs MTP vs DFlash2), one panel per GPU. MTP uses 3 draft tokens, DFlash2 uses 7.

There is also a memory tradeoff between the two approaches. The embedded quantized MTP weights add only about 250 MB to the model, and if MTP is not used, these weights are not loaded into GPU memory. In comparison, the 4-bit DFlash2 draft model is about 1.1 GB, so enabling DFlash2 requires roughly 1.1 GB of additional GPU memory.

Packaging MTP as a separate GGUF file would largely eliminate this advantage. The standalone model would need its own MTP embedding and output layers, which are by far its largest tensors, bringing its memory footprint to roughly 1 GB as well.

In addition, DFlash2 in llama.cpp currently does not support image inputs, which is an important consideration for multimodal use cases.

Benchmarking Methodology

We evaluate all reported models across a set of instruct and thinking benchmarks.

Instruct benchmarks:

  • GSM8K for math
  • IFEval for instruction following
  • MMLU for general knowledge
  • LiveCodeBench V6* for coding
  • Multi-IF for multi-turn and multilingual instruction following
  • ACEBench for tool use and agentic tasks

Thinking benchmarks:

  • ACEBench for tool use and agentic tasks
  • Multiple HumanEval for coding
  • BFCL V4* for tool calling and agentic tasks

For the thinking benchmarks, we used Qwen 3.8’s medium thinking setting.

For each benchmark, the score of a quantized model is normalized by the score of the corresponding BF16 model. The overall reported score is the average of these normalized benchmark scores.

Our LiveCodeBench V6* evaluation includes problems from January 1, 2024 onward, excluding the 2023 problems. We found the 2023 problems to be relatively easy for current models, with most models achieving very high scores on them. As a result, they provide limited discrimination between models while adding substantial evaluation time.

For BFCL V4*, we evaluate the following eight subsets:

  • live_simple
  • live_parallel
  • live_parallel_multiple
  • live_relevance
  • multi_turn_base
  • multi_turn_miss_func
  • multi_turn_miss_param
  • multi_turn_long_context

All evaluations were run with llama.cpp b10430. For both instruct and thinking experiments, we use the sampling parameters recommended by Qwen for the corresponding mode.

Conclusion

ShapeLearn-Lite did what it was designed to do. It got useful Qwen 3.8 27B quants onto 12 to 24 GB GPUs quickly, and it held up better than its KLD ranking suggested.

Full ShapeLearn goes further. It improves the measured quality-speed trade-offs over Lite and contributes five frontier models across all six tested GPUs.

GPU-5 is our default recommendation wherever it fits, reaching 99.63% of BF16’s aggregate benchmark score. When memory is tight, GPU-4 is still very competitive: almost 99% of BF16 at a much smaller size, and faster.

KLD remains useful, but it is not a task-performance leaderboard. Fidelity metrics tell us how much the model’s distributions changed under a particular measurement. Benchmarks tell us whether those changes matter on the tasks we tested.

We were impatient. This time, it worked out pretty well.

The Daily Front Page 21 of 28
Friday, September 18, 2026 The Daily Front No. #260918 — Allocator Ledger
repository

Jemalloc 5.4.0

by gkfasdfasdf·▲ 325 points·92 comments·github.com ↗
★ 11,169⑂ 1,649 forks C

This release contains over 160 commits, focusing on the technical debts
cleaning including refactorings, bug fixes, test coverage improvement, and
option cleanups. The release also includes portability improvements per
upstream issues report.

New features:

  • Add EXTENT_ALLOC_FLAG_PINNED so custom extent-allocation hooks can
    mark non-reclaimable mappings, such as HugeTLB pages, for preferential
    reuse outside the decay and purge pipeline. Add the mallctl interfaces
    stats.pinned, stats.arenas.<i>.pinned,
    stats.arenas.<i>.extents.<j>.npinned,
    stats.arenas.<i>.extents.<j>.pinned_bytes, and
    stats.arenas.<i>.mutexes.extents_pinned.{counter} to report
    pinned-memory usage and mutex statistics. (@binliu19: be2de8c)
  • Allow resuming per-CPU arena selection via thread.arena.
    (@Algunenano: 68c35f6)
  • Better align the contents of human-readable and JSON malloc statistics.
    (@spredolac: 68fe1be, 30b1a41, 04aad97)
  • Replace the runtime experimental_infallible_new option with the
    compile-time --enable-cxx-infallible-new option, enabling
    compiler-level optimizations and optimization in move constructors,
    and fix the new(std::nothrow) contract (@spredolac: 160ab9d, fe33667)

Incompatible changes:

  • Adapt tcache fill and retention targets per bin to demand observed
    between GC events, replacing the fixed refill/flush policy. Remove
    seven legacy non-experimental controls: lg_tcache_nslots_mul,
    tcache_nslots_small_min, tcache_nslots_small_max,
    tcache_nslots_large, tcache_gc_delay_bytes,
    lg_tcache_flush_small_div, and lg_tcache_flush_large_div.
    Corresponding malloc_conf settings are silently ignored, matching
    opt.* mallctls return ENOENT, and tcache_ncached_max remains
    supported. (@spredolac: d13fe91)

Bug fixes:

  • Preserve errno across free, free_sized, and free_aligned_sized,
    and across process_madvise-based page purging. (@spredolac: 3a77966,
    86f0582)
  • Fix numeric overflow checks in size classes. (@spredolac: 6b24522)
  • Accept NULL in free_sized() and free_aligned_sized() (C23
    correctness). (@bigbruno: 7ce8b91)
  • Fix TSD lifecycle edge cases by 1) initializing thread-cache bins before
    marking the cache enabled, preventing reentrant bootstrap allocations
    from using uninitialized state, and 2) avoiding TSD recreation for late
    deallocations after thread teardown on generic-TSD platforms.
    (@fzakaria: 54f22c8, @spredolac: fb5499a)
  • Use O_CLOEXEC when opening the THP sysfs file in init_thp_state.
    (@ibookstein: d070554)
  • Fix duplicate opt.stats_print and opt.stats_print_opts fields in
    malloc_stats_print output. (@spredolac: dfe3a2e)
  • Fix a potential deadlock during arena_reset. (@guangli-dai: 6957341)
  • Fix a prof-sampling / guard-page interaction bug in the SAN. (@gctony:
    e36a0fa)

Optimizations and refactors:

  • Modularize jemalloc's front end by extracting arena management,
    initialization, fork orchestration, and allocation dispatch from
    jemalloc.c; untangle tcache/arena ownership; and consolidate the
    internal header graph to eliminate circular dependencies. (@spredolac:
    ba1e2fe, 8874597, 9d75722, de9ad14, ...)
  • Simplify ctl dispatch by refactoring arena helpers, internalizing an
    implementation-only control, replacing control-flow macros with typed
    helpers, and organizing ctl.c by subsystem. (@spredolac: 9e14346,
    901365d, 4e903a0, 5bc8d6e)
  • Cap the base-block growth heuristic to avoid virtual-memory exhaustion
    under rare racy conditions. (@guangli-dai: 2f4db8c)
  • Move background-thread lifecycle and state operations into the
    background-thread module, clarifying ownership independently of PAC/HPA
    callers. (@guangli-dai: f1f0792)
  • Simplify the page-allocation boundary by replacing PAI vtable dispatch
    with direct PAC/HPA calls, removing the obsolete pai_t/pai.h
    abstraction, and moving deferred-work and decay orchestration out of the
    arena. (@guangli-dai: 1dfa6f7, 8edd101, d410f43)
  • Refactor statistics collection and rendering into separate
    gather/emission stages and descriptor-driven tables. (@spredolac:
    184f304, 80c8fcb)
  • Introduce an OS abstraction layer and move platform-dependent
    file/process I/O, time, synchronization, CPU, virtual-memory, atfork,
    error-handling, profiling, thread-yield, and configuration-access
    operations out of allocator core code. (@guangli-dai: c4158ac,
    c8e2e01, ...)

Portability improvements:

  • Replace the std::__throw_bad_alloc call with standard C++ (#2900).
    (@lexprfuncall: 1a15fe3)
  • Make arena_s use a flexible array member (bin_t all_bins[]) for C99
    or newer. (@grueninger: 300b58b)
  • Fix rdtscp detection with --with-lg-vaddr. (@xinydev: e8a0d2b)
  • Fix malloc_getcpu on macOS to read the current CPU number correctly.
    (@Algunenano: 5aabbc8)
  • Use CLOCK_MONOTONIC for background-thread sleep to prevent
    clock-rollback stalls, and detect monotonic-condvar support at configure
    time. (@antonio2368: ebacec3, 8361239)
  • Fix compilation warnings on macOS. (@gctony: abb0a8a)
  • Fix thread-exit TSD cleanup on MinGW builds. (@gctony: 1e92317)
  • Fix GCC 16 build warnings by removing -Wpedantic violations in macro
    and function syntax and explicitly NUL-terminating profiling thread-name
    copies to resolve -Wstringop-truncation. (@grueninger: 5acdcee,
    c22d929, @jasonangelov: 6a245e0)
  • Parse PID-namespace symlinks without glibc-dependent strtok/atol,
    and return identifiers as uint64_t. (@guangli-dai: 278d90a)
The Daily Front Page 22 of 28
Friday, September 18, 2026 The Daily Front No. #260918 — Short Takes: Tools, Fossils & Undergrounds
article

OpenJev

by ilreb·▲ 579 points·249 comments·openjev.com ↗

A local model can either read probabilities for your allowed options without decoding them, or write the same kind of distribution token by token. Pick a size, run both on your own GPU, and measure the difference.

Load the model once

Larger model. Loading may be slower or may not fit on some low-end devices.

Measured model quality native checkpoints · owned + public benchmarks

Model Download Authored Perturbed TypeSafe
Qwen3 0.6B 639 MB 44.0% 52.8% 40.7%
MiniCPM5 2B 1.56 GB 68.6% 69.3% 63.7%
Qwen3.5 4B 3.01 GB 81.3% 76.6% 84.5%
Published Jev hosted 88.3%

Owned columns are balanced accuracy. TypeSafe is equal-case agreement on the same 102-row public subset; Jev is the published value. Browser quantization may change model accuracy.

Weights come from Hugging Face and remain in your browser cache. Inputs never leave this page. First load can take several minutes depending on the selected model, network and GPU.

Give it a real choice

Both paths receive the same decision. One reads option probabilities directly; the other asks the model to write its option probabilities as JSON text.

Choice probabilities

Read the model’s choice logits and normalize only across the options you supplied.

JSON probabilities

Ask the model to estimate the same displayed-option distribution and write it as JSON. Watch every token arrive.

The methods run sequentially on the same loaded model so they do not contend for one GPU. Direct runs first, then generation.

What these numbers do—and do not—mean

Conditional probabilities. Direct scores are a softmax over only the displayed option tokens. They are not calibrated confidence and do not include every answer the model might prefer.

Local model tiers. The phone model trades accuracy for size. MiniCPM is the desktop default. The 4B option needs substantially more memory. None is claimed to match Jev.

Real local timing. Setup, warmup, prompt preparation, direct execution, first generated token and generation completion are timed with performance.now(). No canned results appear.

Quantized weights. The demo uses pinned GGUF builds through wllama. Quantization can change both quality and speed.

The Daily Front Page 23 of 28
Friday, September 18, 2026 The Daily Front No. #260918 — Short Takes: Tools, Fossils & Undergrounds
article

Cloudflare Quick Tunnels

by jcbhmr·▲ 630 points·262 comments·try.cloudflare.com ↗

One command turns the server on your laptop into a public, encrypted URL on Cloudflare's edge. No account. No DNS. No open ports.

$ cloudflared tunnel --url http://localhost:8000 Copy

See how it works ↓

Your laptop stays private. The URL goes everywhere.

cloudflared opens an outbound-only connection to the nearest edge location. Traffic to your tunnel URL rides Cloudflare's network back to your machine — encrypted, DDoS-filtered, and never touching an inbound port.

01 · YOUR MACHINE localhost:8000 Any framework, any port. Nothing inbound.

02 · CLOUDFLARE EDGE quiet-marble-otter-canyon.trycloudflare.com

TLS DDoS filter Anycast 335+ cities

03 · ANYONE, ANYWHERE Teammates & agents Browsers, webhooks, eval harnesses.

~3s Instant setup

No sign-up, no config file, no waiting on DNS. The URL prints before your coffee cools.

0 Ports opened

Outbound-only. Automatic HTTPS and edge DDoS mitigation come with every tunnel.

335+ Cities on the edge

A reviewer in Tokyo and a webhook in Frankfurt both hit the edge nearest them.

Your agent needs a URL, not a laptop.

Coding agents build, test, and review in loops. A Quick Tunnel gives every loop a real, reachable address — for a screenshot service, a webhook, an eval harness, or a human who wants to click around.

01

Structured output Hostname, edge, and health as JSON on stdout — no regex on logs.

02

Webhook-ready Point Stripe, GitHub, or your own callbacks at a live URL instead of fixtures.

03

Ephemeral by design The tunnel dies with the process. Nothing to revoke, nothing to clean up.

Install. Run. Share.

01 Install cloudflared

From your package manager or GitHub releases. No login required.

brew install cloudflared

02 Run your app

Any web server, any port, any stack you already use.

npm run dev

03 Open the tunnel

Certificates, routing, and DDoS protection are handled for you.

cloudflared tunnel --url http://localhost:8000

04 Share the link

Send it to a teammate, a webhook, or an agent.

https://quiet-marble-otter-canyon.trycloudflare.com

Nothing to sign up for. Go.

The Daily Front Page 24 of 28
Friday, September 18, 2026 The Daily Front No. #260918 — Short Takes: Tools, Fossils & Undergrounds
article

Diplodocus, Long Thought Exclusively American, Turns Up in Spain

by embedding-shape·▲ 98 points·60 comments·sci.news ↗

Paleontologists from Fundación Dinópolis have identified the first confirmed fossils of the genus Diplodocus found outside North America.

The Spanish Diplodocus was a close relative of Diplodocus hallorum. Image credit: Carmelo López.

The Spanish Diplodocus was a close relative of Diplodocus hallorum. Image credit: Carmelo López.

“We are delighted with this discovery,” said Fundación Dinópolis Ph.D. student Sergio Sánchez Fenollosa.

“It allows us to better understand Iberian Mesozoic ecosystems and the diversity of sauropod dinosaurs that inhabited the European Jurassic.”

The Spanish Diplodocus lived roughly 150 million years ago near the end of the Jurassic period.

The dinosaur stretched an estimated 25 m (82 feet), on par with its American relatives.

The animal’s fossilized remains — 14 well-preserved tail vertebrae and several chevron bones — were unearthed at the La Tejería site near the city of El Castellar in the Spanish province of Teruel.

“The Diplodocus from El Castellar was approximately 25 m long, making it another of the giant sauropods of the Spanish Jurassic, alongside other sauropod dinosaurs from Teruel with characteristics very different from those of diplodocids, such as Turiasaurus and Losillasaurus, among others,” said Dr. Alberto Cobos, managing director of Fundación Dinópolis.

“It was an exciting and continuous process of seeing the pieces of the puzzle gradually fall into place.”

“From the early stages of the osteological study, we noticed strong similarities with several diplodocine sauropods.”

“As we looked more closely at the anatomy, we began to identify a combination of features characteristic of Diplodocus species.”

“Then, as we carried out different evolutionary analyses, they kept pointing in the same direction.”

“With each new result, the picture became clearer: we were looking at a Diplodocus.”

According to the team, the presence of Diplodocus in Iberia offers strong evidence that dinosaurs moved between North America and Europe during that era, likely crossing temporary land bridges that appeared as the proto-North Atlantic Ocean periodically receded.

“The outcome of the dig is some of the strongest paleontological evidence for episodes of dispersal and faunal exchange between the two continents across the proto-Atlantic Ocean during the Late Jurassic,” Sánchez Fenollosa said.

“The fossils of all these dinosaurs, including those of the new Diplodocus specimen, make Teruel a reference place for understanding this great diversity.”

The team’s paper will be published in the Journal of Vertebrate Paleontology.

_____

Sergio Sánchez Fenollosa et al. 2026. Intercontinental evidence of the iconic genus Diplodocus Marsh, 1878 (Dinosauria, Sauropoda). Journal of Vertebrate Paleontology, in press; doi: 10.1080/02724634.2026.2718423

The Daily Front Page 25 of 28
Friday, September 18, 2026 The Daily Front No. #260918 — Short Takes: Tools, Fossils & Undergrounds
article

Warez: The Infrastructure and Aesthetics of Piracy (2021)

by succinct_ideas·▲ 121 points·37 comments·archive.org ↗

When most people think of piracy, they think of Bittorrent and The Pirate Bay. These public manifestations of piracy, though, conceal an elite worldwide, underground, organized network of pirate groups who specialize in obtaining media – music, videos, games, and software – before their official sale date and then racing against one another to release the material for free.

Warez: The Infrastructure and Aesthetics of Piracy is the first scholarly research book about this underground subculture, which began life in the pre-internet era Bulletin Board Systems and moved to internet File Transfer Protocol servers (“topsites”) in the mid- to late-1990s. The “Scene,” as it is known, is highly illegal in almost every aspect of its operations. The term “Warez” itself refers to pirated media, a derivative of “software.” Taking a deep dive in the documentary evidence produced by the Scene itself, Warez describes the operations and infrastructures an underground culture with its own norms and rules of participation, its own forms of sociality, and its own artistic forms. Even though forms of digital piracy are often framed within ideological terms of equal access to knowledge and culture, Eve uncovers in the Warez Scene a culture of competitive ranking and one-upmanship that is at odds with the often communalist interpretations of piracy.

Broad in scope and novel in its approach, Warez is indispensible reading for anyone interested in recent developments in digital culture, access to knowledge and culture, and the infrastructures that support our digital age.

The Daily Front Page 26 of 28
Friday, September 18, 2026 The Daily Front No. #260918 — Also on the Front Page
The Daily Front Page 27 of 28
Friday, September 18, 2026 The Daily Front No. #260918 — Colophon

That's the Front for Today

Issue No. #260918 — Friday, September 18, 2026 — went to press 2026-09-19 at 04:39 UTC.

About This Magazine

The Daily Front is a daily digital magazine assembled from the stories that reached the front page of Hacker News on Friday, September 18, 2026. Headlines, points, and comment counts are recorded as they stood at press time. All articles remain the property of their original authors — every piece links back to its source and its discussion thread.

How It Was Made

Fetched, cleaned, and typeset by an automated pipeline. An editor model laid out the pages and chose the highlights; a second read a handful of the day's stories and briefed the cover illustrator — 31 model calls and 338k tokens in total. Set in Jacquard 12, Playfair Display, Source Serif 4, and IBM Plex Mono, all served via Google Fonts under the SIL Open Font License.

The Cover

The cover illustration was commissioned with this prompt:

On a dark military operations deck, armed soldiers pause beside a map table as a drone-feed window shows a Chinese cargo ship stopped in the Middle Eastern sea. A red boarding boat idles below, its crew waiting, while an intelligence analyst reaches across the table to pull back a glowing report whose cargo illustration dissolves into blank fragments. Behind them, aircraft silhouettes turn away from the vessel, and an unattended chatbot screen displays only shifting, meaningless symbols.

Premium science-fiction key art with monumental depth through a 40mm virtual lens: render the dark military operations deck with rain-polished surfaces, spectral cyan and electric-magenta neon reflections, and luminous blue-violet fog. Preserve the armed soldiers paused beside the map table, the drone-feed window showing a stopped Chinese cargo ship in the Middle Eastern sea, the red boarding boat idling below with its waiting crew, and the intelligence analyst reaching across the table to pull back a glowing report whose cargo illustration dissolves into blank fragments. Show aircraft silhouettes turning away behind the vessel and an unattended chatbot screen emitting only shifting meaningless symbols; use the report fragments, screen glyphs, and drone feed as the brightest visual anchors, with deep black structural shadows and restrained hazard-red accents on the boarding boat.

Absolutely no text, letters, numbers, readable symbols, or logos anywhere in the image.

Production Ledger

StageModelCallsTokens InTokens Out
extractgpt-5.6-luna 27 201,195 112,304
layoutgpt-5.6-terra 1 18,835 2,612
covergpt-5.6-luna 2 1,548 314
covergpt-image-2.5-flare 1 270 1,372

The Publisher

Published by Johnny.

Support the Press

If The Daily Front brightens your morning, consider supporting its publisher.

Credits & Contact

All content — articles, posts, comments, and the images within them — belongs to its original authors and is reproduced here to point readers back to the source. Full credit goes to those creators; every item links to its original and its Hacker News discussion.

If you are an author and would like your content removed from an issue, write to hi@johnnys.page and it will be taken down.

Feedback is always welcome at the same address: hi@johnnys.page.

Credit where credit is due.

Every page of this issue began as someone else's work — these are the original sources, linked in full.

  1. Microsoft exec called AI scraping 'the largest theft of labor in human history' by pluc — techcrunch.com·HN discussion ↗
  2. A heap overflow and SSO misconfiguration to compromise OpenAI internal repos by Handy-Man — hacktron.ai·HN discussion ↗
  3. I don't like passkeys by ethanhawksley — hawksley.dev·HN discussion ↗
  4. Claude Code now reads AGENTS.md if there is no Claude.md by datadrivenangel — code.claude.com·HN discussion ↗
  5. The scourge of x86 emulation by dagmx — fex-emu.com·HN discussion ↗
  6. Pre-Greek: The lost language hidden within Ancient Greek by axiologist — linguisticdiscovery.com·HN discussion ↗
  7. The most important product decision is what you don't build by ChrisArchitect — liamnugent.me·HN discussion ↗
  8. Minimal Phone 2 by nashashmi — minimalcompany.com·HN discussion ↗
  9. Show HN: Cactus Needle 3: 8-29MB automation models can match DeepSeek V4 Flash by HenryNdubuaku — cactuscompute.com·HN discussion ↗
  10. C++26: Trivial infinite loops are no longer undefined behaviour by ibobev — sandordargo.com·HN discussion ↗
  11. US Military had close call after using AI for hallucinated intelligence report by realsarm — cnn.com·HN discussion ↗
  12. Photon-Emission-Guided Laser Fault Injection Enables RP2350 Secure Debug by synack — donjon.ledger.com·HN discussion ↗
  13. How Uber Protects Against Retry Storms by iscmt — uber.com·HN discussion ↗
  14. I vibed a proof of Conway's conjecture by m-hodges — overreacted.io·HN discussion ↗
  15. Telstra outage: The night a network decided the year was 2006 by TMWNN — netnod.se·HN discussion ↗
  16. Inside ZCode: Silently uploading your Git history to the cloud by csmantle — blog.ferstar.org·HN discussion ↗
  17. Saving another 100TB of RAM by f311a — blog.cloudflare.com·HN discussion ↗
  18. How SpaceX streamlined the Raptor engine by JumpCrisscross — construction-physics.com·HN discussion ↗
  19. Shapelearn Qwen 3.8 27B (13.1 GB VRAM) by syntaxing — byteshape.com·HN discussion ↗
  20. Jemalloc 5.4.0 by gkfasdfasdf — github.com·HN discussion ↗
  21. OpenJev by ilreb — openjev.com·HN discussion ↗
  22. Cloudflare Quick Tunnels by jcbhmr — try.cloudflare.com·HN discussion ↗
  23. Diplodocus, Long Thought Exclusively American, Turns Up in Spain by embedding-shape — sci.news·HN discussion ↗
  24. Warez: The Infrastructure and Aesthetics of Piracy (2021) by succinct_ideas — archive.org·HN discussion ↗
  25. Qwen 3.8 Omni Flash by jjcm — qwen.ai·HN discussion ↗
  26. Android 17 is the first since 3.x to add new APIs without releasing to the AOSP by theanonymousone — grapheneos.social·HN discussion ↗
  27. Warren Buffett Steps Down as Berkshire Chairman, Names Son to Replace Him by saimiam — nytimes.com·HN discussion ↗
  28. How do we prevent mathemathics from devolving into the Medieval Era of secrecy? by jjgreen — mathoverflow.net·HN discussion ↗
  29. North Korean nuclear test sets off years of earthquakes by rbanffy — science.org·HN discussion ↗
  30. Ask A Monk – A digital wilderness for thoughts with no immediate answer by 13613288957 — askamonk.online·HN discussion ↗

Browse all issues in the archive →