Cover illustration

TheDaily Front

Issue No. #260901 Tuesday, September 1 2026 #260901 — TUESDAY, SEPTEMBER 1, 2026
Platforms tighten the gates while machines, models, and memories find other ways through.
Tuesday, September 1, 2026 The Daily Front No. #260901 — Contents
30stories
9,783points
5,233comments
299kllm tokens
Assembled with 32 model calls — 197,619 tokens read, 101,028 written.

Highlights

AnkiDroid: Google Play no longer allowing Open Collective donation link

AnkiDroid faces removal from Google Play after its Open Collective donation link was rejected.

Claude Fable 5.1 and Claude Mythos 5.1

Claude’s new Fable and Mythos releases put capability, safeguards, and pricing under the same bright lamp.

I trained a small transformer in 1.5hrs and it beats many LLMs

A small open transformer claims a notable ARC-AGI result after just 1.5 hours of training.

Evidence of Fraud in an Influential Study About Procrastination

A replication failure renews scrutiny of a famous study on deadlines and procrastination.

Show HN: Running 104GB Qwen3.8-Flash-Next on 48GB Mac with at ~12 tok/s

Slotstream demonstrates how a 104 GB model can be streamed from SSD to run on a 48 GB Mac.

From the Editor

The machines have had a busy day: some grew cleverer, some grew smaller, and some found themselves locked out by the very storefronts meant to distribute them. Meanwhile, the old crafts—aviation maintenance, vacuum tubes, rotary telephones, and careful skepticism—continue to make an excellent case for themselves.

  1. AnkiDroid: Google Play no longer allowing Open Collective donation link3
  2. Claude Fable 5.1 and Claude Mythos 5.14
  3. I trained a small transformer in 1.5hrs and it beats many LLMs5
  4. How accurate have Ed Zitron's AI skeptic predictions been?6
  5. Show HN: Running 104GB Qwen3.8-Flash-Next on 48GB Mac with at ~12 tok/s7
  6. Atlas: A World Model for Spatial Intelligence8
  7. GPU World9
  8. The ChatGPT/Codex app bundles a full copy of LibreOffice10
  9. Launch HN: Nori Robotics (YC S26) – A low-cost humanoid robot for development10
  10. Fastpotify11
  11. Play Store blocks AuroraStore, hurting GrapheneOS users11
  12. Introducing Ad Blocker for Firefox on iOS12
  13. Evidence of Fraud in an Influential Study About Procrastination13
  14. American Airlines mechanic Azriel “Al” Blackman has died14
  15. Lion-man15
  16. Borges Labyrinth in Venice reopens to the public16
  17. Magic eye tube17
  18. Refurbishing a Tektronix TDS7104 Oscilloscope18
  19. RotaryCell: Making an unmodified rotary phone work over LTE with an ESP32-S319
  20. 2004 RuneScape fit a multiplayer RPG into 56k dial-up20
  21. Flat vs. segmented memory – it's recursive21
  22. Movie Scene Map – 13,312 films, series, games, anime and manga22
  23. Run macOS Software on Linux23
  24. The creator of Jujutsu has joined ERSC24
  25. Restroom Archive25
  26. Terence Tao explains 6 essential mathematical concepts [video]26
  27. Ambient CSS v3 – Blender meets CSS26
  28. Tmp.0ut Volume 526
  29. Ask HN: Who is hiring? (September 2026)27
  30. Ask HN: Who wants to be hired? (September 2026)27
The Daily Front Page 2 of 28
Tuesday, September 1, 2026 The Daily Front No. #260901 — Play Store at the Gate
repository

AnkiDroid: Google Play no longer allowing Open Collective donation link

by hexa555·▲ 858 points·254 comments·github.com ↗
Unless resolved, AnkiDroid will be removed from Google Play on 11 September worldwide.

TL;DR: What we need

From Google: clarification on whether an IRS 501(c)(6) determination satisfies "tax exempt donations".

From readers: please share this post, especially to anyone at Google who may be able to help. Please do not contact Google support directly.


Caution

Since 28 August, Google has rejected updates to AnkiDroid on the Play Store. Unless resolved, AnkiDroid will be removed from Google Play on 11 September worldwide (except for India and Russia).

Summary

AnkiDroid is a free, open-source flashcard app for Android with over 10 million installs, used worldwide for various purposes, primarily in medical education and language learning.

AnkiDroid's donations go to Open Source Collective, a US non-profit which acts as our fiscal host. It holds an IRS determination letter stating that it is tax-exempt under 501(c)(6). This determination letter has been provided to Google.

  • Google's payments policy states: "Other than the conditions described in Section 3 ..., apps may not lead users to a payment method other than Google Play's billing system"
  • Section 3 of the policy states Play billing "must not be used in cases where payments include … tax exempt donations".
  • Google's 2026-08-06 email states donations may only be collected for "a validated tax-exempt organization (for example, a validated 501(c)(3) charitable organization in the United States or the local equivalent)".
  • Google's 2026-08-07 email, [after receiving an IRS determination letter for 501(c)(6) status], states AnkiDroid "allows users to contribute donations to an organization that is not tax-exempt".

Google support's last reply is as follows:

As mentioned in the previous email, we reviewed the documentation you submitted but found that your app still violates the Payments policy.

  • Specifically, your app allows users to contribute donations to an organization that is not tax-exempt.

You can refer to the attached screenshot for additional information.

Image

Google Support have not explained why a 501(c)(6) determination is insufficient.

Note

501(c)(6) is a tax-exempt status; donations are not tax-deductible for the donor. Google's communications explicitly state "tax-exempt".

What we will do

We are forced to remove donation links immediately from our Play Store build in 2.24.X, and are doing so under protest, as otherwise we would be unable to distribute AnkiDroid to the majority of our users who use the Play Store.

Who are we

AnkiDroid is the Anki client for Android, developed and maintained by volunteers. Nothing is sold in the app and our Open Collective is the sole source of funding for maintenance and development. AnkiDroid is independent from Anki, AnkiWeb, AnkiMobile, and AnkiHub.

All funding goes to our Open Collective, where our ledger is public.

Relevant links

Timeline

  • 2026-07-20: Initial notice [automated]

    Translated from Japanese

    AnkiDroid Open Source Team Developers

    After reviewing your app, AnkiDroid Flashcards (package name com.ichi2.anki), we have found that it does not comply with one or more Google Play policies. Therefore, your app's status is affected. Status: Further Action Required If you do not fix the issue by the following deadline, users in some countries/regions may not be able to use your app.

    Publishing Status

    Status: Further Action Required

    If you do not fix the issue by the following deadline, users in some countries/regions may not be able to use your app.

    For more information about this issue and how to fix it, please refer to the Google Play Console.

    Go to Google Play Console


    An issue was found. Violation of Payment Policy

    Your app contains content that does not comply with the Payment Policy.

    Your app will be published to Google Play users in India and Russia starting August 03, 2026. If this issue is not resolved, the app will be removed from Google Play in other countries/regions.

    Details of the Issue

    The following issues were found:

    • In-app experience: Please see the attached screenshot IN_APP_EXPERIENCE-5899.png

    Here's how to resolve this issue:

    • If your donation is to a tax-exempt organization, please provide documentation proving your organization is tax-exempt (e.g., an Internal Revenue Service (IRS) decision for a U.S. corporation) in your reply to the email you receive after creating your appeal case.
    • If your organization is not tax-exempt, you must do one of the following:
    1. Remove the donation feature from your app.
    2. Collect donations using Google Play's in-app purchase system.
    3. Enroll in the U.S. Alternative Payment System program (if you plan to offer only in-app payment options to U.S. users).
    4. Enroll in Google's External Content Links program (if you plan to direct U.S. users outside your app for promotional purposes, etc.).
    • Submit your changes to Google for review. Go to [Publication Summary].

    Your app may be subject to country-specific requirements. Also, in certain countries (such as India), you may be able to offer your app for a limited time without making changes by paying a specified service fee. You can also change the countries or regions where your app is distributed by following the steps in the Google Play Console Help Center article "Distributing your app release to specific countries."

    Regarding Payment Policies

    For more information, please refer to the Payment Policies and Google Play Payment Policies pages in the Help Center. For more information about alternative billing options available in specific countries or regions, please refer to this FAQ.

    Submit an Appeal

    If you believe Google's decision was incorrect, you can submit an appeal. It may take up to 7 days to receive a response (in exceptional cases, it may take longer).

    Submit an Appeal

    Once you have completed fixing the issue, submit your app changes for review on the Google Play Console Publication Summary page.

    More Details

    For more information about the enforcement process, please refer to the Help Center article "Enforcement Process." For the latest information on Google Play policies, please see #PolicyBytes in the Android for Developers section. You can change the audio track in the settings to select your preferred language.

    Thank you for your cooperation in our efforts to make Google Play a great experience for developers and users.

    Details

    AnkiDroid Open Source Team デベロッパー各位

    お客様のアプリ AnkiDroid Flashcards(パッケージ名 com.ichi2.anki)を審査した結果、Google Play の 1 つ以上のポリシーに準拠していないことが判明いたしました。このため、アプリのステータスが影響を受けています。ステータス: さらなる対応が必要以下の期日までに問題を修正していただけない場合、一部の国 / 地域のユーザーはお客様のアプリを利用できなくなる可能性があります。

    Publishing Status

    ステータス: さらなる対応が必要

    以下の期日までに問題を修正していただけない場合、一部の国 / 地域のユーザーはお客様のアプリを利用できなくなる可能性があり ます。

    この問題とその修正方法について詳しくは、Google Play Console をご参照ください。

    Google Play Console に移動


    問題が見つかりました。 お支払いに関するポリシーへの違反

    お客様のアプリには、 お支払いに関するポリシーに準拠していないコンテンツが含まれて います。

    お客様のアプリは、 August 03, 2026 よりインド, ロシアの Google Play ユーザーに公開されます。この問題を修正していただけない場合、 アプリはその他の国 / 地域の Google Play から削除されます。

    問題の詳細

    次の項目で問題が見つかりました。

    • アプリ内エクスペリエンス: 添付のスクリーンショット IN_APP_EXPERIENCE-5899.png をご覧ください

    この問題を解決する方法は次のとおりです

    • 非課税の対象となる組織への寄付である場合は、 非課税団体であることを証明できる資料(米国内の法人に対する Internal Revenue Service(IRS)の決定書)を、 再審査請求 ケースの作成 後に届くメールへの返信にてご提供ください。

    • 非課税団体ではない組織の場合は、 次のいずれかを行っていただく必要があります。

      1. アプリから寄付の機能を削除する
      2. Google Play のアプリ内課金システムを使用して寄付金を集める
      3. 米国の代替の課金システム プログラム に登録する( 米国のユーザーにアプリ内決済オプションのみを提供する予定があ る場合)
      4. Google の 外部コンテンツ リンク プログラム に登録する(特典のプロモーションなどの目的で、 米国のユーザーをアプリ外に誘導する予定がある場合)
    • 変更内容を審査のために Google に送信します。[ 公開の概要] に移動します。

    アプリには国別の要件が適用される場合があります。また、 特定の国(インドなど)では、 所定のサービス料金を支払うことで、 変更を加えずにアプリを一定の期間提供できる場合があります。 Google Play Console ヘルプセンター記事「特定の国にアプリのリリースを配信する」 の 手順に沿って、 アプリを配信する国や地域を変更することもできます。

    お支払いに関するポリシー について

    詳しくは、ヘルプセンターの お支払い に関するポリシー、および Google Play のお支払いに関するポリシーについて のページをご参照ください。 特定の国や地域で利用できる代替の課金オプションについて詳しく は、こちらの よくある質問 をご参照ください。

    再審査請求を送信する

    Google の決定に誤りがあると思われる場合は、再審査請求を送信できます。回答を受け取るまでに 7 日ほどかかることがあります(例外的にもっとかかる場合もあります)。

    再審査請求を送信する

    この問題の修正が完了しましたら、Google Play Console の 公開の概要 ページで、アプリの変更を審査のために送信してください。

    詳細

    違反措置の適用プロセスについて詳しくは、ヘルプセンター記事「違反措置の適用プロセス」をご参照ください。Google Play のポリシーに関する最新情報については、デベロッパー向け Android の #PolicyBytes をご覧ください。設定で音声トラックを変更すると、ご希望の言語をお選びいただけます。

    デベロッパーとユーザーの皆様に Google Play を快適にご利用いただくための取り組みにご協力いただき、ありがとうございます。

  • 2026-07-20 - [informational] AnkiDroid's request to Open Source Collective for IRS determination letter

    Hello there -

    I am the administrator of a collective that uses Open Source Collective as a fiscal host - https://opencollective.com/dashboard/ankidroid - and we have an app (AnkiDroid) that is listed in the Google Play Store (https://play.google.com/store/apps/details?id=com.ichi2.anki).

    That app has a UI element with a link to our Open Collective page for donations. Google Play Store notified us that we must present proof of tax-exempt status in order to maintain that link - or in the absence of proof we will be removed from the store until the link is removed.

    I believe we should be able to prove tax-exempt status, but we if we could have formal proof from you all to deliver to them, I believe that would help.

    This is what they say specifically in the "How to fix" area that I believe applies to us:

    If the donations are for an eligible tax-exempt organization, please provide verifiable documentation that indicates the organization’s tax-exempt status (for example, Internal Revenue Service determination letter for entities in the United States) to the email you receive after the appeal case has been created.

    Can you send me an IRS determination letter or similar verifiable documentation showing tax-exempt status for Open Source Collective, such that I may open an appeal with a good chance of success with the Play Store?

    Thanks -

    -Mike

  • 2026-07-21: Open Source Collective determination of 501(c)(6) status; provided to Google

    To whom it may concern,

    This letter is to confirm that Ankidroid is a member collective (project) of Open Source
    Collective, a 501(c)(6) non-profit organization, registered in the state of California.

    Donations to Ankidroid made through the Open Collective platform at
    https://opencollective.com/ankidroid are received and held by Open Source Collective and
    are exempt from tax.

    I have attached a copy of our determination.

    Benjamin Nickolls
    President


    Dear Applicant:

    We're pleased to tell you we determined you're exempt from federal income tax under Internal Revenue Code (IRC) Section 501(c) (6). This letter could help resolve questions on your exempt status. Please keep it for your records.

    If we indicated at the top of this letter that you're required to file Form 990/990-EZ/990-N, our records show you're required to file an annual information return (Form 990 or Form 990-EZ) or electronic notice (Form 990-N, the e-Postcard). If you don't file a required return or notice for three consecutive years, your exempt status will be automatically revoked.

    If we indicated at the top of this letter that an addendum applies, the enclosed addendum is an integral part of this letter.

    For important information about your responsibilities as a tax-exempt organization, go to www.irs.gov/charities. Enter "4221-NC" in the search bar to view Publication 4221-NC, Compliance Guide for Tax-Exempt Organizations (Other than 501(c) (3) Public Charities and Private Foundations), which describes your recordkeeping, reporting, and disclosure requirements.

    We sent a copy of this letter to your representative as indicated in your power of attorney.

  • 2026-07-22: Google Reply [we'll contact you as soon as we have more information to share]

    Hi Mike,

    Thanks for your reply.

    Please kindly take note that all future communication regarding the Payments policy issue of your app, AnkiDroid Flashcards (com.ichi2.anki), will be done via this current ticket ( #9-2777000041594 ).

    We're currently looking into your appeal and we'll contact you as soon as we have more information to share.

    Thanks for waiting while we look into your app.
    Regards,
    [Name Redacted]
    The Google Play Team

    Please visit the Google Play Developer Policy Center and Google Play's Academy for App Success to learn more about building policy compliant and high quality apps. You can also visit the Android Developers Blog for the latest Android and Google Play news for app and game developers.

  • 2026-08-02: AnkiDroid follow-up

    Hello there!

    I just wanted to note that the google play store still has a banner on our account indicating that this case will result in our app having restricted listing starting tomorrow Aug 3rd, but we haven't heard back.

    We believe, subject to your review of course, that our app falls squarely into the non-profit regime that allows for a donations link, and we have provided evidence of same, which is of course this very appeal.

    It is our hope that while the appeal is in process in good faith with everyone involved, that the app store listing won't be restricted. Can you confirm that is the case while we are on appeal, or will the listing be restricted tomorrow?

    Thanks -

    -Mike

  • 2026-08-03: Google Response [thank you for waiting]

    Hi Mike,

    Thanks for reaching out to us.

    We're still looking into your appeal and we'll contact you as soon as we have more information to share.

    Thanks again for waiting while we look into your app.
    Regards,
    [Redacted]
    The Google Play Team

  • 2026-08-06: Google Reply: Your app still violates Google Play policy

    Hi Mike,

    Thanks for your patience.

    We reviewed the documentation you have submitted. However, we found that your app still violates Google Play policy.

    Please resolve the issue described below within the timeframe displayed in your Play Console account or your app may be removed from or distribution limited on Google Play.

    Step 1: Fix the policy violation with your app

    During our review, we found that your app violates the Payments policy. Specifically, Google Play's billing system must not be used in cases where payments include tax exempt donations. You can read through the Payments policy page for more details.

    For example, your app currently leads users to make donations through a payment system other than Google Play's billing system.

    Donations may only be collected within an app under certain conditions:

    • The donations are for a validated tax-exempt organization (for example, a validated 501(c)(3) charitable organization in the United States or the local equivalent), and
    • The donations are collected through a secure payment system.

    You can learn more about in-app billing in the Android Developers Help Center. Note that all of these conditions need to be fulfilled for us to reinstate your app.

    Please update your app to fix this issue.

    You may also want to double check that your app complies with all other policies listed in the Developer Policy Center as additional enforcement could occur if there are further policy violations.

    Eligible developers have the following options in select countries/regions:

    • Offer an alternative billing system alongside Google Play’s billing system as part of our user choice billing pilot. Please visit our FAQ for more details.
    • Offer alternative billing without user choice to their users in the European Economic Area (EEA). Please visit our FAQ for more details.
    • Offer alternative billing to their users in South Korea, Please visit our FAQ for more details.
    • Offer alternative billing to their users in India, Please visit our FAQ for more details.
    • Enroll in our external offers program to lead users in the European Economic Area (EEA) outside the app, including to promote offers. Please visit our FAQ for more details.

    Step 2: Submit your proof of eligibility and/or update your app

    If the donations are for a validated tax-exempt organization, please reply to this email and attach proof of the organization’s tax-exempt status (for example, Internal Revenue Service determination letter for entities in the United States).

    Alternatively, if the organization is not a tax-exempt organization, you must remove the donation functionality from your app, or use Google Play's billing system when collecting donations.

    To submit an updated app bundle or APK:

    • Prepare your updates.
    • Create a new release using the compliant app bundle or APK. Be sure to create the new release on the same track(s) as the noncompliant app bundle or APK, increment the version number, and set the release to 100% rollout.
    • Follow the on-screen instructions to add APKs or app bundles, then review and roll out your release.
    • Remember that your app may be subject to country-specific requirements. You can make changes to the in-app experience to comply with these requirements, and you can change which countries your app is distributed in by following the instructions in this Help Center article.

    Kindly note that your changes aren't sent for review automatically. You must go to the Publishing overview page and click Send for review to submit your changes.

    If you are located in the EU, you may have additional redress options. Learn more about those potential options in the EU Out-of-Court Dispute Resolution Help Center. Routing ID: ZLFS

    Please let us know if you have any other questions. Thanks for working with us to fix the policy issue and for your continued support of Google Play.
    Regards,
    [Redacted]
    The Google Play Team

    Please visit the Google Play Developer Policy Center and Google Play's Academy for App Success to learn more about building policy compliant and high quality apps. You can also visit the Android Developers Blog to learn more about the latest Android and Google Play news for app and game developers

  • 2026-08-07: AnkiDroid reply requesting clarification: why is 501(c)(6) documentation insufficient

    Hi [Redacted] -

    I must admit I am having a difficult time understanding this determination.

    Looking at the exact wording in your communication I note these two items:

    On 2026.08.06 7:10 PM, googleplay-developer-support@google.com wrote:

    Step 1: Fix the policy violation with your app

    • During our review, we found that your app violates the Payments policy. Specifically, Google Play's billing system must not be used in cases where payments include tax exempt donations. You can read through the Payments policy page for more details.
    • For example, your app currently leads users to make donations through a payment system other than Google Play's billing system.
    • Donations may only be collected within an app under certain conditions:
    • The donations are for a validated tax-exempt organization (for example, a validated 501(c)(3) charitable organization in the United States or the local equivalent), and
      AnkiDroid is in fact part of a validated tax-exempt organization, specifically it is a 501(c)(6) non-profit and we attached proof of same. Did you not receive that? Or is there something missing in the documentation we provided?
      The donations are collected through a secure payment system.
      Donations are indeed collected solely on https://opencollective.com/ankidroid which collects payments using secure systems

    So, I guess I'm confused. We seem to be the exact case that the policy is designed to allow. We have always done our best to follow Play Store policy and we certainly intended to with our donation link.

    Can you please confirm that you received the tax-exempt organization documentation and/or what's missing if it's not sufficient?

    Many thanks -

    -Mike

  • 2026-08-07 Google Reply: [we reviewed the documentation you submitted but found that your app still violates the Payments policy]

    Hi Mike,

    Thanks for your reply.

    As mentioned in the previous email, we reviewed the documentation you submitted but found that your app still violates the Payments policy.

    • Specifically, your app allows users to contribute donations to an organization that is not tax-exempt.

    You can refer to the attached screenshot for additional information.

    Donations can only be contributed to tax-exempt organizations, and any organizations or entities that are not verifiable as tax-exempt are not allowed to collect donations through a payment method other than Google Play's billing system.

    You may reply to this email and attach documentation that demonstrates the entity receiving donations in your app is tax-exempt. Otherwise, you may remove the donation functionality from your app or use Google Play' billing system when collecting donations.

    You may refer to my previous email for more details.

    Thanks for your continued support of Google Play.
    Regards,
    [Name Redacted]

  • 2026-08-27: AnkiDroid 2.25.0alpha3 submitted to Play Store

  • 2026-08-28 App status: Rejected

    Translated from Japanese

    To all AnkiDroid Open Source Team developers

    After reviewing your app, AnkiDroid Flashcards (package name com.ichi2.anki), we found that it does not comply with one or more Google Play policies. Therefore, your app's status has been affected.

    Publishing Status

    App status: Rejected

    Your app changes were not published due to an issue with the following policy. If you have an older version of your app, it will continue to be published on Google Play.

    For more information about this issue and how to fix it, please refer to the Google Play Console.

    Go to Google Play Console

    A problem has been found. Violation of payment policy.

    Your app contains content that does not comply with our payment policies.

    Your app is directing users to payment methods other than the Google Play billing system.

    Problem details

    Problems were found in the following items.

    Version Code 122400300: In-App Experience : See attached screenshot IN_APP_EXPERIENCE-2431.png

    Here's how to solve this problem:

    Please remove any links that direct users to make payments using a system other than Google Play's billing system.

    If you plan to direct U.S. users outside of your app for promotional purposes, such as offering special offers, please register for Google's External Content Links Program. For more information, please see our FAQ.

    Submit your changes to Google for review. Go to [ Publish Summary ].

    Apps may be subject to country-specific requirements. Additionally, in certain countries (such as India), you may be able to offer your app according to applicable conditions. You can also change the country or region where your app is distributed by following the steps in this article in the Google Play Console Help Center .

    About our payment policy

    For more information, please see the payment policies in the Help Center and the Google Play payment policies page. For more information about alternative billing options and external linking options available in some countries and regions, please see our FAQ .

    Submit a request for reconsideration

    If you believe Google's decision was incorrect, you can submit an appeal. It may take up to 7 days to receive a response (in exceptional cases, it may take longer).

    Submit a request for reconsideration

    Once you have finished fixing this issue, please submit the app changes for review on the Google Play Console's Publication Summary page.

    detail

    For more information on the enforcement process, see the Help Center article " Enforcement Process ." For the latest information on Google Play policies, see #PolicyBytes for Android Developers. You can change the audio track in Settings to select your preferred language.

    Thank you for your cooperation in our efforts to make Google Play a great experience for developers and users alike.

    Please answer the two questions in this survey . Your responses will help us improve Google Play services.

    Google Play Team

    Japanese Original

    ご対応のお願い: Google Play のポリシーにアプリが準拠していません - (AnkiDroid Flashcards) (Request for Action: The app does not comply with Google Play policies - (AnkiDroid Flashcards))

    AnkiDroid Open Source Team デベロッパー各位

    お客様のアプリ AnkiDroid Flashcards(パッケージ名 com.ichi2.anki)を審査した結果、Google Play の 1 つ以上のポリシーに準拠していないことが判明いたしました。このため、アプリのステータスが影響を受けています。

    Publishing Status

    アプリのステータス: 否承認

    お客様のアプリの変更は、以下のポリシーに関する問題により公開されませんでした。古いバージョンのアプリがある場合は、引き続き Google Play で公開されます。

    この問題とその修正方法について詳しくは、Google Play Console をご参照ください。

    Google Play Console に移動

    問題が見つかりました。 お支払いに関するポリシーへの違反

    お客様のアプリには、お支払いに関するポリシーに準拠していないコンテンツが含まれています。

    お客様のアプリは、Google Play の課金システム以外のお支払い方法にユーザーを誘導しています。

    問題の詳細

    次の項目で問題が見つかりました。

    バージョン コード 122400300: アプリ内エクスペリエンス: 添付のスクリーンショット IN_APP_EXPERIENCE-2431.png をご覧ください

    この問題を解決する方法は次のとおりです。

    Google Play の課金システム以外のシステムで支払いを行うようユーザーを誘導しているリンクを削除してください。

    または

    特典のプロモーションなどの目的で、米国のユーザーをアプリ外に誘導する予定がある場合は、Google の外部コンテンツ リンク プログラムにご登録ください。詳しくは、よくある質問をご参照ください。

    変更内容を審査のために Google に送信します。[公開の概要] に移動します。

    アプリには国別の要件が適用される場合があります。また、特定の国(インドなど)では、適用される条件に従ってアプリを提供できる場合があります。Google Play Console ヘルプセンターのこちらの記事の手順に沿って、アプリを配信する国または地域を変更することもできます。

    お支払いに関するポリシー について

    詳しくは、ヘルプセンターのお支払いに関するポリシー、および Google Play のお支払いに関するポリシーについてのページをご参照ください。一部の国や地域で利用できる代替の課金オプションと外部リンク オプションについて詳しくは、こちらのよくある質問をご参照ください。

    再審査請求を送信する

    Google の決定に誤りがあると思われる場合は、再審査請求を送信できます。回答を受け取るまでに 7 日ほどかかることがあります(例外的にもっとかかる場合もあります)。

    再審査請求を送信する

    この問題の修正が完了しましたら、Google Play Console の [公開の概要] ページで、アプリの変更を審査のために送信してください。

    詳細

    違反措置の適用プロセスについて詳しくは、ヘルプセンター記事「違反措置の適用プロセス」をご参照ください。Google Play のポリシーに関する最新情報については、デベロッパー向け Android の #PolicyBytes をご覧ください。設定で音声トラックを変更すると、ご希望の言語をお選びいただけます。

    デベロッパーとユーザーの皆様に Google Play を快適にご利用いただくための取り組みにご協力いただき、ありがとうございます。

    こちらのアンケートの 2 つの質問へのご回答をお願いいたします。いただいた回答は Google Play のサービス向上に役立てさせていただきます。

    Google Play チーム

    TODO: IN_APP_EXPERIENCE-2431.png

  • 2026-08-28: Status: Further action required [Automated Play Store Email]

    Translated from Japanese AnkiDroid Open Source Team Developers After reviewing your app, AnkiDroid Flashcards (package name com.ichi2.anki), we have found that it does not comply with one or more Google Play policies. Therefore, your app's status is affected.

    Publishing Status

    Status: Further Action Required

    If you do not fix the issue by the following deadline, your app may become unavailable to users in some countries/regions.

    For more information about this issue and how to fix it, please refer to the Google Play Console.

    Go to Google Play Console

    Issues Found: Violation of Payment Policies

    Your app contains content that does not comply with payment policies.

    Your app directs users to payment methods other than the Google Play billing system.

    Your app will be released to Google Play users in India and Russia on September 11, 2026. If you do not fix this issue, your app will be removed from Google Play in other countries/regions.

    Issue Details

    Issues were found in the following areas:

    Version Code 122400300: In-App Experience: See attached screenshot IN_APP_EXPERIENCE-4384.png

    Here's how to resolve this issue:

    Remove any links directing users to make payments using systems other than Google Play's billing system.

    Alternatively,

    If you plan to direct US users outside your app for promotional purposes, such as offering incentives, register for Google's External Content Links Program. See the FAQ for more information.

    Submit your changes to Google for review. Go to [Publish Summary].

    Your app may be subject to country-specific requirements. Also, in certain countries (such as India), you may be able to offer your app under applicable conditions. You can also change the country or region where your app is distributed by following the steps in this article in the Google Play Console Help Center.

    About Payment Policies

    For more information, see the Payment Policies page in the Help Center and the Google Play Payment Policies page. For more information about alternative billing and external linking options available in some countries and regions, see the FAQ.

    Submit a Review Request

    If you believe Google's decision was incorrect, you can submit a review request. It may take up to 7 days to receive a response (in exceptional cases, it may take longer).

    Submit a Review Request

    Once you have completed fixing this issue, submit your app changes for review on the [Publish Summary] page in the Google Play Console.

    Learn More

    For more information on the enforcement process, see the Help Center article "Enforcement Process." For the latest information on Google Play policies, see #PolicyBytes for Android Developers. You can change the audio track in Settings to select your preferred language.

    Thank you for helping us make Google Play a great experience for developers and users.

    Please answer the two questions in this survey. Your responses will help us improve Google Play.

    The Google Play Team

    Japanese Original AnkiDroid Open Source Team デベロッパー各位 お客様のアプリ AnkiDroid Flashcards(パッケージ名 com.ichi2.anki)を審査した結果、Google Play の 1 つ以上のポリシーに準拠していないことが判明いたしました。このため、アプリのステータスが影響を受けています。

    Publishing Status

    ステータス: さらなる対応が必要

    以下の期日までに問題を修正していただけない場合、一部の国 / 地域のユーザーはお客様のアプリを利用できなくなる可能性があります。

    この問題とその修正方法について詳しくは、Google Play Console をご参照ください。

    Google Play Console に移動

    問題が見つかりました。 お支払いに関するポリシーへの違反

    お客様のアプリには、お支払いに関するポリシーに準拠していないコンテンツが含まれています。

    お客様のアプリは、Google Play の課金システム以外のお支払い方法にユーザーを誘導しています。

    お客様のアプリは、September 11, 2026よりインド, ロシアの Google Play ユーザーに公開されます。この問題を修正していただけない場合、アプリはその他の国 / 地域の Google Play から削除されます。

    問題の詳細

    次の項目で問題が見つかりました。

    バージョン コード 122400300: アプリ内エクスペリエンス: 添付のスクリーンショット IN_APP_EXPERIENCE-4384.png をご覧ください

    この問題を解決する方法は次のとおりです。

    Google Play の課金システム以外のシステムで支払いを行うようユーザーを誘導しているリンクを削除してください。

    または

    特典のプロモーションなどの目的で、米国のユーザーをアプリ外に誘導する予定がある場合は、Google の外部コンテンツ リンク プログラムにご登録ください。詳しくは、よくある質問をご参照ください。

    変更内容を審査のために Google に送信します。[公開の概要] に移動します。

    アプリには国別の要件が適用される場合があります。また、特定の国(インドなど)では、適用される条件に従ってアプリを提供できる場合があります。Google Play Console ヘルプセンターのこちらの記事の手順に沿って、アプリを配信する国または地域を変更することもできます。

    お支払いに関するポリシー について

    詳しくは、ヘルプセンターのお支払いに関するポリシー、および Google Play のお支払いに関するポリシーについてのページをご参照ください。一部の国や地域で利用できる代替の課金オプションと外部リンク オプションについて詳しくは、こちらのよくある質問をご参照ください。

    再審査請求を送信する

    Google の決定に誤りがあると思われる場合は、再審査請求を送信できます。回答を受け取るまでに 7 日ほどかかることがあります(例外的にもっとかかる場合もあります)。

    再審査請求を送信する

    この問題の修正が完了しましたら、Google Play Console の [公開の概要] ページで、アプリの変更を審査のために送信してください。

    詳細

    違反措置の適用プロセスについて詳しくは、ヘルプセンター記事「違反措置の適用プロセス」をご参照ください。Google Play のポリシーに関する最新情報については、デベロッパー向け Android の #PolicyBytes をご覧ください。設定で音声トラックを変更すると、ご希望の言語をお選びいただけます。

    デベロッパーとユーザーの皆様に Google Play を快適にご利用いただくための取り組みにご協力いただき、ありがとうございます。

    こちらのアンケートの 2 つの質問へのご回答をお願いいたします。いただいた回答は Google Play のサービス向上に役立てさせていただきます。

    Google Play チーム
    Play academy

The Daily Front Page 3 of 28
Tuesday, September 1, 2026 The Daily Front No. #260901 — Claude’s New Guardrails
article

Claude Fable 5.1 and Claude Mythos 5.1

by denysvitali·▲ 1,050 points·977 comments·anthropic.com ↗
Claude Fable 5.1 and Claude Mythos 5.1 are the same model, but with different levels of safeguards.

We’re introducing Claude Fable 5.1 and Claude Mythos 5.1. They’re the world’s most advanced models for coding and knowledge work—and their research capabilities offer an early glimpse of how AI models will contribute to scientific progress.

Claude Fable 5.1 and Claude Mythos 5.1 are the same model, but with different levels of safeguards. Fable 5.1 is generally available, while Mythos 5.1 is available only through our trusted access programs; its safeguards are specifically designed to support work in cybersecurity and the life sciences.

Alongside its increased capabilities, Fable 5.1 takes important steps towards addressing the feedback we’ve received from customers on price, data retention, and safeguards.

Price.** Fable 5.1 will cost an estimated 25% less than Fable 5 for typical workloads, wherever usage is billed by token. This is because we’re reducing our pricing on cache reads (where the model reads inputs that have already been processed and stored). For highly agentic work, the savings will often be much larger—up to approximately 45%.

Data retention.** Our new system of Enterprise Frontier Safeguards (EFS) gives customers complete privacy (the same as a zero data retention policy) while still being state-of-the-art at preventing adversarial use. EFS works by storing data in cloud infrastructure controlled entirely by the customer, not Anthropic. It will be made available to enterprise customers in phases, beginning later this fall. Until EFS is available, eligible customers will be able to use Fable 5.1 with zero data retention.

Safeguards.** We’ve improved our safeguards to reduce false positives (where the system flags benign content). In cybersecurity, our newest safeguards block 60% fewer false positives than before. In part, this is because Fable 5.1 can now be used to discover software vulnerabilities—though not to develop exploits for them. In biology, we’ve established an access program, developed in partnership with the US government, to enable access to Claude Mythos 5.1’s advanced biology capabilities. We expect to open enrollment for scientists soon.

A new performance frontier

Claude Fable 5.1 sets a new standard for coding, knowledge work, and long-running problem-solving tasks. The charts below show that Fable 5.1 is capable of much higher performance than its predecessor, Fable 5. And when set to Low or Medium effort, Fable 5.1 achieves results similar to or better than Fable 5’s at a much lower cost. (Note that Fable 5.1 defaults to High effort in Claude Code, and to Medium in Claude Cowork and on Claude.ai.)

Terminal-Bench-Science 0.1: The standard error is ±3.5–4.5 pts per model. The public leaderboard (3 trials/task, Claude Code harness) reports Claude Opus 5 at 30.0% and Claude Fable 5 at 21.4%; our setup reproduces them at 29.0% and 24.7%, respectively, both within noise.

Terminal-Bench 4.0 scores by cost (log scale), at each effort level. Claude Fable 5.1 and Claude Mythos 5.1 are the same underlying model; the gap between them reflects the tasks on which our earlier, less precise cyber safeguards intervened. With the improvements we’re making to these safeguards today, we expect the difference between the models to be much smaller.

Humanity’s Last Exam scores by cost (log scale), at each effort level.

CursorBench 3.2.0 by cost (log scale), at each effort level.

Fable 5.1 avoids shortcuts that result in poorer-quality work, and it’s smart enough to fix the root causes of software issues. For example, in testing by the investment firm Millennium, Fable 5.1 found the cause of a rare crash in its internal systems that none of its engineers (or any other model) had been able to explain after several years of trying.

Here, you can see how Fable 5.1 compares across various benchmarks:

Fable 5.1 Fable 5 Opus 5 GPT-5.6 Sol
Agentic scientific research — Terminal-Bench-Science 0.1 [1] 52.6% 24.7% 29.0% 22.4%
Agentic coding — Terminal-Bench 4.0 55.8% 60.9% (Mythos 5.1) 42.0% 52.3%
Knowledge work — GDPval-AA v2 1853 1723 1824 1711
Computer use — OSWorld 2.0 [2], partial 77.9% partial 72.9% partial
Computer use — OSWorld 2.0, strict 41.7% strict 36.1% strict
Multidisciplinary reasoning — Humanity’s Last Exam, no tools 60.9% 57.8% 56.6% no tools
Multidisciplinary reasoning — Humanity’s Last Exam, with tools 65.0% 63.8% 63.6% with tools
Business workflows — AutomationBench 31.4% 17.1% 26.9% 19.6%
Agentic coding — CursorBench 3.2.0 73.4% 70.5% 70.0% 67.2%

Fable 5.1 was evaluated with its production safeguards enabled. On tasks where these safeguards intervened, Fable 5.1 and Fable 5 scored a zero on OSWorld 2.0, and Fable 5 scored a zero on AutomationBench. In all other interventions from our safeguards, cybersecurity tasks were completed by Claude Opus 4.8, and biology tasks were completed by Claude Opus 5. This likely reduces the performance of Fable 5.1 and Fable 5 on these benchmarks.

Our early-access partners noticed these performance upgrades, and also picked up on more qualitative improvements in the model’s outputs. Here’s what they told us:

“In internal benchmarks, Claude Fable 5.1 solves more of our coding problems than Fable 5 or Opus 5, and achieves state of the art on trading intuition. While prior models became hard to follow the longer they worked, Fable 5.1 remains readable over long, multi-step tasks.”

“We’re moving our Opus 5 traffic in Devin to Claude Fable 5.1 on launch day. It matched or edged out Fable 5 in our testing at a lower cost per task, and with the new cache read pricing a Fable-class model is finally economical for the workloads we’d kept on Opus, starting with code review.”

“A particular piece of code had an extremely rare crash, about one in a million runs, that nobody on our team had explained in four to five years. Every model I tried, including Fable 5, missed it. Claude Fable 5.1 was the first to find it. It disassembled an external vendor library, matched it against the core dump, and traced the crash to a bug in that library. The time it would have taken to conduct that analysis is hard to justify.”

“Claude Fable 5.1 built a complex prototype in about three days. It did initial research across all of our services code and documentation to produce a novel and extensible design. It then ran for hours unattended, with strong verification loops, to implement the entire prototype. I would wake up in the morning to the next phase finished, with a full visual walkthrough of what it built and clear evidence of its success.”

“It’s friendly Fable. Fable-level intelligence, Opus-level price, Sonnet-speed. In our tests it was about twice as fast as Opus 5 and used half as many tokens, so for anyone used to using Opus as their daily driver it’s an obvious upgrade.”

“On our research suite, Claude Fable 5.1 set new best scores. On one task it came up with a novel solution along a completely different axis than we’d seen from other models or from human researchers in the past, which took its results well above the previous plateau. It's better at creative problem solving and getting that flash of insight you need to solve a difficult problem.”

“As part of our ongoing evaluation of AI models, Claude Fable 5.1 delivered impressive results in our tests. Using Claude Code, it correctly identified the root cause of every broken build we tested, across all the effort levels. It also communicates more effectively than earlier Anthropic models, with updates that are more concise and easier to follow.”

“We asked Claude Fable 5.1 to review a clinical research project for Rakuten Medical that three other frontier models had signed off on. It found a gap none of them had seen and insisted on testing it further. It then proposed a completely new hypothesis, turning a dataset we had written off into a new research direction in one afternoon. It’s the first time a frontier model like Claude has empowered us to explore new research in this way.”

“Claude Fable 5.1 is very smart. On our 30-day simulated run-a-business eval, where the model gets full access to simulated Square tools, customers, employees, and vendors, it was far more efficient per token than Opus 5. We plan to use it to work through our most complex scenarios, the kind that used to take days of whiteboarding, so our engineering teams can keep moving fast.”

“For anything research, greenfield or long-horizon, I would absolutely use Claude Fable 5.1 as the orchestrator. One unattended 38-hour run on a machine learning problem diagnosed a prior result as a label artifact, made the correction, kicked off six parallel experiments that ran overnight, and returned with a result and next steps. Given an open-ended prompt to find the highest-leverage problem nobody owned, it surfaced an unowned alert tied to a production outage, pulled the logs and prescribed the fix.”

“The standout in Claude Fable 5.1 is the writing: more understandable, more meaningful, and it follows our writing guidance better. In blind tests against Fable 5, I preferred its writing and output. And in Canva Code it built a rhythm game with real music and on-beat gameplay matched to the level it generated, something no other model we tested delivered.”

“On our PowerPoint eval, Claude Fable 5.1 produced the best decks of any model we’ve tested, both in slide craft and in fully answering our research topic. That same completeness showed up in our Citations eval, where it had the best fact recall over financial documents. And on complex, multi-part questions, it’s the first model we’ve seen answer every part.”

“We had a change that touched more than eight services across three codebases. Claude Fable 5.1 mapped the whole workflow end to end, in extremely fine detail, from the incoming service call down to the individual function and the database tables and rows, and it was accurate all the way down. We appreciated the opportunity to test the model and provide feedback, helping us prepare for a new frontier where we can increasingly rely on these tools to take on bolder initiatives.”

“Across our evaluation sets, our judges preferred Claude Fable 5.1’s answers roughly 2-to-1 over Fable 5 on everyday knowledge questions, high-intent research, and drafting and artifact work, all with improved response grounding over Fable. These results have made Fable 5.1 our go-to recommendation wherever Fable 5 was previously the choice.”

“On the hardest problems we work on, Claude Fable 5.1 separates strongly from any other model we’ve tried. On a grand challenge-tier problem we‘ve used as a testbed for 18 months, it actually produced material progress. Rather than being trapped in stamp collecting, it made clear white-space connections I have yet to see elsewhere. It also optimized a compute kernel that Fable 5 had tapped out on by about 35%.”

“On our hardest browser-agent benchmark, Claude Fable 5.1 completed 82% of tasks in about 10 minutes each, against 74% for Opus 5 and 57% for Fable 5, while using fewer tokens than either. It feels stronger than Fable 5 in every dimension we test. It also did exactly the right amount of work: never crossed a critical stop point across hundreds of measured tasks.”

“On our internal Finance benchmark, Claude Fable 5.1 matches Fable 5 on accuracy while using 20% fewer tokens. We also saw big gains in slide generation, with Fable 5.1 showing improvements in explaining complex data in simpler English and translating it into more banker-grade visuals.”

“Claude Fable 5.1 is more comfortable with long, unattended work than Fable 5. I’ve had workflows run for a long stretch without losing the plot: it keeps its own records, reprioritizes as things change, and picks up where it left off.”

“Compared to Fable 5, Claude Fable 5.1 was a massive improvement on RedlineBench, our contract redlining benchmark, improving from 47.9 to 57.0. Most of the gain came on first-turn quality, where it doubled the previous score, with substantial gains on counterparty acceptance as well. Its edits were also more concise, with smaller changes on average to the documents.”

“Claude Fable 5.1 is a leading model for our incident investigation evals, which use real production incidents to assess how effectively our agent, Bits Investigation, can produce root cause analyses. We evaluate our agent’s output against root causes identified by our engineers. In these evaluations, it has demonstrated stronger reasoning than Opus 5 and has successfully diagnosed the most complex production incidents we’ve tested.”

“Claude Fable 5.1 is the most capable model we’ve run on CursorBench 3.2, scoring 73.4% at max effort. We found it especially skilled at verifying its own work, allowing it to take on difficult coding tasks from start to finish.”

“We evaluate models and systems on real-world investor workflows. On FrontierFinance, our latest finance benchmark, Claude Fable 5.1 shows a clear gain over Fable 5, with a 55.9% rubric score compared to 49.2%. The gain comes from its ability to dig harder into grounded, authoritative sources: on one earnings question, it went straight to the call transcript and captured the exact figures management cited, where other models leaned on secondary coverage.”

Scientific research

We tested the scientific research capabilities of Claude Fable 5.1 and Claude Mythos 5.1 across a wide range of domains. What we found—which includes the early examples we share below—adds to the evidence that AI models will soon make important contributions to scientific discovery.

Molecular design.** Many modern medicines work by binding to targets within the body to block, activate, or deliver something to them. High-affinity binders are necessary for drugs to work at lower doses; designing one is the first step in the development process for many common drug modalities. To see how well Claude Mythos 5.1 could do at this task, we gave the model access to open-source protein design and folding tools and sent its designs to two external organizations for experimental validation. Mythos 5.1 proved able to design very high-affinity binders. On three targets, [3] its binding affinities were 10 times higher than the best designs submitted to Adaptyv Bio’s protein design competitions. Its hit rate (that is, the number of designs that were viable binders) was the strongest we’ve measured to date: it reached nearly 50% across 12 targets. (Hit rates of 10–15% are typical in protein design today.)

Claude-designed protein binders (orange) for each of 12 targets (grey). Every design in the video was confirmed to bind in the lab. Structures shown are ESMFold2 predictions.

Computational analysis and modeling. Claude Fable 5.1 trained a neural network to create a new, high-resolution elevation map of a third of the planet Venus. Its work was based on radar images taken by NASA’s Magellan mission more than 30 years ago and a map that already existed for one-fifth of the planet. Claude’s new map now reveals details down to two to three kilometers, rather than 10 to 20, and shows heights up to 25% more accurately than before.

We’re releasing this map under a Creative Commons license in advance of upcoming NASA VERITAS and ESA EnVision missions, in hopes that it might help them determine which geologic features to target for future observation.

Radar image (Magellan): bright cone, radian lava flows

Magellan radar

Altimetry 10-20km footprint

Altimetry 10-20km footprint

New DEM (300m) a volcano 15km across

New DEM (300m) a volcano 15km across

A small shield volcano on Venus

Computational biology.** In computational biology, it’s common to run task-specific machine learning models on GPUs. The speed of these models is therefore a bottleneck to research progress. Mythos 5.1 provided one solution to this problem: by writing custom GPU kernels and caching their intermediate results, it sped up seven open-source deep learning models by up to 2.5 times (with identical outputs).

The benefits of such speed-ups accumulate quickly. In any given experiment, biologists might run these models thousands of times (for example, testing every possible mutation near every human gene). On analyses like these, the optimized models cut estimated GPU costs by 30–60%. This kind of optimization would normally take a team of performance engineers weeks, and is often unaffordable for academic labs. Mythos 5.1 was able to do it in just days, using the publicly available source code alone. We plan to open-source these optimizations soon.

Inference speedup for seven open-source protein and genomics models on an NVIDIA H100

Estimated GPU cost of three genome-wide analyses before and after optimization, at cloud list price. Evo 2 40B saves more on a whole job (2.3x) than per forward (1.4x) because some of its optimizations only pay off across many sequences.

As our models’ scientific capabilities improve, our investment in scientific progress is also growing. Last week, we previewed the Model Hardware Standard, which allows Claude to directly and safely operate laboratory equipment. We’ve also recently expanded our support for scientists through our AI for Science program, which provides free credits to researchers working on high-impact scientific projects, and we are offering steeply discounted usage through our new Claude Team plan for scientists.

Safety, security, and alignment

AI models’ agentic capabilities have become much more powerful over the past two years. But as we’ve documented, greater autonomy comes with new risks. Work on safety, security, and alignment needs to advance at the same pace as AI capabilities. Yesterday, we published a report describing how we are improving our own alignment and security efforts

Prior to releasing Claude Fable 5.1 and Claude Mythos 5.1, we (and, in some cases, external researchers) subjected the models to extensive testing for risks across many areas. We describe these efforts in full in our System Card; below is a brief summary.

Chemical and biological risks. We tested the extent to which Claude Mythos 5.1 could help create chemical or biological weapons. This involved expert red-teaming, automated evaluations, and a tabletop exercise that paired PhD-level biologists with AI experts, testing whether the models could match human specialists’ performance. Mythos 5.1’s capabilities are greater than those of Mythos 5. However, our evaluations indicate that it still falls short of the next risk tier defined in our Responsible Scaling Policy. We are therefore deploying Mythos 5.1 with the same safeguards that we applied to Mythos 5, which restrict access to research biology capabilities.

Cyber risks. We ran a suite of evaluations to assess the cyber capabilities of Claude Mythos 5.1 (with cybersecurity safeguards off). Overall, the model demonstrates the strongest cyber capabilities of any model we’ve released, though it still falls within the lower category of risk in our Frontier Compliance Framework. We also performed extensive stress-testing of our cybersecurity safeguards for Fable 5.1: in addition to our own dynamic evaluation of their robustness, we commissioned external testing from two organizations, along with automated testing by Gray Swan. As with Fable 5 and Opus 5, we have not found evidence of a critical-severity jailbreak for these safeguards.

Agentic safety. We ran evaluations of how Claude Mythos 5.1 responds to malicious requests and prompt injections (adversarial instructions hidden within content processed by AI models). It refused malicious agentic coding and computer use requests at a comparable rate to Mythos 5, Sonnet 5, and Opus 5, and it is our most robust model to date on an external prompt injection benchmark.

Alignment.** We tested the model’s behavior through static and interactive behavioral evaluations, analyses of its internal thinking using natural language autoencoders, misalignment-related capability evaluations, a review of our training data, and analyses of our internal pilot use. We also received reports from external testing.

Our automated behavioral audit found that Claude Mythos 5.1 is better aligned across most metrics than its predecessor, Mythos 5. The model is significantly less likely than Mythos 5 to try to access resources outside of its test environment when assigned an otherwise impossible task. It is also less likely than Mythos 5 to use motivated reasoning to justify its actions (for instance, by reasoning that the situation is a simulation or evaluation), and it is less likely to ignore explicit constraints in pursuit of users’ goals. From our review of its training data, Mythos 5.1 both attempts reward hacking (or cheating), and succeeds at it, at a lower overall rate than Mythos 5.

Though generally our alignment evaluations showed improvements, our testing found the model can still sometimes bypass approvals and auto-mode classifiers (as we discuss in more detail in our System Card). There are also limitations to the coverage provided by our alignment assessment. Currently, our automated behavioral audit provides less visibility into very long-context work and multi-agent settings. We also have less coverage of impossible tasks (which can elicit more abnormal and misaligned behavior) than we’d like, although we’ve recently made improvements in this domain and are working hard to continue doing so.

We have also improved our safeguards so that they allow our models to be more useful without compromising on safety. We describe these changes below.

Automated safeguards for enterprises.** Enterprise Frontier Safeguards (EFS) allows us to detect and respond to misuse of our models while still providing our enterprise customers the privacy of a zero data retention agreement. With EFS, customers store their data on their own cloud infrastructure, rather than on Anthropic’s systems; any human review is, by default, done by the customer themselves, rather than Anthropic. We developed EFS in close collaboration with more than 100 customers across industries like financial services, healthcare, manufacturing, telecom, law, retail, and the public sector, and with our cloud partners at Amazon Web Services, Google Cloud, and Microsoft Azure.

EFS will be supported on Claude Code, Claude Enterprise, the Claude Platform, Amazon Bedrock, Claude Platform on AWS, Google’s Agent Platform, and Microsoft Foundry. It’s rolling out in phases, starting this fall. As noted above, customers who are eligible for EFS can use Fable 5.1 (and Fable 5) with zero data retention until EFS is ready. You can read more about EFS here; to request access, please complete this form.

More precise safeguards for biology and cybersecurity.** In the past few months, we’ve made progress in making our safeguards for Fable 5.1 more precise: ensuring that they’re less likely to flag benign content (like queries about medical issues or cyberdefenders using the model to make their systems safer), but still ensuring they provide robust protection against genuine threats.

As we recently shared, our latest biology safeguards for Fable 5.1 and Fable 5 fire 85% less often for benign requests related to elementary biology and medical questions (relative to those that launched with Fable 5). However, queries related to research and development in the life sciences will still be directed to our Opus models. We’re making the model’s life sciences capabilities available to professionals through an access program for Claude Mythos 5.1 that we’ve developed in partnership with the US government, which we discuss below.

With Fable 5.1, we’re updating our cybersecurity safeguards to be more precise. We’re also now allowing Fable 5.1 to be used for identifying software vulnerabilities—that is, to conduct the kind of defensive work that improves software security. As a result of these changes, Claude Code users can expect an average of around 60% fewer interventions per session from our cyber safeguards, relative to the previous safeguards on Fable 5. Our safeguards do, however, still redirect several kinds of dual-use cybersecurity tasks (tasks that might have helpful or harmful applications) to our Opus models. This includes penetration testing, exploit generation, and binary-based vulnerability scanning.

Anti-distillation mechanisms.** Distillation is a method used to extract the capabilities of advanced models. It is often employed on an industrial scale, using thousands of fake accounts. Distillation is a safety risk, since the distilled capabilities can subsequently be released without adequate safeguards. Fable 5.1 comes with strengthened mechanisms to make distillation attacks harder. For example, it is no longer possible for new API accounts (those created from today onwards) to manually edit Claude’s prior context in a multi-turn conversation while preserving the transcript of Claude’s prior thinking. This closes off a common, publicly documented distillation technique, which allowed distillers to illicitly extract Claude’s thinking. We’re rolling out the change gradually, to minimize disruption: existing accounts are not currently affected by this change, though it will apply to all users with future model releases. A small number of customers’ custom integrations will then be affected. Our Help Center article explains more about this change and the adjustments that developers can make.

Trusted access for Claude Mythos 5.1

Claude Mythos 5.1 is identical to Fable 5.1, but it offers more permissive safeguards for vetted individuals and organizations whose work is affected by the cybersecurity and life sciences restrictions outlined above. It will be available through two trusted access programs:

  • Cyber Verification Program: The CVP currently provides access to certain Opus- and Sonnet-class models with reduced cyber safeguards for defensive security work. In the near future, this program will also include access to Claude Mythos-class models. Apply to join the CVP here.
  • Life Sciences Verification Program: The LSVP is designed so that life sciences professionals can use Claude Mythos 5.1 with safeguards designed for professional research and development activities (while all other safeguards remain in place). In partnership with the US government, we have enrolled our first participants, and we plan to expand access to this program to the broader life sciences community.

In addition to these trusted access programs, Claude Security, our product that scans codebases for vulnerabilities and suggests patches for human review, is now also powered by Claude Mythos 5.1.

Compliance with the EU AI Act

In July 2026, Anthropic (along with 190 other signatories, including several other major AI model providers) signed the EU AI Act’s Code of Practice on Transparency of AI-Generated Content.

This required us to add a watermark—a numerical way of determining the likelihood that Claude was involved in writing a piece of text—to the outputs of models released after August 2, 2026. As we recently explained, this watermark is invisible to anyone who does not have the detection API. It has no practical impact on the quality or content of Claude’s outputs and contains no information about the user, their organization, or their conversations with Claude.

The Act also required us to provide a way for users to tell whether a text likely contains the watermark. We are thus rolling out a detection API in private preview. It is currently available to eligible organizations (such as regulators, law enforcement, media, fact-checkers, independent researchers, educational organizations, and EU civil society groups) as required under EU law. It is also available for enterprises that are similarly obligated to verify watermarking for their own compliance with the Act. We plan to expand access to the detection API over time. You can register interest in access here.

Cost and availability

Claude Fable 5.1 is available today on all platforms, including Amazon Web Services, Google Cloud, and Microsoft Azure. Developers can get started with claude-fable-5-1 on the Claude API.

As mentioned above, we have reduced the price of Fable 5.1’s cache reads (where the model reuses context it has already processed) wherever usage is billed by token, such as on our API. Cache reads now cost 75% less, or $0.25 per million tokens.

This change leads to a substantial reduction in the overall cost of running the model. For typical workloads, costs are reduced by around 25% relative to Fable 5. For complex coding and highly agentic tasks, the savings could be up to around 45%. The graph below illustrates why this change makes such a big difference:

Indexed cost of running the same workloads on Fable 5 and Fable 5.1, at usage-based pricing measured at default effort over four weeks of actual usage in August 2026. Typical workload covers Fable usage across Claude Enterprise, Claude Code, and the API. Highly agentic workload covers context-heavy, tool-heavy work, where cache reads make up most of the cost.

Fable 5.1’s pricing is otherwise the same as Fable 5’s: $10 per million input tokens and $50 per million output tokens. In parallel, we’re continuing our work to bring many of the improvements of Fable 5.1 to the rest of our model family.

As discussed above, Claude Mythos 5.1 is available to vetted cyberdefenders and life scientists. Currently, it is only available to a set of US organizations, though we’re coordinating with the US government to expand access to a broader set of domestic and international partners as quickly as possible. To register interest in access to Claude Mythos 5.1 for cyberdefense through the CVP, head here.

Footnotes

  1. Terminal-Bench-Science 0.1: The standard error is ±3.5–4.5 pts per model. The public leaderboard (3 trials/task, Claude Code harness) reports Claude Opus 5 at 30.0% and Claude Fable 5 at 21.4%; our setup reproduces them at 29.0% and 24.7%, respectively, both within noise.

  2. OSWorld 2.0: Scores are on the benchmark authors’ August 2026 task release; Fable 5 and Opus 5 were re-run under the same conditions. Because the task files differ from earlier releases, these numbers aren't directly comparable to previously published OSWorld 2.0 results, which is why no competitor score is shown.

  3. These three targets are (EGFR, Nipah G, 15-PGDH) and come from Adaptyv Bio’s protein design competitions. The Nipah G comparison is against de novo designs targeting the receptor-binding site on the G head (best: ~8–12 nM, N1032). A de novo entry from Nick Boyd/Escalante Bio that targets a different region (the stalk) reached ~1.4 nM (design_7), comparable to our best binder.

The Daily Front Page 4 of 28
Tuesday, September 1, 2026 The Daily Front No. #260901 — Reasoning on 67 Cents
article

I trained a small transformer in 1.5hrs and it beats many LLMs

by porridgeraisin·▲ 597 points·155 comments·mvakde.github.io ↗
I trained a small transformer from scratch in 1.5hrs on a 5090.

I trained a small transformer from scratch in 1.5hrs on a 5090
Beats many LLMs, and scores the same as TRM/HRM

This is an upgrade to my previous model
Faster, better, cheaper and still open source.

Also gets 7% on ARC-2

Discussion on Twitter, Code on github

ARC-1 Public Eval

Performance on ARC-1 public eval. I only compare against models that do similar test time training

This is the 3rd blog in a series of works on ARC-AGI. Prev: Blog 2, Blog 1.

Many ppl thought the prev result was impossible. It got attention from top researchers and went viral on X. Eg: Discussions by Lucas Beyer, Jeremy Howard, Rohan Anil, and comments by many others.

Why work on this?

I think sample efficiency is the most important problem in AI today and I want to solve it.

The intention behind this work is to (1) find the limits of sample efficiency when restricted to transformers / today’s deep learning methods and (2) reduce costs so iteration is much faster and cheaper.

ARC is a great benchmark to test this:

  • Very few samples (only a 1000 puzzles) in a high dimensional space
  • Its a metalearning benchmark, so each puzzle uses a different rule, with some common concepts
  • Very few priors needed: every concept needed in the eval set is present in the train set
  • It is incredibly easy for humans to solve, and accessible to even poor AI researchers
  • Benchmark is still unsaturated (for data efficiency, ignore LLMs and approaches that use tons of synthetic data or human inductive biases)

Next, I’ll work on new research ideas to break these limits. I’ll try to keep costs low so that anyone in the world can work on this.

Tech details

How does it work?

The overall approach is similar to last time (full technical details here), but I added a bunch of upgrades. Here’s a quick summary of the approach:

  • Each input-output pair is converted to a sequence of tokens. These sequences are autoregressively trained on by a small transformer. This is done from scratch at test time on both the train set and eval set puzzles (test labels hidden).
  • To enable cross-task learning, each puzzle is given a separate additive embedding (learnt). Since each sequence has two 2D grids, positional are learnt using 3D RoPE embeddings.
  • The sequences are augmented with color and dihedral permutations. During inference, the test inputs are augmented, and the inverse aug is applied on the outputs produced. The 2 most common outputs are submitted (AAIVR).

Changes since last time

The main goal was to find improvements to the architecture / algorithm that improve the sample efficiency of the model.

The biggest increases in scores were due to

  • Modern architecture (SwiGlu instead of GELU, RMSnorm not layernorm, etc.)
  • More data diversity, better shuffling of data
  • scaling up: 8 layers instead of 4,

Biggest decreases in cost were due to:

  • Way fewer augmentations (more sample efficient!)
  • AdamW -> Normuon
  • flash attention with varlen training + flex attention kernels for inference

A major change is that I don’t train on input tokens anymore. This means the loss function only includes output tokens (which makes the approach supervised). This. performs slightly better 40% -> 44% but I don’t understand why. Perhaps finite model capacity

I also increased the training data by adding the non-overlapping tasks from ARC-2. I did this very carefully to ensure no leakage. You can remove the extra data if you don’t like it and it will still score ~40%, but it will need ~double the compute.

Context: ARC-2 contains 773 ARC-1 puzzles and 347 new puzzles. Most eval puzzles of ARC-1 are repeated, so if you naively train on ARC-2, then its a dataleak and you will score 100%. I avoid this by carefully filtering out the 773 repeated puzzles (so no leak!)

There are many other changes that gave incremental improvements in performance or speed. Find the full list of changes here.

Interesting behaviour

Since I am no longer training on inputs, this approach is now supervised. What’s weird is that the test loss is now worse, yet it scores better! Also it is more stable and there’s less variance in scores.

Many ppl today are working on sample efficiency by aiming for the lowest val loss on a small dataset. I think that’s great, but this points out a failure mode in such an approach

I do think the unsupervised style training will be better in some scenarios, and I am evaluating this.

Before NorMuon, I tried vanilla Muon. Obviously it trained much faster than AdamW, but the loss (and scores) would loiter at the end instead of converging. I found that cranking down the momentum and/or LR drastically at this point helped, but I didn’t want to make manually changes like this. When I switched to NorMuon, the problem disappeared

Ablations

The biggest contribution to performance seems to be good representations (3D RoPE + per-task embedding).

Ablating RoPE and per-task embeddings

Removing 3D RoPE or the per-task embedding gives a steep drop. Both ablations saturate at 25%

  • Training on inputs performs slightly worse -> ~39%
  • Restricting training set to ARC-1+ConceptARC only performs about the same: ~40%
  • Switching from 3D RoPE to 1D drops score to ~24%
  • Removing the per-task embeddings drops score to ~24%
  • Running the model CompressARC style (training from scratch on each task separately, and unsupervised), gives a drops performance down to ~18%
  • CompressARC but supervised gets ~15%

Other ablations, best scores

Finding the best scores on other ablations. Comparing costs makes little sense here as all but the first ablation requires a lot more compute

How can others contribute?

The code is open source. Feel free to modify it and improve score or reduce cost. (Pls don’t increase training data)

Try reaching 65% – you won’t need many modifications. Evidence: I took the union of all solved tasks from multiple runs, and got 55%. Also a bunch of other tasks are “almost” solved. Some ideas:

  • RoPE mixes positional and content information, which probably worsens performance. PoPE should perform on par or better. Or maybe invent a new pos embedding
  • The architecture can definitely be modernised further

Costs can probably be reduced 10x with handmade GPU code. There are architectural changes that can also do this.

Lastly, figure out how to remove data augmentations. (I hate that I used it, ignore everyone who thinks its okay). There are a few obvious ways to do so, but the challenge is keeping training costs low.

Misc

TBH, I didn’t expect to reach 45% with just the transformer, I thought this would need new ideas. I certainly didn’t expect to reach it at such low costs/flops. The ablations show that a surprising amount of perfomance is retained even without augmentations or synthetic data. Now I’m pretty sure 65% can be reached within the transformer framework

I don’t understand why others didn’t figure this out. Its just a transformer with the most obvious representation. This benchmark has been open for 6 years, was high profile, and had a million dollar prize! Maybe researchers underestimate deep learning? Maybe the cost of experimentation was high enough that they couldn’t run ablations properly? Blindsided by LLMs or using harnesses?

Appendix

Prev criticism/validation on my approach from famous researchers

My old result went viral on X and many experienced researchers debated about it, both for and against. Threads by Jeremy, Lucas, Susan, Andreas, Yoav, and many more. I’m listing all the criticisms here with my answers.

Training on the eval puzzles is cheating / “training on test”

  • No this is false. “Training on test” specifically means training on the labels of test data. The labels were not trained on.

  • Also, ARC is a metalearning benchmark, so you’re supposed to learn from the eval puzzles.

    • Jargon: ARC has a set of train puzzles and a set of eval puzzles. Each puzzle has example pairs and test pairs. A pair consists of an input grid + output grid.
    • The ARC, the label is only the test pair’s output grid in an eval puzzle.
    • These labels were not trained on. They are hidden. You can delete it beforehand if you wish

Training on the inputs of eval puzzles leaks information

  • No, this is false. Such an approach is called transductive reasoning and has been studied since the time of Vapnik.
  • Also, this dogma of ignoring eval inputs doesn’t make sense in a world trying to solve continual learning
  • Other approaches train a metalearning algorithm and then deploy it to learn by running a CoT or by modifying latents through a recurrent loo. My approach or what I did here is directly metalearn by modifying the weights of a single forward function is no different than learning by
  • Note: in the new 44% result, training on inputs has been removed as it scores slightly worse

Even if training on eval puzzle inputs is allowed, the test input specifically should be forbidden

  • No, this is false. The same “transduction” argument applies here

  • A metalearning benchmark can be transductive in 2 ways:

    • train puzzle $\to$ eval puzzles
    • within the eval puzzle, example pair $\to$ test pair
  • This criticism is specifically answered by the latter

This is against testing policy

  • No this is false.

  • The policy says “test taker must not know what the test will be”. People interpreted this as saying TTT is banned. But it actually refers to the human designing the AI system, not the AI system itself.

    • Eg: to discourage designing inductive biases based on the eval set.
  • To anyone active in the ARC community, this has always been clear since test time training has been allowed and encouraged. Steven and Chew’s comments clarify this and other concerns.

  • TTT also follows the spirit of a metalearning benchmark, so its fine!

You are not including training costs

  • No, this is false. I show the entire lifetime compute. This is the cost of training the model from init + the total cost of running inference on all tasks. Yes it totally amounts to 67 cents. Check the prices of a 5090 for 2hrs on vast.ai

Test time training is traditionally done one task at a time. Training on all test tasks at once is unrealistic

  • Yes, this criticism makes sense. But it’s nuanced
  • I agree that its rare to see to face problem sets in real life where every problem is given at once. Even if it is (like an exam), humans can usually only attempt one at a time
  • But just because humans don’t have a capability shouldn’t mean it invalidates building an AI model with that capability. Otherwise we could say LLMs are unrealistic since humans can’t train on the entire internet / can’t read tokens as fast
  • Also, it is unclear if humans are limited to one might be able to train on different data from multiple sensory at a time, exactly like

Providing cost per task amortises cost of training since all test tasks are trained on at once. So comparing other models is unfair

  • Yeah this is fair. In my defense:

    • That’s how the organisers compare every model, including TRM which also trains on all test tasks at once
    • I was also more generous by including training and inference costs while LLMs and other models exclude pre-training/offline training costs.
  • I have now switched to (a) showing lifetime compute cost instead of per-task, (b) comparing only with TRM, HRM and CompressARC and not with LLMs / other methods and (c) I added ablations with comparable training styles

Answering criticism about ARC-AGI itself

When I posted last time, there was a lot of debate about ARC-AGI itself. Some were valid, but a lot of them were questions Chollet has answered many times before:

  • What does ARC even test for? (fluid intelligence)
  • Why should we care about ARC? (fluid intelligence isn’t fully solved)
  • Solving ARC-AGI will not lead to AGI (no one claimed that)
  • ARC keeps shifting goalposts / its adversarially constructed for LLMs (Both are false)

Chollet’s paper and these tweets1 are good sources. Summing up his stance: The benchmark intended to test fluid intelligence, which he considers necessary but not sufficient for AGI. Solving ARC-1 / 2 implies non-zero fluid intelligence, but it isn’t an upper bound. The benchmarks don’t signal AGI is reached, they intend to point out the right research questions to ask. There were no goalposts moved: ARC-1 precedes LLMs, ARC-2 was announced pre-chatGPT and ARC-3 was announced before ARC-2 was saturated. He’s also happy about progress on ARC since it documents progress in AI.

I mainly care about ARC since it can be used to test for sample efficiency which is an important unsolved problem today! It’s also a well constructed meta-learning benchmark, and is accessible to GPU poor peeps. Historically, its been great at pointing out the strengths and flaws of LLMs. I also think its cool that the benchmark stood unsaturated for 6 years, despite being high profile / having a large cash prize since we now know DL can perform extraordinarily well on ARC-1/2.

There are some valid criticisms IMO:

  • They should disallow synthetic data for ARC-1/2

    • Its against the spirit of the benchmark and yet most top scores today rely on large amounts of it

      • synthetic data lowers the bar of fluid intelligence needed to solve puzzles
      • It only made sense till 2024 when DL scores sucked.
      • We now know LLMs/DL can learn anything given enough training data
    • This would also make the benchmark a great test for sample efficiency. It would complement ARC-3 very well

    • Question is how to prevent synthetic data? Simple:

  • Ban offline training/pretraining. Models must train from scratch after submission

    • Previously this was considered impossible so rule. My model shows this is possible
    • Guarantees no synthetic data can be used
    • It makes the comparison fair across differet models. Otherwise some models like LLMs can benchmaxx ARC by using ungodly amounts of offline training. (Since the benchmark has been around a long time, many ARC-like datasets have been created)
  • A single leaderboard graph comparing multiple types of models doesn’t make sense. It brings the following 3 problems (solution: separate charts)

    • The x-axis is cost/task. But it only counts online compute cost. Some of these models (like LLMs) have massive offline pretraining phases whose costs arent counted. You can use infinite training compute to effectively bring the test set into distribution, so these models should be evaluated separately.
    • Dividing cost by number of tasks makes no sense for the models that train on all test tasks at once (like mine, TRM & HRM)
    • Comparing LLMs on the public eval set makes no sense since the answers to the public puzzles are available on the internet
  • The organisers drew premature conclusions from TRM and HRM and attributed success to recursive loops+deep supervision. I think this bias is because they assume pure deep learning can’t solve ARC (eg: base LLMs still suck at ARC-2). I disagree

  • The wording of the testing policy can be improved to remove confusion. (Explained here)

Mistakes that I think other approaches are making

Assuming recursion is the next big thing (Eg: HRM, TRM, Arcprize blog)
I do see the appeal, but there aren’t enough ablations to prove this. And my model shows you can reach the same performance without recursion. The only confirmed benefit of recursion is allowing you to increase compute without increasing memory movement.

Misleading advertising by HRM/TRM: I also don’t like that TRM advertised itself as a 7M model when there are O(100M+) embedding weights being trained. It is misleading, makes it more like a lookup table, and calls into question what causes the performance. Worst case it should have been called 7M “active” weights. Same for HRM. Both didn’t mention this anywhere!

LLM based approaches on ARC aren’t showing new capabilities anymore:
Watching LLMs climb the ARC leaderboard has been extremely useful as explained below, but I don’t think there’s much to learn from their ARC-1/ARC-2 scores anymore:

  • Increases in LLM scores are now mainly driven by post training (evidence in next section) and are probably a function of amount of synthetic data. They are learning to solve ARC tasks, not learn general abstract reasoning
  • There’s also too many confounding factors to glean anything from new scores. Comparing LLMs based on benchmarks is bad science in general (eg: differing amounts of training data aimed at a benchmark)
  • For LLMs, only private scores should count. Their scores on the public leaderboard are useless as the answers are available on the internet, and are trained on.
  • Using harnesses on top of LLMs to improve performance makes little sense to me. All the post-training magic is happening inside the frontier labs, and they can build harnesses themselves. I think its unlikely continual learning will be solved by a harness.

Anti-bitter lesson cheats
I have already argued before that synthetic data and augmentations are bad. Designing inductive biases into the model is also bad. The fact that we can’t scale this benchmark without cheating like this shows that there are still breakthroughs waiting. I hope more people try to reduce such tricks that are anti-bitter lesson.

Learnings from LLMs on ARC-AGI

LLMs have now saturated v1 and v2 of this benchmark. Here’s what I infer from their progress:

ARC-AGI predates LLMs. They performed terribly on the benchmarks initially, showing that pretraining doesn’t confer general reasoning capabilities and that LLMs can suck at tasks that are incredibly easy for humans

OpenAI’s O1 getting 75% was a big win for LLMs. It suggested that given enough data, LLMs can learn any task during post-training. I assume this is what Sholto Douglas often argues about.

When ARC-2 came out, it reset progress of all LLMs, including the thinking ones. This suggests even post-training doesn’t confer general reasoning capabilities, otherwise a model that performs well on ARC-1 would automatically perform well on ARC-2.

(Basically, the models are learning how to solve ARC puzzles, not general abstract reasoning and its scores on a task are dependent on how well it is represented in its training data. Also, I’m not sure whether “general reasoning” even exists in the first place? Maybe humans are specialised too)

Since then, thinking LLMs have made steady progress on ARC-2. People often think this means models are better at general reasoning BUT what they don’t notice is that the base models are stuck at single digits. Taken with other evidence, this suggests:

  • Scores on ARC-2 are driven by post-training. (Probably largely depend on amount of synthetic ARC data?)
  • Labs are benchmaxxing (probably coz customers care about benchmark performance?)
  • LLMs are not sample efficient in any way.

Don’t get me wrong, I am very bullish on LLMs. The trends on ARC-2 show that performance will keep improving with increase in compute and data. Its also incredible to see the reduction in inference costs.

Full list of changes

Changes that modify training dynamics

  1. Optimizer changed from AdamW-only to NorMuon + auxiliary AdamW
  2. LR schedule changed from warmup+cosine to WSD schedule (warmup %, hold, then linear decay to floor).
  3. LayerNorm was replaced by RMSNorm
  4. FFN changed from Linear -> GELU -> Linear to SwiGLU-style gated FFN (chunk + SiLU gate).
  5. Weight decay changed from “non-attention linear only” to explicit group-wise decay: attention weights, token embeddings, and task/dihedral embeddings each have their own WD knobs.
  6. Training objective changed from outputs["loss"] (unsupervised style input+output LM loss) to outputs["output_loss"] (supervised style) only.
  7. Training batching changed from smart bucketing based on length to true random batching (bucketing retained for inference paths).
  8. Straggler/incomplete batches are now dropped in training (drop_last=True enforced).
  9. Dataset construction now supports/uses broader sources (ARC-1, ARC-2, ConceptARC, optional filtered cross-dataset tasks, submission/private modes), changing train data composition.
  10. Color augmentation changed from one global epoch-level permutation to per-example augmentation tuples (color + dihedral).
  11. Color permutation domain changed to excludes output-only colors, instead of blind 1..9 permutations.
  12. Augmentation generation now deduplicates transformed inputs via hashing across a task (higher unique-sample diversity).
  13. Augmentation selection is now epoch-cycled with shuffled candidate order (without-replacement per cycle behavior)
  14. Changes in hyperparams: optimizer/hparams, epochs, augment cap/type, depth (n_layers), and dataset path.
  15. A new dihedral_embedding was added and is now summed into token conditioning. (Only a very mild performance increase)

Speed increases without changing training dynamics:

  1. Training batches changed from padded [B,S] to packed token stream with cu_seqlens (no pad tokens in train path).
  2. Attention path changed from padded SDPA masking to packed varlen flash-attention support (cu_seqlens), plus flex-attention decode kernels.
  3. Dihedral augmentation moved from offline dataset expansion to online augmentation selection at collate time.
  4. Build-time training split changed from ("train","test") to ("train",)

Misc.

  1. Resume behavior changed: optimizer-switch/hparam-change detection now can reset/rewarm schedule, altering resumed-run dynamics.
  2. Scheduler stepping changed to fractional epoch progress when training, instead of pure per-step cosine progression.

TODO: ADD CITATIONS

  1. Chollet’s original paper about the ARC benchmark, some of his tweets explaining what it intends to test, and two tweets explaining the timeline of how the benchmark has evolved. 
The Daily Front Page 5 of 28
Tuesday, September 1, 2026 The Daily Front No. #260901 — The Skeptic’s Ledger
article

How accurate have Ed Zitron's AI skeptic predictions been?

by jatins·▲ 548 points·635 comments·danluu.com ↗
I looked at how his predictions panned out.

I was curious how well the predictions of the most widely cited AI skeptic I've seen (Ed Zitron) have done, so I looked at how his predictions panned out. To disclose my own biases, I've never had a particularly strong pro or anti AI progress position. For example, in 2022, I did a comprehensive look at predictions Futurists made, including well-respected folks like Kurzweil and found them to be generally wrong on both the prediction results as well as the reasoning. On the flip side, in 2015, I wrote about how people were underestimating AI's ability to displace humans in jobs and have repeatedly been on the record as saying that many people are underestimating AI's ability to displace humans from jobs. My position on AI has been extremely boring and is basically, "if something is currently happening, the people who are saying that it's impossible that it will ever happen are probably wrong".

One comment I've seen from a lot of AI skeptics when someone responds to an AI skeptic is that all of the people who are saying that AI isn't fake are self-interested liars. Personally (to my obvious detriment), I have no particular financial interest in AI companies. I own whatever the standard share of them is via boring index funds. I have some seed stage investments, but just due to the timing and what's gotten big, that part of my portfolio is underweight on AI. I don't work at an AI lab or a company that supplies AI labs. I've mentioned being hilariously bad at interviews before, and I did interview at an AI lab a number of years ago and failed the phone screen in a performance that was the kind of performance that must've inspired Jeff Atwood's famous Why Can’t Programmers... Program? where he concludes that there must be a lot of fake programmers out there because nobody could fail a coding interview that badly if they knew how to program. I don't benefit in any particular way if AI does well, except insofar as anyone who holds broad index funds benefits, but I do care about accuracy.

2024: Meta, Google, and Microsoft are dying

Because there are quite a few prediction results, let's look at one in detail before the complete list to get an idea of the kind of reasoning Zitron uses. We'll arbitrarily look at this November 2024 talk where Zitron says, among other things, the major tech companies (like Meta and Google) are dying and they're thrashing around on AI because they don't know how to grow.

Zitron specifically named Meta as a company that's dying ("it's a dying product, and it's kind of a dying company"). Meta's revenue and profit (GAAP operating income) have been

PeriodRevenueProfit Amount%Amount% 2023$135B16%$47B62% 2024$165B22%$69B48% 2025$201B22%$83B20% First half 2026$117B30%$42B10%

When he talked about companies not knowing how to grow ("none of these companies anymore really know how to grow ... in the desperation to try to reignite growth in a dying ecosystem the tech industry is going to shove this [AI] shit into everything"), he named Google and then Microsoft. Alphabet (Google's parent company) has had the following revenue and profit numbers:

PeriodRevenueProfit Amount%Amount% 2023$307B9%$84B13% 2024$350B14%$112B33% 2025$403B15%$129B15% First half 2026$230B23%$80B30%

And Microsoft's numbers have been (note that, for consistency, all numbers are calendar year numbers and not fiscal year numbers):

PeriodRevenueProfit Amount%Amount% 2023$228B12%$101B21% 2024$262B15%$118B17% 2025$305B17%$143B21% First half 2026$173B18%$79B19%

Although this wouldn't be in the spirit of Zitron's statement, one could argue that Meta is actually dying, it just hasn't died yet. However, the reasoning in Zitron's argument is incorrect here—the Meta, Google, and Microsoft ecosystems are not dying. Given how fast these companies are growing (in terms of revenue and profit), it doesn't seem that AI is, as Zitron implied, some kind of desperation move they're reaching for because "they don't know how to grow" and are all out of ideas. I don't think it's worth spending this much text on each prediction, but the pattern Zitron used here is illustrative.

To make the case that these things are dying, he pulls on minor issues that are not positioned to cause the very large changes he suggests are about to occur. For Meta, he cited some kind of alleged MAU drop for Facebook. Rather than use Meta's own MAU figures or any kind of revenue or profit numbers, he seems to have used numbers from Similarweb. My experience with 3rd party tracking numbers like this is that they're quite inaccurate and generally useless for anything other than a rough order of magnitude comparison, making the Zitron's cited decline meaningless (Meta's reported numbers show no such sustained decline).

For Google, he cites Prabhakar Raghavan, who he calls truly evil and "a computer scientist class traitor that sided with the management consultancy sect", as having done some kind of grievous damage to Google search. In his rants about Raghavan, he never credibly establishes that Raghavan is doing severe harm to Google search, and the Google search engineers who've commented on his rant don't seem to agree with the Raghavan as sole or even major reason for search issues hypothesis.

But even if we posit that Zitron is right and the villain Prabhakar Raghavan defeated the hero Ben Gomes, causing some kind of issue for Google search, this still doesn't make the case that Google revenue growth is in trouble at large because they have a number of other major products (such as YouTube and Google Cloud) that could drive growth even if search wasn't growing.

Every significant part of the chain of reasoning here is not only incorrect, it's not plausible if you know anything about Google or big companies in general. I'll be the first person to say that Google search quality has some serious problems and that Google has been increasing the relative priority of revenue over the user experience over time. This was a source of consternation for a number of user-focused engineers at Google when I was there in 2013.

For one of the issues Zitron cites, ads being confusing to users, in 2013, I asked a search engineer about Google changing the background color of ads to look more like search results because there was a previous study that showed that more an ad looked like a search result, the more users got confused over whether a result was an ad or a real search result, and I'd heard that Google deliberately made the ads not look like search results to avoid user confusion. The search engineer said that because some people didn't want users to get confused, it was impossible to make ads nearly identical to search results in a single change because it would be too obvious what's going on.

The way this was going to happen was that every time you A/B test tweaking ads to look a bit closer to search results, you make a lot more money, so the change would happen over multiple years in multiple parts, each small enough that the people who want to fight back against this kind of thing would have a hard time making a case. That happened just as this engineer predicted, but it was going to happen whether or not Raghavan ended up overseeing search. And, of course, that kind of thing happening doesn't cause Google to run out of room to grow and become desperate to reignite growth in a dying ecosystem. Whether or not you think Google should do it, it's something that makes Google more money.

How do people cite Zitron?

From what I can tell of how people cite Zitron, they cite him as an authority so they can say that this guy who looked at the numbers has made this claim, so their claim is backed up by the numbers. It turns out that if you look at the claims Zitron makes and know anything about the topic, the claims don't make sense, but I don't think that's the point. The point is one can say that someone looked at the numbers. The other point seems to be that this guy is angry, which is a good way to drive engagement.

But when people bring him up, they're of course not generally citing his anger; they're saying here's this guy who's looked at the numbers and, if you're angry about AI, he's right there with you being angry about AI, and he's got numbers on his side. Like I said above, I don't want to go into this level of detail on each claim; this is just an illustrative example about how the claims below look. For any of his posts that I read, while there are numbers thrown around, the numbers don't actually connect to a coherent argument. In many cases, as we saw above, the numbers don't even really support his argument (such as an MAU decline in Facebook causing Meta financial problems which would then cause Meta to spuriously insert AI in places it doesn't belong). I suspect he's relying on people's eyes glazing over when they see numbers and just not thinking about what the numbers mean.

With the predictions below, someone could have the exact same prediction record and have completely reasonable reasons that just didn't pan out. Or someone could be correct in every case and also be wrong because all of their reasons are wrong. Someone like the latter person might have some kind of intuition that they're unable to articulate, or perhaps they're someone who just got lucky. Fortunately for us, we don't have to make this difficult judgement call because Zitron is wrong on the predictions and also wrong on the reasoning.

People with attention to detail on Zitron

Since I've been living under a rock for years and am just catching on the AI discourse, I hadn't actually read or watched anything by Zitron or any of the big AI commentators, but on looking up what people who have good judgement say, they also seem to find that Zitron's use of numbers is just sleight of hand, such as this comment by Juho Snellman:

His writing is certainly flamboyant, but the aggression and expletives seem more targeted at hyping up people who already believe the things he writes, not for making people change their minds. He found a niche in anti-tech grift, and is now exploiting the niche for all he can. But you might want to actually fact-check a few of the things he says that convince you, because at least for his written articles basically everything is made up or misrepresented. There's plenty of links to sources, sure, but if you follow them down to the primary source what they're saying is very different from what Zitron is implying

Here's an example where commenters seem to assume that Zitron's analysis is good for some reason, to which Juho Snellman replies:

The key problem is that his economic analysis is absolute trash. I used to think he was just totally incompetent at it, but given the bias in the errors, it is pretty clearly intentional deception. But it's often pretty hard to address that, because every article he writes is a 10k word gish gallop. I've tried debunking key points a few times in HN comments for just one of the intentional mistakes he makes, and people complain about the reply being too long.

For example, when Timothy B. Lee looked at a spreadsheet that Zitron used to create a projection of Anthropic's revenue, he found

He doesn't count February 1-10, counts March 1-10 twice, counts August 21-October 21 as one month instead of two, and doesn't count October 21-November 1. [another commenter notes that his spreadsheet also contains February 30] ... Ed claims he tried to compute Anthropic's revenue for 2025 and came up with $3.6 billion, suggesting some funny business [but the numbers work out once you fix the errors]

Some Zitron predictions

  • Feb 2024: "I believe we're reaching the upper limits about what generative AI can do and how accurate its outputs can be."

    • Wrong
  • March 2024: "Have We Reached Peak AI?"; another prediction that hallucinations mean that AI progress is limited to then-current levels

    • Wrong
  • April 2024: "As I previously warned, artificial intelligence companies are running out of data ..."; another prediction that models can't improve because there's no more data

    • Wrong
  • June 2024: OpenAI growth is stalling (with the implication it will continue to stall), which will lead to some kind of collapse of OpenAI

    • Wrong (it could be the case that OpenAI will collapse but, if so, it won't be due to any kind of growth stall from 2024)
  • July 2024: "Generative AI, as I said back in March, is peaking, if it hasn't already peaked. It cannot do much more than it is currently doing, other than doing more of it faster with some new inputs"

    • Wrong
  • July 2024: "Generative AI models aren’t getting more energy-efficient, nor are they getting more “powerful” in a way that would increase their functionality"

    • Wrong (models continued to get more powerful)
  • August 2024: "generative AI is a dead-end technology that has peaked”

    • Wrong
  • August 2024: re-iteration that the AI bubble has 3 quarters to prove itself (from March 2024) or there will be a collapse

    • Wrong (Bartek Ogryczak notes, arguably Right because AI proved itself, but Zitron also argues no improvement, so Wrong by Zitron's accounting)
  • September 2024: "o1 shows that OpenAI is both desperate and out of ideas", with a re-iteration of the idea that models can't improve due to lack of data

    • Wrong
  • Oct 2024: OpenAI's forecast of $3.7B revenue in 2024 and $11.6B in 2025 and $100B in 2029 are absurd, "a statement so egregious that I am surprised it's not some kind of financial crime to say it out loud"

    • Wrong (2025 goal exceeded, 2029 TBD but not an egregious financial crime level of implausible)
  • Oct 2024: "[OpenAI revenue] growth is already slowing, and will slow dramatically as we enter the new year"

    • Wrong (OpenAI exceeded the forecasts and contiued to grow quickly)
  • Dec 2024: "I also warned you in March that generative AI had already peaked.”

    • Wrong (also, bizarrely, implying no progress since March 2024)
  • Jan 2025: "I believe we’re at peak AI"

    • Wrong
  • Jan 2025: "DeepSeek has commoditized the [LLM]"

    • Wrong (OpenAI and Anthropic had and still have significant pricing power and can maintain prices well above DeepSeek)
  • February 2025: Anthropic making $34.5B in revenue 2027 is "is laughable on many levels, chief of which is that OpenAI, which made around twice as much revenue as Anthropic did in 2024, barely made a billion dollars from API calls in the same year."

    • Wrong (whether or not they make that in 2027, their 2026 ARR greatly exceeding that makes the 2027 estimate non-laughable)
  • February 2025: "Sundar Pichai wants Gemini to be 'used by 500 million people before the end of 2025, 'a number so unrealistic that someone at Google should have been fired, and that someone is Sundar Pichai."

    • Wrong (Gemini hit 750 M users)
  • February 2025: "Sam Altman deputizing Orion from GPT-5 to GPT-4.5 suggests that OpenAI has hit a wall with making its next model, requiring him to lower expectations";

    • Wrong (GPT-5 was a substantial improvement over GPT-4.5)
  • February 2025: "I will keep writing this stuff until I’m proven wrong."

    • Wrong (Zitron continues to write despite repeatedly being proven wrong)
  • March 2025: "In my years writing this newsletter I have come across few companies as rotten as CoreWeave ..." Zitron goes on to say that the company will not be able to survive for six months except with fundraising, though $4B raised might by them a year

    • Wrong (CoreWeave still exists and it's currently at more than double its IPO price as of this writing; CoreWeave only raised $1.5B at IPO)
  • April 2025: Zitron calls the bubble again and says "We're about to find out if I'm right."

    • Wrong (in that Zitron implied momentous events were about to happen which would prove him right and no such events happened)
  • April 2025: "It also, at this point, is pretty obvious that generative AI isn't going to do much more than it does today."

    • Wrong
  • May 2025: "I do not know how you come away from this story and not think Cohere is going to die. Their projections are so far off from reality."

    • Technically unfalsifiable because there's no end date, but implied claim is wrong
  • July 2025: "I am not trying to be dramatic, but it's pretty easy to come to the conclusion that Cursor is going to die"

    • Wrong (Cursor gets a $60B exit)
  • August 2025: "These models have clearly hit a wall where training is hitting diminishing returns"

    • Wrong
  • August 2025: Zitron says Cursor is dying and expects that it will sell for a firesale price; a price as high as $10B is not plausbie: "Is Cursor worth $10 billion? Nope! No matter how good its product may or may not be, it is not good enough to be sold at a price that doesn’t require Cursor to incinerate hundreds of millions of dollars with no end in sight."

    • Wrong
  • October 2025: In response to the question, “If you had to guess, what is the timeline we are looking at for the AI bubble to pop?”, Zitron answers, "No later than Q2 2026"

    • Wrong
  • Nov 2025: "the fact we're running out of high quality training data and we're hitting the walls of scaling laws, in the training paradigm, these models aren't getting better. What we're seeing today is pretty much what they're always gonna be like"

    • Wrong

After this point, most further predictions that I saw were either non-falsifiable or resolve in the future. Note that I didn't attempt to catalogue statements that are nonsensical or were simply factually incorrect statements at the time, such as his December 2024 claim that “Generative AI's products have effectively been trapped in amber for over a year.” January 2026 claim that "[models are] basically the same as they were a year ago. They have the same efficacy". Zitron has not only made forward-looking statements that AI capabilities will not improve, he's also consistently made backwards-looking statements that capabilities have not improved which, while obviously false at the time, seem to play well to his base (along with his other false statements). If you connect all his statements together, it's implied that AI had the same capabilities in January 2026 as they did in December 2023 (and if you connect later statements, it's actually implied that capabilities in August 2026 are the same as in December 2023, though to be fair to Zitron he frequently contradicts himself and has also admitted to limited improvement in mid 2026).

To be fair, we could say that Zitron is speaking colloquially, so we when he says things like "have effectively been trapped in amber for over a year", that doesn't mean there's actually be no change December 2023, so the statements aren't transitive. Even if you assume a kind of colloquial sloppiness here, the collection of statements still implies that, from December 2023 to August 2026, improvements have been minimal (perhaps except, as noted above, when he contradicts himself and admits there have been limited improvements in some areas).

Comparing to respected Futurists

If we compare to how futurists did in our analysis of futurists, on style, Zitron relies much more heavily on anger than any of the futurists we looked at. On the quality of reasoning, he was probably about average compared to the futurists. Despite being wrong on roughly everything, he's not more unreasonable than someone like Buckminster Fuller, who suggested we'll be able to send people by radio because atoms have frequencies and radio waves have frequencies so it will be possible to pick up all of our frequencies and send them by radio.

In terms of the style of reasoning, of the futurists reviewed, he's probably closest to Kurzweil, in that he uses numbers to give a kind of aura of credibility, but if you know something about the topic he's discussing or look at the numbers, the reasoning falls apart. Zitron's reasoning isn't worse than Kurzweil's, who (for example) continually made new predictions of extremely fast progress that didn't pan out (such as, in 2001, predicting unbounded lifespans by 2011). Continually predicting that AI progress will stop for reasons that are incorrect is just taking the flip side of the bet on progress. Instead of having infinite progress, we're going to have no progress. Every time that prediction is proven wrong, you can just make another similar prediction and then move the date forward a bit. Michał Zalewski (lcamtuf) has some thoughts on why this happens:

The surest way to build [a] popular following is to articulate positions that are crisp, strong, and leave no room for doubt. You can't get too many podcast or TV appearances out of "well, the market could go either way", "both political parties make good points", "there's some merit but also some hype to AI". Or, to tap into the example in the post, "Harry Potter is an OK book".

In fact, there's a positive feedback loop. If you take a provocative, edgy stance, you get more attention and likes, so you sort of... self-radicalize? At some point, it's no longer an opinion that can be changed. It's an identity, a personal brand.

It's ... why Ed Zitron has a blockbuster blog about how it's all just one big scam. If you take a more nuanced view, you will at best get no reaction, or at worst, you'll invite scorn from both sides.

For anoyone looking for well-reasoned anti-AI takes, I find whitequark to be quite good (not that I agree, but I think the reasoning is sound and I could see how someone would agree if they have slightly different premises than I do), but of course whitequark doesn't draw the kind of big audience that Zitron does.

How long can you maintain an incorrect position for?

I'm curious what people do after being on the wrong side of a set of failed predictions about progress like this. For the futurists, even the ones who were nearly completely wrong (which was every single one reviewed here), they can still make some kind of case like "a quarter of the things I said would happen happened, it just took two to twenty times longer than I expected" and if they're not so stuck on accuracy, they can round this up to "the things I said would happen happened", which is often what they've done. That seems to have served them well as nobody really cares to look at the details anyway.

But what happens to someone like Paul Ehrlich, who predicted imminent catastrophe when this clearly was not happening as he was writing and then did not happen? Just looking at Ehrlich's Wikipedia page, we have

A common criticism is that Ehrlich's predictions routinely failed to come true; for instance, Ronald Bailey of Reason magazine has termed him an "irrepressible doomster ... who, as far as I can tell, has never been right in any of his forecasts of imminent catastrophe." On the first Earth Day in 1970, he warned that "[i]n ten years all important animal life in the sea will be extinct. Large areas of coastline will have to be evacuated because of the stench of dead fish."

In a 1971 speech, he predicted that: "By the year 2000 the United Kingdom will be simply a small group of impoverished islands, inhabited by some 70 million hungry people." "If I were a gambler," Professor Ehrlich concluded before boarding an airplane, "I would take even money that England will not exist in the year 2000."

When this scenario did not occur, he responded that "When you predict the future, you get things wrong. How wrong is another question. I would have lost if I had had taken the bet. However, if you look closely at England, what can I tell you? They're having all kinds of problems, just like everybody else."

Ehrlich wrote in The Population Bomb that, "India couldn't possibly feed two hundred million more people by 1980." In 1967, Ehrlich called to cut off emergency food aid to India as "hopeless". This position was later criticized, as India's food production subsequently skyrocketed through the Green Revolution in India, and its per capita caloric intake rose significantly in the following decades, even as its population doubled.

A large increase in global food production since the 1960s and a slowing of population growth have, within the current context of continued depletion of non-renewable resources, averted the scale of food shortage, famine and catastrophe foretold by the Ehrlichs.

Canadian journalist Dan Gardner, in his 2010 book Future Babble, argues that Ehrlich has been insufficiently forthright in acknowledging errors he made, while being intellectually dishonest or evasive in taking credit for things he claims he got "right". For example, he rarely acknowledges the mistakes he made in predicting material shortages, massive death tolls from starvation (as many as one billion in the publication Age of Affluence) or regarding the disastrous effects on specific countries. Meanwhile, he is happy to claim credit for "predicting" the increase of AIDS or global warming.

In the case of disease, Ehrlich had predicted the increase of a disease based on overcrowding, or the weakened immune systems of starving people, so it is "a stretch to see this as forecasting the emergence of AIDS in the 1980s." Similarly, global warming was one of the scenarios that Ehrlich described, so claiming credit for it, while disavowing responsibility for failed scenarios is a double standard. Gardner believes that Ehrlich is displaying classical signs of cognitive dissonance, and that his failure to acknowledge obvious errors of his own judgement render his current thinking suspect.

Barry Commoner has criticized Ehrlich's 1970 statement that "When you reach a point where you realize further efforts will be futile, you may as well look after yourself and your friends and enjoy what little time you have left. That point for me is 1972." Gardner has criticized Ehrlich for endorsing the strategies proposed by William and Paul Paddock in their book Famine 1975!. They had proposed a system of "triage" that would end food aid to "hopeless" countries such as India and Egypt. In Population Bomb, Ehrlich suggests that "there is no rational choice except to adopt some form of the Paddocks' strategy as far as food distribution is concerned." Had this strategy been implemented for countries such as India and Egypt, which were reliant on food aid at that time, they would almost certainly have suffered famines. Instead, both Egypt and India have greatly increased their food production and now feed much larger populations without reliance on food aid

Amazingly, following the series of incorrect predictions Ehrlich made in and after writing The Population Bomb in 1968, he followed this up with The Population Explosion in 1990 and has continued saying that we have global overpopulation that is causing or will cause a dire crisis unless we cut worldwide population. He has said the same thing this century and even this decade. It appears the only reason he's not saying that today is that he died earlier this year.

If I didn't look it up, I would've guessed that his recent position would be something like "well, I got some things wrong, but it was only due to these actions that were inspired by my work that crisis was averted" or "while crisis was averted, it was a lucky roll of the dice and, in most universes, the agricultural advancements that staved off the mass starvation deaths I was predicting don't happen", not "just you wait, the crisis is happening now and I'm about to be proven right"; in 2015, referring to his incorrect 1968 book, he said "[m]y language would be even more apocalyptic today". That's the pattern we've seen from Zitron, but I wouldn't have guessed that the one person I looked up would've kept that up for 50 more years. Maybe we'll get 50 more years of Zitron predicting the end of AI progress.

Some reactions to Zitron

In one of the quotes from Juho Snellman, above, Snellman says that he writes a large amount of gish gallop, which is a term for when someone floods you with so much cheap (as in cheap to produce) nonsense that no one would want to take the time to bother to refute it. In discussing one small part of Zitron's talk in detail, we spent more than 1000 words explaining why Zitron has an incorrect understanding of how corporations work and how Zitron got the reasoning wrong. Someone can read that and then say, "but you didn't address X" in the talk, which is true. When I first watched the talk, I actually closed the tab after 90 seconds because there was so much nonsense that it didn't seem worth the time to go any further. I could write 5k words on the first 90 seconds of the video. Because Zitron is just saying a bunch of nonsense, he can do that very cheaply and it would take 30-60 minutes to refute 90 seconds of his nonsense if I had all the facts at hand. With time to look up the exact right information, it probably would take double or triple the amount of time. When someone who has good judgement sees something like this, they tend to immediately write the person off. Just for example, I mentioned to a friend of mine that I'm writing this post and they said

I was listening to this podcast with the guy and I couldn't get through it. My heart rate was going up because he would just say this false thing and then the interviewer, who was reasonable, would ask about it, "what about X?", and then we would just jump to another falsehood ...

... before I ducked out, he talks about how LLMs haven't gotten a lot better over the past year, and the interviewer says people use them and they've definitely gotten a lot better in the past year, and Zitron denies it and says 'have they?', and the interviewer is just like, "yes..." At that point, I'm just like, why am I listening to this conversation?

We mostly discussed predictions and not incorrect statements about the past or present, but everything I've read or watched by Zitron is also full of things like this. Many people will look at something like this and decide the guy is a crank and stop paying attention. But many other people will look at something like this, see someone refute a set of things, and then say, "but you didn't refute X" and, in general, the person doing the refuting may respond to a couple of these, but they eventually give up because the gish gallop method has the same properties as an amplification DoS attack. It's very cheap to generate new nonsense, but it takes some effort to refute it.

BTW, I was curious what this interview was, so I put the above quote into ChatGPT and asked it to find the interview. It was able to identify an interview with the relevant exchange (it actually identified multiple, as this appears to be a common question and response pattern by Zitron) and the timestamp of each relevant statement in the interview (the start of the general argument is here and a "have they" response is here. Prior to the "have they?" comment, the interviewer tries to establish a baseline that agents have improved in capability. Zitron denies that this has happened, and then when the interviewer notes that people who use these things for their jobs Zitron denies this with the "have they?" comment (he actually makes multiple contradictory statements in the sequence).

Another thing to note here is Zitron's extremely high level of stated confidence. Some that we noted were OpenAI's forecast that is "a statement so egregious that I am surprised it's not some kind of financial crime to say it out loud" (which they've achieved so far) and his claim that Google's forecast for Gemini users is "a number so unrealistic that someone at Google should have been fired, and that someone is Sundar Pichai" (they managed to exceed the forecast by 50% when Zitron's claim was that it would be completely absurd for them to reach the number at all).

I've made quite a few predictions, and quite a few of those predictions are wrong. When I'm really making a prediction, I attach a confidence level to the prediction just for my own sake, so I can look back at these things and see how well calibrated the predictions are. I have never been wrong about a prediction that has anywhere near the confidence Zitron gives to some of his predictions. Given the stated level of confidence, even a single incorrect prediction would be a sign of an extremely high degree of overconfidence. One should effectively never be wrong about a prediction delivered with that level of confidence but Zitron is routinely wrong about predictions he makes with what is rhetorically pretty much the highest possible degree of confidence.

BTW, a funny thing about Gemini hitting 500M users being "so unrealistic that someone at Google should have been fired, and that someone is Sundar Pichai" is that Zitron has also (incorrectly) said that Google doesn't know how to grow, and that as a result they're shoving AI everywhere. Dennis Snell pointed out that, if Zitron takes his own statement seriously, Google can make Gemini's user numbers go to any number it wants by doing the exact thing Zitron said they would do, sticking AI everywhere.

You can't actually take Zitron's statement about Google's lack of growth leading to AI desperation seriously and also take it seriously when he says that Sundar is committing some kind of gross malpractice by naming a number like 500M users. This is another thing that is immediately obvious on watching one of his talks or reading his writing. There are a bunch of disconnected statements that don't fit together, except insofar as they're statements about how AI companies and people and companies that are using AI are evil and bad. The actual numbers and logic of the statements are contradictory. It seems to be whatever comes to mind that can be used to paint the villains as evil. And, funnily enough, the 750M user number Gemini hit shows that both of Zitron's statements were incorrect. If Google were as desperate to juice the numbers as Zitron claimed, they could've easily gotten the number above 1B by sticking Gemini everywhere, and of course 750M > 500M.

BTW, the point at which I stopped the talk for the first time was

a market obsessed with year-over-year revenue growth. And this progression was natural. It was horrible. You can blame Marc Andreessen. He's a horrible man. You can blame many horrible men. There are so many guys to be mad at the moment.

That last sentence really sums up Zitron's position. "There are so many guys to be mad at the moment". In this talk, he throws in this jab at Andreesen and blames Andreesen for Meta, Google, and Microsoft pursuing growth. In reality, if Marc Andreesen had never existed, Meta, Google, and Microsoft would almost certainly still be trying to grow so we of course cannot actually blame Andreesen for these companies trying to grow. There's just this thing that he says is bad, and in his usual style, he pulls some person and says they're the evil villain that's to blame for this, and then moves on to the next non sequitur.

How can people take this seriously?

Because I'm a masochist, I actually went and read a bunch of Zitron discussions (I believe I read every major discussion on HN and lobsters, and a bunch of other ones as well) to see what people who take Zitron seriously are saying. One common defense was the one above, sure, you refuted some points, but you didn't cover X. A more common defense is to say, just in general, people attack Zitron because of Y (usually his style), but they never address his points, "which tells me everything I need to know" (or something along those same lines). Based on the timestamps of the messages, just scoping to the stories that were being discussed, there were generally already comments discussing Zitron's actual errors, but Zitron's defenders would ignore this and just claim that people were unable to point to mistakes Zitron had made. This is a very Zitronian move and it makes sense that people who like his style would also use this move. After all, who would find Zitron convincing? Someone who thinks this kind of thing is valid reasoning.

The next most common "move" was to simply deny that Zitron said something that was refuted. When people would mention that Zitron was repeatedly on the record in 2024 and 2025 as having said LLMs couldn't improve further for fundamental reasons, Zitron's defenders would say that he never said that, and likewise for previous predictions or factually incorrect statements.

Another class of defense I saw were comments like "but what about all the AI hypists who are wrong?". Like I said before, I wrote a 34k word post about how a bunch of the most respected futurists have been wrong, not just because they made incorrect predictions, but their methods and reasoning were wrong. But a bunch of people who hype the future being wrong doesn't make people like Ed Zitron or Paul Ehrlich any less wrong. Zitron and Ehrlich are still exactly as wrong as they would be if those futurists never existed.

A friend of mine also noted this about comments on cases where people point out that Zitron was wrong about models not improving from 2023 to 2026 (and yes, this is specifically on stories or comments that discuss Zitron's disproven statements on capabilities not improving):

It's incredible to see so many people saying, "Zitron isn't wrong, he's just early!" I guess the implication is that we'll eventually realize that the models we have in 2026 are actually no better than the ones we had in 2024 or ??

Future predictions

Although Zitron's past predictions have generally been wrong, maybe he'll be right about something in the future. Perhaps some of these companies will have valuations decline for some reason. But, even if there's some kind of massive AI crash and OpenAI and Anthropic go to zero, in terms of the societal impact, if on top of that, some other event occurs that prevents further progress in models beyond whatever AI labs have internally right now, that's still going to result in a fair amount of change. Which companies are successful will change who gets rich, but particular companies failing won't stop changes that fall out of current or next generation model capabilities from happening; it just moves around who benefits the most.

Personally, it doesn't matter to me if folks at one company vs. another get rich. If one company does something better (in some abstract sense) than another, that's of some interest to me, but I have some skepticism about any particular company's claims that they'll do more of "the right thing" than another company (I could be convinced on this one, but I don't find the public claims that I know of very convincing).

If Zitron ends up being right about some company or other collapsing, that's pretty uninteresting to me compared to how capabilities have developed and will develop, where he's been wrong to date. It also happens that he's been wrong about the financial predictions he's made to date, but that doesn't really interest me, though I included a number of financial predictions for completeness.

Thanks to Yossi Kreinin, Juho Snellman, Dennis Snell, Nick Bergson-Shilcock, Bartek Ogryczak, Jamie Brandon, and Shriram Krishnamurthi for comments/corrections/discussion.

Appendix: Ed Zitron on why people don't like Ed Zitron

While looking for discussions about Zitron's work, the #2 hit on reddit was this comment by Zitron:

... some men don't like me because emotional honesty and introspection are difficult for them. Feelings are something that men are told to repress or compress. I refuse, and I find it disgusting when anyone tells me to do so ...

... Let's start with emotions, because it's the most obvious one. People really do not like that I am how I am, and think that I am "getting mad as a bit," or even go as far as to describe me as psychotic, out-of-control, and so on and so forth. This is a common reaction, I find, from anyone who themselves is emotionally repressed, especially in their own work. It is hard to be emotional and have well-done opinions ...

... I also have not taken the route you are "meant to take" to get here. You are "meant" to be an establishment writer from a big outlet, or an analyst, or in finance, or any number of other different "true paths" where you are "worthy" of whatever it is you're meant to get. I did not "earn my stripes" in the traditional sense, and those that have believe I did not earn my way here ...

... My work is also thorough, which is frustrating for people that do not do thorough work. Notice how many people still claim "it's just like Uber" or "it's just like the dot com boom." It's much easier to just assume shit without ever checking if it's true! Having some asshole who comes along with thoroughly and with passion is frustrating. It reflects badly on your work ...

... I do a good photo shoot, I do a good interview, and I capitalize on events, and I do so without being craven, because I usually show up with a few thousand words of thoughts or an episode about a thing. I believe there are some that would like this level of attention or prestige, but they do not want to do the work to get it, and that chafes ...

... I love big, I love hard, I am who I am, I have never been made to feel welcome by any "in" group. I work my ass off, I write more than anybody else, I show up. With whatever space I create I will fight back against "in groups" or cliques. I hate them, and they hate me right back. And I fundamentally know why I believe what I believe. That upsets people who do not.

I have no idea if he means any of that or not (if this Wired profile is accurate, one would have to lean towards not), but Zitron seems to be very good at saying what his audience wants to hear, so this proably gives some kind of insight into his audience, at a minimum.

One thing to note about the bit about cliques and "in groups", if you just search his name on reddit commenters note that if you post anything indicating that AI has improved on his subreddit (such as link to benchmarks), you get banned for it, resulting in a highly clique-y echo chamber. I'm on the record as having said that METR's progress benchmark isn't meaningful and that you're better off going on vibes than leaning on a misleading analysis and that widely cited AI evals are frequently flawed, so it's not like I think that benchmarks are generally good, but the picture I got from reading comments was that you get banned pretty quickly if you don't hew to the party line, which is the opposite of the picture painted above. This isn't anything unique to Zitron; when looking up another influencer a while back, if you disagreed with that influencer on their reddit, they would write a comment thanking you for your comment and saying how much they loved getting feedback from people and how the world is some kind of great peace and love fest and we should all love each other while simultaneously banning you from their reddit.

I also found Zitron's comments on how people don't like his work because they dislike thorough work to be interesting for a couple reasons.

One is that my own work is frequently positively cited as being rigorous and thorough. There are plenty of people who dislike my work as well, but not only do I not know of anyone who's said they dislike it because it's thorough, I would be surprised if there was anyone who secretly dislikes it because it's thorough. In general, just doesn't seem like a reason that people dislike things.

The second thing is that, I wouldn't personally consider my work to be thorough. The same thing I mentioned here about not feeling that my work is good also applies to not feeling my work is thorough. I do some amount of checking of my work. I don't know that I'd say that it's more than most in terms of time spent, but in terms of effectiveness, I suspect the combination of methods and time spent works better than average. But I always have a dissatisfaction with my work when I published it because I could keep checking more thoroughly forever and never publish anything, so I force myself to publish at a level that I suspect is above average on thoroughness, but well short of thorough. If I compare my work to the work of someone I consider thorough, like Gary Bernhardt, I don't know how I could call my work thorough. I have a few friends who produce Bernhardt-quality work and I make the choice to produce much more but also lower quality work. I think this is a fine place to sit in the quality-speed tradeoff space, but that doesn't make my work thorough. This goes double for everything I've published since starting to write publicly again this July since I'm experimenting with pushing things out the door with much less checking and editing than usual. And yet, it would seem that my fact checking process is a lot more thorough than Zitron's.

Appendix: why write this?

No good reason, really. I got four hours of sleep and my brain wasn't good for much of anything and I saw someone posted a screenshot of a reddit post dunking on Ed Zitron's prediction record. When I wrote this review of futurist prediction accuracy, I tried to make sure that I didn't bias what I was reviewing in any way. It's not obvious from the post if the redditor who reviewed Zitron's predictions was pulling predictions in an unbiased fashion or if they were biased in some way (since AI has become a culture war issue, it wouldn't be surprising if someone pulled biased predictions), so I decided to read some Zitron in my spare time while poking at agents to get them to do an unrelated task I wanted them to do. For the futurist post, I read multiple entire books to pull predictions and generally only stopped when someone was being repetitive and kept saying the same thing over and over again. In this case, all Zitron does is be repetitive, so the methodology in the futurist review would mean that I review a few predictions and then stop immediately. To overcome this, I had ChatGPT give me a list of predictions (with no attempted tilt towards correct or incorrect predictions) and then I skimmed/read the posts that ChatGPT linked to. There were some cases where I thought ChatGPT's reading of the post was incorrect (these were generally cases where it flagged a prediction that would be incorrect if its reading was correct, but I disagreed with its reading) and I also removed predictions which weren't falsifiable or seemed pointless because they were tautological (I noted something similar to this in the futurist post).

If I really thought about it, I probably could've found something better to do with the time, but here we are; I sometimes have tasks on my todo list for when I'm too tired to do real work, but I didn't have one. I don't think they cherry picked particularly bad predictions, although they did pick some that are among the more absurd sounding. However, if you go and look into the details of ones that aren't such ironclad "dunks" (like saying that Gemini hitting 500M by EOY users is so absurd Sundar should be fired for the idea, when Gemini actually hit 750M by EOY), these are just as wrong as claims that Cursor has no realistic buyer with the implication they won't even sell for $10B when "everybody" (who cares about AI exits) knows they sold for $60B.

The redditor picked the high-profile failed predictions, but Zitron's prediction corpus has many more failures and, as noted above, the bigger issue is his reasoning.

Another thing about the reddit comment is, whether or not the comment is unbiased, one might have the suspicion of a kind of bias because it was posted to r/accelerate by someone who apparently is an r/accelerate believer. On looking at the actual predictions they are consistent with some bias (they would also be consistent with an honest mistake as there's no way to distinguish these from the record). For example, one of the "refutations" is a statement by Zitron that OpenAI will collapse in 12-24 months. OpenAI didn't collapse, so this would appear on the surface to be a great way to show that Zitron was wrong, but if you read Zitron's post, Zitron's actual claim was that OpenAI will either collapse or raise a lot more money and they raised a lot more money. I disagree with Zitron's implications that this is inevitable just leading to a later collapse but his stated prediction was not falsified.

This prediction wasn't in the set of predictions scored in this post. Some would argue that this should be scored in the post. The reason this wasn't scored is because the prediction seems meaningless except insofar as it contributes to Zitron's broader point (that OpenAI is doomed and must collapse).

If we think about predictions one could make, a tautological prediction (if you write out all the edge cases I'll elide for space reasons) that has to be true is OpenAI has enough money to operate or it doesn't, and if it doesn't, it must raise the money somehow. I could make a million such tautological predictions, but if one were scoring my prediction record, it wouldn't make sense to include these because they're meaningless. In general, a company that's alive will cover its costs. If it does not, it will try to raise money. If it fails to do that, it will shut down or get acquired. A prediction that a company will either cover its costs or it will not cover its costs says nothing.

OpenAI's own projections were that it would not yet be profitable and its costs would exceed its revenue. That seemed nearly certain, so if you assume that this nearly certain thing is true, then you have the nearly tautological prediction that OpenAI will either collapse or it will raise money to cover its costs. It would have been reasonable to make a prediction like this at very high confidence (99.9% or above). If you use any kind of prediction scoring methodology, such as Brier score, these predictions contribute essentially nothing except when they're wrong as long as Zitron has a significant number of high-confidence incorrect predictions.

And, as we noted above, Zitron is repeatedly incorrect on predictions he gives the highest possible confidence (given his wording, I would rate a number of these at 6 9s or above), so on any kind of scoring mechanism like Brier score, Zitron's record is very poor. And a summary metric like this really understates how meaningless predictions like this are. Hypothetically, let's say Zitron made an unbounded number of correct 99.99% certainty near tautological predictions, which would make the score from the bounded number of other predictions he made meaningless on something like Brier score. This would still give you zero confidence for any of his non-near tautological predictions, and those are the predictions people generally talk about (AI progress is done, AI companies must collapse and this will bring down major tech companies as well, etc.).

Back the topic of the reddit commenter's potential bias vs. mine, as noted above, I don't have a particular bias towards a view that rapid progress is inevitible and have called out cases where people are overly optimistic, as evidenced by this post on futurist predictions. I'm also not someome who needs to or has any desire to farm engagement by manufacturing reasons that someone is wrong or bad and don't consistently rate every predictor as bad, as evidenced by this review of Steve Yegge's prediction record, in which I note that he scored well and also actually performed much better than the raw score indicated because the predictions are generally well reasoned and directionally correct even if the precise prediction was incorrect. I think it's actually awesome if someone has good insight in the future and shares it publicly, so I'm happy to call these cases out when I noticed them. It's just that, in this case, Zitron is a kind of anti-Yegge: someone with a poor prediction record whose predictions are actually worse than they seem from the record alone.

Appendix: errors in this post

I think it's almost certain that this post has multiple errors. In general, I find it very difficult to read a long stream of incorrect reasoning and then not get sloppy when looking for errors in it. I had this exact same problem when reviewing futurist predictions. It reminds me of when you're programming for some system where the compiler is very buggy and you hit compiler bugs all day every day (not uncommon when working with embedded systems, at least pre-LLM; now you can fix the bugs relatively easily). I find it hard not to get sloppy and think "hmm, this might be a compiler bug" even though, every once in a while, it will actually be your bug and not a compiler bug. The problem is much worse when looking at predictions from these kinds of predictions since the compiler still generally basically works and is often right, whereas when reading text like discussed here, you're just constantly drowning in nonsense that is occasionally punctuated by a good and accurate point.

I think, to do this well, you'd either need to find someone with very unusually high endurance for trudging through this stuff (I mean, much more than me, and I seem to have a somewhat above average endurance for this kind of thing) or have a team of people who independently rate and score things, but who would want to spend that kind of effort when any surface-level reading immediately reveals many things that indicate that these folks are pretty much totally wrong?

I did ask ChatGPT (web interface, Pro) and Claude (web interface, Fable 5) to fact check this post. They both found some minor errors that were fixed before publication.

One year ago, I found fact checks like this nearly useless, but they're halfway decent now and, contra Zitron, I would expect them to continue to get better. For people who are curious about the two, ChatGPT was much more thorough than Claude in this case and found more errors as well as finding every error that Claude found. However, it was overzealous and cited a number of non-errors, such as suggesting that tongue-in-cheek comments were incorrect, and that a number of statements that were generally true should be re-phrased in some more literal way (complete with AI-styled text).


  1. But, even if it were the case that the accusations against Raghavan are true (I'm not sure how they could be, as how could one be a class traitor to computer scientists in the first place, but let's posit that, whatever it means, it's true), Zitron's contention is that "this shithead [points to an image of Raghavan] took over Google search in 2020" and then prioritized certain metrics over search quality. I'm not sure why one would name a particular person for this as this is something that was a long-standing fight with many people involved on all sides but, if we posit that this is all true, then we posit that the "management consultancy sect" will move metrics that will cause engagement and/or revenue to increase at the cost of search quality. This would have the opposite of the effect Zitron needs here to make his case that Google growth is done and they're so desperate for growth they have to put AI everywhere in some kind of crazed last-ditch attempt to save Google. Perhaps one could make the argument that this will eventually cause Google search to decline, but Zitron's argument was that, in 2024, they were desperate, not that users will eventually leave Google search, which will later cause a decline.

    Anyone who's read a lot of Zitron will recognize a standard "move" of his, turning the situation into some kind of hero-villain narrative (for search, the alleged hero is Ben Gomes and the villain is Prabhakar Raghavan); it's as if his mental model of how companies works comes from movies about companies. If you ever watch a movie that's allegedly about some events and then read about it, you'll find that things get oversimplified into a hero-villain narrative and that almost all of the nuance is stripped out of the situation. And then if you're ever personally involved in something or talk to people who are personally involved and compare what happened to the books that get written about it, the same thing happens again; in general, the major causal factors are not identified in books about what happened in tech and many of the most instrumental people involved in some of the key decisions aren't even named because journalists talking to people about what happened aren't really able to piece together a plausibly correct story about what happened to someone who understands the underlying mechanics and has good information. Anyway, without knowing anything about the situation, if someone tells you a hero-villain narrative of the kind Zitron likes to spin, you can already be a bit skeptical.

  2. BTW, I don't think his anger really comes across in the video. I mean, he explicitly says he's angry and he swears and insults people, just like in his writing, but he doesn't really read as angry to me. It reminds me of this test on emotion recognition I took with a bunch of folks recently.

    I found the test fairly difficult and spent maybe 5 minutes on the first question because the person had a huge fake smile on their face and also looked a bit uncomfortable and anxious. I couldn't tell if you were supposed to say that the person is happy or uncomfortable/anxious. Is it supposed to be a very easy test or is it supposed to be a test that has a bit of subtlety? Based on what the test looked like, after thinking about it for a while, I chose "happy". Luckily, the test actually tells you if you got the question right or not, so I realized the test was about the fake exaggerated expression being made and not the person's actual expression and most the rest of the questions were easy. One was difficult because they were faking one particular emotion with what is a textbook display, as in, the kind of thing one sees in a textbook, but in a very specific way that was less complete and more unrealistic than the other textbook displays; it was as if someone had read a description of what a contemptuous sneer is, and then was trying to make the facial expression based on the textual description. I had to think about that one for a couple minutes to get the correct answer.

    Anyway, to me, Zitron seems like someone who's playacting anger and not someone who's actually angry. The tone of voice, facial expression, body language, style of movement, etc., just don't seem angry to me. I think this anger positioning works better in his writing than in his speeches because the cues he uses (swearing, saying he's angry, showing a lot of contempt, insulting people, etc.) are about as good as it gets for anger cues in writing. When you have audio and video, these are fairly weak cues; if the stronger cues don't really indicate anger, the person just doesn't seem angry. It's possible he has a non-standard way of showing anger or I just wasn't paying enough attention, but after watching some videos of him where he talks like he writes but didn't seem angry, the writing just doesn't feel angry to me anymore.

  3. If you want to see an example of what it looks like when someone tries to discuss the numbers, here's a thread where Juho Snellman pushes back on someone who insists that people have done the math. As I've been catching up AI discussions, I've seen many discussions like this where one side has someone who's actually looked at the numbers and the other side waves around some kind of vague insistence that numbers have been looked at. This never really goes anywhere because, for one of the sides, the point isn't that you can understand something from the numbers, it's that they have a piece of evidence they can wield because someone has looked at the numbers.

  4. Zitron's argument at the time was that hallucinations were as good as they were going to get, which meant that AI performance is capped at 2024 levels. Both the overall prediction and the mechanism were wrong. This one seemed wrong at the time, in that I noted here in 2024 that you can make AI code halfway decently by just putting it in a loop and having it run until the code compiles and tests pass; I wasn't a heavy AI user at the time, but anyone who was using AI could see that there were ways that you could mitigate the hallucination rate which weren't being widely applied (this was before coding agents like codex and claude executed code and would check that tests pass, etc.).

  5. This is another one that also seemed untrue at the time. I'm not an ML person, but the moment someone told me what an RL environment was, within minutes, I thought of a bunch of ways one could generate synthetic data for improved training. I'm sure none of these were novel and they're things that AI labs are doing; my point is just that anyone who thinks about it for a few minutes can come up with a lot of ways that models could be improved even if there were no new data to find on the internet (not to mention that more effort could be used to get data that isn't just reddit comments or whatever the easiest to scrape content on the internet is).

  6. Note that this only scores predictions that have resolved. In that particular post Zitron also states that progress towards AGI will never happen, which is still both fuzzy and difficult to adjudicate and also one that you can never really reliably resolve as positive. Similarly, a prediction in a previous post that some company would have to add subscriptions isn't listed because it's open ended and not really resolvable as a negative (it was implied to have to happen soon, so is arguably wrong, but if one wanted to weasel out of it one could say that it will happen in the future).

  7. Here, Zitron also said, "I’ve realized now that it isn’t super useful to attach things to time (though I stand by my prediction) and thus I think it’s more useful to suggest what the terms of the bubble popping actually are". After this point, Zitron makes relatively fewer dated statements after this point and makes many more open-ended unfalsifiable statements. Perhaps a reaction to being wrong so frequently with his previous predictions?

  8. In a small piece of optimism, I'll say that this blog seems to have done ok despite not leaning into extremist positions and generally trying to avoid clickbait. This often means that, when I look at some data, I'll see something that looks like it would make for a really interesting viral hit piece, but then on looking more closely, it's actually a boring negative result, like when I ran this quick and dirty programming language eval, which originally appeared to show a very interesting result, which went away once I fixed the obvious eval bugs. Oh well. I'd like it if people published more boring negative results, so I published the boring negative result.

    I wouldn't be surprised if this blog is within an order of magnitude of traffic as Zitron's substack (server-side stats show 540k uniques for me in the past month, but who knows how many of those are bots with some but minimal Cloudflare bot blocking) despite Zitron writing much more frequently than me and pulling out every clickbait trick in the book, while I just occasionally post something when I feel like writing something up. Although my goal obviously isn't to get traffic, if we adjust for the level of time or effort, I don't think this blog does terribly compared to Zitron. Could Zitron have 5.4M monthly uniques? It's not impossible and it's hard to tell what these numbers mean with bot traffic, but for reference, The Economist has about 1.3M subs and the NYT has about 13M digital subs. If we hypothesize that 3/4 of uniques will be bot traffic, having an order of magnitude more traffic than this blog would put Zitron into the same class as The Economist, which doesn't really seem plausible.

    Ceteris paribus, I think Zalewski is right on the incentives, and I've seen a lot of people become caricatures of themselves as they lean into what drives the most engagement, but I think doing the opposite can work ok.

    For example, with a style that could be described as the opposite of clickbait, Simon Willison has written what I suspect is the most widely read blog among programmers for the past 3-4 years (in the same way that, at various times in the past, Joel Spolsky or Jeff Atwood or Steve Yegge seemed to be the most widely read programmer among programmers). Among programmers and other serious users of AI, I would guess that Willison has a larger audience than Zitron.

    However, it's true that Zitron has a kind of audience that Willison can never really get with his style. In the body of this post, we looked at common defenses of Zitron on forums where people use AI. That was pulled from forums where people use AI. If we look at the world at large, the comments look fairly different. For example, on the video that my friend mentioned, where Zitron repeatedly denies reality and the interviewer pushes back, the top comments at the moment are all in support of Zitron and they also just deny reality and claim that the places where the interviewer pushes back with a piece of reality are the interviewer being biased or just not knowing what he's talking about. Among the top comments, there seems to be little to no engagement with the facts of the matter; it's all mood affiliation. The comments remind me of what supporters say about politicians who use the gish gallop strategy and just say a bunch of outrageous nonsense. I could imagine Zitron running for office one day on the strength of his reality-denying popularity or becoming a demagogue who's a right-hand-man of someone in office, so Zalewski is right in that Zitron's appeal is not one someone is going to get by accurately describing what's happening in AI.

    But, while I don't know Willison and this could be totally wrong, my impression is that, like me, he's doing something he wants to do anyway and the audience just sort of happened despite him not trying to maximize his audience. When I say it works ok, I mean that he seems to be able to support himself working as a full-time open source developer due to the sponsorships he's gotten (which I would presume are generally because he has such a large audience), which seems like a good outcome even if this doesn't create the kind of mass appeal someone like Zitron can generate.

The Daily Front Page 6 of 28
Tuesday, September 1, 2026 The Daily Front No. #260901 — The Model That Lives on Disk
show hn

Show HN: Running 104GB Qwen3.8-Flash-Next on 48GB Mac with at ~12 tok/s

by carloslfu·▲ 182 points·91 comments·github.com ↗
Run Qwen3.8-Flash-Next on a Mac that can't hold it.

Run Qwen3.8-Flash-Next on a Mac that can't hold it. The model is a 125B-parameter mixture-of-experts, 104 GB on disk at 4-bit; slotstream streams it from SSD and runs it in whatever memory you give it. It's one Swift binary, no Python. It speaks the Ollama and OpenAI chat APIs, so your existing tools work unchanged.

on a 48 GB M5 Pro Warm decode ~12 tok/s Engine start ~2 s (only the 3.8 GB trunk loads) Peak memory 32 GB (auto-sized; you can cap it) Weights on disk 104 GB

Will it run on my Mac?

You need Apple Silicon, macOS 14+, and ~110 GB of free disk. Disk bites first: whatever your memory, a 512 GB Mac is the realistic minimum.

Auto-sizing never takes the whole machine. What each tier gets:

your Mac slotstream takes warm decode 8 GB 8.1 GB, the floor ~3 tok/s, and doctor warns it will page 16 GB 10 GB ~4 tok/s 24 GB 16 GB ~8 tok/s 32 GB 22 GB ~9 tok/s 48 GB and up 33 GB (more buys nothing; see Memory) ~12 tok/s

These rows come straight from slotstream doctor --sim-ram N, so you can reproduce them. Only the 48 GB row is measured on real hardware; the others are estimates from its curve, and smaller Macs also have slower SSDs. The middle column assumes nothing else is holding memory: with a browser open, auto takes less and says so in the plan it prints at startup (see Memory). Run slotstream doctor before downloading anything: it prints your machine's plan and whether the disk can hold the weights.

Install

curl -fsSL https://raw.githubusercontent.com/carloslfu/slotstream/main/install.sh | sh

Installs the latest release to ~/.slotstream/bin and puts it on your PATH. Re-run the same line to upgrade. To uninstall, rm -rf ~/.slotstream and remove the /usr/local/bin/slotstream wrapper or the PATH line the installer told you it added.

Releases are built by CI from the tagged commit with signed provenance, so you can verify an asset instead of trusting the download:

gh attestation verify slotstream-arm64.tar.gz --repo carloslfu/slotstream

Or build main yourself. Command Line Tools are enough, no Xcode needed:

git clone https://github.com/carloslfu/slotstream && cd slotstream
make build

The 104 GB download

The binary is small; the weights are not: 103.8 GB across 24 files, one time. serve and run offer the download on first use, or slotstream pull does it directly. Before transferring anything it prints the size, the destination, and your free disk, waits for a yes, and refuses outright if the disk can't hold it.

Your link sets the pace. pull opens eight TCP connections; a full install from a 1 Gbit/s datacenter link measured 112 MB/s, 16 minutes for the whole thing, which is the port. At 100 Mbps plan on ~2 h 20; at 25 Mbps, ~9 h. One connection alone is bounded by the round trip to Hugging Face — about 70 MB/s from a datacenter, 25 to 40 from a home link 100 ms away — which is why the count matters and why pull prints how many it is actually using. (Through 0.2.0 it ran on one connection whatever the flag said; see the changelog.)

Interrupting is safe: pull picks up where it stopped, redoing at most the few chunks that were in flight, and all 24 files are checked against sha256 hashes compiled into the binary, so a truncated or corrupted download can't reach the engine. The files come from a mirror of the pinned revision, with the original repo as fallback; the hashes are the same either way. pull --verify re-hashes an existing copy any time (8 s here).

Use it

First taste, no server:

slotstream run --prompt "why is the sky blue?"

For everything else, serve listens on port 11434 and implements the chat/generate subset used by Ollama clients and OpenAI SDKs:

slotstream serve
curl localhost:11434/api/chat -d '{
  "model": "qwen3.8-flash-next:4bit",
  "messages": [{"role": "user", "content": "hello"}]
}'

Open WebUI and the OpenAI SDKs are tested against this subset (the Ollama CLI is not there yet; see Status). Streaming, CORS, and the usual sampling options all work. What isn't supported (tools, images, JSON-schema output, logprobs) returns a clear 400 instead of being silently ignored. Every endpoint, field, default, and error is in docs/API.md.

Speed

Decode is the easy part: ~12 tok/s warm on a 48 GB Mac, and the tier table above says what smaller ones get. The slow axis is the prompt. All of it is processed before the first token appears, so 8,000 tokens wait about a minute on a 48 GB Mac and over three on a 16 GB one. Prompt plus completion is capped at 32,768 tokens (--max-context).

Within a conversation you only pay that once. Follow-up turns prefill just what's new, so time to first token stays flat as the chat grows: over eight turns at a 16 GB target, 6.0 s on the last turn instead of 25.8 s. Reused state isn't bit-identical to recomputing it, so a reply can occasionally differ where two tokens were nearly tied; --no-prefix-cache turns it off if you need exact reproducibility.

Decode has one more gear on machines with room to spare, and it is a small one. The model ships a draft head that predicts the token after next; with --mtp (default auto, new in 0.2.0) slotstream drafts the next token and verifies it in one two-token pass, and the draft is right 86% of the time (measured). It only pays where the expert cache is already near its best. On the dev Mac the 0.2.0 build, which drafted four tokens, lost at every cache size that fit, from ×0.55 at 20 experts per layer to ×0.96 at 57. One draft, the default now, reads ×1.13 at 57 and, at the 122 experts per layer a quiet 48 GB Mac runs with the head on, ×1.17 (10.1 → 11.8 tok/s, five pairs; two drafts ×1.13 there, four ×0.88). So auto turns it on only at that size, where the 1.6 GB it takes would otherwise buy experts past the plateau and costs nothing, and keeps it off below a 28 GB target. The ceiling is measured too: with every expert resident a two-token verify pass costs 1.17 single passes, which caps one-draft speculation at about ×1.4. The auto ceiling becomes 34.6 GB with the head on. It needs a one-time conversion that pulls 4.9 GB from the official release and writes a 1.5 GB mtp.safetensors next to the weights (Tools/mtp_convert.py, run from a clone with the repo's Python environment); without the file, everything runs with it off.

Memory

With no flags, slotstream sizes itself to your machine and tells you what it chose. This is a 48 GB Mac; it reads 52 because everything here counts in decimal GB:

slotstream memory plan (auto)
  device: 52 GB RAM (36.0 GB reclaimable now), 40.2 GB Metal working set
  target: 33.0 GB total for this process   (override: --memory-gb N | --max-ram-percent P)
  cache:  ~152 of 512 experts per layer  (7280 global slots = 20.1 GB pool)
  expect: ~32.0 GB peak, ~12 tok/s warm decode (est. from M5 Pro anchors)
  prefill: 4096 tokens per pass (~125 tok/s here; costs ~5.3 GB of the target)
  reuse:  up to 32768 tokens across 4 conversations (~1.2 GB), so a follow-up turn re-prefills only what is new

Auto takes the lowest of three limits (33 GB, 70% of RAM, and 2 GB under the Metal working-set limit) and sizes down further while other apps are actually holding memory. The 33 GB cap is the knee of the measured curve, not politeness: in a GB-at-a-time sweep, nothing between 34 and 84 GB decoded or prefilled any faster, so a 128 GB Mac gets the same plan a 48 GB one does. While running, slotstream re-checks every 15 s and resizes the cache between requests, shrinking under pressure and growing back once the pressure passes. Output is byte-identical across resizes.

To cap it yourself, --memory-gb G sets the total for the process (minimum 8.1, and it will go past 33 if you want to experiment). --max-ram-percent P moves the 70% share, and --experts-per-layer / --pool-gb size the cache directly. slotstream doctor prints the plan any of these would produce without loading anything.

How it works

Almost all of the model's bytes sit in two places: 68 GB of routed experts (512 per layer, 10 active per token) and a 32 GB n-gram table. The dense trunk is only 3.8 GB and stays resident. Experts are read with pread into a fixed pool of cache slots shared by all 48 layers, so hot layers borrow slots from cold ones.

Cache size changes speed, never output. Greedy decoding is byte-identical between a 4 GB cache and a 24 GB one, and that equivalence is a standing test.

Why not just mmap the file? MLX (Apple's ML framework) can't materialize part of a memory-mapped tensor: a top-10 expert gather evaluates all 512 experts of that layer, so an mmap path loads ~100 GB and dies. The stock mlx_lm.load() route took this 48 GB machine into 48 GB of swap without producing a token.

Status and limits

Working, and measured on one machine, an M5 Pro with 48 GB. The smaller tiers are estimates from its curve, not runs on real hardware.

  • One model, one process. v0 runs exactly qwen3.8-flash-next:4bit; the engine is built around its geometry, and pull knows no other name. A per-user lock allows one model process at a time.
  • macOS 14 and 15 have only had the installer exercised, not the runtime.
  • The Ollama CLI can't connect in 0.2.0. Its requests carry fields the release's strict validator rejects (empty name, system, template, options, and Ollama's empty-prompt "load" request), so ollama run stops before the first message. Fixed on main and verified with a real ollama run in both modes; it ships in the next release. curl, Open WebUI, and the OpenAI SDKs work today.

Docs

  • docs/API.md: every endpoint, accepted field, sampling default, and deliberate 400.
  • docs/TROUBLESHOOTING.md: port clashes, paging, slow decode, moving or verifying the weights.
  • docs/CLI.md: every command and flag, the memory knobs and their precedence, environment variables, where files live. (slotstream <command> --help carries the same text with more discussion.)
  • CHANGELOG.md: what each release changed.
  • PLAN.md: the design and the milestone tracker.
  • MEASUREMENTS.md: every number here with its method, including the experiments that failed.
  • llms.txt: a map of all of this for AI agents, with the commands, memory knobs, and API essentials inline; llms-full.txt is every doc above in one file.

Testing

Tools/verify.sh is the acceptance battery: weight provenance, goldens against a version-matched Python reference, byte-equality across cache sizes and live resizes, the speculative-decode gates, and a serving-robustness suite of inputs that used to crash the server. Tools/e2e_release.sh tests the other thing users actually touch: the curl | sh install and the binary it leaves behind. The parts that need no weights run in CI on every release build.

License

MIT. Sources/SlotstreamCore/Vendored/GatedDelta.swift is ported from mlx-swift-lm (MIT), and Tools/reference/ vendors the community qwen4_exp.py used as the test oracle. Weights come from pipenetwork/Qwen3.8-Flash-Next-MLX-4bit and remain under the Qwen community license.

The Daily Front Page 7 of 28
Tuesday, September 1, 2026 The Daily Front No. #260901 — Worlds in the Machine
article

Atlas: A World Model for Spatial Intelligence

by johnsutor·▲ 182 points·43 comments·worldlabs.ai ↗
World models generate, reconstruct, and simulate any possible world.

World models generate, reconstruct, and simulate any possible world. They understand how worlds appear, behave, and evolve so that we can render imagined worlds for creative users, simulate the real world in high fidelity, and help robots plan actions. At World Labs, we build these general purpose world models in pursuit of spatial intelligence.

Today we are introducing Atlas, our next-generation world model. Atlas is an omni model that we pretrained from scratch to natively operate on text, images, video, and 3D. It is a multimodal autoregressive diffusion transformer: all inputs are combined into a shared spatial context. Atlas uses that context to generate what comes next, staying consistent in 3D with everything it has seen and imagining what lies beyond it. Atlas is built to scale: its performance improves with increased training compute, and we expect this trend to hold as we continue scaling.

Atlas can perform a broad range of tasks spanning world generation, reconstruction, and simulation:

  • Camera-Controlled Generation: Atlas generates images and videos from one or more images with pixel-perfect camera control, outputting up to 1 minute of video at 1440p.
  • Spatial Reconstruction: Atlas reconstructs real world scenes from one to dozens of input images. It generates both image frames from novel views and explicit 3D outputs, outperforming state-of-the-art models specialized for 3D reconstruction.
  • Space-Time Simulation: Atlas models space and time from input videos, reframing videos for dramatic visual effects and enabling Real-to-Sim workflows for robotics.
  • Image Generation: Atlas generates images and 360 panoramas from text; it can follow complex prompts, render text, and generate a wide variety of visual styles.

Atlas will power future versions of Marble and other products from World Labs.

Request early access to Atlas

Camera-Controlled Generation

Atlas takes one or more reference images and generates new views at any camera position and angle you specify. Generated views match the content and geometry of the input images, smoothly extrapolating beyond them to imagine parts of the scene not visible in the inputs.

Atlas handles a broad range of scene types, visual styles, and camera motions.

Videos are generated from one to six input images with manually designed camera paths

Pixel-Perfect Camera Control

Atlas uses precise camera geometry as a native input type, going beyond coarse text-based instructions for camera control. This lets you frame every shot and control every motion.

In the examples here, Atlas generates a complete scene from a single input image. It uses the content of the input image along with its broad world knowledge to imagine what the scene should look like from new angles. For example, it generates the back side of the robot, and it guesses that there should be a grassy lawn next to the pool.

From a single image, Atlas generates views from any angle. Drag to change the view.

Generating with Spatial Context

Similar to an LLM, Atlas first encodes its inputs into a context, then generates outputs conditioned on the context. However, Atlas is unique because each image is grounded at a 3D position in space; this forms a spatial context.

Managing this spatial context unlocks entirely new kinds of creative control. For example, you can place two unrelated reference images in the context and position them in 3D space; Atlas then generates a world that smoothly interpolates between them.

These examples demonstrate the model's world knowledge and creativity; it imagines doorways, hallways, nooks, and other transitions between otherwise unrelated image pairs.

Select left and right frames to populate the spatial context, and Atlas stitches them together

Controllable Long Videos

Atlas lets you generate long videos with precise control by combining camera movement and spatial context management. You design every scene and every camera angle. This puts you in the director's chair: you are staging the scene, not pulling the lever of a slot machine.

In the example below, we generate a 1 minute video at 1440p resolution using a small number of reference images. We hand-design a camera path through the scene, and Atlas generates a coherent world. The rest of the videos on this page have been compressed to optimize page performance.

Spatial Reconstruction

Atlas reconstructs real-world spaces from one or more input images. It does not require special capture equipment or hundreds of dense views to faithfully reconstruct objects and scenes. We believe Atlas is a major step forward toward solving the problem of novel view synthesis from sparse input images, a decades-old fundamental problem in 3D computer vision.

Reconstructing from Multiple Images

Atlas can take a variable number of input views of a scene. When parts of the world are not visible in the input views, Atlas imagines a plausible way to fill in the gaps by drawing from its rich world knowledge.

But sometimes you do not want imagination; you might want an exact reconstruction of a real-world location. Passing more input images gives Atlas more context: the more it sees, the less it imagines. Atlas typically gives faithful reconstructions with as few as two or three images, outperforming state-of-the-art results by models specially trained only for 3D reconstruction. However, Atlas can also make use of over a hundred input images in its spatial context, allowing for faithful recreation of real world environments.

In the first example below, Atlas generates an aerial view of the scene from just a single ground-level photo. The garden visible in the single input photo is accurately recreated in the model output, but the rest of the scene is imagined. After adding a second real-world input image of the cottage next to the garden, the model's output shows both the garden and the cottage, but it still imagines the house to the left. After adding a third input image of the main house, the entire scene is accurately depicted.

In the second example, we build up Stanford's Main Quad piece by piece, beginning with the grassy main entrance and ending with the colorful mosaics decorating the facade of Memorial Church. Though Atlas only receives two to twenty-five ground-level input images, it can generate paths from aerial views flying far above the campus.

Reconstructing Diverse Paths

Atlas can generate many different trajectories through the same scene, giving new perspectives on the same input images. No matter how many times you change the camera path, the scene stays consistent.

In the example below, we show that given a small set of input images, Atlas can generate multiple camera paths through the same scene. Different camera paths can emphasize various parts of the scene or change moods by varying in speed, length, or complexity.

Atlas can generate many different paths through the same scene using a small set of input images.

Explicit 3D Outputs

In the results above you have seen Atlas output 2D images and videos, which are sufficient for some applications. But workflows in robotics, gaming, design, VFX, and beyond often require explicit 3D outputs. Atlas natively operates on both 2D image frames and 3D depth maps, enabling it to output worlds as point clouds or 3D Gaussian splats.

From one image, Atlas generates new views and 3D geometry, then converts to 3D Gaussian splats

From a single input image, Atlas produces a full 3D world by jointly generating new views and estimating their geometry. From a video of a real space, it predicts the depth of every frame and combines them into a 3D reconstruction. In either case, Atlas fills in regions that no camera ever saw.

Atlas can reconstruct 3D point clouds from input videos

Point clouds estimate a scene's geometry, but 3D Gaussian splats make it usable. Atlas fills the remaining gaps and turns the point cloud into a complete splat scene that renders on-device at high resolution and frame rates. This is the same representation used in Marble, enabling Atlas to integrate naturally with the rest of our products.

Space-Time Simulation

Atlas serves as a world simulator. It understands both the spatial structure of the world and how the world evolves over time. Combining its spatial and temporal abilities leads to new applications for VFX, robotics, and beyond.

Reframing Video

Atlas turns a handful of ordinary cameras into a "bullet time" multiview capture studio. With footage from as few as three cameras, Atlas can freeze time and reframe shots, letting you view events from impossible angles.

Real-world videos can be reframed from new camera angles without an expensive capture studio

Notably, these shots did not require professional photographers or specialized equipment. Each of them was filmed by a few engineers and researchers with ordinary cell phones on tripods and clamps that fit in a backpack. Atlas reconstructs the scene from three to five camera views, after which you can reframe shots however you like.

Behind the scenes: the clips above were captured using just a few cell phones and action cameras

Robotics Simulation

Atlas opens up new ways to scale Real-to-Sim for both navigation and manipulation.

You have already seen Atlas reconstruct a space in explicit 3D from a few images. For robotics, reconstruction is only half the job: as a simulated robot moves through space, Atlas also generates the RGB and depth data its sensors would observe along the way. The world and the robot's view of it come from the same model.

In these examples, we captured two large environments with a cell phone video, using 24 frames each for reconstruction. Scanning spaces like these traditionally requires elaborate and expensive equipment. We then simulate different kinds of robots navigating different paths and use Atlas to generate images from the perspective of the robot's body-mounted cameras.

Atlas reconstructs spaces and aids in simulating robot navigation

Robotic manipulation goes a step further. From a few casual recordings, Atlas aids in building a simulation that also captures how objects move and interact. Once a task is simulated, you can vary it easily: change the objects, their positions, the robot's motion, the lighting, and the background. The result is diverse training data and testing environments for robotics at scale.

Atlas enables Real-to-Sim from just a few real-world recordings, recreating physical interactions with rigid, articulated, and deformable objects while supporting controllable variations.

Image Generation

The primary focus of Atlas is world modeling, and every image is a window to a possible world. Though image generation is not its primary focus, Atlas is a capable image generator: it follows complex prompts, renders text, and generates a wide variety of visual styles.

Atlas also generates 360 images from text or image prompts, where again it can generate a wide variety of scene types and visual styles.

Technical Details

Model Architecture

Atlas is an omni model designed to handle many tasks and many kinds of input and output data in a single unified architecture, putting spatial control at the heart of the model. These goals require us to depart from standard architectures used by both LLMs and video models, and design a new base architecture to serve as the foundation of future world models.

Atlas is a multimodal autoregressive diffusion transformer. Its inputs are grounded in 3D space to form a spatial context, and it generates multimodal outputs conditioned on its context.

Atlas is a multimodal autoregressive diffusion transformer. It operates on multimodal sequences, generating each new element of the sequence one at a time. These architectural properties work together to achieve our goals, and taken together they enable a new paradigm of generation based on a spatial context. We unpack these ideas in turn:

  • Multimodal: Atlas can natively process many different data types. At present it can operate on text, images, camera poses, and 3D depth maps; videos are represented as sequences of images. Each image and depth map is conditioned on an explicit camera pose, making spatial control a central component of the architecture.
  • Autoregressive: Atlas operates on sequences of elements, where each element is one of the multimodal data types above. Each output is generated one at a time, conditioned on earlier parts of the sequence. This flexible design naturally adapts to a wide variety of tasks: each task is just a different kind of sequence, where inputs are followed by outputs.
  • Diffusion: Atlas is a rectified flow model that generates outputs by gradually denoising them. Diffusion models excel at modeling high-dimensional continuous data like images and video, and can naturally trade off speed and quality by varying the number of denoising steps used during inference.
  • Transformer: The transformer architecture consists primarily of large matrix multiply operations and is well-adapted to modern hardware. It is a robust backbone for world modeling.

Atlas is a blend of ideas from modern LLMs and video models. It can benefit from architectural, algorithmic, and systems advances used in both types of models.

Like an LLM, it is an autoregressive transformer, so it can take advantage of innovations used to serve and accelerate LLMs including KV-caching, cache-aware routing, disaggregated serving, and more. Like a modern image or video model, it is a latent diffusion model and can make use of algorithms such as diffusion distillation, classifier-free guidance, shifted noise schedules, and advances in VAE design.

Benchmarks

Atlas is an omni model for world modeling that performs many tasks. There is thus no single benchmark that fully captures its generality. We highlight quantitative evaluations of Atlas on two key tasks: camera-conditioned generation and 3D reconstruction. On both tasks it outperforms more specialized models.

We compare against a selection of top-performing video models for camera-conditioned generation. In each trial, we pair a single input image with a sequence of one to three cinematic camera motions (pan, truck, crane, etc.).

We prompt each model with a single input image and a target camera path. For Atlas, we encode the camera path using its native camera input format. Other models do not accept cameras as a native input format, so we describe the camera path in the input text prompt, using standard cinematic terms. It is possible that more sophisticated prompt engineering or creative multimodal prompts could improve camera following for some models, but we use text as it is the most common input modality for describing camera motions.

Third-party human raters judge which model better follows the intended camera path. These results confirm that Atlas outperforms recent video models at camera-controlled generation, and this advantage grows as camera trajectories become more complex.

We additionally evaluate Atlas on the task of 3D reconstruction from sparse input views. In each trial, the model receives a set of images and their camera poses, and predicts a 3D point corresponding to each input pixel. This problem has attracted much interest in the academic community, and many specialist reconstruction models have been developed in recent years.

Atlas is an omni model which performs both generation and reconstruction. Despite its generality, Atlas outperforms the best specialized open-source reconstruction models.

We evaluate on several state-of-the-art benchmarks for this task, reproducing the results for all baselines to ensure a common and fair evaluation protocol across all methods.

Model Scaling

Most progress in modern AI has been driven by scaling. Models improve in large part by scaling up simple algorithms to make use of more data and compute.

We see strong evidence that Atlas will continue to improve with scale. We pretrained Atlas from scratch on a large diverse corpus of multimodal data. Over the course of development, we trained a series of models of increasing size and training compute, and found that each new level of compute unlocked new model capabilities. We are confident that our future world models will follow this trend, dramatically improving their capabilities as we continue to scale.

Build with Atlas

Atlas is entering early access with select partners. If you would like to build with it, request access below and we will reach out. We are excited to see what you build, and to work with you to make Atlas the go-to world model for generating, reconstructing, and simulating any world.

Request early access to Atlas

We are also hiring across research and engineering to advance spatial intelligence.

This post was produced by the World Labs team. Please cite as:

@article{worldlabs2026atlas,
    author = {World Labs Team},
    title = {Atlas: A World Model for Spatial Intelligence},
    journal = {World Labs Blog},
    year = {2026},
    note = {https://www.worldlabs.ai/blog/atlas},
}
The Daily Front Page 8 of 28
Tuesday, September 1, 2026 The Daily Front No. #260901 — Compute, Robots, and Heavy Apps
article

GPU World

by simonpure·▲ 395 points·272 comments·gpuworld.org ↗

“The future is already here, it’s just unevenly distributed.”

— William Gibson

In our story contest, we ask people to imagine the future, evenly distributed.

The AI revolution has only just begun to affect humanity, and billions of humans have yet to so much as talk to a frontier LLM like Fable or Sol. This is in large part because compute limitations make it impossible to serve the highest-quality AIs to more than a relative handful of users: humanity is GPU-poor. Only a few million GPUs capable of efficiently serving frontier models are manufactured annually.

This will increase, however, as both hardware and software are scaled and optimized. Someday, such as in 2040, there may be available, for every human being, the performance equivalent of 'a B300 GPU for contemporary LLMs'.

What would this world be like?

What will our world be like when (not if) every human being has access to the equivalent of a Fable or Sol LLM 24/7/365? Will this lead to a panopticon of indefatigable AI surveillance? Will education be revolutionized by infinitely patient tutors? Will social media cease to exist as we know it? Will healthcare be revolutionized by world class AI doctors and personalized medicines? What will happen in the oft-ignored developing world?

AI as we know it already holds the potential to be far more transformative than the smartphone or perhaps even the Internet itself.

We invite you to imagine this 'mundane' future.

Premise

Imagine that AI frontier progress stops as of 1 September 2026: AI becomes faster and cheaper, but it never becomes superhuman or improves considerably across the board.

So the Singularity never happens—but GPUs keep getting made. By 2040, there may be the equivalent of 8 billion GPUs globally and everyone has access to a frontier LLM.

What happens in this 'business as usual' future?

Submit Your Story

Prizes ($100,000)

  • First Prize

    $40,000

  • Second Prize

    $20,000

  • Third Prize

    $12,000

  • Finalist Prizes

    7 prizes of $4,000 each

Selection Committee

The top 10 submissions after pre-screening will be read by the judges.

Timeline

  • August

    Submissions open

  • October

    Submissions close on October 31, 2026

  • December

    Winners announced

The three winning pieces will be published on gpuworld.org, paradigm.xyz, and gwern.net.

Submission Guidelines

  • Eligibility

    Open to everyone. One entry per person.

  • Genre

    Fiction or nonfiction.

  • Length

    1000 to 5000 words.

  • Format

    Markdown or PDF.

  • Copyright

    Submissions must be licensed CC BY-NC or under a freer license, so they may be republished.

  • Deadline

    11:59 PM PT, October 31, 2026.

  • AI Use

    LLM use is permitted, but discouraged; we remind participants that LLM use tends to reduce originality and writing quality, and the flaws are especially obvious when LLM outputs are read as a group---unskillful use of LLMs will reduce the odds of being the best entry. We request disclosure of AI use.

The Daily Front Page 9 of 28
Tuesday, September 1, 2026 The Daily Front No. #260901 — Compute, Robots, and Heavy Apps
article

The ChatGPT/Codex app bundles a full copy of LibreOffice

by timpera·▲ 310 points·146 comments·simonwillison.net ↗

I was poking around in my ~/.cache/ folder using OmniDiskSweeper when I spotted something interesting. The OpenAI Codex desktop app (since rebranded to just ChatGPT) has 1.7GB of stuff in there in a folder called codex-primary-runtime, including a full Python installation, a full Node.js installation, and native binaries for Poppler, git, and the LibreOffice open source office suite (which forked from OpenOffice.org in 2010):

Screenshot of a macOS disk usage app window in column view, titled "/Users/simon/.cache - 442.1 GB". First column: 356.8 GB huggingface, 82.5 GB uv, 1.7 GB codex-runtimes (selected), 609.0 MB datasette-sqlite, 298.8 MB rod. Second column: 1.7 GB codex-primary-runtime (selected). Third column: 1.7 GB dependencies (selected), 6.3 MB plugins, 4.1 kB runtime.json. Fourth column: 771.0 MB native (selected), 446.4 MB node, 440.6 MB python, 28.7 kB bin. Fifth column: 429.7 MB libreoffice-headless (selected), 187.9 MB poppler, 148.1 MB git, 4.7 MB libheif, 679.9 kB jxrlib.

The ~/.cache/codex-runtimes/codex-primary-runtime/plugins/openai-primary-runtime/plugins/documents folder includes skills which tell Codex how to find and use those binaries.

article

Launch HN: Nori Robotics (YC S26) – A low-cost humanoid robot for development

by AntonioLi·▲ 127 points·42 comments·norirobotics.com ↗

The most capable robot for $1,688

Everyday tasks

Nori can support you with day-to-day home tasks.

Teach Nori to support you in every way

Skills Marketplace

Train your Nori at home, share its skills anywhere.

Affordable and capable

App

The Nori Lab laptop app helps you train, operate, and manage your robot.

App

Arms

7+1 DOF, 1.5kg payload per-arm.

Arms

Lidar

12m range, 8-12Hz scanning frequency. Angular resolution 0.72° at 10Hz.

Lidar

Cameras

x4 720p RGB cameras, up to 30 fps, mounted on the grippers, head, and neck.

Cameras

Audio

Speaker and microphone for spoken commands.

Audio

Battery

6-8 hours battery life.

Battery

A line of NORI robots in front of an American flag in the San Francisco workshop

Based in the USA

Assembled in San Francisco

$1688

Full price, no deposit

Ships Fall 2026

The Daily Front Page 10 of 28
Tuesday, September 1, 2026 The Daily Front No. #260901 — The App Store Question
article

Fastpotify

by nreece·▲ 813 points·534 comments·fastpotify.rocks ↗

A lightweight Spotify client with local playback, library access, and Spotify Connect controls for Linux, macOS, and Windows.

Download

What is Fastpotify?

GitHub

Fastpotify showing the Late night focus playlist with the queue panel open, a track playing, and the library in the sidebar

Lightweight

A native binary with no browser engine. It starts in well under a second and typically uses 100–250 MB of RAM.

Spotify Connect

Play locally, gapless and at up to 320 kbps, or control playback on a speaker, phone, or TV from the same window.

Library and search

Browse playlists, Liked Songs, albums, artists, and podcasts. Search the catalogue and edit playlists you own.

Themes

Use light, dark, or system mode. Pages and the player bar can take their colour from the album art.

Winamp mini player

Ctrl+M opens a small player for classic Winamp 2 skins, with a spectrum analyser, equalizer, and playlist.

See it in action

MilkDrop

Run projectM’s MilkDrop visualiser in its own window, with fullscreen, preset packs, and keyboard controls.

Open the guide

Desktop controls

Keyboard shortcuts, MPRIS media controls on Linux, and a tray option that keeps music playing after you close the window.

Open source

MIT-licensed Rust built with egui and librespot. The docs explain its connections and stored credentials.

Read the source

repository

Play Store blocks AuroraStore, hurting GrapheneOS users

by erikvanoosten·▲ 489 points·203 comments·gitlab.com ↗

Description

Currently, Aurora Store, including Nightly (2026-08-31), returns a “&$Server busy, please try again later.” error when attempting to install an application via an anonymous account. It should be noted that, at least in my case, this error persists regardless of whether I use a VPN, clear the cache, refresh the anonymous account, force-close the app, or restart the device. Since I do not have a Google account, I am unsure if this issue is specific to anonymous accounts.

Expected Behaviour

I would expect Aurora Store to install the application as it normally does.

Actual Behaviour

Instead, all applications fail to install, and Aurora Store presents an error that reads: “&$Server busy, please try again later.”

Steps to Reproduce

  1. Search for an application
  2. Click the application's search result
  3. Click "Install"

Environment

  • Device model & codename: Fairphone 5
  • Android version: 16
  • Aurora Store version: 4.8.4
  • Nightly date: 2026-08-31
  • Account Type: Anonymous
  • Installation method: Session
  • OS: CalyxOS 7.2.4.20
The Daily Front Page 11 of 28
Tuesday, September 1, 2026 The Daily Front No. #260901 — The App Store Question
article

Introducing Ad Blocker for Firefox on iOS

by HieronymusBosch·▲ 365 points·125 comments·blog.mozilla.org ↗

Firefox

Introducing Ad Blocker for Firefox on iOS: More Control, Fewer Distractions

There’s only so much room on your screen. Pop-ups, overlays, and ads can take over fast, getting between you and what you came to do.

That’s where Ad Blocker for Firefox on iOS comes in: a built-in option that blocks many third-party ads and ad-related trackers before they load, helping reduce clutter and distractions while you browse.

How it works

Ad Blocker uses Apple’s WebKit Content Blocker technology and the EasyList filter list to determine what gets blocked. There’s no separate extension to install, and you can turn it on in Settings > Browsing > Ad Blocker. It’s off by default, so you decide whether to use it.

Ad Blocker won’t block every ad. Ads served directly by the site you’re visiting and ads shown in search results will still appear. Sponsored shortcuts and other sponsored content shown by Firefox when you open a new tab are separate from ads on the web pages you visit, so Ad Blocker doesn’t affect them.

Ad Blocker works alongside the privacy protections already built into Firefox, including Enhanced Tracking Protection, which blocks many trackers and limits tracking across the web.

More control over how you experience the web

On Desktop and Android, Firefox already supports a strong ecosystem of ad-blocking and privacy extensions, giving people the flexibility to choose the tools that work best for them. We value that ecosystem and will keep supporting it.

iOS works differently. Extensions aren’t available in the same way, and we know people want more options. Bringing ad blocking to Firefox on iOS meant building it directly into the browser.

Giving people choice in how they experience the web is important to us. Advertising helps fund much of the open web, supporting the publishers, creators and websites people rely on. We also know that ads can sometimes crowd the screen or interrupt what you’re trying to do.

That’s why Ad Blocker is optional: you decide whether it’s part of how you browse. It’s part of a broader approach across Firefox to give you more control over your experience, from the extensions you use to how AI shows up in your browser.

Try it

To turn on Ad Blocker, go to Settings > Browsing > Ad Blocker.

If you find an ad you expected to be blocked, a site that behaves strangely or something we should improve, let us know on Mozilla Connect.

Visit our Support page for more details on Ad Blocker for Firefox on iOS. For more on Firefox’s built-in privacy protections, check out How Firefox Protects Your Data.

The Daily Front Page 12 of 28
Tuesday, September 1, 2026 The Daily Front No. #260901 — Deadlines on Trial
article

Evidence of Fraud in an Influential Study About Procrastination

by Anon84·▲ 351 points·261 comments·datacolada.org ↗
A new paper in Psychological Science reports a failure to replicate Study 2.

A new paper in Psychological Science (.htm) reports a failure to replicate Study 2 of Ariely and Wertenbroch’s influential article entitled, “Procrastination, Deadlines, and Performance: Self-Control by Precommitment.” The original study, published in Psychological Science in 2002 (.htm), found that people performed better on a set of tasks when each task had its own externally imposed deadline than when people set their own deadlines or faced a single last-day deadline for all tasks. The paper has had a lasting influence. It has been assigned reading in many economics and psychology courses, and has more than 2,100 citations on Google Scholar.

Because this paper has been so influential, it is worthwhile to take a close look at the original study to try to understand why it did not replicate. We did that. This post – and the next one – is about what we found.

About 20 years ago, on April 20, 2006, one of the authors of the forthcoming replication, Kyle Hyndman, received the original data files in an email sent from [email protected] [1]. And about 3 years ago, on August 9, 2023, a week after Francesca Gino sued us for $25 million, we received an out-of-the-blue email from Hyndman in which he sent those files to us. We performed quick analyses of the data, and then had a conversation with Hyndman and his co-author, Alberto Bisin. In that conversation, they told us they were going to conduct a replication, and, finding ourselves busy with the lawsuit, we left it at that.

We recently learned that their replication was forthcoming in Psychological Science. And upon reading Footnote 14 of their paper, we also learned this:

. . . In October 2024, at the request of the editors, we shared with Dan Ariely an analysis of the contents from the file purportedly for their Study 2 and asked for permission to include a summary of it in the paper. Dan Ariely denied our request, arguing, among other things, that the files we received may not be the actual data. He did not subsequently provide us with any additional data from the original paper. Consequently, we are unable to supplement our replication exercise with any additional analysis of the files we received in 2006 or any other data.

This motivated us to return to this paper and fully analyze the original data for the two main studies. We conclude that the data in Studies 1 and 2 were tampered with. In two posts, we present the evidence that led us to this conclusion. Today’s post focuses on the study that failed to replicate (Study 2), and our next post is about Study 1.

Our assessment that the data were tampered with are based entirely on the analyses presented in our posts. Readers can review the evidence and draw their own conclusions.

To the best of our knowledge, Klaus Wertenbroch has never had access to any version of the data for any of the studies. And, we believe it is thanks to him that we do. When Kyle Hyndman reached out to the authors back in 2006, Klaus replied with this email [2]:

Our ResearchBox contains the data and code to reproduce all of the results in this post.

Finally, it should be noted that when we shared these posts with Ariely and Wertenbroch a few weeks ago, they reached out to Psychological Science to request that the article be retracted. As of this writing, that process is ongoing.

The Study That Did Not Replicate: Study 2 of Ariely and Wertenbroch (2002)

As noted above, Ariely and Wertenbroch explored how deadlines influence performance. In a context in which people had multiple tasks to perform, the authors hypothesized that people would perform better in the face of evenly spaced deadlines for those tasks, rather than when they were all due at the end.

The experiment involved an incentivized proofreading task. Each participant received three 10-page documents, each containing 100 “grammatical and spelling errors” (p. 222). Participants were tasked with finding and correcting those errors.

Sixty participants were randomly assigned to one of three conditions, exactly 20 participants in each condition:

Condition 1. Evenly Spaced Deadlines. One document was due each week, so after 7, 14, and 21 days.
Condition 2. Set Your Own Deadlines. Participants chose their own deadlines (within 21 days).
Condition 3. Last Day Deadline. All three documents were due on the final (21st) day.

The results perfectly and strongly supported the authors’ hypothesis. Participants given evenly spaced deadlines did much better, in terms of performance, delays, and earnings [3].

Do We Have The Original Data?

As a reminder, in 2023 Hyndman sent us files he received from [email protected] in 2006. There were three Excel files – data for a pilot study, for Study 1, and for Study 2 – all with file properties indicating that the data were “Last saved by” “Dan Ariely”.

With these files we are able to reproduce all nine means and all nine standard errors shown in the figure above, as shown visually in this footnote: [4]. We also successfully reproduce the six other means reported in the text [5].

Red Flags

In our analyses we identified four major red flags. We discuss each in turn.

Red Flag #1: The Effect Is Too Big

As shown in the reprinted figure above, Ariely and Wertenbroch report a perfect pattern of results, for all three dependent variables, with a sample size of only 20 per condition. The effects are also large. Extremely, implausibly large.

Consider the proofreading performance results. Participants with Evenly Spaced Deadlines made an average of 136.1 corrections, whereas those with the Last Day Deadline made an average of only 71.1 corrections, about half as many. This effect has a Cohen’s d = 2.5, indicating that the condition means are 2.5 standard deviations apart. The correlation between experimental condition and number of corrections is r = .79.

To appreciate that this effect is just too big, consider it in the context of other effect sizes. An effect size of d = 2.5 is larger than obvious effects we notice in everyday life, effects that can easily be seen with the naked eye. For example, it is much larger than the effect of gender on height (men are taller: d ≈ 1.8) and on number of shoes owned (women own more shoes: d ≈ 1.2; see Colada[18]). It is also larger than some manipulation checks. For example, Petty and Cacioppo (1984) report that participants exposed to messages containing nine arguments said that they encountered more arguments than people exposed to messages containing three arguments. This has to be true. And it was true, but only to the tune of d = 1.49 [6]. It is not plausible that deadlines influence proofreading performance more strongly than the number of arguments influences the perceived number of arguments.

Effect sizes greater than or equal to 2.5 are not impossible – they are sometimes observed with manipulation checks – but they are extraordinarily rare for non-obvious psychological findings, particularly for a measure like proofreading error detection, which is likely to be noisy, and highly variable across people.

Another way to appreciate the enormousness of this effect is to look at the distribution of the dependent variable across conditions. The figure below shows that they barely overlap. For instance, whereas nobody in the Last Day Deadline condition made more than 100 corrections, 90% of the participants in the Evenly Spaced Deadlines condition did:

Red Flag #2: Duplicate Observations

If looking at Figure 2 you thought, “wait, why are there so many red bars with 2s?”, good catch. That is weird. The 2s represent people who found exactly the same total number of corrections made across three tasks. But it’s actually weirder than that. These participants found not just the same number of corrections in total, but also made the same number of corrections for each of the three separate proofreading tasks.

Here is a screenshot of the original data file, formatted and sorted to be easier to digest:

We see that 18 of the 20 participants in the Last Day Deadline condition had a “Corrections Twin”, another participant who found exactly the same number of errors for each of the three proofreading tasks. Interestingly, these twins have ID numbers that are exactly 10 positions apart (e.g., subject S1 and subject S11 are twins; so are S7 and S17; etc.). (There were no error twins in the other two conditions.)

The existence of so many of these twins – and all of them in only one condition – is inconsistent with these data being real.

Red Flag #3: Things That Should Be Very Highly Correlated Aren’t Correlated At All

At the end of their study, Ariely and Wertenbroch purportedly “asked participants to evaluate their overall experience [of the proofreading task] on five attributes: how much they liked the task, how interesting it was, how good the quality of the writing was, how good the grammatical quality was, and how effectively the text communicated the ideas contained in it” (p. 223). These questions were answered on scales ranging from 0 to 100. The replicators asked the same questions to their participants.

You might expect these judgments to be correlated. For example, if someone says they liked the task, you might also expect them to say that it was interesting.

In the replication, this was (super) true. Controlling for experimental condition, the partial correlation between liking and interest was, quite sensibly, close to perfect [7]:

But in the original data, this relationship was not only imperfect; it was not there at all. Participants who said they liked the task more did not say that they found the task to be more interesting:

In total, there are five subjective measures. In the replication, the (partial) correlations among these five measures range from +.63 to +.92. They are all large and very highly significant (ps < 0.0000024). In the original data, these correlations range from -.29 to +.18, and none of them are both positive and significant. This is very strange.

The problem is not limited to these subjective measures. Consider the fact that people did three very similar proofreading tasks, each with 100 mistakes. Surely, we’d expect people who do better on one task to also do better on another, nearly identical task. That simple fact should manifest in extremely large correlations between performance on one task and performance on another. And in the replication data it does, as the correlations range from +.74 to +.90. But in the original data it doesn’t, as the correlations range from +.03 to +.27.

Finally, consider that participants were asked to report how many minutes they spent on each of the three tasks. Again, we’d expect those who said they spent more time on one task to be more likely to say they spent more time on another, nearly identical task. And so we’d expect these variables to be very highly correlated. Once again, within the replication data they were – the correlations ranged from +.79 to +.95 – and within the original data they were not – the correlations ranged from +.05 to +.17.

The correlations we have reviewed in this section are essentially just sanity checks. Does liking correlate with interest? Does performance correlate with performance? Does reported time spent correlate with reported time spent? Sane data pass these checks. Insane data do not. The replication data are sane. The original data are not.

Red Flag #4: No Rounding In Self-Reported Minutes

As you’ll recall from a minute ago, Ariely and Wertenbroch (2002) purportedly asked participants to “estimate how much time they had spent on each of the three tasks” (p. 223). When people provide estimates like this, they tend to round. They usually say “20 minutes” or “30 minutes” instead of “17 minutes” or “32 minutes”. And, indeed, when the replicators asked people to report how many minutes they spent on each of the three tasks, 85% of them gave a round number:

This is what we’d expect humans to do.

But in the original data, they did not do that. Only 11.7% of estimated minutes were round, consistent with the 10% you’d expect by chance alone:

This is not what we’d expect humans to do.

Conclusion

We are unable to generate a benign explanation for all of the anomalies presented here. The original findings are too large and yet they do not replicate; there are duplicated observations; correlations that should be very strong are often non-existent; and values that should be rounded are not rounded. Based on this evidence, we believe the data for Study 2 of Ariely and Wertenbroch (2002) were severely tampered with or fabricated to produce the desired results.

In our next post, we will share analyses of the Study 1 data file that Hyndman received from [email protected]. That experiment is quite different. Our analyses are quite different. But our conclusions are quite similar.

Author Feedback

About 6 weeks ago, on July 20th, 2026, we shared drafts of our posts with the original authors (Dan Ariely and Klaus Wertenbroch), the replication authors (Kyle Hyndman and Alberto Bisin), and the editor-in-chief of Psychological Science (Simine Vazire).

Klaus Wertenbroch sent us a response in which he begins by thanking Hyndman and Bisin for having done the replication. He restates that he never had access to the data for any of the studies. He distinguishes between demand for precommitment*,* a finding that was replicated by Hyndman and Bisin and which is consistent with earlier work by him and others, and the effectiveness of such precommitments in these specific studies, which did not replicate. And he indicated that he has asked the editor to retract the paper.

You can read his response in full (PDF).

Dan Ariely did not reply to any of the three emails we sent him. But on August 7th, he wrote on LinkedIn (htm) and on his personal website (htm):

“. . . Recently, I was made aware that data underlying a 2002 paper about deadlines and procrastination that I co-authored contained serious anomalies. The documentary record I have at my disposal today about those experiments isn’t sufficient to answer the questions that have been raised, and more than two decades, and hundreds of experiments later, my memory is similarly insufficient. Moving forward, my responsibility lies in ensuring accuracy – in updating the record on these experiments and, along with my co-author, cooperating with the journal that first published our paper to support their reviews and retraction processes.”

Neither LinkedIn nor Dan’s website allowed archive.org to save copies; so we screen recorded both pages (mp4).

Kyle Hyndman and Alberto Bisin asked us to include this statement:

“As stated in the posts, in April 2006, we received three data files attached to an email sent from Dan Ariely’s MIT email account, with no stated restrictions on their use. In August 2023, we provided those files to Uri Simonsohn, Joe Simmons and Leif Nelson to obtain their professional assessment. We did not participate in Data Colada’s analysis or in drafting the posts. Our independent replication relies on newly collected data and stands on its own methodological findings. Questions concerning the provenance or integrity of the historical files should be addressed by Data Colada, Dan Ariely, and the institutions with appropriate responsibility for those questions.”

Simine Vazire indicated that she is only allowed to say that Psychological Science is considering “best next steps regarding the 2002 paper in accordance with COPE guidelines.”

Footnotes

  1. PDF copies of this email – and related emails (including the one posted below) – are available in both the replicatiors’ ResearchBox (3063) and this blogpost’s ResearchBox (7135). Throughout this post, we refer to these emailed files as the “original” data files, though previous (unaltered) versions may have existed.

  2. As mentioned in our previous footnote, this email is included in Hyndman and Bisin’s ResearchBox (.htm). Wertenbroch gave us permission to include this email in our post.

  3. Ariely and Wertenbroch refer to their performance measure as “errors detected”, but in this post we refer to it as “corrections made”, and we have altered the figure below accordingly. We made this change because readers may confuse participants’ “error detection” in a proofreading task with our own “error detection” in the original dataset. This nomenclature hopefully prevents any such confusion. We have also altered the the figures below so they contain our preferred and more intuitive condition names.

  4. We reproduce the published article’s figures with the spreadsheets we received. For example, here is Panel A of Figure 2:

  5. We do not, however, reproduce all of the F values in the published paper, the statistical results purportedly obtained when testing the differences among those means. We were puzzled as to why until we discovered a previous version of the manuscript still posted to INSEAD’s website (.htm). It is INSEAD working paper “2001/09/MKT” and the title page indicates that it is “Under review, Psychological Science.” It turns out that while some of the means changed between the working and published versions of the paper, the F values remained the same. But when means change, so do F values that compare those means. Therefore, the F values are necessarily wrong in at least one of the two versions of the paper. We believe the authors changed the means reported in the paper but forgot to update the F values. We provide evidence of this in this Appendix: (.pdf).

  6. From Petty and Cacioppo (1984, p. 74): Participants “were asked, ‘About how many arguments did the author put forth in favor of the advocated proposal?’ Subjects were free to record any number they wanted, and those exposed to the nine-argument messages claimed that there were significantly more arguments in their messages (M = 6.60) than did subjects exposed to the three-argument messages (M = 3.68), F(1*,* 158) = 87.17, p < .0001.” You can compute Cohen’s d from this F-value: d = 2*sqrt(F/dferror) = 2*sqrt(87.17/158) = 1.49.

  7. In this section, we examine and report partial correlations that control for experimental condition. A genuine association between liking a task and finding it interesting should emerge within conditions, not merely across them. By contrast, someone fabricating data to create condition differences could inadvertently induce positive raw correlations simply by assigning higher values to both variables in one condition than another.

The Daily Front Page 13 of 28
Tuesday, September 1, 2026 The Daily Front No. #260901 — The Last Great Mechanic
article

American Airlines mechanic Azriel “Al” Blackman has died

by NaOH·▲ 376 points·150 comments·simpleflying.com ↗
His life defined aviation maintenance for the whole airline.

sky harbor airport 12-28-2025 Phoenix, AZ USA  American Airlines Boeing 777-200 N751AN with 75 nose art dedicated to AL Blackman, landing on runway 26 at Phoenix Sky Harbor Intl. Airport. Credit: Shutterstock

The aviation industry is mourning the loss of one of its most remarkable figures after American Airlines icon American Airlines mechanic Azriel "Al" Blackman passed away at the age of 100. His life defined aviation maintenance for the whole airline. According to announcements shared by American Airlines and members of the aircraft maintenance community on social media , Blackman died on the evening of Friday, July 24, bringing to a close an extraordinary career that spanned more than 80 years.

Blackman's story was unlike any other in commercial aviation. As the Guinness World Records holder for the longest career as an airline mechanic, he dedicated his entire professional life to a single airline, witnessing the industry's transformation from flying boats to modern twin-engine widebodies. His career became synonymous with American Airlines itself, inspiring generations of aircraft maintenance technicians.

The Man Behind Aviation's Longest Career

American Airlines team member Azriel “Al” Blackman, an Aviation Maintenance Technician (AMT) Crew Chief based at New York’s John F. Kennedy International Airport (JFK), celebrated 75 years of service with the airline today. At a ceremony at JFK today, the airline surprised him by dedicating a Boeing 777 in his honor. Credit: American Airlines

Blackman's aviation journey began during WWII on July 17, 1942, when he joined American Export Airlines, the predecessor of American Airlines, at just 16 years old. Fresh out of Aviation High School in Manhattan, he needed his mother's permission to accept the job and started as an apprentice in the sheet metal shop, earning just 50 cents per hour.

His first assignments reflected a very different era of aviation. Working at New York's LaGuardia Marine Air Terminal, Blackman helped maintain Sikorsky flying boats, even wading into Flushing Bay to secure seaplanes before they were hauled into the hangar for maintenance. Over the following decades, he would work on virtually every aircraft type operated by American Airlines, culminating with the Boeing 777 fleet.

American Airlines paid tribute to Blackman following his passing, describing him as a "true aviation legend" whose commitment, professionalism and mentorship influenced countless colleagues throughout the company. The airline also recognized that his impact extended well beyond maintenance, making him one of the most respected employees in its history.

Azriel "Al" Blackman loved his job, as he described it on the AA website, alluding to a famous saying about passion that makes labor feel light:

"When you like what you do, it's not work."

From Flying Boats To Boeing 777s

The Sikorsky S-43 and its military cousin, the JRS-1 were designed in the late 1930s as airliners and military personnel transports. Fifty-three were built between 1937 and 1941 for civilian customers and the U.S. Navy. A total of three S-43/JRS-1s survive today. Sometimes called the "Baby Clipper," S-43s served with Pan American Airlines and other airlines on shorter routes for which the larger flying boats were not needed. The Navy purchased seventeen of these aircraft with two of them going t Credit: Wikimedia Commons

Very few people in history experienced aviation's technological evolution as closely as Blackman. He began his career before the jet age, when American still operated flying boats across the Atlantic. Over the next eight decades, he witnessed the arrival of piston-powered airliners, the dawn of jet travel, the introduction of widebody aircraft, and eventually modern long-haul twinjets.

His longevity was equally remarkable inside the airline. While countless aircraft types entered and left American's fleet, Blackman remained a constant presence, serving as an Aviation Maintenance Technician Crew Chief based at New York John F. Kennedy International Airport (JFK). His knowledge became invaluable to younger mechanics, many of whom regarded him as both a mentor and a living history book.

Even in his later years, Blackman's routine reflected an unwavering dedication to his work. Although his official shift began at 5:00 a.m., he typically arrived at the hangar before 3:00 a.m. His late wife, Delores, reportedly joked that he should "go to work" and "play with your friends," highlighting how much he genuinely enjoyed spending time alongside fellow mechanics, as reported by the Aviation Circle.

The Boeing 777 That Carries His Legacy

Incheon, South Korea - Jul 27 2025: A Boeing 777-200ER of American Airlines touches down at Incheon International Airport in South Korea. Credit: Shutterstock

American Airlines also honored Blackman's unprecedented career in spectacular fashion during his 75th anniversary with the company in 2017. In what the airline said was a first for one of its mechanics, it dedicated a Boeing 777-200 to him, registration N751AN, complete with a commemorative plaque recognizing what it described as the longest career in aviation history as an Aviation Maintenance Technician.

The aircraft quickly became one of the airline's most recognizable widebodies among maintenance employees. Mechanics affectionately nicknamed it "7BK" in tribute to Blackman, while Guinness World Records officials attended the dedication ceremony to formally recognize his record-breaking achievement. American also arranged for a US Capitol flag to be flown in his honor, underscoring the national significance of his accomplishment.

Blackman's milestones did not stop there. He continued working for another five years, celebrating an astonishing 80 years with the airline in 2022. Few employees in any industry have remained with a single organization for so long, making his achievement virtually unique in commercial aviation history.

The Daily Front Page 14 of 28
Tuesday, September 1, 2026 The Daily Front No. #260901 — The First Imagined Being
article

Lion-man

by gurjeet·▲ 140 points·84 comments·en.wikipedia.org ↗
The Lion-man of Hohlenstein-Stadel is a prehistoric sculpture.

The Löwenmensch figurine, also called the Lion-man of Hohlenstein-Stadel, is a prehistoric sculpture discovered in Hohlenstein-Stadel, a German cave, part of the Caves and Ice Age Art in the Swabian Jura UNESCO World Heritage Site, in 1939. The German name, Löwenmensch, meaning "lion-person" or "lion-human", is used most frequently because it was discovered and is exhibited in Germany. It is an anthropomorphic figurine combining a human-like body with the head of a cave lion (Panthera spelaea).

Determined by carbon dating of the layer in which it was found to be between 35,000 and 41,000 years old, it is one of the oldest-known examples of an artistic representation and the oldest confirmed statue ever discovered.[1] Its age associates it with the archaeological Aurignacian culture of the Upper Paleolithic.[2] An example of zoomorphic art, it was carved out of mammoth ivory using a flint stone knife. Seven parallel, transverse, carved gouges are on the left arm.

After several reconstructions that have incorporated newly found fragments, the figurine stands 31.1 cm (12.2 in) tall, 5.6 cm (2.2 in) wide, and 5.9 cm (2.3 in) thick. It is currently displayed in the Museum Ulm, in the city of Ulm.

Side view showing the transverse gouges on the left arm

History

Systematic excavations at Hohlenstein-Stadel cave began in 1937 under the direction of historian Robert Wetzel [de].[3]

The discovery of a fragmented mammoth ivory figurine was made on 25 August 1939 by geologist Otto Völzing [de].[4] The start of World War II just one week later meant that the fieldwork was left incomplete and analysis of the finds was not undertaken. The excavation trenches were back-filled with the same soil in which the ivory had been found.[5] For approximately thirty years, the fragments lay forgotten at the nearby Museum Ulm. It was not until archaeologist Joachim Hahn started an inventory and assembly of more than 200 fragments that a figurine with animal and human features began to emerge.[5]

Wetzel continued to spend summers digging at the site until 1961,[6] and further finds of ivory were made on the cave floor in the 1970s. In 1982, paleontologist Elisabeth Schmid combined the new fragments with Hahn's reconstruction, correcting some errors and adding pieces of the nose and mouth which emphasized the figurine's feline characteristics.[5][a]

In 1987, a comprehensive restoration began in the workshops of the Landesmuseum Württemberg by Ute Wolf in cooperation with Schmid. During the work, which took more than six months, it was realized that the figurine was only about two-thirds complete. The back was severely damaged and the legs were missing some ivory lamellae. The ears, eye-holes, two-thirds of the mouth and nose, and the back of the head were preserved. To fill gaps in the head and body, a reversible substance consisting of a mixture of beeswax, artificial wax, and chalk was used.[9]

In 2008, further excavations were carried out in the cave. All layers were sifted systematically, which led to many minute fragments being discovered. The first new adjustments were simulated virtually so that fragments could be added without having to disassemble the original recreation.[10][b]

In 2012, a second restoration was begun in the workshops of the State Office for the Preservation of Historical Monuments in Esslingen under the leadership of Nicole Ebinger-Rist. The figurine was disassembled into its individual parts and newly discovered fragments were added along with the old ones, allowing further completion of areas of the head, back, and right side of the body, and artificial additions used during the first restoration were discarded.[12] The Löwenmensch figurine grew in height from 296 to 311 millimetres.[13] Work was completed in late 2013.[12]

Interpretation

Some researchers have ascribed sexual characteristics to the object. Initially, the figurine was classified as male by Hahn who suggested a plate on the abdomen could be a flaccid penis. Schmid later classified this feature as a pubic triangle;[4] however, from examination of new parts of the sculpture, she proposed that the figurine was that of a woman with the head of a female cave lion.[14][15] Male European cave lions appear to have largely or completely lacked the distinctive manes of their African counterparts, so the absence of a mane could not determine categorically that the figurine was that of a lioness, and a debate about its sex ensued among some involved in the research and the popular press. Kurt Wehrberger, of the Museum Ulm, stated that the statue had become an "icon of the feminist movement".[4]

After the 2012–2013 restoration, it was realized that the triangular platelet in the genital area was processed all around, separating it from the figurine. A fracture point suggests that originally it may have been square in shape, which most commonly could be interpreted as a stylized male sex organ.[16] Debate continues, even though an objective determination of the sex of the Löwenmensch figurine may be impossible.

The Löwenmensch figurine lay in a chamber almost 30 metres (98 ft) from the entrance of the Stadel cave, accompanied by many other objects. Bone tools and worked antlers were found, along with jewellery consisting of pendants, beads, and perforated animal teeth. The chamber was probably a special place, possibly used as a storehouse, hiding-place, or maybe as an area for cultic rituals.[17]

A similar but smaller lion-headed human figurine was found in Hohle Fels.[18] Archaeologist Nicholas Conard suggested that "the occupants of Hohle Fels in the Ach Valley and Hohlenstein-Stadel in the Lone Valley must have been members of the same cultural group and shared beliefs and practices connected with therianthropic images of felids and humans" and that "the discovery of a second Löwenmensch lends support to the hypothesis that Aurignacian people practised a form of shamanism."[18]

The figurine shares certain similarities with later French cave paintings, which also show hybrid creatures with human-like lower bodies and animal heads, such as the "Sorcerer" from the Trois Frères in the Pyrenees or the "Bison-man" from the Grotte de Gabillou in the Dordogne.[19][20]

Debate exists as to whether the figurine depicts a lion or human-lion hybrid figure at all; with similarities to a standing bear, and the unreliable nature of the reconstructions cited.[21]

Manufacture

The carving of the figurine from hard mammoth tusk would have been a complex and time-consuming task.[c] A similarly sized tusk found in the same cave has marks that "indicate that the skin and thin bone around the tooth cavity of the upper jaw were cut through to the surface of the tooth, which was then exposed for detachment with a hammer. The tip was harder and had to be removed by wedging and splitting."[23]

Wulf Hein and Kurt Wehrberger conducted an experimental replication with the kinds of stone tools available at the time. Removing the base of the tusk took ten hours. The body was carved with a steep-fronted scraper; the burins requiring regular resharpening. Several tools were needed to separate the torso from the insides of the arms while shaping the head and shoulders, which involved difficult cutting across the grain of the ivory, often requiring two hands on the tool. The basic shaping is estimated to have taken around 200 hours, and in total the recreation likely took more than 370 hours.[d] Jill Cook, Curator of Palaeolithic collections at the British Museum, suggests that "unless the sculpture was created slowly at odd moments over several months, someone as skilled as an artist may have been excused from other subsistence tasks to work specially on this piece."[23]

In his October 2017 BBC Radio 4 series Living with the Gods, Neil MacGregor asked Cook

... so why would a community living on the edge of subsistence, whose primary concerns were finding food, keeping that fire going, protecting children from predators, allow someone to spend so much time away from those tasks?[24]

She replied that it was about

... a relationship to things unseen, to the vital forces of nature, that you need to perhaps propitiate, perhaps connect to, in order to ensure your successful life.[24]

Exhibitions

In 1956 Robert Wetzel, leader of the original excavation expedition, endowed all his archaeological finds from the Lonetal to the city Ulm. They are still owned by the city today.[25] The original figurine is part of the permanent archaeological exhibition of Museum Ulm since its first re-assembly in the 1970s.

Special exhibitions

  • 2009/2010 Der Löwenmensch – Das Experiment (“The lion-man – The experiment”), Museum Ulm: The focus of this exhibition was a reproduction of the figurine carved under historically authentic conditions. The facsimile figurine was made by archaeotechnician Wulf Hein using flintstone tools in 320 hours. In 2010 this exhibition was shown in Stadtmuseum Erlangen together with additional paleolithic ivory artifacts from the University of Erlangen collection.[26]
  • 2013/2014 Die Rückkehr des Löwenmenschen. Geschichte – Mythos – Magie (“The Return of the Lion Man: History – Myth – Magic”), Museum Ulm: exhibition of the newly re-assembled figurine. An accompanying booklet was published in English.[27]
  • 2025/2026 Fabelhaft! Der Löwenmensch und seine Nachfahren, Kunsthalle Weishaupt: exhibition of the reassembled figurine together with portrayals of other hybrid creatures from human cultural history.[28]
The Daily Front Page 15 of 28
Tuesday, September 1, 2026 The Daily Front No. #260901 — A Garden of Forking Paths
article

Borges Labyrinth in Venice reopens to the public

by gone35·▲ 118 points·35 comments·wallpaper.com ↗
Inspired by Borges’ short story ‘The Garden of Forking Paths’, he vowed to create a labyrinth.

The Labirinto Borges, part of the Giorgio Cini Foundation, reopens in Venice this week after an extensive renovation

Labirinto borges maze giorgio cini foundation

(Image credit: Courtesy Fondazione Giorgio Cini)

In 1979 the press secretary of the British embassy in Buenos Aires, Randoll Coate, woke from a dream where his friend, the Argentine writer Jorge Luis Borges had died. Writing to Susana Bombal (who had originally introduced them in the 1950s), he unveiled his idea: inspired by Borges’ short story ‘The Garden of Forking Paths’, he vowed to create a labyrinth in his honour whenever the day came.

True to his word, after the writer’s death in 1986, the maze maker’s design came to life in Argentina. Due to its success, Borges’ widow Maria Kodama began to propose Coate’s project to places dear to the writer. Venice, itself a winding labyrinth of a city, held a special place in his heart.

Labirinto Borges at the Giorgio Cini Foundation

the Labirinto Borges Maze, Venice

(Image credit: Courtesy Fondazione Giorgio Cini)

In 2011, twenty-five years after Borges' death, the labyrinth opened on the island of San Giorgio Maggiore. Set in a weathered cordon steel skeleton to protect it from the elements, vivid green boxwood plants wind in long articulated lines, revealing the form of an open book, the writer’s surname splayed out in elegant type on each 'page', mirrored and intertwined.

the Labirinto Borges Maze, Venice

(Image credit: Courtesy Fondazione Giorgio Cini)

Looking closer, images emerge: a question mark, a cane, a pair of eyes. 'The labyrinth’s perimeter asks us with its symbols and signs to remember the reasons why it was planted,' says Pedro Memelsdorff, the musical director of the Giorgio Cini Foundation. '[It’s] a monumental tribute to Borges for the future and for future generations.' Today, forty years after the writer’s death, thanks to the generous support of PwC Italia the labyrinth has been fully restored.

Restoring the original maze

the Labirinto Borges Maze, Venice

The Labirinto Borges under construction, 2011

(Image credit: Courtesy Fondazione Giorgio Cini)

'We might think of it as a mosaic, with the plants as the tiles that live or die,' says Renata Codello, the Secretary General of the Foundation. The process of creating and maintaining the maze is 'a constant action … always living and present.' 165 dead or severely damaged boxwood plants were removed, while the irrigation system responsible for maintaining the verdant green was inspected and repaired where faults were found. Only time will tell if the older plants accept the new ones; like a body after an organ transplant, things need to unite. In the 15 years since it was first planted change was inevitable, with this revisitation the 1,1150 metres of winding path have been trimmed and unified. Weeds painstakingly removed by hand, the soil revamped, refreshed, and enriched with additional fertilization.

Labirinto borges maze giorgio cini foundation

(Image credit: Courtesy Fondazione Giorgio Cini)

Thanks to a collaboration with the Italian Union of the Blind and Visually Impaired, the restoration also includes the addition of a tactile map at the start of the maze. Borges lived for the last years of his life in darkness, his sight gradually deteriorating over time until he went blind at the age of 55. So this small but powerful gesture creates another moving link to the author, allowing visitors to understand the structure of the labyrinth and feel its twists and turns before they take that first step into the unknown, and experience it for themselves.

The Labyrinth is now open to the public
visitcini.com

The Daily Front Page 16 of 28
Tuesday, September 1, 2026 The Daily Front No. #260901 — The Glow of Signal
article

Magic eye tube

by peter_d_sherman·▲ 90 points·23 comments·en.wikipedia.org ↗
A magic eye tube gives a visual indication of the amplitude of an electronic signal.

EM34 tuning eye

EM84 tuning indicator

A magic eye tube or tuning indicator, in technical literature called an electron-ray indicator tube,[1] is a vacuum tube which gives a visual indication of the amplitude of an electronic signal, such as an audio output, radio-frequency signal strength, or other functions.[1] The magic eye (also called a cat's eye, or tuning eye in North America) is a specific type of such a tube with a circular display similar to the EM34 illustrated. Its first broad application was as a tuning indicator in radio receivers, to give an indication of the relative strength of the received radio signal, to show when a radio station was properly tuned in.[1]

The magic eye tube was the first in a line of development of cathode ray type tuning indicators developed as a cheaper alternative to needle movement meters. It was not until the 1960s that needle meters were made inexpensively enough in Japan to displace indicator tubes.[2] Tuning indicator tubes were used in vacuum tube receivers from around 1936 to 1980, before vacuum tubes were replaced by transistors in radios.[3] An earlier tuning aid which the magic eye replaced was the "tuneon" neon lamp.[3][4]

History

The magic eye tube (or valve) for tuning radio receivers was invented in 1932 by Allen B. DuMont (who spent most of the 1930s improving the lifetime of cathode ray tubes, and ultimately formed the DuMont Television Network).[5][6][7]

The RCA 6E5 from 1935 was the first commercial tube.[8][9]

The earlier types were end-viewed (EM34), usually with an octal or side-contact base. Later developments featured a smaller side-viewed noval B9A based all-glass type with either a fan type display or a band display (EM84). The end-viewed version had a round cone-shaped fluorescent screen together with the black cap that shielded the red light from the cathode/heater assembly. This design prompted the contemporary advertisers to coin the term magic eye, a term still used.

There was also a sub-miniature version with wire ends (Mullard DM70/DM71, Mazda 1M1/1M3, GEC/Marconi Y25) intended for battery operation, used in one Ever Ready AM/FM battery receiver with push-pull output, as well as a small number of AM/FM mains receivers, which lit the valve from the 6.3 V heater supply via a 220 ohm resistor or from the audio output valve's cathode bias. Some reel-to-reel tape recorders also used the DM70/DM71 to indicate recording level, including a transistorized model with the valve lit from the bias-oscillator voltage.

The function of a magic eye can be achieved with modern semiconductor circuitry and optoelectronic displays. The high voltages (100 volts or more) required by these tubes are no longer in modern devices, so the magic eye tube is obsolete.

Method of operation

Schematic diagram of a magic eye indicator tube: a = anode, k = cathode, g = grid, b = deflection

A magic eye tube is a miniature cathode ray tube, usually with a built-in triode signal amplifier. It usually glows bright green, (occasionally yellow in some very old types, e.g., EM4) and the glowing ends grow to meet in the middle as the voltage on a control grid increases. It is used in a circuit that drives the grid with a voltage that changes with signal strength; as the tuning knob is turned, the gap in the eye becomes narrowest when a station is tuned in correctly.

Internally, the device is a vacuum tube consisting of two plate electrode assemblies, one creating a triode amplifier and the other a display section consisting of a conical-shaped target anode coated with zinc silicate or similar material. The display section's anode is usually directly connected to the receiver's full positive high tension (HT) voltage, whilst the triode-anode is usually (internally) connected to a control electrode mounted between the cathode and the target-anode, and externally connected to positive HT via a high-value resistor, typically 1 megaohm.

When the receiver is switched on but not tuned to a station, the target-anode glows green due to electrons striking it, with the exception of the area by the internal control-electrode. This electrode is typically 150–200 V negative with respect to the target-anode, repelling electrons from the target in this region, causing a dark sector to appear on the display.

The control-grid of the triode-amplifier section is connected to a point where a negative control voltage dependent on signal strength is available, e.g. the automatic gain control (AGC) line in an AM superheterodyne receiver, or the limiter stage or FM detector in an FM receiver. As a station is tuned in the triode-grid becomes more negative with respect to the common cathode.

Use in radios

6G5 Magic eye tube

The purpose of magic eye tubes in radio sets is to help with accurate tuning to a station; the tube makes peaks in signal strength more obvious by producing a visual indication, which is better than using the ear alone. The eye is especially useful because the AGC action tends to increase the audio volume of a mistuned station, so the volume varies relatively little as the tuning knob is turned. The tuning eye was driven by the AGC voltage rather than the audio signal.

When, in the early 1950s, FM radio sets were made available on the UK market, there were many different types of magic eye tubes with differing displays, but they all worked the same way. Some had a separate small display to light up indicating a stereo signal on FM.

The British Leak company used an EM84 indicator as a very precise tuning-indicator in their Troughline FM tuner series, by mixing the AGC voltages from the two limiter valve grids at the indicator sensing-grid. By this means accurate tuning was indicated by a fully open sharp shadow, whilst off-tune the indicator produced a partially closed shadow.

Common types

In U.S. made radios, the first type issued was the type 6E5 single pie shaped image, introduced by RCA and used in their 1936 line of radios. Other radio makers used the 6E5 as well until, soon after, the less sensitive type 6G5 was introduced. Also, a type 6AB5 aka 6N5 tube with lower plate voltage was introduced for series filament radios. Type number 6U5 was similar to the 6G5 but had a straight glass envelope. Zenith Radio used a type 6T5 in their 1938 model year radios with "Target tuning" indicator (resembling a camera iris), but was abandoned after a year, with Ken-Rad manufacturing a replacement type. All these types use a 6-pin base with two larger pins for filament connection.

Several other "eye tubes" were introduced in U.S. radios and also used in test equipment and audio gear, including the octal-based types 6AF6GT, 6AD6GT and 1629. The latter was an industrial type with 12 volt filament looking identical to type 6E5. Later U.S. made audio gear used European tubes like EM80 (equivalent to 6BR5), EM81 (6DA5), EM84 (6FG6), EM85 (6DG7) or EM87 (6HU6).

Other applications

Magic eye tubes were used as the recording level indicator for tape recorders (for example in the Echolette [de]), and it is also possible to use them (in a specially adapted circuit) as a means of rough frequency comparison as a simpler alternative to Lissajous figures.

A magic eye tube acts as an inexpensive uncalibrated (and not necessarily linear) voltage indicator, and can be used wherever an indication of voltage is needed, saving the cost of a more accurate calibrated meter.

At least one design of capacitance bridge uses this type of tube to indicate that the bridge is balanced.

The magic eye tube appears on the cover of My Morning Jacket's 2011 album Circuital. The tube is shown almost fully lit.

References

  1. 1 2 3 Spangenberg, Karl R. (1948). Vacuum Tubes (PDF). New York: McGraw-Hill Book Co. pp. 723–724.
  2. "Precise dB Monitoring With Eye Tubes" (PDF). Archived from the original (PDF) on 2012-04-17. Retrieved 2011-07-29.
  3. 1 2 Erb, Ernest (2009-07-27). "History of tuning indicators; meters, graphs, Magic Eye, LED". Radio History forum. RadioMuseum.com. Retrieved 2014-08-23.
  4. Radio Museum: Tuneon
  5. US 2098231, DuMont, Allen B., "Cathode ray device", published 1932-05-28, issued 1937-11-09
  6. US 2163256, DuMont, Allen B., "Cathode ray tube", published 1934-09-18, issued 1939-06-20
  7. David Weinstein, The Forgotten Network: DuMont and the Birth of American Television. Temple University Press, 2006, p. 11
  8. Herbert M. Wagner of R.C.A. invented the familiar form of the "magic eye" tuning vacuum tube: US 2051189, Wagner, H. M., "Tuning indicator tube", published 1935-06-27, issued 1936-08-18
  9. Radio Museum: 6E5

Further reading

The Daily Front Page 17 of 28
Tuesday, September 1, 2026 The Daily Front No. #260901 — Bench Repair Department
article

Refurbishing a Tektronix TDS7104 Oscilloscope

by jwise0·▲ 103 points·50 comments·tomverbeure.github.io ↗
The unit was in excellent cosmetic shape, but the price tag of $700 was way out of line.

Introduction

A little over a month ago, I ran into a Tektronix TDS7104 at the Silicon Valley Flea Market, where else?

TDS7104 in the trunk of my car

Other than some dirty buttons, a few smudges here and there, and the usual assortment of calibration and asset tracking tags, the unit was in excellent cosmetic shape, but the price tag of $700 was way out of line: as I write this a try-before-you-buy TDS7104 can be had on Craigslist for the same price.

But Paul, the seller/liquidator, has a habit of saying “I’ll make you a deal” and he did before I even asked: $300. That’s still a lot by flea market standards, but a pretty good price for a TDS7104… if you can get it to work.

At home, the scope powered up right away and it booted straight into the main scope application. Other than a screen that was way too dim, everything seemed fine.

TDS7104 at first power up

But when I tried it again a few hours later, it got stuck at the BIOS screen with a CMOS battery error.

TDS7104 bootup error

In this blog post, I go over the steps I took to get the scope back in top shape.

The TDS7104

The TDS7104 is a 4-channel oscilloscope with 1 GHz bandwidth and a maximum sample rate of 10Gs/s, though that’s only possible when using 1 channel. The sample rates drop to 5 Gs/s for 2 channels and 2.5 Gs/s for 3 or 4. Even by today’s standards, the specs exceed those of hobbyist class oscilloscopes, think Rigol and Siglent, though there’s a price to pay in terms of weight, 39 pounds, and volume: they’re as wide and deep as the earlier TDS700 series, for example, and much taller. The TDS7054 is its little brother, figuratively speaking only. In the same chassis, it has a 500 MHz bandwidth and 5 Gs/s.

Unlike more advanced TDS7xxx models, the 7104 and 7054 have BNC connectors instead of custom Tektronix ones that require probes or adapters with prices that exceed today’s price of the scope itself.

Introduced mid 2000, these scopes initially ran Windows 98 but they must have upgraded soon after to Windows 2000 Pro Embedded, because that’s what mine has and it has components with a late 2000 timestamp.

The PC motherboard has the little-used NLX form factor. Mine was a RadiSys SF810 with a Socket 370 and a 100 MHz front-side bus. Originally, these scopes shipped with a dog slow 550 MHz Celeron, I got lucky with a 850 MHz Celeron. The fastest compatible Celerons with 100 MHz FSB go up to 1.4 GHz, but they’re pricy. You should be able to find 1.1 GHz versions for around $20 on eBay.

Unlike my Agilent 54831, the 640x480 LCD screen has resistive touch control which makes it possible to use the advanced scope features without the need to connect a mouse.

In addition to the PC motherboard, there is a PowerPC-based controller board that runs VxWorks like many other Tektronix products of that time, and a large acquisition board.

TDS7104 with advanced jitter analysis license

In addition to a few hardware options such as a 4M/channel sample memory, up from a 500k default, there are plenty of software options for advanced measurements: jitter testing, USB certification testing, etc. Both the software and hardware options can be enabled with a license key. To the suprise of no one, that protection scheme was hacked long time ago…

According to the labels on the chassis, my scope came from the PSD lab at Cypress Semiconductor, where it was used for things like measuring high-bandwidth signals such as the battery current on the Apple TV Remote. :o)

TDS7104 Apple TV Remove measurements

Common Failures

As always, you’ll find a bunch of hobbyists trying to revive this kind of scope on the EEVblog forum, Youtube and some blogs. Here are the most common failures:

  1. PC motherboard CMOS backup battery dead
  2. PowerPC backup battery dead
  3. Hard drive dead
  4. Power supply capacitors leaking

I was lucky and only had to deal with issues 1 and 3, sort of.

A dead PowerPC backup battery will give you considerably more work than what’s described in this blog post. After booting up the TekScope application will show the splash screen, but it will hang there forever. You will need to:

  • Take apart the scope even more and take out all the PC components: floppy, HD, CDROM drive, motherboard.
  • Replace the top cap of the Dallas DS9034 NVRAM with a new battery.
  • Connect with RS-232 to the PowerPC controller board.
  • Enter a bunch of values to store in the NVRAM.

You can detailed step-by-step instructions here. You should also check out this repair video by Feedbackloop.

A dead power supply is another common problem. It often will prevent the scope from booting up at all. There are plenty of discussions about this on the Eevblog forum, here is one that has the reverse engineered power supply schematic attached. Often, all that’s needed is to replace some leaking capacitors.

I didn’t have to do any of that…

Make an Image of the Hard Drive

Whether the machine boots or not, your first step should always be to make an image of the hard drive, a 6 GB IBM Travelstar in my case. Like my Agilent 54831, I thought that I’d have to open the case to access the drive, but you can just push on the spring-loaded black cover in the back and pull the drive sled out1. Nice!

TDS7104 hard drive sled

Remove the drive from the sled, plug it into a USB-to-IDE adapter, and extract the data. On Windows, I use HDD Raw Copy Tool.

TDS7104 HD out of sled

I often use Linux for this kind of maintenance, but since this scope is a Windows 2000 machine, I ended up needing a bunch of Windows-only tools.

The Travelstar HD was running on fumes, because HDD Raw Copy Tool ran into a number of corrupt sectors during the copying operation. I was lucky, the impacted files were related to the French Windows 2000 manual, but it shows the importance of making an image of the drive ASAP.

CR2032 Backup Battery Replacement and Display Brightness

You’ll need to open up the case to get to the PC motherboard CR2032 backup battery. See the next 2 sections for that.

CR2032 on motherboard

After installing the new CR2032, the scope booted back up again, but the TekScope window had some weird corruption and waveforms didn’t render right. This was because the Chips & Technologies 69000 graphics card settings had been changed to a 256 color palette mode. It needs to be set to True Color 24-bit mode.2

Display Settings

Notice the presence of 2 video cards: an Intel 810 integrated graphics card and the Chips & Technologies 69000. The latter is responsible for driving the LCD screen. It has special hardware to render oscilloscope waveforms in overlay mode: they are sent by the acquisition board to the video memory through DMA3 without CPU intervention.

While we’re on the topic of the display: after installing the CR2032, the LCD display was still very dim, to the point that I was researching replacement CCFL backlight tubes. That turned out to be entirely unnecessary: the TDS7104 doesn’t have a way to control the intensity of the LCD backlight. The previous users must have used it in a dark lab and dialed down the brightness by adjusting the gamma settings in Windows:

Windows gamma settings control

Do NOT Remove the Front Panel

I’m putting this section before the Disassembly one to make sure those with a low attention span get the message: chances are high that you don’t need to remove the front panel.

And that’s good because, unlike the TDSnnn series scopes, the front panel has some plastic tabs that are very easy to break. That said, even if you do break them (I did!), the result is not catastrophic and you should be able to put the panel back firmly where it belongs with no one noticing a thing.

Service manual figure 6-3: Trim Removal (Click to Enlarge)

The TDS7000 Series Service Manual makes it sound easy enough:

To remove the trim ring, slide the flat end of a soldering aid into the side slot on the trim ring. Press in, then lift up to hook it underneath, then pry up.

And from the pictures, it’s as if you can remove the front panel without removing anything else. That just didn’t work…

The front panel consists of multiple click tabs: 1 on the left side, 1 on the right and then a bunch at the top and the bottom. So far so good. However, the left and right side also have 2 slide tabs that go into the metal rails. If you lift the left and right tabs too much, these plastic slide tabs break off.

So you need to be very careful to make sure that you don’t lift the plastic trim too much, and that you slide the panel out while it stays parallel with the display.

Or… you don’t touch it: you can do all PC maintenance, including replacing the floppy drive, without removing the front panel.

Scope Disassembly

I will continue my tradition of documenting the disassembly of test equipment in too much detail because nobody else does it. Even though the service manual technically describes how to do it, a few pictures go a long way to make it easier.

To access the inside of the scope, you need to remove more than 30 screws. On the plus side, they’re all Torx-15 screws and they’re all the same length, so you don’t need to worry about keeping track of which screw goes where.

Still, it takes a while and it validated my recent purchase of this cordless screwdriver, recommended by Shrirar over at The SignalPath.

Unbutton the accessory bag

Remove the accessory bag

This took me longer to figure out than I want to admit: you can just unclick the bag from the chassis, but the buttons can be very tight and if you’re not careful the fabric can tear. Use a flat-head screwdriver right next to each button to lever it off.

Put the scope upright on its back feet

It’s an unusual arrangement, but the easiest way to dismantle the scope is by putting it on its back feet: you don’t need to remove any screw from the back!

TDS7104 on its back feet

Let me once again sing the praises of a sturdy equipment cart: it’s so much easier to walk around the cart than to muscle around bulky, heavy test equipment on a table.

Remove the top panel

4 screws through the accessory bag buttons (“snap studs”) fix the top panel to the chassis.

TDS7104 remove top panel

After removing this panel, you could remove the side panels already, but I found it much easier to remove the bottom panel next.

Remove the bottom panel and loosen the black front connector trim

Next, remove the 5 screws of the bottom panel as well as 3 screws that keep the black trim of the front BNC connectors in place.

TDS7104 bottom panel and connector enclosure

The black trim doesn’t need to be completely removed, only loosened because otherwise it will soon be in the way of some other screws.

The bottom panel shall now be removed though. Just slide it down a bit and take it off.

In the picture above, you can see 2 screws that aren’t marked in red. That’s because they don’t keep the bottom panel in place. But if you feel like it, you might as well remove them now too.

Remove the handle and side panels

TDS7104 side panel with handle

With the bottom panel gone, the side panels are a breeze to remove after unscrewing the handle.

I lied: these 2 screws are different than the others. But they’re a different color and impossible to get wrong.

Remove the 2 sheet metal parts

With the outer covers removed, you’re now staring at the sheet metal RF protection enclosure. It consists of 2 parts, each part covers 2 sides. Remove all the screws, take off the bottom part and then the top.

TDS7104 sheet metal top

TDS7104 sheet metal right

Note how some of the bottom screws are hidden underneath the BNC connector cover. That’s why you had to remove its 3 screws of the black trim.

TDS7104 sheet metal bottom

TDS7104 sheet metal left

Congratulations! For those who didn’t keep track: you’ve removed 32 screws!

After removing the panels, you now have access to the acquisition board at the bottom and the PC motherboard at the top:

TDS7104 acquisition board

TDS7104 PC motherboard

One side has nothing but cooling fans, but from the other side you can see the power supply and an RS-232 port that you will need to connect if the backup battery of the PowerPC controller board expires.

TDS7104 right side

If you need access to those items, you’ve only done the easy disassembly part. On my unit, both the PSU and the controller backup battery were fine so I was done.

Note on the picture above that the front panel has been removed. You do NOT have to do this for pretty much all restoration cases! And you really shouldn’t.

Reinstall the bottom sheet metal cover

All of my work on the scope was on the PC motherboard and I had to put the scope back in its horizontal position. To make sure that I didn’t accidentally damage the acquistion board, I put the bottom sheet metal cover back in its place.

TDS7104 bottom sheet metal back in place

A Failed Attempt at Switching over to an SSD

I’ve been using CompactFlash cards in the past to replace ailing hard drives. They work, but unless you buy a more expensive “industrial” card, they don’t have built-in wear leveling support. That is not a problem on a Rohde AMIQ that runs DOS, but on an OS like Windows with swap space, it could be4. So this time, I chose a 64 GB mSATA SSD ($35) and an mSATA SSD to IDE 44 Pin 2.5” adapter ($15)5.

64 GB mSATA SSD and IDE converter

The standard way to move away from a failing hard drive to an SSD is to once again use HDD Raw Copy Tool to write the image to the SSD and that is that. I tried that with the 64 GB SSD, and while the scope got past the first-stage boot process, it errored out during the second stage when it tries to bring up the Windows GUI with a STOP: c0000218 {Registry File Failure} error.

Registry File Failure

Older systems often had issues with partitions larger than 32 GB, so I bought a 32 GB mSATA SSD instead, $3 cheaper for half the capacity, but I got the same error.

Just copying the drive image to an SSD worked fine for others, but for me it was a dead-end that I spent many hours trying to get around. I eventually decided to reinstall all the software from scratch, which was a whole other adventure.

Reinstalling from Scratch: Windows 2000 Pro or Windows XP?

I had wanted to avoid reinstalling the OS from scratch because I expected to run into a bunch of driver issues, but in the end I had no choice. While a number of people have reported that Windows XP can work on some of the TDS7104 motherboards, I decided to stick with Windows 2000 Pro because I know that works and I didn’t have a pressing need for more functionality, whatever that might be.

The scope has Windows 2000 Pro Embedded, but I wasn’t able to find an installation disk for that and the regular version works fine too. The ISO file can be downloaded from the Internet Archive.

The license key that’s printed on the back to the scope does not work with the regular Windows 2000 Pro. The Internet Archive one has a key that works, and other valid keys are just a Google away, but I didn’t even need one: I was never asked for a license key during the Win2k installation on the scope.

The standard way to install Win2k Pro is with a CDROM drive. Unfortunately, the drive didn’t work which meant I had to open the whole machine again to install a replacement drive.

Not All TEAC CD-224E Drives are the Same

The TEAC CD-224E laptop drive in my TDS7104 got detected just fine by the BIOS and in Windows, but when you inserted a disc in the drive, neither the BIOS nor Windows could read from it.

Since the RadiSys motherboard doesn’t support booting from USB stick, I decided to replace the TEAC CD-224E laptop drive with a ‘new’ one that I got from eBay for $20.

Unlike the hard drive, the CDROM drive can’t be removed without opening up the TDS7104, but once the case is open, the effort is minimal. I first removed the floppy drive to have a bit more maneuvering freedom with the cables, but it’s not really necessary.

Unplug the CDROM IDE cable

CDROM IDE cables

Remove 2 screws

CDROM screws

The CDROM drive sits in a metal enclosure with a small adapter PCB that converts the CD-224E 50-pin slimline IDE connector to a standard PATA/IDE connector.

Old and new CDROM drive and converter PCB

I tested the broken drive with the adapter PCB and my USB-to-IDE dongle on my laptop to make sure the issue was with the drive and not the CDROM disc, and that didn’t work, as expected. With the new CD-224E/dongle combo, my laptop could read the installation CD just fine, but when I installed the new drive in the TDS7104, the BIOS couldn’t even detect the drive! I tried every BIOS setting under the sun, but no luck.

There are many versions of the CD-224E, all with the same dimensions and slimline IDE interface, but clearly they don’t all behave the same. The version of the broken one is version A93 (2000), the new one is CD0 (2005). You can find A93 drives on eBay, but $69 is way too high for something that I’d be using exactly once.

Installing Windows 2000 Pro on an Old Machine through a Virtual Machine

(Another dead-end)

It is allegedly possible to install Windows 2000 Pro on an old machine without CDROM and USB port by using a virtual machine. The process is convoluted:

  • mount the installation CDROM ISO and the SSD onto the virtual machine.
  • go through the first phase of the installation process until asked to reboot.
  • now move the SSD to the old the machine (the scope) and proceed with the installation there.

I once again spent a few hours getting this to work, but the scope never managed to make it to the Windows installation GUI.

Burning the Windows 2000 Pro Installation Disk onto a USB Stick

Alright, so I’m running out of options and USB is about the only storage interface left. The scope can’t boot from a USB stick directly but there is a way around that.

Let’s first create a bootable USB stick with the Win2k installation ISO on it.

Most of the time, you can use a utility like Balena Etcher to burn a CDROM ISO onto a USB stick, but of course that doesn’t work for the Windows 2000 Pro installation CDROM.

Instead, you need to use WinSetupFromUSB to prepare the USB stick:

  • Download, install, launch
  • Select the USB stick as target
  • Select Auto format with FBinst and use the FAT32 file system
  • Add to USB disk: Windows 2000/XP/2003 Setup
  • Select the mounted Win2K Pro ISO drive as source
  • Press “GO” to copy Win2K Pro onto the USB stick

Booting from USB Stick with a Plop Boot Manager

Plop Boot Manager makes it possible to boot from a USB stick on machines that don’t support it.

It goes like this:

  • copy the plpbt.img image from the plpbt-5.0.15.zip archive to a floppy disk with a tool like WinImage, Rawrite32, or RawWrite for Windows.6
  • boot the Plop Boot Manager from floppy disk.
  • the boot manager has a USB mass storage device driver
  • select USB as boot device

I tried hard to avoid the floppy disk route because my experience with floppy drives on old test equipment has been abysmal: none of them worked. Having no choice, I tried to copy the boot manager image with my USB floppy drive and… that didn’t work either. All these years the USB floppy drive, freshly bought from Amazon, was the culprit!

Since the scope still worked fine with the IBM HD, I used its own floppy drive to put the image onto the floppy disc and that worked.

Plop boot manager selection menu

After setting the BIOS to allow booting from floppy, the scope booted into the Plop Boot Manager just fine and it was able to boot the USB stick with the Windows 2000 Installation ISO.

Plop doesn’t support USB hubs. The RadiSys motherboard has only 1 USB port which will be occupied by the USB stick, so you’ll at least need a PS/2 keyboard to do anything.

Installing Windows 2000 Pro

With the empty 32GB SSD plugged into the scope, the installation of Windows 2000 Pro was uneventful. There are 2 phases: the first one uses text mode and primarily copies all the necessary drivers onto the SSD. The machine then reboots and continues the installation in Windows GUI mode from the SSD, though the USB stick is still needed in a later stage.

The TDS7104 has a bunch of specialty hardware that needs dedicated drivers, but those are not needed to get the OS up and running.

At long last, I was able to see this image:

Windows 2000 Professional installation complete

Installing Special TDS7104 Drivers

There is a great GitHub repo with a bunch of TDS7000-series software, including this Drivers directory. The README.md says that the driver should work for Windows 98 and XP, but the Chips and Technologies video driver definitely did not work for Win2k!7

I used Driver Collector to extract drivers from the original hard drive and that worked fine. You can find these drivers here.

Device manager missing drivers

The 4 specialty drivers are for these components:

  • Front panel

    This is the USB Device that’s listed under “Other Devices”

  • Texas Instruments PCI-1225 CardBus Controller

    You need to install this driver twice, once for each port. Windows installed a default PCI-1225 driver for this, but that one doesn’t work, hence the exclamation mark next to it. The name of the driver .inf file is unsup.inf, for unsupported? Confusing, but that’s the one to use.

  • PCI2PCI bridge

    That’s the Other PCI Bridge Device.

    Other PCI driver

  • Chips and Technologies 69000 video driver

    The default Windows driver for the C&T 69000 is what makes the screen work when running Windows, but it’s not sufficient to render measured signals in the TekScope application. For that, you need to update to the Chips and Technologies (Asiliant) 69000 driver.

    C&T driver selection

Installing Tektronix Firmware

The TDS7104 and TDS7054 firmware v2.5.5 can be freely downloaded from the Tektronix website. The installation was painless, just launch the executable.

The TDS7104 has a convoluted architecture where the PowerPC on the controller board can access files on the hard drive of the regular PC that are located in the c:\vxboot directory. Since the controller backup battery on my scope was still in good condition, I didn’t have to do anything special: the vxboot directory was created automatically during the firmware installation.

Installing TekFonts

The Tektronix scope application uses custom TrueType fonts to render some of the symbols screen, e.g. the rising edge trigger symbol. Without those fonts, it will show some Greek characters instead.

To fix that, you need to download the tekfonts.zip file, unzip it, and install the 3 fonts.

Despite rendering those Greek characters, those font files were already installed on the new system, so I had to delete them first and reinstall the new file. Things looked good after that.

To delete or install the fonts, do Start -> Settings -> Control Panel -> Fonts.

Install fonts

The Scope is Working!

And with that, I finally had a working TDS7104 with SSD!

TDS7104 with IBM Travelstar in front

The time from pressing the power button to having a waveform on the screen was much lower too: from 2min50s down to 1min35s.

Re-enabling the Existing License

One thing was missing, though: the advanced jitter license option.

The same GitHub repo that I mentioned earlier also has an unlock options directory with scripts to enable and validate license key features. On the Eevblog forum, plenty of people have been able to use it, but it’s not as user-friendly as other license key schemes.

Most of the time, license keys are additive, with one license key per feature that must be enabled. On the TDS7104, there is 1 license key that enables all features at once.

The validate script shows how that works with the license key and serial number of my scope:

./validate.py BREHZ9885D3MNKXHHYQCQRGQRW7C
E1 91 73 BF F7 7B E4 C5 52 3D C7 3A E1 9E 71 8F 76 01
44 2F 54 00 00 C0 1B 79 48 00 00 00 00 00 00 00 00 A8 16 30 00 10 00 00 00 00 00
This key is for UID 1BC00000542F (S/N 21551, model TDS/DSA/DPO7104):
CRC: 4879
Key is valid, active options:
00 00 00 00 00 00 00 00 08 00 00 00 00 00 00 00 00 00

We can see how that long string of gibberish contains:

  • the serial number 21551
  • the model number TDS/DSA/DPO7104
  • a UID that is really just a combination of the serial number and the model
  • a CRC
  • an 18-byte or 144-bit bitmask

I can recreate the license key by feeding these parameters back in the generation tool:

./gen.py tds7104 B021551 000000000000000008000000000000000000
XBGDV-K8GDM-KH7X3-979Y9-ZZ593-9ZRZZ-4837X-9VV5Z-T9HB

I had to join the 18 bytes into one 72-digit hex number.

The license key that comes out doesn’t match the original one, but after entering it into my scope, it worked just the same:

New Jitter Analysis - Advanced license

The scope is very forgiving about the license keys: upper case, lower case, dash or no dash, it all seems to work. You can even reduce the number of hex digits in the license enable mask to a certain extent, and the license key will still work:

./gen.py tds7104 B021551 000000000000000008000000000000
7GWUZ-RRRMK-59LYT-978Y8-GZD93-8ZQGZ-C836X-8CVD

What remains is the question which bit maps to which feature? This post in the eevblog forum has you partially covered here:

########################################################################
4   
# options masks/names/descriptions, conversion functions
5   
6   
# 01 - 1M
7   
# 02 - 2M
8   
# 04 - 3M
9   
# 06 - 2M 2A
10   
# 08 - 4M
11   
# 00 00 00 00 00 00 04 - USB
12   
# 00 00 00 00 00 00 20 - JT3
13   
# 00 00 00 00 00 00 00 80 - ET3
14   
# 00 00 00 00 00 00 00 00 08 - JA3
15   
16   
# 00 00 05 00 00 00 00 00 00 10 - ASM DDRA DJA
17   
# 00 44 00 00 00 00 02 08 - SM ST J1 J3E
18   
# 04 40 00 00 00 00 06 C0 10 - 3M JT2 USB2 ST
19   
# 04 44 FF 03 00 00 8D A3 EF FF 17 - 10XL, MTH, PTH1, ASM, LT, DDRA, SLE, EQ, TDSDDM2, TDSUSB2, YDSCPM2, RTE, IBA, PCI, TDSDVI, TDSET3, SAS, TDSHT3, TBD, JA3, TDSPTD, TDSVNM, DPOPWR, TDSHT3v1.3, 73, 74, DJE, DJA, 77, 78, 79, SVE, SVP, SVM, SLA

Note how JA3, advanced jitter analysis, indeed has bit 14 set to 1.

Some people just use a mask of FFFFFFF....FFFF.

Some of these analysis tools can once again be found in the same GitHub repo, or on the Tektronix website.

Jitter Analysis - Advanced tool

This is all theoretical, of course. I don’t think I’ll ever have a hobbyist need for any of this…

Cleaning Up

The final act is cleaning. This scope was in exceptional condition, except for the knobs on the control panel.

Dirty knob and less dirty one

The knobs have a thin anti-slip layer on them that is a finger grease magnet. Removing that layer with isopropyl alcohol makes the knobs look like new without a noticeable difference in control. Just be careful about using 99% isopropyl, I think it attacks the plastic. 90% was fine.

Peeling dirty knob

If some knobs are missing or cracked, the ones of a TDS220 are identical. You can buy knobs new or on eBay, but they’re expensive. If you really need a few, you might be better off buying a donor TDS220 instead.

TDS220 on top of TDS7104

The End

And with that, the scope is ready to be deployed to a shelf in my garage. One day I’ll need something with this kind of firepower but for everything else, a small scope with lower specs is way more practical. I like the scope better than the Agilent 54831 so that will probably hit Craigslist at some point.

All words in this blog posts were written by a human.

References

Footnotes

  1. I obviously only figured this out after removing the enclosure… 
  2. I didn’t try the 16-bit not-so-true color mode. 
  3. The Agilent 54831 uses a similar overlaying method. 
  4. In reality, I will never use this scope enough to ever run into an issue like this. 
  5. You can still find native 2.5” IDE 44 laptop SSDs, like this one, but you pay $30 more for the same capacity. 
  6. RawWrite for Windows is the one to use on a 32-bit Windows system, like the Win2k OS on the IBM Travelstar of the scope. 
  7. If you install the incorrect driver, the scope will still boot with a working LCD screen, but once the Windows GUI starts, it will move its business to the Intel integrated GPU. You need a VGA monitor to follow what’s happening. Even if you later select the right driver, Windows somehow thinks that the old driver is good enough and just doesn’t do it, without any feedback. I had to manually delete the bad driver files from the SSD to finally make it work. 
The Daily Front Page 18 of 28
Tuesday, September 1, 2026 The Daily Front No. #260901 — Dial Tone, Reversible
repository

RotaryCell: Making an unmodified rotary phone work over LTE with an ESP32-S3

by jombib·▲ 146 points·50 comments·github.com ↗
★ 122⑂ 3 forks C++

A portable, battery-powered cellular conversion for classic rotary telephones—fully reversible, with no modification to the original phone.

RotaryCell converts a traditional rotary telephone into a self-contained, battery-powered, portable cellular telephone without modifying the original telephone.

The design continues to use the original handset, rotary dial, switch-hook, mechanical ringer, network block, and existing jacks. The added electronics mount reversibly inside the case; no original telephone parts need to be drilled, cut, or permanently altered.

The working prototype can be carried and operated away from a fixed telephone connection, making the original desk telephone usable at meetings, demonstrations, or anywhere compatible cellular service is available.

Working RotaryCell prototype: an original black Western Electric Model 500 rotary telephone

This is a working engineering archive rather than a finished construction release. The hand-wired prototype operates, while the first integrated PCBs are currently awaiting assembly and validation.

Reproduce the working prototype by hand

The complete point-to-point wiring and component reference is the primary starting point for recreating the proven hand-wired prototype:

RotaryCell complete prototype wiring and component schematic

Open or download the full-resolution printable PDF. It covers the LilyGO, protected 21700, passive audio components, AG1171, GPIO connections, and original Model 500 circuitry without using the new PCBs. A hand-wired installation fits inside the telephone, but arranging and insulating all of the loose components and wiring is challenging; expect repeated dry-fitting and careful routing.

Current baseline

This repository records the project as it stood on August 28, 2026:

  • Firmware v0.10.4 is the current prototype-tested software.
  • The Audio and Reset A4 PCB was ordered from JLCPCB on August 27, 2026.
  • The AG1171 Carrier Through-Hole PCB was ordered from JLCPCB on August 28, 2026.
  • Both ordered PCB packages are archived exactly as submitted and have not yet been validated as assembled boards.

See STATUS.md for the distinction between tested prototype behavior and hardware awaiting validation.

System overview

  • A LilyGO T-A7670G-S3 Standard board supplies the ESP32-S3 controller, A7670 cellular modem, battery charging, and cellular audio interface.
  • A Silvertel AG1171 subscriber-line interface operates the telephone line circuitry, senses the switch-hook, and drives the mechanical ringer.
  • The Audio and Reset A4 PCB provides adjustable transmit/receive audio conditioning and a hardware power-cycle circuit for recovery when software-only modem reset is insufficient.
  • A single protected 21700 cell connects to the LilyGO battery pads through a harness in place of the original 18650 holder and directly supplies the AG1171 carrier VPWR input.

The telephone's RJ11 line jack is used only to deliver regulated 5 V to the LilyGO charging input on the designated pins. It does not power the AG1171 directly and is not used as a telephone-line interface.

Repository layout

Path Contents
firmware/current Current Arduino sketch and source files
firmware/prebuilt Current application OTA binary and source ZIP
firmware/archive Historical firmware snapshots
hardware/audio-reset-a4 Exact Audio and Reset A4 source and manufacturing package
hardware/ag1171-carrier-through-hole Exact through-hole carrier source and Gerber package
hardware/prototype Material associated with the working hand-wired prototype
hardware/experimental Unfinalized schematics, layouts, libraries, and alternatives
hardware/legacy Older hardware documentation retained for reference
docs Architecture, bring-up, and historical documentation
site Draft project-page copy for evilroot.net

For a hand-wired build, start with the complete prototype wiring reference. For the newer PCB implementation, continue with STATUS.md, current hardware wiring, the master BOM, the assembly guide, and BUILDING.md.

Current functions

  • Rotary pulse dialing and switch-hook detection
  • Incoming and outgoing cellular calls
  • Physical bell ringing through the AG1171
  • North American dial, reorder, and receiver-off-hook warning tones
  • Bidirectional handset audio with adjustable levels
  • Battery monitoring
  • USB diagnostics and a temporary maintenance Wi-Fi dashboard
  • Browser/USB AT-command terminal and persistent event log
  • Cellular-network clock synchronization and application OTA updates

Dial service code 0000 starts maintenance Wi-Fi. Service code 9999 performs modem diagnostics and software recovery, but it did not recover the field-observed modem lockup described in STATUS.md.

Important cautions

  • Never connect prototype Tip/Ring wiring or the repurposed charging jack to the public telephone network or energized premises telephone wiring.
  • Clearly label the charging jack and verify its regulated voltage, polarity, pin assignment, and protection before use.
  • Lithium-ion cells require suitable protection, charging, fusing, insulation, and mechanical restraint.
  • The maintenance access point uses the development password rotarycell. Change WIFI_AP_PASSWORD in Config.h before use around untrusted people.
  • The August 2026 PCB files are as ordered, not yet production-tested. Create a new revision rather than silently replacing an as-ordered package.

Repository policy

This is a public engineering and development archive. It is intended to preserve a durable, reproducible baseline, make the working hand-wired prototype available to other builders, and document progress toward a more integrated implementation.

The repository should not be mistaken for a finished construction kit or production release. Files under hardware/audio-reset-a4 and hardware/ag1171-carrier-through-hole record the exact board candidates ordered in August 2026; their assembled operation has not yet been validated. Tested behavior, known failures, and remaining documentation gaps are tracked in STATUS.md.

License

Code and original documentation in this repository are licensed under the MIT License. Third-party datasheets, vendor names, trademarks, and historical telephone designs remain the property of their respective owners.

The Daily Front Page 19 of 28
Tuesday, September 1, 2026 The Daily Front No. #260901 — Five Kilobytes per Second
article

2004 RuneScape fit a multiplayer RPG into 56k dial-up

by fagnerbrack·▲ 135 points·83 comments·jkm.dev ↗
The answer, however, is a sustained, almost obsessive exercise in not wasting bytes.

In 2004 I played too much RuneScape on a 56k modem that died the moment Mum picked up the phone. A 3D world, up to a couple of thousand players on a server, dozens on screen at once - in the browser, on 5 kilobytes per second. It worked. Let’s follow a single step and see how.

As a child I was too preoccupied with picking flax and killing goblins to think about how this worked. The answer, however, is a sustained, almost obsessive exercise in not wasting bytes. So, let’s click one tile north of where we’re standing, and trace every byte that crosses the wire from that click, to the server, to the screen of another player.

Central fountain, Varrock Square

Central fountain, Varrock Square

Methodology

The detail in this post comes from a decompiled 2004 RuneScape 2 client. Snippets are rough translations from that decompile, tidied up in places for readability but with the logic intact.

The core principles aren’t identical across versions, but most of them run all the way from RuneScape Classic (2001) to present-day RuneScape 3 and, of course, Old School RuneScape.

If you played RuneScape in the early 2000s and are still in possession of a hard drive from that era, please check out the RuneScape Archive Project. They do great work to catalogue historic RuneScape versions which are otherwise lost to time, and every contribution is valuable.

Constraints

Let’s look at some of the constraints that Jagex were working with at the time.

  • Bandwidth. A 56k modem syncs at 56 kilobits per second downstream, and less upstream, minus any protocol overheads and line noise. Call it 5 KB/s down and a lot less up. Broadband was available in British homes by 2000, but it wasn’t until the late 2000s that the majority of UK households had a broadband connection, so plenty of players were on dial-up.
  • Java applet, in a browser, in 2004. Java applets ran in a security sandbox, which meant no raw native sockets and no UDP. Every byte travelled over a single TCP connection, in-order and with per-segment overhead.
  • A 600ms server cycle. The RuneScape game server advances in discrete cycles (or ticks) of roughly 600 milliseconds. Every cycle, for every player, the server has to work out everything that player can now see and ship it before the next one.

The cipher layer, briefly

After the login handshake completes, before any game packets are sent, a small encryption layer is set up. This one’s not about saving bytes; it’s the only encryption in the stack (outside of some RSA encryption in the login handshake), and it’s here because the opcode it protects is the very thing every later section depends on.

Every packet begins with an “opcode” byte: a small integer saying what kind of packet this is. That opcode (and only that opcode) is enciphered with a stream cipher called ISAAC. There are two streams in play - one for traffic from client to server, and one for the reverse direction. Both sides need both streams: the client enciphers what it’s about to send and deciphers what just arrived, and the server does the same in mirror image (per connected player).

Both streams are seeded from a shared four-integer key. The client generates two of those integers itself; the other two come from the server as part of the handshake. The server-to-client stream then uses the same seed with 50 added to each word - enough to keep the two directions from sharing a keystream:

this.outboundCipher = new ISAAC(seed);

for (int index = 0; index < 4; index++) {
    seed[index] += 50;
}

this.inboundCipher = new ISAAC(seed);

Enciphering on the way out is one line:

public void putOpcode(int opcode) {
    this.putByte(opcode + this.outboundCipher.value());
}

And on the way in, the mirror image:

this.currentOpcode = (this.currentOpcode - this.inboundCipher.value()) & 0xFF;

So the packet body isn’t encrypted, only the opcode. As we’ll see later, the opcode is what tells you how to read the rest of the packet, and where one packet ends and the next begins. Without it, the body is just a wall of bytes, so enciphering that one byte was the cheapest possible defence against third-party packet parsers.

Sending a walk request

We’re going to look at what happens when you click on a tile one square north, and how that gets transmitted to the server.

Before any networking occurs, the client runs a breadth-first search using the local collision map to build a path from where you are to where you clicked (an easy search, in this case), and then writes the packet for the server to read. The pathfinding is standard so I won’t go into it here.

The first part of the packet is the opcode, followed by a single byte containing the length of the packet body. As you’ll see, the number of bytes contained in the packet is dependent on the size of the path, so this “length” byte allows the server to know how far to read. Not all packets have this length byte, only packets which contain some variably sized body.

The start position takes 4 bytes (two shorts), each subsequent waypoint delta takes 2 bytes, and there’s a final byte for whether the Ctrl key is held. So the body length is 4 + 2 * (pathLength - 1) + 1.

this.outboundStream.putOpcode(ClientToServerOpcodes.WALK_TILE);
this.outboundStream.putByte(4 + 2 * (pathLength - 1) + 1);

The packet contains the absolute position of the first waypoint in the path (x and z sent as a two-byte “short” each), followed by the delta of each waypoint in the path against the first one - one signed byte per axis, which fits comfortably within the byte’s range of -128 to 127, as a single click can only ever land so far away.

The decision to send only a delta here, as 2 bytes per step, rather than absolute coordinates as 4 bytes per step is the first example we’ve seen of Jagex’s networking frugality. In absolute terms it only saves a few bytes for a single walk packet, but every additional waypoint costs 2 bytes instead of 4 - a 50% saving per waypoint.

int firstX = pathX[0];
int firstZ = pathZ[0];

this.outboundStream.putShort(this.playerPositionX + firstX);
this.outboundStream.putShort(this.playerPositionZ + firstZ);

for (int i = 1; i < pathLength; i++) {
    this.outboundStream.putByte(this.pathX[i] - firstX);
    this.outboundStream.putByte(this.pathZ[i] - firstZ);
}

It’s worth noting here that besides the opcode and the length byte, which make up the packet header, the order in which the individual parts of the packet body are written varied between versions as an anti-cheating measure.

For instance, in this revision this packet was written as (x, z, ...path), in other revisions - such as #317 - it was written as (x, ...path, z), and in others as (...path, x, z).

Additionally, byte-mangling transformations would be applied: writing with different endianness, adding constant values to individual bytes, and negating values - again as an anti-cheating measure, again varying between revisions.

I’ve reordered the snippet above, and removed all byte-mangling, for clarity. In reality these measures would be applied on both the client and the server, and would be applied to most parts of the protocol.

Another frugal decision here is that pathX and pathZ do not contain every tile in the path, just the corners. Walking ten tiles in a straight line only sends one waypoint: the destination. The server already knows where you started, so it walks the line itself and validates against its own collision map.

The last part of this packet is a single byte to indicate whether the Ctrl key is held. In early versions of the game, this was used to force “run mode”, in later versions it inverts the current movement mode (runs to your clicked destination if “run” is off, or walks if it’s on):

this.outboundStream.putByte(this.keyStatus[Keys.CTRL] == 1 ? 1 : 0);

So we can see that our single step north takes seven bytes, including our opcode and length marker:

A seven-byte client-to-server walk packet for a single step: one opcode byte, one length byte (value 5), a two-byte destination x short, a two-byte destination z short, and one run-toggle byte. The opcode and length form the header; the remaining five bytes form the body, whose size equals the length byte.

As our path only contained a single step, we don’t enter the loop to send the “delta” waypoints, so we can cross-check our 5-byte payload against the length marker:

  • 4 + 2 * (pathLength - 1) + 1 = 4 + 2 * 0 + 1 = 5

Once the snippets above have run, the packet is in the client’s outbound stream. That stream is drained to the network roughly every 20ms.

Server receives the request

The server’s main loop wakes roughly once every 600ms. On each wake, it drains every player’s inbound buffer, runs whatever handlers the packets call for, and composes the outbound player updates that we’ll look at next. A packet that arrives just before a cycle is processed almost instantly; one that arrives just after waits nearly a full 600ms.

That 600ms cycle time sets the granularity for latency. The 20ms client flush and any other networking overheads all swim well under this time. That’s why the rest of this post is about bytes, not time: there is no latency to save.

Once the inbound buffer has been drained by the server, reading the packet is roughly the process above, but in reverse:

int opcode = player.inboundStream.takeOpcode();

if (opcode == ClientToServerOpcodes.WALK_TILE) {
    int length = player.inboundStream.takeByte();

    int deltaCount = (length - 4 - 1) / 2;

    int[] firstWaypoint = new int[2];
    firstWaypoint[0] = player.inboundStream.takeShort();
    firstWaypoint[1] = player.inboundStream.takeShort();

    int[][] waypointDeltas = new int[deltaCount][2];
    for (int i = 0; i < deltaCount; i++) {
        waypointDeltas[i][0] = player.inboundStream.takeByte();
        waypointDeltas[i][1] = player.inboundStream.takeByte();
    }

    boolean holdingCtrl = player.inboundStream.takeByte() == 1;

    player.processWalkTile(firstWaypoint, waypointDeltas, holdingCtrl);
}

As you can see, once we’ve identified the opcode, we can read the length byte and reverse the write logic to extract the number of deltas.

I mentioned earlier that not all packets contain this length byte. In fact, most don’t; the majority of packets have a fixed-length body. Reading those is even simpler. Take, for instance, the “item on item” packet - sent when a player “uses” one item in their inventory with another:

if (opcode == ClientToServerOpcodes.USE_ITEM_ON_ITEM) {
    int sourceItemId = player.inboundStream.takeShort();
    int sourceInterfaceId = player.inboundStream.takeShort();
    int sourceInterfaceSlot = player.inboundStream.takeShort();

    int targetItemId = player.inboundStream.takeShort();
    int targetInterfaceId = player.inboundStream.takeShort();
    int targetInterfaceSlot = player.inboundStream.takeShort();

    player.processUseItemOnItem(/* ... */);
}

This packet has a fixed length of 12 bytes (6 shorts). The server is aware of this constant length, so there is no need to transmit a length marker as part of this packet.

The server cycle

There are a number of steps that make up a RuneScape server cycle, and the parts we are interested in happen in the following order:

  • read incoming packets
  • process players (queued actions, triggers, movement, etc)
  • build player updates (more on this in the next section)
  • flush outbound packets

There are other steps omitted here, such as processing NPCs, logins and logouts, and technicalities around when exactly actions occur (some happen when the packet is read, others are queued and happen later in the cycle). NPCs are updated via a structurally identical, slightly smaller update block in a separate packet.

The overall principle is clear: read, then do, then write.

Player updates

Before tracing the packet, it’s worth being explicit about the protocol’s foundation: the client holds its own mirror of every player it can see. A tracked list of nearby players, each with their last-known position, appearance, animation and chat state - plus the local player’s own state. The player update packet’s job is to keep that mirror in sync with the server’s authoritative version - which means, almost always, that an update is a delta against what the client already knows. “No change” is so cheap precisely because the client already has the data; the server just confirms it’s still valid.

Every cycle, the server sends each player a single composite “player update packet”. This single packet describes everything the client needs to know about every player it can see - including itself. The receiving client tears this information apart in four steps, and the order of those steps is as follows:

private void readPlayerUpdates(Packet packet) {
    packet.accessMode(PacketAccess.BITS);

    this.readLocalPlayer(packet);

    // other players already tracked by the client
    this.readOtherPlayers(packet);

    // players newly in range, which the client should start tracking
    this.readNewPlayers(packet);

    packet.accessMode(PacketAccess.BYTES);

    // detailed changes about players
    this.readPlayerDetails(packet);
}

The first three steps are bit-packed - the stream is read a few bits at a time, not byte by byte. Only the fourth step in this sequence is byte-aligned. This split is deliberate: movement and registration are high-frequency, and tiny, so they get bits; the less frequent rich updates (a player changed equipment, swung a sword, or said something) get bytes.

Step 1: Local player

The logic to read a local player is simple, so I will let you read it and we can analyse it after:

private void readLocalPlayer(Packet packet) {
    int updated = packet.takeBits(1);

    // no local movement and no local detail changes
    if (updated == 0) {
        return;
    }
    
    int movementType = packet.takeBits(2);

    // type 1: a walk
    if (movementType == 1) {
        int direction = packet.takeBits(3);

        this.localPlayer.step(direction, false);

        int detailUpdated = packet.takeBits(1);
        if (detailUpdated == 1) {
            this.trackPlayerDetails(this.localPlayer.id);
        }
    }
    // type 0: no move, but a detail update follows
    // type 2: a run - two directions back-to-back
    // type 3: a teleport
}

Read that first if statement again. If the local player didn’t move, and nothing about them changed this cycle, their entire presence in the update packet is a single bit. Not a byte. A bit. The most common state of any given player on any given cycle - “no change” - was made the cheapest possible transmission.

If the local player did move, it’s a 1 bit, two bits to represent the type, three bits for the direction and a single bit for the “is there more detail coming?” flag. Seven bits, less than a single byte, for “I took a step.” Excluding the first “update required” flag and the movement type, it fits in four bits.

The other types are cheap, too. Excluding the three bit headers:

  • type 0 (no move, but details to come): no payload. Zero bits.
  • type 2 (a run): two 3-bit directions, and a “more detail” flag bit. Seven bits.
  • type 3 (a teleport): the height plane (2 bits), the x and z coordinates (7 bits each), the “more detail” flag bit, and a “jump” bit (used to tell the client whether it should attempt to animate this movement). Slightly more expensive, but still only eighteen bits - slightly over two whole bytes.

Step 2: Tracked players

This is the same idea as above, applied to the crowd of already-tracked players.

One thing to note is that reading individual bits here continues immediately from the “local player” section above. That is to say, if the local player section is only 1 bit, the section below will begin reading from the 2nd bit - there’s no empty space to pad full bytes.

private void readOtherPlayers(Packet packet) {
    int count = packet.takeBits(8);

    for (int i = 0; i < count; i++) {
        int updated = packet.takeBits(1);

        if (updated == 0) {
            continue;
        }

        // read movementType etc as above
    }
}

An 8-bit count, then one bit per known player to say whether anything happened to them. Stand in a crowd of forty players where nobody’s moving, and that’s forty-eight bits (six bytes) to confirm that the entire scene is static. Any player who did take a step costs the same seven bits as the local player did in step 1.

This is the core trick. The default - “nothing changed” - is a single bit, the cheapest possible representation. Real bits are only spent on the things that actually moved. The server and the client share, baked in at compile time, an identical understanding of the protocol - including what the default is, and what counts as changed. Neither end ever has to detail “no change”; the absence of detail, gated behind the zero bit, is the message.

Step 3: New players in range

When someone walks into (or otherwise arrives in: logging in, teleporting, etc) your view for the first time, the server has to introduce them - who they are and where, relative to you:

private void readNewPlayers(Packet packet) {
    // room for an 11-bit player id
    while (packet.bitsRemaining > 10) {
        int playerId = packet.takeBits(11);

        // sentinel: no more players
        if (playerId == 2047) {
            break;
        }

        Player otherPlayer;
        // ... allocate or look up the player ...

        int updated = packet.takeBits(1);
        if (updated == 1) {
            this.trackPlayerDetails(playerId);
        }

        int teleported = packet.takeBits(1);

        int deltaX = packet.takeBits(5);
        if (deltaX >= 16) { deltaX -= 32; } // signed 5-bit value: -16 to +15

        int deltaZ = packet.takeBits(5);
        if (deltaZ >= 16) { deltaZ -= 32; }

        otherPlayer.move(localPlayer.x + deltaX, localPlayer.z + deltaZ, teleported == 1);
    }
}

An 11-bit player id (2047 is reserved as the “stop” sentinel, so the list doesn’t need a length header), one bit for whether a “more details” update is coming later, one bit for whether they teleported in, and then 10 bits for the position. The position is one of the details I love the most about this section.

Relative coordinates

A player’s absolute world coordinates are a pair of values in the thousands - RuneScape’s map is very large (thousands of tiles on each axis). Two 16-bit numbers, 32 bits total, to place someone anywhere on that map.

But the player update logic above doesn’t need a global position. It only needs to know where they are relative to the local player, because that’s all that can be seen. Another player who’s in range to be drawn is at most about fifteen tiles away. Fifteen fits nicely in a signed 5-bit number (-16 to +15). So a newly-visible player’s location costs ten bits - five per axis - instead of thirty-two. The coordinate space is recentered on the local player, and clipped to what’s visible. The encoding is sized to exactly that clipped range and not a single bit more. The same logic appears in step 1’s teleport branch, where coordinates are expressed as two 7-bit values (enough to address the ~104-tile loaded area) rather than full world coordinates.

This is the pattern repeated everywhere: figure out the smallest set of values that could possibly be needed, then use exactly enough bits to represent that set.

The bit cursor

At the start of this section, I mentioned that steps 1 through 3 read individual bits, while step 4 reads whole bytes. All of the “a few bits at a time” reading is one small method doing the bookkeeping. The convention is that bits fill each byte from the top down - the first bit sits at position 7, the last at position 0:

public int takeBits(int count) {
    int value = 0;
    for (int n = 0; n < count; n++) {
        int bytePos = this.bitPosition / 8;
        int bitInByte = 7 - (this.bitPosition % 8);

        int bitValue = (this.buffer[bytePos] >> bitInByte) & 1;
        value = (value << 1) | bitValue;

        this.bitPosition++;
    }
    return value;
}

The above implementation is somewhat different to that shipped in the RuneScape client - I’ve simplified it to highlight the concept more clearly. The version shipped with the client is slightly more complex as it’s optimised for performance.

As you can see, the method above walks the buffer one bit at a time. Without this, every “three bits per direction” and “one bit per idle player” would need to be read as a byte, taking most of the protocol’s frugality with it - eight idle players would need eight bytes rather than one.

When the bit-packed steps finish, the cursor is rounded up to the next whole byte and step 4 takes over with conventional byte reads.

Step 4: Player detail changes

This fourth step is responsible for any detailed player updates, generally related to the appearance of the player. It only touches players flagged as “more detail to come” in one of the earlier steps.

The full list of update flags is:

  • facing entity
  • facing tile
  • forced public chat
  • animation
  • appearance changed: equipment, etc (more on this below)
  • took a hit
  • normal public chat
  • graphical effect
  • forced movement along a path

Looking at the layout, the bottom two flags in the list are always (as far as I can tell) represented by bits in the high byte of the update type. These also tend to be the rarer updates, and I believe the assignment is a deliberate economic choice: only rare events require the second byte of the update type to be transmitted.

Later revisions add a “took a second hit this cycle” update - this is also always represented by a bit in the high byte, as further evidence that only rarer events require this extra byte for the update type.

Every player in the array of “more detail” updates is iterated over, and an “update type” flag is read:

private void readPlayerDetails(Packet packet) {
    for (int i = 0; i < moreDetailPlayerCount; i++) {
        int updateType = packet.takeByte();

        if ((updateType & 0b1000_0000) != 0) {
            updateType |= packet.takeByte() << 8;
        }
        
        // ...
    }
}

We can see another byte efficiency trick in use here. The nine flags we just listed are too many to fit in a single byte when each flag is an individual bit, so the full update type needs two bytes to address. Rather than reading two bytes per player (using takeShort), seven flags are packed into the first byte with a single marker bit, the most significant bit. When this marker bit is set, a second byte is read, shifted left by one byte and combined with the first to give a 16-bit value (of which 10 bits are meaningful: the 9 flags plus the marker).

The bits used for each flag, including the marker, are varied between versions as an anti-cheating measure.

After obtaining the full update type, it is checked for the presence of individual flags to apply certain details. Some of these are illustrated below:

if ((updateType & 0b0000_0100) != 0) {
    // player is facing an entity (npc or another player)
    player.targetEntityId = packet.takeShort();
}

if ((updateType & 0b0010_0000) != 0) {
    // player is facing a tile
    player.targetTileX = packet.takeShort();
    player.targetTileZ = packet.takeShort();
}

if ((updateType & 0b0000_0010) != 0) {
    // player is performing an animation
    player.animationId = packet.takeShort();
    player.animationDelay = packet.takeByte();
}

// ... other flags ...

// check the least significant bit of the high byte
if ((updateType & (0b0000_0001 << 8)) != 0) {
    // a graphical effect is playing on the player
    player.graphicalEffectId = packet.takeShort();
    player.graphicalEffectHeight = packet.takeShort();
    player.graphicalEffectDelay = packet.takeShort();
}

In the few examples above, you can see a number of the tricks we’ve seen so far. Multiple flag values are packed into the 8-bit or 16-bit update type. Different update mechanisms have different body sizes, as part of the agreed protocol between client and server. The smallest data type appropriate for the values being represented is used. All of these decisions were made with the aim of minimising the amount of data required to transmit this information.

Appearance update

I won’t go into full detail around the “appearance” part of this packet, but it’s the only expensive one in the list. It contains:

  • name
  • combat level
  • body part information, including equipped items and NPC transmogs
  • body part colour
  • stand / walk animations
  • gender
  • head icons (prayer icons, PK skull)

In total the appearance section costs between 44 and 80 bytes per player.

Why no bit packing?

It might seem inconsistent that the protocol abandons bit-level frugality just as it reaches the largest part of the packet, but step 4 is actually following the same rule as the rest - just landing on the other side of it. Bit packing trades CPU for bytes: you pay the cost of a bit cursor to reclaim the slack between a value’s real width and the byte it would otherwise sit in. It’s worth that trade only where the slack actually exists and repeats.

In steps 1, 2 and 3 it does, many times over. The default state - “no change” - is a single bit, and it repeats across every visible player every cycle, so the saving compounds across dozens of entities. Step 4 has neither half of that. There is no tiny default: a player either has no update at all (already gated by a single bit upstream) or a real one, whose smallest field, “facing an entity”, is already a two-byte short. A short has no slack to reclaim - it fills both its bytes - so bit packing would save nothing while still charging the cursor cost. The multiplier is gone too: step 4 only ever contains the handful of players who changed this cycle, not the whole crowd, so even if there were bits to save there’s almost nothing to multiply them by. The one place the trick still pays off is the update-type byte itself, with the marker bit buying a second byte only when needed - bit-packed within a byte, exactly where slack still exists.

The second reason is how the server composes this part of the packet, and it’s really the same point seen from the server’s side. Many fields in step 4 aren’t recomputed each cycle - I believe the appearance buffer, for example, is built once per player per change and held as a byte buffer the server splices into outgoing packets for any observer who needs it. The client certainly caches it that way, reusing it when a tracked player leaves visible range and re-enters; it would be strange for the server not to mirror that. What makes the splice cheap is that a byte-aligned blob is position-independent: wherever it lands in a given observer’s packet, it’s the same sequence of bytes, so inserting it is a plain array copy. Bit-align it and its offset would depend on everything written before it - which differs for every observer and every cycle - so the same cached blob would need a fresh shift-and-mask for every observer, every cycle, and the cache stops being worth keeping.

So the two halves of the packet are tuned for two different scarce resources. The bit-packed front is cheap to compute, impossible to cache, and exists to spare the client’s downstream dial-up. Nothing in it can be shared between observers: each sees a different crowd, positioned relative to itself. The byte-aligned back is expensive to compute but rarely changes, so it’s built once and spliced wherever it’s needed - and here the binding constraint isn’t the wire at all, but the server’s budget to assemble up to two thousand of these before the next cycle. The protocol switches representation at exactly the point where that constraint flips.

The bytes on the wire

Let’s add it up for the actual scenario: you take one step north, and we count what a nearby player’s client receives in that cycle’s player update packet. Say there are twenty other players in their view and, this cycle, only you moved.

The downstream player-update payload for one tick, laid out as six rows of eight bits (one row per byte). Bit 0 is pass 1, the local player, who did not move. Bits 1 to 8 are the pass 2 player count, spilling across the first byte boundary. Bits 9 to 15 are your seven-bit step. Bits 16 to 34 are nineteen idle players at one bit each, running across three rows. Bits 35 to 45 are the eleven-bit pass 3 new-player sentinel. Bits 46 and 47 are byte-alignment padding. Forty-eight bits total, six bytes, plus a one-byte opcode and two-byte length make nine bytes on the wire.

Add the opcode byte and a length marker (two bytes, rather than the single-byte marker used for our walk packet - the length of the player update block can be greater than 255), and you’re at roughly nine bytes for the complete answer to “what did everyone around me just do?” on a cycle where one person took one step in a crowd of twenty-one. Your upstream walk packet was seven bytes; the update echoed back to you is about nine. Sixteen bytes, round trip, for a step - and the server sends that same nine-byte answer to every other player who can see you. At 5 KB/s you have headroom for hundreds of those per second, which is exactly the point - combat, crowds and chat all have to fit in the same pipeline.

For scale: the text of this post is a touch over 32 KB - more than three thousand of those nine-byte scene updates. The write-up describing the frugality is far larger than anything the protocol actually sends.

The complete round trip, end to end:

A sequence diagram with three participants: your client, the server, and another player's client. Your client pathfinds and writes a walk packet, then sends a seven-byte WALK_TILE packet up to the server, flushed roughly every 20 milliseconds. The server runs a roughly 600 millisecond cycle: read, process, build updates, flush. At the end of the cycle it sends an approximately nine-byte player-update packet down to the other player's client, which renders your step, and a copy of about nine bytes back to your own client, making the round trip. The net cost is seven bytes up, about nine bytes down per observer, and one 600 millisecond cycle of latency.

The general lesson

The RuneScape client and the server it communicated with are not two systems exchanging messages. They work together as one system, which happens to be split across a TCP connection. Every economy in this protocol depends on both ends sharing knowledge that is never transmitted:

  • Both ends run the same pathfinder over the same collision map, so the client can send corners and the server can simply validate the path.
  • Both ends agree, at compile time, that the default state of a player is “didn’t change”, so “didn’t change” can cost only one single bit.
  • Both ends agree that visible means “within ~15 tiles”, so a position can be five bits per axis instead of sixteen.
  • Both ends agree on a fixed table of what things can change, so a bitmask can stand in for a schema.

None of this shared understanding is sent over the wire. It’s in the design. The protocol is small because the two programs were written together, by people treating the network as an implementation detail of a single application rather than a boundary separating two.

It’s tempting to read this as a relic - the way things had to be built before bandwidth became cheap. But the dividing line was never old versus new; it’s what the system is for, and which constraint is actually binding. A modern web service is built the opposite way on purpose: loosely coupled, self-describing, versioned, verbose - the same scene update as JSON over HTTP would run to hundreds of bytes, its headers alone dwarfing the nine. That heft isn’t waste; it’s what buys the ability to change one side without redeploying the other, to serve many different clients, and to debug by reading the wire. Those are the right defaults when the thing pressing on you is teams and change velocity, not bytes.

What’s easy to miss is how much software written today still lives on RuneScape’s side of that line. A competitive shooter, a rollback fighting game, a market-data feed - anywhere both ends ship together and every byte is contested - reach for the same tightly co-designed, bit-packed, schema-baked-in approach. The decoupled style isn’t a feature of modern design - it’s a response to independent deployability. You move toward it or away from it depending on which constraint binds.

Push the other way - make every byte genuinely matter - and you get this instead: a data model and wire format co-designed so tightly that they exist as one artifact. One where the cleverness lives in everything you’ve arranged not to send. Studying this protocol is studying what engineering looks like under a hard, absolute limit.

Thanks

Thank you to Jagex for building something that not only has stood the test of time, but that is good enough to be worth taking apart and learning from twenty years later.

Thank you to the many, many members of the preservation and reverse-engineering communities I’ve worked with over the last fifteen years to build the understanding I have today.

The variable names and structure come from a decompiled and cleaned up RuneScape client; the design is entirely Jagex’s.

The Daily Front Page 20 of 28
Tuesday, September 1, 2026 The Daily Front No. #260901 — Memory, Segmented
article

Flat vs. segmented memory – it's recursive

by signa11·▲ 52 points·1 comments·humprog.org ↗
The decline of fine-grained memory protection.

Flat vs segmented memory -- it's recursive

My recent forays in x86 segmentation (1 2) made me notice a trend in the evolution of x86: the decline of fine-grained memory protection, both across the move to 64-bit and indeed before that in the fast system calling features. These both partially hobbled the sophisticated segmentation system, which had been a hallmark of the architecture since the 286. The explanation is possibly Unix: the dominance of Unix and its preference for flat address spaces, rather than segmented ones, arguably inherited from the PDP-11 or indeed PDP-7, meant that “nobody wanted” a 64-bit version of the segmentation features.

Meanwhile, I like to joke among WebAssembly enthusiasts that “segmented memory is coming back”. WebAssembly is rather like a return to the programming model of OS/2 or some other non-Unix OSes, where flatness did not reign so supreme.

Such non-flatness is still highly relevant to safety and security of course. I recently revisited some of the ever-enjoyable writings of Poul-Henning Kamp, who framed CHERI as a reaction against flat memory models that he describes as “unsafe at any speed”. Kamp observes that the first thing software does on any flat memory is impose some subdivision structure on it.

One way to look at the question is as about to what extent the hardware should know about this subdivision... CHERI says yes, whereas earlier hardware had said no—except, of course, in certain cases, for a bunch of segmentation stuff!

  • Of course segmentation as realised in x86 doesn't have the semantics needed for fine-grained confinement, i.e. confinement within appropriate corners of the non-flat address space. Unprivileged code can reload segment registers and thereby reach any defined segment, modulo a very coarse-grained four-ring privilege model. So, the non-flatness of traditional segments was more for fault isolation than for security: it was secure only up to coarse-grained distinctions like user vs system, and otherwise protected only against incompetence not malice.

I think that “yes or no” is the wrong way to look at it, though. It's recursive! When we've imposed some subdivision structure on a flat memory, we like to do so again. Think arenas or memory pools, but also think about fields within structures (within structures). The question is not about flatness or not—the programmer's mental abstraction is never flat—but somehow how we square a fundamentally recursive phenomenon practised by programmers, namely subdivision, with hardware—which is very much non-recursive. Hardware is conceptually finite-state, and its engineering practices tend to prefer fixed structures and bounded depths; maybe an ultra-CISC CPU will provide some iteration in microcode, but that's about the limit.

If hardware is non-recursive but software is recursive, how can we square those differences? A naive approach is just to bound the depth: say hardware knows up to N levels of decomposition, probably with N=1, and the rest is on software. But that is not satisfactory; it is “the hardware washing its hands” of what the programmer is doing. It guarantees non-unformity and the loss, at higher N, of any hardware-added value. CHERI does not do this; it keeps the flexibility to deal with recursive decomposition because software still handles the recursive steps: bounds can be arbitrarily narrow (-ish), but are narrowed by software and passed around explicitly. What you can address at any point, therefore, is determined not so much by the state of the hardware but by an emergence of the software: what has flowed within reach of the currently executing code, i.e. the memory you can access the transitive closure of reachable capabilities. (This emergence naturally opens up an obvious auditing difficulty, although one which the right tools could address.)

CHERI also buys this flexibility by, ironically, a restriction: addressing is constrained to be over an unbroken chain of capability derivation operations with monotonically decreasing bounds. I've always felt some discomfort about this bargain, because, software being software, some programs will choose to go their own way, e.g. by performing funky non-monotonic address calculations. Our several-decades' legacy of programs expressing their traversal of a somehow-subdivided address space in software, just the way they like it, set up friction with any new, more opinionated hardware. Overcoming this is a mere matter of development effort, but making that effort will only become normalised, a.k.a. its cost “successfully” externalised across the industry, if CHERI (or something similar) “wins”. (That cost would come with great benefits in return of course. But “winning”, in the sense of achieving hardware ubiquity, is a high-stakes game.)

With liballocs, I have been “happily” much less concerned with security and therefore largely free not to get into the business of prescribing rules for how addresses may be derived. I've instead been much more concerned with capturing descriptively whatever structure real software may have come up with, as it recursively subdivides the flat address space it starts out with. It has a recursive abstraction at its heart: allocations nest within other allocations, forming a tree. There's also no “level-N cut-off” or hardware/software divide: it's fundamentally software, and it wants to capture the structure all the way down using reflective abstractions that are as uniform as possible despite their many and heterogeneous implementations within the system. This “homogeneous interface, heterogeneous implementation” idea is of course often associated with object-orientation, and rarely with hardware.

  • While liballocs itself is not opinionated about how programs use the recursively subdivided structure that liballocs keeps track of, it could certainly be used to build added-security mechanisms that impose some opinions—although secure against malice if, and only if, the underlying hardware provides useful primitives for securing those mechanisms themselve, within the same address space. Annoyingly, x86-style segmentation would have been a near-sufficient basis, if the OS actually exposed it to userland. For roughly what I'd like, I'm constantly reminded of the amusingly-titled “Lord of the x86 Rings” paper.

Incidentally, to finish on another object-oriented note, the classical language-VM approach to subdivision punts in completely the opposite way to hardware: everything is near-maximally subdivided, into tiny objects and an enormous explicit interreferencing (pointer) relation between them. The programmer no doubt has coarser-grained structures in their head, but they stay there: the system doesn't offer to structure storage around them. As a result, these systems also punt on spatial locality—the hardware's heuristic of grouping together bytes or words into larger units, hence the longstanding performance disadvantages of such approaches.

The Daily Front Page 21 of 28
Tuesday, September 1, 2026 The Daily Front No. #260901 — The Map Room
article

Movie Scene Map – 13,312 films, series, games, anime and manga

by Flightmussy·▲ 182 points·29 comments·moviescenemap.com ↗
A free interactive map of 15,565 real filming locations in 166 countries.

Movie Scene Map is a free interactive map of 15,565 real filming locations in 166 countries: the studios, castles, streets and landscapes where films and television series were shot. Pan and zoom the map above, click any point for a photograph and the productions shot there, or open a place’s page for everything filmed in it. Beside the 9,287 films and series with a page of their own sit 2,153 video games, 407 anime and 365 manga, placed by where their stories are set, because nothing is filmed in a drawn world.

Browse by kind

Films & series 9,287Video games 2,153Anime 407Manga 365Franchises 653

Browse by kind of place

🎬 Studios & sets 134🏰 Castles & palaces 597🏛 Landmarks & buildings 4,114🛣 Streets & squares 830🏞 Landscapes & nature 2,155🏙 Towns & cities 7,536🌍 Countries & regions 168🎭 Fictional places 31

Filming locations by country

The countries with the most places on the map:

All 94 countries with a page → · Filming locations near you: day trip guides for 400 cities →

The most filmed places

Cities and regions ranked by everything shot inside them, the one number a single filming location statement cannot give you:

The most filmed cities and regions on earth, ranked → · The 150 most famous scenes ever filmed on location →

Famous productions to start with

The most widely covered productions on the map, by number of Wikipedia language editions:

Every film and series on the map →

How this map is built

Movie Scene Map is built from open data. The backbone is the filming location statements on Wikidata, joined to each place’s coordinates, its photograph on Wikimedia Commons and its Wikipedia article. Beside them sit places named in a production’s own Wikipedia article, or by the setting categories its editors file it under, and those are labelled per Wikipedia wherever they appear, because a sentence is weaker evidence than a statement and the two are never mixed. Nothing is scraped from listicles and nothing is generated.

The atlas is curated, not complete. A production earns a page once enough Wikipedia editions cover it, and a country that looks empty means sparse Wikidata coverage, never that nothing was filmed there. Missing a film you know? Add a filming location statement to its Wikidata item, with a source, and it appears here at the next rebuild and in every other project reading that data. The five steps are on the missing page, and the work list names the productions where one edit would do it.

The whole atlas is free to download as GeoJSON or CSV, CC0, and the same page describes the read only MCP endpoint that answers these questions live for AI assistants. Every change, including the mistakes, is in the changelog, and the about page says what the data is and what it leaves out.

Frequently asked questions

What is Movie Scene Map?

Movie Scene Map is a free interactive map of 15,565 real filming locations in 166 countries: the studios, castles, streets and landscapes where films and television series were shot, plus the real places video games, anime and manga are set in. Search a title to see where it was made, or a place to see what was made there.

How many filming locations are on the map?

15,565 places across 166 countries, attached to 9,287 films and series, 2,153 video games, 407 anime and 365 manga with a page of their own. The count moves with every rebuild from Wikidata, and the changelog records each one.

Where does the data come from?

Filming location statements on Wikidata, joined to each place’s coordinates, photograph and Wikipedia article. Beside them, and labelled per Wikipedia wherever they appear, sit places named in a production’s own Wikipedia article or by the setting categories its editors file it under. Nothing is scraped from listicles and nothing is generated.

What does per Wikipedia mean on a page?

That the claim rests on a sentence or a category in Wikipedia rather than on a Wikidata statement. A sentence is weaker evidence than a statement, so the two are never mixed, and every such row names the exact section or category it came from.

Are video games, anime and manga filmed somewhere?

No. A game is rendered and an anime or a manga is drawn, so none of them was filmed anywhere. The atlas places them by where their story is set, from Wikidata’s narrative location and from Wikipedia’s setting categories, and every page about one says set in rather than filmed in.

Why is a film I know missing, or placed only in a country?

Coverage follows Wikidata, and it is uneven. Some famous productions carry no filming location statement at all, and some carry one so coarse that it names a whole country, which this atlas reports as text rather than placing a pin. The work list on the gaps page names the productions where one edit would fix it.

Can I add a filming location?

Yes, upstream. Add a filming location statement to the production’s Wikidata item, with a source, and the place appears here at the next rebuild and in every other project reading that data. The five steps are on the missing page.

Is Movie Scene Map free to use?

Yes. No account, no advertising, no paywall. The whole atlas is also downloadable as GeoJSON and CSV under CC0, and a read only MCP endpoint answers the same questions live for AI assistants.

The Daily Front Page 22 of 28
Tuesday, September 1, 2026 The Daily Front No. #260901 — Compatibility and Version Control
article

Run macOS Software on Linux

by Bluestein·▲ 280 points·85 comments·darlinghq.org ↗

Darling is a translation layer that lets you run macOS software on Linux

  • Fast

    Darling runs macOS software directly without using a hardware emulator.

  • Free

    Like Linux, Darling is free and open-source software.
    It is developed openly on GitHub and distributed under the GNU GPL license version 3.

  • Compatible

    Darling implements a complete Darwin environment. Mach, dyld, launchd — everything you'd expect.

  • Easy to use

    Darling does most of the setup for you. Sit back and enjoy using your favorite software.

  • Native

    We aim to fully integrate apps running under Darling into the Linux desktop experience by making them look, feel and behave just like native Linux apps.

~ $ uname
Linux
~ $ darling shell
Darling [~]$ uname
Darwin
  • That sounds a lot like Wine

    And it is! Wine lets you run Windows software on Linux, and Darling does the same for macOS software.

  • Does it support GUI apps?

    Almost! This took us a lot of time and effort, but we finally have basic experimental support for running simple graphical applications.

  • Does it violate Apple's EULA?

    No! We only directly use those parts of Darwin that are released as fully free software.

  • Does the name Darling mean anything?

    The name Darling is a combination of “Darwin” and “Linux”. Darwin is the core operating system macOS and iOS are based on.

  • Can I run Darling on Windows using WSL?

    With WSL 2, yes! See the documentation for more details.

  • Do you know about opensource.apple.com, GNUstep, The Cocotron and other projects?

    We do, and in fact, Darling is largely based on the original Darwin source code published by Apple. We use The Cocotron as a basis for our Cocoa implementation, along with the Apportable Foundation and various bits of GNUstep.

  • Do you have plans for supporting iOS apps?

    Yes, in the long run, we'd like to be able to run iOS apps on ARM devices (like most Android phones). A significant challenge here would be to write our own implementation of UIKit. Come talk to us if you're interested in working on this!

  • How do I contribute?

    Start by reading the documentation and our blog to get familiar with Darling internals. Then, come and join us on GitHub. It's great if you have experience in developing for macOS or iOS, but it's absolutely not required to start contributing.

The Daily Front Page 23 of 28
Tuesday, September 1, 2026 The Daily Front No. #260901 — Compatibility and Version Control
article

The creator of Jujutsu has joined ERSC

by steveklabnik·▲ 209 points·153 comments·ersc.io ↗

NEW YORK (September 1, 2026), FOR IMMEDIATE RELEASE

East River Source Control has appointed Martin von Zweigbergk, creator of the Jujutsu version control system, as chief technology officer, to lead engineering on the company’s next generation version control platforms.

von Zweigbergk began Jujutsu as a side project in late 2019 and turned it into his full-time work at Google. The project has more than 30,000 stars on GitHub and ships under the Apache 2.0 license.

Before Jujutsu, he worked on Fig, the Mercurial client that gave Google engineers a distributed workflow on top of Piper, the monorepo holding most of the company’s code. He also contributed to Git, which according to the 2022 Stack Overflow developer survey, 96% of developers use professionally.

The hard problems many engineering teams are just beginning to face are ones he’s already been working on for over a decade. Having Martin lead our engineering affords ERSC a wholly different level of technical capability as an organization.

— Benjamin Brittain
     co-founder and chief executive officer of East River Source Control.

The company is building tools to help organizations manage the exponentially increasing needs of their source code management and collaboration tools brought about by the way AI is re-shaping the software industry. ERSC Storage will be entering private beta later this month.

von Zweigbergk will continue to be a core maintainer of JJ as an open source project under the Apache 2.0 license.

Jujutsu improves the part of version control that sits on your laptop. But the remote server is still Git, which has a ceiling that comes fast for products at scale. We think the storage layer has to change to match the model, and that work can be better supported by a company than an open source project.

— Martin von Zweigbergk
     chief technology officer of East River Source Control.


About East River Source Control

East River Source Control is building the next generation of version control platforms for humans and machines. The company launched in 2025, backed by Amplify Partners. For more information, visit ersc.io.

The Daily Front Page 24 of 28
Tuesday, September 1, 2026 The Daily Front No. #260901 — The Curiosity Cabinet
article

Restroom Archive

by jcalx·▲ 367 points·83 comments·restroomarchive.com ↗

This compact stall has an open bottom and open top with a colorful mosaic of tiles against the far wall and a black tile floor. A door hook allows visitors to conveniently hang their bags.

The Daily Front Page 25 of 28
Tuesday, September 1, 2026 The Daily Front No. #260901 — Also on the Front Page
The Daily Front Page 26 of 28
Tuesday, September 1, 2026 The Daily Front No. #260901 — Appointments & Situations Wanted
ask hn

Ask HN: Who is hiring? (September 2026)

Please state the location and include REMOTE for remote work, REMOTE (US) or similar if the country is restricted, and ONSITE when remote work is not an option.

Please only post if you personally are part of the hiring company—no recruiting firms or job boards. One post per company. If it isn't a household name, explain what your company does.

Please only post if you are actively filling a position and are committed to replying to applicants.

Commenters: please don't reply to job posts to complain about something. It's off topic here.

Readers: please only email if you are personally interested in the job.

Searchers: try https://nthesis.ai/public/hn-who-is-hiring, https://dheerajck.github.io/hnwhoishiring/, http://nchelluri.github.io/hnjobs/, https://hnjobs.emilburzo.com.

Don't miss this other fine thread: Who wants to be hired? https://news.ycombinator.com/item?id=49522896

View the position →

ask hn

Ask HN: Who wants to be hired? (September 2026)

Share your information if you are looking for work. Please use this format:

Location: Remote: Willing to relocate: Technologies: Résumé/CV: Email:

Please only post if you are personally looking for work. Agencies, recruiters, job boards, and so on, are off topic here.

Readers: please only email these addresses to discuss work opportunities.

Searchers: try https://nthesis.ai/public/hn-wants-to-be-hired, https://www.wantstobehired.com.

View the position →

The Daily Front Page 27 of 28
Tuesday, September 1, 2026 The Daily Front No. #260901 — Colophon

That's the Front for Today

Issue No. #260901 — Tuesday, September 1, 2026 — went to press 2026-09-02 at 05:18 UTC.

About This Magazine

The Daily Front is a daily digital magazine assembled from the stories that reached the front page of Hacker News on Tuesday, September 1, 2026. Headlines, points, and comment counts are recorded as they stood at press time. All articles remain the property of their original authors — every piece links back to its source and its discussion thread.

How It Was Made

Fetched, cleaned, and typeset by an automated pipeline. An editor model laid out the pages and chose the highlights; a second read a handful of the day's stories and briefed the cover illustrator — 32 model calls and 299k tokens in total. Set in Jacquard 12, Playfair Display, Source Serif 4, and IBM Plex Mono, all served via Google Fonts under the SIL Open Font License.

The Cover

The cover illustration was commissioned with this prompt:

In a university archive, a researcher stands before a half-open steel gate, holding a transparent folder of mismatched experiment sheets and scattered deadline cards. Behind the gate, a compact laptop sits beside an external solid-state drive, its cooling fan blowing loose paper strips through the narrow opening like a stream. A row of sealed filing cabinets recedes into darkness, while one cabinet door hangs ajar, revealing duplicated data tables and a nest of disconnected cables.

Pastel-and-charcoal drawing on deep slate-blue colored paper, using powdery marks and rubbed shadows under a low amber sidelight; depict the researcher before the half-open steel gate, holding a transparent folder of mismatched experiment sheets and scattered deadline cards, with the laptop and external solid-state drive behind it blowing loose paper strips through the opening, receding sealed filing cabinets fading into darkness, and one ajar cabinet exposing duplicated data tables and a nest of disconnected cables.

Absolutely no text, letters, numbers, readable symbols, or logos anywhere in the image.

Production Ledger

StageModelCallsTokens InTokens Out
extractgpt-5.6-luna 28 176,250 92,951
layoutgpt-5.6-terra 1 19,446 2,328
covergpt-5.6-luna 2 1,716 261
covergpt-image-2 1 207 5,488

The Publisher

Published by Johnny.

Support the Press

If The Daily Front brightens your morning, consider supporting its publisher.

Credits & Contact

All content — articles, posts, comments, and the images within them — belongs to its original authors and is reproduced here to point readers back to the source. Full credit goes to those creators; every item links to its original and its Hacker News discussion.

If you are an author and would like your content removed from an issue, write to hi@johnnys.page and it will be taken down.

Feedback is always welcome at the same address: hi@johnnys.page.

Credit where credit is due.

Every page of this issue began as someone else's work — these are the original sources, linked in full.

  1. AnkiDroid: Google Play no longer allowing Open Collective donation link by hexa555 — github.com·HN discussion ↗
  2. Claude Fable 5.1 and Claude Mythos 5.1 by denysvitali — anthropic.com·HN discussion ↗
  3. I trained a small transformer in 1.5hrs and it beats many LLMs by porridgeraisin — mvakde.github.io·HN discussion ↗
  4. How accurate have Ed Zitron's AI skeptic predictions been? by jatins — danluu.com·HN discussion ↗
  5. Show HN: Running 104GB Qwen3.8-Flash-Next on 48GB Mac with at ~12 tok/s by carloslfu — github.com·HN discussion ↗
  6. Atlas: A World Model for Spatial Intelligence by johnsutor — worldlabs.ai·HN discussion ↗
  7. GPU World by simonpure — gpuworld.org·HN discussion ↗
  8. The ChatGPT/Codex app bundles a full copy of LibreOffice by timpera — simonwillison.net·HN discussion ↗
  9. Launch HN: Nori Robotics (YC S26) – A low-cost humanoid robot for development by AntonioLi — norirobotics.com·HN discussion ↗
  10. Fastpotify by nreece — fastpotify.rocks·HN discussion ↗
  11. Play Store blocks AuroraStore, hurting GrapheneOS users by erikvanoosten — gitlab.com·HN discussion ↗
  12. Introducing Ad Blocker for Firefox on iOS by HieronymusBosch — blog.mozilla.org·HN discussion ↗
  13. Evidence of Fraud in an Influential Study About Procrastination by Anon84 — datacolada.org·HN discussion ↗
  14. American Airlines mechanic Azriel “Al” Blackman has died by NaOH — simpleflying.com·HN discussion ↗
  15. Lion-man by gurjeet — en.wikipedia.org·HN discussion ↗
  16. Borges Labyrinth in Venice reopens to the public by gone35 — wallpaper.com·HN discussion ↗
  17. Magic eye tube by peter_d_sherman — en.wikipedia.org·HN discussion ↗
  18. Refurbishing a Tektronix TDS7104 Oscilloscope by jwise0 — tomverbeure.github.io·HN discussion ↗
  19. RotaryCell: Making an unmodified rotary phone work over LTE with an ESP32-S3 by jombib — github.com·HN discussion ↗
  20. 2004 RuneScape fit a multiplayer RPG into 56k dial-up by fagnerbrack — jkm.dev·HN discussion ↗
  21. Flat vs. segmented memory – it's recursive by signa11 — humprog.org·HN discussion ↗
  22. Movie Scene Map – 13,312 films, series, games, anime and manga by Flightmussy — moviescenemap.com·HN discussion ↗
  23. Run macOS Software on Linux by Bluestein — darlinghq.org·HN discussion ↗
  24. The creator of Jujutsu has joined ERSC by steveklabnik — ersc.io·HN discussion ↗
  25. Restroom Archive by jcalx — restroomarchive.com·HN discussion ↗
  26. Terence Tao explains 6 essential mathematical concepts [video] by matthewsinclair — youtube.com·HN discussion ↗
  27. Ambient CSS v3 – Blender meets CSS by kikkupico — ambientcss.vercel.app·HN discussion ↗
  28. Tmp.0ut Volume 5 by ghuntley — tmpout.sh·HN discussion ↗
  29. Ask HN: Who is hiring? (September 2026) by whoishiring — news.ycombinator.com·HN discussion ↗
  30. Ask HN: Who wants to be hired? (September 2026) by whoishiring — news.ycombinator.com·HN discussion ↗

Browse all issues in the archive →