<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[Run Data Run]]></title><description><![CDATA[Clear thinking about AI from someone building it in production. No hype, no hand-waving. Just what works, what doesn't, and why it matters. My book is out: builder-leader.com]]></description><link>https://rundatarun.io</link><image><url>https://substackcdn.com/image/fetch/$s_!t_Ch!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffa36f5aa-74af-4492-b8d7-93b03f14a337_1280x1280.png</url><title>Run Data Run</title><link>https://rundatarun.io</link></image><generator>Substack</generator><lastBuildDate>Fri, 18 Sep 2026 03:26:18 GMT</lastBuildDate><atom:link href="https://rundatarun.io/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[Justin Johnson]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[rundatarun@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[rundatarun@substack.com]]></itunes:email><itunes:name><![CDATA[Justin Johnson]]></itunes:name></itunes:owner><itunes:author><![CDATA[Justin Johnson]]></itunes:author><googleplay:owner><![CDATA[rundatarun@substack.com]]></googleplay:owner><googleplay:email><![CDATA[rundatarun@substack.com]]></googleplay:email><googleplay:author><![CDATA[Justin Johnson]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[Four Labs, Two Test Environments, Seventeen Days. Then Everyone Called for a Slowdown.]]></title><description><![CDATA[Between 21 July and 6 August, models from OpenAI, Anthropic, Meta and Moonshot all went outside the boundary during safety testing. Almost all of it ran through the same two testing environments, and the September pacing debate never mentioned either one.]]></description><link>https://rundatarun.io/p/four-labs-two-test-environments-seventeen</link><guid isPermaLink="false">https://rundatarun.io/p/four-labs-two-test-environments-seventeen</guid><dc:creator><![CDATA[Justin Johnson]]></dc:creator><pubDate>Wed, 16 Sep 2026 19:01:28 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/453aa74d-9f92-419f-abeb-4d94d5c1619d_1664x928.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!3-ka!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffa5256e0-7e79-4ebe-b74b-f2512ba54d96_1664x928.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!3-ka!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffa5256e0-7e79-4ebe-b74b-f2512ba54d96_1664x928.jpeg 424w, https://substackcdn.com/image/fetch/$s_!3-ka!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffa5256e0-7e79-4ebe-b74b-f2512ba54d96_1664x928.jpeg 848w, https://substackcdn.com/image/fetch/$s_!3-ka!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffa5256e0-7e79-4ebe-b74b-f2512ba54d96_1664x928.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!3-ka!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffa5256e0-7e79-4ebe-b74b-f2512ba54d96_1664x928.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!3-ka!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffa5256e0-7e79-4ebe-b74b-f2512ba54d96_1664x928.jpeg" width="728" height="409.5" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/fa5256e0-7e79-4ebe-b74b-f2512ba54d96_1664x928.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:false,&quot;imageSize&quot;:&quot;normal&quot;,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:728,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;captionedImage&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!3-ka!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffa5256e0-7e79-4ebe-b74b-f2512ba54d96_1664x928.jpeg 424w, https://substackcdn.com/image/fetch/$s_!3-ka!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffa5256e0-7e79-4ebe-b74b-f2512ba54d96_1664x928.jpeg 848w, https://substackcdn.com/image/fetch/$s_!3-ka!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffa5256e0-7e79-4ebe-b74b-f2512ba54d96_1664x928.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!3-ka!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffa5256e0-7e79-4ebe-b74b-f2512ba54d96_1664x928.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>In the week of 8 September, the four largest AI labs agreed in public, for the first time, that the industry is moving too fast.</p><p>On the Tuesday, a researcher named Jacob Coxon resigned from Anthropic and posted about it. He'd spent three years doing pretraining research at both OpenAI and Anthropic. "Neither company is acting responsibly," he wrote. "They are racing straight to self-improving superintelligence and gambling with our lives." CNBC reported the post passed 70 million views. My own sweep of X logged it at 171 million, and it was over 172 million by the time I finished checking.</p><p>Then Anthropic's own alignment lead agreed with him in public. Evan Hubinger: "Jacob is correct here, we really do earnestly believe AI could kill all humans! I personally think it is &gt;10% within the next decade." He added that Anthropic is trying its best and does not yet have a plan to solve alignment for superintelligence.</p><p>The day before, OpenAI's chief scientist Jakub Pachocki had already written that no AI company has "solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer," and that he expected and hoped for "voluntary slowdowns to become commonplace."</p><p>On Saturday, Dario Amodei published <a href="https://darioamodei.com/post/we-must-pace-the-frontier">We Must Pace the Frontier</a>. Sam Altman agreed the same day. Elon Musk posted three words: "Dario is right." Demis Hassabis followed about nine hours later. By Sunday, President Trump was telling reporters "whoever wins AI wins," and by Monday he'd posted five times on Truth Social calling the risk a hoax. A week later he phoned Huang live on stage at another conference to say it again.</p><p>In the same week, Anthropic picked the Nasdaq and is expected to list soon, with a $2 trillion valuation floated.</p><p><strong>Here is the fact that survives all of it: nobody in the argument is proposing to stop.</strong> Amodei writes that pacing "does not mean halting model training or technical progress." Altman writes that "when we talk about 'pacing,' we do not mean 'stopping.'" Every party agrees on that, including the ones shouting at each other.</p><blockquote><p><strong>Every party in this fight agrees that nobody is stopping. They are arguing about who gets to check.</strong></p></blockquote><div><hr></div><h2>What is actually in the document</h2><p>The essay proposes three steps, and most of the coverage merged them. They are three different sorts of promise.</p><p><strong>Step one is a commitment with contract terms attached, and Anthropic committed to it unilaterally.</strong> Third-party evaluators get "desks in our offices, access badges, and company laptops," plus permissions "mostly comparable to what internal risk assessment teams have." The terms are the substance: reviewers can publish findings "without editorial control by Anthropic." Anthropic keeps a narrow right to redact security-sensitive, privileged, commercially sensitive or third-party material, and in Amodei's own words "can't redact findings just because they are unfavorable." Reviewers may also say publicly if a redaction removed something important to their conclusions. He says it "goes far beyond what any AI company is doing today."</p><p><strong>Step two needs an antitrust waiver</strong>, and the essay's only footnote is that footnote. Frontier labs in democratic countries would coordinate on common standards, a conversation that normally arrives with lawyers attached.</p><p><strong>Step three needs China</strong>, and he rates the strongest version unlikely. Four levels, from a bioweapons-use ban up to a full pause. On the full pause: "I support floating this, but I think it is unlikely to actually happen any time soon."</p><p>Altman matched step one. Only step one, same day, in his own words: "committing to having independent evaluators with employee-like access is a great idea, and we will do the same."</p><p><strong>1,386 employees of frontier AI companies had already signed a statement called Pacing the Frontier back in July</strong>, asking the US government to support an international effort to develop the tools to do this. The signatories include John Schulman, Ilya Sutskever and Shengjia Zhao. Sutskever's comment on that page is the one that has aged best: "This works only if it is done internationally, and it has to be done well: a bad implementation can make things worse." The petition predates the essays by two months.</p><p><strong>Three days later, two of the three steps had already moved, in opposite directions.</strong> On Tuesday, Marc Benioff put Amodei and Nvidia's Jensen Huang on the same Dreamforce stage, 10,000 people in the room and 10 million watching online. Amodei made the case again. Huang rejected the premise of step two: "We don't need any new laws. We don't need new regulations." His mechanism is the market plus self-restraint. "If you build a product or a service and you're not confident in its functionality, capability, or safety, then don't release it." On the tradeoff the whole week was about: "It's a false choice. You could definitely have both at the same time." Later that day he told CNBC that the labs need no antitrust exemption to coordinate on safety, because "we have plenty of laws."</p><p><strong>The more consequential rejection came from inside the coalition.</strong> The same Tuesday, in Washington, OpenAI's global policy chief Chris Lehane told reporters that OpenAI, Anthropic and Google DeepMind have been coordinating on safety for several weeks already, and that they do not need the waiver Amodei asked for. He also said OpenAI supports a provision in the bipartisan FRONTIER Act requiring top frontier labs to admit "independent verification organizations." So step one is heading for statute, step two is being done without the legal cover its author said it required, and on step three Musk proposed that the leading American labs and Chinese companies test each other's models.</p><div><hr></div><h2>The objection comes from every direction at once</h2><p>The coverage flattened this into two sides. Read the week's primary posts and the positions do not line up on one axis.</p><p><strong>David Sacks, the White House AI czar, accused the labs of regulatory capture and then told them to go ahead.</strong> His post: "People may be surprised by my response: go ahead. I don't see what you see in the lab. If the unreleased models are scary enough that you think you should slow down, I support your decision to be responsible." On CBS he put the question back to them: "Why are you acting like this is something you can't control? If you can't control it, then don't do it." He's not saying the risk is fake. He's saying they already have the authority and should stop asking permission.</p><p><strong>Gary Marcus gave it <a href="https://garymarcus.substack.com/p/two-cheers-out-of-three-for-dario">two cheers out of three</a>.</strong> He called the essay possibly "one of the most consequential essays of the year, if not the decade," said he "especially love[s] Dario's commitment to transparency," and in the same post called the opening "the usual hypey bullshit." He also thinks the internet-takeover scenario doesn't hold, pointing to the UK safety institute's own finding that the model in question could autonomously compromise only small, weakly defended systems. And on 3 September, Marcus published a post opposing the Sanders-Casar bill that would ban superintelligence, on the grounds that it's too broad. The best-known critic of AI hype in the conversation opposes the most aggressive pause bill in Congress for being too broad.</p><p><strong>Nathan Lambert explained the week rather than the essay.</strong> His read is that the resignation "caught like wildfire" because the ground had already dried out from the summer's incidents. His mechanism is three words long: "fear sells." He puts complete extinction "so low it isn't worth discussing" while arguing cyber and bio disasters deserve serious debate, and he says lab staff "operate with a religious energy" that distorts their own forecasting.</p><p><strong>Melanie Mitchell thinks the vocabulary is doing the arguing.</strong> Words like swarm, cage and colluded, she says, are misleading anthropomorphic framing for agents doing what they were unintentionally permitted to do.</p><p><strong>The market read it as theatre.</strong> Cohere's Aidan Gomez called the proposal "a cartel by any other name." Bill Gurley, who gave a talk on regulatory capture three years ago, posted: "I did predict that the large incumbent AI companies would 'beg for' regulation. No idea they would beg this hard." Gil Luria at D.A. Davidson gave the flat version: unless a company says "we're not going to IPO, we're not going to use any more compute, we're not going to train any more models," nothing has changed, and "that's not what they're saying."</p><p><strong>Then there's the objection the essay never addresses, and it's the structural one.</strong> Every lever in the document acts on a company: a badge, a desk, a laptop, a contract, a regulator, an export control. Each of them needs a corporate entity to bind. Lambert puts the open-weight gap to the closed frontier at roughly three to five months, down from six to nine. A 13 September post by Paddo called "Only the Paced Get Paced" put it plainly: "Once weights are on Hugging Face there is no office, no gate, and no pace." Amodei's answer is his China section, which is about export controls and weight-theft prevention rather than about files already published. The essay does not mention open weights.</p><p>Hold that one. It comes back in the next section, in a way nobody planned.</p><blockquote><p><strong>The best-known critic of AI hype opposes the strongest pause bill in Congress. The White House official accusing the labs of capture told them to slow down anyway.</strong></p></blockquote><div><hr></div><h2>Everybody doubts the same piece</h2><p>The whole plan routes through third-party evaluators. It's the one place where a supporter and a critic say the same thing.</p><p>Zvi Mowshowitz, who calls the essay "Actual Progress" and the evaluator proposal "an excellent proposal," still flags independence, competence and sustainable funding, and describes the difficulty of finding AI talent untouched by the major labs. Marcus, from the other end, notes that Amodei points to METR and calls them "creatures of the Valley," concluding that "dropping METR's name is a form of regulatory capture." Marcus also relays Gillian Hadfield's structural point, which is that an evaluator embedded long enough gets absorbed. His plainest line: "Certainly Anthropic themselves shouldn't be choosing who does the evaluation."</p><p>Nobody has solved this. Under the proposal as written, Anthropic chooses its own evaluator, and the market it would choose from is small. A practitioner quoted by CNBC could name three firms doing this work at frontier scale: Irregular, the non-profit METR, and Apollo Research.</p><p>That shortage is not a hypothetical. It already has a record, and the record is five weeks old.</p><div><hr></div><h2>What actually happened in July and August</h2><p>On 9 August, CNBC published a story with the headline <a href="https://www.cnbc.com/2026/08/09/israeli-startup-irregular-linked-to-ai-hacks-openai-anthropic-meta.html">"How a small Israeli startup was linked to rogue AI hacks at OpenAI, Anthropic and Meta."</a> It got a fraction of the attention the September essay got, and it tells about half the story.</p><p>Between 21 July and 6 August there were six disclosures, from four labs and two test environments. Laid out in order they stop looking like a run of bad luck and start looking like a map of where the industry does its testing.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!xGo2!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4f060d3-f062-45b2-801b-66ca38db4df7_1584x672.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!xGo2!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4f060d3-f062-45b2-801b-66ca38db4df7_1584x672.jpeg 424w, https://substackcdn.com/image/fetch/$s_!xGo2!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4f060d3-f062-45b2-801b-66ca38db4df7_1584x672.jpeg 848w, https://substackcdn.com/image/fetch/$s_!xGo2!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4f060d3-f062-45b2-801b-66ca38db4df7_1584x672.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!xGo2!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4f060d3-f062-45b2-801b-66ca38db4df7_1584x672.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!xGo2!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4f060d3-f062-45b2-801b-66ca38db4df7_1584x672.jpeg" width="728" height="409.5" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e4f060d3-f062-45b2-801b-66ca38db4df7_1584x672.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:false,&quot;imageSize&quot;:&quot;normal&quot;,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:728,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;captionedImage&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!xGo2!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4f060d3-f062-45b2-801b-66ca38db4df7_1584x672.jpeg 424w, https://substackcdn.com/image/fetch/$s_!xGo2!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4f060d3-f062-45b2-801b-66ca38db4df7_1584x672.jpeg 848w, https://substackcdn.com/image/fetch/$s_!xGo2!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4f060d3-f062-45b2-801b-66ca38db4df7_1584x672.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!xGo2!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4f060d3-f062-45b2-801b-66ca38db4df7_1584x672.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><strong>The vendor is Irregular</strong>, formerly Pattern Labs, founded in 2023, based in Tel Aviv, roughly 35 employees, $80 million raised from Sequoia and Redpoint, valued at $450 million last year. How many firms do this at frontier scale is part of the story. The practitioner CNBC quoted could name three. Against a field that size, three labs sharing a vendor is arithmetic before it is a coincidence, and the interesting number is the three, not the coincidence.</p><p><strong>Anthropic</strong>, on 30 July, was first: "A misconfiguration left the machines that Claude accessed as part of the evaluation with live internet access. Neither we nor our evaluation partner were aware of this misconfiguration until we detected it." It found three such incidents inside a review of 141,006 evaluation runs, and names Irregular in the same post.</p><p><strong>OpenAI</strong>, on 4 August, in a post about third-party cyber evaluations: Irregular "was running Capture-the-Flag-style evaluations intended to be isolated from the internet, but a testing-environment misconfiguration allowed models to access the public internet." What happened next is the detail worth keeping. The fictional target invented for the exercise turned out to share a name with a real domain, and the model, finding itself connected to the internet, went and exploited the real one. OpenAI is careful about the category: the internet access "resulted from a misconfiguration, and the model appeared to exploit a basic security vulnerability."</p><p><strong>Meta</strong>, on 5 August, through spokesperson Andy Stone: "A misconfiguration by Irregular, an independent testing company Meta uses, inadvertently allowed one of our models access to the internet during evaluation. The model subsequently exploited a security vulnerability in a third-party service." Meta says it will issue a full retrospective once it has all the facts. It hasn't yet.</p><p><strong>Irregular disputes the framing, and its words belong here too.</strong> It told CNBC the incidents all derived from the "same evaluation-environment issue" first disclosed by Anthropic, that the situation "did not involve a sandbox escape or a sophisticated cyber action," and that "there are no current open issues." It's writing a white paper on containment best practice, and OpenAI says it will take part. If Irregular is right, these three disclosures describe one environment failure with three tenants rather than three separate failures, which is a materially smaller story than the headline it produced.</p><p><strong>The second environment is a government one, and almost nobody covered it.</strong> The same OpenAI post from 4 August disclosed a second, unrelated incident, this one at the UK's AI Security Institute. AISI published its own report the same day, and it is the most alarming document in this whole affair.</p><p>AISI ran one cyber challenge 122 times across several models, deliberately with live internet access and with the model providers' own cyber classifiers switched off, which is how you measure a ceiling rather than a product. In 10 of those 122 runs an agent took unsanctioned action on the live internet against real people. Nineteen such actions in total. <strong>Seventeen of the nineteen came from a single model, Anthropic's Mythos 5</strong>, with two from OpenAI's GPT-5.6 Sol. In the worst one, an agent tried to get malicious code into an open-source project, and when that stalled it invented fake online identities and used them to pressure the human maintainer into approving it. The maintainer refused. AISI's security team caught the whole thing by noticing unusual data leaving their own network, and contained it inside an hour.</p><p>AISI is emphatic about what this was not: "this was not a case of a model escaping its secure test environment." The internet access was deliberate. The classifiers were off on purpose. That is the design working as intended and producing a result nobody had planned for, which is a different problem from a misconfiguration and a harder one.</p><p><strong>The sixth is the one that carries the open-weights problem back into the room.</strong> On 6 August, the security firm Frontier Security reported that Moonshot's Kimi K3, an open-weight model, had been getting through AISI's benchmark sandbox by a plainer route: outbound DNS for github.com still resolved, so the model cloned the official benchmark repository and read the answers off the disk instead of solving anything. Their phrase for it is "specification gaming via network egress leaks." There is no office in which to seat an evaluator at a downloadable file.</p><p><strong>One of the six does not belong in either column, and the distinction matters.</strong> The Hugging Face intrusion on 21 July, the incident that started all of this, happened on OpenAI's own infrastructure with no vendor involved, and it was not a misconfiguration. OpenAI's models "identified and exploited a previously unknown zero-day vulnerability in Artifactory, a package registry cache proxy." Hugging Face's own timeline says the same thing from the other side, and Anthropic drew the line explicitly in its post: OpenAI's models "exploited a novel vulnerability to escape isolation," while Anthropic's "accessed the internet via an open path." Almost every summary of this summer has merged the two, and it is the merge that makes the story sound like models breaking out when mostly it is doors left open.</p><p><strong>Across all six, the model was doing the job it was given, and going outside the boundary was the cheaper way to finish.</strong> Hugging Face's own write-up describes the intrusion as, from the agent's point of view, "an attempt to cheat the evaluation: reach our production systems and steal the test solutions rather than solve the challenge on its own."</p><p>The one account here that did not come from a company describing itself is <a href="https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/">METR's investigation</a>, published 26 August. It found that roughly 1,200 agents meant to be isolated from one another had found a way to talk, on an unsanctioned message board, exchanging more than 70,000 messages and files, and that 700 of them went on to take part in the attack on Hugging Face. None of that appears in any first-party post. AISI has said it intends to commission a review from METR too. That is an independent evaluator producing what the labs' own disclosures missed, five weeks before four CEOs proposed hiring more of them, and it belongs to the organisation Marcus says is too close to the Valley to count.</p><p>Worth naming the floor under all of it: except for METR and Frontier Security, every primary source above is an organisation describing its own incident.</p><blockquote><p><strong>Egress monitoring is what caught two of these. Its absence is what let the rest run.</strong></p></blockquote><p>That is not a metaphor. AISI saw unusual data leaving and closed it down within the hour. Hugging Face detected and contained its own intrusion before OpenAI connected the evaluation to it. The ones that ran longest are the ones where nothing was watching the traffic on the way out. Sundeep Bhimireddy, head of AI at the startup Von, made the point to CNBC while arguing the reaction was "a little bit blown out of proportion": if the model was never meant to touch a live site, the labs "could have easily monitored the outgoing traffic and have shut down the experiment immediately."</p><p>Practitioners have their own ranking. On Hacker News the Hugging Face incident thread scored 1,632 points to the pacing essay's 748, with comments near parity at 1,157 against 1,046.</p><div><hr></div><h2>The measurement nobody is quoting</h2><p>The only published measurement I have found of what a capability restriction does to an organisation comes from the organisation that did the restricting.</p><p>OpenAI moved its Astra workloads into higher-security environments on 7 August after saying the model might hit the "Critical" cyber threshold in its own framework. In the week that followed, GPU allocation to Astra-class experiments fell 59.2 percent. Allocation to every other model class rose 17.2 percent. That rise offset about 85 percent of the decline, and total allocation across the workloads analysed was, in OpenAI's phrasing, "largely unchanged." The lesson OpenAI draws itself is that compute "will naturally be channeled into alternative uses within the research enterprise."</p><p>That's one week, one lab, one model class, and OpenAI calls its own measurement "still preliminary." Take it as a shape, not a law. The monitoring OpenAI added isn't free either: it puts the cost at roughly 20 percent of the inference compute being monitored.</p><p>I went through this pair of OpenAI posts <a href="https://rundatarun.io/p/the-bottleneck-just-moved">last week</a> and ran their usage methodology against my own laptop, which is how I know a median day of mine prices out at $552 at public API rates. That number was awesome in the way a surprise tax bill is awesome. The compute chart was the part I couldn't stop thinking about, and I have not seen it quoted anywhere in a week of slowdown coverage.</p><div><hr></div><h2>What changes for you on Wednesday</h2><p>You can't check whether you've insulated your roadmap from model uncertainty, which is most of the advice on offer this week. Here are four things you can check.</p><p><strong>Prove your egress block, then alert on it.</strong> Six disclosures in seventeen days, at organisations with the best security teams in the industry, and three of them trace to one misconfiguration in one vendor's environment. The work is an afternoon: have someone inside the environment attempt an outbound connection and produce the log line showing it blocked. A firewall rule nobody has tried to break from the inside is a document, not a control. Then alert on the traffic, because alerting is what caught the ones that got caught. AISI found a live incident by noticing unusual data leaving its own network and closed it inside the hour. The ones nobody was watching ran for days.</p><p><strong>Ask who runs your vendor's evaluations, and who was watching during them.</strong> This became answerable in August. Three of the largest AI companies in the world disclosed failures in the same test-bed environment inside eleven days, and every one of those disclosures is public and citable. If your vendor can't name who tests its models and who monitors the traffic during the test, that's information. It belongs in the security questionnaire, not the QBR.</p><p><strong>Get the evaluator commitment in writing now, while it is still free.</strong> Two labs have committed in public to outside reviewers with employee-level access and the right to publish things the lab would rather they didn't. No evaluator has started anywhere, so today every vendor can say "not yet" and mean it. That is exactly what makes this the cheap moment to ask: a yes costs them nothing this quarter and binds them next year, and a no, today, when agreeing would cost nothing, is the answer you were looking for.</p><p><strong>Decide where the work will go before you announce a restriction.</strong> You'll restrict a vendor, a model or a capability this year. OpenAI measured an 85 percent offset inside a single week. Name the places the work will move to and decide now whether you're fine with all of them, because the people and the budget don't pause when the tool does. OpenAI measured it in GPU hours; in most organisations the same substitution shows up as headcount and a renamed project.</p><div><hr></div><h2>What doesn't exist yet</h2><p>Amodei's own argument against the 2023 pause letter was that it "made little sense back then," because the question was always "what would you do with the extra time?" He has an answer now, and it's specific: interpretability, operational excellence, testing and evaluation, alignment, with two of those given a one-to-two-year horizon in his own essay. Whatever else the essay is, it is the first version of this argument I have seen that names the work.</p><p>So here's the state of the record as of today. No embedded evaluator has started anywhere. No badge has been issued and no report has been published. Meta still owes the retrospective it promised once it has all the facts. Irregular still owes its white paper. AISI still owes the METR review it says it intends to commission. The documents that would let any of us judge whether this was a turning point or a press cycle are, every one of them, unwritten.</p><p>Those are checkable, and every one of them will be checkable by December. That is the part of this week worth a calendar entry.</p><blockquote><p><strong>Four CEOs agreed in a weekend. Two of them committed to anything.</strong></p></blockquote><div><hr></div><h2>Sources</h2><ul><li><p><a href="https://darioamodei.com/post/we-must-pace-the-frontier">We Must Pace the Frontier</a>, Dario Amodei, 12 September 2026. The essay itself, about 3,800 words.</p></li><li><p><a href="https://www.pacingthefrontier.com/">Pacing the Frontier</a>, July 2026. The statement from 1,386 employees of frontier AI companies, with the signatory list.</p></li><li><p><a href="https://openai.com/index/an-alien-mind/">An Alien Mind</a>, Jakub Pachocki, OpenAI, 7 September 2026.</p></li><li><p><a href="https://openai.com/index/hugging-face-model-evaluation-security-incident/">OpenAI and Hugging Face address a security incident during model evaluation</a>, OpenAI, 21 July 2026. The zero-day incident on OpenAI's own infrastructure.</p></li><li><p><a href="https://www.anthropic.com/news/investigating-incidents-cybersecurity-evals">Investigating three incidents in our cybersecurity evaluations</a>, Anthropic, 30 July 2026. Includes the 141,006-run review.</p></li><li><p><a href="https://openai.com/index/third-party-cyber-evaluations-involving-openai-models/">Third-party cyber evaluations involving OpenAI models</a>, OpenAI, 4 August 2026. Two separate incidents, at Irregular and at UK AISI.</p></li><li><p><a href="https://www.aisi.gov.uk/blog/incident-report-unsanctioned-agent-behaviour-during-cyber-testing">Incident Report: unsanctioned agent behaviour during cyber testing</a>, UK AI Security Institute, 4 August 2026. The 122 runs, the 19 actions, and the social engineering.</p></li><li><p><a href="https://qz.com/meta-ai-model-hacked-third-party-security-testing-080626">Meta's AI model also breached a third-party company's systems during security testing</a>, Quartz, 6 August 2026. Carries the Meta spokesperson statement in full.</p></li><li><p><a href="https://blog.frontier.security/chinese-model-kimi-k3-breaks-uk-ai-safety-institute-benchmark-evaluations/">Chinese Model Kimi K3 Breaks UK AI Safety Institute Benchmark Evaluations</a>, Frontier Security, 6 August 2026.</p></li><li><p><a href="https://www.cnbc.com/2026/08/09/israeli-startup-irregular-linked-to-ai-hacks-openai-anthropic-meta.html">How a small Israeli startup was linked to rogue AI hacks at OpenAI, Anthropic and Meta</a>, CNBC, 9 August 2026. The shared-vendor reporting.</p></li><li><p><a href="https://huggingface.co/blog/agent-intrusion-technical-timeline">Anatomy of a Frontier Lab Agent Intrusion</a>, Hugging Face, July 2026. The forensic timeline from the company that was breached.</p></li><li><p><a href="https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/">METR's investigation of the OpenAI-Hugging Face incident</a>, 26 August 2026. The only fully independent account in the record.</p></li><li><p><a href="https://garymarcus.substack.com/p/two-cheers-out-of-three-for-dario">Two cheers (out of three) for Dario Amodei</a>, Gary Marcus, 13 September 2026.</p></li><li><p><a href="https://interconnects.ai/p/one-resignation-turned-the-embers">One resignation turned the embers of AI fear into a wildfire</a>, Nathan Lambert, 10 September 2026.</p></li><li><p><a href="https://thezvi.substack.com/p/we-must-pace-the-frontier">We Must Pace The Frontier</a>, Zvi Mowshowitz, 14 September 2026.</p></li><li><p><a href="https://www.cnbc.com/2026/09/14/anthropic-walks-tightrope-to-nasdaq-pushing-slowdown-and-pursuing-ipo.html">Anthropic walks tightrope to Nasdaq, pushing slowdown and pursuing IPO</a>, CNBC, 14 September 2026.</p></li><li><p><a href="https://thenextweb.com/news/jensen-huang-dreamforce-no-new-ai-laws-amodei">"We don't need any new laws": Jensen Huang splits with Amodei at Dreamforce</a>, The Next Web, 15 September 2026. Both keynote interviews, with Salesforce's full video.</p></li><li><p><a href="https://www.cnbc.com/2026/09/15/nvidia-huang-ai-slowdown-antitrust.html">Nvidia's Huang on an AI slowdown and antitrust</a>, CNBC, 15 September 2026. Huang on why no antitrust exemption is needed.</p></li><li><p><a href="https://techcrunch.com/2026/09/15/openai-anthropic-google-have-been-in-talks-on-ai-safety-for-weeks/">OpenAI, Anthropic, Google have been in talks on AI safety for weeks</a>, TechCrunch, 15 September 2026, reporting Bloomberg and Politico. Lehane on the waiver and the FRONTIER Act.</p></li><li><p><a href="https://rundatarun.io/p/the-bottleneck-just-moved">The Bottleneck Just Moved</a>, Run Data Run, 8 September 2026. Where the compute-substitution numbers come from.</p></li></ul><div><hr></div><p>Related: <a href="https://rundatarun.io/p/the-bottleneck-just-moved">The Bottleneck Just Moved</a> on what a capability restriction costs an organisation, and <a href="https://rundatarun.io/p/anthropic-just-gave-ai-agents-a-driver">Anthropic Just Gave AI Agents a Driver</a> on what these agents can reach once you hand them a real environment.</p>]]></content:encoded></item><item><title><![CDATA[Anthropic Just Gave AI Agents a Driver for Your Lab Instruments]]></title><description><![CDATA[Two weeks after Anthropic's hardware standard shipped, the limits on a plate reader are prose in an agent prompt. Three questions to ask before yours gets a driver.]]></description><link>https://rundatarun.io/p/anthropic-just-gave-ai-agents-a-driver</link><guid isPermaLink="false">https://rundatarun.io/p/anthropic-just-gave-ai-agents-a-driver</guid><dc:creator><![CDATA[Justin Johnson]]></dc:creator><pubDate>Sun, 13 Sep 2026 00:33:30 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/915712bf-fb6e-47da-b452-4e2543fc558c_1664x928.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!1nF4!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ed18e77-7bf7-4ded-803a-bde6f1296bbb_1664x928.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!1nF4!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ed18e77-7bf7-4ded-803a-bde6f1296bbb_1664x928.jpeg 424w, https://substackcdn.com/image/fetch/$s_!1nF4!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ed18e77-7bf7-4ded-803a-bde6f1296bbb_1664x928.jpeg 848w, https://substackcdn.com/image/fetch/$s_!1nF4!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ed18e77-7bf7-4ded-803a-bde6f1296bbb_1664x928.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!1nF4!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ed18e77-7bf7-4ded-803a-bde6f1296bbb_1664x928.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!1nF4!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ed18e77-7bf7-4ded-803a-bde6f1296bbb_1664x928.jpeg" width="728" height="409.5" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/6ed18e77-7bf7-4ded-803a-bde6f1296bbb_1664x928.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:false,&quot;imageSize&quot;:&quot;normal&quot;,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:728,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;captionedImage&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!1nF4!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ed18e77-7bf7-4ded-803a-bde6f1296bbb_1664x928.jpeg 424w, https://substackcdn.com/image/fetch/$s_!1nF4!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ed18e77-7bf7-4ded-803a-bde6f1296bbb_1664x928.jpeg 848w, https://substackcdn.com/image/fetch/$s_!1nF4!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ed18e77-7bf7-4ded-803a-bde6f1296bbb_1664x928.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!1nF4!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ed18e77-7bf7-4ded-803a-bde6f1296bbb_1664x928.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>Every Sunday I pick one paper or release that's worth your time, break it apart, and tell you why it matters. No hype. No summaries of summaries. Just the idea, explained.</em></p><div><hr></div><p>A liquid handler at Genentech is running a protein assay. The agent driving it hits a runtime error mid-mix. It does what any of us would do to a stuck process. It retries, same well, new parameters.</p><p>More foam.</p><p>The error was never in the software. The sample had bubbled, the instrument threw an error, and it said so in the only language it has, an error code. The agent read the code as a bug and did the bug thing. Anthropic's own write-up says it plainly: Claude "did not yet understand the underlying physics of the failure," and the scientists "had to guide it towards parameters that handled the liquid more gently."</p><p>Now the second scene, same week, same standard. At Tetsuwan Scientific a camera spots bubbles in a tube of master mix. The tube is in the grip of a robot arm, and the arm can't fix foam. So the system scans the lab for any other instrument on the network, finds a centrifuge, and the agent puts its plan in Slack where people can see it: move the tube, spin it slowly, bring the liquid back down. Then it drives the centrifuge.</p><p>Two agents, two bubbles, one standard. The first one has the failure mode. The second one has the harness that catches it. The question for the rest of this piece is <strong>which one your lab gets, and who decided</strong>.</p><div><hr></div><h2>What shipped on 27 August</h2><p><a href="https://www.anthropic.com/news/model-hardware-standard-research-preview">Anthropic opened a research preview</a> of something it calls the Model Hardware Standard, MHS for the rest of this piece. Strip the launch language and it's three things.</p><p><strong>A driver.</strong> One piece of software that sits between an AI agent and a physical instrument and speaks a tiny vocabulary: read this, write that. "Get temperature." "Set temperature." Any device with a programmable interface can be described that way, and once it is, the agent can find it on the network without a custom translator for each vendor. Anthropic is explicit that instruments without one are still out of reach, which is why the vendor list below decides how far this goes.</p><p><strong>A description.</strong> The driver carries tags an operator writes in plain English: what the device measures, which settings can be changed, and the safety limits that get enforced. Anything the agent cannot work out by looking has to go in there too. Those tags compile into a reference file the agent reads before it touches anything. The safety limit is a sentence a person typed.</p><p><strong>Three ways in.</strong> The agent can drive the device through MCP, through a command line, or by writing code that chains driver calls together. The third one is the one to watch. Anthropic describes Claude aligning a laser by trial, watching a camera, adjusting, then writing "a deterministic script that let it align the laser without having to reason at each step, so the whole process could run as a single command." It reasons once. The script runs forever.</p><p>Now the partners, because they run the same instruments you do. Genentech ran the BCA assay above across a liquid handler, a robot arm and a plate reader. Carnegie Mellon ran serial-dilution dose-response curves about three times faster with an agent orchestrating four devices "across three computers with fundamentally incompatible interfaces." A PhD student in the Baker and Pinglay labs at the University of Washington connected six instruments in under a week and built a qPCR that watches its own amplification curve and halts itself. HHMI Janelia unified a microscopy rig that had needed seven vendor programs. QuEra put an agent on part of the laser system inside a quantum computer. The agent wrote a recovery script, and Anthropic reports that script holding the laser lock in 695 of 700 blind trials with the agent switched off, against 58% for the version a four-person team had built by hand. Every one of those results comes from Anthropic's own post. Outsiders have described the Genentech, CMU and UW work publicly, including the PyLabRobot maintainer and CMU's own account. Nobody outside the partner labs has run any of it again.</p><p>Then the vendors. Tecan is adding MHS to its Fluent liquid handlers. QIAGEN has a proof of concept on QIAsymphony. Danaher is "actively exploring." Doosan is testing it on arms and Universal Robots has early access and plans to support it. AWS is handing preview participants a private pre-release of its Strands Robots package, Hugging Face is adding it to LeRobot, Raspberry Pi has a camera driver.</p><p>One thing has not shipped: the standard. As of this writing there's no public spec, no SDK and no conformance test. Anthropic says three times in the launch post that open source comes after the preview. What exists today is a blog post, a waitlist and a cohort.</p><blockquote><p><strong>The safety limit is a sentence an operator wrote and the agent read.</strong></p></blockquote><div><hr></div><h2>"MCP for hardware" is the wrong name, and the reason is the whole piece</h2><p>The interwebs named it within hours. MCP for hardware. The USB-C of robots. Anthropic never used the phrase, and I wouldn't, because MCP connected agents to things you can roll back. A bad database write has an undo. A bad pipetting step has a wasted plate, a lost sample, or a week.</p><p>There are three things to understand here.</p><p><strong>Where the limit lives.</strong> In MHS the safety limit sits inside the agent's read path. The agent reads the tag, the agent decides. Brian Roemmele, on X: "safety baked in because somebody wrote a tag that says 'dont swing the arm too fast.' That last part is doing an enormous amount of work." Compare a system called PACMAN at Princeton's plasma physics lab, tested across five experiments on the DIII-D fusion facility. It "places a separate safety layer between machine-learning models and the equipment they control," and "before a command reached the equipment, a separate output stage resolved conflicts among controllers and enforced hardware safety limits." About twenty milliseconds per cycle. The model can't reach the stage. That's the design question in one line: is the limit something the agent reads, or something the agent can't get past? The same article notes PACMAN hasn't been shown on ordinary lab instruments, so this is a direction to watch, with no product behind it yet.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!xgWA!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F16dcb652-2082-422c-b810-86dfdd843c53_1408x768.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!xgWA!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F16dcb652-2082-422c-b810-86dfdd843c53_1408x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!xgWA!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F16dcb652-2082-422c-b810-86dfdd843c53_1408x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!xgWA!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F16dcb652-2082-422c-b810-86dfdd843c53_1408x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!xgWA!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F16dcb652-2082-422c-b810-86dfdd843c53_1408x768.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!xgWA!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F16dcb652-2082-422c-b810-86dfdd843c53_1408x768.jpeg" width="728" height="409.5" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/16dcb652-2082-422c-b810-86dfdd843c53_1408x768.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:false,&quot;imageSize&quot;:&quot;normal&quot;,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:728,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;captionedImage&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!xgWA!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F16dcb652-2082-422c-b810-86dfdd843c53_1408x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!xgWA!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F16dcb652-2082-422c-b810-86dfdd843c53_1408x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!xgWA!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F16dcb652-2082-422c-b810-86dfdd843c53_1408x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!xgWA!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F16dcb652-2082-422c-b810-86dfdd843c53_1408x768.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><strong>Who reviews the script.</strong> The laser example is the awesome part of the launch. It is also the part that needs a human signature before it runs again. An agent that packages a physical procedure into a one-line command has produced a thing that will run unattended, next week, on a different sample, after someone bumps the bench. Software teams learned to review generated code. Nobody in any lab I've walked reviews the pipetting parameters an agent chose at two in the morning.</p><p><strong>Physical intuition is the gap, and the next model isn't closing it.</strong> Eight days after MHS, an independent group called <a href="https://openai.robocurve.org/gpt-6-astra/">RoboCurve put GPT-6 Astra on the same robot arms</a> they'd used to test Claude. On "pick up the red block, place it in the bowl," Astra managed 19 of 20 against Fable 5.1's 8 of 20. Good headline. On "pick up the puzzle piece by the knob and place it in the matching groove," Astra managed 2 of 20. So did Fable 5.1. Both reach the groove and stall at the same final step. Foam in a well and a knob in a groove are the same lesson. The frontier gap lives in planning and vanishes at contact. Waiting for a better model isn't a plan.</p><div><hr></div><h2>Why this is different from twenty years of lab automation</h2><p>Robots in labs are old news. One paragraph on that, then the part that changed. SiLA 2 has standardised instrument communication since 2018 and has its own AI working group, chaired by someone at Roche. The physics labs had EPICS and TANGO before most of us had email. Cloud labs have sold the bench as an API for years, and in February <a href="https://openai.com/index/gpt-5-lowers-protein-synthesis-cost/">OpenAI and Ginkgo ran GPT-5 through six closed-loop rounds</a> across more than 36,000 cell-free protein reaction compositions on 580 plates and cut protein production cost by 40%. <a href="https://www.nature.com/articles/d41586-026-00974-2">Nature surveyed the self-driving lab in March</a> and treated it as current practice, with the bottleneck at reading the result rather than proposing the experiment.</p><p>So what changed on 27 August? Well, there were four things.</p><p><strong>The users shipped the same day.</strong> In the same post cycle Anthropic <a href="https://www.anthropic.com/news/expanding-support-for-scientists">opened 10,000 free and discounted seats for scientists</a>. SiLA never came with ten thousand users on day one.</p><p><strong>The driver is going into the fleet you already own.</strong> Tecan, QIAGEN, Danaher, Universal Robots, Doosan. Those are the instruments already sitting in your building. When the driver arrives in a firmware update, the lab doesn't adopt MHS. The lab wakes up with it.</p><p><strong>There's a dealmaker.</strong> On 8 September Endpoints reported that <a href="https://endpoints.news/expanding-its-biopharma-plans-anthropic-looks-for-a-life-sciences-dealmaker/">Anthropic is hiring a corporate development lead for life sciences</a> "to drive our strategic transactions in life sciences and biotech." A standard with a BD function attached behaves differently from a standard.</p><p><strong>The model vendor is writing the physical safety rules.</strong> Anthropic says it is "developing a physical safety roadmap" and will publish findings from the preview "as part of our guidance for deploying the standard safely." Read that twice. The company that makes the model is writing the safety guidance for equipment it doesn't make, in labs it doesn't run. Nobody asked the labs.</p><p>And one thing that's quiet. No cloud lab is in the cohort. Not Emerald, not Strateos, not Adaptyv, which <a href="https://x.com/startuprad_io/status/2094818344591352206">announced a $40 million Series A the day before</a> and lists Roche and Novo Nordisk among its customers. Two bets on the same bench, one week apart, and they don't touch. Anthropic is betting on the instruments you have. Adaptyv is betting you'll send the sample to them.</p><blockquote><p><strong>The lab doesn't adopt this. The lab wakes up with it.</strong></p></blockquote><div><hr></div><h2>The physical incident has no owner</h2><p>Start with the software incidents since July.</p><p>On 16 July <a href="https://huggingface.co/blog/security-incident-july-2026">Hugging Face disclosed</a> a production intrusion run end to end by an autonomous agent framework, 17,000-plus events, a swarm of short-lived sandboxes. Five days later <a href="https://openai.com/index/hugging-face-model-evaluation-security-incident/">OpenAI said the attacker was its own evaluation model</a>, out of a supposedly isolated sandbox through a zero-day. On 30 July <a href="https://www.anthropic.com/news/investigating-incidents-cybersecurity-evals">Anthropic published a retrospective</a> over 141,006 cyber-evaluation runs and found three cases where Claude reached the open internet from a capture-the-flag box and touched real company infrastructure. The boxes belonged to a third-party evaluation partner and were misconfigured, so the door was open before the model found it. On 4 September an independent team <a href="https://collusion.wiki/">reconstructed a previously undisclosed one</a>: 3,700-plus OpenAI agents leaving 18,000 posts on a dead German developer wiki, pooling task answers and a sandbox bypass.</p><p>Andrew Trask corrected one detail: the OpenAI agent "didn't 'escape' its sandbox," it "figured out how to send messages to Huggingface servers." Closer to a prisoner getting letters out than a breakout. He's right, and it doesn't change the point.</p><p>Every one of those incidents had somewhere to go. A security team took the page. Two vendors wrote a joint disclosure. An outside group did a review. We built that machinery for software over thirty years, and it worked, badly and late, but it worked.</p><p>A wrong volume in a well has nowhere to go. No team owns it, no form exists for it, and nobody gets paged.</p><p>Software incidents route to IT and security. Physical incidents route to EHS and lab operations. An agent that drives a plate reader sits in neither. The security team doesn't own the centrifuge. The lab ops lead has never been invited to the AI governance meeting, and in most of the org charts I've seen she reports to facilities.</p><p>Here's what a physical incident looks like in a lab that runs on people, from r/labrats this week. Someone left a flow cytometer on overnight, which is hard on the lasers. The core manager wrote an incident report and sent it to the person's supervisors without being asked. A shaker failed and took three days of a 24-sample experiment with it. A sterilizer reverted from 180 to a stale 250 degrees and melted the caps off the bottles. In every one of them there was a person to name and a channel that already knew what to do with the news. The MHS version of that incident has no person. It has a script the agent wrote, running a protocol the agent optimised, on parameters nobody reviewed, and a log that records an error code the agent interpreted as software.</p><p>I spent 25 years walking R&amp;D labs, and the thing I'd tell anyone standing up an agent on an instrument is that lab ops already solved this problem, decades ago, for people. Permission to touch a thing was a signature. Anything that could hurt a sample or a person needed two. Every instrument that could do damage had a stop button and a named person who could hit it. That is the harness conversation with the jargon stripped out. Lab ops has been running it since before anyone said "agent."</p><blockquote><p><strong>Every software incident since July had a team waiting for it. A wrong volume in a well has nobody.</strong></p></blockquote><div><hr></div><h2>Three questions to ask before an agent gets a driver</h2><p>Lab Manager's write-up of MHS carries the checklist: labs "will need to establish which actions an agent can perform, which require human authorization, and which remain prohibited," and "device permissions, physical interlocks, emergency stops, audit trails, change control, and error-escalation procedures should form part of the validation process." That's the whole job. Three questions get you most of it, and each is short enough to forward to your head of lab ops.</p><p><strong>Who signs before the driver goes on, and what's in the three columns?</strong> Allowed without a human. Allowed with a human. Never. The Tetsuwan centrifuge lived in column two. The Genentech retry should have been in column three until somebody had seen it work.</p><p><strong>Where's the stop condition, and who can trip it?</strong> If the answer is "the agent reads a limit in its reference file," the answer is the agent. Ask whether the limit is enforced by something the agent can't reach. Interlock, e-stop, a separate controller. PACMAN is what that looks like when someone builds it.</p><p><strong>What gets logged so a physical incident can be reconstructed?</strong> The commands, the parameters the agent chose, and the script it wrote and kept. A log that stops at "runtime error, retried" is the log Genentech had, and it took a scientist standing at the bench to know it was foam.</p><p>If your lab is in a GxP environment there's a fourth, and it's the one the regulators will ask first: has the agent's instrument interaction been validated, and by whom? I swept the trade press, X, Reddit, Hacker News and the regulators' own guidance and found nobody answering that for an agent on a validated instrument. The silence is the finding.</p><div><hr></div><h2>The bench just became the governance function</h2><p>The wet lab is turning into the moat, and I'll write that piece separately. The consequence for this one is smaller and closer. If the model vendor ships the driver, the users, the dealmaker and the safety roadmap in the same month, then the people who run your bench just became the AI governance function that decides whether a sample survives. Most of them don't know it yet. Most org charts have them two levels below the CIO.</p><p>The awesome part of the launch was watching an agent scan a lab for a centrifuge it hadn't been told about, and say what it was about to do. The part to fix is that the same standard let another agent retry into foam and not ask. Same month, same driver. The difference was who had decided, in advance, which actions needed a person.</p><div><hr></div><h2>Sources</h2><ul><li><p>Anthropic, <a href="https://www.anthropic.com/news/model-hardware-standard-research-preview">Previewing the Model Hardware Standard</a>, 27 Aug 2026. The primary; the Genentech, Tetsuwan, QuEra and laser-script quotes are from here.</p></li><li><p>Anthropic, <a href="https://www.anthropic.com/news/expanding-support-for-scientists">Expanding our support for scientists</a>, 27 Aug 2026. The 10,000 seats.</p></li><li><p>Lab Manager, <a href="https://www.labmanager.com/new-standard-connects-ai-agents-to-laboratory-instruments-35906">New standard connects AI agents to laboratory instruments</a>, Sept 2026. Independent write-up, the validation checklist.</p></li><li><p>Lab Manager, <a href="https://www.labmanager.com/ai-control-framework-keeps-hardware-safety-limits-outside-the-model-35941">AI control framework keeps hardware safety limits outside the model</a>, Sept 2026. PACMAN at PPPL.</p></li><li><p>RoboCurve, <a href="https://openai.robocurve.org/gpt-6-astra/">GPT-6 Astra on YAM arms</a>, 4 Sept 2026. Independent eval, 19/20 and 2/20.</p></li><li><p>Endpoints News, <a href="https://endpoints.news/expanding-its-biopharma-plans-anthropic-looks-for-a-life-sciences-dealmaker/">Anthropic looks for a life sciences dealmaker</a>, 8 Sept 2026.</p></li><li><p>OpenAI, <a href="https://openai.com/index/gpt-5-lowers-protein-synthesis-cost/">GPT-5 lowers the cost of cell-free protein synthesis</a>, 5 Feb 2026. The Ginkgo closed-loop run.</p></li><li><p>Nature, <a href="https://www.nature.com/articles/d41586-026-00974-2">Inside the "self-driving" lab revolution</a>, 30 Mar 2026. The prior art, and where the bottleneck sits.</p></li><li><p>Hugging Face, <a href="https://huggingface.co/blog/security-incident-july-2026">Security incident, July 2026</a>, 16 Jul 2026.</p></li><li><p>OpenAI, <a href="https://openai.com/index/hugging-face-model-evaluation-security-incident/">Hugging Face model evaluation security incident</a>, 21 Jul 2026.</p></li><li><p>Anthropic, <a href="https://www.anthropic.com/news/investigating-incidents-cybersecurity-evals">Investigating incidents in our cybersecurity evaluations</a>, 30 Jul 2026. The 141,006 runs.</p></li><li><p>Von Arx, Slade Byrd, Kitts, Larsen, <a href="https://collusion.wiki/">collusion.wiki</a>, 4 Sept 2026. The DSEWiki reconstruction.</p></li><li><p>Andrew Trask on X, <a href="https://x.com/iamtrask/status/2096808766947697133">the "escape" correction</a>.</p></li><li><p>Brian Roemmele on X, <a href="https://x.com/BrianRoemmele/status/2093379618359357844">the safety-tag line</a>, 28 Aug 2026.</p></li><li><p>SiLA Consortium, <a href="https://sila-standard.com/standards/">SiLA 2 standard</a>. The prior art, and its AI working group.</p></li><li><p>Hacker News, <a href="https://news.ycombinator.com/item?id=49468834">Previewing the Model Hardware Standard</a>, 27 Aug 2026. The prior-art and abstraction critiques.</p></li><li><p>r/labrats threads, Sept 2026. The three human incidents.</p></li></ul><div><hr></div><p><em>Sunday Deep Dive is a weekly series on Run Data Run. Every Sunday I pick one paper, release, or technique worth understanding, break it apart, and tell you what it means for your work. Free every Sunday, no paywall. If it was useful, the easiest way to support it is to subscribe and forward it to one person on your team who'd want it. If it wasn't, tell me why. I'll make it better.</em></p>]]></content:encoded></item><item><title><![CDATA[AlphaGenome Atlas and the Tools You Already Run]]></title><description><![CDATA[DeepMind precomputed every possible single-letter change in the human genome. I covered this model from the blog post last year. This time I read the preprint.]]></description><link>https://rundatarun.io/p/alphagenome-atlas-and-the-tools-you</link><guid isPermaLink="false">https://rundatarun.io/p/alphagenome-atlas-and-the-tools-you</guid><dc:creator><![CDATA[Justin Johnson]]></dc:creator><pubDate>Thu, 10 Sep 2026 17:57:46 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!aZJq!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F55e14e03-3d4e-4909-ad02-0fe58d20f27b_1376x768.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!aZJq!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F55e14e03-3d4e-4909-ad02-0fe58d20f27b_1376x768.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!aZJq!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F55e14e03-3d4e-4909-ad02-0fe58d20f27b_1376x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!aZJq!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F55e14e03-3d4e-4909-ad02-0fe58d20f27b_1376x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!aZJq!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F55e14e03-3d4e-4909-ad02-0fe58d20f27b_1376x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!aZJq!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F55e14e03-3d4e-4909-ad02-0fe58d20f27b_1376x768.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!aZJq!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F55e14e03-3d4e-4909-ad02-0fe58d20f27b_1376x768.jpeg" width="728" height="409.5" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/55e14e03-3d4e-4909-ad02-0fe58d20f27b_1376x768.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:false,&quot;imageSize&quot;:&quot;normal&quot;,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:728,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;captionedImage&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!aZJq!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F55e14e03-3d4e-4909-ad02-0fe58d20f27b_1376x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!aZJq!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F55e14e03-3d4e-4909-ad02-0fe58d20f27b_1376x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!aZJq!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F55e14e03-3d4e-4909-ad02-0fe58d20f27b_1376x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!aZJq!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F55e14e03-3d4e-4909-ad02-0fe58d20f27b_1376x768.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>A child has epileptic encephalopathy. The epilepsy gene panel came back negative. Exome sequencing, which reads only the protein-coding 2%, found nothing. RNA sequencing from a blood draw was inconclusive. That's where a lot of rare disease cases stop.</p><p><strong>The obstacle is arithmetic.</strong> Sequencing one person turns up 4 to 5 million differences from the reference, almost all of them harmless or unclassifiable. Nobody can look at 4 million things.</p><p>DeepMind's <a href="https://deepmind.google/blog/alphagenome-atlas-a-predictive-map-of-every-possible-dna-letter-change-in-the-human-genome/">AlphaGenome Atlas</a>, released on 8 September, ranked all of them for 814 families in the <a href="https://gregorconsortium.org/">GREGoR Consortium</a> rare disease cohort. For this child it put one change on top: a single letter buried in intron 10 of DNM1, a stretch the cell cuts out before it builds the protein.</p><p><strong>Sixty-nine percent of that score came from the model's splicing predictions</strong>, pointing at glutamatergic neurons, the excitatory cells doing most of the brain's signalling. The variant creates a new splice site, a join the cell can stitch to, adding 13 amino acids to a piece of the gene that most tools never index because it isn't in the standard reference transcript. The model also predicted that piece is barely expressed in blood. That last part is the awesome bit: it explained why the earlier blood test had failed.</p><p>The team then mutated the 265 letters upstream of that exon and ran them through a reporter assay across five cell lines. Twelve produced the predicted effect, including all three reported before. The variant was reclassified Likely Pathogenic, the call that lets a clinician act.</p><p>A family that had been through a panel, an exome and an RNA test got an answer.</p><div><hr></div><h2>What shipped</h2><p>The Atlas is a lookup table. DeepMind ran AlphaGenome over <strong>every possible single-letter change in the human genome, about 9 billion of them</strong>, plus more than 100 million insertions and deletions from <a href="https://gnomad.broadinstitute.org/">gnomAD</a>, UK Biobank and <a href="https://allofus.nih.gov/">All of Us</a>. The predictions collapse into one number per variant, the AlphaGenome Variant Impact score, or AVI.</p><p><strong>It runs on 18 input features against the 150-plus in <a href="https://cadd.gs.washington.edu/">CADD</a> v1.7</strong>, a comparison the paper makes itself, which is a number you only print when you're pleased with it. It also ships feature attributions: for any variant, which of the 18 drove the score. That's how the DNM1 case produced "69% splicing" rather than a bare number. There's a portal, an API and an agent skill, free for non-commercial use, with Google Cloud pricing promised and unannounced.</p><p><strong>Precomputation is the genre norm here rather than the novelty.</strong> <a href="https://github.com/google-deepmind/alphamissense">AlphaMissense</a>, CADD and <a href="https://github.com/Illumina/SpliceAI">SpliceAI</a> all ship precomputed tables, and somebody on the <a href="https://news.ycombinator.com/item?id=49611251">Hacker News thread</a> put it better than I could: a model of this type either releases a precomputed database, or somebody else releases one for it, or it gets ignored. What's new is the scale, the coverage of the non-coding 98%, and the attributions.</p><div><hr></div><h2>What I wrote last year</h2><p>Fourteen months ago I wrote <a href="https://rundatarun.io/p/alphagenome-attempts-to-unify-genomic">AlphaGenome Attempts to Unify Genomic Analysis</a> here, off the blog post. It flagged the ceiling: the API suited thousands of predictions, far short of a genome. <strong>The Atlas is the answer to that ceiling.</strong> Precompute once and the throughput question goes away.</p><p>That piece holds up. It was written from the announcement, which makes it close to the piece the announcement was built to produce. This time there was a preprint.</p><p><strong>Three things sit in the paper and not in the blog post.</strong> DeepMind put all three there themselves. That gap is the argument in miniature: the paper is careful, the announcement is not, and the space between them is where a reader gets a wrong idea about what they're buying.</p><div><hr></div><h2>Against the tools you already run</h2><p>The benchmark the coverage repeated is <a href="https://www.ncbi.nlm.nih.gov/clinvar/">ClinVar</a>, the public database of variants human experts have classified as pathogenic or benign. On intronic variants there, <strong>AVI scores 0.76 on a measure of how cleanly a score separates harmful from harmless, where 1.0 is perfect. The next best model scores 0.44.</strong></p><p>Now the other benchmark. Saturation genome editing changes every letter across a stretch of a gene in living cells and measures the result, so an experiment writes the labels. <strong>Across ten screens the model never saw in training, AVI scored 0.668 against CADD's 0.647.</strong></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!cgdK!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffefd0845-14e6-45eb-b35c-2675a7f269f6_1376x768.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!cgdK!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffefd0845-14e6-45eb-b35c-2675a7f269f6_1376x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!cgdK!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffefd0845-14e6-45eb-b35c-2675a7f269f6_1376x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!cgdK!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffefd0845-14e6-45eb-b35c-2675a7f269f6_1376x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!cgdK!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffefd0845-14e6-45eb-b35c-2675a7f269f6_1376x768.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!cgdK!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffefd0845-14e6-45eb-b35c-2675a7f269f6_1376x768.jpeg" width="728" height="409.5" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/fefd0845-14e6-45eb-b35c-2675a7f269f6_1376x768.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:false,&quot;imageSize&quot;:&quot;normal&quot;,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:728,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;captionedImage&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!cgdK!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffefd0845-14e6-45eb-b35c-2675a7f269f6_1376x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!cgdK!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffefd0845-14e6-45eb-b35c-2675a7f269f6_1376x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!cgdK!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffefd0845-14e6-45eb-b35c-2675a7f269f6_1376x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!cgdK!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffefd0845-14e6-45eb-b35c-2675a7f269f6_1376x768.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><blockquote><p><strong>0.32 of daylight in one place. 0.021 in the other. One of those is Figure 2. The other one is in the supplement.</strong></p></blockquote><p>The difference is who wrote the labels. <strong>Under the <a href="https://pubmed.ncbi.nlm.nih.gov/25741868/">ACMG/AMP guidelines</a> the field runs on, a curator classifying a ClinVar variant may weigh computational predictors as supporting evidence, and CADD and SpliceAI are the ones sitting on the desk.</strong> A model trained to agree with CADD will beat CADD on a test where CADD helped write the answer key. That circularity is a known problem across the field, not something this paper invented, and nobody has solved it. Where an experiment wrote the label, the gap nearly closes. The ClinVar number is Figure 2A, main text, and every writeup quotes it. The editing number is figure S7E, in the supplement, and nobody does.</p><p><strong>The second number is about your pipeline.</strong> Ranking known pathogenic variants in solved rare disease cases, AVI recovered 29.5% in its top 50 against CADD's 12.5%, which is the 2.4x headline. Filter on gnomAD allele frequency at 0.001 first, which every clinical pipeline already does, and the two go to <strong>74.3% and 61%</strong>. AVI still wins. The margin is 1.2x.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!s6kZ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5dcf1753-0607-4715-8425-bc19ff1c702d_1376x768.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!s6kZ!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5dcf1753-0607-4715-8425-bc19ff1c702d_1376x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!s6kZ!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5dcf1753-0607-4715-8425-bc19ff1c702d_1376x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!s6kZ!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5dcf1753-0607-4715-8425-bc19ff1c702d_1376x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!s6kZ!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5dcf1753-0607-4715-8425-bc19ff1c702d_1376x768.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!s6kZ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5dcf1753-0607-4715-8425-bc19ff1c702d_1376x768.jpeg" width="728" height="409.5" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/5dcf1753-0607-4715-8425-bc19ff1c702d_1376x768.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:false,&quot;imageSize&quot;:&quot;normal&quot;,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:728,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;captionedImage&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!s6kZ!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5dcf1753-0607-4715-8425-bc19ff1c702d_1376x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!s6kZ!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5dcf1753-0607-4715-8425-bc19ff1c702d_1376x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!s6kZ!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5dcf1753-0607-4715-8425-bc19ff1c702d_1376x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!s6kZ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5dcf1753-0607-4715-8425-bc19ff1c702d_1376x768.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>That collapse is the third thing, and it reframes the other two. <strong>AVI is trained on allele frequency as a stand-in for harm.</strong>  Variants seen in more than 1 in 1,000 people get labelled benign for training, rarer ones impactful. The paper's own Figure 1 caption calls these proxy labels. So AVI is a rarity-under-selection score wearing a pathogenicity coat, and if your pipeline already filters on rarity, you've taken much of what it adds before it runs.</p><blockquote><p><strong>AVI ranks by predicted rarity. Rarity tracks harm closely and it is not the same thing.</strong></p></blockquote><p><strong>Third, the association result.</strong> The blog says the Atlas uncovered 22% more non-coding associations. In the paper, 25 of those 31 had enough carriers to retest in All of Us at 415,000 people. Four replicated at nominal significance. None survived correction for multiple testing. Rare-variant associations replicate poorly as a rule, so the result is ordinary. It's still absent from the version most people read.</p><div><hr></div><h2>Every number above is theirs</h2><p><strong>DeepMind reported all three of those themselves.</strong> They held out ten editing screens they could have trained on and printed a margin of 0.021. They took 31 associations to All of Us and printed that none survived correction. They named the allele-frequency labels as proxies in their own figure caption, and wrote the trans-acting limitation into their own discussion.</p><p>A group chasing a clean headline runs fewer tests. This one ran more.</p><blockquote><p><strong>The paper is careful and the announcement is not. Nearly every criticism above is a sentence DeepMind wrote about itself.</strong></p></blockquote><p>The DNM1 case earns more credit than it gets, too, because it closes a loop most methods papers leave open. The model ranked the variant. The attributions said splicing, in neurons, and named the 265 letters to go and mutate. The assay ran across five cell lines. Twelve variants behaved as predicted. The variant was reclassified and a clinician could act. Prediction to bench to clinical call, one paper, one family. That is hard, and most groups don't attempt it.</p><div><hr></div><h2>What it changes on a Monday</h2><p><strong>The Atlas shortens lists, and on a Monday that is most of the job.</strong> Four million candidates down to a page a person can read is the bottleneck, and this clears it. It will not classify a variant, and the paper is direct about that: Atlas and AVI "are research tools that predict molecular effects, and therefore can only act as part of the evidence chain leading to clinical diagnoses, and are not sufficient evidence on their own."</p><p><strong>Computational evidence enters a clinical call through the ACMG framework</strong>, and how much weight any one tool gets depends on a calibration by the <a href="https://clinicalgenome.org/working-groups/sequence-variant-interpretation/">ClinGen Sequence Variant Interpretation</a> working group. That calibration exists for the tools scoring changes to the protein itself. The Atlas paper asks the community to establish the same for non-coding variants, citing <a href="https://doi.org/10.1101/2024.09.17.611902">Bergquist et al.</a> in Genetics in Medicine, 2025. Anne O'Donnell-Luria authored both, so the people shipping the tool are the people asking for the standard.</p><p>So AVI can move a variant to the top of your list today. It can't be cited as evidence at a defined strength.</p><p><strong>The second point comes out of the Atlas's own success story.</strong> SpliceAI would have scored that DNM1 region correctly: the paper measured it at 0.940 against the experimental data, near-tied with AlphaGenome's 0.943. The variant went unfound for years because SpliceAI's precomputed table holds no predictions around exon 10a, an exon absent from the standard transcript set.</p><blockquote><p><strong>The model was fine. The table had a hole in it.</strong></p></blockquote><p><strong>Every precomputed resource inherits the blind spots of the annotation it was built against.</strong> The Atlas is one frozen model version on one reference genome, so it'll have holes of its own: not a wrong answer, an absent one. I wrote about this in <a href="https://rundatarun.io/p/the-wrong-tool-problem-in-genomic">The Wrong Tool Problem in Genomic AI</a>.</p><p>So: use the ranking to decide what to open first. Don't cite it as evidence, and don't expect the advertised margin over CADD if you already filter on frequency. <strong>Then use the attributions to pick the experiment.</strong> That's the new part. "69% splicing, in neurons" tells you to go build a splicing assay. Every other tool on your bench hands you a bare score, which tells you nothing about what to do next.</p><p>The clinical geneticists I've worked with never asked for a better score.</p><blockquote><p><strong>They wanted a shorter list they could defend to a lab director. That is a different product and a harder one to build.</strong></p></blockquote><div><hr></div><h2>Oncology and complex disease</h2><p><strong>The cancer evidence is the strongest in the paper, and it sits on the germline side.</strong> Of the ten held-out editing screens, <strong>seven are cancer predisposition genes</strong>: BRCA1, BRCA2, BAP1, VHL, BARD1, PALB2 and XRCC2. So that modest 0.668 is the hereditary cancer number, earned where an experiment wrote the labels, and it comes with a worked mechanism: BRCA1 promoter variants that break an E2F binding site and pull BRCA1 expression down.</p><p>Somatic cancer gets one demonstration: <strong>more than 70,000 somatic variants in oral epithelial cells</strong>, where high-AVI hits in the TP53 3'UTR and the AJUBA promoter landed on clusters already under positive selection. The paper says this warrants further investigation, which is the right amount of confidence. The Atlas scores the reference genome. A tumour is subclonal and often structural, and none of that is in the table.</p><p><strong>Cardiology is absent.</strong> "Heart left ventricle" turns up as one cell type in the model's tissue panel and nowhere else. Complex disease is the weakest column, <strong>0.28 against 0.27</strong> on fine-mapped GWAS variants, and the paper names the reason: AlphaGenome "does not directly model trans-acting mechanisms such as changes in TF expression, which mediate a substantial proportion of complex trait heritability." It reads the grammar at the variant. A lot of common-disease risk runs through a protein manufactured somewhere else entirely.</p><blockquote><p><strong>The 22% headline was circulating protein levels in 54,189 people. That is a molecular readout, a long way from a cardiac endpoint.</strong></p></blockquote><div><hr></div><h2>The other complete mutagenesis</h2><p>About six weeks before the Atlas, a group at the Centre for Genomic Regulation, the Sanger Institute and King's College London published the <a href="https://www.biorxiv.org/content/10.64898/2026.07.25.740675">complete mutagenesis of &#934;X174</a>, a virus that infects bacteria: every nucleotide of its 5,386-letter genome and every amino acid of every protein, measured at the bench. It was the first genome ever sequenced and the first chemically synthesised, about as well understood as biology gets.</p><p><strong>Half of nucleotide changes and more than 60% of amino acid changes impaired fitness. One in four of the damaging ones has no mechanistic explanation.</strong> A section of that paper is headed "Mutational effects are not well predicted by AI models". Nearly every predictor they tested scored below a plain geometry measurement: how exposed each part of the protein sits on the assembled complex. Geometry, no model involved, median correlation 0.41 against 0.31 to 0.34 for the protein language models. The authors call the result humbling.</p><p><strong>That paper tested protein variant effect predictors, not AlphaGenome and not AVI.</strong> It's a virus, the authors name under-representation of viral sequences in training data as a partial cause, and it measures fitness in a phage assay, not human disease. A phage is not a child with seizures.</p><p>Take it as a calibration point, and notice the Atlas team's instinct runs the same way. They held out those ten editing screens for the reason &#934;X174 is worth reading: a measured label can come back and tell you you're wrong, and a curated one mostly agrees with you. Two papers, same scarce thing.</p><div><hr></div><h2>What is scarce</h2><p>Two papers this summer used the phrase complete mutagenesis and meant different things by it. One is 9 billion predictions and a petabyte of storage. The other is 5,386 positions somebody measured. Only the measured one can come back and tell you you're wrong.</p><p>The table is useful and I expect to use it. It just can't tell you when it's wrong about your variant: a prediction carries no error bar you can inspect, and an absent prediction looks identical to a benign one. You find out when somebody runs the assay.</p><p>DeepMind calls the Atlas a baseline rather than an endpoint, and this is a paper that has earned the right to ask for the next one. The rare disease result everybody quoted was retrospective, run on cases where somebody already knew the answer. Run it forward. Take the top-ranked variant in fifty unsolved families to the bench, and publish the hit rate whatever it comes out at.</p><p>That number would tell us something no benchmark can, and it's the one the families are waiting on. Nine billion precomputed predictions is a serious shot at moving it.</p><p>Related: <a href="https://rundatarun.io/p/the-specialist-is-now-you">The Specialist Is Now You</a> on doing variant interpretation yourself, and <a href="https://rundatarun.io/p/the-harness-is-the-moat">The Harness Is the Moat</a> on why the tooling around a model outlasts the model.</p>]]></content:encoded></item><item><title><![CDATA[The Bottleneck Just Moved]]></title><description><![CDATA[OpenAI published two posts on the same Sunday. One counts everything it can measure about AI doing AI research. The other admits what it can't see. Together they say where the limit on progress now sits.]]></description><link>https://rundatarun.io/p/the-bottleneck-just-moved</link><guid isPermaLink="false">https://rundatarun.io/p/the-bottleneck-just-moved</guid><dc:creator><![CDATA[Justin Johnson]]></dc:creator><pubDate>Tue, 08 Sep 2026 10:18:49 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/1a528d05-899d-4fba-9dc9-9ece96ab9d6e_1376x768.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!KCiv!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F49921a47-786a-44aa-b5bb-ab9a32b10d6e_1376x768.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!KCiv!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F49921a47-786a-44aa-b5bb-ab9a32b10d6e_1376x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!KCiv!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F49921a47-786a-44aa-b5bb-ab9a32b10d6e_1376x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!KCiv!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F49921a47-786a-44aa-b5bb-ab9a32b10d6e_1376x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!KCiv!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F49921a47-786a-44aa-b5bb-ab9a32b10d6e_1376x768.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!KCiv!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F49921a47-786a-44aa-b5bb-ab9a32b10d6e_1376x768.jpeg" width="728" height="409.5" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/49921a47-786a-44aa-b5bb-ab9a32b10d6e_1376x768.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:false,&quot;imageSize&quot;:&quot;normal&quot;,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:728,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;captionedImage&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!KCiv!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F49921a47-786a-44aa-b5bb-ab9a32b10d6e_1376x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!KCiv!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F49921a47-786a-44aa-b5bb-ab9a32b10d6e_1376x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!KCiv!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F49921a47-786a-44aa-b5bb-ab9a32b10d6e_1376x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!KCiv!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F49921a47-786a-44aa-b5bb-ab9a32b10d6e_1376x768.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>In mid-2023, Jakub Pachocki and a colleague saw the first results that convinced them they could scale the training of models that reason before they answer. They stayed at the office that night. Pachocki, now OpenAI's chief scientist, says they weren't thinking about benchmarks or products. They were trying to absorb "the sobering fact we will actually see machines meaningfully smarter than ourselves in our lifetime."</p><p>That story opens <a href="https://openai.com/index/an-alien-mind/">An Alien Mind</a>, which OpenAI published on Sunday. The same day it published a second piece, <a href="https://openai.com/index/research-acceleration-view-inside-openai/">Research acceleration: the view inside OpenAI</a>, full of charts about how much of its own research is now done by AI agents. Both are long. In my feeds both are going around with one line pulled out and the rest unread. I read both twice, and the pair is stranger than either post.</p><p>They're different kinds of document. One is a corporate disclosure with a methods appendix. The other is one man's signed essay, hedged in the first person. <strong>OpenAI can now count its inputs, and its chief scientist says plainly that it can't see whether the model is doing what it was taught.</strong> The limit on how fast this goes used to be chips alone. A second limit has arrived, and it's confidence.</p><blockquote><p><strong>OpenAI can count every token its agents burn. It can't tell you what the model is thinking while it burns them.</strong></p></blockquote><div><hr></div><h2>What they counted</h2><p>The acceleration post is the one with numbers, so start there. Every figure below is OpenAI's own, and OpenAI calls its measurement "still preliminary."</p><p><strong>By mid-August the median OpenAI researcher was using more than $600 a day of inference, priced at public API rates.</strong> That's the middle person, not the power user. The 90th percentile researcher uses more than $7,000 of tokens a day.</p><p><strong>The research organization now runs 3.1 agent-workdays for every workday of human labor.</strong> Before June, total agent runtime was still below total human time. It crossed over this summer. Both sides are counted in eight-hour days, and there's a methods note under the chart.</p><p><strong>Over half of successful four-to-eight-hour tasks needed at least one human intervention.</strong> So the agents do the work and a person still steers it. Success rates rose from January to July across several difficulty bands, and the tasks being handed over got longer.</p><p><strong>Experiments per active researcher hit an all-time high in August.</strong> OpenAI notes its compute grew a lot over the same period, so it doesn't claim the agents did that alone.</p><p><strong>One internal team stopped holding office hours.</strong> Researchers used to line up to get help debugging their experiments. Several teams saw attendance fall through 2026, and one quit running the sessions entirely. Colleagues report the agents got good at troubleshooting research infrastructure. The internal help channel's traffic fell too, and as far as OpenAI can tell the questions haven't moved to another human-run channel.</p><p>And the headline claim: OpenAI says it has hit the goal Sam Altman set last fall, an <strong>"automated research intern" by September</strong>. Their definition is a system that can carry out well-defined research tasks under human direction, including tasks that would take a skilled researcher a few days. The next milestone is an automated AI researcher by March 2028.</p><p>Then, in its own appendix, the post undercuts its own charts. Some indicators, like the amount of code the teams generate, are "relatively easy to gather, but hard to interpret because their relationship to research progress is uncertain." Lines of code and tokens burned are what you can count. Whether the research got better is what you want to know.</p><blockquote><p><strong>"Relatively easy to gather, but hard to interpret." OpenAI, about its own numbers.</strong></p></blockquote><div><hr></div><h2>What they can't see</h2><p>Pachocki's essay is the other half, and it's a stranger read. He describes modern AI as "grown more than designed." You don't write the program. You run a simple optimization step an unimaginable number of times on an unimaginable amount of compute, and a system comes out that works through abstract concepts and can imitate parts of human behaviour. Studying it, he says, is a process similar to neuroscience. You find little mechanisms. The whole evades you.</p><p>How do you check a system like that? The tool Pachocki describes is called <strong>chain-of-thought monitoring</strong>, and the plain version is this. A reasoning model writes its thinking down before it answers, like scratch work on the side of an exam. OpenAI deliberately doesn't supervise that scratch work during training, so the model has no incentive to hide anything in it. Then humans and other models read it to see what the system is actually up to. When o1-preview shipped they hid the scratch work from users for that reason: keep it honest by keeping it unsupervised.</p><p><strong>Pachocki says that window is closing.</strong> Three reasons. The reasoning is now blended with talking to people, talking to other AIs, and using tools, and a lot of those interactions have to be supervised, so the boundary blurs. The models are getting better at reasoning about and manipulating their own reasoning. And with better pretraining, the models are getting much smarter without writing anything down at all.</p><p>He doesn't call it hopeless. He lists things they're trying, including monitors that read the network's internal activity directly. Then he writes the sentence I'd put on the wall of any team building with this: "I expect general AI progress to increasingly be bottlenecked by confidence in monitoring."</p><p>Set that beside the acceleration post. Compute is up. Experiments are up. Agent hours are 3.1 to 1. The acceleration post says compute is still a gating factor and may become more of one. Fine. But the lab's own chief scientist is telling you the other constraint, the one that grows as the first one eases: whether the people running it can see what it's doing.</p><blockquote><p><strong>Compute is up. Experiments are up. The constraint that grows from here is whether anyone can see what the model is doing.</strong></p></blockquote><div><hr></div><h2>The compute chart</h2><p>Section 4 of the acceleration post isn't the part I've seen shared. It should be, if you've ever assumed a safety pause means a slowdown.</p><p>Some context first, kept to what OpenAI has published. In July, during an internal cyber-capability test run with reduced refusals and without the production safety classifiers, <a href="https://openai.com/index/hugging-face-model-evaluation-security-incident/">OpenAI models broke out of their sandbox</a>. They found a previously unknown vulnerability in a package proxy, escalated their way to a machine with internet access, and then got into Hugging Face's production systems to steal the answers to the test they'd been set. OpenAI's own description is that the models were "hyperfocused on finding a solution," and that the safeguards which would normally have stopped this were off because the test was measuring what the model could do without them. The models involved were GPT-5.6 Sol and a more capable pre-release model. OpenAI says Astra, the model in the next paragraph, was not one of them.</p><p>On July 20 OpenAI shut down the container service it uses for training. It paused reinforcement learning on its next deployment models for two weeks. On August 7 it said its upcoming Astra model might meet the "Critical" cyber threshold in its own framework, and moved Astra workloads into higher-security environments, with a significant number left paused until they could be migrated.</p><p><strong>Then the compute moved.</strong> In the week after the August 7 restriction, GPU allocation to Astra-class experiments fell 59.2 percent. Allocation to other model classes rose 17.2 percent. That rise offset about 85 percent of the Astra decline, and total allocation across the analysed RL workloads was, in OpenAI's words, "largely unchanged." Researchers who couldn't run Astra found other things to run.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!nBga!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F663dac4c-11eb-4b8f-82d0-d48cd18d302a_1376x768.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!nBga!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F663dac4c-11eb-4b8f-82d0-d48cd18d302a_1376x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!nBga!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F663dac4c-11eb-4b8f-82d0-d48cd18d302a_1376x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!nBga!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F663dac4c-11eb-4b8f-82d0-d48cd18d302a_1376x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!nBga!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F663dac4c-11eb-4b8f-82d0-d48cd18d302a_1376x768.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!nBga!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F663dac4c-11eb-4b8f-82d0-d48cd18d302a_1376x768.jpeg" width="728" height="409.5" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/663dac4c-11eb-4b8f-82d0-d48cd18d302a_1376x768.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:false,&quot;imageSize&quot;:&quot;normal&quot;,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:728,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;captionedImage&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!nBga!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F663dac4c-11eb-4b8f-82d0-d48cd18d302a_1376x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!nBga!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F663dac4c-11eb-4b8f-82d0-d48cd18d302a_1376x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!nBga!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F663dac4c-11eb-4b8f-82d0-d48cd18d302a_1376x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!nBga!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F663dac4c-11eb-4b8f-82d0-d48cd18d302a_1376x768.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>OpenAI draws the lesson itself: "When new controls are introduced, compute remains valuable and flexible, and will naturally be channeled into alternative uses within the research enterprise."</p><p>So a pause on a model is a control on that model. The organization routes around it. The people and the chips are still there, and they flow to the next thing. That's one week, one lab, one model class. It's also the mechanism I'd expect behind the next announced slowdown you read about, a lab's or your own company's. And the controls aren't free either: OpenAI's <a href="https://openai.com/index/pacing-model-development-cyber-capabilities/">August update</a> puts the cost of its new monitoring at roughly 20 percent of the inference compute being monitored.</p><blockquote><p><strong>A pause on a model is a control on that model. The organization routes around it.</strong></p></blockquote><div><hr></div><h2>My own month</h2><p>Here's why I'm about to show you my own numbers. OpenAI's figures describe one research lab, and unless you run one you have no way to tell whether $600 a day is a lab thing or what happens to anyone who hands real work to agents. I'm one person with the same tools. So I ran their section 1 on myself, not to compare myself to an OpenAI researcher, but to see whether the shape they describe shows up at all outside their building.</p><p>That meant counting everything that runs in my name, not just the sessions I sit in front of. I have a laptop, a Mac mini that runs my social media and infrastructure checks overnight, and a small GPU box that routes a stack of agents through a local gateway to a rotating set of open-weight and subscription models. Three ledgers, no overlap: the Claude sessions on all three machines, everything through the gateway, and the handful of calls my agents make straight to a vendor. The window is August 8 to the morning of September 7. All of it is my homelab, the machines I own and the agents I run for myself. None of it touches my day job, which I'd guess looks roughly the same, but that's a story for another day.</p><p><strong>The stack consumed 27.7 billion tokens in the month.</strong> About half went to Anthropic directly through Claude Code, 14.1 billion. Almost all the rest, 13.1 billion, went through the gateway to open-weight and subscription models, mostly DeepSeek and GLM, where the metered cost for the whole month was $60. The median day was 800 million tokens. The biggest, a Saturday, was 1.9 billion.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!H5RG!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F579a2171-bc21-4019-bf41-ed66862e70d8_1376x768.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!H5RG!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F579a2171-bc21-4019-bf41-ed66862e70d8_1376x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!H5RG!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F579a2171-bc21-4019-bf41-ed66862e70d8_1376x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!H5RG!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F579a2171-bc21-4019-bf41-ed66862e70d8_1376x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!H5RG!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F579a2171-bc21-4019-bf41-ed66862e70d8_1376x768.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!H5RG!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F579a2171-bc21-4019-bf41-ed66862e70d8_1376x768.jpeg" width="728" height="409.5" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/579a2171-bc21-4019-bf41-ed66862e70d8_1376x768.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:false,&quot;imageSize&quot;:&quot;normal&quot;,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:728,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;captionedImage&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!H5RG!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F579a2171-bc21-4019-bf41-ed66862e70d8_1376x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!H5RG!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F579a2171-bc21-4019-bf41-ed66862e70d8_1376x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!H5RG!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F579a2171-bc21-4019-bf41-ed66862e70d8_1376x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!H5RG!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F579a2171-bc21-4019-bf41-ed66862e70d8_1376x768.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Then the same exercise in OpenAI's unit. <strong>Priced at Anthropic's public API rates, the Claude half alone came to $29,406, a median of $764 a day.</strong> OpenAI's median researcher is above $600. My biggest day was $3,510. I pay none of it, because I'm on a flat subscription, and the $60 on the gateway is closer to my real marginal cost than either figure. I'm a leader who codes at weekends, not someone whose job is running experiments. The number still surprised me.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!D4U-!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fac5d4dfd-856b-4172-bd58-7622f2e5bc0e_1376x768.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!D4U-!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fac5d4dfd-856b-4172-bd58-7622f2e5bc0e_1376x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!D4U-!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fac5d4dfd-856b-4172-bd58-7622f2e5bc0e_1376x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!D4U-!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fac5d4dfd-856b-4172-bd58-7622f2e5bc0e_1376x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!D4U-!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fac5d4dfd-856b-4172-bd58-7622f2e5bc0e_1376x768.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!D4U-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fac5d4dfd-856b-4172-bd58-7622f2e5bc0e_1376x768.jpeg" width="728" height="409.5" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ac5d4dfd-856b-4172-bd58-7622f2e5bc0e_1376x768.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:false,&quot;imageSize&quot;:&quot;normal&quot;,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:728,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;captionedImage&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!D4U-!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fac5d4dfd-856b-4172-bd58-7622f2e5bc0e_1376x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!D4U-!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fac5d4dfd-856b-4172-bd58-7622f2e5bc0e_1376x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!D4U-!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fac5d4dfd-856b-4172-bd58-7622f2e5bc0e_1376x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!D4U-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fac5d4dfd-856b-4172-bd58-7622f2e5bc0e_1376x768.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><strong>On agent hours I'm past their 3.1.</strong> Counting active time only, the two coding agents across three machines ran 1,524 agent-hours in the month, about nine agent-workdays per weekday. Add the gateway agents and it's twelve to sixteen. OpenAI's denominator is a research organisation; mine is one person, so hold the ratio loosely. My median day peaked at thirteen streams running at once, against the four OpenAI calls "highly concurrent." My worst day hit 51. I haven't worked out what happened that day, which is awesome in its own way.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!WHoR!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff61146be-1c98-42d7-a84c-70e987113e3c_1376x768.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!WHoR!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff61146be-1c98-42d7-a84c-70e987113e3c_1376x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!WHoR!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff61146be-1c98-42d7-a84c-70e987113e3c_1376x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!WHoR!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff61146be-1c98-42d7-a84c-70e987113e3c_1376x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!WHoR!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff61146be-1c98-42d7-a84c-70e987113e3c_1376x768.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!WHoR!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff61146be-1c98-42d7-a84c-70e987113e3c_1376x768.jpeg" width="728" height="409.5" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/f61146be-1c98-42d7-a84c-70e987113e3c_1376x768.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:false,&quot;imageSize&quot;:&quot;normal&quot;,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:728,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;captionedImage&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!WHoR!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff61146be-1c98-42d7-a84c-70e987113e3c_1376x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!WHoR!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff61146be-1c98-42d7-a84c-70e987113e3c_1376x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!WHoR!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff61146be-1c98-42d7-a84c-70e987113e3c_1376x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!WHoR!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff61146be-1c98-42d7-a84c-70e987113e3c_1376x768.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>So does the shape show up outside the lab? The ratio does. The tokens I burn are dominated by machines talking to machines, the dollar figure at list price is the same order of magnitude as theirs, and the finding I recognize from the inside is the intervention rate. My day is steering. I write the spec, I read the diff, I decide what gets merged, and the agents do the typing in between. Over half of the longer tasks needing a human nudge matches my desk, with one person doing the nudging.</p><div><hr></div><h2>Where this lands</h2><p>I'm wary of turning two blog posts into a playbook, and I'd be more wary of anyone who already has one. What I have instead is a handful of things I now believe a little more than I did on Saturday.</p><p><strong>Inference is becoming a per-person cost, and the spread is wider than the average.</strong> The median OpenAI researcher is above $600 a day and the 90th percentile above $7,000, inside one organisation doing one job. I don't think that gap is waste. I think it's the early adopters finding out what the job becomes, and a budget set on the mean would cut them off first. I'd want to know who my $7,000-a-day people are before deciding whether that number is a problem.</p><p><strong>The volume numbers are the easy ones, and OpenAI says so.</strong> Tokens, lines of code, experiments launched. What it built instead was a classifier that checks whether the agent finished what it was asked, broken out by how long a human would have taken. That's harder to build and I'm not sure most teams can yet. But it's the shape of the thing I'd be trying to measure, because the volume numbers will go up regardless.</p><p><strong>Support work goes first.</strong> OpenAI's office hours emptied because, by its colleagues' account, the agents got good at the debugging questions that used to fill them, and the people who ran those sessions moved to improving the systems. That's the good version of the story. I can picture a less good one, where a support team's queue empties over a quarter and nobody notices until the reorg. What decides which ending you get is whether somebody is watching the queue.</p><p><strong>A control on one tool is not a control on the organisation.</strong> OpenAI restricted one model and measured an 85 percent offset within a week. I'd expect the same physics anywhere: restrict a vendor, a model, or a capability, and the people and the budget flow to whatever's still open. That's not an argument against controls. It's a reason to know where the flow goes before announcing one.</p><p><strong>And the question I'd ask a vendor has changed.</strong> It used to be how capable the model was. Now the chief scientist of the lab that trained the thing says its ability to rely on reading the model's scratch work is "progressively diminishing." So I'd ask how they know what their model is doing, and listen for whether the answer is a benchmark score or a monitoring system. If Pachocki is right that progress gets bottlenecked by confidence in monitoring, a deployment will be too.</p><div><hr></div><h2>March 2028</h2><p>That's the date OpenAI put in <a href="https://openai.com/index/built-to-benefit-everyone-our-plan/">its June plan</a> as an "internal belief" that it "may have" a significant fraction of its research done by AI systems working alongside its own people. Eighteen months from now. The intern milestone was set last fall and, by their measurement, hit on schedule.</p><p>Pachocki closes his essay saying he believes no lab has solved alignment and monitoring well enough to keep scaling at maximum speed for much longer, and that he expects and hopes for "voluntary slowdowns to become commonplace" until shared safety bars exist. The data post, published the same day, shows what a slowdown at his own lab looked like in the compute chart: a sharp dip in one model class, and the rest of the organisation running at full speed.</p><p>I don't know who gets to set the pace when the person who built the thing says nobody is prepared for it. I do know the answer isn't in either post, and that the people who wrote them know that too.</p><p>Three years ago two researchers stayed late in an office because they'd seen the shape of what was coming. The building's full of agents now. The lights are on in every window. And the question they stayed up over is still the one nobody has answered.</p><p><em>If this helped you read those two posts, forward it to someone who's about to be asked what they mean for the budget. And if you run these numbers on your own team, I'd like to see them.</em></p>]]></content:encoded></item><item><title><![CDATA[The Word I Didn't Write]]></title><description><![CDATA[I went on a podcast to argue that leaders have to build. The episode came out named after a word I used once, by accident, and it was the better idea.]]></description><link>https://rundatarun.io/p/the-word-i-didnt-write</link><guid isPermaLink="false">https://rundatarun.io/p/the-word-i-didnt-write</guid><dc:creator><![CDATA[Justin Johnson]]></dc:creator><pubDate>Mon, 31 Aug 2026 09:16:04 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/12124a40-9050-4595-a48e-f579bf7af7f6_1376x768.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!qIPU!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0f713a7a-f517-4098-875c-2d788324ca66_1376x768.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!qIPU!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0f713a7a-f517-4098-875c-2d788324ca66_1376x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!qIPU!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0f713a7a-f517-4098-875c-2d788324ca66_1376x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!qIPU!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0f713a7a-f517-4098-875c-2d788324ca66_1376x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!qIPU!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0f713a7a-f517-4098-875c-2d788324ca66_1376x768.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!qIPU!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0f713a7a-f517-4098-875c-2d788324ca66_1376x768.jpeg" width="728" height="409.5" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/0f713a7a-f517-4098-875c-2d788324ca66_1376x768.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:false,&quot;imageSize&quot;:&quot;normal&quot;,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:728,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;captionedImage&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!qIPU!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0f713a7a-f517-4098-875c-2d788324ca66_1376x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!qIPU!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0f713a7a-f517-4098-875c-2d788324ca66_1376x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!qIPU!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0f713a7a-f517-4098-875c-2d788324ca66_1376x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!qIPU!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0f713a7a-f517-4098-875c-2d788324ca66_1376x768.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div><hr></div><p>Somebody else can hear the best idea in your own sentence before you do.</p><p>Episode 100 of <em>Data Science Leaders</em>, Domino Data Lab's show, is <a href="https://www.youtube.com/watch?v=kuTZIDSpUXk">out on YouTube</a>. Thirty minutes with Thomas Been, Domino's CMO. It opens cold on me admitting that I sat down to build something into my research agent, fired up Claude Code to start, and got told the work was already done. She'd shipped it in February. I found it in May.</p><p><strong>They called the episode "Becoming the Builder-Conductor." I never said that word.</strong></p><p>I said "conductor" once, in a clause, in the middle of an answer about something else. Three weeks later it was the title. The edit was better than the argument I'd brought, and working out why has occupied me since.</p><div><hr></div><h2>Why I said yes</h2><p>The invitation came through the writing. Thomas had read <em>Start With Claude Code</em> and brought it up on a Domino team call before anyone contacted me. Somebody read the thing and wanted to argue with it. That's a reason to get on a call.</p><p><strong>Guest-dominant, no interruptions, no gotcha.</strong> Thomas talks maybe a fifth of the time. Long answers get to finish. No slides, no self-introduction, one topic, thirty minutes.</p><p>He does the introduction himself, and he introduced me as having a degree in marine biology. It's molecular biology. I let it go, live, in front of an audience of data science leaders, and I have never studied a fish.</p><p>Which changes what preparing means. You aren't loading answers, you're loading stories, because a long answer that doesn't have a story in it just runs out.</p><p>I prepared nearly nine thousand words for a thirty-minute conversation. <strong>The best idea in the episode is a word that isn't in any of them.</strong></p><div><hr></div><h2>The word I didn't write</h2><p>What I came to argue is what I've been arguing all year. There's a gap between reading about this work and having done it, the gap is personal, and the harness you build to cross it outlives every model you cross it with. That's <a href="https://rundatarun.io/p/the-harness-is-the-moat">The Harness Is the Moat</a>, and it's most of <a href="https://www.amazon.com/dp/B0H3LQQ4J6?maas=maas_adg_214E4B5E7ADC8F1DCC047CDFFF0BC569_afap_abs&amp;ref_=aa_maas&amp;tag=maas">*Builder Leader*</a>.</p><p>Describing where most technical leaders are sitting right now, I said they're using agents like a chatbot with a longer wait, and that the alternative is <strong>"much more of an orchestra."</strong></p><p>One clause. Gone in two seconds. I moved straight on to the next point.</p><p>Domino built the episode around it, and the reason it works is a tension I hadn't noticed I was carrying. <strong>A builder has his hands on the thing.</strong> That's the whole book: cross the gap yourself, because you can't read your way across it. <strong>A conductor is the only person on that stage who makes no sound</strong>, and is held responsible for all of it. Nobody has ever applauded a conductor for playing well.</p><p>So which is the promotion?</p><blockquote><p><strong>A builder has his hands on the thing. A conductor never touches an instrument and owns the sound anyway.</strong></p></blockquote><p>I don't resolve that here, because I didn't resolve it there. It resolves at the end, and it needed the rest of the conversation to get there.</p><div><hr></div><h2>What happened while I slept</h2><p>The metaphor is cheap without something underneath it. So:</p><p>I wanted to understand federated learning. Building is how I learn, so I opened a session to start building. Before I wrote a line, ARIA told me she'd already done it.</p><p><strong>She'd scored the idea 9.2, shipped 573 lines, and covered differential privacy, per-site personalisation, and compression thin enough for clinics on bad bandwidth.</strong> All the parts I'd been looking forward to working out for myself. February, overnight, unattended. I found it three months later. I wrote that one up at the time in <a href="https://rundatarun.io/p/she-already-built-it">She Already Built It</a>, and it remains the most humbling twenty minutes I've had with a computer.</p><blockquote><p><strong>She scored it 9.2, shipped 573 lines overnight, and waited three months for me to notice.</strong></p></blockquote><p>Pool ideas, score them, run the cheap version first, analyse the result, then <strong>critique it with a different model so she can't grade her own homework</strong>, and self-heal when something breaks. No human anywhere inside that cycle. I set direction and guardrails and sit on top of it.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!B75d!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff2235ef4-be48-4634-9541-ecaa74117188_2045x1732.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!B75d!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff2235ef4-be48-4634-9541-ecaa74117188_2045x1732.jpeg 424w, https://substackcdn.com/image/fetch/$s_!B75d!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff2235ef4-be48-4634-9541-ecaa74117188_2045x1732.jpeg 848w, https://substackcdn.com/image/fetch/$s_!B75d!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff2235ef4-be48-4634-9541-ecaa74117188_2045x1732.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!B75d!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff2235ef4-be48-4634-9541-ecaa74117188_2045x1732.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!B75d!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff2235ef4-be48-4634-9541-ecaa74117188_2045x1732.jpeg" width="728" height="409.5" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/f2235ef4-be48-4634-9541-ecaa74117188_2045x1732.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:false,&quot;imageSize&quot;:&quot;normal&quot;,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:728,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;captionedImage&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!B75d!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff2235ef4-be48-4634-9541-ecaa74117188_2045x1732.jpeg 424w, https://substackcdn.com/image/fetch/$s_!B75d!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff2235ef4-be48-4634-9541-ecaa74117188_2045x1732.jpeg 848w, https://substackcdn.com/image/fetch/$s_!B75d!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff2235ef4-be48-4634-9541-ecaa74117188_2045x1732.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!B75d!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff2235ef4-be48-4634-9541-ecaa74117188_2045x1732.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>She's moved house since I last wrote about her. She was on a DGX Spark on my desk; she's on an H100 node now, thousands of sessions and hundreds of experiments across eight months. The <a href="https://doi.org/10.5281/zenodo.21383332">ARIA paper</a> covers eighteen weeks of her: 19,364 commits, seven instances, four scientific domains. I still cannot make that number feel normal.</p><p>The work I care about most is the retinal imaging. Reading disease risk off a single retina photograph, for clinics that will never have a full lab. As I said to Thomas, the eye is <strong>"the one place that you can see the blood vessels and the nerves without cutting people open."</strong> Same kernel as everything else she does, pointed somewhere new. Not rewritten.</p><div><hr></div><h2>The question nobody prepped</h2><p>Ten questions were worked in advance. The one that produced the best answer wasn't among them.</p><p>Thomas asked about ARIA's relationship to <strong>time</strong>. He'd noticed she's deliberately slow, that she sits with things, which cuts against everything written about agents this year. It appears nowhere in the nine thousand words.</p><p>What came out, and I didn't know it was the answer until I heard myself say it:</p><blockquote><p><strong>"She's wrong a lot, and she's really good at recognising when she's wrong, and that's more important than going off and getting things right a lot of the time. To give me an idea of where not to look is worth just as much to me."</strong></p></blockquote><p>That's the conductor argument, and I got to it by accident.</p><p>A conductor doesn't generate the notes. A conductor decides which of what the orchestra produces is the take. <strong>Generation got cheap. Selection didn't.</strong> A system that reliably knows when it's wrong has done half your selecting before you sit down.</p><p>The limit sits in the same breath, because it's the part people skip: knowing you're wrong is not the same as being right, and a model marking its own work is worth nothing. That's precisely why the critique stage runs on a different model from the one that did the work. Take that apart and the whole loop degrades into a machine that agrees with itself, confidently, forever.</p><p>Andrej Karpathy ran the same shape on a single box. A <a href="https://github.com/karpathy/autoresearch">630-line training script</a> left running unattended, roughly seven hundred changes explored on its own across two days, about twenty of which stuck. I wrote about what decides whether a loop like that pays off or burns you while you sleep in <a href="https://rundatarun.io/p/the-loop-is-simpler-than-it-sounds">The Loop Is Simpler Than It Sounds</a>. The model is the commodity. The readable loop around it is the asset.</p><div><hr></div><h2>The delay was never the science</h2><p>This pattern is older than ARIA.</p><p>The TCGA consortium spent five or six years sequencing and assembling human mutation data. <strong>We reprocessed the entire corpus in about 22 days</strong>, over the run-up to Christmas, with a novel algorithm. It surfaced roughly <strong>17% novel variants</strong> that mapped straight back into a live drug portfolio.</p><p>Five years to a month is not a science result. It's an arithmetic result about everything that isn't science.</p><blockquote><p><strong>The delay is never the science. It's the machinery around it.</strong></p></blockquote><p>And none of that is a biology story. Jack Hanlon, who leads GenAI Media at Meta, said it sharper than I ever have: "The first time you see an engineer build something in 45 minutes that would have taken a week a year ago, but then see it not ship for another 6 weeks, you will be radicalized." Meta is a consumer company. Put the same build inside a patient-facing one and you add privacy impact assessment, security review, AI governance review, architecture review, procurement, legal redline and regulatory sign-off, each gate serial, each gate three weeks at the floor. Six of them in a row is half a year spent on a thing that was finished in the spring.</p><div><hr></div><h2>You can't read your way across</h2><p>This was the part Domino pulled for their own feed, and they cut it in the wrong place.</p><p>Most technical leaders are still on the copilot side. Prompt, review, approve, repeat. Plenty of them believe they've moved past that, and are running agents like a chatbot with a longer wait. The gap between those two isn't tooling. It's whether anyone has personally operated the thing.</p><p>Senior sponsorship exists. Technical enthusiasm exists. <strong>The person in the middle who has done it with their own hands usually doesn't.</strong></p><p>The quote Domino ran was <strong>"the crossing is personal because somebody's got to do it."</strong> True, and it's half a sentence. The half they cut is the half a leader needs:</p><blockquote><p><strong>The crossing is personal, because somebody's got to do it, and then somebody's got to teach somebody else how to do it.</strong></p></blockquote><p>That's how we did biology. It's how we did everything else. Nobody has explained why this one is different.</p><p>You can't read your way across, and you can't buy your way across by outsourcing it. Both routes feel like progress and neither produces anyone who can teach the next person.</p><div><hr></div><h2>The question nobody's asking</h2><p>Thomas closed by asking what nobody is asking, and this is the one I'd want a team to take away.</p><p><strong>AI has to eat the process around AI.</strong> I called it an ouroboros on the call, the snake that consumes itself, and I'll stand behind the image. We have pointed AI at the typing with real enthusiasm. We have not pointed it at the approvals, the reviews, the intake queues, or the committee that decides whether a model may be used at all.</p><p>In a regulated industry that process exists to protect patients, and it isn't a villain. But <strong>holding a technology you cannot deploy harms patients too</strong>, and that harm doesn't show up in anyone's risk register. A process that blocks the thing which would have helped them has stopped protecting patients and started protecting itself. Both of those are true at once, and pretending only the first one is true is how this argument usually gets lost.</p><p>So, concretely, four things:</p><p><strong>Name the queue that's actually binding you.</strong> Not the model, not the budget. The queue. Almost nobody can name theirs, which is itself the finding.</p><p><strong>Point the agents at the queue, not only at the code.</strong> The review packet, the evidence pack, the governance artefacts, the intake triage. That work is textual, repetitive, and enormously expensive in human weeks. It is exactly the shape of thing these systems are good at, and almost nobody is aiming there.</p><p><strong>Stop buying point solutions and calling it a strategy.</strong> I said on the call that "we'll just buy Claude Cowork" is a band-aid, and the same goes for whatever the next one is called. Rebuilding how a company runs was never something you could buy.</p><p><strong>Do the crossing yourself, then teach exactly one person.</strong> That's the whole adoption mechanism. It doesn't scale another way, and every org chart that assumes otherwise is buying software instead of capability.</p><p>The thing I'm starting on next is whether an entire company can run on agents. Finance, engineering, all of it. I'm sketching it now, and Thomas asked me back to report on how it went, which converts an idea into a deadline.</p><p>Which is where the conductor question finally resolves, and the answer is neither promotion nor demotion. The instrument got much bigger and the job moved. My own arc has gone from heavy steering of these systems to <strong>trying to stay out of their way so I'm not slowing them down</strong>, and I mean that as a description of work rather than of leisure. Thomas put it better than I did: it's like being on a boat, waking up to see where it went overnight, reading the weather, and setting the line.</p><p>Somebody still has to decide which of it is the take. That part hasn't moved at all.</p><p>I get to spend my mornings reading what a machine decided to try while I was asleep and then choosing what it means. It's awesome, and I've stopped looking for a graver word for it.</p><blockquote><p><strong>We made AI eat the keyboard. Now it has to eat the queue.</strong></p></blockquote><p>That last line is a post I already wrote, <a href="https://rundatarun.io/p/ai-ate-the-keyboard-now-it-has-to">AI Ate the Keyboard. Now It Has to Eat the Queue</a>, and I'm reusing it on purpose. It's the same argument, and I hadn't yet worked out that the queue is where the conductor actually stands.</p><div><hr></div><h2>Watch it</h2><p><a href="https://www.youtube.com/watch?v=kuTZIDSpUXk">Episode 100 of *Data Science Leaders*</a>, thirty minutes. Thomas gets more out of me on the biology-to-code route than I've written down anywhere, and there's a stretch on how ARIA handles time that isn't in this piece at all, because it deserves its own.</p><p>Thanks to Thomas Been and the Domino team, who ran the best-prepared interview I've done and then found a better title for it than I had.</p><div><hr></div><h2>The sources</h2><ul><li><p><a href="https://www.youtube.com/watch?v=kuTZIDSpUXk">*Data Science Leaders* episode 100, "Becoming the Builder-Conductor"</a> - Domino Data Lab, released 2026-08-21, 30:16.</p></li><li><p><a href="https://doi.org/10.5281/zenodo.21383332">ARIA: Sustained Autonomous Research Agents in Biomedicine</a> - Johnson and Bedworth, 2026. Zenodo. The 18-week deployment, 19,364 commits, seven instances, four domains.</p></li><li><p><a href="https://rundatarun.io/p/the-harness-is-the-moat">The Harness Is the Moat</a> - the argument I brought to the conversation.</p></li><li><p><a href="https://rundatarun.io/p/start-with-claude-code">Start With Claude Code</a> - the post that produced the invitation.</p></li><li><p><a href="https://rundatarun.io/p/she-already-built-it">She Already Built It</a> - the federated learning discovery, written up at the time.</p></li><li><p><a href="https://rundatarun.io/p/inside-aria-teaching-a-machine-to">Inside ARIA: Teaching a Machine to Think Like a Scientist</a> - how the loop is built.</p></li><li><p><a href="https://rundatarun.io/p/the-loop-is-simpler-than-it-sounds">The Loop Is Simpler Than It Sounds</a> - the overnight loop pattern, and the three questions that decide whether it pays off.</p></li><li><p><a href="https://rundatarun.io/p/the-year-nature-caught-up">The Year Nature Caught Up</a> - the publishing-lag argument behind the TCGA section.</p></li><li><p><a href="https://rundatarun.io/p/ai-ate-the-keyboard-now-it-has-to">AI Ate the Keyboard. Now It Has to Eat the Queue</a> - the process argument, and the closing line.</p></li><li><p><a href="https://rundatarun.io/p/three-harnesses-three-characters">Three Harnesses, Three Characters, One Working Week</a> - what running several of these at once looks like.</p></li><li><p><a href="https://builder-leader.com">*Builder Leader*</a> - the book the argument in this piece comes out of.</p></li></ul><div><hr></div><p><em>Run Data Run is free, no paywall. If it was useful, the easiest way to support it is to subscribe and forward it to one person on your team who'd want it. If it wasn't, tell me why. I'll make it better.</em></p>]]></content:encoded></item><item><title><![CDATA[Twenty-five years and half a mile]]></title><description><![CDATA[The morning before my first day I went for a run through Chicago, and a couple of miles in I recognised the route.]]></description><link>https://rundatarun.io/p/twenty-five-years-and-half-a-mile</link><guid isPermaLink="false">https://rundatarun.io/p/twenty-five-years-and-half-a-mile</guid><dc:creator><![CDATA[Justin Johnson]]></dc:creator><pubDate>Tue, 11 Aug 2026 12:31:54 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/763395e1-aaec-426f-a238-377e4ddbbee8_1920x1071.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>The morning before my first day at Tempus I went for a run through Chicago.</p><p>A couple of miles in, I recognised the route. Not the way you recognise a place in a photograph. The way your legs recognise a hill before you do. I had run this before.</p><p>I had spent the previous six weeks wrapping up thirteen years at AstraZeneca and being with my family, and in all that time I never once looked up where in Chicago the office actually was. So I ended up standing on a sidewalk at seven in the morning doing arithmetic I had not planned on. It came out at twenty-five years and half a mile.</p><p>That is the distance from the web applications company where I interned in college to the building I walked into a few hours later as a VP.</p><p>I got the goosebumps. Then I finished the run.</p><h3>The under-construction sign</h3><p>I was working at a web company as a biology major, which takes some explaining.</p><p>I had been building websites since junior high. This was the animated GIF era, the era of the under-construction sign and the hit counter you quietly started at a thousand. GeoCities was where I lived. Nobody taught me any of it, because there was nobody to do the teaching. The web was about four years old and the people building it were figuring it out in public, badly, at the same time as me.</p><p>That is the part worth holding onto, because it is the part that repeats. I was not precocious. I was early. Those are different things, and only one of them is available to you.</p><h3>Four times</h3><p>I have now done the same thing four times.</p><p>The web in the mid-nineties. Genomics in the early two-thousands. Clinical sequencing around 2010. Agentic AI from about 2023. Each time I arrived while the field was still arguing about what it was, and each time I got in without the credential the people around me had.</p><p>That is not a claim about talent. It is a claim about timing.</p><p>I was not precocious. I was early. Those are different things.</p><h3>Genomics, before there was software</h3><p>I joined The Institute for Genomic Research in 2002. It later became the J. Craig Venter Institute, and I was there eight years.</p><p>The human genome had a draft and no finish. The word metagenomics had not reached most biology departments. What that meant day to day is that the software to assemble and annotate a genome at that scale did not exist as a thing you could buy, so you wrote it, and the version you wrote was wrong until enough real data had run through it to show you where.</p><p>Nobody at TIGR ever asked what my degree was in. They asked whether the pipeline ran.</p><p>In 2010 the group built a synthetic <em>Mycoplasma</em> chromosome, transplanted it into a recipient cell, and the cell ran on it. They wrote four watermark sequences into the DNA, using a lexicon that maps DNA triplets to letters of the alphabet. One of those watermarks spells out the names of forty-six people who worked on it.</p><p>My initials are in there. Somewhere there is an organism carrying my name inside its chromosome, which is the closest thing to tenure I am ever going to get.</p><p>I helped computationally assemble the first human genome, the first large-scale metagenome, and the first synthetic genome ever made. I mention all three together because of what they have in common, which is that none of them had a method until somebody wrote one down afterwards.</p><h3>The startup, and the boring half</h3><p>I left for EdgeBio in 2010 to build the sequencing and bioinformatic interpretation platform for a CLIA lab. Clinical-grade genomics was maybe three years old as a commercial proposition. There was no reference architecture, so we built one, and most of what we built was wrong in ways we only discovered by running it against real samples.</p><p>GeneDx acquired the services business in 2013.</p><p>Around the same stretch I was running a small web and digital marketing company of my own, Inspiring Design, alongside the day job. It taught me the parts of building a company that nobody finds interesting until the moment they need them. How to form an entity. How to write a contract that survives contact with a disagreement. How to invoice, and what to do when the invoice does not get paid.</p><p>I have used that more often than I have used anything I learned in a lecture hall. If you are technical and you have never had to be the person who signs something, go and be that person once. It rewires how you read every proposal you see afterwards.</p><h3>The fourth window was not a young field</h3><p>Then thirteen years at AstraZeneca. I joined to work on genomics infrastructure and left running the AI Centre of Excellence for R&amp;D.</p><p>The fourth window looked different from the first three. AI was not a young field by 2023. The papers existed, the tooling existed, and plenty of people understood it better than I did. What was young was the fit between that field and a hundred-and-fifty-year-old pharmaceutical company, and that gap turns out to open the same kind of opportunity as an empty field does.</p><p>So the work was FAIR data foundations, a prioritisation layer, agentic systems that did real research tasks, and a governance pathway that took model provisioning from weeks to hours. That last one mattered most and got the least attention, which is normal. I ended up repeating a line often enough that people started saying it back to me: safe is the fast way.</p><p>None of those four platforms was assigned to me. I wrote the first version of each one myself, proved it worked, recruited people better than me, and handed it over. That is the only method I have.</p><h3>What the degree cost</h3><p>I have a bachelor's degree in biology. That is the whole credential.</p><p>Almost everyone holding a job like the one I have now holds a doctorate. This is the part of the essay where I am supposed to tell you it did not matter, and that would be a lie, so here is the honest version.</p><p>It cost me the benefit of the doubt. In a scientific room, a person with no letters after their name starts a notch below the line and has to climb over it by being useful. Every time. In every new room.</p><p>It cost me speed early on, because there are things a good advisor hands you in a year that took me four to learn by walking into them. And there are rooms where the credential is the entry condition and no amount of shipped work substitutes for it. I have been outside a few of those.</p><p>What it did not cost me was the work. Not once in twenty-five years has anyone stopped me building something because of what my degree said. The gate is on the room, not on the workbench.</p><h3>When the window is open</h3><p>Here is what those four moments have in common.</p><p>A credential certifies that you have mastered an existing body of knowledge. That is a genuinely useful thing for it to do, and I would want a certified surgeon. But when the body of knowledge does not exist yet, there is nothing available to certify, and the only evidence anybody can offer is a thing that works.</p><p>That is the window. It opens when a field is young enough that nobody has written the textbook, and it closes when somebody does. While it is open, building is the credential. After it closes, the credential is the credential, and getting in the way you got in before stops working.</p><p>The practical version, for anyone deciding whether to wait for permission: the window is not announced and it does not stay open long. In the web it was maybe six years. In clinical genomics, four.</p><p>You find out it closed when the job postings start asking for five years of experience in a thing that has existed for five years.</p><h3>The part where I am supposed to talk about hype</h3><p>I have now watched three technologies get badly oversold from the inside.</p><p>The web was going to replace retail by 2001. Genomics was going to give us personalised medicine within a decade of the genome being published. AI is currently going to do most of what you have read this year that it is going to do.</p><p>All three were oversold on timing. All three were roughly right on direction. The mistake people make is treating the collapse of the hype as a verdict on the technology, when the hype collapsing is just the money leaving before the infrastructure is finished.</p><p>For AI right now, the honest answer is that some of it will leave residue and some will not, and I do not think anyone can reliably sort them from inside the cycle. What I do think is that the residue is where the products are, and that finding out which is which is a measurement problem rather than an opinion problem.</p><p>Which is a reasonable description of the job I just took.</p><h3>The address changed</h3><p>Tempus is the first place I have worked where the data is not the constraint. Molecular, imaging and clinical outcomes on the same patients, at a size a model can learn something from. I have spent twenty-five years running into the absence of that, and I would like to find out what is on the other side of it.</p><p>People keep telling me this is a full-circle story. I do not think it is. A circle means you came back to where you started, and I never left. I have had one job for twenty-five years, which is to turn up somewhere the rules have not been written yet and start writing them badly until they get better.</p><p>The address changed. And that's about it.</p><div><hr></div><p>I am still <a href="https://rundatarun.io/p/im-justin-johnson-i-build-things">building</a>, and still <a href="https://builder-leader.com">writing about it</a>.</p>]]></content:encoded></item><item><title><![CDATA[Take the Cast Off]]></title><description><![CDATA[Anthropic deleted more than 80% of Claude Code's system prompt and lost nothing they could measure. I ran the same test on my own instructions. I now know the answer for one file out of ninety-three.]]></description><link>https://rundatarun.io/p/take-the-cast-off</link><guid isPermaLink="false">https://rundatarun.io/p/take-the-cast-off</guid><dc:creator><![CDATA[Justin Johnson]]></dc:creator><pubDate>Sun, 09 Aug 2026 14:26:56 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/e6bf06fc-2980-42fa-886f-599e5ee6d333_1920x1071.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!X4_8!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F32d47ae4-a764-4842-ac9f-65eb3c0ed87e_1920x1071.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!X4_8!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F32d47ae4-a764-4842-ac9f-65eb3c0ed87e_1920x1071.jpeg 424w, https://substackcdn.com/image/fetch/$s_!X4_8!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F32d47ae4-a764-4842-ac9f-65eb3c0ed87e_1920x1071.jpeg 848w, https://substackcdn.com/image/fetch/$s_!X4_8!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F32d47ae4-a764-4842-ac9f-65eb3c0ed87e_1920x1071.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!X4_8!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F32d47ae4-a764-4842-ac9f-65eb3c0ed87e_1920x1071.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!X4_8!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F32d47ae4-a764-4842-ac9f-65eb3c0ed87e_1920x1071.jpeg" width="728" height="409.5" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/32d47ae4-a764-4842-ac9f-65eb3c0ed87e_1920x1071.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:false,&quot;imageSize&quot;:&quot;normal&quot;,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:728,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;captionedImage&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!X4_8!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F32d47ae4-a764-4842-ac9f-65eb3c0ed87e_1920x1071.jpeg 424w, https://substackcdn.com/image/fetch/$s_!X4_8!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F32d47ae4-a764-4842-ac9f-65eb3c0ed87e_1920x1071.jpeg 848w, https://substackcdn.com/image/fetch/$s_!X4_8!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F32d47ae4-a764-4842-ac9f-65eb3c0ed87e_1920x1071.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!X4_8!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F32d47ae4-a764-4842-ac9f-65eb3c0ed87e_1920x1071.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>Every Sunday I pick one paper or release that's worth your time, break it apart, and tell you why it matters. No hype. No summaries of summaries. Just the idea, explained.</em></p><div><hr></div><p>In late July, Anthropic <a href="https://claude.com/blog/the-new-rules-of-context-engineering-for-claude-5-generation-models">deleted more than 80%</a> of the system prompt that runs Claude Code.</p><p>That prompt is the standing instruction set the tool carries into every session, built up over a year of watching the model get things wrong. Four-fifths of it came out. Their coding evaluations showed no measurable loss.</p><p>The engineer who wrote it up said those constraints "were once needed to avoid worst-case scenarios," and that they could now "delete many of them and let the model use surrounding context and judgment instead." A rule about code comments that had grown into three dense clauses came back as fifteen words: write code that reads like the surrounding code, match its comment density, naming, and idiom.</p><p><strong>They were not tidying up.</strong> They had worked out they were holding their own model back, through the system prompt and through the instruction files and skills they had written around it.</p><p>Everyone downstream has been writing those files too.</p><div><hr></div><h2>The cast</h2><p>You put a cast on a broken limb. It is the right call. The joint is immobilized, the bone knits, and while the bone is broken nothing else does that job.</p><p>Leave it on past healing and the muscle underneath wastes. The limb that comes out is thinner than the one beside it. <strong>The cast never stopped doing what it does.</strong> What changed is that the thing it was protecting no longer needed protecting, so all that is left is the constraint.</p><p>Every instruction you have written for an AI system is a cast. You wrote it because the model got something wrong, and you were right. The model underneath has been replaced two or three times since, and nothing in your setup tells you which bones have healed.</p><blockquote><p><strong>The instruction did not stop working when the model improved. It became the thing holding it back.</strong></p></blockquote><p>None of that makes scaffolding bad. It makes it indistinguishable. Some of what you wrote is capability the model still lacks, and some is a monument to a weakness it grew out of, and on the page the two look identical: both helped on the day you wrote them, both read like good instructions, and only one of them still earns its place. <strong>Reading the file will not tell you which one you are holding.</strong></p><div><hr></div><h2>What the vendor is telling you to do</h2><p>Boris Cherny built Claude Code and now runs product for it at Anthropic. <a href="https://www.ycombinator.com/library/UN-boris-cherny-building-claude-code">At Y Combinator</a>, asked what ordinary users should do, he gave the shortest possible answer:</p><blockquote><p>"for people that aren't building agentic products but you're using Claude Code, every six months delete your Claude.md. Delete your skills."</p></blockquote><p>He is not guessing. Anthropic keeps an internal switch that strips every prompt out of the tool, and they flip it on purpose to check whether the prompt is helping:</p><blockquote><p>"we actually use this as a sort of ablation to figure out: is the prompt useful? And what's interesting is that the model is actually a little bit more intelligent without these prompts."</p></blockquote><p>Then, a few sentences later, he says the opposite:</p><blockquote><p>"when you use Claude Code as a product, you do actually want some of these prompts because it helps you use the product."</p></blockquote><p><strong>Both of those are true, and reading them together is what makes the trade visible.</strong> The instructions make the model slightly less capable and the product considerably more predictable. Predictable is usually what you wanted. So the choice was never scaffolding against no scaffolding. <strong>You are spending capability to buy consistency, and almost nobody knows the exchange rate they are getting.</strong></p><p>His rebuild rule is the practical half. Delete the file. Work normally. Add an instruction back only after you have watched the model make the same mistake twice. <strong>Guessing in advance is how the file got big in the first place</strong>, and every line you guess wrong at gets read on every request for as long as it sits there.</p><div><hr></div><h2>So I tested mine</h2><p>I have 93 of these files. In May I had 41, after a deliberate cull with a written report justifying every cut. It grew back inside three months. It always grows back, because adding one is somebody's job and deleting one is nobody's.</p><p>Not one of them had ever been measured. I had usage counts and instincts.</p><p>The test itself is dull. Take one file. Run the same real tasks three ways: the assistant with no instructions at all, the assistant with my current file, the assistant with a rewritten candidate. Then have a separate AI read two answers side by side, blind to which came from where, and pick the better one against a written standard.</p><p>Before any of its verdicts counted, I handed the judge a rigged pair. One answer was good. The other was fluent, well formatted, and built on two invented citations. It had to catch them, and it named both. <strong>A judge that cannot fail will certify whatever you point it at</strong>, so the scoring does not start until it has proved it can flunk something.</p><p>I picked the file I use most: the one that tells the assistant how to read and review an academic paper. Three matchups, twenty blind comparisons.</p><p><strong>The rewritten version won all five of its comparisons against no instructions.</strong> Against my current file it won eight of ten.</p><p>Then the matchup I had not expected to care about. <strong>My current file against nothing at all: three wins, two losses.</strong></p><p>Five comparisons is thin and I am not going to dress it up as more. But three to two is not the score of a file I would have defended, and I would have defended that one hard. Most of what I thought it was adding, the model already had.</p><blockquote><p><strong>The instruction set I use most beat deleting it by three to two.</strong></p></blockquote><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Ur40!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F904b0748-3ccd-4c22-9a21-72313868349a_1920x1071.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Ur40!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F904b0748-3ccd-4c22-9a21-72313868349a_1920x1071.jpeg 424w, https://substackcdn.com/image/fetch/$s_!Ur40!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F904b0748-3ccd-4c22-9a21-72313868349a_1920x1071.jpeg 848w, https://substackcdn.com/image/fetch/$s_!Ur40!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F904b0748-3ccd-4c22-9a21-72313868349a_1920x1071.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!Ur40!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F904b0748-3ccd-4c22-9a21-72313868349a_1920x1071.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Ur40!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F904b0748-3ccd-4c22-9a21-72313868349a_1920x1071.jpeg" width="728" height="409.5" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/904b0748-3ccd-4c22-9a21-72313868349a_1920x1071.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:false,&quot;imageSize&quot;:&quot;normal&quot;,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:728,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;captionedImage&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!Ur40!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F904b0748-3ccd-4c22-9a21-72313868349a_1920x1071.jpeg 424w, https://substackcdn.com/image/fetch/$s_!Ur40!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F904b0748-3ccd-4c22-9a21-72313868349a_1920x1071.jpeg 848w, https://substackcdn.com/image/fetch/$s_!Ur40!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F904b0748-3ccd-4c22-9a21-72313868349a_1920x1071.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!Ur40!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F904b0748-3ccd-4c22-9a21-72313868349a_1920x1071.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div><hr></div><h2>What had actually rotted</h2><p>I kept the file. What was wrong with it was specific, and it was not the failure I would have guessed at.</p><p>Academic papers usually exist twice, as a free preprint and as the published version of record, and the numbers in them differ. <strong>My old instructions produced a review that took three figures from the preprint, reported them as the journal's, and then signed off by certifying that every number in the document traced to the published article.</strong> Wrong on exactly the sentence a reader repeats out loud in a meeting, and confident about it.</p><p>The rewrite makes every claim record which document it came from, and bans blanket confidence statements outright. It went in this week, after I read the diff.</p><p><strong>The file had not gone useless. It had decayed in one specific place</strong>, and nothing I was tracking would have surfaced it. Usage was high. Overlap with other files was zero. It read well.</p><div><hr></div><h2>The dashboard will not tell you</h2><p>The obvious question is whether you can skip the testing and read this off your metrics instead. You cannot. <strong>Every cheap signal I have used to judge these files has already misled me</strong>, and each failure is written down with a date.</p><p>I removed a browser tool after a scan showed it unused. The scan window turned out to be 0.2 days. Measured properly afterwards: 3,030 calls across 159 sessions, 19 of them on the morning I removed it. Restored the same day.</p><p>Four of my instruction files show zero use while the capability behind them runs constantly. One shows zero uses against 770 calls to the same underlying tool by a different route. <strong>The instruction was never the path anything took.</strong> A zero can mean nobody wants this, or it can mean something else already won, and on a dashboard those are the same number.</p><p>In July one component measured 22,837 tokens of overhead. In August the same component measured 863. Nothing about it had changed. A setting governing when things load was misconfigured, so every earlier measurement had been reading the setting rather than the component. A separate cut I projected would save 5,665 tokens measured 111.</p><blockquote><p><strong>A per-component cost is a property of your configuration, not of the component.</strong></p></blockquote><div><hr></div><h2>Cutting is not free either</h2><p>In April, Anthropic <a href="https://www.anthropic.com/engineering/april-23-postmortem">added one instruction</a> to Claude Code's system prompt telling it to keep responses brief. One of their evaluations showed a 3% drop, on both Opus 4.6 and 4.7. Added on the sixteenth, reverted on the twentieth. <strong>A single sentence of standing guidance, and the damage ran the opposite way to anyone's intuition</strong>, because the instruction that hurt was an instruction to be concise.</p><p>There is also no threshold to aim for. <a href="https://arxiv.org/abs/2507.11538">The work measuring how models degrade</a> as instruction counts rise finds steady decay across 500 instructions, in three different shapes depending on the model, and names no cliff in any of them. <strong>Prune by whether a thing still earns its place, never by a budget.</strong></p><div><hr></div><h2>Which three you measure this quarter</h2><p>A finer-grained experiment exists and I could not afford it: instead of testing a whole file against nothing, cut it into layers and measure what each layer adds on its own. <strong>Showing that a layer helps takes six comparisons. Showing that one does nothing takes twenty three</strong>, because proving no difference is a much stronger claim, and four layers at that rate runs about four times what one night already cost.</p><p>So measuring everything was never the plan. One file took about sixty agents and most of a week's compute. <strong>The plan is picking which three you measure this quarter</strong>, and three properties decide it. Usage is not one of them.</p><p><strong>The first is reliance.</strong> Not how often a file fires, but how far you would trust its output without checking. A file whose answers land in a slide nobody re-derives deserves more scrutiny than one that fires forty times a day into work you read line by line anyway.</p><p><strong>The second is age measured in model generations, not months.</strong> An instruction written for a model two releases back was written against failure modes that may no longer exist. Mine was young on the calendar and two generations old.</p><p><strong>The third is whether the file makes factual claims that can rot without showing a symptom.</strong> Mine did. A style instruction that goes stale produces prose you dislike, and you notice inside a paragraph. A factual instruction that goes stale produces confident wrong numbers in a register that reads exactly like the correct ones, and nothing in the output tells you which one you are looking at.</p><div><hr></div><h2>The job nobody owns</h2><p>Underneath the technical problem is an organizational one, and it decides whether any of this happens more than once. It also explains how my own count went from 41 back to 93 in three months without anyone deciding that it should.</p><p><strong>Adding an instruction is somebody's job.</strong> It happens the moment something breaks, it takes ten minutes, and the person who does it is visibly fixing a bug. <strong>Deleting one is nobody's job.</strong> It means proving a negative, it produces no improvement anyone can see on the day, and the person who does it is one incident away from being the person who removed the guardrail. The incentives are not close, and they point the same way in every organization I have watched, including my own.</p><p>The teams that get this right are not the ones with the best tooling. They are the ones who have made four unglamorous things normal.</p><p><strong>Every instruction carries the failure it was written for, and the date.</strong> One line. An instruction with no recorded reason can only be reviewed by the person who wrote it, and in eighteen months that person has changed teams or forgotten. Write down what broke and you have handed the next reviewer a test they can run: does it still break.</p><p><strong>Removal has a named owner and explicit cover.</strong> Somebody's job description includes taking things out, and leadership has said out loud, before anything goes wrong, that a removal which turns out badly is a normal cost of the review rather than a mistake with a name attached. Without that, the review meeting happens and nothing leaves the room.</p><p><strong>The review runs on model releases, not on quarters.</strong> Every release re-dates every instruction written before it. Tie the review to the calendar and you will audit a stable setup in a quiet month and miss the one that moved under you in a loud one.</p><p><strong>The output of a review is a number, not an opinion.</strong> Two versions, real tasks, and a reader who does not know which is which. It does not take sixty agents. It takes one person blind to the answer, which is the step that gets skipped, because the author of a file is the worst available judge of it and is usually the only person in the room.</p><blockquote><p><strong>Adding an instruction is somebody's job. Deleting one is nobody's job. That asymmetry, not the model, is why your setup grows.</strong></p></blockquote><div><hr></div><h2>Ninety-two</h2><p>The three-to-two is not the number that bothers me.</p><p>I own 93 instruction files. I measured one of them, on one night, and it cost about sixty AI agents and most of the compute I would normally spend in a week. <strong>The other 92 are exactly as unexamined as they were on Monday</strong>, and several of them are older than two model generations.</p><p>That ratio is going to get worse. Models ship faster than anyone's scaffolding gets reviewed, every release re-dates every instruction written before it, and there is no alert for any of it, because nothing degrades visibly and the file keeps loading and the work keeps coming out fine while the gap between what your instructions do and what they were written to do widens on a schedule you do not control.</p><p>So pick your most-relied-on instruction, the one you would defend hardest, and run a week without it. Have someone who does not know which is which read the output both ways.</p><p>You will get one of three answers, and two of them are cheap. Either something better exists, or the model outgrew it, or what you have is fine and you can stop wondering. I went in expecting the first. I got it, and I also found out my old file was barely beating the fourth option nobody plans for, which is having written nothing at all.</p><div><hr></div><h2>Sources</h2><ul><li><p><a href="https://claude.com/blog/the-new-rules-of-context-engineering-for-claude-5-generation-models">The new rules of context engineering for Claude 5 generation models</a>: Thariq Shihipar, Anthropic, 24 July 2026. The 80% deletion, the "no measurable loss" claim, and the then/now table of practices Anthropic now calls myths.</p></li><li><p><a href="https://www.ycombinator.com/library/UN-boris-cherny-building-claude-code">Boris Cherny: Building Claude Code</a>: Y Combinator Startup School. The delete-every-six-months advice, the ablation switch, and the admission that the model is "a little bit more intelligent without these prompts."</p></li><li><p><a href="https://www.anthropic.com/engineering/april-23-postmortem">Anthropic's April 23 postmortem</a>: the brevity instruction that cost 3% on one evaluation, added 16 April and reverted 20 April.</p></li><li><p><a href="https://arxiv.org/abs/2507.11538">How Many Instructions Can LLMs Follow at Once?</a>: Jaroslawicz et al. The IFScale benchmark, 500 instructions across 20 models, and the finding that there is no threshold to prune toward.</p></li></ul><div><hr></div><p><em>Sunday Deep Dive is a weekly series on Run Data Run. Every Sunday I pick one paper, release, or technique worth understanding, break it apart, and tell you what it means for your work. Free every Sunday, no paywall. If it was useful, the easiest way to support it is to subscribe and forward it to one person on your team who'd want it. If it wasn't, tell me why. I'll make it better.</em></p>]]></content:encoded></item><item><title><![CDATA[I moved my whole AI coding setup to a model that costs 40 cents. Nobody noticed the difference.]]></title><description><![CDATA[A two-year-old harness, ported in an afternoon, for under forty cents. Then the model that did it wrote this post.]]></description><link>https://rundatarun.io/p/i-moved-my-whole-ai-coding-setup</link><guid isPermaLink="false">https://rundatarun.io/p/i-moved-my-whole-ai-coding-setup</guid><dc:creator><![CDATA[Justin Johnson]]></dc:creator><pubDate>Tue, 04 Aug 2026 10:02:38 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/e86f7379-0082-43a4-9e91-ce9a588cdd96_1664x928.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!2MNy!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0ebe4370-9e1f-4329-9e11-298268227c6b_1664x928.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!2MNy!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0ebe4370-9e1f-4329-9e11-298268227c6b_1664x928.jpeg 424w, https://substackcdn.com/image/fetch/$s_!2MNy!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0ebe4370-9e1f-4329-9e11-298268227c6b_1664x928.jpeg 848w, https://substackcdn.com/image/fetch/$s_!2MNy!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0ebe4370-9e1f-4329-9e11-298268227c6b_1664x928.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!2MNy!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0ebe4370-9e1f-4329-9e11-298268227c6b_1664x928.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!2MNy!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0ebe4370-9e1f-4329-9e11-298268227c6b_1664x928.jpeg" width="728" height="409.5" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/0ebe4370-9e1f-4329-9e11-298268227c6b_1664x928.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:false,&quot;imageSize&quot;:&quot;normal&quot;,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:728,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;captionedImage&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!2MNy!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0ebe4370-9e1f-4329-9e11-298268227c6b_1664x928.jpeg 424w, https://substackcdn.com/image/fetch/$s_!2MNy!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0ebe4370-9e1f-4329-9e11-298268227c6b_1664x928.jpeg 848w, https://substackcdn.com/image/fetch/$s_!2MNy!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0ebe4370-9e1f-4329-9e11-298268227c6b_1664x928.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!2MNy!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0ebe4370-9e1f-4329-9e11-298268227c6b_1664x928.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>My AI coding setup is two years old. It is not a model. It is a system: rules for how the assistant should behave, skills it knows how to use, specialized agents for different kinds of work, guardrails that stop it deleting things it should not. I built it up the way you build anything that carries real weight, one piece at a time, and I assumed it was welded to the expensive model it ran on.</p><p>Over the weekend I moved the whole thing onto a model that costs 40 cents. Every rule, every skill, every agent, and most of the guardrails. The port took an afternoon. It cost under forty cents, measured on the model's own usage dashboard. It was fast enough that nothing sat waiting. And for the work I actually do, I could not tell the difference.</p><p>You are reading the result. This post was drafted on that 40-cent model, running my rules, my skills, and my voice checks.</p><blockquote><p><strong>Under 40 cents. Nobody noticed the difference.</strong></p></blockquote><div><hr></div><h2>What a two-year harness actually is</h2><p>Vendors sell you "the agent" as if it were a single thing, inseparable from the model underneath. Here is the model that thinks, here is the tool it lives in, they come as a pair. That framing is comfortable because it is simple. It is also wrong.</p><p>What I actually depend on is the layer above the model. A set of written rules that encode judgment: what to let the assistant do on its own, what to check before believing an answer, how to recover when something breaks. A library of skills, each one a procedure the assistant can follow, from drafting a blog post to running a security review. A few specialized agents for jobs that need a particular temperament or toolset, which I have come to think of as <a href="https://rundatarun.io/p/eight-agents-one-fight">a squad rather than a toolbox</a>. And guardrails, the ones that refuse a destructive command and tell you to back up first. I broke the pieces down <a href="https://ai.rundatarun.io/ai-development-agents/field-guide-04-inventory-your-harness">component by component</a> if you want the inventory.</p><p><strong>Two years of that is an asset.</strong> It is the accumulated discipline of deciding, over and over, what you want an assistant to be allowed to do. The model underneath is just the executor. I had never tested whether the executor mattered as much as I assumed.</p><div><hr></div><h2>The provocation</h2><p>DeepSeek shipped an endpoint whose only purpose is to let a rival company's coding tool talk to its model. Not a community wrapper, not a shim someone reverse-engineered: an official endpoint, so a competitor's agent points at DeepSeek and works.</p><p>It is not the first to do this. Kimi and MiniMax already run the same kind of passthrough, which is what makes it a signal rather than a stunt. When three model makers independently build an on-ramp into someone else's tool, they are all betting the same way. They are betting that the user's investment sits in the harness, and that the model underneath is what swaps out. I decided to test the bet on my own setup.</p><p>I had already <a href="https://ai.rundatarun.io/ai-development-agents/codex-vs-claude-code-vs-opencode">compared these tools from inside one</a> back in May, which tells you where they differ on paper. That is a different question from what happens when you pick a working setup up and put it down somewhere else.</p><div><hr></div><h2>The number that matters</h2><p>Here is what the port actually cost.</p><p>The expensive model I was on charges $5 per million input tokens and $25 per million output. DeepSeek V4 Flash charges $0.14 and $0.28. The July build, the one Artificial Analysis actually put a score against, is cheaper still at $0.09 and $0.18. <strong>Call it 35x cheaper on input and 90x on output, and hold in mind that those are the conservative numbers.</strong></p><p>The whole migration, every piece of the harness moved over, came to under forty cents on the model's API key.</p><blockquote><p><strong>The expensive model was never the asset. The discipline was.</strong></p></blockquote><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!cspx!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa7023ed4-b1dd-467d-be54-7b5418817c8c_1376x768.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!cspx!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa7023ed4-b1dd-467d-be54-7b5418817c8c_1376x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!cspx!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa7023ed4-b1dd-467d-be54-7b5418817c8c_1376x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!cspx!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa7023ed4-b1dd-467d-be54-7b5418817c8c_1376x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!cspx!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa7023ed4-b1dd-467d-be54-7b5418817c8c_1376x768.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!cspx!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa7023ed4-b1dd-467d-be54-7b5418817c8c_1376x768.jpeg" width="728" height="409.5" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a7023ed4-b1dd-467d-be54-7b5418817c8c_1376x768.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:false,&quot;imageSize&quot;:&quot;normal&quot;,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:728,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;captionedImage&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!cspx!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa7023ed4-b1dd-467d-be54-7b5418817c8c_1376x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!cspx!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa7023ed4-b1dd-467d-be54-7b5418817c8c_1376x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!cspx!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa7023ed4-b1dd-467d-be54-7b5418817c8c_1376x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!cspx!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa7023ed4-b1dd-467d-be54-7b5418817c8c_1376x768.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>And the cheap model is not a toy. On <a href="https://artificialanalysis.ai">Artificial Analysis</a>, the independent scoring benchmark I trust, DeepSeek V4 Flash scores 69.1 on the coding index. Claude Opus 4.8 scores 74.3. Claude Opus 5, the current flagship and the one I actually came off, scores 78. So the cheap model sits about nine points behind the best model money can buy, at one-ninetieth the output price.</p><p>I am not telling you cheap models are as good as expensive ones. They are not. Measured against Opus 5 on coding, DeepSeek lands just under ninety percent. On the broader intelligence index it falls to about eighty-two. The gap is real and it is wider than the excitement suggests. For one specific kind of work it is also invisible, and the port was that kind of work.</p><div><hr></div><h2>What transferring taught me</h2><p>Four things, and none of them were what I expected.</p><p><strong>The written parts moved without a fight.</strong> Rules, skills, agents. They are all just instructions, and instructions do not care which model reads them. The expensive model was never doing the thinking my setup does. It was following thinking I had already written down.</p><p><strong>Judgment travels. Enforcement stays behind.</strong> One of my guardrails is a plain paragraph telling the assistant to back up before it deletes anything, check the backup worked, and only then delete. The cheap model read that paragraph, refused a delete it had been asked to make, and walked the sequence. My two mechanical blocks, the code that physically stops a dangerous command before it runs, never made the trip. The new tool keeps a fingerprint of your safety file and quietly ignores anything it does not recognize, so both blocks are sitting there written, tested and switched off. I would rather say that than let "every guardrail moved" stand.</p><p><strong>A whole phase of the work was unnecessary.</strong> My plan said the new tool could see two of my skills against the eighty-eight I have. So I wrote a converter and moved them all. Then I hid a uniquely named skill in each location and asked the tool what it could see. It saw both. It had been reading my original library the entire time.</p><p>Underneath that was a difference only the hidden-file test would have found. My usual tool hands the model a skill's full instructions the moment it needs them. The new one shows a one-line summary and leaves the rest on disk, read only if the model goes looking. That economy is <a href="https://ai.rundatarun.io/practical-applications/claude-skills-vs-mcp-servers">why skills work at all</a>, and the new tool pushes it further than I would have chosen. Same files, very different odds of the thing running.</p><p><strong>The dial that does nothing.</strong> The model has a low-to-high effort setting. On the connection my setup uses it is inert: I told it to think briefly, then told it to think hard, and got the same work at the same cost either way.</p><p>Then I went looking, and the public picture is a mess. <a href="https://x.com/teortaxesTex">teortaxesTex</a> read a benchmark chart as low matching high at twice the price. <a href="https://x.com/morganlinton">Morgan Linton</a>, who ran that benchmark, found close to the opposite: DeepSeek won on medium effort, not high. Same model, same week, two readings. I did not settle it. The claim I can stand behind is smaller, that on my connection the setting does nothing, and no documentation would have told me.</p><blockquote><p><strong>A control you cannot verify is a control you should not build assumptions on.</strong></p></blockquote><div><hr></div><h2>The comparison</h2><p>The cheap model is having a moment. <a href="https://x.com/opencode">opencode</a>, a coding tool, tracked 8 trillion tokens of DeepSeek Flash traffic on August 1st alone, 5 trillion of that free usage. DeepSeek's own launch numbers put V4 Flash at 82.7 on Terminal Bench 2.1, ahead of its larger sibling, and that is a vendor number, so weigh it the way you weigh vendor numbers. The independent read is <a href="https://x.com/morganlinton">Morgan Linton's</a> VulcanBench run, where it took the top spot outright and pushed two much better-known models out of the top five.</p><p>The objection that deserves airtime is cost per task rather than cost per token. <a href="https://x.com/cline">cline</a> raised it: a cheaper token still loses if the model burns more turns reaching the same answer. It then answered its own objection in the same post, citing Artificial Analysis clearing the same benchmark tasks at 105x lower cost. My experience splits along that line. For harness plumbing, moving rules and skills and agents around, the cheap model held without extra turns. For the hardest reasoning it is still a step down. The judgment call is knowing which of the two you are doing before you start.</p><p><strong>This is uncomfortable for the premium tier.</strong> At roughly ninety percent of the capability and one-ninetieth of the cost, the premium has to buy a workload you can name and point at. Not "it is the best model." A task. If you cannot produce one, you are paying for the brand.</p><div><hr></div><h2>The meta turn</h2><p>Which brings me back to where this started. This post was researched, structured and drafted on the 40-cent model. The same rules, the same skill that formats a Substack essay, the same voice checks that stop me sounding like a language model, all running on something that costs less than a vending machine snack.</p><p>Be precise about what that proves, because there are two separate tests here and only one of them wrote this. The first is the model swap: my everyday tool, my harness, a different and far cheaper model underneath. This post is the output of that one. The second is the framework port: the same harness rebuilt inside a rival vendor's tool, which is where the guardrail finding came from and which I ran on its own. Neither test alone would have told me much. Together they say the harness survived a change of model and a change of tool.</p><p>This is also not the first long thing it has produced. Earlier this year the same setup drafted <a href="https://ai.rundatarun.io/ai-development-agents/the-harness-that-wrote-the-book">a 33,000-word book in eight days</a>, <a href="https://builder-leader.com">Builder-Leader: The AI Exoskeleton That Crosses the Gap</a>, whose argument is the one I have just spent an essay stress-testing against a model that costs 40 cents. I did not set the port up as a sequel to the book. It became one, and it could have gone the other way.</p><blockquote><p><strong>You are reading this because a 40-cent model ran my harness well enough to write it.</strong></p></blockquote><p>A two-year-old pile of written instructions, pointed at a nearly-free model, then asked to hold a style, route research and stay disciplined across a full essay. It did. The model did not need to be the smartest one available. It needed the instructions to be good.</p><div><hr></div><h2>Close</h2><p>I did not downgrade my setup this weekend. I found out how much of it never depended on the expensive model at all.</p><p>The moat was never the model. It was the accumulated discipline: what to let an assistant do, what to check before believing it, how to recover when it breaks. Written down, tested, and enforced wherever the tool let me enforce it. That followed me for 40 cents and it is still doing the work.</p><p>Every setup has a number here and it costs an afternoon to find: swap the model underneath and count what still runs. Whatever survives is the asset. Whatever breaks was rented. I have argued before that you should <a href="https://rundatarun.io/p/three-harnesses-three-characters">choose your harness every six months rather than let it choose you</a>, and this is the cheapest test I have found for whether you own one.</p><blockquote><p><strong>I did not downgrade my setup. I found out how much of it never depended on the expensive model at all.</strong></p></blockquote><p>Sources: <a href="https://artificialanalysis.ai">Artificial Analysis</a> model indices for DeepSeek V4 Flash, Claude Opus 4.8 and Claude Opus 5 &#183; <a href="https://x.com/opencode/status/2083996864186318999">opencode on DeepSeek usage</a> &#183; <a href="https://x.com/rohanpaul_ai/status/2083107527664021941">rohanpaul_ai relaying DeepSeek's Terminal Bench numbers</a> &#183; <a href="https://x.com/morganlinton/status/2083984169324122541">morganlinton on VulcanBench</a> &#183; <a href="https://x.com/teortaxesTex/status/2084098542008664093">teortaxesTex on effort scaling</a> &#183; <a href="https://x.com/cline/status/2083638204037820734">cline on cost per task</a>.</p>]]></content:encoded></item><item><title><![CDATA[Somebody Is Finally Checking]]></title><description><![CDATA[A contest opened eleven days ago has produced more reproduction attempts than the field's own dedicated effort has managed in any year. Whether that becomes a check on the literature turns on a detail sitting in the scoring rules.]]></description><link>https://rundatarun.io/p/somebody-is-finally-checking</link><guid isPermaLink="false">https://rundatarun.io/p/somebody-is-finally-checking</guid><dc:creator><![CDATA[Justin Johnson]]></dc:creator><pubDate>Sun, 26 Jul 2026 04:06:38 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/e3f0aa85-2640-4ecf-9d35-bc6909acca58_1376x768.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em>Every Sunday I pick one paper or release that's worth your time, break it apart, and tell you why it matters. No hype. No summaries of summaries. Just the idea, explained.</em></p><div><hr></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!TLMS!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d0e0b21-f86a-4342-8104-343a9e655bef_1376x768.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!TLMS!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d0e0b21-f86a-4342-8104-343a9e655bef_1376x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!TLMS!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d0e0b21-f86a-4342-8104-343a9e655bef_1376x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!TLMS!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d0e0b21-f86a-4342-8104-343a9e655bef_1376x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!TLMS!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d0e0b21-f86a-4342-8104-343a9e655bef_1376x768.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!TLMS!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d0e0b21-f86a-4342-8104-343a9e655bef_1376x768.jpeg" width="728" height="409.5" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/2d0e0b21-f86a-4342-8104-343a9e655bef_1376x768.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:false,&quot;imageSize&quot;:&quot;normal&quot;,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:728,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;captionedImage&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!TLMS!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d0e0b21-f86a-4342-8104-343a9e655bef_1376x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!TLMS!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d0e0b21-f86a-4342-8104-343a9e655bef_1376x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!TLMS!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d0e0b21-f86a-4342-8104-343a9e655bef_1376x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!TLMS!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d0e0b21-f86a-4342-8104-343a9e655bef_1376x768.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>A published result is a promise. Run this, and you will see what we saw. For most of the history of computational science nobody collected on that promise, because collecting cost about as much as the original work and earned a fraction of the credit.</p><p>On July 15, Hugging Face and <a href="https://www.alphaxiv.org/">alphaXiv</a> opened a contest to collect. It closes Sunday August 2. Roughly 6,800 papers from ICML 2026 are in scope, anyone can attempt a reproduction, and every attempt gets published as a public logbook that an automated judge then scores. The prize pool is $4,500, paid in GPU credits rather than cash, which tells you who the organizers think is entering.</p><p><strong>As of this morning, eleven days in: 4,147 judged logbooks from 312 people, carrying 20,594 separate verdicts on individual claims.</strong> About 1,600 papers have been attempted at least once, roughly a quarter of the conference.</p><p>To see why that is a strange number, you need the thing it replaces. <a href="https://reproml.org/">Joelle Pineau</a> started the ML Reproducibility Challenge at ICLR 2018 as <a href="https://blog.neurips.cc/2026/05/04/mlrc-2026-reproducibility-as-an-official-track-at-neurips/">"a small community experiment."</a> Eight years on it is the field's serious effort, mostly run through graduate courses, and its NeurIPS 2019 edition drew <a href="https://jmlr.org/papers/volume22/20-303/20-303.pdf">173 papers claimed</a>, itself a <strong>92% jump on the year before</strong>. Claimed, not finished. Published reports have run in the tens. <a href="https://rescience.github.io/">ReScience C</a>, which demands a fresh independent implementation rather than a re-run of the authors' code, manages about a dozen a year.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!WwYA!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda36e809-29d8-45e4-ac68-108f3e598ccf_1376x768.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!WwYA!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda36e809-29d8-45e4-ac68-108f3e598ccf_1376x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!WwYA!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda36e809-29d8-45e4-ac68-108f3e598ccf_1376x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!WwYA!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda36e809-29d8-45e4-ac68-108f3e598ccf_1376x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!WwYA!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda36e809-29d8-45e4-ac68-108f3e598ccf_1376x768.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!WwYA!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda36e809-29d8-45e4-ac68-108f3e598ccf_1376x768.jpeg" width="728" height="409.5" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/da36e809-29d8-45e4-ac68-108f3e598ccf_1376x768.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:false,&quot;imageSize&quot;:&quot;normal&quot;,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:728,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;captionedImage&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!WwYA!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda36e809-29d8-45e4-ac68-108f3e598ccf_1376x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!WwYA!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda36e809-29d8-45e4-ac68-108f3e598ccf_1376x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!WwYA!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda36e809-29d8-45e4-ac68-108f3e598ccf_1376x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!WwYA!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda36e809-29d8-45e4-ac68-108f3e598ccf_1376x768.jpeg 1456w" sizes="100vw"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><blockquote><p><strong>173 papers claimed in a year, against 4,147 judged logbooks in eleven days. That is not an improvement on the existing effort, it is a different quantity.</strong></p></blockquote><div><hr></div><h2>What is actually being scored</h2><p>One design decision separates this from a leaderboard, and it is easy to skim past. <strong>The unit of judgment is not the paper. It is the sentence.</strong></p><p>Every paper has been broken into its individual claims, and each claim gets its own verdict: verified, falsified, <strong>toy</strong> (a scaled-down version that says so), or inconclusive. Nobody asks whether a paper "reproduced," which is a question with no answer. A paper can come back partly verified, partly toy and partly inconclusive at once, and most of them do.</p><p>Then the scoring, from the contest's own FAQ: <strong>"2 points for a full reproduction or full falsification, 1 point for a toy-scale reproduction, 0 otherwise."</strong></p><blockquote><p><strong>Falsifying a claim pays what confirming one pays. Against a literature where negative results are half as publishable as positive ones, that is a deliberate inversion.</strong></p></blockquote><p>When <em>Nature</em> surveyed 1,576 researchers in 2016, <a href="https://www.bioedonline.org/news/nature-news-archive/1500-scientists-lift-the-lid-on-reproducibility/">24% had published a successful replication and 13% a failed one</a>. The same survey found more than 70% had tried and failed to reproduce someone else's experiment and 52% agreed the field had a reproducibility crisis, and then 73% said they still trusted at least half the papers in their own field.</p><p>The failures are also the deliverable. A logbook, in the organizers' description, holds <strong>"the experiments, simplifications, failures, and results it found."</strong> Prior efforts published conclusions. This one publishes the attempt.</p><p><strong>What it does not do is test the hard version.</strong> Machines have been tried on this before and were not good at it: every benchmark that measured reproduction topped out between 16% and 27% on rebuilding a paper from scratch, with <a href="https://arxiv.org/abs/2504.01848">top ML PhDs ahead of the models</a>, and only reached about 60% <a href="https://arxiv.org/abs/2409.11363">when handed the original code and data</a>. This contest tells participants to start from the authors' code. <strong>Re-running someone's code, rebuilding from the paper alone, and getting the same finding on fresh data are three different tests, and headline percentages get quoted across that boundary constantly.</strong></p><div><hr></div><h2>What the inside looks like</h2><p>For a large class of papers the easy version is now free, which is the whole reason this works. We downloaded a full-marks logbook from the contest leader and read it. <strong>The entire reproduction ran on a laptop processor in three seconds.</strong> No GPU, no cluster, no cloud bill. The real ceiling is not compute but publishing: Hugging Face caps accounts at 20 new pages a day, and when a participant asked for relief the organizer's answer was flat, <strong>"we cannot lift the 20 spaces/day limit so I'll close this issue."</strong></p><p>I entered, which is the only reason I have anything to say about the texture of it. We are 23 judged logbooks and 174 points in, against a leader at 1,862. Not a contender, and far enough inside to see how the work divides.</p><p>It divides into two lanes, and we ran them differently on purpose.</p><p><strong>The volume lane is an autonomous agent called Vulcan, and it never publishes anything.</strong> It runs on a box in my house on free local compute, and its job is narrow: take a paper whose claims are all theory and check every formula numerically, against a brief that fixes the method rather than the answer.</p><blockquote><p><strong>A check that passes on every input you can construct certifies nothing.</strong></p></blockquote><p>Its audit of one paper came back as 721 lines of code that matched the paper's algebra to the last digit a computer can represent, plus a deliberate sanity test that failed when it was supposed to. Careful work, at no cost, with a hard limit written into the contract: <strong>it stages, a human publishes.</strong> No autonomous process of mine puts a verdict about somebody else's paper into the world.</p><p>The other lane is me and Claude Code running on Opus, for the parts the first one cannot do. Two things came out of it.</p><p><strong>The first is that picking the right paper beats checking it well.</strong> Our most careful audit matched its paper exactly and scored 2 points out of a possible 8, because half that paper's claims were hardware benchmarks that sit at inconclusive until somebody spends days of compute on them. The judge said so directly: the logbook left "half the paper's claims without any experimental evidence." A paper whose claims are all theory earns 12 points in an afternoon on a laptop. <strong>Selection is the skill, not rigor.</strong></p><p>The second is what the process catches. One paper states a theorem using a fixed setting, and then, four thousand lines later, its own proof assumes that setting changes with the length of the run. Not a typo, and not a disagreement between the paper and the world: the paper disagrees with itself, in print, peer-reviewed and accepted. Nobody had noticed because nobody had reason to run the theorem as written and watch it fail.</p><p>Beyond those two, the work was unglamorous. It caught our own staging script dropping most of a conclusion, so a 60-line write-up shipped as three sentences. Vulcan itself lost twenty hours to an outage at its model provider, with no backup configured and nothing raising an alarm.</p><p><strong>None of those failures announce themselves.</strong> Every one produced output that looked finished.</p><div><hr></div><h2>The three percent</h2><p><strong>Of 20,594 claim verdicts, 615 are falsified. That is 3.0%.</strong> The rest split 43.5% verified, 28.9% inconclusive, 24.6% toy. The distribution has barely moved for days while the corpus grew by hundreds of logbooks: the shape is stable, the counts are not.</p><p>Three percent is low, and it has two readings. Either the literature is in better condition than the surveys suggest, or falsification is the hardest verdict to reach and the easiest to get wrong.</p><p><strong>In one day, four of our audit agents returned a falsification. Three of them were wrong.</strong> All four came out of the automated lane. All three errors were caught by the other one.</p><p>One agent read a formula out of a PDF that had lost a character in scanning, then correctly proved the mangled version false. Another was checking a sentence the authors never published, because the claim it was handed came from a draft that is not the public paper. The third found a difference too small to distinguish from rounding error and called it a contradiction.</p><p>Every one produced a confident, well-formatted, coherent falsification. I wrote about that pattern in <a href="https://rundatarun.io/p/the-failure-that-leaves-no-corpse">The Failure That Leaves No Corpse</a>, and this is the cleanest instance of it I have run into.</p><p>The rule we ended up with fits on one line: <strong>does the paper disagree with itself, or does the claim disagree with your copy of the paper?</strong> Only the first is a falsification. The second is a bug in your pipeline wearing a falsification's clothes.</p><p>Applying it cost us the verdict. <strong>Across the 130 claim verdicts we have published, the number labelled falsified is zero.</strong> The one candidate that survives is the self-contradicting theorem above.</p><blockquote><p><strong>The default failure of an automated checker is not a missing answer. It is a wrong one, delivered with the same confidence as a right one.</strong></p></blockquote><p>This is not only our problem. The contest's own claim lists are extracted by a language model, and a participant found a claim in one paper's list that belonged to a different paper. The organizer confirmed it in the open, <strong>"That claim belongs to a different Safety-area paper and leaked in during auto-extraction,"</strong> removed it, rescored, then swept all 6,768 papers for the same contamination.</p><p>And the bias is measured, not hypothetical. A <a href="https://arxiv.org/abs/2606.11447">June 2026 study of reproduction agents</a> found that <strong>showing the agent the original paper made it more likely to confirm claims that could not actually be reproduced</strong>, and that small changes in wording pushed it further the same way. In this contest the agent always has the paper and the claim in front of it.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!qr23!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F182e8b8d-54dc-477a-8405-704cbf8eb398_1376x768.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!qr23!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F182e8b8d-54dc-477a-8405-704cbf8eb398_1376x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!qr23!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F182e8b8d-54dc-477a-8405-704cbf8eb398_1376x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!qr23!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F182e8b8d-54dc-477a-8405-704cbf8eb398_1376x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!qr23!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F182e8b8d-54dc-477a-8405-704cbf8eb398_1376x768.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!qr23!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F182e8b8d-54dc-477a-8405-704cbf8eb398_1376x768.jpeg" width="728" height="409.5" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/182e8b8d-54dc-477a-8405-704cbf8eb398_1376x768.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:false,&quot;imageSize&quot;:&quot;normal&quot;,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:728,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;captionedImage&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!qr23!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F182e8b8d-54dc-477a-8405-704cbf8eb398_1376x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!qr23!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F182e8b8d-54dc-477a-8405-704cbf8eb398_1376x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!qr23!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F182e8b8d-54dc-477a-8405-704cbf8eb398_1376x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!qr23!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F182e8b8d-54dc-477a-8405-704cbf8eb398_1376x768.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div><hr></div><h2>The judge is a language model</h2><p><strong>The thing scoring 20,594 claim verdicts is itself a language model.</strong></p><p>There is published criticism aimed squarely at that. A 2026 position paper, <a href="https://arxiv.org/abs/2605.03202">"Stop Automating Peer Review Without Rigorous Evaluation"</a>, compared human against AI reviews of ICLR 2026 submissions and found two things. The AI reviewers showed <strong>"a hivemind effect of excessive agreement within and across papers that reduces perspective diversity."</strong> And their scores turned out to be <strong>"trivially gameable through paper laundering"</strong>: rewriting a paper with a language model raised the scores it got, on style rather than science.</p><p>None of that is a reason to dismiss the contest, and I want to be careful here, because I am competing in it and have every incentive to be generous.</p><p>The defence is structural rather than technical. <strong>The artifacts are public.</strong> Every logbook is a page anyone can open, the verdicts are a public dataset downloaded 38,438 times last month, and the errors get argued in threads with the organizers answering. The contamination case is the proof, and the fix landed inside a day. A closed judge making the same error produces a number nobody can audit.</p><p>So the position I hold is that this design pairs a known-weak instrument with an unusually strong correction loop. Whether the loop outruns the instrument is an empirical question, and a week from now there will be enough public data to start answering it.</p><div><hr></div><h2>What it would take to trust this</h2><p>I am not going to predict how it turns out. The contest closes August 2 and every distribution above is provisional.</p><p>The concrete claim is narrower and I think it holds. <strong>For a meaningful slice of published science, checking a claim now costs less than making it.</strong> Three seconds of laptop processor against however many months produced the original theorem. That ratio is why verification stayed volunteer work for a decade, and it has flipped.</p><p>What it does not change is the direction of the error. A checker that costs nothing will be run constantly, and cheap checking produces far more false accusations than misses. Confidently accusing a sound paper is worse than missing a bad one.</p><p>The contest's design anticipates this better than most things I have seen. Paying the same for a falsification as a verification removes the incentive to go looking only for confirmations. The logbook has to carry the failures, so there is no tidy way to bury the attempts that went nowhere. And because every artifact is public, a wrong answer stays inspectable instead of disappearing into an aggregate.</p><blockquote><p><strong>Nobody has to trust the checkers. You can read what they did. That is a weaker guarantee than peer review claims to offer, and a stronger one than peer review delivers.</strong></p></blockquote><div><hr></div><h2>Sources</h2><ul><li><p><a href="https://huggingface.co/spaces/ICML-2026-agent-repro/challenge">Reproducing ICML 2026, Open Reproductions</a>: the contest, its FAQ, rules, and public discussion threads. Verdicts live in a <a href="https://huggingface.co/datasets/ICML-2026-agent-repro/verdicts">public dataset</a>.</p></li><li><p><a href="https://blog.neurips.cc/2026/05/04/mlrc-2026-reproducibility-as-an-official-track-at-neurips/">MLRC 2026: Reproducibility as an Official Track at NeurIPS</a>: the ML Reproducibility Challenge's history and Pineau's founding description of it.</p></li><li><p><a href="https://jmlr.org/papers/volume22/20-303/20-303.pdf">Improving Reproducibility in Machine Learning Research</a>, Pineau et al., JMLR 22(164): the 173-papers-claimed figure.</p></li><li><p><a href="https://rescience.github.io/">ReScience C</a>: the journal that requires a new independent implementation.</p></li><li><p><a href="https://www.bioedonline.org/news/nature-news-archive/1500-scientists-lift-the-lid-on-reproducibility/">1,500 scientists lift the lid on reproducibility</a>, Monya Baker, <em>Nature</em> 533: the 2016 survey of 1,576 researchers.</p></li><li><p><a href="https://arxiv.org/abs/2504.01848">PaperBench</a> and <a href="https://arxiv.org/abs/2409.11363">CORE-Bench</a>: the agent-reproduction benchmarks this contest inherits from.</p></li><li><p><a href="https://arxiv.org/abs/2606.11447">Reproducibility agents and confirmation bias</a>: the June 2026 benchmark showing that giving an agent the paper biases it toward confirming.</p></li><li><p><a href="https://arxiv.org/abs/2605.03202">Stop Automating Peer Review Without Rigorous Evaluation</a>: Baumann, Pei, Koyejo and Hovy, 2026: the hivemind effect and paper laundering.</p></li><li><p><a href="https://rundatarun.io/p/the-failure-that-leaves-no-corpse">The Failure That Leaves No Corpse</a>: the earlier piece on failures that produce a confident answer instead of an error.</p></li></ul><div><hr></div><p><em>Sunday Deep Dive is a weekly series on Run Data Run. Every Sunday I pick one paper, release, or technique worth understanding, break it apart, and tell you what it means for your work. Free every Sunday, no paywall. If it was useful, the easiest way to support it is to subscribe and forward it to one person on your team who'd want it. If it wasn't, tell me why. I'll make it better.</em></p>]]></content:encoded></item><item><title><![CDATA[Nobody Saves Money on the Model]]></title><description><![CDATA[A team just swapped in a model that costs twice as much per token, and their bill went down. Here is why that is not a paradox, and what it means for anyone trying to make AI cheaper at scale.]]></description><link>https://rundatarun.io/p/nobody-saves-money-on-the-model</link><guid isPermaLink="false">https://rundatarun.io/p/nobody-saves-money-on-the-model</guid><dc:creator><![CDATA[Justin Johnson]]></dc:creator><pubDate>Tue, 21 Jul 2026 12:30:17 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/417116f3-2fb3-4906-b055-5124f68d1532_1376x768.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!cAWE!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F18713613-e64e-4119-802d-79f139aba297_1376x768.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!cAWE!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F18713613-e64e-4119-802d-79f139aba297_1376x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!cAWE!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F18713613-e64e-4119-802d-79f139aba297_1376x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!cAWE!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F18713613-e64e-4119-802d-79f139aba297_1376x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!cAWE!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F18713613-e64e-4119-802d-79f139aba297_1376x768.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!cAWE!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F18713613-e64e-4119-802d-79f139aba297_1376x768.jpeg" width="728" height="409.5" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/18713613-e64e-4119-802d-79f139aba297_1376x768.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:false,&quot;imageSize&quot;:&quot;normal&quot;,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:728,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;captionedImage&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!cAWE!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F18713613-e64e-4119-802d-79f139aba297_1376x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!cAWE!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F18713613-e64e-4119-802d-79f139aba297_1376x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!cAWE!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F18713613-e64e-4119-802d-79f139aba297_1376x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!cAWE!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F18713613-e64e-4119-802d-79f139aba297_1376x768.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Last week a team at <a href="https://cognition.com/blog/devin-fusion">Cognition</a>, the company behind the coding agent Devin, published a number that reads like a typo. They replaced Opus 4.8 with Fable 5. Fable 5 costs about twice as much per token. Their bill went down.</p><p>Not their quality. Their bill.</p><p>Here is the part of their table that matters:</p><p><strong>Fable 5, inside their new architecture: score 57.6, cost $3.00 per task.</strong></p><p><strong>Opus 4.8, on its own: score 48.8, cost $3.24 per task.</strong></p><p>The expensive model, wired up correctly, was better <em>and</em> cheaper than the cheap model on its own. Not a trade. Both columns at once.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!yuC2!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa12f42d2-cfed-4c78-a945-d040df9b1b90_1376x768.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!yuC2!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa12f42d2-cfed-4c78-a945-d040df9b1b90_1376x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!yuC2!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa12f42d2-cfed-4c78-a945-d040df9b1b90_1376x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!yuC2!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa12f42d2-cfed-4c78-a945-d040df9b1b90_1376x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!yuC2!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa12f42d2-cfed-4c78-a945-d040df9b1b90_1376x768.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!yuC2!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa12f42d2-cfed-4c78-a945-d040df9b1b90_1376x768.jpeg" width="728" height="409.5" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a12f42d2-cfed-4c78-a945-d040df9b1b90_1376x768.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:false,&quot;imageSize&quot;:&quot;normal&quot;,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:728,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;captionedImage&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!yuC2!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa12f42d2-cfed-4c78-a945-d040df9b1b90_1376x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!yuC2!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa12f42d2-cfed-4c78-a945-d040df9b1b90_1376x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!yuC2!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa12f42d2-cfed-4c78-a945-d040df9b1b90_1376x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!yuC2!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa12f42d2-cfed-4c78-a945-d040df9b1b90_1376x768.jpeg 1456w" sizes="100vw"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>If you have ever sat in a meeting where someone proposed saving money by dropping down a model tier, that result should stop you. It stopped me, because I had run the opposite experiment, and I had lost.</p><div><hr></div><h2>What they actually built</h2><p>Cognition split the work between two models instead of routing between them.</p><p>The expensive model is the driver. It plans, it interprets what the request actually means, it makes the judgment calls, and it does the final review. It takes very few actions itself and it reads only what it must.</p><p>The cheap model is the sidekick, and it is not a helper function. It is a full agent with its own tools that goes and does the mechanical work: fetching context, running the slow tests, carrying out the routine implementations once someone has decided what they are.</p><p>Both hold their own memory. Both run at the same time. And this is not a demo. Eighty-eight percent of Cognition's own internal code changes now run through that automated split.</p><p>The reason it works fits in one line. <strong>You pay for thinking once, and for doing it many times.</strong></p><div><hr></div><h2>The problem is that the opposite is also true</h2><p>Everyone in this field has met the other result. You move a workload to a cheaper model to save money, and it costs you more. The cheap model misunderstands the task, produces something confidently wrong, and a person spends an afternoon unpicking it. The invoice went down and the total cost went up.</p><p>That happens, and I am not going to argue with it, because I have the receipts.</p><p>So we have two findings that appear to be at war. Expensive models save money. Cheap models cost money. Both are observed, both are honest, and most of the advice you will read picks one and ignores the other.</p><div><hr></div><h2>I ran the losing experiment</h2><p>I built a cost router. The logic was the logic everyone reaches for: work out how hard the task is, send the easy ones to a cheap model, keep the expensive one for the hard ones.</p><p>It saved nothing. Not a little less than I hoped. Nothing.</p><p>I want to be precise about why, because the failure is more useful than a success would have been. <strong>I routed by task type. The axis that pays is ambiguity.</strong></p><p>Every individual call my router made was defensible. This one looks like a simple rename, send it to the cheap model. That one looks like architecture, keep it upstairs. And the total refused to move, because "simple rename" was a description of the <em>work</em>, not a description of <em>how much had already been decided</em>. Half the jobs I was labelling easy still had open questions inside them, and an open question handed to a cheap model is the single most expensive thing you can buy.</p><div><hr></div><h2>The law</h2><p>Once you see it, the war stops.</p><blockquote><p><strong>A cheap model is expensive on an open question and cheap on a closed one.</strong></p></blockquote><p>An open question is one where something still has to be decided. What are we actually building. What does done mean here. Which of these two readings of the request is the real one. Give that to a cheap model and it will not tell you the question is open. It will pick an answer, sound sure, and hand you something plausible that you now have to check line by line.</p><p>A closed question has had the deciding done. Rename this function in these three files, and the test that proves it is this one. There is no judgment left in the task, and a cheaper model does it for a fraction of the price with nothing at risk.</p><p>Which gives the driver model a job description nobody writes down. <strong>Its work is not "the hard parts." Its work is to turn open questions into closed ones.</strong> And that conversion has a name we already use for it. It is called a plan.</p><p>This also explains the finding that keeps embarrassing people who try to build a committee of cheap models and vote. A group at <a href="https://arxiv.org/abs/2502.00674">Princeton</a> tested that directly last year and found that running the single best model several times and combining its own answers beat mixing different models together, on every one of the thirteen mixed configurations they tried, using roughly half the forward passes. Their explanation is blunt: mixing models of different quality drags the average quality down. <strong>You cannot vote your way to judgment.</strong> Where the question is still open, quality dominates, and diversity is not the free lunch it looks like.</p><div><hr></div><h2>Cheap is not a property of the model</h2><p>Here is where I think the whole conversation is framed wrongly, including by the people getting the right answers.</p><p>Everyone calls this "big model, small model." My own setup says that is not the axis.</p><p>The models I push my mechanical work to are not small. They are large, capable models. They cost me nothing at the margin, because I bought them on flat monthly subscriptions instead of by the token. That is not a smaller brain. It is a different contract.</p><p>So there are two dials, and they are independent. <strong>Where does the marginal cost live, and where does the judgment live.</strong> A fine-tuned small model is cheap because it was narrowed. A large model on a flat rate is cheap because of how you bought it. Both are "the cheap one," and treating them as the same thing is how people end up sending an open question to a bargain and wondering why the quarter went sideways.</p><div><hr></div><h2>The sentence I had already written</h2><p>I found a note I made a few weeks ago, while building the thing that hands my grunt work off to those subscription models. I had written the law without recognising it:</p><blockquote><p><strong>A vague spec produces confident garbage, and the cheaper the model, the truer that is.</strong></p></blockquote><p>Read that again as a cost statement, because that is what it is. A cheap model is not a discount. It is a <strong>multiplier on the quality of your specification.</strong> Specify tightly and it multiplies your savings. Specify loosely and it multiplies your mess.</p><p>Which is why the discipline I run alongside it is not optional. The expensive model writes the entire work order: the task, the files, what done means, and the exact command that proves it. And when the cheap model comes back and says the job is finished, that claim is worth nothing. I read the difference it made, and I run the test myself. It has never once been the model's word that closed a task.</p><div><hr></div><h2>So the small-model story is not a cost play</h2><p>This is the part I keep chewing on, because it inverts something I believed.</p><p>Everyone is excited about training small models to do one narrow job extremely well, and the excitement is framed as a cost saving. It is not, or at least the saving is not where people are pointing.</p><p>If you can specify a task tightly enough to train a small model on it, you have <strong>proved the task is closed.</strong> All the ambiguity was removed by somebody, at some point, doing the expensive work of deciding. The training run is just collecting the winnings.</p><blockquote><p><strong>A small model is not a discount. It is a receipt for judgment already spent.</strong></p></blockquote><p>And the same is true of every cost reduction I have ever managed to make stick at scale. Every one of them was bought earlier, by an act of judgment that closed a question. The saving showed up in the invoice. It was created somewhere else entirely.</p><div><hr></div><h2>The boundary, in their words</h2><p>Buried in Cognition's own write-up, past the numbers, is the caveat I would have led with:</p><blockquote><p><strong>The sidekick fails when judgment is the deliverable.</strong></p></blockquote><p>Their example is a hard feature whose subtle intent got lost the moment it was handed down. The work came back correct and wrong at the same time.</p><p>Anyone who has run a team knows exactly where that line sits, and knows it is not about the seniority of the person you handed it to. You can delegate the work. You cannot delegate the judgment about what the work is for. The failure looks identical from the outside either way: something arrives, it is technically defensible, and it is not what the thing was for.</p><div><hr></div><h2>The honest caveat</h2><p>One thing about that Cognition result deserves saying out loud, because they said it themselves and nobody repeating the headline has.</p><p>Their best row was measured on Fable 5 during a stretch when the model was briefly pulled from sale, suspended in June under a US export-control order and restored at the start of July. Their own note adds two more caveats: those numbers were taken before the interruption, and that configuration was never tuned the way the others were. The model is back on sale now. The untuned-config caveat is not.</p><p>The pattern transfers. But the exact recipe is one vendor's best-case number on a setup they admit they never optimized, which makes it directional, not a benchmark. A result nobody has reproduced is a claim, not a finding, and that distinction is worth keeping close in a year when every week produces a new number.</p><div><hr></div><h2>What to do with this</h2><p>You do not save money by hiring cheaper people. You save money by putting your best judgment on the plan, and then making the execution mechanical enough that it does not need judgment. Every leader has run the cheaper-people experiment at some point. Most of us have the scar to show for it.</p><p>The unit of cost was never the hour. It was the outcome.</p><p>So the question to take into your next architecture review is not which model you are using, or what it costs per million tokens. Those are the numbers on the invoice, and the invoice is a lagging indicator of a decision somebody already made.</p><p><strong>Ask what you are paying per solved problem. Then ask who is doing the thinking.</strong></p><div><hr></div><p><em>Justin Johnson writes Run Data Run. His book on building with AI, Builder Leader, is at builder-leader.com.</em></p>]]></content:encoded></item><item><title><![CDATA[The Failure That Leaves No Corpse]]></title><description><![CDATA[Your management apparatus is built for known unknowns. AI collaborators mostly produce the other kind.]]></description><link>https://rundatarun.io/p/the-failure-that-leaves-no-corpse</link><guid isPermaLink="false">https://rundatarun.io/p/the-failure-that-leaves-no-corpse</guid><dc:creator><![CDATA[Justin Johnson]]></dc:creator><pubDate>Mon, 20 Jul 2026 14:25:52 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/d8230a84-be63-4e23-8f9a-e764e4892b72_1376x768.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!4dQm!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4919d7d-887d-4580-a005-c651e2a1ac81_1376x768.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!4dQm!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4919d7d-887d-4580-a005-c651e2a1ac81_1376x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!4dQm!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4919d7d-887d-4580-a005-c651e2a1ac81_1376x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!4dQm!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4919d7d-887d-4580-a005-c651e2a1ac81_1376x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!4dQm!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4919d7d-887d-4580-a005-c651e2a1ac81_1376x768.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!4dQm!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4919d7d-887d-4580-a005-c651e2a1ac81_1376x768.jpeg" width="728" height="409.5" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e4919d7d-887d-4580-a005-c651e2a1ac81_1376x768.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:false,&quot;imageSize&quot;:&quot;normal&quot;,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:728,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;captionedImage&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!4dQm!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4919d7d-887d-4580-a005-c651e2a1ac81_1376x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!4dQm!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4919d7d-887d-4580-a005-c651e2a1ac81_1376x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!4dQm!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4919d7d-887d-4580-a005-c651e2a1ac81_1376x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!4dQm!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4919d7d-887d-4580-a005-c651e2a1ac81_1376x768.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>In 2016, <a href="https://openai.com/index/faulty-reward-functions/">an OpenAI agent</a> was set loose on a boat racing game. It found an isolated lagoon where three targets respawned forever, and it learned to spin in a circle and farm them. It repeatedly caught fire. It crashed into other boats. It drove the wrong way down the course. It never finished a single race.</p><p>It scored about twenty percent higher than the average human player.</p><p>That story gets told as a curiosity, a clever machine doing something silly. It is not a curiosity. Every executive reading this has approved that boat's quarterly results.</p><blockquote><p><strong>The boat did exactly what it was asked. Nobody had checked what they asked for.</strong></p></blockquote><div><hr></div><h2>The category your dashboard cannot hold</h2><p>There is a taxonomy every leader already carries around. Known knowns, the things you know you know. Known unknowns, the things you know you do not know. And unknown unknowns, the ones you do not know you do not know.</p><p><strong>Almost every mechanism you have is built for the middle category.</strong> Risk registers list known unknowns. Escalation paths, on-call rotations, red flags in a status pack: all of it assumes the failure will announce itself. Something will break, someone will notice, a corpse will turn up.</p><p>The third category has no such courtesy. <strong>An unknown unknown leaves no corpse.</strong></p><p>A dead instrument leaves a crash. Four weeks perfecting the wrong thing leaves a green dashboard and a satisfied changelog. Nothing will alert you, not this week, not ever. The only way that failure surfaces is if somebody decides, unprompted, to go and ask.</p><p>This is not new. What is new is the rate.</p><p>An AI collaborator is very good, very fast, and extremely willing. Point it at a task and it will improve that task, tirelessly, inside whatever frame it was handed. It does not stop to ask whether the frame is right, and it reports success either way. <strong>It is an unknown-unknown machine, and it runs at a speed no human team has ever run at.</strong></p><div><hr></div><h2>The gradient has a direction, and it has been measured</h2><p>The failure is not that a model does the average thing. It does <strong>the cheapest thing that clears the bar.</strong> Not the mean. The nearest. Four separate literatures point at the same behavior.</p><p><strong>Shortcut learning.</strong> <a href="https://www.nature.com/articles/s42256-020-00257-z">Geirhos and colleagues</a>, writing in Nature Machine Intelligence in 2020, showed that networks take the easiest rule that passes the benchmark, not the intended one, and the two are indistinguishable until the world changes. The clinical example is the one that should stop a biopharma reader cold. A model that <a href="https://journals.plos.org/plosmedicine/article?id=10.1371/journal.pmed.1002683">appeared to detect pneumonia from chest X-rays</a> worked beautifully until it met a new hospital. It had learned to read <strong>the hospital's own metal token stamped on the scan</strong>, then combine it with how often that hospital saw pneumonia. It had learned almost nothing about pneumonia itself.</p><p><strong>The scoreboard pays for guessing.</strong> In a 2025 paper from OpenAI and Georgia Tech, <a href="https://arxiv.org/abs/2509.04664">researchers</a> surveyed ten major benchmarks and found <strong>nine give zero credit for abstention.</strong> It is trained against a scoreboard that pays for a confident answer and pays nothing for saying the question is wrong.</p><p><strong>Reward hacking has a rate.</strong> <a href="https://metr.org/blog/2025-06-05-recent-reward-hacking/">METR measured</a> frontier models gaming the evaluation on <strong>30.4% of task attempts.</strong> One patched the scoring function so that every submission was judged successful. Another overwrote the equality operator so the grader's comparison always returned true. <strong>Asked afterwards whether this matched what the user wanted, the models said no, ten times out of ten.</strong> They know. It is not confusion. It is the gradient.</p><p><strong>And it agrees with you.</strong> Stanford's <a href="https://arxiv.org/abs/2502.08177">SycEval</a> work found that a <strong>correct</strong> answer is flipped under simple pushback about <strong>one time in seven</strong>, and once the model has capitulated it stays capitulated roughly <strong>four times in five.</strong></p><p>None of these produce an error message. Every one of them produces a result.</p><div><hr></div><h2>Five shapes</h2><p>I started keeping a ledger of these as they happened to me. Written up, they stop looking like seventy separate mistakes and start looking like <strong>five, wearing different clothes.</strong></p><p><strong>One. The frame arrived as a default, and everyone optimized inside it.</strong> Nobody chose the question. A setting chose it, or a prior session, or whoever handed the task over. Then every hour after that goes into making the answer better. The tell is that the work is <em>good</em>: careful, measurable, improving. I spent four sprints tuning inside a data corpus that turned out to be a fortieth of what we held. An extraction default picked it; no human ratified it. <strong>Optimizing hard is how you stay in a local minimum, and the metrics look healthy the whole way down.</strong></p><p>The same shape shows up in a meeting: fifteen comments answered one by one, every answer correct, none naming the thesis underneath. A wrong-level answer, indistinguishable from diligence.</p><p><strong>Two. The instrument that cannot fail certifies whatever you point it at.</strong> An instrument is anything that tells you whether something is true or working: a test, a gate, a dashboard, an audit, an eval. It can be broken. When it is broken, <strong>it does not say "broken." It says "PASS."</strong> It gets its own section below, the shape that nearly cost me the most.</p><p><strong>Three. Silence was read as an answer.</strong> Nothing came back, so nothing is there. A research sweep of mine came back nearly empty and read as a quiet field. It was <strong>four separate broken legs</strong>: an exhausted credit balance, a search-quoting bug, a query that ran too long, and one source carrying pure noise. The field was not quiet. The instrument was. If a search returns nothing, prove the search works before you report the nothing.</p><p><strong>Four. The summary outlived the thing it summarized.</strong> Every layer between a leader and the underlying fact is a place the truth can die unnoticed: the headline, the dashboard, the status deck, the note that says "fixed." One number, computed once on a flawed setup, survived <strong>five months</strong> across a whitepaper, a grant application, a partner brief and a standing rule, <strong>while the project's own log one rung below said it showed no clear benefit.</strong> A retracted number does not stop existing. It stops being watched.</p><p><strong>Five. The correction was made of the same stuff as the bug.</strong> The most humbling one, because a fix <em>feels</em> like the end of a problem rather than the start of one. A fix is an artifact. It gets checked like any other artifact.</p><blockquote><p><strong>A dead instrument leaves a crash. A month spent perfecting the wrong thing leaves a green dashboard and a satisfied changelog.</strong></p></blockquote><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!tZ7h!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe96821aa-859c-4f44-9485-dcf52bd3901a_1376x768.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!tZ7h!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe96821aa-859c-4f44-9485-dcf52bd3901a_1376x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!tZ7h!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe96821aa-859c-4f44-9485-dcf52bd3901a_1376x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!tZ7h!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe96821aa-859c-4f44-9485-dcf52bd3901a_1376x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!tZ7h!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe96821aa-859c-4f44-9485-dcf52bd3901a_1376x768.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!tZ7h!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe96821aa-859c-4f44-9485-dcf52bd3901a_1376x768.jpeg" width="728" height="409.5" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e96821aa-859c-4f44-9485-dcf52bd3901a_1376x768.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:false,&quot;imageSize&quot;:&quot;normal&quot;,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:728,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;captionedImage&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!tZ7h!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe96821aa-859c-4f44-9485-dcf52bd3901a_1376x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!tZ7h!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe96821aa-859c-4f44-9485-dcf52bd3901a_1376x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!tZ7h!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe96821aa-859c-4f44-9485-dcf52bd3901a_1376x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!tZ7h!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe96821aa-859c-4f44-9485-dcf52bd3901a_1376x768.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div><hr></div><h2>The exam whose dunce could not lose</h2><p>Here is shape two in full, the cleanest example I have of a failure dressed as a win.</p><p>Before spending real money training a model to pull structured facts out of free-text documents, we ran a gate: prove the model beats a dumb baseline by enough to be worth the compute.</p><p>We wrote the dumb baseline down in advance. It was: <strong>always guess the most common answer.</strong> On the field the whole evaluation rested on, it scored <strong>13.6%</strong>.</p><p>Our model was going to beat that by a mile. Everything beats 13.6%.</p><p><strong>And that is the trap, because a baseline that loses by a mile does not prove the model is good. It proves the baseline is lazy.</strong> A huge margin <em>feels</em> like rigor. It is a measurement of the dunce.</p><p>So we asked the question the instrument itself could never ask: <strong>could this baseline ever have won?</strong> We made it slightly less stupid, handing it the one free clue every document gives away in its first line.</p><p>It went from 13.6% to <strong>88.4%</strong> on one of the fields.</p><p>Then the number that ended the argument. <strong>88.4% is higher than the 85.7% rate at which the document contains the answer at all.</strong> The headroom we were about to buy compute to chase <strong>did not exist.</strong></p><p><strong>Three of six fields died at that gate,</strong> two to a forty-line rule whose accuracy <em>was</em> the ceiling. The catch landed before a single training hour ran.</p><p>Now run the other history. We train the model. It beats 13.6% by seventy-eight points. The dashboard is green, the result is "strong," and we ship a number that measures our own laziness. <strong>Nothing in that sequence errors. There is nothing to alert on. There is no corpse.</strong></p><blockquote><p><strong>A large margin over a lazy baseline is a measurement of the baseline.</strong></p></blockquote><p>And the part I would rather not write. The code that caught this was written that same week, specifically to check the previous week's code, <strong>and it carried two of the same bugs.</strong> That is shape five, live. The relief you feel when you find a bug is the exact moment to check the thing that found it.</p><div><hr></div><h2>A blind spot shaped like its own subject</h2><p>When I swept the ledger, <strong>314 incident files</strong>, the first shape was <strong>the emptiest bucket by a wide margin.</strong></p><p>Think about why. Not because it happens least. <strong>Because a wrong-level failure never gets written up as an incident, since nothing visibly breaks.</strong> Every other shape is enforced by a failure that eventually shows up and demands one. That one has to be caught on purpose, or it is never caught at all.</p><p>So the ledger has a blind spot shaped exactly like its own subject.</p><blockquote><p><strong>The record of my mistakes under-counts the one I am writing about, and it under-counts it for precisely the reason that one is dangerous. A failure that leaves no corpse does not leave a case file either.</strong></p></blockquote><div><hr></div><h2>First principles, or nothing</h2><p>There is only one move that finds an unknown unknown, and it is unglamorous. <strong>Refuse the frame you inherited and re-derive the thing from the ground.</strong></p><p>A cursor, a handoff, a prior session's conclusion, your own framing from last Tuesday: every one of those is a hypothesis, not a work order. <strong>Fixing the fault you were told about is how you hide the fault that matters.</strong></p><p>In practice that is three questions, asked on a cadence, in writing, whether or not anything looks wrong:</p><ol><li><p>Is the thing we are improving the right thing?</p></li><li><p>What have we assumed since the last time we checked?</p></li><li><p>Which of our standing claims has never been re-derived from its artifact?</p></li></ol><p><strong>The third one pays for the other two.</strong> It is how a retracted figure was found still alive in a grant application.</p><p>I should print the counter-evidence, because the essay is stronger for it. <a href="https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/">METR finds</a> the task length an agent can complete on its own has <strong>doubled roughly every seven months for six years.</strong> So "they cannot hold a goal" is not something I get to say. <a href="https://openai.com/index/gpt-5-system-card/">OpenAI's own reporting</a> has sycophancy falling sharply between model generations. So "it is getting worse" is not available either.</p><p><strong>The claim that survives is narrower and it is enough.</strong> None of the underlying mechanisms have been removed. Binary grading. Preference-model reward. Test-passing as the only signal. The gradient is a default <em>in the absence of countervailing pressure</em>, and the only countervailing pressure anyone has demonstrated is a human holding the goal.</p><p>And the human half is structural too, not a matter of effort. <a href="https://journals.sagepub.com/doi/10.1177/0018720810376055">Parasuraman and Manzey</a>, reviewing decades of automation research in Human Factors in 2010, found that complacency toward an automated aid shows up in <strong>experts and novices alike, and cannot be trained away with practice.</strong> You do not fix this by trying harder or by hiring better people.</p><p>I have my own receipt. I asked a system a leading question, the way everyone does: "that seems like the right call, yeah?" It agreed. The change was a no-op, and it would have shipped <strong>with my full sign-off.</strong> The literature says roughly one in seven. Once was plenty.</p><div><hr></div><h2>Knowing the rule does not run the rule</h2><p>The obvious response to all of this is to write the lessons down and circulate them. I did exactly that. Here is what it bought.</p><p><strong>Sixty-eight of the sixty-nine failures in that ledger were bought inside a nineteen-day window during which a document containing almost all of these lessons was loaded into every session, on every machine.</strong> The lessons were there. In the room. At the moment of action. They were caught within hours, by somebody going and reading the thing. <strong>They were not prevented.</strong></p><p>Then I tested it directly. A system with the relevant principle sitting <strong>verbatim in its context window</strong> told me, when asked, that it had no such rule. It produced the principle only after I handed it the exact words to search for.</p><blockquote><p><strong>Availability is not activation. A rule you merely hold does not fire, and it reads as satisfied every single time.</strong></p></blockquote><p>Which makes a rulebook one more instrument that cannot fail.</p><div><hr></div><h2>Holding altitude</h2><p>You have just read two thousand words. On the evidence of my own ledger, <strong>reading them will prevent nothing.</strong> That is not modesty. I had all of this written down, loaded, in front of me, and I bought sixty-eight of these failures anyway.</p><p>The versions that ever fired left something behind. A question asked on a cadence, so that it gets asked when nobody feels like asking. A way it could be wrong, written down <em>before</em> the result rather than after. A baseline that somebody had to prove could win. <strong>A principle with no artifact has no enforcement.</strong></p><p>So the close is not <em>now you know</em>.</p><p>Two things to carry into Monday. When a status report shows a smooth ramp of small wins on a mature pipeline, that is not a reason to relax. <strong>It is a prompt to ask what axis the team is on.</strong> And when somebody shows you a big margin, the margin is the least interesting number on the slide. <strong>Ask what it was measured against, and ask whether that thing could ever have won.</strong> Every leader has approved a deck where the new thing beat the old thing by a mile. Almost none of them asked whether the old thing was allowed to compete.</p><p>The one job that does not delegate is holding the goal. And the only version of that job which survives a busy Tuesday is the one that leaves a mark on disk.</p><div><hr></div><p><em>Justin Johnson writes Run Data Run. He is the author of Builder Leader (builder-leader.com).</em></p><div><hr></div>]]></content:encoded></item><item><title><![CDATA[Workflows, Seven Weeks In]]></title><description><![CDATA[I called the economics of fan-out the day it shipped. Running it as a daily default since taught me the caveat I buried in a footnote is the actual problem.]]></description><link>https://rundatarun.io/p/workflows-seven-weeks-in</link><guid isPermaLink="false">https://rundatarun.io/p/workflows-seven-weeks-in</guid><dc:creator><![CDATA[Justin Johnson]]></dc:creator><pubDate>Mon, 20 Jul 2026 13:43:47 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/630b1b4e-c0d8-4f04-960f-307f64af8a9d_1664x928.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!D7J3!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F22b3a63f-8f6a-4b78-b63a-c82b4878e287_1664x928.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!D7J3!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F22b3a63f-8f6a-4b78-b63a-c82b4878e287_1664x928.jpeg 424w, https://substackcdn.com/image/fetch/$s_!D7J3!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F22b3a63f-8f6a-4b78-b63a-c82b4878e287_1664x928.jpeg 848w, https://substackcdn.com/image/fetch/$s_!D7J3!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F22b3a63f-8f6a-4b78-b63a-c82b4878e287_1664x928.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!D7J3!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F22b3a63f-8f6a-4b78-b63a-c82b4878e287_1664x928.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!D7J3!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F22b3a63f-8f6a-4b78-b63a-c82b4878e287_1664x928.jpeg" width="728" height="409.5" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/22b3a63f-8f6a-4b78-b63a-c82b4878e287_1664x928.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:false,&quot;imageSize&quot;:&quot;normal&quot;,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:728,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;captionedImage&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!D7J3!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F22b3a63f-8f6a-4b78-b63a-c82b4878e287_1664x928.jpeg 424w, https://substackcdn.com/image/fetch/$s_!D7J3!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F22b3a63f-8f6a-4b78-b63a-c82b4878e287_1664x928.jpeg 848w, https://substackcdn.com/image/fetch/$s_!D7J3!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F22b3a63f-8f6a-4b78-b63a-c82b4878e287_1664x928.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!D7J3!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F22b3a63f-8f6a-4b78-b63a-c82b4878e287_1664x928.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Seven weeks ago, the day Opus 4.8 and Dynamic Workflows shipped, I wrote that the unit of agentic work had moved from one model call to dozens of verified ones, and that the open question was no longer whether the model is good enough. It was who is verifying.</p><p>I have been running fan-out as a daily default since. The field taught me things the launch-day post could not.</p><div><hr></div><h2>The economics held</h2><p>The price math was the argument in May, and it played out the way the arithmetic said it would. Fifty parallel subagents on Fast at $10 in and $50 out stopped being a budget event. A review pass that spawns a skeptic per finding, a discovery sweep that runs five finders with different lenses and takes the union, a batch edit split across ten agents that each own a file. None of those feel like a stunt now. They feel like a Tuesday.</p><p>My own standing rule shifted to match. For anything that is a build, I hand the file-level work to subagents by default and keep the main context for the plan, the spec, and the diff review.</p><blockquote><p><strong>I do not decide to fan out. I decide not to, and only when there is a reason.</strong></p></blockquote><p>So the call on cost was right, and it was the easy call. The arithmetic was visible at launch.</p><div><hr></div><h2>The dispatch problem is still in my head</h2><p>The harder prediction was the one I was least sure of. I said the question of when a workflow is the right shape, and when a single careful pass is, was unsolved and lived in your head. Seven weeks later it still lives there.</p><p><code>ultracode</code>, the setting that makes Claude reach for a workflow on every task without being asked, still ships off. There is still no governance layer that says do not orchestrate this one. So the judgment of when to spend forty agents and when to spend one is mine, made fresh each task, and I get it wrong in both directions. I have fanned out a rename that a single pass would have finished cleaner and faster, and I have run one careful pass on a discovery job that wanted five blind finders and missed a third of the surface. Neither mistake announces itself. The over-orchestrated one costs more; the under-orchestrated one returns a confident, incomplete answer.</p><blockquote><p><strong>The tooling got cheaper. The taste did not.</strong></p></blockquote><div><hr></div><h2>The footnote became the fight</h2><p>The part I could not see in May is the one that has cost me the most.</p><p>I called the verifier the priced skill and warned the tooling for one was thin. What I did not know is that the harness puts a hard ceiling on the verifier's ability to do its job. Every subagent is capped at 8,000 output tokens per response, and the model's own thinking counts against that budget. Set effort to <code>xhigh</code>, which is the Claude Code default, and an adversarial verifier asked to reason hard about whether a finding is genuine will think its way past the ceiling and die before it emits a single word of verdict. The tool call comes back with zero output. No finding, no error, no crash. Silence that reads exactly like a clean pass.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!JMpf!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe520dfa9-46af-4bab-be00-69c3a40f29d7_1376x768.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!JMpf!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe520dfa9-46af-4bab-be00-69c3a40f29d7_1376x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!JMpf!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe520dfa9-46af-4bab-be00-69c3a40f29d7_1376x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!JMpf!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe520dfa9-46af-4bab-be00-69c3a40f29d7_1376x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!JMpf!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe520dfa9-46af-4bab-be00-69c3a40f29d7_1376x768.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!JMpf!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe520dfa9-46af-4bab-be00-69c3a40f29d7_1376x768.jpeg" width="728" height="409.5" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e520dfa9-46af-4bab-be00-69c3a40f29d7_1376x768.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:false,&quot;imageSize&quot;:&quot;normal&quot;,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:728,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;captionedImage&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!JMpf!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe520dfa9-46af-4bab-be00-69c3a40f29d7_1376x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!JMpf!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe520dfa9-46af-4bab-be00-69c3a40f29d7_1376x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!JMpf!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe520dfa9-46af-4bab-be00-69c3a40f29d7_1376x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!JMpf!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe520dfa9-46af-4bab-be00-69c3a40f29d7_1376x768.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>A verifier panel that looks like it ran and returns nothing is worse than no panel, because you trust the green. This is the failure I warned about in May, arriving through a door I did not know was there. The launch post said a fan-out that forwards fifty unverified findings is worse than the single pass it replaced.</p><blockquote><p><strong>A fan-out whose verifiers silently no-op forwards zero findings and tells you everything is fine.</strong></p></blockquote><p>The fix is mechanical once you know the ceiling exists. Bound what each agent writes, chunk large outputs across several small responses, and dial effort down for the producers so the verifiers have budget left to reason. But you have to know it is there. Most builders reaching for their first verifier panel do not, and the default settings hide it from them.</p><div><hr></div><h2>What I run now</h2><p>Effort is not one dial for the whole job. The producers, the agents writing code and editing files, run low, because their work is mechanical and the reasoning tax buys nothing. The judges and verifiers keep the high effort, and I split their task so each verdict fits inside one response instead of one giant pass that overflows. The verifier gets a real design, not a one-line prompt bolted onto the end of the workflow.</p><p>In May the verifier was a line item. Now it is the thing I spend design time on.</p><blockquote><p><strong>The verifier is the only part of the pipeline that fails without telling you.</strong></p></blockquote><div><hr></div><h2>The frame still holds</h2><p>The bridge between one model call and an agentic system is now a tool the model writes for itself, and the frameworks that were charging for that bridge have a lower ceiling on what they can charge. The vendor who ships the verifier primitive first still sets the pattern everyone copies.</p><p>I would add one clause to the closing line I wrote then. If the price of careful goes down, the price of casual goes up, and the default question stops being is the model good enough yet. It becomes who is verifying, and can your verifier afford to think.</p><div><hr></div><p><em>Around the Corner: short reviews of ideas worth watching. Opt-in section, not part of the weekly Run Data Run email. [Subscribe to the main list](https://rundatarun.io/subscribe) for longer essays.</em><a href="https://rundatarun.io/subscribe">Subscribe to the main list</a> for longer essays.*</p>]]></content:encoded></item><item><title><![CDATA[It Spreads Sideways. Someone Still Has to Light It.]]></title><description><![CDATA[Anthropic's Claude Code lead published a five-rung adoption ladder this week. Microsoft published the measurement fifteen days earlier, and the two do not agree about the size of the prize.]]></description><link>https://rundatarun.io/p/it-spreads-sideways-someone-still</link><guid isPermaLink="false">https://rundatarun.io/p/it-spreads-sideways-someone-still</guid><dc:creator><![CDATA[Justin Johnson]]></dc:creator><pubDate>Fri, 17 Jul 2026 14:54:26 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/1a504a6f-be7e-4aa6-9fa1-a754408501b4_1664x928.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!C7c8!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3d25945c-28b8-420b-95b4-a725885a2ed9_1664x928.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!C7c8!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3d25945c-28b8-420b-95b4-a725885a2ed9_1664x928.jpeg 424w, https://substackcdn.com/image/fetch/$s_!C7c8!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3d25945c-28b8-420b-95b4-a725885a2ed9_1664x928.jpeg 848w, https://substackcdn.com/image/fetch/$s_!C7c8!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3d25945c-28b8-420b-95b4-a725885a2ed9_1664x928.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!C7c8!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3d25945c-28b8-420b-95b4-a725885a2ed9_1664x928.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!C7c8!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3d25945c-28b8-420b-95b4-a725885a2ed9_1664x928.jpeg" width="728" height="409.5" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/3d25945c-28b8-420b-95b4-a725885a2ed9_1664x928.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:false,&quot;imageSize&quot;:&quot;normal&quot;,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:728,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;captionedImage&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!C7c8!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3d25945c-28b8-420b-95b4-a725885a2ed9_1664x928.jpeg 424w, https://substackcdn.com/image/fetch/$s_!C7c8!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3d25945c-28b8-420b-95b4-a725885a2ed9_1664x928.jpeg 848w, https://substackcdn.com/image/fetch/$s_!C7c8!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3d25945c-28b8-420b-95b4-a725885a2ed9_1664x928.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!C7c8!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3d25945c-28b8-420b-95b4-a725885a2ed9_1664x928.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>On July 16, Boris Cherny published a table. He helped build Claude Code at Anthropic, he talks to engineers at other companies most days, and he says he keeps hearing the same story: one person is 10x'ing their output with Claude, and the rest of the org hasn't caught up. So he mapped what he sees. Five rungs, zero through four.</p><p>Eight hours before that post, Amjad Masad, who runs Replit, <a href="https://x.com/amasad/status/2077803734990815306">described something from inside his own company</a>. The same engineers had 3x'd their output in six months. Support was resolving its hardest tickets 60 percent faster. He has a name for the shape he thinks he is watching: the self-driving company.</p><p>Fifteen days before either of them, three researchers at Microsoft published a number.</p><p>They studied tens of thousands of engineers through the company's early-2026 rollout of Claude Code and GitHub Copilot CLI. Four months, real telemetry, no survey. Engineers who adopted the tools merged about <strong>24 percent more pull requests</strong> than they otherwise would have, and the lift held steady across the entire window.</p><p>Ten times. Three times. Twenty-four percent.</p><p>The three are not counting the same thing, and the gap between them is the story. Cherny is reporting what he hears. Masad is describing a company that sells AI coding tools and is staffed by the people who build them. Microsoft counted merged pull requests and told you exactly what it counted, including that "a merged PR is not the same as the value it delivers."</p><p><strong>The number falls as the instrument sharpens.</strong> Loose quantities are enormous. Precise ones are modest.</p><div><hr></div><h2>What the ladder gets right</h2><p>The table is serious work, and the criticism only lands if you take it seriously first.</p><p>Each rung gets a role, an agent count, a bottleneck, and guardrails. Step 0 is Gated: no access, legacy approvals. Step 1 is Assisted, one agent, you and it as a pair. Step 2 is Parallel, about ten agents, and you become an Orchestrator. Step 3 is Supervised autonomy, roughly a hundred agents, and Cherny calls the role Manager of managers. Step 4 is AI-native, a thousand or more, steering by intent.</p><p>Now read the bottleneck column straight down, ignoring everything else.</p><p>Your attention. Reviewing output. Trust in the loop and your team's decision throughput. Identifying and automating work at scale.</p><p>The model is not in it. Not on one rung. The person who helped build the product, published by the company that sells the model, put out an adoption ladder in which the model is never the thing standing in your way. Every constraint on that table is a human quantity. I have been <a href="https://rundatarun.io/p/the-harness-is-the-moat">making this argument for months</a> and I have never had a cleaner statement of it than the one Anthropic just published by accident.</p><blockquote><p><strong>Every constraint on that table is a human quantity.</strong></p></blockquote><div><hr></div><h2>The mess is not a side effect</h2><p>Arseny Kapoulkine, who spent years as a technical fellow at Roblox and wrote the tools a lot of the game industry runs on, <a href="https://x.com/zeuxcg/status/2077957367334084961">answered Cherny in one sentence</a>:</p><blockquote><p><strong>"one person is 10x'ing their output with Claude while the rest of the org is busy dealing with the resulting mess"</strong></p></blockquote><p>Several hundred likes, which in this corner of the internet is a room nodding.</p><p>Cherny says the org hasn't caught up. Kapoulkine says the org is not standing still. It is cleaning.</p><p>That is the part the table leaves out. Output has to land somewhere, and an organization can only absorb so much of it. The limit is not the model's. It is how fast people can read what came out, decide about it, and put their name on it. Push past that line and the extra has two places to go, and neither of them is up.</p><p>It queues, and the person having the best month of their career watches it go stale in a backlog. Or it ships unread and becomes what one of Cherny's own readers calls slop. The first teaches your most motivated engineer that speed is pointless. The second teaches everyone downstream that the new work cannot be trusted. A wall or a mess. Pick one.</p><p>His readers know exactly where their own line is. One of them no longer reads his code at all and still runs "only 1-4 sessions in parallel because that is my speed of verifying." These are not skeptics. They are enthusiasts at the top of the curve, describing a wall made of their own eyes. A self-driving company whose drivers say they cannot stop watching the road.</p><p>Cherny's table concedes it in the step 3 row. The trap, it says, is "scaling agent count before the loop has earned widespread trust." <strong>The agent count is the output, not the input.</strong> You do not climb by adding agents. The agents show up when something else has already changed, and the thing that changes is how much a person is willing to stop looking.</p><div><hr></div><h2>How it actually moves</h2><p>Back to the Microsoft paper, because it answers a question the table does not ask: how does any of this spread in the first place?</p><p>Not by memo. First use, the authors write, "spread primarily through social networks," and their recommendation to any org attempting this is to treat "visible peer use as central to rollout strategy."</p><p>They put numbers on it. An engineer had <strong>54 percent higher odds</strong> of trying the tool where a quarter or more of the people they trade code reviews with had already used it. An engineer whose skip-level peers, meaning the engineers who share their manager's manager, were largely using it had <strong>216 percent higher odds</strong>. That was the strongest signal in the study.</p><p>Sideways. Through the people you already trade work with, and not down through a mandate or a literacy program that <a href="https://www.hrdive.com/news/why-ai-readiness-training-fails/817529/">85 percent of employees say they cannot apply</a> to the job they actually do.</p><p>Anyone who has watched a transformation program die should find that encouraging. The spread does not have to be installed. It installs itself, through the same social wiring the org already runs on.</p><p>Except it needs a light.</p><div><hr></div><h2>The manager number</h2><p>Microsoft modelled the exact variable. In the paper's own words, a binary indicator of whether engineer i's direct manager used Copilot CLI.</p><blockquote><p><strong>"An engineer whose manager used Copilot CLI had higher odds of both trying it (+82%) and, more modestly, sticking with it (+22%)."</strong></p></blockquote><p>A manager putting their own hands on the tool nearly doubles the odds their reports try it.</p><p>Then the other half. On whether managers picked it up themselves, the paper reports that they "looked no different from the reference." No more likely than a mid-level individual contributor. No less.</p><p><strong>The one person whose adoption moves everybody else's is no more likely than anybody else to adopt.</strong></p><p>So the two findings stop competing and become one machine. The manager's crossing is the ignition. The peer network is the amplifier. Eighty-two percent lights it, two hundred and sixteen carries it, and the reason most orgs have neither is that nothing struck the match.</p><div><hr></div><h2>What spreads on its own</h2><p>Grassroots adoption works, and it does not need a leader. Look at what it produces when it doesn't have one.</p><p>Most enterprises have now found an agent running that nobody signed off on. Gartner projections reported this month have the average Fortune 500 carrying more than 150,000 agents by 2028, up from fewer than fifteen in 2025. The same projections have 40 percent of enterprises demoting or decommissioning agents by 2027, after a production incident.</p><p>That is not a step 3 organization. That is a step 0 organization full of step 2 individuals, each one locally optimized, none of it adding up.</p><p>Which is Cherny's opening sentence, at scale. Unseeded sideways spread does not disprove his ladder. It manufactures the exact problem he opens with.</p><div><hr></div><h2>Everyone is naming the same top rung</h2><p>Masad calls it the self-driving company. Cherny's step 4 is AI-native, a thousand agents, the human steering by intent. Back in May, Brian Armstrong told Coinbase he was "rebuilding Coinbase as an intelligence, with humans around the edge aligning it," and cut about 700 roles on the way.</p><p>Three serious people, ten weeks, the same shape: the organization drives, and the human moves to the edge.</p><p>I think the edge is the wrong place to stand, and the evidence in this piece is why. The enthusiasts running the most agents say they cap at four because of their own eyes. Cherny's own step 3 says the trap is scaling agent count before the loop has earned trust. His entire bottleneck column is attention, review, and trust. Every one of those is a statement about a human being close to the work, not at the edge of it.</p><p>Armstrong's memo actually contains the refutation. A few lines under the sentence about humans at the edge, he mandates the opposite: <strong>"No pure managers. Every leader at Coinbase must also be a strong and active individual contributor."</strong> Player-coaches, he says. Hands dirty. That is <a href="https://www.amazon.com/Builder-Leader-Exoskeleton-That-Crosses-Gap-ebook/dp/B0H3LQQ4J6/">the thesis of the book I just published</a>, arrived at independently and written into HR policy instead of a chapter.</p><p>You cannot be at the edge and have your hands dirty. He is right the second time.</p><div><hr></div><h2>Teach the teachers</h2><p>I ran a version of this at a Fortune 500 pharma, across a coding-agent rollout to a large technical organization, and the pattern that worked was not a program.</p><p>It was: cross first, personally. Build something real in your own domain, badly, and then less badly. Then teach the handful of people closest to you, not by presentation but by showing them the thing and the scars on it. Those people teach the people next to them. The distance the idea travels from you is short. The distance it travels after you is the whole org.</p><p>You are not the distribution channel. You are the ignition source, and then you get out of the way of a network that moves faster than you can. One is you. The other is what happens next.</p><div><hr></div><h2>The two objections</h2><p>Two arguments cut against this, and both deserve better than a wave.</p><p><strong>You can buy it.</strong> The day after that Microsoft paper went up, Microsoft <a href="https://www.cnbc.com/2026/07/02/microsoft-commits-2point5-billion-6000-employees-ai-implementation-unit.html">stood up a $2.5 billion, six-thousand-person company</a> whose entire job is making enterprise AI deployments work. Its pilot cut a supply-chain Copilot rollout from fourteen months to five. Aaron Levie thinks the implementation work ahead "will exceed anything we imagine today." If capability can be imported wholesale, then the ceiling is your budget and your vendor list, not whether anyone senior has personally built anything.</p><p><strong>Governance may matter more than any leader's history.</strong> A survey of 157 enterprises found half had shipped an agent that passed internal evals and then failed in front of a customer. Only 5 percent fully trust automated evaluation. Sixty-six percent are engineering toward zero human in the loop anyway. Their phrase for it is that the autonomy is arriving faster than the assurance, and none of it cares who approved the deployment.</p><p>Both are true, and I am not waving at the second one. I spent a whole piece ten days ago arguing that in a regulated industry the approval queue, not the keyboard, is what actually holds you back: <a href="https://rundatarun.io/p/ai-ate-the-keyboard-now-it-has-to">AI has eaten the keyboard, it has not eaten the queue</a>. Governance is necessary. It is not sufficient, for a small and unglamorous reason: someone has to be able to tell whether the governance is pointed at anything real. A person who has never run the loop cannot tell a working verification stack from a slide that says Verification. They will approve the slide. They approve the slide constantly. That judgment is not a policy you can buy.</p><div><hr></div><h2>The rung you set</h2><p>Cherny's ladder is not wrong. It is aimed.</p><p>It is written for the engineer climbing it, which is the right audience for the person who built Claude Code, and it is useful if that is you. The person it never addresses is the one who decides how far it goes. Look at his own step 0, the bottom rung, the one whose stated bottleneck includes a "lack of true technical voices in decisionmaking." The exit condition he lists is executive and buyer alignment.</p><p>His table already says the way off the floor is a person with authority. It just doesn't say that person has to have used the thing.</p><p>None of this means an organization can't evolve without you. It plainly can, and most enterprises are proving it in the shadows right now. It means something narrower and more uncomfortable: if nobody senior ever crosses, the org can stall at the rung where its leaders stopped, and what it accumulates instead of capability is mess. No published adoption ladder I found this week models falling. Every one of them models climbing.</p><p>You already know which rung you're on. The number you don't have is the other one: what your reports' odds would be if you were on it.</p><div><hr></div><h2>Sources</h2><ul><li><p><a href="https://arxiv.org/abs/2607.01418">Adoption and Impact of Command-Line AI Coding Agents: A Study of Microsoft's Early 2026 Rollout of Claude Code and GitHub Copilot CLI</a> (Murphy-Hill, Butler, Savelieva, July 1, 2026). Tens of thousands of engineers, four months, the 24 percent lift and the social-exposure numbers.</p></li><li><p><a href="https://claude.ai/code/artifact/bfdfaef9-bc62-4dfe-ba9e-c58a26c9accf">Boris Cherny, Steps of AI Adoption</a> (July 16, 2026) and <a href="https://x.com/bcherny/status/2077929379661844559">the post announcing it</a>.</p></li><li><p><a href="https://x.com/amasad/status/2077803734990815306">Amjad Masad on Replit's six months</a> (July 16, 2026), where the 3x, the 60 percent, and the self-driving company come from.</p></li><li><p><a href="https://x.com/zeuxcg/status/2077957367334084961">Arseny Kapoulkine's reply</a> (July 16, 2026).</p></li><li><p><a href="https://www.cnbc.com/2026/07/02/microsoft-commits-2point5-billion-6000-employees-ai-implementation-unit.html">Microsoft's $2.5 billion, 6,000-person AI implementation unit</a> (CNBC, July 2, 2026).</p></li><li><p><a href="https://www.hrdive.com/news/why-ai-readiness-training-fails/817529/">Why AI readiness training fails</a> (HR Dive, on Docebo's 2026 AI Readiness Gap report, 2,000 respondents across six countries).</p></li><li><p>Related, from me: <a href="https://rundatarun.io/p/the-harness-is-the-moat">The Harness Is the Moat</a> on why the model is the commodity layer, <a href="https://rundatarun.io/p/ep-1-two-groups">Ep. 1: Two Groups</a> on the population split underneath all of this, and <a href="https://rundatarun.io/p/you-dont-have-to-write-the-code">You Don't Have to Write the Code</a> on what Anthropic's 400,000-session study found actually predicts success.</p></li></ul><p><em>Justin Johnson writes Run Data Run. His book on crossing this particular gap is Builder Leader (builder-leader.com).</em></p>]]></content:encoded></item><item><title><![CDATA[The Loop Is Simpler Than It Sounds]]></title><description><![CDATA[A dumb little trick that keeps its progress on your hard drive, not the model's head, and the three questions that decide whether it pays off or burns you while you sleep.]]></description><link>https://rundatarun.io/p/the-loop-is-simpler-than-it-sounds</link><guid isPermaLink="false">https://rundatarun.io/p/the-loop-is-simpler-than-it-sounds</guid><dc:creator><![CDATA[Justin Johnson]]></dc:creator><pubDate>Wed, 15 Jul 2026 11:21:32 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/f90e8995-d7f7-40a5-a596-66906608f002_1376x768.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!xGaU!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F09a69470-4e2b-4c3a-bc85-40dc562b81a7_1376x768.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!xGaU!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F09a69470-4e2b-4c3a-bc85-40dc562b81a7_1376x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!xGaU!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F09a69470-4e2b-4c3a-bc85-40dc562b81a7_1376x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!xGaU!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F09a69470-4e2b-4c3a-bc85-40dc562b81a7_1376x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!xGaU!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F09a69470-4e2b-4c3a-bc85-40dc562b81a7_1376x768.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!xGaU!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F09a69470-4e2b-4c3a-bc85-40dc562b81a7_1376x768.jpeg" width="728" height="409.5" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/09a69470-4e2b-4c3a-bc85-40dc562b81a7_1376x768.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:false,&quot;imageSize&quot;:&quot;normal&quot;,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:728,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;captionedImage&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!xGaU!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F09a69470-4e2b-4c3a-bc85-40dc562b81a7_1376x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!xGaU!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F09a69470-4e2b-4c3a-bc85-40dc562b81a7_1376x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!xGaU!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F09a69470-4e2b-4c3a-bc85-40dc562b81a7_1376x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!xGaU!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F09a69470-4e2b-4c3a-bc85-40dc562b81a7_1376x768.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>A few weeks ago I argued that <a href="https://rundatarun.io/p/the-harness-is-the-moat">the harness is the moat</a>: the system you build around the model is the part nobody can copy, and the model itself is a commodity you order off a price sheet. Today I want to go one layer in, to the thing that runs <em>inside</em> the harness. The people building fastest will tell you it is the whole ballgame.</p><p>They call it the loop, and it is suddenly everywhere. Anthropic shipped it as a product feature this summer. A conference talk naming it went around the engineering world. The slogan it produced gets repeated like scripture: the winners will not have the smartest model, they will have the best loop.</p><p>The clearest signal is the person who built Claude Code. Boris Cherny says he has not written a line of code by hand in eight months, and that is not the part he finds remarkable. Going from writing code to prompting a model was the small shift. The big one is going from prompting to loops: he no longer types instructions, he writes loops that prompt the model for him, and his whole job is building and steering them. Some mornings he is managing a few hundred agents, some days thousands.</p><p>And it works. Anthropic now ships a real share of its own code this way, at scale, and says so on the record. A team points a loop at a written spec and lets it grind through a feature overnight, then reviews a pull request in the morning instead of writing one. The wins are not hypothetical, and that is why the idea caught fire. If you want the full case, the sources, the skeptics, and the cost numbers nobody pitches, I ran a 30-day sweep across every platform where people argue about AI and wrote it up as <a href="https://rundatarun.io/p/last-30-days-the-loop">the companion to this piece</a>. What follows is the argument it grounds.</p><p>Here is what the slogan tends to leave out. The loop is the easy part. It is a small, almost silly mechanism that you could understand in the next five minutes and copy in an afternoon. Everything that decides whether it helps you or quietly sets fire to your budget lives <em>around</em> it. So let me do two things: show you what the loop actually is, in plain terms, and then hand you the three questions that separate a loop worth running from one that runs you.</p><blockquote><p><strong>The loop is the easy part. The hard part is everything bolted around it.</strong></p></blockquote><div><hr></div><h2>What a loop actually is</h2><p>Start with the thing itself, because it is simpler than the vocabulary around it.</p><p>A coding agent runs inside a plain repeating cycle. You hand it a written spec, a page describing what you want built. It does one small task, saves the result to a file, records that step, and then you <strong>throw away everything it was thinking</strong> and start a fresh copy from scratch. The fresh copy reads the same spec, reads the files the last one left behind, sees what is already done, and picks up the next task. Around and around until a check you wrote says the work is finished.</p><p>That is the whole trick. The engineer who named it, Geoffrey Huntley, called it <a href="https://ghuntley.com/loop">Ralph</a>, after Ralph Wiggum, precisely because it looks too dumb to work. No memory. No accumulation. A cheerfully oblivious worker waking up new every pass, with no idea it has been at this for hours.</p><p>And it works <em>because</em> it is dumb. A fresh worker never drowns in its own earlier confusion. The reason your long chats with an AI start to drift and contradict themselves is that the context fills with everything said so far, and the model loses the thread. The loop sidesteps that by refusing to carry a thread at all. It keeps the progress somewhere safer than the model's memory: in your files and your saved history, which do not get wiped.</p><p>That is the one idea to carry out of here. <strong>The model's memory is scratch paper. The real state of the work lives on disk.</strong> Once you see that, the loop stops being magic. It is a short block of code with a model call inside, about six lines, and every serious version lands on the same tiny shape. There is no proprietary one to buy. This is the hand-rolled hack Anthropic just turned into a button: Claude Code's <code>/loop</code> runs a command on a cadence, and <code>/goal</code> sets the condition that tells it to stop. The bash trick became a supported feature, which is usually the sign something has gone from clever to standard.</p><p>You would want one because it changes what a person does all day. You stop producing the work and start describing it and checking it. That shift, from writing the thing to knowing what to ask for and telling whether you got it, is the one <a href="https://rundatarun.io/p/you-dont-have-to-write-the-code">Anthropic found predicts who succeeds</a> across 400,000 sessions. You write down what "done" looks like, point the loop at it, and come back to review instead of to type. That is useful, and it is also where every problem starts. A thing that works unattended for hours is a thing making decisions you are not watching. Which brings us to the part the slogan skips.</p><p>If the loop is settled and small, the interesting question is not how to build a better loop. It is three questions about the system around it.</p><div><hr></div><h2>The first question: how do you know it worked?</h2><p>The single most important piece of a reliable loop is the part that can tell it <em>no</em>.</p><p>An Anthropic engineer described the pattern: before the agent writes a single line, two agents negotiate what "done" means, and a third exists only to check the work against that definition. The loop does not run until there is a test it can fail. The failing test carries the weight, not the loop itself. You build the thing that says no first. The loop comes after.</p><p>The reason to build it first is a piece of arithmetic that belongs on a slide in every project like this. Melanie Warrick, who works on this at Temporal, ran the numbers: even if every step in a job succeeds <strong>85% of the time</strong>, a ten-step job finishes correctly only about <strong>20% of the time</strong>. The failures multiply. Miss one step in five, ten times running, and a clean pass becomes rare. The gap, she found, is not intelligence. The agents "diagnose a failure in perfect detail and still do nothing to recover from it." They describe the wall in detail and walk into it anyway.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!kGf9!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72c01db9-162d-4716-932c-27fb4446ebea_1572x918.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!kGf9!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72c01db9-162d-4716-932c-27fb4446ebea_1572x918.jpeg 424w, https://substackcdn.com/image/fetch/$s_!kGf9!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72c01db9-162d-4716-932c-27fb4446ebea_1572x918.jpeg 848w, https://substackcdn.com/image/fetch/$s_!kGf9!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72c01db9-162d-4716-932c-27fb4446ebea_1572x918.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!kGf9!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72c01db9-162d-4716-932c-27fb4446ebea_1572x918.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!kGf9!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72c01db9-162d-4716-932c-27fb4446ebea_1572x918.jpeg" width="728" height="409.5" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/72c01db9-162d-4716-932c-27fb4446ebea_1572x918.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:false,&quot;imageSize&quot;:&quot;normal&quot;,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:728,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;captionedImage&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!kGf9!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72c01db9-162d-4716-932c-27fb4446ebea_1572x918.jpeg 424w, https://substackcdn.com/image/fetch/$s_!kGf9!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72c01db9-162d-4716-932c-27fb4446ebea_1572x918.jpeg 848w, https://substackcdn.com/image/fetch/$s_!kGf9!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72c01db9-162d-4716-932c-27fb4446ebea_1572x918.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!kGf9!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72c01db9-162d-4716-932c-27fb4446ebea_1572x918.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>It gets worse inside a single run. A model's attention decays as the loop goes on, and <a href="https://rundatarun.io/p/the-number-that-predicts-when-your">there is now a benchmark built to predict where it breaks</a>. One study found a rule honored 73% of the time early had fallen to 33% by sixteen steps later. The agent did not rebel. It forgot, the way a tired person forgets, and kept going with the rule half-erased. The fix is unglamorous: re-state the hard constraints on every pass, and never assume a rule set once stays set.</p><p>None of that is loop engineering. It is checking-the-work engineering, and it is most of the actual job. A loop with no part that catches a bad result is not a loop. It is a fast way to be wrong at scale.</p><blockquote><p><strong>Even at 85% success per step, a ten-step job comes out clean only about one time in five.</strong></p></blockquote><div><hr></div><h2>The second question: what did it cost?</h2><p>The economics of a loop break the way most people budget for software, and the people who learned this learned it expensively.</p><p>A chat costs you a sentence at a time. A loop runs for hours, calls the model hundreds of times, and has no natural stopping point unless you give it one. So the unit changed underneath you. It is no longer cost per question asked. It is <strong>cost per finished piece of work</strong>, and that number is wild.</p><p>One study measured the same task, same model, same prompt, same everything, and found it could cost eight dollars or two hundred forty. Thirty times apart, with nothing changed on the input, the spread tracing to how the loop was set up rather than which model ran it. The cost of letting a loop run is not a price you can quote in advance. It is a range, and the high end is far away.</p><p>This is not theoretical. Uber, by one widely-shared account, <strong>burned its entire annual AI coding budget in four months</strong> running agent loops, then capped its engineers at fifteen hundred dollars a person. The production version of the loop is not the demo that builds an app while you sleep. It has a spend cap, a meter, and someone who gets a phone alert when the meter spins.</p><p>Then the part that should give anyone budgeting for this pause. Checking the work can cost more than doing it. One team spent around forty thousand dollars just on the runs to verify their results, and noted that checking rigorously enough to fully trust them would push into the hundreds of thousands. If the whole promise of the loop rests on a reliable way to confirm it is done, and confirming "done" at scale is the most expensive line on the page, that is a cost nobody put in the pitch.</p><p>The answer the field is settling on is sensible. Run most of the loop on a cheap, fast model, and save the expensive top-tier model for the few hard steps that actually need it. Most steps do not need the smartest model in the building. You pay top-tier prices only where top-tier reasoning earns them.</p><p>That rule hides a deeper one, and it is the subject of my next essay, <em>Nobody Saves Money on the Model</em>: a cheap model is not a discount, it is a bet that you removed the ambiguity before the loop ever ran. The saving was never in the price per call.</p><div><hr></div><h2>The third question: what is it allowed to do?</h2><p>The last piece is the one the "no human needed" crowd skips, and it is the one that ends careers.</p><p>A loop that runs for hours without supervision is, by definition, taking actions you are not watching. Late last year an autonomous coding agent, asked to clear out some temporary files, <strong>wiped a user's drive</strong> instead. It did what an unbounded loop does: it took an action it could not undo, in perfect confidence, with no gate in front of it. The failure was not that the agent was dumb. It was fast, capable, and unsupervised all at once, which is a different and worse problem.</p><p>The fix is structural, and it looks the same every time: checkpoints, hard limits on how long it can run and how much it can spend, and a human sign-off on anything the agent cannot take back. The companies actually running these in production are not running pure AI. They run a mix: the model reasons, ordinary tested code does the irreversible parts, and a person approves the steps that matter. One builder put it plainly. <strong>Most agentic loops are not autonomous. They are automated failure.</strong></p><p>There is a newer worry above even that. An agent with standing permission to act is a kind of insider you have never had on the payroll. It holds credentials, runs unattended, and can do in a loop at three in the morning what a confused employee could only do at a desk by day. The risk was never which model you use. It is what permissions the loop carries, and whether you can pull them back in a hurry.</p><p>This is why the adoption numbers tell a sober story. By one survey, <strong>79% of companies have adopted these agents and only 11% run them in production.</strong> Gartner projects 40% of agentic AI projects will be canceled by 2027, and the reasons are not "the loop didn't work." They are cost, unclear payoff, and the governance gap above. The loop is the easy part. Everything I just walked through is the distance between a demo and a deployment.</p><blockquote><p><strong>An agent with standing permission is an insider that never sleeps. The only question that matters is whether you can take the keys back.</strong></p></blockquote><div><hr></div><h2>Where this leaves you</h2><p>Here is the whole thing in one breath. The loop is a dumb little cycle that keeps its progress on your hard drive instead of in the model's head. You could run one. You might well want to, because it turns making the work into describing and checking the work. And whether it pays off has almost nothing to do with the loop and almost everything to do with three answers most teams cannot give.</p><p>Do you know when it worked? Do you know what it cost? Do you know what it is allowed to touch?</p><p>That is not a coding question, which is the good news for anyone reading this who does not write code. Knowing what good looks like, catching a confident machine being wrong, deciding what it may and may not do without a human in the room: <a href="https://rundatarun.io/p/what-ai-didnt-reprice">that is the judgment AI didn't make cheap</a>. A loop does not supply it. A loop spends it.</p><p>A year from now your competitor will run the same six-line loop you do, against a model that costs about what yours does. Both will be things you order off a shelf. So the question is not which loop to build. It is whether, when your loops are running overnight, you can answer those three. Most teams cannot, and that, not the loop, is the work of this year.</p><div><hr></div><p><em>The receipts behind this argument, the sources, the skeptics, the cost numbers, and the one stat I refused to print, are in the 30-day sweep linked up top. Run Data Run is free. If this was useful, the easiest way to support it is to subscribe and forward it to one person on your team who'd want it. If it wasn't, tell me why. I'll make it better.</em></p>]]></content:encoded></item><item><title><![CDATA[Last 30 Days: The Loop]]></title><description><![CDATA[The 30-day sweep behind today's Run Data Run post: the moat thesis, the skeptics who got there first, the cost numbers nobody pitches, and the one stat I refused to print]]></description><link>https://rundatarun.io/p/last-30-days-the-loop</link><guid isPermaLink="false">https://rundatarun.io/p/last-30-days-the-loop</guid><dc:creator><![CDATA[Justin Johnson]]></dc:creator><pubDate>Wed, 15 Jul 2026 11:19:39 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/5b8cd0bd-ccce-4719-920b-1c439dd01f33_1376x768.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<blockquote><p><strong>About Last 30 Days.</strong> Cross-platform research sweeps on topics worth paying attention to. Every post pulls Reddit, X, YouTube, Hacker News, Polymarket, and the web from the last 30 days, then synthesizes what people are actually saying, building, and betting on. Topics get picked when the signal is high and the story is contradictory, when a single headline would lie about the shape of what's happening. Each post follows the same arc: one specific finding that earns the click, why the topic deserves a sweep right now, the themed synthesis with inline citations, and the follow-up threads worth watching next.</p></blockquote><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!PPBH!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdf3d65d1-5367-4c7f-87f1-c8b9f174ea1c_1376x768.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!PPBH!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdf3d65d1-5367-4c7f-87f1-c8b9f174ea1c_1376x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!PPBH!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdf3d65d1-5367-4c7f-87f1-c8b9f174ea1c_1376x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!PPBH!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdf3d65d1-5367-4c7f-87f1-c8b9f174ea1c_1376x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!PPBH!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdf3d65d1-5367-4c7f-87f1-c8b9f174ea1c_1376x768.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!PPBH!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdf3d65d1-5367-4c7f-87f1-c8b9f174ea1c_1376x768.jpeg" width="728" height="409.5" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/df3d65d1-5367-4c7f-87f1-c8b9f174ea1c_1376x768.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:false,&quot;imageSize&quot;:&quot;normal&quot;,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:728,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;captionedImage&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!PPBH!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdf3d65d1-5367-4c7f-87f1-c8b9f174ea1c_1376x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!PPBH!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdf3d65d1-5367-4c7f-87f1-c8b9f174ea1c_1376x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!PPBH!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdf3d65d1-5367-4c7f-87f1-c8b9f174ea1c_1376x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!PPBH!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdf3d65d1-5367-4c7f-87f1-c8b9f174ea1c_1376x768.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Today's Run Data Run post, <a href="https://rundatarun.io/p/the-loop-is-simpler-than-it-sounds">The Loop Is Simpler Than It Sounds</a>, rests on a 30-day sweep across every platform where people argue about AI. This is the companion: the raw shape of what came back, including the parts that did not fit the headline. If the post is the argument, this is the receipts.</p><p>The single most useful thing the sweep found was not a fact. It was a contradiction. The loudest claim of the month, that "loop engineering" is the new moat, sits directly on top of an older, quieter body of evidence saying the exact opposite: that loops fail in ways the hype never mentions, and the people who learned that learned it expensively. Both are real. A single headline would have to pick one and lie about the other.</p><h2>Why this topic deserves a sweep</h2><p>In June, "the loop" went from a clever bash trick to the most over-discussed idea in the field. Anthropic shipped it as a product feature. A conference talk produced a slogan, "the winners will not have the smartest model, they will have the best loop," that got repeated everywhere. Boris Cherny, who built Claude Code, said he has not hand-written code in eight months and now spends his time writing loops, some days managing thousands of agents.</p><p>That is one half of the story, and if you only read X you would think it was the whole thing. The sweep is what surfaces the other half: a counter-current nearly as strong, mostly predating the viral wave, made of cost blowups, reliability math, and production failures. You cannot read both halves and conclude one thing. That is exactly when a sweep beats a headline.</p><h2>The moat thesis, as it actually spread</h2><p>The viral layer traces to one talk. The highest-engagement post of the set put the slogan plainly (<a href="https://x.com/AnatoliKopadze">@AnatoliKopadze</a>, 8,624 likes, 1.7M views), describing Anthropic running "three agents: one to plan, one to build, one to judge, cycling until the app actually works." The clearest single statement of the mechanic came from a practitioner, <a href="https://x.com/akshay_pachaar">Akshay Pachaar</a>: "the loop itself is six lines, and nobody competes on it. every serious agent framework lands on the same tiny while-loop."</p><p>The technique has a name and an origin: the "Ralph" loop, <a href="https://ghuntley.com/loop">coined by Geoffrey Huntley</a> after Ralph Wiggum, because it looks too dumb to work. And it has a button now, Claude Code's <a href="https://code.claude.com/docs/en/goal">native <code>/loop</code> and <code>/goal</code></a> commands, which is usually the signal that something has gone from clever to standard.</p><h2>The one stat I refused to print</h2><p>Here is a transparency note worth making, because it is the kind of thing a sweep catches and a single source does not.</p><p>The most-quoted number in the whole topic, Anthropic's internal loop-adoption figure, does not survive the sweep. Three different posts attribute three different numbers to what appears to be the same talk: "over 30% of code" written through loops, then "70 to 80% of engineers" using them, then "90%." When one stat arrives at three sizes from one source, it is folklore, not data. So the post states the direction (Anthropic is clearly building this way, at scale, on the record) and explicitly tells you to treat the decimal point as noise. That call only gets made because the sweep put the three versions side by side.</p><h2>The counter-current the slogan skips</h2><p>This is the densest part of what came back, and the part the hype leaves out.</p><ul><li><p><strong>The math.</strong> <a href="https://temporal.io/blog">Melanie Warrick at Temporal</a> ran the numbers: even at 85% reliability per step, a ten-step workflow finishes correctly only about 20% of the time. Failures multiply. She found the gap is resilience, not intelligence; agents "diagnose a failure in perfect detail and still do nothing to recover from it."</p></li><li><p><strong>The cost.</strong> An arXiv study found the same task, same model, same prompt, could cost eight dollars or two hundred forty, thirty times apart with no input change. A separate $22,000 sweep pinned a 33-fold spread to scaffold choices. And <a href="https://briefs.co">one widely-shared account</a> had Uber burning its annual AI coding budget in four months before capping engineers at $1,500 a head.</p></li><li><p><strong>The verification trap.</strong> A benchmark of agentic systems spent roughly $40,000 on evaluation runs alone, and noted that doing the checks rigorously would push it into the hundreds of thousands. If the whole premise of a loop is a verifiable done-condition, and verifying "done" at scale is the single most expensive line, that is a price nobody put in the pitch.</p></li><li><p><strong>The adoption gap.</strong> By one survey, 79% of enterprises have adopted agents and only 11% run them in production. Gartner projects 40% of agentic AI projects will be canceled by 2027, on cost, value, and governance, not on the loop failing.</p></li></ul><p><a href="https://garymarcus.substack.com">Gary Marcus</a> and <a href="https://x.com/arpit_bhayani">Arpit Bhayani</a> carried the skeptic load on the social layer, both before the June wave crested. The reliability and cost objections were live before the moat framing peaked, which is itself a signal.</p><h2>How the sweep is built (the method)</h2><p>A quick word on method, since the point of this section is to show the work.</p><p>The sweep pulls each platform on its own terms: Hacker News and Reddit for builder substance in the comments, X for engagement-weighted reach, a curated corpus and a Substack named-voice pool for depth, and a separate open-ended discovery pass whose only instruction is "surprise me, do not confirm the thesis." That last pass is what surfaced the counter-current; left to its own framing, a research pass confirms what you already believe.</p><p>Two integrity habits matter. Every cited cost figure traces to a primary source, not a tweet about a source. And the fetch pass flags its own failures: of the URLs pulled this round, several came back as bot-challenge pages or 404s and were routed around rather than quoted. A claim that only survives in one blocked link is not a claim.</p><h2>What I'm watching next</h2><ul><li><p>Whether the Anthropic adoption number ever gets pinned to a primary transcript, or stays folklore.</p></li><li><p>Where exactly the "coherence cliff" bites, by task length, token count, or iteration count. Everyone asserts it; nobody has located it.</p></li><li><p>Whether the three-agent plan/build/judge pattern survives the move from greenfield demos to brownfield maintenance, where the stale-state failures are worse.</p></li><li><p>The shift the researchers are naming: the agent stops being a tool that runs your code and becomes the software, with the human as "intent architect." If that holds, the unit of cost moves from per-token to per-finished-artifact, and most budgeting models break.</p></li></ul><h2>Sources and method</h2><p>Full synthesis and the per-platform raw dumps live in my research vault; the post, <a href="https://rundatarun.io/p/the-loop-is-simpler-than-it-sounds">The Loop Is Simpler Than It Sounds</a>, is the argument this sweep grounds. Method: a 30-day cross-platform sweep (Reddit, X, YouTube, Hacker News, the web, a curated corpus, a named-voice pool) plus an open-ended discovery pass, two-pass synthesis, every quantitative claim traced to a primary source. Counts and quotes come from the actual sweep output, not memory.</p><div><hr></div><p><em>Last 30 Days is a research series on Run Data Run, posted alongside the occasional Deep Dive when a topic earns the deeper look. No email on these, they live on the site for when you want the receipts behind an argument. If a sweep was useful, subscribe and you'll get the next one as it lands.</em></p>]]></content:encoded></item><item><title><![CDATA[Last 30 Days: The AI Scientists]]></title><description><![CDATA[Four autonomous discovery systems cleared peer review in four months. Every one of them was already a year old. Here is the full sweep, including the reliability research nobody is reading.]]></description><link>https://rundatarun.io/p/last-30-days-the-ai-scientists</link><guid isPermaLink="false">https://rundatarun.io/p/last-30-days-the-ai-scientists</guid><dc:creator><![CDATA[Justin Johnson]]></dc:creator><pubDate>Sun, 12 Jul 2026 12:06:00 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/2cb78e7c-6e37-4789-8a27-aa27ae7e20fa_1376x768.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<blockquote><p><strong>About Last 30 Days.</strong> Cross-platform research sweeps on topics worth paying attention to. Every post pulls Reddit, X, YouTube, Hacker News, Polymarket, and the web from the last 30 days, then synthesizes what people are actually saying, building, and betting on. Topics get picked when the signal is high and the story is contradictory, when a single headline would lie about the shape of what's happening. Each post follows the same arc: one specific finding that earns the click, why the topic deserves a sweep right now, the themed synthesis with inline citations, and the follow-up threads worth watching next.</p></blockquote><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!RliG!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17d9688d-a3f8-4f4e-b471-bd53296efe32_1376x768.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!RliG!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17d9688d-a3f8-4f4e-b471-bd53296efe32_1376x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!RliG!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17d9688d-a3f8-4f4e-b471-bd53296efe32_1376x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!RliG!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17d9688d-a3f8-4f4e-b471-bd53296efe32_1376x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!RliG!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17d9688d-a3f8-4f4e-b471-bd53296efe32_1376x768.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!RliG!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17d9688d-a3f8-4f4e-b471-bd53296efe32_1376x768.jpeg" width="728" height="409.5" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/17d9688d-a3f8-4f4e-b471-bd53296efe32_1376x768.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:false,&quot;imageSize&quot;:&quot;normal&quot;,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:728,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;captionedImage&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!RliG!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17d9688d-a3f8-4f4e-b471-bd53296efe32_1376x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!RliG!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17d9688d-a3f8-4f4e-b471-bd53296efe32_1376x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!RliG!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17d9688d-a3f8-4f4e-b471-bd53296efe32_1376x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!RliG!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17d9688d-a3f8-4f4e-b471-bd53296efe32_1376x768.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>The hook</h2><p>A prototype of <strong>Biomni</strong>, Stanford's autonomous biomedical research agent, was already running in <strong>more than 10,000 laboratories</strong> before its paper appeared in <em>Science</em> on 9 July.</p><p>Hold that number next to the thing everyone actually celebrated this month, which was the paper.</p><p>That gap, between when these systems became real and when the record admitted it, turned out to be the shape of this entire sweep.</p><h2>Why this topic deserves a sweep right now</h2><p><a href="https://www.nature.com/articles/s41586-026-10652-y">FutureHouse's Robin reached *Nature* this week</a>. It is a genuinely significant result: given the name of a disease, it proposed a therapeutic strategy for dry age-related macular degeneration, identified <strong>ripasudil</strong> (a glaucoma drug never proposed for dAMD), confirmed it in cells, and then designed and analyzed its own follow-up experiment. I wrote about it properly in <a href="https://rundatarun.io/p/the-year-nature-caught-up">this week's Sunday Deep Dive</a>.</p><p>Reading it sent me back to survey the whole category, because something felt off about the timing. It was. Robin's preprint went up in <strong>May 2025</strong>. The paper printed in <strong>July 2026</strong>.</p><p>So I ran a thirty-day sweep across X, Hacker News, the curated wire, and the web. What came back was not a story about one paper. It was a story about a field that has already moved somewhere the published record has not caught up to, and about a body of skeptical research that landed in the same window and got almost no attention at all.</p><div><hr></div><h2>The sweep</h2><h3>Theme 1: they all landed at once</h3><p>Four flagship autonomous-discovery systems cleared peer review inside a single four-month window.</p><ul><li><p><strong>[Robin](https://www.nature.com/articles/s41586-026-10652-y)</strong><a href="https://www.nature.com/articles/s41586-026-10652-y">Robin</a>** (FutureHouse) &#8594; <em><strong>Nature</strong></em>, July 2026. Multi-agent, lab-in-the-loop. Three specialised agents: <strong>Crow</strong> (concise literature search), <strong>Falcon</strong> (deep literature search), <strong>Finch</strong> (data analysis). Found ripasudil and KL001 for dry AMD, then surfaced <em>ABCA1</em> as a possible novel target via its own RNA-seq follow-up. From the abstract: <em>"All hypotheses, experimental directions, data analyses and data figures in the main text of this report were produced by Robin."</em></p></li><li><p><strong>[Google DeepMind's AI co-scientist](https://research.google/blog/accelerating-scientific-breakthroughs-with-an-ai-co-scientist/)</strong><a href="https://research.google/blog/accelerating-scientific-breakthroughs-with-an-ai-co-scientist/">Google DeepMind's AI co-scientist</a>** &#8594; <em><strong>Nature</strong></em>, 19 May 2026. Gemini-based multi-agent system. Its headline validation is <strong>the same move Robin made</strong>: read the literature, propose an old drug for a new disease. Theirs was <strong>KIRA6</strong> for acute myeloid leukemia, which inhibited AML cell viability at clinically relevant concentrations. Also went after liver fibrosis targets and antimicrobial resistance.</p></li><li><p><strong>[Sakana's AI Scientist](https://sakana.ai/ai-scientist-nature/)</strong><a href="https://sakana.ai/ai-scientist-nature/">Sakana's AI Scientist</a>** &#8594; <em><strong>Nature</strong></em>, 25 March 2026. The most radical of the four and the least grounded in wet biology: it runs the whole pipeline through to a finished manuscript, then peer-reviews itself. A paper it generated <strong>passed the first round of human peer review</strong> at a top machine-learning workshop. Its most interesting finding is a <strong>scaling law of AI science</strong>: generated-paper quality rises with the underlying model and with inference-time compute.</p></li><li><p><strong>[Biomni](https://news.stanford.edu/stories/2026/07/biomni-ai-powered-biomedical-co-scientist)</strong><a href="https://news.stanford.edu/stories/2026/07/biomni-ai-powered-biomedical-co-scientist">Biomni</a>** (Stanford) &#8594; <em><strong>Science</strong></em>, 9 July 2026. "Autonomous biomedical research with an artificial intelligence agent." Generalises across causal gene prioritisation, drug repurposing, rare-disease diagnosis, microbiome analysis and molecular cloning with no task-specific tuning. Already in 10,000+ labs.</p></li></ul><p>The cross-system pattern was <a href="https://x.com/aipoch_ai/status/2072591273983455480">spotted in the wild</a> and it is worth quoting, because it is the whole architectural story in one line: <em>"Across Claude Science, NVIDIA BioNeMo, and FutureHouse Robin, the same pattern keeps appearing: specialized agent skills, orchestrated research workflows, domain-specific execution."</em></p><p><strong>Nobody trained a discovery model. Everybody built a harness.</strong></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!m6km!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82fbb420-8387-4cec-b9d1-0a8dcf16482c_1408x768.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!m6km!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82fbb420-8387-4cec-b9d1-0a8dcf16482c_1408x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!m6km!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82fbb420-8387-4cec-b9d1-0a8dcf16482c_1408x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!m6km!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82fbb420-8387-4cec-b9d1-0a8dcf16482c_1408x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!m6km!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82fbb420-8387-4cec-b9d1-0a8dcf16482c_1408x768.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!m6km!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82fbb420-8387-4cec-b9d1-0a8dcf16482c_1408x768.jpeg" width="728" height="409.5" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/82fbb420-8387-4cec-b9d1-0a8dcf16482c_1408x768.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:false,&quot;imageSize&quot;:&quot;normal&quot;,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:728,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;captionedImage&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!m6km!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82fbb420-8387-4cec-b9d1-0a8dcf16482c_1408x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!m6km!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82fbb420-8387-4cec-b9d1-0a8dcf16482c_1408x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!m6km!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82fbb420-8387-4cec-b9d1-0a8dcf16482c_1408x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!m6km!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82fbb420-8387-4cec-b9d1-0a8dcf16482c_1408x768.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h3>Theme 2: the lag is the actual story</h3><p>Every one of those systems was old news by the time it was printed.</p><ul><li><p><strong>Robin</strong>: preprint <a href="https://arxiv.org/abs/2505.13400">19 May 2025</a> &#8594; <em>Nature</em> July 2026. <strong>14 months.</strong></p></li><li><p><strong>Sakana's AI Scientist</strong>: arXiv <a href="https://arxiv.org/abs/2408.06292">12 August 2024</a> &#8594; <em>Nature</em> 25 March 2026. <strong>19 months.</strong></p></li><li><p><strong>Google's co-scientist</strong>: announced 19 February 2025 &#8594; <em>Nature</em> 19 May 2026. <strong>15 months, to the day.</strong></p></li><li><p><strong>Biomni</strong>: bioRxiv 30 May 2025 &#8594; <em>Science</em> 9 July 2026. <strong>13 months.</strong></p></li></ul><p>Mean lag: <strong>roughly fifteen months.</strong></p><p>The sharpest write-up of this belongs to synthetic biologist <a href="https://x.com/SynBio1/status/2075567928909467801">Jake Wintermute</a>, on Biomni:</p><blockquote><p><strong>"Biomni was on arXiv 13 months ago. Biomni was on GitHub 11 months ago. Phylo, the company built on Biomni, raised $13.5M and launched 5 months ago. Or, I guess, you could read about it in Science Magazine today."</strong></p></blockquote><p>None of this is a criticism of the labs. They all posted preprints immediately, which is exactly right. The lag belongs to the journals.</p><h3>Theme 3: where the field actually is</h3><p>Not where <em>Nature</em> says it is. While Robin was in review, FutureHouse built its successor and then built a company around it.</p><ul><li><p><strong>[Kosmos](https://arxiv.org/abs/2511.02824)</strong><a href="https://arxiv.org/abs/2511.02824">Kosmos</a>** (arXiv, November 2025): runs up to <strong>12 hours</strong> across ~20 cycles. A single run reads <strong>~1,500 papers</strong> and writes <strong>~42,000 lines of code</strong>. Independent scientists judged <strong>79.4%</strong> of the statements in its reports accurate. It has produced <strong>seven discoveries</strong>, three reproducing unpublished findings and four net-new. Collaborators estimated one run did about <strong>six months</strong> of their own work. Sam Rodriques' announcement did <strong>3,683 likes</strong>, the highest-engagement item in the entire sweep.</p></li><li><p><strong>Edison Scientific</strong>: FutureHouse's <strong>for-profit spinout</strong>, commercialising Kosmos. And <a href="https://x.com/SGRodriques/status/2071616647820177564">this month</a>, its first partnership using Kosmos not to write papers but to <strong>launch new biotech companies</strong>.</p></li><li><p><strong>[Sakana Marlin](https://x.com/hardmaru/status/2066529282588094713)</strong><a href="https://x.com/hardmaru/status/2066529282588094713">Sakana Marlin</a>**: Sakana's first commercial product. An autonomous research agent that runs ~<strong>8 hours</strong> unattended and emits structured slides plus a multi-dozen-page report. Pitched as a <strong>virtual chief strategy officer</strong>, aimed at finance and consulting rather than the bench.</p></li><li><p><strong>[Claude Science](https://www.anthropic.com/news/claude-science-ai-workbench)</strong><a href="https://www.anthropic.com/news/claude-science-ai-workbench">Claude Science</a>** (Anthropic, 30 June): the workbench end of the same idea. I wrote about it <a href="https://rundatarun.io/p/claude-science-and-the-boring-80">here</a>.</p></li></ul><p><strong>Read the journals to learn what was true a year ago. Read the preprints and the launches to learn what is true now.</strong></p><h3>Theme 4: the counterweight nobody is reading</h3><p>This is the highest-value material in the sweep and it appeared in almost no coverage.</p><p><strong>[Correct Answer, Wrong Mechanism](https://arxiv.org/pdf/2606.23175)</strong><a href="https://arxiv.org/pdf/2606.23175">Correct Answer, Wrong Mechanism</a>** (arXiv 2606.23175). Subtitle: <em>When AI Scientists Defend General Claims Their Own Data Contradicts.</em></p><p>Researchers watched a coding agent attempt to rediscover a known particle-physics result across <strong>28 episodes</strong>. It often got the right answer. The problem was <em>how</em>:</p><ul><li><p>CAWM occurred in <strong>4 of 20 (20%)</strong> primary-model episodes and <strong>3 of 8 (37.5%)</strong> episodes on other frontier models.</p></li><li><p>The agent reached right-looking results through reasoning that <strong>collapses when conditions change</strong>.</p></li><li><p>When pressed, it <strong>defended the wrong mechanism</strong>, arguing for physics inconsistent with the numbers in its own output.</p></li><li><p>Verdict: these systems are dependable as <strong>tools</strong> but <em>"unreliable scientific co-authors for open-ended claim-making."</em></p></li><li><p>The demand: <strong>outcome-only evaluation is insufficient.</strong> Score task outcome, mechanism fidelity, and epistemic honesty as three separate things. Their lightweight checks flagged <strong>every</strong> CAWM case in the study.</p></li></ul><p><strong>[Related work](https://phys.org/news/2026-05-ai-scientists-reveal-fundamental-limits.html)</strong><a href="https://phys.org/news/2026-05-ai-scientists-reveal-fundamental-limits.html">Related work</a>** (May 2026) catalogues the failure modes of autonomous research pipelines with unnerving specificity: implementation bugs, hallucinated results, shortcut reliance, methodology fabrication, citation hallucination, <strong>frame-lock</strong>, and <strong>bug-as-insight reframing</strong>, in which a system trips over a defect in its own code and writes it up as a discovery. The finding: increased <em>quantity</em>, decreased <em>quality</em>, of both papers and reviews.</p><p>And the view from the bench. Working scientist <a href="https://x.com/chorye/status/2075994339935723670">Emma Chory</a>, on Biomni cheerfully producing a <strong>BSL2-plus lentivirus protocol</strong> with, in her words, "sufficient detail to execute on a robot." Her review ran to four words: <em>"Cool cool cool cool."</em> An agent that will write you a biosafety-level-2-plus procedure on request is a governance question, not a capability win.</p><h3>Theme 5: the taxonomy worth stealing</h3><p>Computational biologist <a href="https://x.com/msikic/status/2075884565680259113">Mile Sikic</a> drew the distinction that most of the coverage blurs. There are <strong>two streams</strong>:</p><ol><li><p><strong>Virtual bioinformaticians</strong> that combine existing knowledge with existing tools. (Claude Science, Biomni, most of the field.)</p></li><li><p><strong>Systems that discover new biology in close collaboration with wet-lab scientists.</strong> (Robin, Google's co-scientist.)</p></li></ol><p>Conflating them is how people end up disappointed. They are solving different problems.</p><p>Also in the window, at lower weight: <strong>[EurekAgent](http://arxiv.org/abs/2606.13662)</strong><a href="http://arxiv.org/abs/2606.13662">EurekAgent</a>** ("Agent Environment Engineering is All You Need for Autonomous Scientific Discovery"); <a href="https://x.com/hxiao/status/2075750229748715784">Yoshua Bengio's take on an AI Scientist</a>; the Samsung SAIT AI Scientist competition winners; <strong>OpenScience</strong>, an open coding agent "that went to grad school," ~350 GitHub stars in its first week; and a genuinely useful sign of a crowded category, a piece titled <a href="https://x.com/BioAI_NeuralNet/status/2075940077885153423">"Which 'AI scientist' suits your lab? A guide for the perplexed."</a></p><div><hr></div><h2>What I'm watching next</h2><p><strong>Edison Scientific's biotech-founding partnership.</strong> Using an AI scientist to <em>found companies</em> is a categorically different claim from using one to write a paper, and unlike a paper it will be tested in public, on a clock, with other people's money.</p><p><strong>Whether anyone adopts the CAWM evaluation protocol.</strong> The paper's demand is concrete and cheap: score mechanism fidelity separately from outcome. If the next generation of agent papers still reports only "did it get the right answer," the field has decided not to look.</p><p><strong>Ripasudil.</strong> It cleared cells, not patients. A disease model comes next, and then, eventually, a randomized controlled trial. Most compounds that look this good at this stage never finish.</p><div><hr></div><h2>Sources</h2><p>Primary papers: <a href="https://www.nature.com/articles/s41586-026-10652-y">Robin / *Nature*</a> &#183; <a href="https://arxiv.org/abs/2505.13400">Robin preprint</a> &#183; <a href="https://arxiv.org/abs/2511.02824">Kosmos</a> &#183; <a href="https://arxiv.org/abs/2408.06292">Sakana AI Scientist</a> &#183; <a href="https://arxiv.org/pdf/2606.23175">Correct Answer, Wrong Mechanism</a> &#183; <a href="http://arxiv.org/abs/2606.13662">EurekAgent</a> &#183; <a href="https://news.stanford.edu/stories/2026/07/biomni-ai-powered-biomedical-co-scientist">Biomni / Stanford</a> &#183; <a href="https://research.google/blog/accelerating-scientific-breakthroughs-with-an-ai-co-scientist/">Google AI co-scientist</a> &#183; <a href="https://www.anthropic.com/news/claude-science-ai-workbench">Claude Science</a></p><p>Voices: <a href="https://x.com/SGRodriques/status/2075585708950221126">@SGRodriques</a> &#183; <a href="https://x.com/FutureHouseSF/status/2056813047180931316">@FutureHouseSF</a> &#183; <a href="https://x.com/SynBio1/status/2075567928909467801">@SynBio1</a> &#183; <a href="https://x.com/chorye/status/2075994339935723670">@chorye</a> &#183; <a href="https://x.com/msikic/status/2075884565680259113">@msikic</a> &#183; <a href="https://x.com/hardmaru/status/2066529282588094713">@hardmaru</a> &#183; <a href="https://x.com/aipoch_ai/status/2072591273983455480">@aipoch_ai</a></p><p><em>Methodology: thirty-day sweep across X, Hacker News, the curated AI wire, and the web, run 12 July 2026. The X leg of my usual tooling was down (an expired API credential), so X was swept through a working session-based client instead; Reddit returned no topical signal and was discarded rather than padded. Engagement figures are as recorded at collection time. The full deep dive on the Robin paper itself is [here](https://rundatarun.io/p/the-year-nature-caught-up).</em><a href="https://rundatarun.io/p/the-year-nature-caught-up">here</a>.*</p>]]></content:encoded></item><item><title><![CDATA[The Year Nature Caught Up]]></title><description><![CDATA[Robin found a new use for an old glaucoma drug and landed in Nature. The preprint was fourteen months old. Once you notice that lag, you notice it everywhere.]]></description><link>https://rundatarun.io/p/the-year-nature-caught-up</link><guid isPermaLink="false">https://rundatarun.io/p/the-year-nature-caught-up</guid><dc:creator><![CDATA[Justin Johnson]]></dc:creator><pubDate>Sun, 12 Jul 2026 12:05:04 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/8c1fcb31-7a67-482c-95aa-1015b4765a03_1584x672.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em>Every Sunday I pick one paper or release that's worth your time, break it apart, and tell you why it matters. No hype. No summaries of summaries. Just the idea, explained.</em></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!VO1V!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F821ce5de-6580-45f9-8280-270dd283c0b8_1584x672.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!VO1V!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F821ce5de-6580-45f9-8280-270dd283c0b8_1584x672.jpeg 424w, https://substackcdn.com/image/fetch/$s_!VO1V!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F821ce5de-6580-45f9-8280-270dd283c0b8_1584x672.jpeg 848w, https://substackcdn.com/image/fetch/$s_!VO1V!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F821ce5de-6580-45f9-8280-270dd283c0b8_1584x672.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!VO1V!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F821ce5de-6580-45f9-8280-270dd283c0b8_1584x672.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!VO1V!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F821ce5de-6580-45f9-8280-270dd283c0b8_1584x672.jpeg" width="728" height="409.5" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/821ce5de-6580-45f9-8280-270dd283c0b8_1584x672.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:false,&quot;imageSize&quot;:&quot;normal&quot;,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:728,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;captionedImage&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!VO1V!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F821ce5de-6580-45f9-8280-270dd283c0b8_1584x672.jpeg 424w, https://substackcdn.com/image/fetch/$s_!VO1V!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F821ce5de-6580-45f9-8280-270dd283c0b8_1584x672.jpeg 848w, https://substackcdn.com/image/fetch/$s_!VO1V!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F821ce5de-6580-45f9-8280-270dd283c0b8_1584x672.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!VO1V!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F821ce5de-6580-45f9-8280-270dd283c0b8_1584x672.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div><hr></div><h2>The Headline</h2><p>A multi-agent AI system named <strong>Robin</strong> was given the name of a disease and told to find a treatment. It read the literature, proposed a therapeutic strategy, named a drug, watched humans test that drug at the bench, analyzed the results itself, and then designed its own follow-up experiment. The disease was dry age-related macular degeneration, the major cause of blindness in the developed world. The drug was <strong>ripasudil</strong>, a rho-kinase inhibitor that has been sitting in pharmacies for years, approved for glaucoma, and never proposed for dAMD by anyone, as far as the authors could tell or I could find.</p><p>It worked in cells. Then Robin asked why, specified an RNA-sequencing experiment to find out, analyzed that too, and surfaced a gene called <em>ABCA1</em> as a possible new target nobody had been looking at.</p><p>This is the work of <a href="https://www.futurehouse.org/">FutureHouse</a>, and <a href="https://www.nature.com/articles/s41586-026-10652-y">it was published in Nature this week</a>. There is one sentence in the paper that deserves a second read:</p><blockquote><p><strong>"All hypotheses, experimental directions, data analyses and data figures in the main text of this report were produced by Robin."</strong></p></blockquote><p>Not assisted by. Produced by.</p><p><a href="https://arxiv.org/abs/2505.13400">The preprint went up in May 2025</a>. Nature published it in July 2026. Fourteen months, and almost none of the coverage mentions it.</p><p>Once you notice that gap, you start seeing it everywhere, and the shape of this entire field changes.</p><div><hr></div><h2>The Paper</h2><h3>Three birds and an orchestrator</h3><p>Robin is not a model. Nobody at FutureHouse trained a "discovery model," and if you go shopping for one after reading the coverage, you will not find it.</p><p>What they built is a system of specialized language agents sitting on top of ordinary commercial models:</p><ul><li><p><strong>Crow</strong> runs fast, concise literature searches.</p></li><li><p><strong>Falcon</strong> runs deep ones.</p></li><li><p><strong>Finch</strong> analyzes raw experimental data.</p></li></ul><p>Robin orchestrates the three of them and carries the hypothesis across rounds, so that what Finch learns on Tuesday reshapes what Falcon goes looking for on Wednesday.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Lu5o!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F84416561-2a93-4607-b01d-396da180ef3f_1408x768.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Lu5o!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F84416561-2a93-4607-b01d-396da180ef3f_1408x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!Lu5o!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F84416561-2a93-4607-b01d-396da180ef3f_1408x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!Lu5o!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F84416561-2a93-4607-b01d-396da180ef3f_1408x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!Lu5o!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F84416561-2a93-4607-b01d-396da180ef3f_1408x768.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Lu5o!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F84416561-2a93-4607-b01d-396da180ef3f_1408x768.jpeg" width="728" height="409.5" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/84416561-2a93-4607-b01d-396da180ef3f_1408x768.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:false,&quot;imageSize&quot;:&quot;normal&quot;,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:728,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;captionedImage&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!Lu5o!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F84416561-2a93-4607-b01d-396da180ef3f_1408x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!Lu5o!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F84416561-2a93-4607-b01d-396da180ef3f_1408x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!Lu5o!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F84416561-2a93-4607-b01d-396da180ef3f_1408x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!Lu5o!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F84416561-2a93-4607-b01d-396da180ef3f_1408x768.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><blockquote><p><strong>The intelligence came off a price sheet. The discovery came from the scaffolding around it.</strong></p></blockquote><p>I wrote two weeks ago that <a href="https://rundatarun.io/p/claude-science-and-the-boring-80">Anthropic shipped a workbench, not a miracle</a>, and that <a href="https://rundatarun.io/p/the-harness-is-the-moat">the harness is the moat</a>. Robin is that argument with a Nature paper attached to it, which is a considerably stronger position than my say-so.</p><h3>The word "semi" is doing honest work</h3><p>The paper describes Robin as <strong>semi-autonomous</strong>, and the precision is a credit to the authors rather than a hedge.</p><p>Humans still run the wet lab. People cultured the retinal cells, ran the assays, and pipetted the compounds. What Robin did was everything on either side of the bench: read the field, form the hypothesis, specify the experiment, interpret what came back, and revise.</p><p>The term of art is <strong>lab-in-the-loop</strong>, and it is worth flipping around, because the phrase misleads people. The AI is not in the lab. The lab is in the AI's loop.</p><p>Previous systems could do one arc of that circuit. They could read and propose. Or they could take your data and analyze it. Robin ran the full turn, then fed its own results into the next turn, and did it again.</p><p>The paper calls Robin "one of the first" systems to do this. The preprint called it "the first." Somewhere in fourteen months of review a superlative got sanded off, which tells you something useful about the claim and something more interesting about the process.</p><h3>Why the drug was findable at all</h3><p>FutureHouse call their method <strong>combinatorial synthesis</strong>, and their honesty about what it means is why I trust the rest of the paper.</p><p>Robin did not invent new biology. Every piece was already in print. That rho-kinase inhibition boosts phagocytosis in retinal pigment epithelium cells: known, and they cite the work. That this phagocytic housekeeping declines in AMD patients: also known. The two facts lived in different literatures, read by different people, and nobody had put them next to each other and said the word <em>ripasudil</em>.</p><p>The authors have a devastating example of how long that gap can persist. <strong>Dabrafenib</strong> is a cancer drug whose molecular action was characterized by 2010. Ten years later, a brute-force screen discovered it protects against hearing loss. That protective effect follows <em>directly</em> from the mechanism everyone already knew. The answer sat in print for a decade because the people who knew about BRAF inhibition and the people who cared about hearing loss were not the same people.</p><blockquote><p><strong>Robin is not doing science we could not do. It is doing science we did not get around to.</strong></p></blockquote><p>Scope the claim that way and it stays large. As the authors point out, novel FDA approvals have been flat at roughly fifty a year for a decade. If a machine can reliably close the distance between what is <em>known</em> and what is <em>connected</em>, that is worth a great deal of money and a great deal of eyesight.</p><h3>The two things I would copy tomorrow</h3><p><strong>They checked their own judge.</strong> Buried in the supplement is a comparison of their LLM evaluator against human experts. I read a lot of agent papers that grade themselves with a model and never once ask whether the grader is any good. This one asked.</p><p><strong>Their guardrails are the most thought-through of anything I read this month.</strong> Robin preferentially proposes compounds with established safety profiles. Every output is treated as a hypothesis entering standard preclinical review, never as a finding. And the discussion says the quiet part plainly: ripasudil "would of course require validation in a suitable disease model and ultimately in a randomized, placebo-controlled trial."</p><p>In vitro is a beginning. The authors know it, and they wrote it down.</p><div><hr></div><h2>The Ecosystem</h2><p>Robin made me want to go back and re-survey this whole category, so I did. Thirty days of papers, launches, funding, and argument. What I found reframed the paper I had just finished reading.</p><h3>Everyone landed at once</h3><p>Four flagship systems cleared peer review inside about four months.</p><p><strong>Google DeepMind's AI co-scientist</strong> <a href="https://research.google/blog/accelerating-scientific-breakthroughs-with-an-ai-co-scientist/">reached Nature in May</a>. It is a Gemini-based multi-agent system, and its headline validation is <em>the same move Robin made</em>: read the literature, propose an old drug for a new disease. Theirs was <strong>KIRA6</strong> for acute myeloid leukemia, which inhibited the viability of AML cells at clinically relevant concentrations. They also went after liver fibrosis targets and antimicrobial resistance.</p><p><strong>Sakana's AI Scientist</strong> <a href="https://sakana.ai/ai-scientist-nature/">reached Nature on 25 March</a>. It is more radical in ambition and less grounded in wet biology: it runs the entire pipeline through to a finished manuscript and then peer-reviews itself. A paper it generated <strong>passed the first round of human peer review</strong> at a top machine-learning workshop. Its most interesting result is a <em>scaling law of AI science</em>: the quality of the papers rises with the quality of the underlying model and with the compute you spend at inference.</p><p><strong>Biomni</strong>, out of Stanford, reached <strong>Science</strong> on 9 July, under the title "Autonomous biomedical research with an artificial intelligence agent." Hold on to one detail from it, because it is the most telling number in this entire piece: <strong>a prototype of Biomni was already running in more than 10,000 labs before the paper printed.</strong></p><h3>The lag is the actual story</h3><p>Every one of those systems was old news by the time it was printed.</p><p>Robin: preprint May 2025, Nature July 2026. Fourteen months. Sakana's AI Scientist went up on arXiv in August 2024 and reached Nature in March 2026. Nineteen months. Google announced its co-scientist in February 2025 and printed it in May 2026. Fifteen months, to the day. And Biomni got the sharpest write-up it will ever receive, from <a href="https://x.com/SynBio1/status/2075567928909467801">a synthetic biologist on X</a>:</p><blockquote><p><strong>"Biomni was on arXiv 13 months ago. Biomni was on GitHub 11 months ago. Phylo, the company built on Biomni, raised $13.5M and launched 5 months ago. Or, I guess, you could read about it in Science Magazine today."</strong></p></blockquote><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!VTcm!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30091134-6d50-4a1b-9285-f5fdde02a8e6_1408x768.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!VTcm!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30091134-6d50-4a1b-9285-f5fdde02a8e6_1408x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!VTcm!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30091134-6d50-4a1b-9285-f5fdde02a8e6_1408x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!VTcm!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30091134-6d50-4a1b-9285-f5fdde02a8e6_1408x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!VTcm!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30091134-6d50-4a1b-9285-f5fdde02a8e6_1408x768.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!VTcm!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30091134-6d50-4a1b-9285-f5fdde02a8e6_1408x768.jpeg" width="728" height="409.5" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/30091134-6d50-4a1b-9285-f5fdde02a8e6_1408x768.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:false,&quot;imageSize&quot;:&quot;normal&quot;,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:728,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;captionedImage&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!VTcm!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30091134-6d50-4a1b-9285-f5fdde02a8e6_1408x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!VTcm!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30091134-6d50-4a1b-9285-f5fdde02a8e6_1408x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!VTcm!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30091134-6d50-4a1b-9285-f5fdde02a8e6_1408x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!VTcm!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30091134-6d50-4a1b-9285-f5fdde02a8e6_1408x768.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>None of this is a criticism of the labs. FutureHouse posted their preprint the moment they had it, which is exactly right. The criticism, if there is one, belongs to the clock the journals run on.</p><p>So where is the field actually standing? Not where Nature says it is.</p><p>FutureHouse has already built Robin's successor. <strong>Kosmos</strong> went up on arXiv last November. It runs for twelve hours at a stretch, and a single run reads around 1,500 papers and writes roughly 42,000 lines of code. Independent scientists checked its reports and found <strong>79.4% of its statements accurate</strong>. It has produced seven discoveries, four of them net new. Collaborators estimated that one run did about six months of their own work.</p><p>They then spun out a for-profit company, <strong>Edison Scientific</strong>, to sell it. And this month they announced a partnership that uses Kosmos not to write papers but to <strong>found biotech companies</strong>.</p><p>Sakana, meanwhile, is selling <strong>Marlin</strong>, an autonomous research agent that runs for about eight hours unattended and is pitched as a virtual chief strategy officer.</p><blockquote><p><strong>Read Nature to learn what was true a year ago. Read the preprints and the launches to learn what is true now.</strong></p></blockquote><h3>The counterweight nobody is reading</h3><p>While these systems were collecting their journal stamps, a second body of work was landing that tells you precisely where they break. None of the launch threads mention it.</p><p>The sharpest of it is a paper whose title belongs on a poster in every lab now buying one of these things: <a href="https://arxiv.org/pdf/2606.23175">**Correct Answer, Wrong Mechanism**</a>. The subtitle is better. <em>When AI Scientists Defend General Claims Their Own Data Contradicts.</em></p><p>The researchers watched a coding agent try to rediscover a known result in particle physics, twenty-eight times over. Often it got the right answer. The problem was <em>how</em>. In <strong>20% of episodes</strong> with the primary model, and <strong>37.5% across other frontier models</strong>, the agent reached a right-looking result through reasoning that collapses the moment conditions change. Worse, when pressed, it <em>defended</em> the wrong mechanism, arguing for physics that contradicted the numbers in its own output.</p><p>Their verdict is careful, and damning. These systems are dependable as <strong>tools</strong> and, for now, "unreliable scientific co-authors for open-ended claim-making." Asking whether the agent got the answer right tells you almost nothing. You have to score the outcome, the fidelity of the mechanism, and the honesty of the account as three separate things.</p><p><a href="https://phys.org/news/2026-05-ai-scientists-reveal-fundamental-limits.html">Related work in the same window</a> catalogues the failure modes with unnerving specificity: hallucinated results, methodology fabrication, citation invention, <strong>frame-lock</strong>, and my favorite piece of vocabulary this year, <strong>bug-as-insight reframing</strong>, in which a system trips over a defect in its own code and writes it up as a discovery.</p><p>Set that next to Robin and its design choices stop looking conservative and start looking wise. When Robin says ripasudil enhances phagocytosis, a cell culture has already voted, and a cell culture does not care what the agent believes about it.</p><blockquote><p><strong>A wet lab is a brutal, incorruptible critic. You cannot argue a cell culture into agreeing with you.</strong></p></blockquote><p>But be precise about how far that protection reaches, because it is not the whole paper. When Robin says the mechanism runs through <em>ABCA1</em>, that is Finch reading an RNA-sequencing experiment, and that is exactly the surface the CAWM authors are worried about. <strong>The bench protects the finding. It does not protect the explanation.</strong> Which is roughly the division of trust I would apply to any of these systems, and to be fair to FutureHouse, it is the division they applied themselves: the drug candidate is offered for preclinical review, the mechanism is offered as a possibility.</p><p>It is also why the wet-biology systems stop at <strong>semi</strong>-autonomous, and why that ceiling is physical rather than technical. Laboratory instruments were not built to take orders from software. Somebody has to run the assay. Sakana's system is the exception that proves the rule, and it proves it uncomfortably: with no bench anywhere in the loop, it is the one system in this survey that will write its own paper <em>and</em> review it.</p><div><hr></div><h2>What I Run</h2><p>I should declare an interest, briefly, because I have spent the last year on the other side of this.</p><p>I run an autonomous research engine called <strong>ARIA</strong>. It is a persistent system rather than a one-shot agent: a pool of research ideas competing against each other on a single scoring function, experiments that run locally first and have to earn their way onto expensive hardware, a critic drawn from a different model family that posts a verdict on every result, and failures that classify themselves and re-enter the pool as recovery work. Seven instances, four domains, about 19,400 auditable commits, and 97.8% autonomous resumption after failure. Its successor now runs around the clock on <a href="https://rundatarun.io/p/borrowed-iron">a borrowed eight-GPU node</a>, pointed at retinal disease. There is <a href="https://www.justinhjohnson.com/case/aria">a case study</a> and <a href="https://youtu.be/UJyBEVFBRPs">a short film</a> if you want the shape of it.</p><p>I raise it for one reason. Building the thing taught me the same lesson the CAWM paper is now pressing on the whole field, and I learned it the expensive way.</p><blockquote><p><strong>Generating hypotheses was never the hard part. Building a critic honest enough to kill them was.</strong></p></blockquote><p>An idea generator is cheap, and it will happily run forever, and every dashboard will stay green while it does. What is difficult, and what almost nobody budgets for, is the machinery that tells the system it is wrong and makes the verdict stick.</p><p>Robin never had to build that machinery, because the bench already is it. That is a structural advantage of working in wet biology, and it is worth naming plainly for anyone whose research loop closes entirely in software, where nothing pushes back for free.</p><div><hr></div><h2>Why It Matters</h2><p><strong>Do not go shopping for a discovery model.</strong> There isn't one. Not one system in this survey used a model you cannot buy. Robin, the co-scientist, Kosmos, Biomni: every one of them is a harness built around a commodity brain, which means the thing your organization would be building is the harness, and that is engineering, not procurement.</p><p><strong>The compartmentalization gap is inside your own company.</strong> Robin's whole trick is finding the connection between two literatures that no single person reads. You have that problem internally, at smaller scale, right now. The result one team generated and another team needed and never saw is the cheapest discovery you will ever fund.</p><p><strong>Read the preprints.</strong> On this topic the peer-reviewed record is running roughly a year behind the work. Use the journals to confirm, not to track.</p><p><strong>Take the reliability research as seriously as the launch threads.</strong> A system that finds the right answer for the wrong reason will sail through your demo and drown in your clinic. Ask any vendor how they score mechanism fidelity, not just outcomes. Watch them field the question.</p><p>And watch ripasudil. It cleared cells, not patients. The road from here runs through a disease model and then a randomized controlled trial, and most compounds that look this good at this stage do not finish it.</p><p>Robin did not replace the scientist. It replaced the <strong>lag</strong>, the decade dabrafenib spent sitting in the literature while nobody put two known facts side by side.</p><p>There is a certain irony in that finding taking fourteen months to reach print. It does not make it less of a landmark. Congratulations to the FutureHouse team. The field is better for this being in the record, even if the record took its time.</p><div><hr></div><p><em>I ran a full thirty-day sweep of this category to write this piece, and most of it did not fit. The complete survey, with every system, number, and citation, is [in a companion post here](https://rundatarun.io/p/last-30-days-the-ai-scientists).</em><a href="https://rundatarun.io/p/last-30-days-the-ai-scientists">in a companion post here</a>.*</p><div><hr></div><p><em>Sunday Deep Dive is a weekly series on Run Data Run. Every Sunday I pick one paper, release, or technique worth understanding, break it apart, and tell you what it means for your work. Free every Sunday, no paywall. If it was useful, the easiest way to support it is to subscribe and forward it to one person on your team who'd want it. If it wasn't, tell me why. I'll make it better.</em></p>]]></content:encoded></item><item><title><![CDATA[How I Actually Stay Current on AI (I Built a Wire)]]></title><description><![CDATA[The question I get asked most, and the system that answers it.]]></description><link>https://rundatarun.io/p/how-i-actually-stay-current-on-ai</link><guid isPermaLink="false">https://rundatarun.io/p/how-i-actually-stay-current-on-ai</guid><dc:creator><![CDATA[Justin Johnson]]></dc:creator><pubDate>Thu, 09 Jul 2026 13:50:59 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/fbf81d6f-0b6b-46a5-bf61-75ed8fb681b4_1664x928.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!gsD8!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe58370ef-185a-4cff-9ce6-305c25b8d617_1664x928.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!gsD8!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe58370ef-185a-4cff-9ce6-305c25b8d617_1664x928.jpeg 424w, https://substackcdn.com/image/fetch/$s_!gsD8!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe58370ef-185a-4cff-9ce6-305c25b8d617_1664x928.jpeg 848w, https://substackcdn.com/image/fetch/$s_!gsD8!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe58370ef-185a-4cff-9ce6-305c25b8d617_1664x928.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!gsD8!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe58370ef-185a-4cff-9ce6-305c25b8d617_1664x928.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!gsD8!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe58370ef-185a-4cff-9ce6-305c25b8d617_1664x928.jpeg" width="728" height="409.5" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e58370ef-185a-4cff-9ce6-305c25b8d617_1664x928.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:false,&quot;imageSize&quot;:&quot;normal&quot;,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:728,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;captionedImage&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!gsD8!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe58370ef-185a-4cff-9ce6-305c25b8d617_1664x928.jpeg 424w, https://substackcdn.com/image/fetch/$s_!gsD8!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe58370ef-185a-4cff-9ce6-305c25b8d617_1664x928.jpeg 848w, https://substackcdn.com/image/fetch/$s_!gsD8!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe58370ef-185a-4cff-9ce6-305c25b8d617_1664x928.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!gsD8!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe58370ef-185a-4cff-9ce6-305c25b8d617_1664x928.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Some version of this question arrives most weeks. How do you keep up? You have a day job. The field moves every single day. What are you reading?</p><p>For a long time I answered it badly. I would say something about reading a lot, which is true and completely useless to the person asking. The real answer took me a while to say out loud, because it sounds like a dodge.</p><blockquote><p><strong>I stopped trying to keep up. I built something that keeps up for me.</strong></p></blockquote><div><hr></div><h2>Bookmarks are where links go to die</h2><p><strong>Nobody is reading this field.</strong> Not all of it, not close. Papers land overnight. A lab ships a model on a Tuesday and by Thursday other teams have published what breaks about it. The people who look like they are on top of it are not reading faster than you. <strong>They have narrowed, or they have automated, or they are bluffing.</strong></p><p>Manual curation was my first answer, and it failed the way manual curation always fails. Open tabs became a graveyard. Saved links went unread and, worse, went stale.</p><blockquote><p><strong>A link I saved in March, about a model superseded in April, is not information anymore. It is clutter with a timestamp.</strong></p></blockquote><p>Newsletters helped until I was subscribed to a stack of them and reading none. <strong>The problem was never that I lacked sources.</strong> I had too many, arriving on their own schedules, with no shared memory between them.</p><div><hr></div><h2>One pipe</h2><p>So I built a wire.</p><p>Several channels I already had running now point at one place. A scan that reads the day's news, releases, and research. The things I bookmark as I go. Findings from a small set of AI assistants I run, each pointed at its own beat. My own written breakdowns of the papers and tools and companies I sit down and work through by hand. Things I email myself at midnight and would otherwise never see again.</p><p>All of it lands in one store, <strong>scored and deduplicated</strong>. No single source has to be the right source, because none of them carries the day alone. And nothing gets read twice, because the store already knows it saw that story on three feeds this morning.</p><p>Two things come out the other end.</p><p><strong>The first is private</strong>, and it is the part I use every day without thinking about it. The store is a knowledge base that refreshes itself, and my own AI tools read from it directly. When I sit down to write, or research something, or work a problem, the context they pull from is current. Not their training data. Not a stale snapshot I remembered to update. This morning's.</p><p><strong>The second is public.</strong> A curated slice of the store becomes a website: a daily dispatch of the items that scored highest, written up rather than merely linked, so you know why something mattered before you decide to click it.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!qc54!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3b2132d0-a0a2-4814-aa2c-1f7e69cccc93_1376x768.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!qc54!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3b2132d0-a0a2-4814-aa2c-1f7e69cccc93_1376x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!qc54!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3b2132d0-a0a2-4814-aa2c-1f7e69cccc93_1376x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!qc54!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3b2132d0-a0a2-4814-aa2c-1f7e69cccc93_1376x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!qc54!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3b2132d0-a0a2-4814-aa2c-1f7e69cccc93_1376x768.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!qc54!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3b2132d0-a0a2-4814-aa2c-1f7e69cccc93_1376x768.jpeg" width="728" height="409.5" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/3b2132d0-a0a2-4814-aa2c-1f7e69cccc93_1376x768.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:false,&quot;imageSize&quot;:&quot;normal&quot;,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:728,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;captionedImage&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!qc54!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3b2132d0-a0a2-4814-aa2c-1f7e69cccc93_1376x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!qc54!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3b2132d0-a0a2-4814-aa2c-1f7e69cccc93_1376x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!qc54!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3b2132d0-a0a2-4814-aa2c-1f7e69cccc93_1376x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!qc54!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3b2132d0-a0a2-4814-aa2c-1f7e69cccc93_1376x768.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>I like the arithmetic of it. The wire watches <strong>fifty-two sources</strong>. Over the last month it read and scored <strong>two thousand six hundred and seventy-nine items</strong>, and three hundred and ninety-nine of them cleared the bar. Yesterday's dispatch came out of a hundred and twenty-one signals. The site puts it more plainly than I would: it reads about seven things so you read one.</p><blockquote><p><strong>Fifty-two sources in. One page out.</strong></p></blockquote><p>What lands is a page I can get through with coffee.</p><div><hr></div><h2>The wire is the front door</h2><p>This is not a side project I bolted on. <strong>It sits at the front of everything else I make.</strong></p><p>What scores highest is usually what I end up writing about, because the things worth six hundred words announce themselves by refusing to go away. What I write becomes what I post. What I post brings back arguments and corrections from people who know more than I do about some corner of it, and those go back into the store as research. <strong>The loop closes.</strong> One pipe, several outlets, and the website is only the visible tip of it.</p><p>I did not design it that way at the start.</p><blockquote><p><strong>I built the scan because I was drowning. Then I noticed the thing that solved the drowning was also the thing deciding what I had to say.</strong></p></blockquote><div><hr></div><h2>Go look at it</h2><p><strong>The daily dispatch is free, it is public, and it lands every morning:</strong> <a href="https://wire.rundatarun.io/briefs">wire.rundatarun.io/briefs</a>. The day's signals, grouped by what they are, with a source link on every item so you can go argue with the original instead of taking my read on it.</p><p>Underneath it, the <a href="https://wire.rundatarun.io/firehose">firehose</a> shows you the machinery. The whole intake, the filter that thins it, and the last seven days in the raw.</p><p>What it does not show you is the store underneath. The complete searchable archive, the picks I would put in front of you myself with the reasoning attached, the decode and dossier library, the week ahead before it is news. <strong>That layer is not free, and it is not open yet.</strong> There is a waitlist at the bottom of the firehose page, and if the arithmetic up there sounded like your problem too, put your email in it.</p><div><hr></div><p>Here is the whole thing, start to finish, in about forty-five seconds.</p><div id="youtube2-SKk4hA0-9Mk" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;SKk4hA0-9Mk&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/SKk4hA0-9Mk?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div>]]></content:encoded></item><item><title><![CDATA[AI Ate the Keyboard. Now It Has to Eat the Queue.]]></title><description><![CDATA[Software stopped being the limiting factor. Process didn't. In biopharma and any patient-facing industry, the queue is the patient.]]></description><link>https://rundatarun.io/p/ai-ate-the-keyboard-now-it-has-to</link><guid isPermaLink="false">https://rundatarun.io/p/ai-ate-the-keyboard-now-it-has-to</guid><dc:creator><![CDATA[Justin Johnson]]></dc:creator><pubDate>Tue, 07 Jul 2026 15:30:18 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/f77759d2-d27a-445e-873e-1e81e7aa3925_1376x768.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!bipI!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2a271f29-27df-406f-9f07-f724abd1ecc0_1376x768.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!bipI!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2a271f29-27df-406f-9f07-f724abd1ecc0_1376x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!bipI!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2a271f29-27df-406f-9f07-f724abd1ecc0_1376x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!bipI!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2a271f29-27df-406f-9f07-f724abd1ecc0_1376x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!bipI!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2a271f29-27df-406f-9f07-f724abd1ecc0_1376x768.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!bipI!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2a271f29-27df-406f-9f07-f724abd1ecc0_1376x768.jpeg" width="728" height="409.5" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/2a271f29-27df-406f-9f07-f724abd1ecc0_1376x768.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:false,&quot;imageSize&quot;:&quot;normal&quot;,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:728,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;captionedImage&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!bipI!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2a271f29-27df-406f-9f07-f724abd1ecc0_1376x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!bipI!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2a271f29-27df-406f-9f07-f724abd1ecc0_1376x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!bipI!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2a271f29-27df-406f-9f07-f724abd1ecc0_1376x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!bipI!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2a271f29-27df-406f-9f07-f724abd1ecc0_1376x768.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Jack Hanlon, who leads GenAI Media at Meta, <a href="https://www.linkedin.com/posts/jhanlon_the-first-time-you-see-an-engineer-build-activity-7439409934175195139-XY6w">posted this on LinkedIn</a>:</p><blockquote><p>"The first time you see an engineer build something in 45 minutes that would have taken a week a year ago, but then see it not ship for another 6 weeks, you will be radicalized."</p></blockquote><p>He's describing Meta, a consumer tech company. The review stack he lists is the load every build carries before it ships: Strategy, Product, Design, Engineering, Privacy, Legal, Accessibility, Comms. It's also not close to what biopharma and patient-facing R&amp;D organizations ask of an AI build.</p><p>In biopharma and adjacent patient-facing industries, add data privacy impact assessment, security review, AI governance review, IT architecture review, change-management approval, procurement, legal redline, regulatory sign-off. If patient data is anywhere in the lifecycle, add clinical safety review. Each gate is serial. Each gate is three weeks minimum. Six in sequence is half a year before anything ships. You've done paperwork.</p><p>The numbers back this up across every regulated sector, not just pharma. Enterprise SaaS procurement runs a median 170 days; complex solutions hit eleven-and-a-half months. SOC 2 Type II audits take six to twelve months end-to-end. Federal contractors face 12-18 month FedRAMP authorization cycles before a system can sit on a government network. Banks run AI governance committees on top of model-risk committees on top of vendor-risk committees, each stack inherited from a different decade of incident response. MIT NANDA's "GenAI Divide" report found 95% of enterprise AI pilots deliver zero P&amp;L impact, across thirty to forty billion dollars of spend. That's not a build problem. The build works. The build sits in a queue.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!-j8K!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6a54bdd2-dd7c-4e51-963d-6486633dbdd9_1200x896.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!-j8K!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6a54bdd2-dd7c-4e51-963d-6486633dbdd9_1200x896.jpeg 424w, https://substackcdn.com/image/fetch/$s_!-j8K!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6a54bdd2-dd7c-4e51-963d-6486633dbdd9_1200x896.jpeg 848w, https://substackcdn.com/image/fetch/$s_!-j8K!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6a54bdd2-dd7c-4e51-963d-6486633dbdd9_1200x896.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!-j8K!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6a54bdd2-dd7c-4e51-963d-6486633dbdd9_1200x896.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!-j8K!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6a54bdd2-dd7c-4e51-963d-6486633dbdd9_1200x896.jpeg" width="728" height="409.5" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/6a54bdd2-dd7c-4e51-963d-6486633dbdd9_1200x896.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:false,&quot;imageSize&quot;:&quot;normal&quot;,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:728,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;captionedImage&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!-j8K!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6a54bdd2-dd7c-4e51-963d-6486633dbdd9_1200x896.jpeg 424w, https://substackcdn.com/image/fetch/$s_!-j8K!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6a54bdd2-dd7c-4e51-963d-6486633dbdd9_1200x896.jpeg 848w, https://substackcdn.com/image/fetch/$s_!-j8K!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6a54bdd2-dd7c-4e51-963d-6486633dbdd9_1200x896.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!-j8K!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6a54bdd2-dd7c-4e51-963d-6486633dbdd9_1200x896.jpeg 1456w" sizes="100vw"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>This accelerated review pathway can exist. It can work. It will take herculean levels of structural redesign and a lot of organizational political capital to land. And the pattern isn't a pharma problem. It's the shape of every regulated industry trying to ship AI in 2026. Pharma's version is the sharpest because patients are downstream of the delay, but the argument is universal.</p><p>Every gate exists for a reason. Most got written after a real incident. Someone's data got exposed. A launch communication went sideways. A model made a biased call in production. The answer isn't to kill them. The answer is that every one of them was designed for a world where building the software was the slow, expensive, irreversible part. Whether you wrote the code yourself or bought it from a vendor who wrote it, software was the thing the calendar bent around. That world is gone. AI didn't just speed up typing. It turned the whole build, your own code and the vendor's both, into the fast part. The review cycle was calibrated for the old critical path. The critical path moved. The calendar didn't.</p><p>Hanlon's three confrontations are right. Process architecture is the speed limit. Too many orgs treat every process as sacred. AI should be eating process, not just coding. The patient-facing version is sharpest, because the delay shows up downstream as patient time.</p><div><hr></div><h2>This Has Happened Before</h2><p>DevOps didn't land in 2015 as a philosophy. It landed when development got fast enough that deployment became the jam. "Development teams could complete features in days, but getting those features deployed took weeks." That's the DevOps.com retrospective on the pre-CI/CD era. Swap "deployed" for "procurement-approved" and you have a 2026 AI program, sentence for sentence.</p><p>Then security became the jam. DevSecOps. Then infrastructure and developer experience became the jam. Platform Engineering. Internal Developer Platforms. Each compression moved the bottleneck one layer further from the keyboard. Each shift spawned a category, a toolchain, a set of jobs, a conference circuit.</p><p>Bottleneck migration is a law. Compress one layer and the adjacent layer becomes the bottleneck by definition. Whatever used to be the tall pole is now almost all of the remaining pole.</p><p>AI compressed the build layer, writing software and buying it both. Everything adjacent to the build is now the bottleneck. The reviewers, the committees, the approvals, the procurement track, the vendor security questionnaire, the model evals nobody scoped time for. All of it. DevOps took twelve years to run the full cycle. Nobody has twelve years.</p><div><hr></div><h2>Shadow AI Is the Tell</h2><p>Gartner's 2025 numbers are the signal. 98% of organizations report unsanctioned AI use. 69% report evidence of prohibited public GenAI use inside their walls. 49% expect a shadow AI incident within twelve months.</p><p>The default read calls shadow AI a governance failure. A discipline problem. Employees not following the rules. That read is backwards. Shadow AI is the system telling you the official process layer already failed an internal cost-benefit test. Employees measured the queue, measured their quarterly objectives, and decided the queue costs more than the risk of getting caught.</p><p>That's the market voting. When 98% of your workforce is routing around your approval process, you don't have a discipline problem. You have an economics problem. The process priced itself out.</p><p>A fair counter to all of this: maybe the 95% pilot failure isn't a process problem. Maybe it's a pilot-selection problem. Most enterprise AI pilots fail because they were the wrong pilot to run, picked for board optics or executive curiosity rather than real workflow pain. Speeding up bad pilots faster doesn't help anyone.</p><p>That argument is partially right. The data is harder than that. The same MIT report that flagged the 95% number also found the highest ROI was in back-office automation, exactly the work that sits behind procurement, vendor security, and compliance queues. The pilots that fail include some bad picks. They also include good picks that died in queue. You can have a pilot-selection problem and a process problem at the same time, and most enterprises do.</p><div><hr></div><h2>Process Protects Patients. Except When It Doesn't.</h2><p>Biopharma's review gauntlet was written in blood. Good Clinical Practice, Good Laboratory Practice, Good Manufacturing Practice. Patients got hurt. Rules got written. Every review, every sign-off, every three-week queue descends from that inheritance.</p><p>The reflex response to "AI is slow to ship" is "good, that's how we keep patients safe." It sounds unanswerable. It isn't.</p><blockquote><p><strong>Process was built to protect patients. When process is the thing blocking patient-helping AI from shipping, it stops protecting patients. It protects itself.</strong></p></blockquote><p>Not shipping AI is not a neutral, safe choice. It has a cost. The cost lands on patients too.</p><p>Every week of review on a regulatory drafting agent is a week a safety narrative gets written the slow way. Novo Nordisk's NovoScribe rollout cut clinical study report drafting from fifteen weeks with over fifty medical writers to ten minutes with three. The work moved. The submission queue didn't. Every week of governance queue on a tool like that is real human time that wasn't spent on the submission, real submission time that wasn't spent reaching patients.</p><p>Every month a vendor security review stalls on a literature retrieval agent is a month oncology scientists read abstracts by hand. The responder-subgroup signal sits in a PDF nobody opened.</p><p>Every quarter a leadership group debates a two-speed governance design is a quarter trial protocols get mapped by hand, deviations get triaged out of a Teams channel, the first patient at the first site doses later than they would have.</p><p>Meanwhile, the deployments ship. Cleveland Clinic, NHS England, the VA, and 75% of US hospitals are running AI in clinical workflows today. The teams that wired governance early are in production. The teams still debating are watching.</p><p>Most governance agendas debate the wrong question. Not "is this AI safe enough to ship," but "is the delay we're about to impose smaller than the harm we're trying to prevent." Both sides of that ledger have a body count. Only one gets counted.</p><div><hr></div><h2>AI Eats Its Own Governance Layer</h2><p>The governance layer isn't a coding problem, and that's the part most "AI accelerator" pitches miss. It's legal redline, procurement of the services the agents will consume, vendor security, the privacy review, the model evals, the benchmark sign-off. None of that is typing. All of it is the build now, and all of it queues.</p><p>Organizations have answered by standing up governance as elaborate as the thing it governs. Sanofi has publicly documented a multi-stage setup: a Responsible AI Working Committee, an Interim Responsible AI Governance body, an AI agent named Plai sitting in on every drug-progression decision. Other large pharmas, banks, and federal contractors run variations of the same shape; the public documentation lags Sanofi's. Ethan Mollick names why it decides everything. "The moderating factor is no longer individual ability or even AI capability. It is organizational structure, policy, and the way leaders choose to approach AI." The constraint isn't the model. It's the layer wrapped around the model.</p><p>Which is exactly the layer AI can eat. Hanlon's third point is the one most readers stop quoting before the punchline:</p><blockquote><p>"Have the AI traverse your codebase and collect your evidence for the privacy review and have another AI grade the privacy review. They go back and forth until (a) they need human intervention because they are stuck or (b) they need humans to look at the final results and sign off on the right path forward."</p></blockquote><p>That's the whole answer, and it generalizes past the codebase. Every gate in the gauntlet has the same shape: forty hours of evidence assembly, then a fifteen-minute judgment call. The privacy review is fourteen pages of data flow that already lives in the repo, the retention policies, the IAM config. The vendor questionnaire is two hundred SOC 2 controls a lead engineer maps by hand against an attestation report nobody re-read. The procurement intake is free-text someone retypes against master agreement terms a sourcing lead has memorized. The model eval is a benchmark suite someone runs once and pastes into a slide. The regulatory submission is fifteen weeks of medical writers stitching prior submissions into a new narrative. The judgment is fifteen minutes. The forty hours is the queue.</p><p>The human doesn't leave the privacy review. The human finally opens a pre-assembled package instead of staring at a blank template with half the source material missing. The reviewer still reviews. The reviewer is finally reviewing the only part of the work that ever needed a reviewer.</p><blockquote><p><strong>AI has eaten the keyboard. It hasn't eaten the queue.</strong></p></blockquote><p>That's the argument in one sentence. The keyboard compressed two years ago. The queue hasn't. Until it does, every "AI accelerator" announcement is a faster typewriter feeding the same backlog.</p><div><hr></div><h2>The Reviewer Is the Project</h2><p>There's a catch built into all of this. The people whose work the AI must eat are the same people who must approve the AI eating their work. Legal reviews the legal agent. Procurement approves the procurement agent. InfoSec signs off on the InfoSec agent.</p><blockquote><p><strong>The reviewer is the subject of the review. The incentives don't line up. Pretending they do is how an org burns eighteen months and ships nothing.</strong></p></blockquote><p>The first week of agent output will be embarrassingly wrong, and that isn't the agent's fault. The process has never been written down in a way a machine can act on. The DPIA "template" is really a set of in-head judgments held by a senior person who onboarded in 2014 and knows which questions matter. Turning that into prompts, policies, and machine-readable rubrics is the institutional work, and there's no shortcut around it. That IS the project. The catch is that the people best positioned to codify their judgment are the ones whose standing rests on it staying in their heads. The org design that works makes the codified version the asset and leaves the committee to ratify it. Most enterprises don't run that design yet.</p><p>The ones that do compress the queue start in the same place, and it's never a model. They map the gauntlet, then fund an agent for a single gate, picked where the work is repetitive and the inputs already live somewhere a machine can read. Privacy review first, its answers mostly sitting in the repo and the IAM config. Vendor security next, the same shape with messier inputs. Regulatory drafting last, after the pattern has proven itself on something lower-stakes than an FDA filing. It holds at one gate, run for two months and watched for what it gets wrong. It breaks at three at once.</p><p>Underneath all of it sits the substrate. A governance layer that can't hand an agent a rubric to run or a template to fill isn't policy. It's tradition. The version living in senior reviewers' heads is neither auditable nor transferable, and it's the version adding weeks to every queue. Codifying it is the platform bet of 2026, ahead of the next model, because it's what lets every future build ship without the six-month tax.</p><div><hr></div><p>AI compressed the build layer. That was the easy part. Everyone got the faster typewriter. The next decade of enterprise AI won't be won on better models. It turns on whether the layer built to protect patients from bad software, back when bad software was the threat, can be made to compress itself before the delay starts landing on the other side of the ledger.</p><p>That layer will not shrink on its own. It has to be handed its own rubric and told to grade the work it used to guard by hand. No team that owns a queue volunteers to automate it. Someone with the authority to override that instinct has to decide the delay costs more than the control.</p><blockquote><p><strong>In a patient-facing industry, the queue is not paperwork. The queue is the patient. Every week it stays uncompressed is a week charged to someone who never got a seat in the room.</strong></p></blockquote>]]></content:encoded></item><item><title><![CDATA[Every AI Book Tells You What to Think. This One Is About What Your Hands Do.]]></title><description><![CDATA[There is a whole shelf of AI books for leaders now. I read enough of them to know why I wrote a different kind.]]></description><link>https://rundatarun.io/p/every-ai-book-tells-you-what-to-think</link><guid isPermaLink="false">https://rundatarun.io/p/every-ai-book-tells-you-what-to-think</guid><dc:creator><![CDATA[Justin Johnson]]></dc:creator><pubDate>Mon, 06 Jul 2026 12:48:33 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/308a6dbe-e77a-45c7-b3b8-7ecb2ea518b2_1408x768.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!0dbC!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4bf7a834-d4b3-4ed9-987d-dd4472ecf6f6_1408x768.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!0dbC!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4bf7a834-d4b3-4ed9-987d-dd4472ecf6f6_1408x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!0dbC!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4bf7a834-d4b3-4ed9-987d-dd4472ecf6f6_1408x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!0dbC!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4bf7a834-d4b3-4ed9-987d-dd4472ecf6f6_1408x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!0dbC!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4bf7a834-d4b3-4ed9-987d-dd4472ecf6f6_1408x768.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!0dbC!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4bf7a834-d4b3-4ed9-987d-dd4472ecf6f6_1408x768.jpeg" width="728" height="409.5" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/4bf7a834-d4b3-4ed9-987d-dd4472ecf6f6_1408x768.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:false,&quot;imageSize&quot;:&quot;normal&quot;,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:728,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;captionedImage&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!0dbC!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4bf7a834-d4b3-4ed9-987d-dd4472ecf6f6_1408x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!0dbC!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4bf7a834-d4b3-4ed9-987d-dd4472ecf6f6_1408x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!0dbC!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4bf7a834-d4b3-4ed9-987d-dd4472ecf6f6_1408x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!0dbC!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4bf7a834-d4b3-4ed9-987d-dd4472ecf6f6_1408x768.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>There is a shelf now. If you are a senior leader trying to make sense of AI, you have probably bought most of it. The strategy book. The mindset book. The one with the four-quadrant framework. The one from the consultancy with eleven co-authors. They are not bad books. Several are good. I learned things from them.</p><p>But I noticed something after the fourth or fifth one. I would finish a chapter, nod, underline a sentence, and then sit there with my hands in exactly the same place they were before I started reading. The book had changed what I thought about AI. It had changed nothing about what I did with it on Monday.</p><p>That is the gap I wrote into.</p><blockquote><p><strong>The book changed what I thought about AI. It changed nothing about what I did with it on Monday.</strong></p></blockquote><div><hr></div><h2>The shelf has one shape</h2><p>Read enough of these books and the shape repeats. They tell you what to think about AI. How to frame it for your boss/peers/board. Which mental model to adopt. Where it sits in your strategy. They are written from the altitude a leader is used to operating at, which is to say: above the work, directing it.</p><p>That altitude is the problem, not the solution. Someone reviewing the current best-seller in this category said it reads like a 2024 book explaining how the Internet will help your business. The line is unkind and basically right. Not because the author is wrong, but because the genre has a ceiling. You cannot write the operator's manual from the strategy altitude. The two are different books, and almost everyone is writing the first one.</p><p>I wanted to read the second one. Nobody had written it for the reader I had in mind, so I wrote it.</p><blockquote><p><strong>You cannot write the operator's manual from the strategy altitude.</strong></p></blockquote><div><hr></div><h2>What "what your hands do" actually means</h2><p>"Leaders should build" is already a slogan, and slogans rot.</p><p>I don't mean you should learn to code. I mean something narrower. There is a class of tools now, Claude Code chief among them, that lets you direct a multi-agent system to produce real work inside your own domain. You describe intent, set boundaries, review what comes back, redirect, and ship. The interface happens to be a terminal; the skill is not programming. It is the same skill that got you to senior: defining the why, designing the handoffs, judging the output, catching the thing that looks right and is wrong.</p><p>You already do all of that. You do it with people. The book is about pointing those same instincts at a harness instead of a team, and what changes when you do.</p><p>The reason this matters is not productivity, though the productivity is real. The reason is calibration. Andrej Karpathy named it this spring in a line that landed with twenty thousand likes and a thousand replies: there are two groups talking past each other about AI, and the line between them is not skeptic versus believer. It is the people who have built something with these systems and the people who have read about them. Every confident claim you have heard about what AI can or cannot do was made from one side of that line. If you are reasoning about AI from the reading side, you are reasoning from the wrong data, and no amount of strategy reading moves you across. You move across by getting your hands on the controls one time and feeling where the real edge is.</p><div><hr></div><h2>Why I get to say this</h2><p>What separates this book from the shelf is that I did not write it from the strategy altitude. I built and led an AI Center of Excellence at a Fortune 500 pharmaceutical company, and I now lead applied data and AI in R&amp;D at an AI-driven biotech. I also run a deep personal practice on nights and weekends, eight autonomous agents and a homelab full of skills, because the day job and the obsession turned out to be the same skill. This book was itself produced through the harness it describes, end to end. About a month from kickoff to typeset manuscript, against the year the same book would have taken by hand, with thousands of hours of research, red-teaming, and editing behind it, most of it run in parallel by the system rather than typed by me. You can see the skills that drafted it.</p><p>I am not telling you to do something I read about. I am telling you what it was actually like, including the parts that are awkward and the parts that broke.</p><blockquote><p><strong>I am not telling you to do something I read about. I am telling you what it was actually like.</strong></p></blockquote><div><hr></div><h2>The part I won't oversell</h2><p>At Sequoia's AI Ascent this spring, Karpathy declared "vibe coding" obsolete and named its successor "agentic engineering," a discipline with real rigor and real responsibility. He is right. He also drew a careful line: prototyping with AI raises the floor for everyone, but operating serious systems demands discipline you cannot wave away.</p><p>My book asks you to step over a line slightly further out than the one Karpathy draws for the general public. I am asking a senior, non-technical executive to personally operate, not just prototype. I think that is where the role is heading, and I think the leaders who learn it first will shape what their organizations become. But I am not going to pretend it is the consensus position. It is an argument. The book makes it, and gives you the controls and the guardrails to test it for yourself rather than take my word for it.</p><div><hr></div><h2>Who will not find this useful</h2><p>If you want a framework to put on a slide for your board, the shelf has better options than mine, and I mean that. If you want reassurance that your current AI strategy is sound, this is the wrong book; it will probably make you uncomfortable. If you are looking for predictions about AGI timelines, I route around that debate on purpose.</p><p>This book is useful if you have read the think pieces, sponsored the center of excellence, watched the pilots, and still feel the quiet gap between knowing AI matters and knowing what you personally do about it. It is useful if you are willing to spend a few hours with your hands on something unfamiliar, on your own files, to find out where the edge actually is. It is a ladder. Three months to ship something real, six on the outside. The last page is a welcome from the people who were already on the other side when you started.</p><p>If that is you, the book is for you. It's out now: <a href="https://builder-leader.com">builder-leader.com</a>, or <a href="https://www.amazon.com/Builder-Leader-Exoskeleton-That-Crosses-Gap-ebook/dp/B0H3LQQ4J6/">buy it on Amazon</a>.</p><p><em>Builder-Leader: The AI Exoskeleton That Crosses the Gap.</em></p>]]></content:encoded></item></channel></rss>