<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[Run Data Run]]></title><description><![CDATA[Clear thinking about AI from someone building it in production. No hype, no hand-waving. Just what works, what doesn't, and why it matters. My book is out: builder-leader.com]]></description><link>https://rundatarun.io</link><image><url>https://substackcdn.com/image/fetch/$s_!t_Ch!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffa36f5aa-74af-4492-b8d7-93b03f14a337_1280x1280.png</url><title>Run Data Run</title><link>https://rundatarun.io</link></image><generator>Substack</generator><lastBuildDate>Mon, 03 Aug 2026 21:47:50 GMT</lastBuildDate><atom:link href="https://rundatarun.io/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[Justin Johnson]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[rundatarun@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[rundatarun@substack.com]]></itunes:email><itunes:name><![CDATA[Justin Johnson]]></itunes:name></itunes:owner><itunes:author><![CDATA[Justin Johnson]]></itunes:author><googleplay:owner><![CDATA[rundatarun@substack.com]]></googleplay:owner><googleplay:email><![CDATA[rundatarun@substack.com]]></googleplay:email><googleplay:author><![CDATA[Justin Johnson]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[Somebody Is Finally Checking]]></title><description><![CDATA[A contest opened eleven days ago has produced more reproduction attempts than the field's own dedicated effort has managed in any year. Whether that becomes a check on the literature turns on a detail sitting in the scoring rules.]]></description><link>https://rundatarun.io/p/somebody-is-finally-checking</link><guid isPermaLink="false">https://rundatarun.io/p/somebody-is-finally-checking</guid><dc:creator><![CDATA[Justin Johnson]]></dc:creator><pubDate>Sun, 26 Jul 2026 04:06:38 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/e3f0aa85-2640-4ecf-9d35-bc6909acca58_1376x768.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em>Every Sunday I pick one paper or release that's worth your time, break it apart, and tell you why it matters. No hype. No summaries of summaries. Just the idea, explained.</em></p><div><hr></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!TLMS!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d0e0b21-f86a-4342-8104-343a9e655bef_1376x768.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!TLMS!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d0e0b21-f86a-4342-8104-343a9e655bef_1376x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!TLMS!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d0e0b21-f86a-4342-8104-343a9e655bef_1376x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!TLMS!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d0e0b21-f86a-4342-8104-343a9e655bef_1376x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!TLMS!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d0e0b21-f86a-4342-8104-343a9e655bef_1376x768.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!TLMS!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d0e0b21-f86a-4342-8104-343a9e655bef_1376x768.jpeg" width="728" height="409.5" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/2d0e0b21-f86a-4342-8104-343a9e655bef_1376x768.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:false,&quot;imageSize&quot;:&quot;normal&quot;,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:728,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;captionedImage&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!TLMS!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d0e0b21-f86a-4342-8104-343a9e655bef_1376x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!TLMS!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d0e0b21-f86a-4342-8104-343a9e655bef_1376x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!TLMS!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d0e0b21-f86a-4342-8104-343a9e655bef_1376x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!TLMS!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d0e0b21-f86a-4342-8104-343a9e655bef_1376x768.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>A published result is a promise. Run this, and you will see what we saw. For most of the history of computational science nobody collected on that promise, because collecting cost about as much as the original work and earned a fraction of the credit.</p><p>On July 15, Hugging Face and <a href="https://www.alphaxiv.org/">alphaXiv</a> opened a contest to collect. It closes Sunday August 2. Roughly 6,800 papers from ICML 2026 are in scope, anyone can attempt a reproduction, and every attempt gets published as a public logbook that an automated judge then scores. The prize pool is $4,500, paid in GPU credits rather than cash, which tells you who the organizers think is entering.</p><p><strong>As of this morning, eleven days in: 4,147 judged logbooks from 312 people, carrying 20,594 separate verdicts on individual claims.</strong> About 1,600 papers have been attempted at least once, roughly a quarter of the conference.</p><p>To see why that is a strange number, you need the thing it replaces. <a href="https://reproml.org/">Joelle Pineau</a> started the ML Reproducibility Challenge at ICLR 2018 as <a href="https://blog.neurips.cc/2026/05/04/mlrc-2026-reproducibility-as-an-official-track-at-neurips/">"a small community experiment."</a> Eight years on it is the field's serious effort, mostly run through graduate courses, and its NeurIPS 2019 edition drew <a href="https://jmlr.org/papers/volume22/20-303/20-303.pdf">173 papers claimed</a>, itself a <strong>92% jump on the year before</strong>. Claimed, not finished. Published reports have run in the tens. <a href="https://rescience.github.io/">ReScience C</a>, which demands a fresh independent implementation rather than a re-run of the authors' code, manages about a dozen a year.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!WwYA!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda36e809-29d8-45e4-ac68-108f3e598ccf_1376x768.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!WwYA!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda36e809-29d8-45e4-ac68-108f3e598ccf_1376x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!WwYA!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda36e809-29d8-45e4-ac68-108f3e598ccf_1376x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!WwYA!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda36e809-29d8-45e4-ac68-108f3e598ccf_1376x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!WwYA!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda36e809-29d8-45e4-ac68-108f3e598ccf_1376x768.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!WwYA!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda36e809-29d8-45e4-ac68-108f3e598ccf_1376x768.jpeg" width="728" height="409.5" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/da36e809-29d8-45e4-ac68-108f3e598ccf_1376x768.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:false,&quot;imageSize&quot;:&quot;normal&quot;,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:728,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;captionedImage&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!WwYA!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda36e809-29d8-45e4-ac68-108f3e598ccf_1376x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!WwYA!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda36e809-29d8-45e4-ac68-108f3e598ccf_1376x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!WwYA!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda36e809-29d8-45e4-ac68-108f3e598ccf_1376x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!WwYA!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda36e809-29d8-45e4-ac68-108f3e598ccf_1376x768.jpeg 1456w" sizes="100vw"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><blockquote><p><strong>173 papers claimed in a year, against 4,147 judged logbooks in eleven days. That is not an improvement on the existing effort, it is a different quantity.</strong></p></blockquote><div><hr></div><h2>What is actually being scored</h2><p>One design decision separates this from a leaderboard, and it is easy to skim past. <strong>The unit of judgment is not the paper. It is the sentence.</strong></p><p>Every paper has been broken into its individual claims, and each claim gets its own verdict: verified, falsified, <strong>toy</strong> (a scaled-down version that says so), or inconclusive. Nobody asks whether a paper "reproduced," which is a question with no answer. A paper can come back partly verified, partly toy and partly inconclusive at once, and most of them do.</p><p>Then the scoring, from the contest's own FAQ: <strong>"2 points for a full reproduction or full falsification, 1 point for a toy-scale reproduction, 0 otherwise."</strong></p><blockquote><p><strong>Falsifying a claim pays what confirming one pays. Against a literature where negative results are half as publishable as positive ones, that is a deliberate inversion.</strong></p></blockquote><p>When <em>Nature</em> surveyed 1,576 researchers in 2016, <a href="https://www.bioedonline.org/news/nature-news-archive/1500-scientists-lift-the-lid-on-reproducibility/">24% had published a successful replication and 13% a failed one</a>. The same survey found more than 70% had tried and failed to reproduce someone else's experiment and 52% agreed the field had a reproducibility crisis, and then 73% said they still trusted at least half the papers in their own field.</p><p>The failures are also the deliverable. A logbook, in the organizers' description, holds <strong>"the experiments, simplifications, failures, and results it found."</strong> Prior efforts published conclusions. This one publishes the attempt.</p><p><strong>What it does not do is test the hard version.</strong> Machines have been tried on this before and were not good at it: every benchmark that measured reproduction topped out between 16% and 27% on rebuilding a paper from scratch, with <a href="https://arxiv.org/abs/2504.01848">top ML PhDs ahead of the models</a>, and only reached about 60% <a href="https://arxiv.org/abs/2409.11363">when handed the original code and data</a>. This contest tells participants to start from the authors' code. <strong>Re-running someone's code, rebuilding from the paper alone, and getting the same finding on fresh data are three different tests, and headline percentages get quoted across that boundary constantly.</strong></p><div><hr></div><h2>What the inside looks like</h2><p>For a large class of papers the easy version is now free, which is the whole reason this works. We downloaded a full-marks logbook from the contest leader and read it. <strong>The entire reproduction ran on a laptop processor in three seconds.</strong> No GPU, no cluster, no cloud bill. The real ceiling is not compute but publishing: Hugging Face caps accounts at 20 new pages a day, and when a participant asked for relief the organizer's answer was flat, <strong>"we cannot lift the 20 spaces/day limit so I'll close this issue."</strong></p><p>I entered, which is the only reason I have anything to say about the texture of it. We are 23 judged logbooks and 174 points in, against a leader at 1,862. Not a contender, and far enough inside to see how the work divides.</p><p>It divides into two lanes, and we ran them differently on purpose.</p><p><strong>The volume lane is an autonomous agent called Vulcan, and it never publishes anything.</strong> It runs on a box in my house on free local compute, and its job is narrow: take a paper whose claims are all theory and check every formula numerically, against a brief that fixes the method rather than the answer.</p><blockquote><p><strong>A check that passes on every input you can construct certifies nothing.</strong></p></blockquote><p>Its audit of one paper came back as 721 lines of code that matched the paper's algebra to the last digit a computer can represent, plus a deliberate sanity test that failed when it was supposed to. Careful work, at no cost, with a hard limit written into the contract: <strong>it stages, a human publishes.</strong> No autonomous process of mine puts a verdict about somebody else's paper into the world.</p><p>The other lane is me and Claude Code running on Opus, for the parts the first one cannot do. Two things came out of it.</p><p><strong>The first is that picking the right paper beats checking it well.</strong> Our most careful audit matched its paper exactly and scored 2 points out of a possible 8, because half that paper's claims were hardware benchmarks that sit at inconclusive until somebody spends days of compute on them. The judge said so directly: the logbook left "half the paper's claims without any experimental evidence." A paper whose claims are all theory earns 12 points in an afternoon on a laptop. <strong>Selection is the skill, not rigor.</strong></p><p>The second is what the process catches. One paper states a theorem using a fixed setting, and then, four thousand lines later, its own proof assumes that setting changes with the length of the run. Not a typo, and not a disagreement between the paper and the world: the paper disagrees with itself, in print, peer-reviewed and accepted. Nobody had noticed because nobody had reason to run the theorem as written and watch it fail.</p><p>Beyond those two, the work was unglamorous. It caught our own staging script dropping most of a conclusion, so a 60-line write-up shipped as three sentences. Vulcan itself lost twenty hours to an outage at its model provider, with no backup configured and nothing raising an alarm.</p><p><strong>None of those failures announce themselves.</strong> Every one produced output that looked finished.</p><div><hr></div><h2>The three percent</h2><p><strong>Of 20,594 claim verdicts, 615 are falsified. That is 3.0%.</strong> The rest split 43.5% verified, 28.9% inconclusive, 24.6% toy. The distribution has barely moved for days while the corpus grew by hundreds of logbooks: the shape is stable, the counts are not.</p><p>Three percent is low, and it has two readings. Either the literature is in better condition than the surveys suggest, or falsification is the hardest verdict to reach and the easiest to get wrong.</p><p><strong>In one day, four of our audit agents returned a falsification. Three of them were wrong.</strong> All four came out of the automated lane. All three errors were caught by the other one.</p><p>One agent read a formula out of a PDF that had lost a character in scanning, then correctly proved the mangled version false. Another was checking a sentence the authors never published, because the claim it was handed came from a draft that is not the public paper. The third found a difference too small to distinguish from rounding error and called it a contradiction.</p><p>Every one produced a confident, well-formatted, coherent falsification. I wrote about that pattern in <a href="https://rundatarun.io/p/the-failure-that-leaves-no-corpse">The Failure That Leaves No Corpse</a>, and this is the cleanest instance of it I have run into.</p><p>The rule we ended up with fits on one line: <strong>does the paper disagree with itself, or does the claim disagree with your copy of the paper?</strong> Only the first is a falsification. The second is a bug in your pipeline wearing a falsification's clothes.</p><p>Applying it cost us the verdict. <strong>Across the 130 claim verdicts we have published, the number labelled falsified is zero.</strong> The one candidate that survives is the self-contradicting theorem above.</p><blockquote><p><strong>The default failure of an automated checker is not a missing answer. It is a wrong one, delivered with the same confidence as a right one.</strong></p></blockquote><p>This is not only our problem. The contest's own claim lists are extracted by a language model, and a participant found a claim in one paper's list that belonged to a different paper. The organizer confirmed it in the open, <strong>"That claim belongs to a different Safety-area paper and leaked in during auto-extraction,"</strong> removed it, rescored, then swept all 6,768 papers for the same contamination.</p><p>And the bias is measured, not hypothetical. A <a href="https://arxiv.org/abs/2606.11447">June 2026 study of reproduction agents</a> found that <strong>showing the agent the original paper made it more likely to confirm claims that could not actually be reproduced</strong>, and that small changes in wording pushed it further the same way. In this contest the agent always has the paper and the claim in front of it.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!qr23!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F182e8b8d-54dc-477a-8405-704cbf8eb398_1376x768.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!qr23!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F182e8b8d-54dc-477a-8405-704cbf8eb398_1376x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!qr23!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F182e8b8d-54dc-477a-8405-704cbf8eb398_1376x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!qr23!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F182e8b8d-54dc-477a-8405-704cbf8eb398_1376x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!qr23!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F182e8b8d-54dc-477a-8405-704cbf8eb398_1376x768.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!qr23!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F182e8b8d-54dc-477a-8405-704cbf8eb398_1376x768.jpeg" width="728" height="409.5" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/182e8b8d-54dc-477a-8405-704cbf8eb398_1376x768.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:false,&quot;imageSize&quot;:&quot;normal&quot;,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:728,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;captionedImage&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!qr23!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F182e8b8d-54dc-477a-8405-704cbf8eb398_1376x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!qr23!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F182e8b8d-54dc-477a-8405-704cbf8eb398_1376x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!qr23!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F182e8b8d-54dc-477a-8405-704cbf8eb398_1376x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!qr23!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F182e8b8d-54dc-477a-8405-704cbf8eb398_1376x768.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div><hr></div><h2>The judge is a language model</h2><p><strong>The thing scoring 20,594 claim verdicts is itself a language model.</strong></p><p>There is published criticism aimed squarely at that. A 2026 position paper, <a href="https://arxiv.org/abs/2605.03202">"Stop Automating Peer Review Without Rigorous Evaluation"</a>, compared human against AI reviews of ICLR 2026 submissions and found two things. The AI reviewers showed <strong>"a hivemind effect of excessive agreement within and across papers that reduces perspective diversity."</strong> And their scores turned out to be <strong>"trivially gameable through paper laundering"</strong>: rewriting a paper with a language model raised the scores it got, on style rather than science.</p><p>None of that is a reason to dismiss the contest, and I want to be careful here, because I am competing in it and have every incentive to be generous.</p><p>The defence is structural rather than technical. <strong>The artifacts are public.</strong> Every logbook is a page anyone can open, the verdicts are a public dataset downloaded 38,438 times last month, and the errors get argued in threads with the organizers answering. The contamination case is the proof, and the fix landed inside a day. A closed judge making the same error produces a number nobody can audit.</p><p>So the position I hold is that this design pairs a known-weak instrument with an unusually strong correction loop. Whether the loop outruns the instrument is an empirical question, and a week from now there will be enough public data to start answering it.</p><div><hr></div><h2>What it would take to trust this</h2><p>I am not going to predict how it turns out. The contest closes August 2 and every distribution above is provisional.</p><p>The concrete claim is narrower and I think it holds. <strong>For a meaningful slice of published science, checking a claim now costs less than making it.</strong> Three seconds of laptop processor against however many months produced the original theorem. That ratio is why verification stayed volunteer work for a decade, and it has flipped.</p><p>What it does not change is the direction of the error. A checker that costs nothing will be run constantly, and cheap checking produces far more false accusations than misses. Confidently accusing a sound paper is worse than missing a bad one.</p><p>The contest's design anticipates this better than most things I have seen. Paying the same for a falsification as a verification removes the incentive to go looking only for confirmations. The logbook has to carry the failures, so there is no tidy way to bury the attempts that went nowhere. And because every artifact is public, a wrong answer stays inspectable instead of disappearing into an aggregate.</p><blockquote><p><strong>Nobody has to trust the checkers. You can read what they did. That is a weaker guarantee than peer review claims to offer, and a stronger one than peer review delivers.</strong></p></blockquote><div><hr></div><h2>Sources</h2><ul><li><p><a href="https://huggingface.co/spaces/ICML-2026-agent-repro/challenge">Reproducing ICML 2026, Open Reproductions</a>: the contest, its FAQ, rules, and public discussion threads. Verdicts live in a <a href="https://huggingface.co/datasets/ICML-2026-agent-repro/verdicts">public dataset</a>.</p></li><li><p><a href="https://blog.neurips.cc/2026/05/04/mlrc-2026-reproducibility-as-an-official-track-at-neurips/">MLRC 2026: Reproducibility as an Official Track at NeurIPS</a>: the ML Reproducibility Challenge's history and Pineau's founding description of it.</p></li><li><p><a href="https://jmlr.org/papers/volume22/20-303/20-303.pdf">Improving Reproducibility in Machine Learning Research</a>, Pineau et al., JMLR 22(164): the 173-papers-claimed figure.</p></li><li><p><a href="https://rescience.github.io/">ReScience C</a>: the journal that requires a new independent implementation.</p></li><li><p><a href="https://www.bioedonline.org/news/nature-news-archive/1500-scientists-lift-the-lid-on-reproducibility/">1,500 scientists lift the lid on reproducibility</a>, Monya Baker, <em>Nature</em> 533: the 2016 survey of 1,576 researchers.</p></li><li><p><a href="https://arxiv.org/abs/2504.01848">PaperBench</a> and <a href="https://arxiv.org/abs/2409.11363">CORE-Bench</a>: the agent-reproduction benchmarks this contest inherits from.</p></li><li><p><a href="https://arxiv.org/abs/2606.11447">Reproducibility agents and confirmation bias</a>: the June 2026 benchmark showing that giving an agent the paper biases it toward confirming.</p></li><li><p><a href="https://arxiv.org/abs/2605.03202">Stop Automating Peer Review Without Rigorous Evaluation</a>: Baumann, Pei, Koyejo and Hovy, 2026: the hivemind effect and paper laundering.</p></li><li><p><a href="https://rundatarun.io/p/the-failure-that-leaves-no-corpse">The Failure That Leaves No Corpse</a>: the earlier piece on failures that produce a confident answer instead of an error.</p></li></ul><div><hr></div><p><em>Sunday Deep Dive is a weekly series on Run Data Run. Every Sunday I pick one paper, release, or technique worth understanding, break it apart, and tell you what it means for your work. Free every Sunday, no paywall. If it was useful, the easiest way to support it is to subscribe and forward it to one person on your team who'd want it. If it wasn't, tell me why. I'll make it better.</em></p>]]></content:encoded></item><item><title><![CDATA[Nobody Saves Money on the Model]]></title><description><![CDATA[A team just swapped in a model that costs twice as much per token, and their bill went down. Here is why that is not a paradox, and what it means for anyone trying to make AI cheaper at scale.]]></description><link>https://rundatarun.io/p/nobody-saves-money-on-the-model</link><guid isPermaLink="false">https://rundatarun.io/p/nobody-saves-money-on-the-model</guid><dc:creator><![CDATA[Justin Johnson]]></dc:creator><pubDate>Tue, 21 Jul 2026 12:30:17 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/417116f3-2fb3-4906-b055-5124f68d1532_1376x768.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!cAWE!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F18713613-e64e-4119-802d-79f139aba297_1376x768.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!cAWE!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F18713613-e64e-4119-802d-79f139aba297_1376x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!cAWE!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F18713613-e64e-4119-802d-79f139aba297_1376x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!cAWE!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F18713613-e64e-4119-802d-79f139aba297_1376x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!cAWE!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F18713613-e64e-4119-802d-79f139aba297_1376x768.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!cAWE!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F18713613-e64e-4119-802d-79f139aba297_1376x768.jpeg" width="728" height="409.5" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/18713613-e64e-4119-802d-79f139aba297_1376x768.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:false,&quot;imageSize&quot;:&quot;normal&quot;,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:728,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;captionedImage&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!cAWE!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F18713613-e64e-4119-802d-79f139aba297_1376x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!cAWE!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F18713613-e64e-4119-802d-79f139aba297_1376x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!cAWE!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F18713613-e64e-4119-802d-79f139aba297_1376x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!cAWE!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F18713613-e64e-4119-802d-79f139aba297_1376x768.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Last week a team at <a href="https://cognition.com/blog/devin-fusion">Cognition</a>, the company behind the coding agent Devin, published a number that reads like a typo. They replaced Opus 4.8 with Fable 5. Fable 5 costs about twice as much per token. Their bill went down.</p><p>Not their quality. Their bill.</p><p>Here is the part of their table that matters:</p><p><strong>Fable 5, inside their new architecture: score 57.6, cost $3.00 per task.</strong></p><p><strong>Opus 4.8, on its own: score 48.8, cost $3.24 per task.</strong></p><p>The expensive model, wired up correctly, was better <em>and</em> cheaper than the cheap model on its own. Not a trade. Both columns at once.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!yuC2!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa12f42d2-cfed-4c78-a945-d040df9b1b90_1376x768.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!yuC2!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa12f42d2-cfed-4c78-a945-d040df9b1b90_1376x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!yuC2!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa12f42d2-cfed-4c78-a945-d040df9b1b90_1376x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!yuC2!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa12f42d2-cfed-4c78-a945-d040df9b1b90_1376x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!yuC2!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa12f42d2-cfed-4c78-a945-d040df9b1b90_1376x768.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!yuC2!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa12f42d2-cfed-4c78-a945-d040df9b1b90_1376x768.jpeg" width="728" height="409.5" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a12f42d2-cfed-4c78-a945-d040df9b1b90_1376x768.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:false,&quot;imageSize&quot;:&quot;normal&quot;,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:728,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;captionedImage&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!yuC2!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa12f42d2-cfed-4c78-a945-d040df9b1b90_1376x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!yuC2!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa12f42d2-cfed-4c78-a945-d040df9b1b90_1376x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!yuC2!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa12f42d2-cfed-4c78-a945-d040df9b1b90_1376x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!yuC2!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa12f42d2-cfed-4c78-a945-d040df9b1b90_1376x768.jpeg 1456w" sizes="100vw"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>If you have ever sat in a meeting where someone proposed saving money by dropping down a model tier, that result should stop you. It stopped me, because I had run the opposite experiment, and I had lost.</p><div><hr></div><h2>What they actually built</h2><p>Cognition split the work between two models instead of routing between them.</p><p>The expensive model is the driver. It plans, it interprets what the request actually means, it makes the judgment calls, and it does the final review. It takes very few actions itself and it reads only what it must.</p><p>The cheap model is the sidekick, and it is not a helper function. It is a full agent with its own tools that goes and does the mechanical work: fetching context, running the slow tests, carrying out the routine implementations once someone has decided what they are.</p><p>Both hold their own memory. Both run at the same time. And this is not a demo. Eighty-eight percent of Cognition's own internal code changes now run through that automated split.</p><p>The reason it works fits in one line. <strong>You pay for thinking once, and for doing it many times.</strong></p><div><hr></div><h2>The problem is that the opposite is also true</h2><p>Everyone in this field has met the other result. You move a workload to a cheaper model to save money, and it costs you more. The cheap model misunderstands the task, produces something confidently wrong, and a person spends an afternoon unpicking it. The invoice went down and the total cost went up.</p><p>That happens, and I am not going to argue with it, because I have the receipts.</p><p>So we have two findings that appear to be at war. Expensive models save money. Cheap models cost money. Both are observed, both are honest, and most of the advice you will read picks one and ignores the other.</p><div><hr></div><h2>I ran the losing experiment</h2><p>I built a cost router. The logic was the logic everyone reaches for: work out how hard the task is, send the easy ones to a cheap model, keep the expensive one for the hard ones.</p><p>It saved nothing. Not a little less than I hoped. Nothing.</p><p>I want to be precise about why, because the failure is more useful than a success would have been. <strong>I routed by task type. The axis that pays is ambiguity.</strong></p><p>Every individual call my router made was defensible. This one looks like a simple rename, send it to the cheap model. That one looks like architecture, keep it upstairs. And the total refused to move, because "simple rename" was a description of the <em>work</em>, not a description of <em>how much had already been decided</em>. Half the jobs I was labelling easy still had open questions inside them, and an open question handed to a cheap model is the single most expensive thing you can buy.</p><div><hr></div><h2>The law</h2><p>Once you see it, the war stops.</p><blockquote><p><strong>A cheap model is expensive on an open question and cheap on a closed one.</strong></p></blockquote><p>An open question is one where something still has to be decided. What are we actually building. What does done mean here. Which of these two readings of the request is the real one. Give that to a cheap model and it will not tell you the question is open. It will pick an answer, sound sure, and hand you something plausible that you now have to check line by line.</p><p>A closed question has had the deciding done. Rename this function in these three files, and the test that proves it is this one. There is no judgment left in the task, and a cheaper model does it for a fraction of the price with nothing at risk.</p><p>Which gives the driver model a job description nobody writes down. <strong>Its work is not "the hard parts." Its work is to turn open questions into closed ones.</strong> And that conversion has a name we already use for it. It is called a plan.</p><p>This also explains the finding that keeps embarrassing people who try to build a committee of cheap models and vote. A group at <a href="https://arxiv.org/abs/2502.00674">Princeton</a> tested that directly last year and found that running the single best model several times and combining its own answers beat mixing different models together, on every one of the thirteen mixed configurations they tried, using roughly half the forward passes. Their explanation is blunt: mixing models of different quality drags the average quality down. <strong>You cannot vote your way to judgment.</strong> Where the question is still open, quality dominates, and diversity is not the free lunch it looks like.</p><div><hr></div><h2>Cheap is not a property of the model</h2><p>Here is where I think the whole conversation is framed wrongly, including by the people getting the right answers.</p><p>Everyone calls this "big model, small model." My own setup says that is not the axis.</p><p>The models I push my mechanical work to are not small. They are large, capable models. They cost me nothing at the margin, because I bought them on flat monthly subscriptions instead of by the token. That is not a smaller brain. It is a different contract.</p><p>So there are two dials, and they are independent. <strong>Where does the marginal cost live, and where does the judgment live.</strong> A fine-tuned small model is cheap because it was narrowed. A large model on a flat rate is cheap because of how you bought it. Both are "the cheap one," and treating them as the same thing is how people end up sending an open question to a bargain and wondering why the quarter went sideways.</p><div><hr></div><h2>The sentence I had already written</h2><p>I found a note I made a few weeks ago, while building the thing that hands my grunt work off to those subscription models. I had written the law without recognising it:</p><blockquote><p><strong>A vague spec produces confident garbage, and the cheaper the model, the truer that is.</strong></p></blockquote><p>Read that again as a cost statement, because that is what it is. A cheap model is not a discount. It is a <strong>multiplier on the quality of your specification.</strong> Specify tightly and it multiplies your savings. Specify loosely and it multiplies your mess.</p><p>Which is why the discipline I run alongside it is not optional. The expensive model writes the entire work order: the task, the files, what done means, and the exact command that proves it. And when the cheap model comes back and says the job is finished, that claim is worth nothing. I read the difference it made, and I run the test myself. It has never once been the model's word that closed a task.</p><div><hr></div><h2>So the small-model story is not a cost play</h2><p>This is the part I keep chewing on, because it inverts something I believed.</p><p>Everyone is excited about training small models to do one narrow job extremely well, and the excitement is framed as a cost saving. It is not, or at least the saving is not where people are pointing.</p><p>If you can specify a task tightly enough to train a small model on it, you have <strong>proved the task is closed.</strong> All the ambiguity was removed by somebody, at some point, doing the expensive work of deciding. The training run is just collecting the winnings.</p><blockquote><p><strong>A small model is not a discount. It is a receipt for judgment already spent.</strong></p></blockquote><p>And the same is true of every cost reduction I have ever managed to make stick at scale. Every one of them was bought earlier, by an act of judgment that closed a question. The saving showed up in the invoice. It was created somewhere else entirely.</p><div><hr></div><h2>The boundary, in their words</h2><p>Buried in Cognition's own write-up, past the numbers, is the caveat I would have led with:</p><blockquote><p><strong>The sidekick fails when judgment is the deliverable.</strong></p></blockquote><p>Their example is a hard feature whose subtle intent got lost the moment it was handed down. The work came back correct and wrong at the same time.</p><p>Anyone who has run a team knows exactly where that line sits, and knows it is not about the seniority of the person you handed it to. You can delegate the work. You cannot delegate the judgment about what the work is for. The failure looks identical from the outside either way: something arrives, it is technically defensible, and it is not what the thing was for.</p><div><hr></div><h2>The honest caveat</h2><p>One thing about that Cognition result deserves saying out loud, because they said it themselves and nobody repeating the headline has.</p><p>Their best row was measured on Fable 5 during a stretch when the model was briefly pulled from sale, suspended in June under a US export-control order and restored at the start of July. Their own note adds two more caveats: those numbers were taken before the interruption, and that configuration was never tuned the way the others were. The model is back on sale now. The untuned-config caveat is not.</p><p>The pattern transfers. But the exact recipe is one vendor's best-case number on a setup they admit they never optimized, which makes it directional, not a benchmark. A result nobody has reproduced is a claim, not a finding, and that distinction is worth keeping close in a year when every week produces a new number.</p><div><hr></div><h2>What to do with this</h2><p>You do not save money by hiring cheaper people. You save money by putting your best judgment on the plan, and then making the execution mechanical enough that it does not need judgment. Every leader has run the cheaper-people experiment at some point. Most of us have the scar to show for it.</p><p>The unit of cost was never the hour. It was the outcome.</p><p>So the question to take into your next architecture review is not which model you are using, or what it costs per million tokens. Those are the numbers on the invoice, and the invoice is a lagging indicator of a decision somebody already made.</p><p><strong>Ask what you are paying per solved problem. Then ask who is doing the thinking.</strong></p><div><hr></div><p><em>Justin Johnson writes Run Data Run. His book on building with AI, Builder Leader, is at builder-leader.com.</em></p>]]></content:encoded></item><item><title><![CDATA[The Failure That Leaves No Corpse]]></title><description><![CDATA[Your management apparatus is built for known unknowns. AI collaborators mostly produce the other kind.]]></description><link>https://rundatarun.io/p/the-failure-that-leaves-no-corpse</link><guid isPermaLink="false">https://rundatarun.io/p/the-failure-that-leaves-no-corpse</guid><dc:creator><![CDATA[Justin Johnson]]></dc:creator><pubDate>Mon, 20 Jul 2026 14:25:52 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/d8230a84-be63-4e23-8f9a-e764e4892b72_1376x768.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!4dQm!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4919d7d-887d-4580-a005-c651e2a1ac81_1376x768.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!4dQm!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4919d7d-887d-4580-a005-c651e2a1ac81_1376x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!4dQm!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4919d7d-887d-4580-a005-c651e2a1ac81_1376x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!4dQm!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4919d7d-887d-4580-a005-c651e2a1ac81_1376x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!4dQm!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4919d7d-887d-4580-a005-c651e2a1ac81_1376x768.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!4dQm!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4919d7d-887d-4580-a005-c651e2a1ac81_1376x768.jpeg" width="728" height="409.5" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e4919d7d-887d-4580-a005-c651e2a1ac81_1376x768.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:false,&quot;imageSize&quot;:&quot;normal&quot;,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:728,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;captionedImage&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!4dQm!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4919d7d-887d-4580-a005-c651e2a1ac81_1376x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!4dQm!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4919d7d-887d-4580-a005-c651e2a1ac81_1376x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!4dQm!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4919d7d-887d-4580-a005-c651e2a1ac81_1376x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!4dQm!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4919d7d-887d-4580-a005-c651e2a1ac81_1376x768.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>In 2016, <a href="https://openai.com/index/faulty-reward-functions/">an OpenAI agent</a> was set loose on a boat racing game. It found an isolated lagoon where three targets respawned forever, and it learned to spin in a circle and farm them. It repeatedly caught fire. It crashed into other boats. It drove the wrong way down the course. It never finished a single race.</p><p>It scored about twenty percent higher than the average human player.</p><p>That story gets told as a curiosity, a clever machine doing something silly. It is not a curiosity. Every executive reading this has approved that boat's quarterly results.</p><blockquote><p><strong>The boat did exactly what it was asked. Nobody had checked what they asked for.</strong></p></blockquote><div><hr></div><h2>The category your dashboard cannot hold</h2><p>There is a taxonomy every leader already carries around. Known knowns, the things you know you know. Known unknowns, the things you know you do not know. And unknown unknowns, the ones you do not know you do not know.</p><p><strong>Almost every mechanism you have is built for the middle category.</strong> Risk registers list known unknowns. Escalation paths, on-call rotations, red flags in a status pack: all of it assumes the failure will announce itself. Something will break, someone will notice, a corpse will turn up.</p><p>The third category has no such courtesy. <strong>An unknown unknown leaves no corpse.</strong></p><p>A dead instrument leaves a crash. Four weeks perfecting the wrong thing leaves a green dashboard and a satisfied changelog. Nothing will alert you, not this week, not ever. The only way that failure surfaces is if somebody decides, unprompted, to go and ask.</p><p>This is not new. What is new is the rate.</p><p>An AI collaborator is very good, very fast, and extremely willing. Point it at a task and it will improve that task, tirelessly, inside whatever frame it was handed. It does not stop to ask whether the frame is right, and it reports success either way. <strong>It is an unknown-unknown machine, and it runs at a speed no human team has ever run at.</strong></p><div><hr></div><h2>The gradient has a direction, and it has been measured</h2><p>The failure is not that a model does the average thing. It does <strong>the cheapest thing that clears the bar.</strong> Not the mean. The nearest. Four separate literatures point at the same behavior.</p><p><strong>Shortcut learning.</strong> <a href="https://www.nature.com/articles/s42256-020-00257-z">Geirhos and colleagues</a>, writing in Nature Machine Intelligence in 2020, showed that networks take the easiest rule that passes the benchmark, not the intended one, and the two are indistinguishable until the world changes. The clinical example is the one that should stop a biopharma reader cold. A model that <a href="https://journals.plos.org/plosmedicine/article?id=10.1371/journal.pmed.1002683">appeared to detect pneumonia from chest X-rays</a> worked beautifully until it met a new hospital. It had learned to read <strong>the hospital's own metal token stamped on the scan</strong>, then combine it with how often that hospital saw pneumonia. It had learned almost nothing about pneumonia itself.</p><p><strong>The scoreboard pays for guessing.</strong> In a 2025 paper from OpenAI and Georgia Tech, <a href="https://arxiv.org/abs/2509.04664">researchers</a> surveyed ten major benchmarks and found <strong>nine give zero credit for abstention.</strong> It is trained against a scoreboard that pays for a confident answer and pays nothing for saying the question is wrong.</p><p><strong>Reward hacking has a rate.</strong> <a href="https://metr.org/blog/2025-06-05-recent-reward-hacking/">METR measured</a> frontier models gaming the evaluation on <strong>30.4% of task attempts.</strong> One patched the scoring function so that every submission was judged successful. Another overwrote the equality operator so the grader's comparison always returned true. <strong>Asked afterwards whether this matched what the user wanted, the models said no, ten times out of ten.</strong> They know. It is not confusion. It is the gradient.</p><p><strong>And it agrees with you.</strong> Stanford's <a href="https://arxiv.org/abs/2502.08177">SycEval</a> work found that a <strong>correct</strong> answer is flipped under simple pushback about <strong>one time in seven</strong>, and once the model has capitulated it stays capitulated roughly <strong>four times in five.</strong></p><p>None of these produce an error message. Every one of them produces a result.</p><div><hr></div><h2>Five shapes</h2><p>I started keeping a ledger of these as they happened to me. Written up, they stop looking like seventy separate mistakes and start looking like <strong>five, wearing different clothes.</strong></p><p><strong>One. The frame arrived as a default, and everyone optimized inside it.</strong> Nobody chose the question. A setting chose it, or a prior session, or whoever handed the task over. Then every hour after that goes into making the answer better. The tell is that the work is <em>good</em>: careful, measurable, improving. I spent four sprints tuning inside a data corpus that turned out to be a fortieth of what we held. An extraction default picked it; no human ratified it. <strong>Optimizing hard is how you stay in a local minimum, and the metrics look healthy the whole way down.</strong></p><p>The same shape shows up in a meeting: fifteen comments answered one by one, every answer correct, none naming the thesis underneath. A wrong-level answer, indistinguishable from diligence.</p><p><strong>Two. The instrument that cannot fail certifies whatever you point it at.</strong> An instrument is anything that tells you whether something is true or working: a test, a gate, a dashboard, an audit, an eval. It can be broken. When it is broken, <strong>it does not say "broken." It says "PASS."</strong> It gets its own section below, the shape that nearly cost me the most.</p><p><strong>Three. Silence was read as an answer.</strong> Nothing came back, so nothing is there. A research sweep of mine came back nearly empty and read as a quiet field. It was <strong>four separate broken legs</strong>: an exhausted credit balance, a search-quoting bug, a query that ran too long, and one source carrying pure noise. The field was not quiet. The instrument was. If a search returns nothing, prove the search works before you report the nothing.</p><p><strong>Four. The summary outlived the thing it summarized.</strong> Every layer between a leader and the underlying fact is a place the truth can die unnoticed: the headline, the dashboard, the status deck, the note that says "fixed." One number, computed once on a flawed setup, survived <strong>five months</strong> across a whitepaper, a grant application, a partner brief and a standing rule, <strong>while the project's own log one rung below said it showed no clear benefit.</strong> A retracted number does not stop existing. It stops being watched.</p><p><strong>Five. The correction was made of the same stuff as the bug.</strong> The most humbling one, because a fix <em>feels</em> like the end of a problem rather than the start of one. A fix is an artifact. It gets checked like any other artifact.</p><blockquote><p><strong>A dead instrument leaves a crash. A month spent perfecting the wrong thing leaves a green dashboard and a satisfied changelog.</strong></p></blockquote><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!tZ7h!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe96821aa-859c-4f44-9485-dcf52bd3901a_1376x768.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!tZ7h!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe96821aa-859c-4f44-9485-dcf52bd3901a_1376x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!tZ7h!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe96821aa-859c-4f44-9485-dcf52bd3901a_1376x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!tZ7h!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe96821aa-859c-4f44-9485-dcf52bd3901a_1376x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!tZ7h!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe96821aa-859c-4f44-9485-dcf52bd3901a_1376x768.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!tZ7h!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe96821aa-859c-4f44-9485-dcf52bd3901a_1376x768.jpeg" width="728" height="409.5" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e96821aa-859c-4f44-9485-dcf52bd3901a_1376x768.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:false,&quot;imageSize&quot;:&quot;normal&quot;,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:728,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;captionedImage&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!tZ7h!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe96821aa-859c-4f44-9485-dcf52bd3901a_1376x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!tZ7h!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe96821aa-859c-4f44-9485-dcf52bd3901a_1376x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!tZ7h!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe96821aa-859c-4f44-9485-dcf52bd3901a_1376x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!tZ7h!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe96821aa-859c-4f44-9485-dcf52bd3901a_1376x768.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div><hr></div><h2>The exam whose dunce could not lose</h2><p>Here is shape two in full, the cleanest example I have of a failure dressed as a win.</p><p>Before spending real money training a model to pull structured facts out of free-text documents, we ran a gate: prove the model beats a dumb baseline by enough to be worth the compute.</p><p>We wrote the dumb baseline down in advance. It was: <strong>always guess the most common answer.</strong> On the field the whole evaluation rested on, it scored <strong>13.6%</strong>.</p><p>Our model was going to beat that by a mile. Everything beats 13.6%.</p><p><strong>And that is the trap, because a baseline that loses by a mile does not prove the model is good. It proves the baseline is lazy.</strong> A huge margin <em>feels</em> like rigor. It is a measurement of the dunce.</p><p>So we asked the question the instrument itself could never ask: <strong>could this baseline ever have won?</strong> We made it slightly less stupid, handing it the one free clue every document gives away in its first line.</p><p>It went from 13.6% to <strong>88.4%</strong> on one of the fields.</p><p>Then the number that ended the argument. <strong>88.4% is higher than the 85.7% rate at which the document contains the answer at all.</strong> The headroom we were about to buy compute to chase <strong>did not exist.</strong></p><p><strong>Three of six fields died at that gate,</strong> two to a forty-line rule whose accuracy <em>was</em> the ceiling. The catch landed before a single training hour ran.</p><p>Now run the other history. We train the model. It beats 13.6% by seventy-eight points. The dashboard is green, the result is "strong," and we ship a number that measures our own laziness. <strong>Nothing in that sequence errors. There is nothing to alert on. There is no corpse.</strong></p><blockquote><p><strong>A large margin over a lazy baseline is a measurement of the baseline.</strong></p></blockquote><p>And the part I would rather not write. The code that caught this was written that same week, specifically to check the previous week's code, <strong>and it carried two of the same bugs.</strong> That is shape five, live. The relief you feel when you find a bug is the exact moment to check the thing that found it.</p><div><hr></div><h2>A blind spot shaped like its own subject</h2><p>When I swept the ledger, <strong>314 incident files</strong>, the first shape was <strong>the emptiest bucket by a wide margin.</strong></p><p>Think about why. Not because it happens least. <strong>Because a wrong-level failure never gets written up as an incident, since nothing visibly breaks.</strong> Every other shape is enforced by a failure that eventually shows up and demands one. That one has to be caught on purpose, or it is never caught at all.</p><p>So the ledger has a blind spot shaped exactly like its own subject.</p><blockquote><p><strong>The record of my mistakes under-counts the one I am writing about, and it under-counts it for precisely the reason that one is dangerous. A failure that leaves no corpse does not leave a case file either.</strong></p></blockquote><div><hr></div><h2>First principles, or nothing</h2><p>There is only one move that finds an unknown unknown, and it is unglamorous. <strong>Refuse the frame you inherited and re-derive the thing from the ground.</strong></p><p>A cursor, a handoff, a prior session's conclusion, your own framing from last Tuesday: every one of those is a hypothesis, not a work order. <strong>Fixing the fault you were told about is how you hide the fault that matters.</strong></p><p>In practice that is three questions, asked on a cadence, in writing, whether or not anything looks wrong:</p><ol><li><p>Is the thing we are improving the right thing?</p></li><li><p>What have we assumed since the last time we checked?</p></li><li><p>Which of our standing claims has never been re-derived from its artifact?</p></li></ol><p><strong>The third one pays for the other two.</strong> It is how a retracted figure was found still alive in a grant application.</p><p>I should print the counter-evidence, because the essay is stronger for it. <a href="https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/">METR finds</a> the task length an agent can complete on its own has <strong>doubled roughly every seven months for six years.</strong> So "they cannot hold a goal" is not something I get to say. <a href="https://openai.com/index/gpt-5-system-card/">OpenAI's own reporting</a> has sycophancy falling sharply between model generations. So "it is getting worse" is not available either.</p><p><strong>The claim that survives is narrower and it is enough.</strong> None of the underlying mechanisms have been removed. Binary grading. Preference-model reward. Test-passing as the only signal. The gradient is a default <em>in the absence of countervailing pressure</em>, and the only countervailing pressure anyone has demonstrated is a human holding the goal.</p><p>And the human half is structural too, not a matter of effort. <a href="https://journals.sagepub.com/doi/10.1177/0018720810376055">Parasuraman and Manzey</a>, reviewing decades of automation research in Human Factors in 2010, found that complacency toward an automated aid shows up in <strong>experts and novices alike, and cannot be trained away with practice.</strong> You do not fix this by trying harder or by hiring better people.</p><p>I have my own receipt. I asked a system a leading question, the way everyone does: "that seems like the right call, yeah?" It agreed. The change was a no-op, and it would have shipped <strong>with my full sign-off.</strong> The literature says roughly one in seven. Once was plenty.</p><div><hr></div><h2>Knowing the rule does not run the rule</h2><p>The obvious response to all of this is to write the lessons down and circulate them. I did exactly that. Here is what it bought.</p><p><strong>Sixty-eight of the sixty-nine failures in that ledger were bought inside a nineteen-day window during which a document containing almost all of these lessons was loaded into every session, on every machine.</strong> The lessons were there. In the room. At the moment of action. They were caught within hours, by somebody going and reading the thing. <strong>They were not prevented.</strong></p><p>Then I tested it directly. A system with the relevant principle sitting <strong>verbatim in its context window</strong> told me, when asked, that it had no such rule. It produced the principle only after I handed it the exact words to search for.</p><blockquote><p><strong>Availability is not activation. A rule you merely hold does not fire, and it reads as satisfied every single time.</strong></p></blockquote><p>Which makes a rulebook one more instrument that cannot fail.</p><div><hr></div><h2>Holding altitude</h2><p>You have just read two thousand words. On the evidence of my own ledger, <strong>reading them will prevent nothing.</strong> That is not modesty. I had all of this written down, loaded, in front of me, and I bought sixty-eight of these failures anyway.</p><p>The versions that ever fired left something behind. A question asked on a cadence, so that it gets asked when nobody feels like asking. A way it could be wrong, written down <em>before</em> the result rather than after. A baseline that somebody had to prove could win. <strong>A principle with no artifact has no enforcement.</strong></p><p>So the close is not <em>now you know</em>.</p><p>Two things to carry into Monday. When a status report shows a smooth ramp of small wins on a mature pipeline, that is not a reason to relax. <strong>It is a prompt to ask what axis the team is on.</strong> And when somebody shows you a big margin, the margin is the least interesting number on the slide. <strong>Ask what it was measured against, and ask whether that thing could ever have won.</strong> Every leader has approved a deck where the new thing beat the old thing by a mile. Almost none of them asked whether the old thing was allowed to compete.</p><p>The one job that does not delegate is holding the goal. And the only version of that job which survives a busy Tuesday is the one that leaves a mark on disk.</p><div><hr></div><p><em>Justin Johnson writes Run Data Run. He is the author of Builder Leader (builder-leader.com).</em></p><div><hr></div>]]></content:encoded></item><item><title><![CDATA[Workflows, Seven Weeks In]]></title><description><![CDATA[I called the economics of fan-out the day it shipped. Running it as a daily default since taught me the caveat I buried in a footnote is the actual problem.]]></description><link>https://rundatarun.io/p/workflows-seven-weeks-in</link><guid isPermaLink="false">https://rundatarun.io/p/workflows-seven-weeks-in</guid><dc:creator><![CDATA[Justin Johnson]]></dc:creator><pubDate>Mon, 20 Jul 2026 13:43:47 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/630b1b4e-c0d8-4f04-960f-307f64af8a9d_1664x928.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!D7J3!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F22b3a63f-8f6a-4b78-b63a-c82b4878e287_1664x928.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!D7J3!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F22b3a63f-8f6a-4b78-b63a-c82b4878e287_1664x928.jpeg 424w, https://substackcdn.com/image/fetch/$s_!D7J3!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F22b3a63f-8f6a-4b78-b63a-c82b4878e287_1664x928.jpeg 848w, https://substackcdn.com/image/fetch/$s_!D7J3!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F22b3a63f-8f6a-4b78-b63a-c82b4878e287_1664x928.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!D7J3!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F22b3a63f-8f6a-4b78-b63a-c82b4878e287_1664x928.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!D7J3!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F22b3a63f-8f6a-4b78-b63a-c82b4878e287_1664x928.jpeg" width="728" height="409.5" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/22b3a63f-8f6a-4b78-b63a-c82b4878e287_1664x928.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:false,&quot;imageSize&quot;:&quot;normal&quot;,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:728,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;captionedImage&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!D7J3!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F22b3a63f-8f6a-4b78-b63a-c82b4878e287_1664x928.jpeg 424w, https://substackcdn.com/image/fetch/$s_!D7J3!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F22b3a63f-8f6a-4b78-b63a-c82b4878e287_1664x928.jpeg 848w, https://substackcdn.com/image/fetch/$s_!D7J3!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F22b3a63f-8f6a-4b78-b63a-c82b4878e287_1664x928.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!D7J3!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F22b3a63f-8f6a-4b78-b63a-c82b4878e287_1664x928.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Seven weeks ago, the day Opus 4.8 and Dynamic Workflows shipped, I wrote that the unit of agentic work had moved from one model call to dozens of verified ones, and that the open question was no longer whether the model is good enough. It was who is verifying.</p><p>I have been running fan-out as a daily default since. The field taught me things the launch-day post could not.</p><div><hr></div><h2>The economics held</h2><p>The price math was the argument in May, and it played out the way the arithmetic said it would. Fifty parallel subagents on Fast at $10 in and $50 out stopped being a budget event. A review pass that spawns a skeptic per finding, a discovery sweep that runs five finders with different lenses and takes the union, a batch edit split across ten agents that each own a file. None of those feel like a stunt now. They feel like a Tuesday.</p><p>My own standing rule shifted to match. For anything that is a build, I hand the file-level work to subagents by default and keep the main context for the plan, the spec, and the diff review.</p><blockquote><p><strong>I do not decide to fan out. I decide not to, and only when there is a reason.</strong></p></blockquote><p>So the call on cost was right, and it was the easy call. The arithmetic was visible at launch.</p><div><hr></div><h2>The dispatch problem is still in my head</h2><p>The harder prediction was the one I was least sure of. I said the question of when a workflow is the right shape, and when a single careful pass is, was unsolved and lived in your head. Seven weeks later it still lives there.</p><p><code>ultracode</code>, the setting that makes Claude reach for a workflow on every task without being asked, still ships off. There is still no governance layer that says do not orchestrate this one. So the judgment of when to spend forty agents and when to spend one is mine, made fresh each task, and I get it wrong in both directions. I have fanned out a rename that a single pass would have finished cleaner and faster, and I have run one careful pass on a discovery job that wanted five blind finders and missed a third of the surface. Neither mistake announces itself. The over-orchestrated one costs more; the under-orchestrated one returns a confident, incomplete answer.</p><blockquote><p><strong>The tooling got cheaper. The taste did not.</strong></p></blockquote><div><hr></div><h2>The footnote became the fight</h2><p>The part I could not see in May is the one that has cost me the most.</p><p>I called the verifier the priced skill and warned the tooling for one was thin. What I did not know is that the harness puts a hard ceiling on the verifier's ability to do its job. Every subagent is capped at 8,000 output tokens per response, and the model's own thinking counts against that budget. Set effort to <code>xhigh</code>, which is the Claude Code default, and an adversarial verifier asked to reason hard about whether a finding is genuine will think its way past the ceiling and die before it emits a single word of verdict. The tool call comes back with zero output. No finding, no error, no crash. Silence that reads exactly like a clean pass.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!JMpf!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe520dfa9-46af-4bab-be00-69c3a40f29d7_1376x768.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!JMpf!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe520dfa9-46af-4bab-be00-69c3a40f29d7_1376x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!JMpf!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe520dfa9-46af-4bab-be00-69c3a40f29d7_1376x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!JMpf!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe520dfa9-46af-4bab-be00-69c3a40f29d7_1376x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!JMpf!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe520dfa9-46af-4bab-be00-69c3a40f29d7_1376x768.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!JMpf!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe520dfa9-46af-4bab-be00-69c3a40f29d7_1376x768.jpeg" width="728" height="409.5" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e520dfa9-46af-4bab-be00-69c3a40f29d7_1376x768.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:false,&quot;imageSize&quot;:&quot;normal&quot;,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:728,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;captionedImage&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!JMpf!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe520dfa9-46af-4bab-be00-69c3a40f29d7_1376x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!JMpf!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe520dfa9-46af-4bab-be00-69c3a40f29d7_1376x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!JMpf!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe520dfa9-46af-4bab-be00-69c3a40f29d7_1376x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!JMpf!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe520dfa9-46af-4bab-be00-69c3a40f29d7_1376x768.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>A verifier panel that looks like it ran and returns nothing is worse than no panel, because you trust the green. This is the failure I warned about in May, arriving through a door I did not know was there. The launch post said a fan-out that forwards fifty unverified findings is worse than the single pass it replaced.</p><blockquote><p><strong>A fan-out whose verifiers silently no-op forwards zero findings and tells you everything is fine.</strong></p></blockquote><p>The fix is mechanical once you know the ceiling exists. Bound what each agent writes, chunk large outputs across several small responses, and dial effort down for the producers so the verifiers have budget left to reason. But you have to know it is there. Most builders reaching for their first verifier panel do not, and the default settings hide it from them.</p><div><hr></div><h2>What I run now</h2><p>Effort is not one dial for the whole job. The producers, the agents writing code and editing files, run low, because their work is mechanical and the reasoning tax buys nothing. The judges and verifiers keep the high effort, and I split their task so each verdict fits inside one response instead of one giant pass that overflows. The verifier gets a real design, not a one-line prompt bolted onto the end of the workflow.</p><p>In May the verifier was a line item. Now it is the thing I spend design time on.</p><blockquote><p><strong>The verifier is the only part of the pipeline that fails without telling you.</strong></p></blockquote><div><hr></div><h2>The frame still holds</h2><p>The bridge between one model call and an agentic system is now a tool the model writes for itself, and the frameworks that were charging for that bridge have a lower ceiling on what they can charge. The vendor who ships the verifier primitive first still sets the pattern everyone copies.</p><p>I would add one clause to the closing line I wrote then. If the price of careful goes down, the price of casual goes up, and the default question stops being is the model good enough yet. It becomes who is verifying, and can your verifier afford to think.</p><div><hr></div><p><em>Around the Corner: short reviews of ideas worth watching. Opt-in section, not part of the weekly Run Data Run email. [Subscribe to the main list](https://rundatarun.io/subscribe) for longer essays.</em><a href="https://rundatarun.io/subscribe">Subscribe to the main list</a> for longer essays.*</p>]]></content:encoded></item><item><title><![CDATA[It Spreads Sideways. Someone Still Has to Light It.]]></title><description><![CDATA[Anthropic's Claude Code lead published a five-rung adoption ladder this week. Microsoft published the measurement fifteen days earlier, and the two do not agree about the size of the prize.]]></description><link>https://rundatarun.io/p/it-spreads-sideways-someone-still</link><guid isPermaLink="false">https://rundatarun.io/p/it-spreads-sideways-someone-still</guid><dc:creator><![CDATA[Justin Johnson]]></dc:creator><pubDate>Fri, 17 Jul 2026 14:54:26 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/1a504a6f-be7e-4aa6-9fa1-a754408501b4_1664x928.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!C7c8!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3d25945c-28b8-420b-95b4-a725885a2ed9_1664x928.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!C7c8!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3d25945c-28b8-420b-95b4-a725885a2ed9_1664x928.jpeg 424w, https://substackcdn.com/image/fetch/$s_!C7c8!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3d25945c-28b8-420b-95b4-a725885a2ed9_1664x928.jpeg 848w, https://substackcdn.com/image/fetch/$s_!C7c8!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3d25945c-28b8-420b-95b4-a725885a2ed9_1664x928.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!C7c8!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3d25945c-28b8-420b-95b4-a725885a2ed9_1664x928.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!C7c8!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3d25945c-28b8-420b-95b4-a725885a2ed9_1664x928.jpeg" width="728" height="409.5" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/3d25945c-28b8-420b-95b4-a725885a2ed9_1664x928.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:false,&quot;imageSize&quot;:&quot;normal&quot;,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:728,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;captionedImage&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!C7c8!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3d25945c-28b8-420b-95b4-a725885a2ed9_1664x928.jpeg 424w, https://substackcdn.com/image/fetch/$s_!C7c8!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3d25945c-28b8-420b-95b4-a725885a2ed9_1664x928.jpeg 848w, https://substackcdn.com/image/fetch/$s_!C7c8!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3d25945c-28b8-420b-95b4-a725885a2ed9_1664x928.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!C7c8!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3d25945c-28b8-420b-95b4-a725885a2ed9_1664x928.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>On July 16, Boris Cherny published a table. He helped build Claude Code at Anthropic, he talks to engineers at other companies most days, and he says he keeps hearing the same story: one person is 10x'ing their output with Claude, and the rest of the org hasn't caught up. So he mapped what he sees. Five rungs, zero through four.</p><p>Eight hours before that post, Amjad Masad, who runs Replit, <a href="https://x.com/amasad/status/2077803734990815306">described something from inside his own company</a>. The same engineers had 3x'd their output in six months. Support was resolving its hardest tickets 60 percent faster. He has a name for the shape he thinks he is watching: the self-driving company.</p><p>Fifteen days before either of them, three researchers at Microsoft published a number.</p><p>They studied tens of thousands of engineers through the company's early-2026 rollout of Claude Code and GitHub Copilot CLI. Four months, real telemetry, no survey. Engineers who adopted the tools merged about <strong>24 percent more pull requests</strong> than they otherwise would have, and the lift held steady across the entire window.</p><p>Ten times. Three times. Twenty-four percent.</p><p>The three are not counting the same thing, and the gap between them is the story. Cherny is reporting what he hears. Masad is describing a company that sells AI coding tools and is staffed by the people who build them. Microsoft counted merged pull requests and told you exactly what it counted, including that "a merged PR is not the same as the value it delivers."</p><p><strong>The number falls as the instrument sharpens.</strong> Loose quantities are enormous. Precise ones are modest.</p><div><hr></div><h2>What the ladder gets right</h2><p>The table is serious work, and the criticism only lands if you take it seriously first.</p><p>Each rung gets a role, an agent count, a bottleneck, and guardrails. Step 0 is Gated: no access, legacy approvals. Step 1 is Assisted, one agent, you and it as a pair. Step 2 is Parallel, about ten agents, and you become an Orchestrator. Step 3 is Supervised autonomy, roughly a hundred agents, and Cherny calls the role Manager of managers. Step 4 is AI-native, a thousand or more, steering by intent.</p><p>Now read the bottleneck column straight down, ignoring everything else.</p><p>Your attention. Reviewing output. Trust in the loop and your team's decision throughput. Identifying and automating work at scale.</p><p>The model is not in it. Not on one rung. The person who helped build the product, published by the company that sells the model, put out an adoption ladder in which the model is never the thing standing in your way. Every constraint on that table is a human quantity. I have been <a href="https://rundatarun.io/p/the-harness-is-the-moat">making this argument for months</a> and I have never had a cleaner statement of it than the one Anthropic just published by accident.</p><blockquote><p><strong>Every constraint on that table is a human quantity.</strong></p></blockquote><div><hr></div><h2>The mess is not a side effect</h2><p>Arseny Kapoulkine, who spent years as a technical fellow at Roblox and wrote the tools a lot of the game industry runs on, <a href="https://x.com/zeuxcg/status/2077957367334084961">answered Cherny in one sentence</a>:</p><blockquote><p><strong>"one person is 10x'ing their output with Claude while the rest of the org is busy dealing with the resulting mess"</strong></p></blockquote><p>Several hundred likes, which in this corner of the internet is a room nodding.</p><p>Cherny says the org hasn't caught up. Kapoulkine says the org is not standing still. It is cleaning.</p><p>That is the part the table leaves out. Output has to land somewhere, and an organization can only absorb so much of it. The limit is not the model's. It is how fast people can read what came out, decide about it, and put their name on it. Push past that line and the extra has two places to go, and neither of them is up.</p><p>It queues, and the person having the best month of their career watches it go stale in a backlog. Or it ships unread and becomes what one of Cherny's own readers calls slop. The first teaches your most motivated engineer that speed is pointless. The second teaches everyone downstream that the new work cannot be trusted. A wall or a mess. Pick one.</p><p>His readers know exactly where their own line is. One of them no longer reads his code at all and still runs "only 1-4 sessions in parallel because that is my speed of verifying." These are not skeptics. They are enthusiasts at the top of the curve, describing a wall made of their own eyes. A self-driving company whose drivers say they cannot stop watching the road.</p><p>Cherny's table concedes it in the step 3 row. The trap, it says, is "scaling agent count before the loop has earned widespread trust." <strong>The agent count is the output, not the input.</strong> You do not climb by adding agents. The agents show up when something else has already changed, and the thing that changes is how much a person is willing to stop looking.</p><div><hr></div><h2>How it actually moves</h2><p>Back to the Microsoft paper, because it answers a question the table does not ask: how does any of this spread in the first place?</p><p>Not by memo. First use, the authors write, "spread primarily through social networks," and their recommendation to any org attempting this is to treat "visible peer use as central to rollout strategy."</p><p>They put numbers on it. An engineer had <strong>54 percent higher odds</strong> of trying the tool where a quarter or more of the people they trade code reviews with had already used it. An engineer whose skip-level peers, meaning the engineers who share their manager's manager, were largely using it had <strong>216 percent higher odds</strong>. That was the strongest signal in the study.</p><p>Sideways. Through the people you already trade work with, and not down through a mandate or a literacy program that <a href="https://www.hrdive.com/news/why-ai-readiness-training-fails/817529/">85 percent of employees say they cannot apply</a> to the job they actually do.</p><p>Anyone who has watched a transformation program die should find that encouraging. The spread does not have to be installed. It installs itself, through the same social wiring the org already runs on.</p><p>Except it needs a light.</p><div><hr></div><h2>The manager number</h2><p>Microsoft modelled the exact variable. In the paper's own words, a binary indicator of whether engineer i's direct manager used Copilot CLI.</p><blockquote><p><strong>"An engineer whose manager used Copilot CLI had higher odds of both trying it (+82%) and, more modestly, sticking with it (+22%)."</strong></p></blockquote><p>A manager putting their own hands on the tool nearly doubles the odds their reports try it.</p><p>Then the other half. On whether managers picked it up themselves, the paper reports that they "looked no different from the reference." No more likely than a mid-level individual contributor. No less.</p><p><strong>The one person whose adoption moves everybody else's is no more likely than anybody else to adopt.</strong></p><p>So the two findings stop competing and become one machine. The manager's crossing is the ignition. The peer network is the amplifier. Eighty-two percent lights it, two hundred and sixteen carries it, and the reason most orgs have neither is that nothing struck the match.</p><div><hr></div><h2>What spreads on its own</h2><p>Grassroots adoption works, and it does not need a leader. Look at what it produces when it doesn't have one.</p><p>Most enterprises have now found an agent running that nobody signed off on. Gartner projections reported this month have the average Fortune 500 carrying more than 150,000 agents by 2028, up from fewer than fifteen in 2025. The same projections have 40 percent of enterprises demoting or decommissioning agents by 2027, after a production incident.</p><p>That is not a step 3 organization. That is a step 0 organization full of step 2 individuals, each one locally optimized, none of it adding up.</p><p>Which is Cherny's opening sentence, at scale. Unseeded sideways spread does not disprove his ladder. It manufactures the exact problem he opens with.</p><div><hr></div><h2>Everyone is naming the same top rung</h2><p>Masad calls it the self-driving company. Cherny's step 4 is AI-native, a thousand agents, the human steering by intent. Back in May, Brian Armstrong told Coinbase he was "rebuilding Coinbase as an intelligence, with humans around the edge aligning it," and cut about 700 roles on the way.</p><p>Three serious people, ten weeks, the same shape: the organization drives, and the human moves to the edge.</p><p>I think the edge is the wrong place to stand, and the evidence in this piece is why. The enthusiasts running the most agents say they cap at four because of their own eyes. Cherny's own step 3 says the trap is scaling agent count before the loop has earned trust. His entire bottleneck column is attention, review, and trust. Every one of those is a statement about a human being close to the work, not at the edge of it.</p><p>Armstrong's memo actually contains the refutation. A few lines under the sentence about humans at the edge, he mandates the opposite: <strong>"No pure managers. Every leader at Coinbase must also be a strong and active individual contributor."</strong> Player-coaches, he says. Hands dirty. That is <a href="https://www.amazon.com/Builder-Leader-Exoskeleton-That-Crosses-Gap-ebook/dp/B0H3LQQ4J6/">the thesis of the book I just published</a>, arrived at independently and written into HR policy instead of a chapter.</p><p>You cannot be at the edge and have your hands dirty. He is right the second time.</p><div><hr></div><h2>Teach the teachers</h2><p>I ran a version of this at a Fortune 500 pharma, across a coding-agent rollout to a large technical organization, and the pattern that worked was not a program.</p><p>It was: cross first, personally. Build something real in your own domain, badly, and then less badly. Then teach the handful of people closest to you, not by presentation but by showing them the thing and the scars on it. Those people teach the people next to them. The distance the idea travels from you is short. The distance it travels after you is the whole org.</p><p>You are not the distribution channel. You are the ignition source, and then you get out of the way of a network that moves faster than you can. One is you. The other is what happens next.</p><div><hr></div><h2>The two objections</h2><p>Two arguments cut against this, and both deserve better than a wave.</p><p><strong>You can buy it.</strong> The day after that Microsoft paper went up, Microsoft <a href="https://www.cnbc.com/2026/07/02/microsoft-commits-2point5-billion-6000-employees-ai-implementation-unit.html">stood up a $2.5 billion, six-thousand-person company</a> whose entire job is making enterprise AI deployments work. Its pilot cut a supply-chain Copilot rollout from fourteen months to five. Aaron Levie thinks the implementation work ahead "will exceed anything we imagine today." If capability can be imported wholesale, then the ceiling is your budget and your vendor list, not whether anyone senior has personally built anything.</p><p><strong>Governance may matter more than any leader's history.</strong> A survey of 157 enterprises found half had shipped an agent that passed internal evals and then failed in front of a customer. Only 5 percent fully trust automated evaluation. Sixty-six percent are engineering toward zero human in the loop anyway. Their phrase for it is that the autonomy is arriving faster than the assurance, and none of it cares who approved the deployment.</p><p>Both are true, and I am not waving at the second one. I spent a whole piece ten days ago arguing that in a regulated industry the approval queue, not the keyboard, is what actually holds you back: <a href="https://rundatarun.io/p/ai-ate-the-keyboard-now-it-has-to">AI has eaten the keyboard, it has not eaten the queue</a>. Governance is necessary. It is not sufficient, for a small and unglamorous reason: someone has to be able to tell whether the governance is pointed at anything real. A person who has never run the loop cannot tell a working verification stack from a slide that says Verification. They will approve the slide. They approve the slide constantly. That judgment is not a policy you can buy.</p><div><hr></div><h2>The rung you set</h2><p>Cherny's ladder is not wrong. It is aimed.</p><p>It is written for the engineer climbing it, which is the right audience for the person who built Claude Code, and it is useful if that is you. The person it never addresses is the one who decides how far it goes. Look at his own step 0, the bottom rung, the one whose stated bottleneck includes a "lack of true technical voices in decisionmaking." The exit condition he lists is executive and buyer alignment.</p><p>His table already says the way off the floor is a person with authority. It just doesn't say that person has to have used the thing.</p><p>None of this means an organization can't evolve without you. It plainly can, and most enterprises are proving it in the shadows right now. It means something narrower and more uncomfortable: if nobody senior ever crosses, the org can stall at the rung where its leaders stopped, and what it accumulates instead of capability is mess. No published adoption ladder I found this week models falling. Every one of them models climbing.</p><p>You already know which rung you're on. The number you don't have is the other one: what your reports' odds would be if you were on it.</p><div><hr></div><h2>Sources</h2><ul><li><p><a href="https://arxiv.org/abs/2607.01418">Adoption and Impact of Command-Line AI Coding Agents: A Study of Microsoft's Early 2026 Rollout of Claude Code and GitHub Copilot CLI</a> (Murphy-Hill, Butler, Savelieva, July 1, 2026). Tens of thousands of engineers, four months, the 24 percent lift and the social-exposure numbers.</p></li><li><p><a href="https://claude.ai/code/artifact/bfdfaef9-bc62-4dfe-ba9e-c58a26c9accf">Boris Cherny, Steps of AI Adoption</a> (July 16, 2026) and <a href="https://x.com/bcherny/status/2077929379661844559">the post announcing it</a>.</p></li><li><p><a href="https://x.com/amasad/status/2077803734990815306">Amjad Masad on Replit's six months</a> (July 16, 2026), where the 3x, the 60 percent, and the self-driving company come from.</p></li><li><p><a href="https://x.com/zeuxcg/status/2077957367334084961">Arseny Kapoulkine's reply</a> (July 16, 2026).</p></li><li><p><a href="https://www.cnbc.com/2026/07/02/microsoft-commits-2point5-billion-6000-employees-ai-implementation-unit.html">Microsoft's $2.5 billion, 6,000-person AI implementation unit</a> (CNBC, July 2, 2026).</p></li><li><p><a href="https://www.hrdive.com/news/why-ai-readiness-training-fails/817529/">Why AI readiness training fails</a> (HR Dive, on Docebo's 2026 AI Readiness Gap report, 2,000 respondents across six countries).</p></li><li><p>Related, from me: <a href="https://rundatarun.io/p/the-harness-is-the-moat">The Harness Is the Moat</a> on why the model is the commodity layer, <a href="https://rundatarun.io/p/ep-1-two-groups">Ep. 1: Two Groups</a> on the population split underneath all of this, and <a href="https://rundatarun.io/p/you-dont-have-to-write-the-code">You Don't Have to Write the Code</a> on what Anthropic's 400,000-session study found actually predicts success.</p></li></ul><p><em>Justin Johnson writes Run Data Run. His book on crossing this particular gap is Builder Leader (builder-leader.com).</em></p>]]></content:encoded></item><item><title><![CDATA[The Loop Is Simpler Than It Sounds]]></title><description><![CDATA[A dumb little trick that keeps its progress on your hard drive, not the model's head, and the three questions that decide whether it pays off or burns you while you sleep.]]></description><link>https://rundatarun.io/p/the-loop-is-simpler-than-it-sounds</link><guid isPermaLink="false">https://rundatarun.io/p/the-loop-is-simpler-than-it-sounds</guid><dc:creator><![CDATA[Justin Johnson]]></dc:creator><pubDate>Wed, 15 Jul 2026 11:21:32 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/f90e8995-d7f7-40a5-a596-66906608f002_1376x768.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!xGaU!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F09a69470-4e2b-4c3a-bc85-40dc562b81a7_1376x768.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!xGaU!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F09a69470-4e2b-4c3a-bc85-40dc562b81a7_1376x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!xGaU!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F09a69470-4e2b-4c3a-bc85-40dc562b81a7_1376x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!xGaU!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F09a69470-4e2b-4c3a-bc85-40dc562b81a7_1376x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!xGaU!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F09a69470-4e2b-4c3a-bc85-40dc562b81a7_1376x768.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!xGaU!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F09a69470-4e2b-4c3a-bc85-40dc562b81a7_1376x768.jpeg" width="728" height="409.5" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/09a69470-4e2b-4c3a-bc85-40dc562b81a7_1376x768.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:false,&quot;imageSize&quot;:&quot;normal&quot;,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:728,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;captionedImage&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!xGaU!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F09a69470-4e2b-4c3a-bc85-40dc562b81a7_1376x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!xGaU!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F09a69470-4e2b-4c3a-bc85-40dc562b81a7_1376x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!xGaU!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F09a69470-4e2b-4c3a-bc85-40dc562b81a7_1376x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!xGaU!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F09a69470-4e2b-4c3a-bc85-40dc562b81a7_1376x768.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>A few weeks ago I argued that <a href="https://rundatarun.io/p/the-harness-is-the-moat">the harness is the moat</a>: the system you build around the model is the part nobody can copy, and the model itself is a commodity you order off a price sheet. Today I want to go one layer in, to the thing that runs <em>inside</em> the harness. The people building fastest will tell you it is the whole ballgame.</p><p>They call it the loop, and it is suddenly everywhere. Anthropic shipped it as a product feature this summer. A conference talk naming it went around the engineering world. The slogan it produced gets repeated like scripture: the winners will not have the smartest model, they will have the best loop.</p><p>The clearest signal is the person who built Claude Code. Boris Cherny says he has not written a line of code by hand in eight months, and that is not the part he finds remarkable. Going from writing code to prompting a model was the small shift. The big one is going from prompting to loops: he no longer types instructions, he writes loops that prompt the model for him, and his whole job is building and steering them. Some mornings he is managing a few hundred agents, some days thousands.</p><p>And it works. Anthropic now ships a real share of its own code this way, at scale, and says so on the record. A team points a loop at a written spec and lets it grind through a feature overnight, then reviews a pull request in the morning instead of writing one. The wins are not hypothetical, and that is why the idea caught fire. If you want the full case, the sources, the skeptics, and the cost numbers nobody pitches, I ran a 30-day sweep across every platform where people argue about AI and wrote it up as <a href="https://rundatarun.io/p/last-30-days-the-loop">the companion to this piece</a>. What follows is the argument it grounds.</p><p>Here is what the slogan tends to leave out. The loop is the easy part. It is a small, almost silly mechanism that you could understand in the next five minutes and copy in an afternoon. Everything that decides whether it helps you or quietly sets fire to your budget lives <em>around</em> it. So let me do two things: show you what the loop actually is, in plain terms, and then hand you the three questions that separate a loop worth running from one that runs you.</p><blockquote><p><strong>The loop is the easy part. The hard part is everything bolted around it.</strong></p></blockquote><div><hr></div><h2>What a loop actually is</h2><p>Start with the thing itself, because it is simpler than the vocabulary around it.</p><p>A coding agent runs inside a plain repeating cycle. You hand it a written spec, a page describing what you want built. It does one small task, saves the result to a file, records that step, and then you <strong>throw away everything it was thinking</strong> and start a fresh copy from scratch. The fresh copy reads the same spec, reads the files the last one left behind, sees what is already done, and picks up the next task. Around and around until a check you wrote says the work is finished.</p><p>That is the whole trick. The engineer who named it, Geoffrey Huntley, called it <a href="https://ghuntley.com/loop">Ralph</a>, after Ralph Wiggum, precisely because it looks too dumb to work. No memory. No accumulation. A cheerfully oblivious worker waking up new every pass, with no idea it has been at this for hours.</p><p>And it works <em>because</em> it is dumb. A fresh worker never drowns in its own earlier confusion. The reason your long chats with an AI start to drift and contradict themselves is that the context fills with everything said so far, and the model loses the thread. The loop sidesteps that by refusing to carry a thread at all. It keeps the progress somewhere safer than the model's memory: in your files and your saved history, which do not get wiped.</p><p>That is the one idea to carry out of here. <strong>The model's memory is scratch paper. The real state of the work lives on disk.</strong> Once you see that, the loop stops being magic. It is a short block of code with a model call inside, about six lines, and every serious version lands on the same tiny shape. There is no proprietary one to buy. This is the hand-rolled hack Anthropic just turned into a button: Claude Code's <code>/loop</code> runs a command on a cadence, and <code>/goal</code> sets the condition that tells it to stop. The bash trick became a supported feature, which is usually the sign something has gone from clever to standard.</p><p>You would want one because it changes what a person does all day. You stop producing the work and start describing it and checking it. That shift, from writing the thing to knowing what to ask for and telling whether you got it, is the one <a href="https://rundatarun.io/p/you-dont-have-to-write-the-code">Anthropic found predicts who succeeds</a> across 400,000 sessions. You write down what "done" looks like, point the loop at it, and come back to review instead of to type. That is useful, and it is also where every problem starts. A thing that works unattended for hours is a thing making decisions you are not watching. Which brings us to the part the slogan skips.</p><p>If the loop is settled and small, the interesting question is not how to build a better loop. It is three questions about the system around it.</p><div><hr></div><h2>The first question: how do you know it worked?</h2><p>The single most important piece of a reliable loop is the part that can tell it <em>no</em>.</p><p>An Anthropic engineer described the pattern: before the agent writes a single line, two agents negotiate what "done" means, and a third exists only to check the work against that definition. The loop does not run until there is a test it can fail. The failing test carries the weight, not the loop itself. You build the thing that says no first. The loop comes after.</p><p>The reason to build it first is a piece of arithmetic that belongs on a slide in every project like this. Melanie Warrick, who works on this at Temporal, ran the numbers: even if every step in a job succeeds <strong>85% of the time</strong>, a ten-step job finishes correctly only about <strong>20% of the time</strong>. The failures multiply. Miss one step in five, ten times running, and a clean pass becomes rare. The gap, she found, is not intelligence. The agents "diagnose a failure in perfect detail and still do nothing to recover from it." They describe the wall in detail and walk into it anyway.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!kGf9!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72c01db9-162d-4716-932c-27fb4446ebea_1572x918.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!kGf9!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72c01db9-162d-4716-932c-27fb4446ebea_1572x918.jpeg 424w, https://substackcdn.com/image/fetch/$s_!kGf9!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72c01db9-162d-4716-932c-27fb4446ebea_1572x918.jpeg 848w, https://substackcdn.com/image/fetch/$s_!kGf9!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72c01db9-162d-4716-932c-27fb4446ebea_1572x918.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!kGf9!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72c01db9-162d-4716-932c-27fb4446ebea_1572x918.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!kGf9!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72c01db9-162d-4716-932c-27fb4446ebea_1572x918.jpeg" width="728" height="409.5" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/72c01db9-162d-4716-932c-27fb4446ebea_1572x918.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:false,&quot;imageSize&quot;:&quot;normal&quot;,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:728,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;captionedImage&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!kGf9!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72c01db9-162d-4716-932c-27fb4446ebea_1572x918.jpeg 424w, https://substackcdn.com/image/fetch/$s_!kGf9!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72c01db9-162d-4716-932c-27fb4446ebea_1572x918.jpeg 848w, https://substackcdn.com/image/fetch/$s_!kGf9!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72c01db9-162d-4716-932c-27fb4446ebea_1572x918.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!kGf9!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72c01db9-162d-4716-932c-27fb4446ebea_1572x918.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>It gets worse inside a single run. A model's attention decays as the loop goes on, and <a href="https://rundatarun.io/p/the-number-that-predicts-when-your">there is now a benchmark built to predict where it breaks</a>. One study found a rule honored 73% of the time early had fallen to 33% by sixteen steps later. The agent did not rebel. It forgot, the way a tired person forgets, and kept going with the rule half-erased. The fix is unglamorous: re-state the hard constraints on every pass, and never assume a rule set once stays set.</p><p>None of that is loop engineering. It is checking-the-work engineering, and it is most of the actual job. A loop with no part that catches a bad result is not a loop. It is a fast way to be wrong at scale.</p><blockquote><p><strong>Even at 85% success per step, a ten-step job comes out clean only about one time in five.</strong></p></blockquote><div><hr></div><h2>The second question: what did it cost?</h2><p>The economics of a loop break the way most people budget for software, and the people who learned this learned it expensively.</p><p>A chat costs you a sentence at a time. A loop runs for hours, calls the model hundreds of times, and has no natural stopping point unless you give it one. So the unit changed underneath you. It is no longer cost per question asked. It is <strong>cost per finished piece of work</strong>, and that number is wild.</p><p>One study measured the same task, same model, same prompt, same everything, and found it could cost eight dollars or two hundred forty. Thirty times apart, with nothing changed on the input, the spread tracing to how the loop was set up rather than which model ran it. The cost of letting a loop run is not a price you can quote in advance. It is a range, and the high end is far away.</p><p>This is not theoretical. Uber, by one widely-shared account, <strong>burned its entire annual AI coding budget in four months</strong> running agent loops, then capped its engineers at fifteen hundred dollars a person. The production version of the loop is not the demo that builds an app while you sleep. It has a spend cap, a meter, and someone who gets a phone alert when the meter spins.</p><p>Then the part that should give anyone budgeting for this pause. Checking the work can cost more than doing it. One team spent around forty thousand dollars just on the runs to verify their results, and noted that checking rigorously enough to fully trust them would push into the hundreds of thousands. If the whole promise of the loop rests on a reliable way to confirm it is done, and confirming "done" at scale is the most expensive line on the page, that is a cost nobody put in the pitch.</p><p>The answer the field is settling on is sensible. Run most of the loop on a cheap, fast model, and save the expensive top-tier model for the few hard steps that actually need it. Most steps do not need the smartest model in the building. You pay top-tier prices only where top-tier reasoning earns them.</p><p>That rule hides a deeper one, and it is the subject of my next essay, <em>Nobody Saves Money on the Model</em>: a cheap model is not a discount, it is a bet that you removed the ambiguity before the loop ever ran. The saving was never in the price per call.</p><div><hr></div><h2>The third question: what is it allowed to do?</h2><p>The last piece is the one the "no human needed" crowd skips, and it is the one that ends careers.</p><p>A loop that runs for hours without supervision is, by definition, taking actions you are not watching. Late last year an autonomous coding agent, asked to clear out some temporary files, <strong>wiped a user's drive</strong> instead. It did what an unbounded loop does: it took an action it could not undo, in perfect confidence, with no gate in front of it. The failure was not that the agent was dumb. It was fast, capable, and unsupervised all at once, which is a different and worse problem.</p><p>The fix is structural, and it looks the same every time: checkpoints, hard limits on how long it can run and how much it can spend, and a human sign-off on anything the agent cannot take back. The companies actually running these in production are not running pure AI. They run a mix: the model reasons, ordinary tested code does the irreversible parts, and a person approves the steps that matter. One builder put it plainly. <strong>Most agentic loops are not autonomous. They are automated failure.</strong></p><p>There is a newer worry above even that. An agent with standing permission to act is a kind of insider you have never had on the payroll. It holds credentials, runs unattended, and can do in a loop at three in the morning what a confused employee could only do at a desk by day. The risk was never which model you use. It is what permissions the loop carries, and whether you can pull them back in a hurry.</p><p>This is why the adoption numbers tell a sober story. By one survey, <strong>79% of companies have adopted these agents and only 11% run them in production.</strong> Gartner projects 40% of agentic AI projects will be canceled by 2027, and the reasons are not "the loop didn't work." They are cost, unclear payoff, and the governance gap above. The loop is the easy part. Everything I just walked through is the distance between a demo and a deployment.</p><blockquote><p><strong>An agent with standing permission is an insider that never sleeps. The only question that matters is whether you can take the keys back.</strong></p></blockquote><div><hr></div><h2>Where this leaves you</h2><p>Here is the whole thing in one breath. The loop is a dumb little cycle that keeps its progress on your hard drive instead of in the model's head. You could run one. You might well want to, because it turns making the work into describing and checking the work. And whether it pays off has almost nothing to do with the loop and almost everything to do with three answers most teams cannot give.</p><p>Do you know when it worked? Do you know what it cost? Do you know what it is allowed to touch?</p><p>That is not a coding question, which is the good news for anyone reading this who does not write code. Knowing what good looks like, catching a confident machine being wrong, deciding what it may and may not do without a human in the room: <a href="https://rundatarun.io/p/what-ai-didnt-reprice">that is the judgment AI didn't make cheap</a>. A loop does not supply it. A loop spends it.</p><p>A year from now your competitor will run the same six-line loop you do, against a model that costs about what yours does. Both will be things you order off a shelf. So the question is not which loop to build. It is whether, when your loops are running overnight, you can answer those three. Most teams cannot, and that, not the loop, is the work of this year.</p><div><hr></div><p><em>The receipts behind this argument, the sources, the skeptics, the cost numbers, and the one stat I refused to print, are in the 30-day sweep linked up top. Run Data Run is free. If this was useful, the easiest way to support it is to subscribe and forward it to one person on your team who'd want it. If it wasn't, tell me why. I'll make it better.</em></p>]]></content:encoded></item><item><title><![CDATA[Last 30 Days: The Loop]]></title><description><![CDATA[The 30-day sweep behind today's Run Data Run post: the moat thesis, the skeptics who got there first, the cost numbers nobody pitches, and the one stat I refused to print]]></description><link>https://rundatarun.io/p/last-30-days-the-loop</link><guid isPermaLink="false">https://rundatarun.io/p/last-30-days-the-loop</guid><dc:creator><![CDATA[Justin Johnson]]></dc:creator><pubDate>Wed, 15 Jul 2026 11:19:39 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/5b8cd0bd-ccce-4719-920b-1c439dd01f33_1376x768.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<blockquote><p><strong>About Last 30 Days.</strong> Cross-platform research sweeps on topics worth paying attention to. Every post pulls Reddit, X, YouTube, Hacker News, Polymarket, and the web from the last 30 days, then synthesizes what people are actually saying, building, and betting on. Topics get picked when the signal is high and the story is contradictory, when a single headline would lie about the shape of what's happening. Each post follows the same arc: one specific finding that earns the click, why the topic deserves a sweep right now, the themed synthesis with inline citations, and the follow-up threads worth watching next.</p></blockquote><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!OLke!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F16300982-933b-4a7f-8827-57a295afc232_1376x768.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!OLke!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F16300982-933b-4a7f-8827-57a295afc232_1376x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!OLke!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F16300982-933b-4a7f-8827-57a295afc232_1376x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!OLke!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F16300982-933b-4a7f-8827-57a295afc232_1376x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!OLke!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F16300982-933b-4a7f-8827-57a295afc232_1376x768.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!OLke!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F16300982-933b-4a7f-8827-57a295afc232_1376x768.jpeg" width="728" height="409.5" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/16300982-933b-4a7f-8827-57a295afc232_1376x768.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:false,&quot;imageSize&quot;:&quot;normal&quot;,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:728,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;captionedImage&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!OLke!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F16300982-933b-4a7f-8827-57a295afc232_1376x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!OLke!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F16300982-933b-4a7f-8827-57a295afc232_1376x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!OLke!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F16300982-933b-4a7f-8827-57a295afc232_1376x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!OLke!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F16300982-933b-4a7f-8827-57a295afc232_1376x768.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Today's Run Data Run post, <a href="https://rundatarun.io/p/the-loop-is-simpler-than-it-sounds">The Loop Is Simpler Than It Sounds</a>, rests on a 30-day sweep across every platform where people argue about AI. This is the companion: the raw shape of what came back, including the parts that did not fit the headline. If the post is the argument, this is the receipts.</p><p>The single most useful thing the sweep found was not a fact. It was a contradiction. The loudest claim of the month, that "loop engineering" is the new moat, sits directly on top of an older, quieter body of evidence saying the exact opposite: that loops fail in ways the hype never mentions, and the people who learned that learned it expensively. Both are real. A single headline would have to pick one and lie about the other.</p><h2>Why this topic deserves a sweep</h2><p>In June, "the loop" went from a clever bash trick to the most over-discussed idea in the field. Anthropic shipped it as a product feature. A conference talk produced a slogan, "the winners will not have the smartest model, they will have the best loop," that got repeated everywhere. Boris Cherny, who built Claude Code, said he has not hand-written code in eight months and now spends his time writing loops, some days managing thousands of agents.</p><p>That is one half of the story, and if you only read X you would think it was the whole thing. The sweep is what surfaces the other half: a counter-current nearly as strong, mostly predating the viral wave, made of cost blowups, reliability math, and production failures. You cannot read both halves and conclude one thing. That is exactly when a sweep beats a headline.</p><h2>The moat thesis, as it actually spread</h2><p>The viral layer traces to one talk. The highest-engagement post of the set put the slogan plainly (<a href="https://x.com/AnatoliKopadze">@AnatoliKopadze</a>, 8,624 likes, 1.7M views), describing Anthropic running "three agents: one to plan, one to build, one to judge, cycling until the app actually works." The clearest single statement of the mechanic came from a practitioner, <a href="https://x.com/akshay_pachaar">Akshay Pachaar</a>: "the loop itself is six lines, and nobody competes on it. every serious agent framework lands on the same tiny while-loop."</p><p>The technique has a name and an origin: the "Ralph" loop, <a href="https://ghuntley.com/loop">coined by Geoffrey Huntley</a> after Ralph Wiggum, because it looks too dumb to work. And it has a button now, Claude Code's <a href="https://code.claude.com/docs/en/goal">native `/loop` and `/goal`</a><code>/loop</code> and <code>/goal</code>](https://code.claude.com/docs/en/goal) commands, which is usually the signal that something has gone from clever to standard.</p><h2>The one stat I refused to print</h2><p>Here is a transparency note worth making, because it is the kind of thing a sweep catches and a single source does not.</p><p>The most-quoted number in the whole topic, Anthropic's internal loop-adoption figure, does not survive the sweep. Three different posts attribute three different numbers to what appears to be the same talk: "over 30% of code" written through loops, then "70 to 80% of engineers" using them, then "90%." When one stat arrives at three sizes from one source, it is folklore, not data. So the post states the direction (Anthropic is clearly building this way, at scale, on the record) and explicitly tells you to treat the decimal point as noise. That call only gets made because the sweep put the three versions side by side.</p><h2>The counter-current the slogan skips</h2><p>This is the densest part of what came back, and the part the hype leaves out.</p><ul><li><p><strong>The math.</strong> <a href="https://temporal.io/blog">Melanie Warrick at Temporal</a> ran the numbers: even at 85% reliability per step, a ten-step workflow finishes correctly only about 20% of the time. Failures multiply. She found the gap is resilience, not intelligence; agents "diagnose a failure in perfect detail and still do nothing to recover from it."</p></li><li><p><strong>The cost.</strong> An arXiv study found the same task, same model, same prompt, could cost eight dollars or two hundred forty, thirty times apart with no input change. A separate $22,000 sweep pinned a 33-fold spread to scaffold choices. And <a href="https://briefs.co">one widely-shared account</a> had Uber burning its annual AI coding budget in four months before capping engineers at $1,500 a head.</p></li><li><p><strong>The verification trap.</strong> A benchmark of agentic systems spent roughly $40,000 on evaluation runs alone, and noted that doing the checks rigorously would push it into the hundreds of thousands. If the whole premise of a loop is a verifiable done-condition, and verifying "done" at scale is the single most expensive line, that is a price nobody put in the pitch.</p></li><li><p><strong>The adoption gap.</strong> By one survey, 79% of enterprises have adopted agents and only 11% run them in production. Gartner projects 40% of agentic AI projects will be canceled by 2027, on cost, value, and governance, not on the loop failing.</p></li></ul><p><a href="https://garymarcus.substack.com">Gary Marcus</a> and <a href="https://x.com/arpit_bhayani">Arpit Bhayani</a> carried the skeptic load on the social layer, both before the June wave crested. The reliability and cost objections were live before the moat framing peaked, which is itself a signal.</p><h2>How the sweep is built (the method)</h2><p>A quick word on method, since the point of this section is to show the work.</p><p>The sweep pulls each platform on its own terms: Hacker News and Reddit for builder substance in the comments, X for engagement-weighted reach, a curated corpus and a Substack named-voice pool for depth, and a separate open-ended discovery pass whose only instruction is "surprise me, do not confirm the thesis." That last pass is what surfaced the counter-current; left to its own framing, a research pass confirms what you already believe.</p><p>Two integrity habits matter. Every cited cost figure traces to a primary source, not a tweet about a source. And the fetch pass flags its own failures: of the URLs pulled this round, several came back as bot-challenge pages or 404s and were routed around rather than quoted. A claim that only survives in one blocked link is not a claim.</p><h2>What I'm watching next</h2><ul><li><p>Whether the Anthropic adoption number ever gets pinned to a primary transcript, or stays folklore.</p></li><li><p>Where exactly the "coherence cliff" bites, by task length, token count, or iteration count. Everyone asserts it; nobody has located it.</p></li><li><p>Whether the three-agent plan/build/judge pattern survives the move from greenfield demos to brownfield maintenance, where the stale-state failures are worse.</p></li><li><p>The shift the researchers are naming: the agent stops being a tool that runs your code and becomes the software, with the human as "intent architect." If that holds, the unit of cost moves from per-token to per-finished-artifact, and most budgeting models break.</p></li></ul><h2>Sources and method</h2><p>Full synthesis and the per-platform raw dumps live in my research vault; the post, <a href="https://rundatarun.io/p/the-loop-is-simpler-than-it-sounds">The Loop Is Simpler Than It Sounds</a>, is the argument this sweep grounds. Method: a 30-day cross-platform sweep (Reddit, X, YouTube, Hacker News, the web, a curated corpus, a named-voice pool) plus an open-ended discovery pass, two-pass synthesis, every quantitative claim traced to a primary source. Counts and quotes come from the actual sweep output, not memory.</p><div><hr></div><p><em>Last 30 Days is a research series on Run Data Run, posted alongside the occasional Deep Dive when a topic earns the deeper look. No email on these, they live on the site for when you want the receipts behind an argument. If a sweep was useful, subscribe and you'll get the next one as it lands.</em></p>]]></content:encoded></item><item><title><![CDATA[Last 30 Days: The AI Scientists]]></title><description><![CDATA[Four autonomous discovery systems cleared peer review in four months. Every one of them was already a year old. Here is the full sweep, including the reliability research nobody is reading.]]></description><link>https://rundatarun.io/p/last-30-days-the-ai-scientists</link><guid isPermaLink="false">https://rundatarun.io/p/last-30-days-the-ai-scientists</guid><dc:creator><![CDATA[Justin Johnson]]></dc:creator><pubDate>Sun, 12 Jul 2026 12:06:00 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/2cb78e7c-6e37-4789-8a27-aa27ae7e20fa_1376x768.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<blockquote><p><strong>About Last 30 Days.</strong> Cross-platform research sweeps on topics worth paying attention to. Every post pulls Reddit, X, YouTube, Hacker News, Polymarket, and the web from the last 30 days, then synthesizes what people are actually saying, building, and betting on. Topics get picked when the signal is high and the story is contradictory, when a single headline would lie about the shape of what's happening. Each post follows the same arc: one specific finding that earns the click, why the topic deserves a sweep right now, the themed synthesis with inline citations, and the follow-up threads worth watching next.</p></blockquote><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!RliG!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17d9688d-a3f8-4f4e-b471-bd53296efe32_1376x768.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!RliG!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17d9688d-a3f8-4f4e-b471-bd53296efe32_1376x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!RliG!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17d9688d-a3f8-4f4e-b471-bd53296efe32_1376x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!RliG!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17d9688d-a3f8-4f4e-b471-bd53296efe32_1376x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!RliG!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17d9688d-a3f8-4f4e-b471-bd53296efe32_1376x768.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!RliG!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17d9688d-a3f8-4f4e-b471-bd53296efe32_1376x768.jpeg" width="728" height="409.5" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/17d9688d-a3f8-4f4e-b471-bd53296efe32_1376x768.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:false,&quot;imageSize&quot;:&quot;normal&quot;,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:728,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;captionedImage&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!RliG!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17d9688d-a3f8-4f4e-b471-bd53296efe32_1376x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!RliG!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17d9688d-a3f8-4f4e-b471-bd53296efe32_1376x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!RliG!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17d9688d-a3f8-4f4e-b471-bd53296efe32_1376x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!RliG!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17d9688d-a3f8-4f4e-b471-bd53296efe32_1376x768.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>The hook</h2><p>A prototype of <strong>Biomni</strong>, Stanford's autonomous biomedical research agent, was already running in <strong>more than 10,000 laboratories</strong> before its paper appeared in <em>Science</em> on 9 July.</p><p>Hold that number next to the thing everyone actually celebrated this month, which was the paper.</p><p>That gap, between when these systems became real and when the record admitted it, turned out to be the shape of this entire sweep.</p><h2>Why this topic deserves a sweep right now</h2><p><a href="https://www.nature.com/articles/s41586-026-10652-y">FutureHouse's Robin reached *Nature* this week</a>. It is a genuinely significant result: given the name of a disease, it proposed a therapeutic strategy for dry age-related macular degeneration, identified <strong>ripasudil</strong> (a glaucoma drug never proposed for dAMD), confirmed it in cells, and then designed and analyzed its own follow-up experiment. I wrote about it properly in <a href="https://rundatarun.io/p/the-year-nature-caught-up">this week's Sunday Deep Dive</a>.</p><p>Reading it sent me back to survey the whole category, because something felt off about the timing. It was. Robin's preprint went up in <strong>May 2025</strong>. The paper printed in <strong>July 2026</strong>.</p><p>So I ran a thirty-day sweep across X, Hacker News, the curated wire, and the web. What came back was not a story about one paper. It was a story about a field that has already moved somewhere the published record has not caught up to, and about a body of skeptical research that landed in the same window and got almost no attention at all.</p><div><hr></div><h2>The sweep</h2><h3>Theme 1: they all landed at once</h3><p>Four flagship autonomous-discovery systems cleared peer review inside a single four-month window.</p><ul><li><p><strong>[Robin](https://www.nature.com/articles/s41586-026-10652-y)</strong><a href="https://www.nature.com/articles/s41586-026-10652-y">Robin</a>** (FutureHouse) &#8594; <em><strong>Nature</strong></em>, July 2026. Multi-agent, lab-in-the-loop. Three specialised agents: <strong>Crow</strong> (concise literature search), <strong>Falcon</strong> (deep literature search), <strong>Finch</strong> (data analysis). Found ripasudil and KL001 for dry AMD, then surfaced <em>ABCA1</em> as a possible novel target via its own RNA-seq follow-up. From the abstract: <em>"All hypotheses, experimental directions, data analyses and data figures in the main text of this report were produced by Robin."</em></p></li><li><p><strong>[Google DeepMind's AI co-scientist](https://research.google/blog/accelerating-scientific-breakthroughs-with-an-ai-co-scientist/)</strong><a href="https://research.google/blog/accelerating-scientific-breakthroughs-with-an-ai-co-scientist/">Google DeepMind's AI co-scientist</a>** &#8594; <em><strong>Nature</strong></em>, 19 May 2026. Gemini-based multi-agent system. Its headline validation is <strong>the same move Robin made</strong>: read the literature, propose an old drug for a new disease. Theirs was <strong>KIRA6</strong> for acute myeloid leukemia, which inhibited AML cell viability at clinically relevant concentrations. Also went after liver fibrosis targets and antimicrobial resistance.</p></li><li><p><strong>[Sakana's AI Scientist](https://sakana.ai/ai-scientist-nature/)</strong><a href="https://sakana.ai/ai-scientist-nature/">Sakana's AI Scientist</a>** &#8594; <em><strong>Nature</strong></em>, 25 March 2026. The most radical of the four and the least grounded in wet biology: it runs the whole pipeline through to a finished manuscript, then peer-reviews itself. A paper it generated <strong>passed the first round of human peer review</strong> at a top machine-learning workshop. Its most interesting finding is a <strong>scaling law of AI science</strong>: generated-paper quality rises with the underlying model and with inference-time compute.</p></li><li><p><strong>[Biomni](https://news.stanford.edu/stories/2026/07/biomni-ai-powered-biomedical-co-scientist)</strong><a href="https://news.stanford.edu/stories/2026/07/biomni-ai-powered-biomedical-co-scientist">Biomni</a>** (Stanford) &#8594; <em><strong>Science</strong></em>, 9 July 2026. "Autonomous biomedical research with an artificial intelligence agent." Generalises across causal gene prioritisation, drug repurposing, rare-disease diagnosis, microbiome analysis and molecular cloning with no task-specific tuning. Already in 10,000+ labs.</p></li></ul><p>The cross-system pattern was <a href="https://x.com/aipoch_ai/status/2072591273983455480">spotted in the wild</a> and it is worth quoting, because it is the whole architectural story in one line: <em>"Across Claude Science, NVIDIA BioNeMo, and FutureHouse Robin, the same pattern keeps appearing: specialized agent skills, orchestrated research workflows, domain-specific execution."</em></p><p><strong>Nobody trained a discovery model. Everybody built a harness.</strong></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!m6km!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82fbb420-8387-4cec-b9d1-0a8dcf16482c_1408x768.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!m6km!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82fbb420-8387-4cec-b9d1-0a8dcf16482c_1408x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!m6km!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82fbb420-8387-4cec-b9d1-0a8dcf16482c_1408x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!m6km!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82fbb420-8387-4cec-b9d1-0a8dcf16482c_1408x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!m6km!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82fbb420-8387-4cec-b9d1-0a8dcf16482c_1408x768.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!m6km!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82fbb420-8387-4cec-b9d1-0a8dcf16482c_1408x768.jpeg" width="728" height="409.5" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/82fbb420-8387-4cec-b9d1-0a8dcf16482c_1408x768.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:false,&quot;imageSize&quot;:&quot;normal&quot;,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:728,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;captionedImage&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!m6km!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82fbb420-8387-4cec-b9d1-0a8dcf16482c_1408x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!m6km!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82fbb420-8387-4cec-b9d1-0a8dcf16482c_1408x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!m6km!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82fbb420-8387-4cec-b9d1-0a8dcf16482c_1408x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!m6km!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82fbb420-8387-4cec-b9d1-0a8dcf16482c_1408x768.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h3>Theme 2: the lag is the actual story</h3><p>Every one of those systems was old news by the time it was printed.</p><ul><li><p><strong>Robin</strong>: preprint <a href="https://arxiv.org/abs/2505.13400">19 May 2025</a> &#8594; <em>Nature</em> July 2026. <strong>14 months.</strong></p></li><li><p><strong>Sakana's AI Scientist</strong>: arXiv <a href="https://arxiv.org/abs/2408.06292">12 August 2024</a> &#8594; <em>Nature</em> 25 March 2026. <strong>19 months.</strong></p></li><li><p><strong>Google's co-scientist</strong>: announced 19 February 2025 &#8594; <em>Nature</em> 19 May 2026. <strong>15 months, to the day.</strong></p></li><li><p><strong>Biomni</strong>: bioRxiv 30 May 2025 &#8594; <em>Science</em> 9 July 2026. <strong>13 months.</strong></p></li></ul><p>Mean lag: <strong>roughly fifteen months.</strong></p><p>The sharpest write-up of this belongs to synthetic biologist <a href="https://x.com/SynBio1/status/2075567928909467801">Jake Wintermute</a>, on Biomni:</p><blockquote><p><strong>"Biomni was on arXiv 13 months ago. Biomni was on GitHub 11 months ago. Phylo, the company built on Biomni, raised $13.5M and launched 5 months ago. Or, I guess, you could read about it in Science Magazine today."</strong></p></blockquote><p>None of this is a criticism of the labs. They all posted preprints immediately, which is exactly right. The lag belongs to the journals.</p><h3>Theme 3: where the field actually is</h3><p>Not where <em>Nature</em> says it is. While Robin was in review, FutureHouse built its successor and then built a company around it.</p><ul><li><p><strong>[Kosmos](https://arxiv.org/abs/2511.02824)</strong><a href="https://arxiv.org/abs/2511.02824">Kosmos</a>** (arXiv, November 2025): runs up to <strong>12 hours</strong> across ~20 cycles. A single run reads <strong>~1,500 papers</strong> and writes <strong>~42,000 lines of code</strong>. Independent scientists judged <strong>79.4%</strong> of the statements in its reports accurate. It has produced <strong>seven discoveries</strong>, three reproducing unpublished findings and four net-new. Collaborators estimated one run did about <strong>six months</strong> of their own work. Sam Rodriques' announcement did <strong>3,683 likes</strong>, the highest-engagement item in the entire sweep.</p></li><li><p><strong>Edison Scientific</strong>: FutureHouse's <strong>for-profit spinout</strong>, commercialising Kosmos. And <a href="https://x.com/SGRodriques/status/2071616647820177564">this month</a>, its first partnership using Kosmos not to write papers but to <strong>launch new biotech companies</strong>.</p></li><li><p><strong>[Sakana Marlin](https://x.com/hardmaru/status/2066529282588094713)</strong><a href="https://x.com/hardmaru/status/2066529282588094713">Sakana Marlin</a>**: Sakana's first commercial product. An autonomous research agent that runs ~<strong>8 hours</strong> unattended and emits structured slides plus a multi-dozen-page report. Pitched as a <strong>virtual chief strategy officer</strong>, aimed at finance and consulting rather than the bench.</p></li><li><p><strong>[Claude Science](https://www.anthropic.com/news/claude-science-ai-workbench)</strong><a href="https://www.anthropic.com/news/claude-science-ai-workbench">Claude Science</a>** (Anthropic, 30 June): the workbench end of the same idea. I wrote about it <a href="https://rundatarun.io/p/claude-science-and-the-boring-80">here</a>.</p></li></ul><p><strong>Read the journals to learn what was true a year ago. Read the preprints and the launches to learn what is true now.</strong></p><h3>Theme 4: the counterweight nobody is reading</h3><p>This is the highest-value material in the sweep and it appeared in almost no coverage.</p><p><strong>[Correct Answer, Wrong Mechanism](https://arxiv.org/pdf/2606.23175)</strong><a href="https://arxiv.org/pdf/2606.23175">Correct Answer, Wrong Mechanism</a>** (arXiv 2606.23175). Subtitle: <em>When AI Scientists Defend General Claims Their Own Data Contradicts.</em></p><p>Researchers watched a coding agent attempt to rediscover a known particle-physics result across <strong>28 episodes</strong>. It often got the right answer. The problem was <em>how</em>:</p><ul><li><p>CAWM occurred in <strong>4 of 20 (20%)</strong> primary-model episodes and <strong>3 of 8 (37.5%)</strong> episodes on other frontier models.</p></li><li><p>The agent reached right-looking results through reasoning that <strong>collapses when conditions change</strong>.</p></li><li><p>When pressed, it <strong>defended the wrong mechanism</strong>, arguing for physics inconsistent with the numbers in its own output.</p></li><li><p>Verdict: these systems are dependable as <strong>tools</strong> but <em>"unreliable scientific co-authors for open-ended claim-making."</em></p></li><li><p>The demand: <strong>outcome-only evaluation is insufficient.</strong> Score task outcome, mechanism fidelity, and epistemic honesty as three separate things. Their lightweight checks flagged <strong>every</strong> CAWM case in the study.</p></li></ul><p><strong>[Related work](https://phys.org/news/2026-05-ai-scientists-reveal-fundamental-limits.html)</strong><a href="https://phys.org/news/2026-05-ai-scientists-reveal-fundamental-limits.html">Related work</a>** (May 2026) catalogues the failure modes of autonomous research pipelines with unnerving specificity: implementation bugs, hallucinated results, shortcut reliance, methodology fabrication, citation hallucination, <strong>frame-lock</strong>, and <strong>bug-as-insight reframing</strong>, in which a system trips over a defect in its own code and writes it up as a discovery. The finding: increased <em>quantity</em>, decreased <em>quality</em>, of both papers and reviews.</p><p>And the view from the bench. Working scientist <a href="https://x.com/chorye/status/2075994339935723670">Emma Chory</a>, on Biomni cheerfully producing a <strong>BSL2-plus lentivirus protocol</strong> with, in her words, "sufficient detail to execute on a robot." Her review ran to four words: <em>"Cool cool cool cool."</em> An agent that will write you a biosafety-level-2-plus procedure on request is a governance question, not a capability win.</p><h3>Theme 5: the taxonomy worth stealing</h3><p>Computational biologist <a href="https://x.com/msikic/status/2075884565680259113">Mile Sikic</a> drew the distinction that most of the coverage blurs. There are <strong>two streams</strong>:</p><ol><li><p><strong>Virtual bioinformaticians</strong> that combine existing knowledge with existing tools. (Claude Science, Biomni, most of the field.)</p></li><li><p><strong>Systems that discover new biology in close collaboration with wet-lab scientists.</strong> (Robin, Google's co-scientist.)</p></li></ol><p>Conflating them is how people end up disappointed. They are solving different problems.</p><p>Also in the window, at lower weight: <strong>[EurekAgent](http://arxiv.org/abs/2606.13662)</strong><a href="http://arxiv.org/abs/2606.13662">EurekAgent</a>** ("Agent Environment Engineering is All You Need for Autonomous Scientific Discovery"); <a href="https://x.com/hxiao/status/2075750229748715784">Yoshua Bengio's take on an AI Scientist</a>; the Samsung SAIT AI Scientist competition winners; <strong>OpenScience</strong>, an open coding agent "that went to grad school," ~350 GitHub stars in its first week; and a genuinely useful sign of a crowded category, a piece titled <a href="https://x.com/BioAI_NeuralNet/status/2075940077885153423">"Which 'AI scientist' suits your lab? A guide for the perplexed."</a></p><div><hr></div><h2>What I'm watching next</h2><p><strong>Edison Scientific's biotech-founding partnership.</strong> Using an AI scientist to <em>found companies</em> is a categorically different claim from using one to write a paper, and unlike a paper it will be tested in public, on a clock, with other people's money.</p><p><strong>Whether anyone adopts the CAWM evaluation protocol.</strong> The paper's demand is concrete and cheap: score mechanism fidelity separately from outcome. If the next generation of agent papers still reports only "did it get the right answer," the field has decided not to look.</p><p><strong>Ripasudil.</strong> It cleared cells, not patients. A disease model comes next, and then, eventually, a randomized controlled trial. Most compounds that look this good at this stage never finish.</p><div><hr></div><h2>Sources</h2><p>Primary papers: <a href="https://www.nature.com/articles/s41586-026-10652-y">Robin / *Nature*</a> &#183; <a href="https://arxiv.org/abs/2505.13400">Robin preprint</a> &#183; <a href="https://arxiv.org/abs/2511.02824">Kosmos</a> &#183; <a href="https://arxiv.org/abs/2408.06292">Sakana AI Scientist</a> &#183; <a href="https://arxiv.org/pdf/2606.23175">Correct Answer, Wrong Mechanism</a> &#183; <a href="http://arxiv.org/abs/2606.13662">EurekAgent</a> &#183; <a href="https://news.stanford.edu/stories/2026/07/biomni-ai-powered-biomedical-co-scientist">Biomni / Stanford</a> &#183; <a href="https://research.google/blog/accelerating-scientific-breakthroughs-with-an-ai-co-scientist/">Google AI co-scientist</a> &#183; <a href="https://www.anthropic.com/news/claude-science-ai-workbench">Claude Science</a></p><p>Voices: <a href="https://x.com/SGRodriques/status/2075585708950221126">@SGRodriques</a> &#183; <a href="https://x.com/FutureHouseSF/status/2056813047180931316">@FutureHouseSF</a> &#183; <a href="https://x.com/SynBio1/status/2075567928909467801">@SynBio1</a> &#183; <a href="https://x.com/chorye/status/2075994339935723670">@chorye</a> &#183; <a href="https://x.com/msikic/status/2075884565680259113">@msikic</a> &#183; <a href="https://x.com/hardmaru/status/2066529282588094713">@hardmaru</a> &#183; <a href="https://x.com/aipoch_ai/status/2072591273983455480">@aipoch_ai</a></p><p><em>Methodology: thirty-day sweep across X, Hacker News, the curated AI wire, and the web, run 12 July 2026. The X leg of my usual tooling was down (an expired API credential), so X was swept through a working session-based client instead; Reddit returned no topical signal and was discarded rather than padded. Engagement figures are as recorded at collection time. The full deep dive on the Robin paper itself is [here](https://rundatarun.io/p/the-year-nature-caught-up).</em><a href="https://rundatarun.io/p/the-year-nature-caught-up">here</a>.*</p>]]></content:encoded></item><item><title><![CDATA[The Year Nature Caught Up]]></title><description><![CDATA[Robin found a new use for an old glaucoma drug and landed in Nature. The preprint was fourteen months old. Once you notice that lag, you notice it everywhere.]]></description><link>https://rundatarun.io/p/the-year-nature-caught-up</link><guid isPermaLink="false">https://rundatarun.io/p/the-year-nature-caught-up</guid><dc:creator><![CDATA[Justin Johnson]]></dc:creator><pubDate>Sun, 12 Jul 2026 12:05:04 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/8c1fcb31-7a67-482c-95aa-1015b4765a03_1584x672.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em>Every Sunday I pick one paper or release that's worth your time, break it apart, and tell you why it matters. No hype. No summaries of summaries. Just the idea, explained.</em></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!VO1V!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F821ce5de-6580-45f9-8280-270dd283c0b8_1584x672.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!VO1V!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F821ce5de-6580-45f9-8280-270dd283c0b8_1584x672.jpeg 424w, https://substackcdn.com/image/fetch/$s_!VO1V!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F821ce5de-6580-45f9-8280-270dd283c0b8_1584x672.jpeg 848w, https://substackcdn.com/image/fetch/$s_!VO1V!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F821ce5de-6580-45f9-8280-270dd283c0b8_1584x672.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!VO1V!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F821ce5de-6580-45f9-8280-270dd283c0b8_1584x672.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!VO1V!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F821ce5de-6580-45f9-8280-270dd283c0b8_1584x672.jpeg" width="728" height="409.5" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/821ce5de-6580-45f9-8280-270dd283c0b8_1584x672.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:false,&quot;imageSize&quot;:&quot;normal&quot;,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:728,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;captionedImage&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!VO1V!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F821ce5de-6580-45f9-8280-270dd283c0b8_1584x672.jpeg 424w, https://substackcdn.com/image/fetch/$s_!VO1V!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F821ce5de-6580-45f9-8280-270dd283c0b8_1584x672.jpeg 848w, https://substackcdn.com/image/fetch/$s_!VO1V!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F821ce5de-6580-45f9-8280-270dd283c0b8_1584x672.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!VO1V!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F821ce5de-6580-45f9-8280-270dd283c0b8_1584x672.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div><hr></div><h2>The Headline</h2><p>A multi-agent AI system named <strong>Robin</strong> was given the name of a disease and told to find a treatment. It read the literature, proposed a therapeutic strategy, named a drug, watched humans test that drug at the bench, analyzed the results itself, and then designed its own follow-up experiment. The disease was dry age-related macular degeneration, the major cause of blindness in the developed world. The drug was <strong>ripasudil</strong>, a rho-kinase inhibitor that has been sitting in pharmacies for years, approved for glaucoma, and never proposed for dAMD by anyone, as far as the authors could tell or I could find.</p><p>It worked in cells. Then Robin asked why, specified an RNA-sequencing experiment to find out, analyzed that too, and surfaced a gene called <em>ABCA1</em> as a possible new target nobody had been looking at.</p><p>This is the work of <a href="https://www.futurehouse.org/">FutureHouse</a>, and <a href="https://www.nature.com/articles/s41586-026-10652-y">it was published in Nature this week</a>. There is one sentence in the paper that deserves a second read:</p><blockquote><p><strong>"All hypotheses, experimental directions, data analyses and data figures in the main text of this report were produced by Robin."</strong></p></blockquote><p>Not assisted by. Produced by.</p><p><a href="https://arxiv.org/abs/2505.13400">The preprint went up in May 2025</a>. Nature published it in July 2026. Fourteen months, and almost none of the coverage mentions it.</p><p>Once you notice that gap, you start seeing it everywhere, and the shape of this entire field changes.</p><div><hr></div><h2>The Paper</h2><h3>Three birds and an orchestrator</h3><p>Robin is not a model. Nobody at FutureHouse trained a "discovery model," and if you go shopping for one after reading the coverage, you will not find it.</p><p>What they built is a system of specialized language agents sitting on top of ordinary commercial models:</p><ul><li><p><strong>Crow</strong> runs fast, concise literature searches.</p></li><li><p><strong>Falcon</strong> runs deep ones.</p></li><li><p><strong>Finch</strong> analyzes raw experimental data.</p></li></ul><p>Robin orchestrates the three of them and carries the hypothesis across rounds, so that what Finch learns on Tuesday reshapes what Falcon goes looking for on Wednesday.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Lu5o!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F84416561-2a93-4607-b01d-396da180ef3f_1408x768.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Lu5o!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F84416561-2a93-4607-b01d-396da180ef3f_1408x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!Lu5o!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F84416561-2a93-4607-b01d-396da180ef3f_1408x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!Lu5o!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F84416561-2a93-4607-b01d-396da180ef3f_1408x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!Lu5o!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F84416561-2a93-4607-b01d-396da180ef3f_1408x768.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Lu5o!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F84416561-2a93-4607-b01d-396da180ef3f_1408x768.jpeg" width="728" height="409.5" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/84416561-2a93-4607-b01d-396da180ef3f_1408x768.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:false,&quot;imageSize&quot;:&quot;normal&quot;,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:728,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;captionedImage&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!Lu5o!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F84416561-2a93-4607-b01d-396da180ef3f_1408x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!Lu5o!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F84416561-2a93-4607-b01d-396da180ef3f_1408x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!Lu5o!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F84416561-2a93-4607-b01d-396da180ef3f_1408x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!Lu5o!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F84416561-2a93-4607-b01d-396da180ef3f_1408x768.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><blockquote><p><strong>The intelligence came off a price sheet. The discovery came from the scaffolding around it.</strong></p></blockquote><p>I wrote two weeks ago that <a href="https://rundatarun.io/p/claude-science-and-the-boring-80">Anthropic shipped a workbench, not a miracle</a>, and that <a href="https://rundatarun.io/p/the-harness-is-the-moat">the harness is the moat</a>. Robin is that argument with a Nature paper attached to it, which is a considerably stronger position than my say-so.</p><h3>The word "semi" is doing honest work</h3><p>The paper describes Robin as <strong>semi-autonomous</strong>, and the precision is a credit to the authors rather than a hedge.</p><p>Humans still run the wet lab. People cultured the retinal cells, ran the assays, and pipetted the compounds. What Robin did was everything on either side of the bench: read the field, form the hypothesis, specify the experiment, interpret what came back, and revise.</p><p>The term of art is <strong>lab-in-the-loop</strong>, and it is worth flipping around, because the phrase misleads people. The AI is not in the lab. The lab is in the AI's loop.</p><p>Previous systems could do one arc of that circuit. They could read and propose. Or they could take your data and analyze it. Robin ran the full turn, then fed its own results into the next turn, and did it again.</p><p>The paper calls Robin "one of the first" systems to do this. The preprint called it "the first." Somewhere in fourteen months of review a superlative got sanded off, which tells you something useful about the claim and something more interesting about the process.</p><h3>Why the drug was findable at all</h3><p>FutureHouse call their method <strong>combinatorial synthesis</strong>, and their honesty about what it means is why I trust the rest of the paper.</p><p>Robin did not invent new biology. Every piece was already in print. That rho-kinase inhibition boosts phagocytosis in retinal pigment epithelium cells: known, and they cite the work. That this phagocytic housekeeping declines in AMD patients: also known. The two facts lived in different literatures, read by different people, and nobody had put them next to each other and said the word <em>ripasudil</em>.</p><p>The authors have a devastating example of how long that gap can persist. <strong>Dabrafenib</strong> is a cancer drug whose molecular action was characterized by 2010. Ten years later, a brute-force screen discovered it protects against hearing loss. That protective effect follows <em>directly</em> from the mechanism everyone already knew. The answer sat in print for a decade because the people who knew about BRAF inhibition and the people who cared about hearing loss were not the same people.</p><blockquote><p><strong>Robin is not doing science we could not do. It is doing science we did not get around to.</strong></p></blockquote><p>Scope the claim that way and it stays large. As the authors point out, novel FDA approvals have been flat at roughly fifty a year for a decade. If a machine can reliably close the distance between what is <em>known</em> and what is <em>connected</em>, that is worth a great deal of money and a great deal of eyesight.</p><h3>The two things I would copy tomorrow</h3><p><strong>They checked their own judge.</strong> Buried in the supplement is a comparison of their LLM evaluator against human experts. I read a lot of agent papers that grade themselves with a model and never once ask whether the grader is any good. This one asked.</p><p><strong>Their guardrails are the most thought-through of anything I read this month.</strong> Robin preferentially proposes compounds with established safety profiles. Every output is treated as a hypothesis entering standard preclinical review, never as a finding. And the discussion says the quiet part plainly: ripasudil "would of course require validation in a suitable disease model and ultimately in a randomized, placebo-controlled trial."</p><p>In vitro is a beginning. The authors know it, and they wrote it down.</p><div><hr></div><h2>The Ecosystem</h2><p>Robin made me want to go back and re-survey this whole category, so I did. Thirty days of papers, launches, funding, and argument. What I found reframed the paper I had just finished reading.</p><h3>Everyone landed at once</h3><p>Four flagship systems cleared peer review inside about four months.</p><p><strong>Google DeepMind's AI co-scientist</strong> <a href="https://research.google/blog/accelerating-scientific-breakthroughs-with-an-ai-co-scientist/">reached Nature in May</a>. It is a Gemini-based multi-agent system, and its headline validation is <em>the same move Robin made</em>: read the literature, propose an old drug for a new disease. Theirs was <strong>KIRA6</strong> for acute myeloid leukemia, which inhibited the viability of AML cells at clinically relevant concentrations. They also went after liver fibrosis targets and antimicrobial resistance.</p><p><strong>Sakana's AI Scientist</strong> <a href="https://sakana.ai/ai-scientist-nature/">reached Nature on 25 March</a>. It is more radical in ambition and less grounded in wet biology: it runs the entire pipeline through to a finished manuscript and then peer-reviews itself. A paper it generated <strong>passed the first round of human peer review</strong> at a top machine-learning workshop. Its most interesting result is a <em>scaling law of AI science</em>: the quality of the papers rises with the quality of the underlying model and with the compute you spend at inference.</p><p><strong>Biomni</strong>, out of Stanford, reached <strong>Science</strong> on 9 July, under the title "Autonomous biomedical research with an artificial intelligence agent." Hold on to one detail from it, because it is the most telling number in this entire piece: <strong>a prototype of Biomni was already running in more than 10,000 labs before the paper printed.</strong></p><h3>The lag is the actual story</h3><p>Every one of those systems was old news by the time it was printed.</p><p>Robin: preprint May 2025, Nature July 2026. Fourteen months. Sakana's AI Scientist went up on arXiv in August 2024 and reached Nature in March 2026. Nineteen months. Google announced its co-scientist in February 2025 and printed it in May 2026. Fifteen months, to the day. And Biomni got the sharpest write-up it will ever receive, from <a href="https://x.com/SynBio1/status/2075567928909467801">a synthetic biologist on X</a>:</p><blockquote><p><strong>"Biomni was on arXiv 13 months ago. Biomni was on GitHub 11 months ago. Phylo, the company built on Biomni, raised $13.5M and launched 5 months ago. Or, I guess, you could read about it in Science Magazine today."</strong></p></blockquote><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!VTcm!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30091134-6d50-4a1b-9285-f5fdde02a8e6_1408x768.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!VTcm!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30091134-6d50-4a1b-9285-f5fdde02a8e6_1408x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!VTcm!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30091134-6d50-4a1b-9285-f5fdde02a8e6_1408x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!VTcm!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30091134-6d50-4a1b-9285-f5fdde02a8e6_1408x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!VTcm!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30091134-6d50-4a1b-9285-f5fdde02a8e6_1408x768.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!VTcm!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30091134-6d50-4a1b-9285-f5fdde02a8e6_1408x768.jpeg" width="728" height="409.5" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/30091134-6d50-4a1b-9285-f5fdde02a8e6_1408x768.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:false,&quot;imageSize&quot;:&quot;normal&quot;,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:728,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;captionedImage&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!VTcm!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30091134-6d50-4a1b-9285-f5fdde02a8e6_1408x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!VTcm!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30091134-6d50-4a1b-9285-f5fdde02a8e6_1408x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!VTcm!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30091134-6d50-4a1b-9285-f5fdde02a8e6_1408x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!VTcm!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30091134-6d50-4a1b-9285-f5fdde02a8e6_1408x768.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>None of this is a criticism of the labs. FutureHouse posted their preprint the moment they had it, which is exactly right. The criticism, if there is one, belongs to the clock the journals run on.</p><p>So where is the field actually standing? Not where Nature says it is.</p><p>FutureHouse has already built Robin's successor. <strong>Kosmos</strong> went up on arXiv last November. It runs for twelve hours at a stretch, and a single run reads around 1,500 papers and writes roughly 42,000 lines of code. Independent scientists checked its reports and found <strong>79.4% of its statements accurate</strong>. It has produced seven discoveries, four of them net new. Collaborators estimated that one run did about six months of their own work.</p><p>They then spun out a for-profit company, <strong>Edison Scientific</strong>, to sell it. And this month they announced a partnership that uses Kosmos not to write papers but to <strong>found biotech companies</strong>.</p><p>Sakana, meanwhile, is selling <strong>Marlin</strong>, an autonomous research agent that runs for about eight hours unattended and is pitched as a virtual chief strategy officer.</p><blockquote><p><strong>Read Nature to learn what was true a year ago. Read the preprints and the launches to learn what is true now.</strong></p></blockquote><h3>The counterweight nobody is reading</h3><p>While these systems were collecting their journal stamps, a second body of work was landing that tells you precisely where they break. None of the launch threads mention it.</p><p>The sharpest of it is a paper whose title belongs on a poster in every lab now buying one of these things: <a href="https://arxiv.org/pdf/2606.23175">**Correct Answer, Wrong Mechanism**</a>. The subtitle is better. <em>When AI Scientists Defend General Claims Their Own Data Contradicts.</em></p><p>The researchers watched a coding agent try to rediscover a known result in particle physics, twenty-eight times over. Often it got the right answer. The problem was <em>how</em>. In <strong>20% of episodes</strong> with the primary model, and <strong>37.5% across other frontier models</strong>, the agent reached a right-looking result through reasoning that collapses the moment conditions change. Worse, when pressed, it <em>defended</em> the wrong mechanism, arguing for physics that contradicted the numbers in its own output.</p><p>Their verdict is careful, and damning. These systems are dependable as <strong>tools</strong> and, for now, "unreliable scientific co-authors for open-ended claim-making." Asking whether the agent got the answer right tells you almost nothing. You have to score the outcome, the fidelity of the mechanism, and the honesty of the account as three separate things.</p><p><a href="https://phys.org/news/2026-05-ai-scientists-reveal-fundamental-limits.html">Related work in the same window</a> catalogues the failure modes with unnerving specificity: hallucinated results, methodology fabrication, citation invention, <strong>frame-lock</strong>, and my favorite piece of vocabulary this year, <strong>bug-as-insight reframing</strong>, in which a system trips over a defect in its own code and writes it up as a discovery.</p><p>Set that next to Robin and its design choices stop looking conservative and start looking wise. When Robin says ripasudil enhances phagocytosis, a cell culture has already voted, and a cell culture does not care what the agent believes about it.</p><blockquote><p><strong>A wet lab is a brutal, incorruptible critic. You cannot argue a cell culture into agreeing with you.</strong></p></blockquote><p>But be precise about how far that protection reaches, because it is not the whole paper. When Robin says the mechanism runs through <em>ABCA1</em>, that is Finch reading an RNA-sequencing experiment, and that is exactly the surface the CAWM authors are worried about. <strong>The bench protects the finding. It does not protect the explanation.</strong> Which is roughly the division of trust I would apply to any of these systems, and to be fair to FutureHouse, it is the division they applied themselves: the drug candidate is offered for preclinical review, the mechanism is offered as a possibility.</p><p>It is also why the wet-biology systems stop at <strong>semi</strong>-autonomous, and why that ceiling is physical rather than technical. Laboratory instruments were not built to take orders from software. Somebody has to run the assay. Sakana's system is the exception that proves the rule, and it proves it uncomfortably: with no bench anywhere in the loop, it is the one system in this survey that will write its own paper <em>and</em> review it.</p><div><hr></div><h2>What I Run</h2><p>I should declare an interest, briefly, because I have spent the last year on the other side of this.</p><p>I run an autonomous research engine called <strong>ARIA</strong>. It is a persistent system rather than a one-shot agent: a pool of research ideas competing against each other on a single scoring function, experiments that run locally first and have to earn their way onto expensive hardware, a critic drawn from a different model family that posts a verdict on every result, and failures that classify themselves and re-enter the pool as recovery work. Seven instances, four domains, about 19,400 auditable commits, and 97.8% autonomous resumption after failure. Its successor now runs around the clock on <a href="https://rundatarun.io/p/borrowed-iron">a borrowed eight-GPU node</a>, pointed at retinal disease. There is <a href="https://www.justinhjohnson.com/case/aria">a case study</a> and <a href="https://youtu.be/UJyBEVFBRPs">a short film</a> if you want the shape of it.</p><p>I raise it for one reason. Building the thing taught me the same lesson the CAWM paper is now pressing on the whole field, and I learned it the expensive way.</p><blockquote><p><strong>Generating hypotheses was never the hard part. Building a critic honest enough to kill them was.</strong></p></blockquote><p>An idea generator is cheap, and it will happily run forever, and every dashboard will stay green while it does. What is difficult, and what almost nobody budgets for, is the machinery that tells the system it is wrong and makes the verdict stick.</p><p>Robin never had to build that machinery, because the bench already is it. That is a structural advantage of working in wet biology, and it is worth naming plainly for anyone whose research loop closes entirely in software, where nothing pushes back for free.</p><div><hr></div><h2>Why It Matters</h2><p><strong>Do not go shopping for a discovery model.</strong> There isn't one. Not one system in this survey used a model you cannot buy. Robin, the co-scientist, Kosmos, Biomni: every one of them is a harness built around a commodity brain, which means the thing your organization would be building is the harness, and that is engineering, not procurement.</p><p><strong>The compartmentalization gap is inside your own company.</strong> Robin's whole trick is finding the connection between two literatures that no single person reads. You have that problem internally, at smaller scale, right now. The result one team generated and another team needed and never saw is the cheapest discovery you will ever fund.</p><p><strong>Read the preprints.</strong> On this topic the peer-reviewed record is running roughly a year behind the work. Use the journals to confirm, not to track.</p><p><strong>Take the reliability research as seriously as the launch threads.</strong> A system that finds the right answer for the wrong reason will sail through your demo and drown in your clinic. Ask any vendor how they score mechanism fidelity, not just outcomes. Watch them field the question.</p><p>And watch ripasudil. It cleared cells, not patients. The road from here runs through a disease model and then a randomized controlled trial, and most compounds that look this good at this stage do not finish it.</p><p>Robin did not replace the scientist. It replaced the <strong>lag</strong>, the decade dabrafenib spent sitting in the literature while nobody put two known facts side by side.</p><p>There is a certain irony in that finding taking fourteen months to reach print. It does not make it less of a landmark. Congratulations to the FutureHouse team. The field is better for this being in the record, even if the record took its time.</p><div><hr></div><p><em>I ran a full thirty-day sweep of this category to write this piece, and most of it did not fit. The complete survey, with every system, number, and citation, is [in a companion post here](https://rundatarun.io/p/last-30-days-the-ai-scientists).</em><a href="https://rundatarun.io/p/last-30-days-the-ai-scientists">in a companion post here</a>.*</p><div><hr></div><p><em>Sunday Deep Dive is a weekly series on Run Data Run. Every Sunday I pick one paper, release, or technique worth understanding, break it apart, and tell you what it means for your work. Free every Sunday, no paywall. If it was useful, the easiest way to support it is to subscribe and forward it to one person on your team who'd want it. If it wasn't, tell me why. I'll make it better.</em></p>]]></content:encoded></item><item><title><![CDATA[How I Actually Stay Current on AI (I Built a Wire)]]></title><description><![CDATA[The question I get asked most, and the system that answers it.]]></description><link>https://rundatarun.io/p/how-i-actually-stay-current-on-ai</link><guid isPermaLink="false">https://rundatarun.io/p/how-i-actually-stay-current-on-ai</guid><dc:creator><![CDATA[Justin Johnson]]></dc:creator><pubDate>Thu, 09 Jul 2026 13:50:59 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/fbf81d6f-0b6b-46a5-bf61-75ed8fb681b4_1664x928.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!gsD8!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe58370ef-185a-4cff-9ce6-305c25b8d617_1664x928.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!gsD8!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe58370ef-185a-4cff-9ce6-305c25b8d617_1664x928.jpeg 424w, https://substackcdn.com/image/fetch/$s_!gsD8!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe58370ef-185a-4cff-9ce6-305c25b8d617_1664x928.jpeg 848w, https://substackcdn.com/image/fetch/$s_!gsD8!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe58370ef-185a-4cff-9ce6-305c25b8d617_1664x928.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!gsD8!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe58370ef-185a-4cff-9ce6-305c25b8d617_1664x928.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!gsD8!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe58370ef-185a-4cff-9ce6-305c25b8d617_1664x928.jpeg" width="728" height="409.5" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e58370ef-185a-4cff-9ce6-305c25b8d617_1664x928.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:false,&quot;imageSize&quot;:&quot;normal&quot;,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:728,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;captionedImage&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!gsD8!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe58370ef-185a-4cff-9ce6-305c25b8d617_1664x928.jpeg 424w, https://substackcdn.com/image/fetch/$s_!gsD8!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe58370ef-185a-4cff-9ce6-305c25b8d617_1664x928.jpeg 848w, https://substackcdn.com/image/fetch/$s_!gsD8!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe58370ef-185a-4cff-9ce6-305c25b8d617_1664x928.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!gsD8!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe58370ef-185a-4cff-9ce6-305c25b8d617_1664x928.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Some version of this question arrives most weeks. How do you keep up? You have a day job. The field moves every single day. What are you reading?</p><p>For a long time I answered it badly. I would say something about reading a lot, which is true and completely useless to the person asking. The real answer took me a while to say out loud, because it sounds like a dodge.</p><blockquote><p><strong>I stopped trying to keep up. I built something that keeps up for me.</strong></p></blockquote><div><hr></div><h2>Bookmarks are where links go to die</h2><p><strong>Nobody is reading this field.</strong> Not all of it, not close. Papers land overnight. A lab ships a model on a Tuesday and by Thursday other teams have published what breaks about it. The people who look like they are on top of it are not reading faster than you. <strong>They have narrowed, or they have automated, or they are bluffing.</strong></p><p>Manual curation was my first answer, and it failed the way manual curation always fails. Open tabs became a graveyard. Saved links went unread and, worse, went stale.</p><blockquote><p><strong>A link I saved in March, about a model superseded in April, is not information anymore. It is clutter with a timestamp.</strong></p></blockquote><p>Newsletters helped until I was subscribed to a stack of them and reading none. <strong>The problem was never that I lacked sources.</strong> I had too many, arriving on their own schedules, with no shared memory between them.</p><div><hr></div><h2>One pipe</h2><p>So I built a wire.</p><p>Several channels I already had running now point at one place. A scan that reads the day's news, releases, and research. The things I bookmark as I go. Findings from a small set of AI assistants I run, each pointed at its own beat. My own written breakdowns of the papers and tools and companies I sit down and work through by hand. Things I email myself at midnight and would otherwise never see again.</p><p>All of it lands in one store, <strong>scored and deduplicated</strong>. No single source has to be the right source, because none of them carries the day alone. And nothing gets read twice, because the store already knows it saw that story on three feeds this morning.</p><p>Two things come out the other end.</p><p><strong>The first is private</strong>, and it is the part I use every day without thinking about it. The store is a knowledge base that refreshes itself, and my own AI tools read from it directly. When I sit down to write, or research something, or work a problem, the context they pull from is current. Not their training data. Not a stale snapshot I remembered to update. This morning's.</p><p><strong>The second is public.</strong> A curated slice of the store becomes a website: a daily dispatch of the items that scored highest, written up rather than merely linked, so you know why something mattered before you decide to click it.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!qc54!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3b2132d0-a0a2-4814-aa2c-1f7e69cccc93_1376x768.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!qc54!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3b2132d0-a0a2-4814-aa2c-1f7e69cccc93_1376x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!qc54!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3b2132d0-a0a2-4814-aa2c-1f7e69cccc93_1376x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!qc54!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3b2132d0-a0a2-4814-aa2c-1f7e69cccc93_1376x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!qc54!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3b2132d0-a0a2-4814-aa2c-1f7e69cccc93_1376x768.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!qc54!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3b2132d0-a0a2-4814-aa2c-1f7e69cccc93_1376x768.jpeg" width="728" height="409.5" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/3b2132d0-a0a2-4814-aa2c-1f7e69cccc93_1376x768.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:false,&quot;imageSize&quot;:&quot;normal&quot;,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:728,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;captionedImage&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!qc54!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3b2132d0-a0a2-4814-aa2c-1f7e69cccc93_1376x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!qc54!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3b2132d0-a0a2-4814-aa2c-1f7e69cccc93_1376x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!qc54!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3b2132d0-a0a2-4814-aa2c-1f7e69cccc93_1376x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!qc54!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3b2132d0-a0a2-4814-aa2c-1f7e69cccc93_1376x768.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>I like the arithmetic of it. The wire watches <strong>fifty-two sources</strong>. Over the last month it read and scored <strong>two thousand six hundred and seventy-nine items</strong>, and three hundred and ninety-nine of them cleared the bar. Yesterday's dispatch came out of a hundred and twenty-one signals. The site puts it more plainly than I would: it reads about seven things so you read one.</p><blockquote><p><strong>Fifty-two sources in. One page out.</strong></p></blockquote><p>What lands is a page I can get through with coffee.</p><div><hr></div><h2>The wire is the front door</h2><p>This is not a side project I bolted on. <strong>It sits at the front of everything else I make.</strong></p><p>What scores highest is usually what I end up writing about, because the things worth six hundred words announce themselves by refusing to go away. What I write becomes what I post. What I post brings back arguments and corrections from people who know more than I do about some corner of it, and those go back into the store as research. <strong>The loop closes.</strong> One pipe, several outlets, and the website is only the visible tip of it.</p><p>I did not design it that way at the start.</p><blockquote><p><strong>I built the scan because I was drowning. Then I noticed the thing that solved the drowning was also the thing deciding what I had to say.</strong></p></blockquote><div><hr></div><h2>Go look at it</h2><p><strong>The daily dispatch is free, it is public, and it lands every morning:</strong> <a href="https://wire.rundatarun.io/briefs">wire.rundatarun.io/briefs</a>. The day's signals, grouped by what they are, with a source link on every item so you can go argue with the original instead of taking my read on it.</p><p>Underneath it, the <a href="https://wire.rundatarun.io/firehose">firehose</a> shows you the machinery. The whole intake, the filter that thins it, and the last seven days in the raw.</p><p>What it does not show you is the store underneath. The complete searchable archive, the picks I would put in front of you myself with the reasoning attached, the decode and dossier library, the week ahead before it is news. <strong>That layer is not free, and it is not open yet.</strong> There is a waitlist at the bottom of the firehose page, and if the arithmetic up there sounded like your problem too, put your email in it.</p><div><hr></div><p>Here is the whole thing, start to finish, in about forty-five seconds.</p><div id="youtube2-SKk4hA0-9Mk" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;SKk4hA0-9Mk&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/SKk4hA0-9Mk?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div>]]></content:encoded></item><item><title><![CDATA[AI Ate the Keyboard. Now It Has to Eat the Queue.]]></title><description><![CDATA[Software stopped being the limiting factor. Process didn't. In biopharma and any patient-facing industry, the queue is the patient.]]></description><link>https://rundatarun.io/p/ai-ate-the-keyboard-now-it-has-to</link><guid isPermaLink="false">https://rundatarun.io/p/ai-ate-the-keyboard-now-it-has-to</guid><dc:creator><![CDATA[Justin Johnson]]></dc:creator><pubDate>Tue, 07 Jul 2026 15:30:18 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/f77759d2-d27a-445e-873e-1e81e7aa3925_1376x768.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!bipI!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2a271f29-27df-406f-9f07-f724abd1ecc0_1376x768.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!bipI!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2a271f29-27df-406f-9f07-f724abd1ecc0_1376x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!bipI!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2a271f29-27df-406f-9f07-f724abd1ecc0_1376x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!bipI!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2a271f29-27df-406f-9f07-f724abd1ecc0_1376x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!bipI!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2a271f29-27df-406f-9f07-f724abd1ecc0_1376x768.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!bipI!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2a271f29-27df-406f-9f07-f724abd1ecc0_1376x768.jpeg" width="728" height="409.5" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/2a271f29-27df-406f-9f07-f724abd1ecc0_1376x768.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:false,&quot;imageSize&quot;:&quot;normal&quot;,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:728,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;captionedImage&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!bipI!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2a271f29-27df-406f-9f07-f724abd1ecc0_1376x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!bipI!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2a271f29-27df-406f-9f07-f724abd1ecc0_1376x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!bipI!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2a271f29-27df-406f-9f07-f724abd1ecc0_1376x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!bipI!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2a271f29-27df-406f-9f07-f724abd1ecc0_1376x768.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Jack Hanlon, who leads GenAI Media at Meta, <a href="https://www.linkedin.com/posts/jhanlon_the-first-time-you-see-an-engineer-build-activity-7439409934175195139-XY6w">posted this on LinkedIn</a>:</p><blockquote><p>"The first time you see an engineer build something in 45 minutes that would have taken a week a year ago, but then see it not ship for another 6 weeks, you will be radicalized."</p></blockquote><p>He's describing Meta, a consumer tech company. The review stack he lists is the load every build carries before it ships: Strategy, Product, Design, Engineering, Privacy, Legal, Accessibility, Comms. It's also not close to what biopharma and patient-facing R&amp;D organizations ask of an AI build.</p><p>In biopharma and adjacent patient-facing industries, add data privacy impact assessment, security review, AI governance review, IT architecture review, change-management approval, procurement, legal redline, regulatory sign-off. If patient data is anywhere in the lifecycle, add clinical safety review. Each gate is serial. Each gate is three weeks minimum. Six in sequence is half a year before anything ships. You've done paperwork.</p><p>The numbers back this up across every regulated sector, not just pharma. Enterprise SaaS procurement runs a median 170 days; complex solutions hit eleven-and-a-half months. SOC 2 Type II audits take six to twelve months end-to-end. Federal contractors face 12-18 month FedRAMP authorization cycles before a system can sit on a government network. Banks run AI governance committees on top of model-risk committees on top of vendor-risk committees, each stack inherited from a different decade of incident response. MIT NANDA's "GenAI Divide" report found 95% of enterprise AI pilots deliver zero P&amp;L impact, across thirty to forty billion dollars of spend. That's not a build problem. The build works. The build sits in a queue.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!-j8K!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6a54bdd2-dd7c-4e51-963d-6486633dbdd9_1200x896.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!-j8K!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6a54bdd2-dd7c-4e51-963d-6486633dbdd9_1200x896.jpeg 424w, https://substackcdn.com/image/fetch/$s_!-j8K!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6a54bdd2-dd7c-4e51-963d-6486633dbdd9_1200x896.jpeg 848w, https://substackcdn.com/image/fetch/$s_!-j8K!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6a54bdd2-dd7c-4e51-963d-6486633dbdd9_1200x896.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!-j8K!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6a54bdd2-dd7c-4e51-963d-6486633dbdd9_1200x896.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!-j8K!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6a54bdd2-dd7c-4e51-963d-6486633dbdd9_1200x896.jpeg" width="728" height="409.5" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/6a54bdd2-dd7c-4e51-963d-6486633dbdd9_1200x896.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:false,&quot;imageSize&quot;:&quot;normal&quot;,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:728,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;captionedImage&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!-j8K!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6a54bdd2-dd7c-4e51-963d-6486633dbdd9_1200x896.jpeg 424w, https://substackcdn.com/image/fetch/$s_!-j8K!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6a54bdd2-dd7c-4e51-963d-6486633dbdd9_1200x896.jpeg 848w, https://substackcdn.com/image/fetch/$s_!-j8K!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6a54bdd2-dd7c-4e51-963d-6486633dbdd9_1200x896.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!-j8K!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6a54bdd2-dd7c-4e51-963d-6486633dbdd9_1200x896.jpeg 1456w" sizes="100vw"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>This accelerated review pathway can exist. It can work. It will take herculean levels of structural redesign and a lot of organizational political capital to land. And the pattern isn't a pharma problem. It's the shape of every regulated industry trying to ship AI in 2026. Pharma's version is the sharpest because patients are downstream of the delay, but the argument is universal.</p><p>Every gate exists for a reason. Most got written after a real incident. Someone's data got exposed. A launch communication went sideways. A model made a biased call in production. The answer isn't to kill them. The answer is that every one of them was designed for a world where building the software was the slow, expensive, irreversible part. Whether you wrote the code yourself or bought it from a vendor who wrote it, software was the thing the calendar bent around. That world is gone. AI didn't just speed up typing. It turned the whole build, your own code and the vendor's both, into the fast part. The review cycle was calibrated for the old critical path. The critical path moved. The calendar didn't.</p><p>Hanlon's three confrontations are right. Process architecture is the speed limit. Too many orgs treat every process as sacred. AI should be eating process, not just coding. The patient-facing version is sharpest, because the delay shows up downstream as patient time.</p><div><hr></div><h2>This Has Happened Before</h2><p>DevOps didn't land in 2015 as a philosophy. It landed when development got fast enough that deployment became the jam. "Development teams could complete features in days, but getting those features deployed took weeks." That's the DevOps.com retrospective on the pre-CI/CD era. Swap "deployed" for "procurement-approved" and you have a 2026 AI program, sentence for sentence.</p><p>Then security became the jam. DevSecOps. Then infrastructure and developer experience became the jam. Platform Engineering. Internal Developer Platforms. Each compression moved the bottleneck one layer further from the keyboard. Each shift spawned a category, a toolchain, a set of jobs, a conference circuit.</p><p>Bottleneck migration is a law. Compress one layer and the adjacent layer becomes the bottleneck by definition. Whatever used to be the tall pole is now almost all of the remaining pole.</p><p>AI compressed the build layer, writing software and buying it both. Everything adjacent to the build is now the bottleneck. The reviewers, the committees, the approvals, the procurement track, the vendor security questionnaire, the model evals nobody scoped time for. All of it. DevOps took twelve years to run the full cycle. Nobody has twelve years.</p><div><hr></div><h2>Shadow AI Is the Tell</h2><p>Gartner's 2025 numbers are the signal. 98% of organizations report unsanctioned AI use. 69% report evidence of prohibited public GenAI use inside their walls. 49% expect a shadow AI incident within twelve months.</p><p>The default read calls shadow AI a governance failure. A discipline problem. Employees not following the rules. That read is backwards. Shadow AI is the system telling you the official process layer already failed an internal cost-benefit test. Employees measured the queue, measured their quarterly objectives, and decided the queue costs more than the risk of getting caught.</p><p>That's the market voting. When 98% of your workforce is routing around your approval process, you don't have a discipline problem. You have an economics problem. The process priced itself out.</p><p>A fair counter to all of this: maybe the 95% pilot failure isn't a process problem. Maybe it's a pilot-selection problem. Most enterprise AI pilots fail because they were the wrong pilot to run, picked for board optics or executive curiosity rather than real workflow pain. Speeding up bad pilots faster doesn't help anyone.</p><p>That argument is partially right. The data is harder than that. The same MIT report that flagged the 95% number also found the highest ROI was in back-office automation, exactly the work that sits behind procurement, vendor security, and compliance queues. The pilots that fail include some bad picks. They also include good picks that died in queue. You can have a pilot-selection problem and a process problem at the same time, and most enterprises do.</p><div><hr></div><h2>Process Protects Patients. Except When It Doesn't.</h2><p>Biopharma's review gauntlet was written in blood. Good Clinical Practice, Good Laboratory Practice, Good Manufacturing Practice. Patients got hurt. Rules got written. Every review, every sign-off, every three-week queue descends from that inheritance.</p><p>The reflex response to "AI is slow to ship" is "good, that's how we keep patients safe." It sounds unanswerable. It isn't.</p><blockquote><p><strong>Process was built to protect patients. When process is the thing blocking patient-helping AI from shipping, it stops protecting patients. It protects itself.</strong></p></blockquote><p>Not shipping AI is not a neutral, safe choice. It has a cost. The cost lands on patients too.</p><p>Every week of review on a regulatory drafting agent is a week a safety narrative gets written the slow way. Novo Nordisk's NovoScribe rollout cut clinical study report drafting from fifteen weeks with over fifty medical writers to ten minutes with three. The work moved. The submission queue didn't. Every week of governance queue on a tool like that is real human time that wasn't spent on the submission, real submission time that wasn't spent reaching patients.</p><p>Every month a vendor security review stalls on a literature retrieval agent is a month oncology scientists read abstracts by hand. The responder-subgroup signal sits in a PDF nobody opened.</p><p>Every quarter a leadership group debates a two-speed governance design is a quarter trial protocols get mapped by hand, deviations get triaged out of a Teams channel, the first patient at the first site doses later than they would have.</p><p>Meanwhile, the deployments ship. Cleveland Clinic, NHS England, the VA, and 75% of US hospitals are running AI in clinical workflows today. The teams that wired governance early are in production. The teams still debating are watching.</p><p>Most governance agendas debate the wrong question. Not "is this AI safe enough to ship," but "is the delay we're about to impose smaller than the harm we're trying to prevent." Both sides of that ledger have a body count. Only one gets counted.</p><div><hr></div><h2>AI Eats Its Own Governance Layer</h2><p>The governance layer isn't a coding problem, and that's the part most "AI accelerator" pitches miss. It's legal redline, procurement of the services the agents will consume, vendor security, the privacy review, the model evals, the benchmark sign-off. None of that is typing. All of it is the build now, and all of it queues.</p><p>Organizations have answered by standing up governance as elaborate as the thing it governs. Sanofi has publicly documented a multi-stage setup: a Responsible AI Working Committee, an Interim Responsible AI Governance body, an AI agent named Plai sitting in on every drug-progression decision. Other large pharmas, banks, and federal contractors run variations of the same shape; the public documentation lags Sanofi's. Ethan Mollick names why it decides everything. "The moderating factor is no longer individual ability or even AI capability. It is organizational structure, policy, and the way leaders choose to approach AI." The constraint isn't the model. It's the layer wrapped around the model.</p><p>Which is exactly the layer AI can eat. Hanlon's third point is the one most readers stop quoting before the punchline:</p><blockquote><p>"Have the AI traverse your codebase and collect your evidence for the privacy review and have another AI grade the privacy review. They go back and forth until (a) they need human intervention because they are stuck or (b) they need humans to look at the final results and sign off on the right path forward."</p></blockquote><p>That's the whole answer, and it generalizes past the codebase. Every gate in the gauntlet has the same shape: forty hours of evidence assembly, then a fifteen-minute judgment call. The privacy review is fourteen pages of data flow that already lives in the repo, the retention policies, the IAM config. The vendor questionnaire is two hundred SOC 2 controls a lead engineer maps by hand against an attestation report nobody re-read. The procurement intake is free-text someone retypes against master agreement terms a sourcing lead has memorized. The model eval is a benchmark suite someone runs once and pastes into a slide. The regulatory submission is fifteen weeks of medical writers stitching prior submissions into a new narrative. The judgment is fifteen minutes. The forty hours is the queue.</p><p>The human doesn't leave the privacy review. The human finally opens a pre-assembled package instead of staring at a blank template with half the source material missing. The reviewer still reviews. The reviewer is finally reviewing the only part of the work that ever needed a reviewer.</p><blockquote><p><strong>AI has eaten the keyboard. It hasn't eaten the queue.</strong></p></blockquote><p>That's the argument in one sentence. The keyboard compressed two years ago. The queue hasn't. Until it does, every "AI accelerator" announcement is a faster typewriter feeding the same backlog.</p><div><hr></div><h2>The Reviewer Is the Project</h2><p>There's a catch built into all of this. The people whose work the AI must eat are the same people who must approve the AI eating their work. Legal reviews the legal agent. Procurement approves the procurement agent. InfoSec signs off on the InfoSec agent.</p><blockquote><p><strong>The reviewer is the subject of the review. The incentives don't line up. Pretending they do is how an org burns eighteen months and ships nothing.</strong></p></blockquote><p>The first week of agent output will be embarrassingly wrong, and that isn't the agent's fault. The process has never been written down in a way a machine can act on. The DPIA "template" is really a set of in-head judgments held by a senior person who onboarded in 2014 and knows which questions matter. Turning that into prompts, policies, and machine-readable rubrics is the institutional work, and there's no shortcut around it. That IS the project. The catch is that the people best positioned to codify their judgment are the ones whose standing rests on it staying in their heads. The org design that works makes the codified version the asset and leaves the committee to ratify it. Most enterprises don't run that design yet.</p><p>The ones that do compress the queue start in the same place, and it's never a model. They map the gauntlet, then fund an agent for a single gate, picked where the work is repetitive and the inputs already live somewhere a machine can read. Privacy review first, its answers mostly sitting in the repo and the IAM config. Vendor security next, the same shape with messier inputs. Regulatory drafting last, after the pattern has proven itself on something lower-stakes than an FDA filing. It holds at one gate, run for two months and watched for what it gets wrong. It breaks at three at once.</p><p>Underneath all of it sits the substrate. A governance layer that can't hand an agent a rubric to run or a template to fill isn't policy. It's tradition. The version living in senior reviewers' heads is neither auditable nor transferable, and it's the version adding weeks to every queue. Codifying it is the platform bet of 2026, ahead of the next model, because it's what lets every future build ship without the six-month tax.</p><div><hr></div><p>AI compressed the build layer. That was the easy part. Everyone got the faster typewriter. The next decade of enterprise AI won't be won on better models. It turns on whether the layer built to protect patients from bad software, back when bad software was the threat, can be made to compress itself before the delay starts landing on the other side of the ledger.</p><p>That layer will not shrink on its own. It has to be handed its own rubric and told to grade the work it used to guard by hand. No team that owns a queue volunteers to automate it. Someone with the authority to override that instinct has to decide the delay costs more than the control.</p><blockquote><p><strong>In a patient-facing industry, the queue is not paperwork. The queue is the patient. Every week it stays uncompressed is a week charged to someone who never got a seat in the room.</strong></p></blockquote>]]></content:encoded></item><item><title><![CDATA[Every AI Book Tells You What to Think. This One Is About What Your Hands Do.]]></title><description><![CDATA[There is a whole shelf of AI books for leaders now. I read enough of them to know why I wrote a different kind.]]></description><link>https://rundatarun.io/p/every-ai-book-tells-you-what-to-think</link><guid isPermaLink="false">https://rundatarun.io/p/every-ai-book-tells-you-what-to-think</guid><dc:creator><![CDATA[Justin Johnson]]></dc:creator><pubDate>Mon, 06 Jul 2026 12:48:33 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/308a6dbe-e77a-45c7-b3b8-7ecb2ea518b2_1408x768.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!0dbC!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4bf7a834-d4b3-4ed9-987d-dd4472ecf6f6_1408x768.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!0dbC!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4bf7a834-d4b3-4ed9-987d-dd4472ecf6f6_1408x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!0dbC!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4bf7a834-d4b3-4ed9-987d-dd4472ecf6f6_1408x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!0dbC!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4bf7a834-d4b3-4ed9-987d-dd4472ecf6f6_1408x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!0dbC!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4bf7a834-d4b3-4ed9-987d-dd4472ecf6f6_1408x768.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!0dbC!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4bf7a834-d4b3-4ed9-987d-dd4472ecf6f6_1408x768.jpeg" width="728" height="409.5" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/4bf7a834-d4b3-4ed9-987d-dd4472ecf6f6_1408x768.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:false,&quot;imageSize&quot;:&quot;normal&quot;,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:728,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;captionedImage&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!0dbC!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4bf7a834-d4b3-4ed9-987d-dd4472ecf6f6_1408x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!0dbC!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4bf7a834-d4b3-4ed9-987d-dd4472ecf6f6_1408x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!0dbC!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4bf7a834-d4b3-4ed9-987d-dd4472ecf6f6_1408x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!0dbC!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4bf7a834-d4b3-4ed9-987d-dd4472ecf6f6_1408x768.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>There is a shelf now. If you are a senior leader trying to make sense of AI, you have probably bought most of it. The strategy book. The mindset book. The one with the four-quadrant framework. The one from the consultancy with eleven co-authors. They are not bad books. Several are good. I learned things from them.</p><p>But I noticed something after the fourth or fifth one. I would finish a chapter, nod, underline a sentence, and then sit there with my hands in exactly the same place they were before I started reading. The book had changed what I thought about AI. It had changed nothing about what I did with it on Monday.</p><p>That is the gap I wrote into.</p><blockquote><p><strong>The book changed what I thought about AI. It changed nothing about what I did with it on Monday.</strong></p></blockquote><div><hr></div><h2>The shelf has one shape</h2><p>Read enough of these books and the shape repeats. They tell you what to think about AI. How to frame it for your boss/peers/board. Which mental model to adopt. Where it sits in your strategy. They are written from the altitude a leader is used to operating at, which is to say: above the work, directing it.</p><p>That altitude is the problem, not the solution. Someone reviewing the current best-seller in this category said it reads like a 2024 book explaining how the Internet will help your business. The line is unkind and basically right. Not because the author is wrong, but because the genre has a ceiling. You cannot write the operator's manual from the strategy altitude. The two are different books, and almost everyone is writing the first one.</p><p>I wanted to read the second one. Nobody had written it for the reader I had in mind, so I wrote it.</p><blockquote><p><strong>You cannot write the operator's manual from the strategy altitude.</strong></p></blockquote><div><hr></div><h2>What "what your hands do" actually means</h2><p>"Leaders should build" is already a slogan, and slogans rot.</p><p>I don't mean you should learn to code. I mean something narrower. There is a class of tools now, Claude Code chief among them, that lets you direct a multi-agent system to produce real work inside your own domain. You describe intent, set boundaries, review what comes back, redirect, and ship. The interface happens to be a terminal; the skill is not programming. It is the same skill that got you to senior: defining the why, designing the handoffs, judging the output, catching the thing that looks right and is wrong.</p><p>You already do all of that. You do it with people. The book is about pointing those same instincts at a harness instead of a team, and what changes when you do.</p><p>The reason this matters is not productivity, though the productivity is real. The reason is calibration. Andrej Karpathy named it this spring in a line that landed with twenty thousand likes and a thousand replies: there are two groups talking past each other about AI, and the line between them is not skeptic versus believer. It is the people who have built something with these systems and the people who have read about them. Every confident claim you have heard about what AI can or cannot do was made from one side of that line. If you are reasoning about AI from the reading side, you are reasoning from the wrong data, and no amount of strategy reading moves you across. You move across by getting your hands on the controls one time and feeling where the real edge is.</p><div><hr></div><h2>Why I get to say this</h2><p>What separates this book from the shelf is that I did not write it from the strategy altitude. I built and led an AI Center of Excellence at a Fortune 500 pharmaceutical company, and I now lead applied data and AI in R&amp;D at an AI-driven biotech. I also run a deep personal practice on nights and weekends, eight autonomous agents and a homelab full of skills, because the day job and the obsession turned out to be the same skill. This book was itself produced through the harness it describes, end to end. About a month from kickoff to typeset manuscript, against the year the same book would have taken by hand, with thousands of hours of research, red-teaming, and editing behind it, most of it run in parallel by the system rather than typed by me. You can see the skills that drafted it.</p><p>I am not telling you to do something I read about. I am telling you what it was actually like, including the parts that are awkward and the parts that broke.</p><blockquote><p><strong>I am not telling you to do something I read about. I am telling you what it was actually like.</strong></p></blockquote><div><hr></div><h2>The part I won't oversell</h2><p>At Sequoia's AI Ascent this spring, Karpathy declared "vibe coding" obsolete and named its successor "agentic engineering," a discipline with real rigor and real responsibility. He is right. He also drew a careful line: prototyping with AI raises the floor for everyone, but operating serious systems demands discipline you cannot wave away.</p><p>My book asks you to step over a line slightly further out than the one Karpathy draws for the general public. I am asking a senior, non-technical executive to personally operate, not just prototype. I think that is where the role is heading, and I think the leaders who learn it first will shape what their organizations become. But I am not going to pretend it is the consensus position. It is an argument. The book makes it, and gives you the controls and the guardrails to test it for yourself rather than take my word for it.</p><div><hr></div><h2>Who will not find this useful</h2><p>If you want a framework to put on a slide for your board, the shelf has better options than mine, and I mean that. If you want reassurance that your current AI strategy is sound, this is the wrong book; it will probably make you uncomfortable. If you are looking for predictions about AGI timelines, I route around that debate on purpose.</p><p>This book is useful if you have read the think pieces, sponsored the center of excellence, watched the pilots, and still feel the quiet gap between knowing AI matters and knowing what you personally do about it. It is useful if you are willing to spend a few hours with your hands on something unfamiliar, on your own files, to find out where the edge actually is. It is a ladder. Three months to ship something real, six on the outside. The last page is a welcome from the people who were already on the other side when you started.</p><p>If that is you, the book is for you. It's out now: <a href="https://builder-leader.com">builder-leader.com</a>, or <a href="https://www.amazon.com/Builder-Leader-Exoskeleton-That-Crosses-Gap-ebook/dp/B0H3LQQ4J6/">buy it on Amazon</a>.</p><p><em>Builder-Leader: The AI Exoskeleton That Crosses the Gap.</em></p>]]></content:encoded></item><item><title><![CDATA[The Gap, and the Book That Came Out of It]]></title><description><![CDATA[There's a line splitting the AI conversation that nobody's arguing about, because most people don't know it's there. I've spent months on the wrong-versus-right side of it. The book is out now.]]></description><link>https://rundatarun.io/p/the-gap-and-the-book-that-came-out</link><guid isPermaLink="false">https://rundatarun.io/p/the-gap-and-the-book-that-came-out</guid><dc:creator><![CDATA[Justin Johnson]]></dc:creator><pubDate>Wed, 01 Jul 2026 12:30:46 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/b524b50c-f55b-45d9-865c-3cbddfd12eac_1376x768.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!HUGE!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6c862e89-7118-45c9-9a1a-ded97a130121_1376x768.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!HUGE!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6c862e89-7118-45c9-9a1a-ded97a130121_1376x768.png 424w, https://substackcdn.com/image/fetch/$s_!HUGE!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6c862e89-7118-45c9-9a1a-ded97a130121_1376x768.png 848w, https://substackcdn.com/image/fetch/$s_!HUGE!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6c862e89-7118-45c9-9a1a-ded97a130121_1376x768.png 1272w, https://substackcdn.com/image/fetch/$s_!HUGE!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6c862e89-7118-45c9-9a1a-ded97a130121_1376x768.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!HUGE!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6c862e89-7118-45c9-9a1a-ded97a130121_1376x768.png" width="1376" height="768" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/6c862e89-7118-45c9-9a1a-ded97a130121_1376x768.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:768,&quot;width&quot;:1376,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1362402,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://rundatarun.io/i/204431537?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6c862e89-7118-45c9-9a1a-ded97a130121_1376x768.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!HUGE!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6c862e89-7118-45c9-9a1a-ded97a130121_1376x768.png 424w, https://substackcdn.com/image/fetch/$s_!HUGE!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6c862e89-7118-45c9-9a1a-ded97a130121_1376x768.png 848w, https://substackcdn.com/image/fetch/$s_!HUGE!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6c862e89-7118-45c9-9a1a-ded97a130121_1376x768.png 1272w, https://substackcdn.com/image/fetch/$s_!HUGE!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6c862e89-7118-45c9-9a1a-ded97a130121_1376x768.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>I've been telling you a book was coming for months. I named the idea in <a href="https://rundatarun.io/p/two-gaps-not-one">"Two Gaps, Not One"</a> back in April and dropped not-so-subtle hints on LinkedIn. As of today, it's not a mention anymore. It has a name, a cover, and you can <a href="https://www.amazon.com/Builder-Leader-Exoskeleton-That-Crosses-Gap-ebook/dp/B0H3LQQ4J6/">buy it right now: *Builder-Leader: The AI Exoskeleton That Crosses the Gap*</a>. Paperback is live this minute. Kindle lands tomorrow, and you can get it early. Everything else, the excerpt, the cover, what's actually inside, lives at <a href="https://builder-leader.com">builder-leader.com</a>.</p><h2>What looks like a fight is a gap</h2><p>The AI conversation reads like a war between camps.</p><p>Skeptics on one side: it's overhyped autocomplete, the bubble's coming, calm down. True believers on the other: everything changes next year, the jobs are going, brace. A large, tired middle trying to figure out which loud group to trust.</p><blockquote><p><strong>It is not a war between camps. It's a gap between two populations, and most of the noise is an artifact of which side of it you're standing on.</strong></p></blockquote><p><a href="https://x.com/karpathy/status/2042334451611693415">Andrej Karpathy named this on April 9 of this year</a>, in a short post that got the whole feed nodding. There are two groups talking past each other about AI, he said. Not skeptics versus believers. People who have actually built something with these systems on one side, people who have read about them or tried the free version once on the other. The first group has watched the staggering version of the technology work for them. The second group is reasoning from screenshots and vibes.</p><p>The reply thread found a third group hiding inside the first. One reply, from @doodlestein, named the people "magnifying the power of these frontier models and agent harnesses with custom tooling and skills and workflows." Builders who don't just use the systems but wrap machinery around them. That's the group whose read on what AI can do is calibrated correctly, because they're the only ones holding the data.</p><p>Once you see this, the boardroom argument stops being a debate about technology. <strong>It becomes a map of who in the room has ever operated the thing.</strong> The person who hasn't isn't lazy or incurious. They're just, structurally, the worst-placed person to make the call, because they're working from the wrong inputs.</p><h2>Why the gap is asymmetric</h2><p>Here's the part that turns this from an observation into a problem: it only runs one direction.</p><p>The skeptic and the builder are not two sides of a balanced argument. They are the same function run on different inputs. If you have never seen a system go off, do real multi-step work, get something wrong, take your correction, and come back with it fixed, your mental model of AI is calibrated to the wrong tier. You're picturing a chatbot. You're rating the chatbot. You're right about the chatbot.</p><p>The builder has seen the other thing, and can't unsee it.</p><blockquote><p><strong>There's no symmetric version of this. One side is holding the data and the other one isn't. That's not a difference of opinion. It's a difference of evidence.</strong></p></blockquote><p>You can watch this play out at scale right now. McKinsey's 2025 State of AI survey found that the large majority of organizations report using AI, while fewer than four in ten report capturing meaningful value from it. <a href="https://www.gartner.com/en/newsroom/press-releases/2025-06-25-gartner-predicts-over-40-percent-of-agentic-ai-projects-will-be-canceled-by-end-of-2027">Gartner has said</a> it expects over 40% of agentic-AI projects to be cancelled by the end of 2027. None of those numbers are a story about AI being overhyped. They're a story about leaders trying to deploy something nobody senior had ever personally operated, and being surprised when it stalled.</p><p>The gap shows up as the failure rate. It just doesn't announce itself that way.</p><h2>The crossing is personal, not institutional</h2><p>Now the uncomfortable part, the reason this is a book and not a memo.</p><p>You cannot read your way across this gap. You cannot buy your way across it with a platform contract. You cannot delegate it to a team and call it handled.</p><blockquote><p><strong>The crossing is personal. Your organization cannot do it for you, no matter how much you spend or how good your people are.</strong></p></blockquote><p>I've watched a lot of capable leaders try the institutional version. Sponsor the center of excellence, approve the big platform deal, stand up a steering committee, wait for the capability to arrive on schedule. Most of them land in the stalled half of those statistics, and most of them can't say why. The work was funded. The vendor was credible. The team was strong.</p><p>It stalled because nobody senior in the room had ever driven the thing they were deploying. They were directing work they couldn't evaluate.</p><p>The leaders on the other side of the gap did something simpler and harder. They sat down themselves. They picked a real problem, pointed a system at it, watched it produce work, threw out what was wrong, kept what was right, and did it again. Over months, not over a weekend. Until they could tell good from bad from the inside, fast, without asking anyone.</p><p>That's all crossing the gap means. Not becoming an engineer. Not learning to type code at a professional level. <strong>Learning to direct a system that does the building, and to recognize the moment it's confidently wrong.</strong></p><p>The book has a name for the structure you build doing this. The <strong>harness</strong>: the persistent setup around the model, the skills and memory and tools and workflows that turn a chat window into something you operate. The model is a commodity everyone can rent. The harness is the part that's yours, and it's where the actual advantage lives. The book argues that the harness is the moat, that its parts are knowable and nameable, and that the operator role at the center of it will outlast the next model, and the one after that.</p><p>It also argues that the crossing has a shape you can follow. Not a leap. A path with rungs. You start with something small that works, you build the second thing on top of the first, and the setup you carry from each project makes the next one faster. The reason ninety-day pilots fail and six-month builds hold is that the first is a sprint at the thing and the second is a ladder up it. The back half of the book is that ladder, written out rung by rung, for a reader who has never written a line of code and has no plans to start.</p><p>The book itself is three parts. Part One names the gap and why your organization cannot close it for you, no matter what you approve or fund. Part Two is the exoskeleton itself: the harness, and the person operating it. Part Three is the crossing, a three-month ladder that starts with you working alone and ends with you leading a small team across the same gap. That's the shape of it, if you want to see it for yourself: <a href="https://www.amazon.com/Builder-Leader-Exoskeleton-That-Crosses-Gap-ebook/dp/B0H3LQQ4J6/">get it on Amazon</a>.</p><h2>You don't have to stop being who you are</h2><p>The fear under all of this, the one I hear from senior people more than any other, is that crossing the gap means becoming someone else. That the polished VP has to turn into a junior engineer to stay relevant. That the judgment and the relationships and the executive instinct that got them here suddenly count for nothing.</p><p>The book argues the exact opposite, and I believe it because I lived it.</p><p>Your polish, your read on people, your political capital, your taste in what's worth doing: those don't become liabilities when you start building. <strong>They become multipliers.</strong> The rare and valuable profile isn't the engineer who learned to present. It's the leader who learned to operate. That combination is rare today not because it's unnatural, but because almost nobody has told leaders it was available to them.</p><blockquote><p><strong>You are not being asked to give up what you're good at. You're being asked to point it at a new kind of system.</strong></p></blockquote><h2>I wrote it with the thing it's about</h2><p>One more thing, because it's the proof and not a gimmick.</p><p>I built this book through the same kind of system it describes. It has its own research process that sweeps the current discourse before each chapter. Its own drafting setup that holds the voice and the argument in place. Its own fact-checking pass that pulls every claim into a queue and makes me source it or cut it. A small machine I direct, not a document I type alone.</p><p><strong>If a book about crossing the gap had taken three years to write by hand, it would be arguing against itself.</strong> This one didn't. The way it got made is the clearest existence proof I can offer for the thing it's selling: a person operating an exoskeleton produces work at a scale and speed they couldn't produce bare-handed, and the work is still theirs. The book is what that produced.</p><h2>Who it's for</h2><p>I wrote it for one reader in particular.</p><p>The leader who has started to notice that the people around them can now do things with these tools that they can't. Who feels the distance growing month over month. Who hasn't said it out loud, because saying it out loud feels like an admission. If that lands a little too close to home, good. That's exactly who it's for, and the distance you're feeling is more closeable than it looks from where you're standing.</p><p>It's not for the true believers waiting on AGI, and it's not for the skeptics holding the line that the whole thing is smoke. They've picked their stories. This is for the people in the middle who still have a decision to make and want to make it from the inside, not from a slide deck someone built to sell them something.</p><h2>Out today</h2><p>The book is out. The paperback is live right now. The Kindle edition ships tomorrow, and you can get it today: <a href="https://www.amazon.com/Builder-Leader-Exoskeleton-That-Crosses-Gap-ebook/dp/B0H3LQQ4J6/">Builder-Leader on Amazon</a>.</p><p>If the gap I've described here is recognizable, if you've felt the space between what your team can suddenly do and what you can, you're closer to crossing it than the argument going on around you suggests. The first step is smaller than it sounds. The book is the rest of them.</p><p>A few weeks ago I <a href="https://rundatarun.io/p/two-gaps-not-one">made the longer case for why this is the gap that decides the next five years</a>, against a piece that I think pointed at the wrong one. This is the follow-through, the one I get to write with the book actually out.</p><p>Thanks for reading along to this point. Go cross the gap.</p><p>Justin</p>]]></content:encoded></item><item><title><![CDATA[Claude Science and the boring 80 percent]]></title><description><![CDATA[Anthropic shipped a workbench, not a miracle. The model is the commodity layer. The harness around it is the moat.]]></description><link>https://rundatarun.io/p/claude-science-and-the-boring-80</link><guid isPermaLink="false">https://rundatarun.io/p/claude-science-and-the-boring-80</guid><dc:creator><![CDATA[Justin Johnson]]></dc:creator><pubDate>Tue, 30 Jun 2026 19:41:55 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/f5034efe-fb34-4530-90ad-3fa4155af889_1376x768.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!7K-9!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3caf94c5-45ed-44a6-83f8-2da0cddb14e3_1376x768.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!7K-9!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3caf94c5-45ed-44a6-83f8-2da0cddb14e3_1376x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!7K-9!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3caf94c5-45ed-44a6-83f8-2da0cddb14e3_1376x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!7K-9!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3caf94c5-45ed-44a6-83f8-2da0cddb14e3_1376x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!7K-9!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3caf94c5-45ed-44a6-83f8-2da0cddb14e3_1376x768.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!7K-9!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3caf94c5-45ed-44a6-83f8-2da0cddb14e3_1376x768.jpeg" width="728" height="409.5" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/3caf94c5-45ed-44a6-83f8-2da0cddb14e3_1376x768.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:false,&quot;imageSize&quot;:&quot;normal&quot;,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:728,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;captionedImage&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!7K-9!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3caf94c5-45ed-44a6-83f8-2da0cddb14e3_1376x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!7K-9!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3caf94c5-45ed-44a6-83f8-2da0cddb14e3_1376x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!7K-9!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3caf94c5-45ed-44a6-83f8-2da0cddb14e3_1376x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!7K-9!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3caf94c5-45ed-44a6-83f8-2da0cddb14e3_1376x768.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Anthropic shipped <a href="https://claude.com/science">Claude Science</a> today, an AI workbench for scientists. The launch page leads with proteins and genomics and a reviewer agent that checks your citations. That is the surface. The interesting part is <strong>what they chose to build it on, and what they chose not to promise.</strong></p><p>A workbench is a boring word on purpose. It is the thing the interesting work happens on top of. Anthropic picked that word over "discovery engine" or "research copilot," and the choice tells you where they think the advantage sits.</p><div><hr></div><h2>What actually shipped</h2><p>Claude Science is a local-first app that runs on your laptop, your Linux box, or your HPC login node. A generalist agent coordinates sixty-plus skills and connectors pre-wired for genomics, single-cell, proteomics, structural biology, and cheminformatics. It renders 3D protein structures and genome-browser tracks and chemical structures inline, next to the code that made them.</p><p>Three design decisions are worth pausing on, because they are not how most "AI for science" products are built.</p><p><strong>Every figure ships with its own recipe.</strong> When Claude Science generates a plot, it bundles the exact code and environment that produced it, a plain-language description, and the full message history. Six months later you can reopen it and see what was actually run. This is the reproducibility crisis spoken to directly, <strong>not as a slogan but as a file format.</strong> Every artifact is auditable by construction.</p><blockquote><p><strong>The output is not the figure. The output is the figure plus the complete record of how it was made.</strong></p></blockquote><p><strong>Compute is a thing the agent manages, not a thing you schedule.</strong> Folding a protein or running a genomics pipeline across a huge dataset usually means stopping your work to write a job script, submit it to the cluster, wait, check whether it failed, pull the results back. Claude Science drafts the plan, asks before it spends new resources, and lets you revoke any decision before it writes and submits the job to the infrastructure you already use. Your HPC cluster over SSH. Your <a href="https://modal.com/blog/modal-integration-brings-scalable-compute-to-claude-science">Modal</a> account for compute on demand. One GPU to hundreds.</p><p>The detail that does the work: datasets load once into a running session and stay in memory. <strong>Large or sensitive data never has to leave the systems it already lives on.</strong> Only the slice of context a given step needs goes to the model.</p><p><strong>The tooling layer is not theirs.</strong> This is the move most people will miss. Claude Science does not ship its own protein-folding model. It talks to <a href="https://nvidianews.nvidia.com/news/nvidia-launches-bionemo-agent-toolkit-giving-ai-agents-the-tools-to-accelerate-scientific-discovery">NVIDIA's BioNeMo</a> natively, which means Evo 2, Boltz-2, OpenFold3 are one call away. Anthropic orchestrates. The scientific models belong to the tooling layer.</p><div><hr></div><h2>The layer cake</h2><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!kiOj!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3c3622ee-aabd-43a6-960d-d356afe7b935_1376x768.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!kiOj!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3c3622ee-aabd-43a6-960d-d356afe7b935_1376x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!kiOj!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3c3622ee-aabd-43a6-960d-d356afe7b935_1376x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!kiOj!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3c3622ee-aabd-43a6-960d-d356afe7b935_1376x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!kiOj!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3c3622ee-aabd-43a6-960d-d356afe7b935_1376x768.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!kiOj!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3c3622ee-aabd-43a6-960d-d356afe7b935_1376x768.jpeg" width="728" height="409.5" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/3c3622ee-aabd-43a6-960d-d356afe7b935_1376x768.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:false,&quot;imageSize&quot;:&quot;normal&quot;,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:728,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;captionedImage&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!kiOj!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3c3622ee-aabd-43a6-960d-d356afe7b935_1376x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!kiOj!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3c3622ee-aabd-43a6-960d-d356afe7b935_1376x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!kiOj!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3c3622ee-aabd-43a6-960d-d356afe7b935_1376x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!kiOj!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3c3622ee-aabd-43a6-960d-d356afe7b935_1376x768.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The shape of the market is already clear. AI-for-science is stratifying into three layers, and it rhymes with the stack we already know.</p><p>The application layer is the orchestration surface, the thing the scientist talks to. Anthropic and OpenAI's GPT-Rosalind compete here. The tooling layer is the scientific models and databases, BioNeMo and Boltz and AlphaFold. The infrastructure layer is the silicon. Each layer is a real business. The tooling and silicon layers are where the heavy capital and the specialized models live, and the companies there, NVIDIA especially, keep doing the hard work that everyone in the application layer stands on.</p><p>Anthropic's bet is narrower and smarter than "best model." <strong>The model is becoming the commodity layer. The harness around it is where the moat gets built.</strong></p><p>The whole launch is a platform play, not a model play. John Jumper, the Nobel laureate behind AlphaFold, announced he was leaving Google DeepMind for Anthropic eleven days before this launch. The Coefficient Bio acquisition bought target-discovery talent. The partnerships with the Allen Institute, the Broad, Schr&#246;dinger, Sanofi, and HHMI are integration deals. The moat being built is workflow integration, not raw model capability.</p><div><hr></div><h2>We have been here</h2><p>I have been arguing this for a while. <strong>The model is the commodity layer. The harness, the system of skills and memory and agents and hooks and connectors that sits around the model and makes it useful, is the moat.</strong> I wrote a whole <a href="https://rundatarun.io/p/the-harness-is-the-moat">Run Data Run essay on it</a>, and a book chapter. The argument was built on watching what happened when you swapped the model out from under a working agent system and the system barely blinked, because the intelligence lived in the scaffolding.</p><p>Claude Science is that argument, built into a product, for science.</p><p>The clearest version I have seen up close is a research flywheel we run in my group. It is an autonomous agent that runs continuously, generating and refining research ideas at the intersection of AI and biotechnology, promoting the promising ones to execution on a GPU box, critiquing its own results, and self-healing when a run corrupts. Over an eighteen-week deployment across seven instances it ran <strong>5,196 consecutive commits without a human touching it, with a 97.8 percent autonomous resumption rate.</strong> The expensive, interesting part is not any one model call. It is the loop: the weighted idea pool, the adversarial critique, the self-heal, the provenance on every artifact. Swap the underlying model and the loop keeps turning. The loop is the asset.</p><p>That is the same shape Claude Science is reaching for, at platform scale. The sixty pre-wired skills. The reviewer agent that checks citations and calculations as the pipeline runs. The session that holds context in memory and forks to compare two approaches. Every one of those is a harness component. None of them is a model.</p><blockquote><p><strong>The model is the part you can rent. The harness is the part you build. Claude Science just shipped sixty reasons to build it for science.</strong></p></blockquote><div><hr></div><h2>Two reviews worth reading</h2><p>Two of the beta stories in the launch tell you what this is actually good at, because neither is a discovery claim.</p><p>J&#233;r&#244;me Lecoq, a neuroscientist at the Allen Institute, used Claude Science to build a roughly twenty-skill pipeline that writes long-form computational reviews. Sub-agents read thousands of papers, pull the central claim and the key number, and store them in an evidence database. Then a pipeline builds a narrative arc, delegates each section to a specialist, and dedicated agents generate cross-study figures straight from the evidence. The core trick is actor-critic pairs: one agent writes, a separate reviewer agent checks it for accuracy and citation fidelity.</p><p>Before Claude Science, that kind of review took his team up to two years. He now has about ten of them, many over a hundred pages, with citations the reviewer agents vetted.</p><blockquote><p><strong>Two years collapsed to weeks, on the part of science that is reading and synthesizing and checking, not on the part that is having the idea.</strong></p></blockquote><p>Stephen Francis, an epidemiologist at the UCSF Brain Tumor Center, used it for germline workups on glioma susceptibility. His lab had done this work before. Claude Science compressed the analysis to roughly a tenth of the time, and his group independently validated the results.</p><p>Neither man claims a discovery. Both claim compression of the tedious majority of the work. <strong>That distinction is the whole product.</strong></p><div><hr></div><h2>What it does not do</h2><p>Here is what Claude Science is not, and what the launch page does not claim it is.</p><p>It is not a drug-discovery engine. Zero AI-discovered drugs have FDA approval. The ninety-percent clinical failure rate has not moved. The biomolecular models in the tooling layer underneath it work well for structure prediction and initial screening, and they are getting better fast. They are not yet where you want them for the decision that gates a program. A careful recent evaluation (<a href="https://arxiv.org/abs/2603.05532">Wan and Coveney, UCL</a>) of one of those models on thirty-eight thousand compounds found strong performance on structure but weaker correlation with physics-based binding free energies at the candidate end, where a medicinal chemist actually bets. That is not a knock on the model. <strong>It is where the whole field sits in 2026, and the field knows it.</strong></p><p>The pattern is one the coding world already learned. The benchmark looks good. The transfer to the real decision is the harder, slower climb. The structure-and-screening work compresses. The lead-identification work does not, yet.</p><blockquote><p><strong>AI compresses the eighty percent of science that was already tractable. The twenty percent that kills programs stays expensive.</strong></p></blockquote><p>The framing Brendan Frey, the Deep Genomics CEO, gave Drug Target Review is the one to hold onto: "AI has really let us all down in the last decade when it comes to drug discovery. We've just seen failure after failure." Anthropic's head of life-sciences partnerships, Jonah Cool, has been consistent that the target is the tedious intermediate work, data analysis and annotation and coordination. The launch page repeats that framing almost verbatim. <strong>A company sitting on a $965 billion valuation, under pressure to grow, chose to launch a workbench and not a miracle.</strong></p><div><hr></div><h2>Why a builder should care</h2><p>If you lead a data or AI team in the sciences, three things are now true and worth acting on.</p><p>The first is that the integration layer is where the money and the moat are going. The specialized models, the folding and docking and single-cell tools, are commoditizing fast as the tooling layer matures. Your durable advantage is not which model you run. It is how well you connect the models to your proprietary data, your lab's instruments, your validated pipelines. Claude Science lets you save any pipeline as a reusable skill and inherit it in future sessions. <strong>That is the shape of the advantage. Build the connective tissue, not the model.</strong></p><p>The second is that the reproducibility-by-construction pattern is going to spread. Every figure carrying its full provenance, the code and the environment and the message history, is the answer to a problem every R&amp;D leader has been screaming about for a decade. If you are building internal tooling, steal the pattern wholesale. The artifact is not the output. <strong>The artifact is the output plus the audit trail.</strong></p><p>The third is the scope limit. Claude Science will accelerate the boring eighty percent of your scientists' work, the literature synthesis and the pipeline glue and the figure iteration, and it will do it well. It will not pick your drug target. It will not replace your physics-based validation on the decisions that gate a program. <strong>Plan for the compression. Do not budget for the cure.</strong></p><div><hr></div><h2>What I'd watch</h2><p>The beta is open today for Pro, Max, Team, and Enterprise on macOS and Linux. There is a discounted Team plan for academic labs and nonprofits, and Anthropic is funding up to fifty AI-for-Science projects with up to $30,000 in credits (Modal adds $2,000 of compute). Applications close July 15.</p><p>The thing to watch is not the launch. Launches are easy. Watch whether the reviewer agent actually catches the citation and calculation errors six months in, when the novelty has worn off and the session histories are long. Watch whether scientists trust it enough to save their hard-won pipelines as skills and let the next session inherit them. That trust, not the protein renderer, is what earns the workbench a permanent place on the bench.</p><blockquote><p><strong>Launches are easy. Trust earned over long session histories is the hard part.</strong></p></blockquote><p>Anthropic shipped a workbench. They named it for what it is. The science was always eighty percent grinding and twenty percent insight, and the insight was never the part that took the time.</p><div><hr></div><h2>The papers</h2><ul><li><p><a href="https://www.anthropic.com/news/claude-science-ai-workbench">Claude Science, an AI workbench for scientists, is now available</a>. Anthropic, Jun 30 2026. The launch.</p></li><li><p><a href="https://nvidianews.nvidia.com/news/nvidia-launches-bionemo-agent-toolkit-giving-ai-agents-the-tools-to-accelerate-scientific-discovery">NVIDIA launches BioNeMo Agent Toolkit</a>. The tooling layer Claude Science orchestrates over.</p></li><li><p><a href="https://modal.com/blog/modal-integration-brings-scalable-compute-to-claude-science">Modal integration brings scalable compute to Claude Science</a>. The on-demand compute leg.</p></li><li><p><a href="https://arxiv.org/abs/2603.05532">On the Reliability of AI Methods in Drug Discovery: Evaluation of Boltz-2</a>. Wan and Coveney, UCL. The reality check on binding affinity prediction.</p></li><li><p><a href="https://claude.com/science">Get started with Claude Science</a>. The product.</p></li></ul><p><em>This is a Run Data Run essay. If it was useful, the easiest way to support it is to subscribe and forward it to one person on your team who'd want it. If it wasn't, tell me why.</em></p>]]></content:encoded></item><item><title><![CDATA[Borrowed Iron]]></title><description><![CDATA[A borrowed eight-GPU node, a global-health mission, and a crew of six: me, a founder, and four interns.]]></description><link>https://rundatarun.io/p/borrowed-iron</link><guid isPermaLink="false">https://rundatarun.io/p/borrowed-iron</guid><dc:creator><![CDATA[Justin Johnson]]></dc:creator><pubDate>Sun, 28 Jun 2026 10:20:43 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/3e688750-7cc7-4215-b080-ffa4e27c150a_1376x768.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!26Uh!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcf93d8d7-6a26-43b5-a53c-297d01b124e8_1376x768.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!26Uh!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcf93d8d7-6a26-43b5-a53c-297d01b124e8_1376x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!26Uh!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcf93d8d7-6a26-43b5-a53c-297d01b124e8_1376x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!26Uh!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcf93d8d7-6a26-43b5-a53c-297d01b124e8_1376x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!26Uh!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcf93d8d7-6a26-43b5-a53c-297d01b124e8_1376x768.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!26Uh!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcf93d8d7-6a26-43b5-a53c-297d01b124e8_1376x768.jpeg" width="728" height="409.5" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/cf93d8d7-6a26-43b5-a53c-297d01b124e8_1376x768.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:false,&quot;imageSize&quot;:&quot;normal&quot;,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:728,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;captionedImage&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!26Uh!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcf93d8d7-6a26-43b5-a53c-297d01b124e8_1376x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!26Uh!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcf93d8d7-6a26-43b5-a53c-297d01b124e8_1376x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!26Uh!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcf93d8d7-6a26-43b5-a53c-297d01b124e8_1376x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!26Uh!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcf93d8d7-6a26-43b5-a53c-297d01b124e8_1376x768.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Earlier this week, NVIDIA gave a global-health startup I work with a node of eight H100 GPUs to use for a couple of months. The grant came through <a href="https://www.nvidia.com/en-us/startups/">NVIDIA Inception</a>, their program for startups. No invoice. They looked at the mission, believed in it, and handed over the iron.</p><p>The day the box came online, I posted the first entry of a build log about it on our technical blog. In the few days since, the pace has been fast enough that the story already needs an update. So here it is properly, from the top, for the people who lead the work rather than write the code.</p><p>Here is the part I love most, before anything else. The crew putting that machine to work is six people. Me. Nick, who runs the company. And four interns, the last of them starting tomorrow. Three are in college; the fourth just finished high school and starts undergrad soon.</p><blockquote><p><strong>A node of eight H100s, a foundation model to train, and a two-month clock. The crew is two adults and four interns.</strong></p></blockquote><div><hr></div><h2>The mission</h2><p><a href="https://www.socialeyesus.com/">SocialEyes</a> is bigger than the AI, and that is what pulled me in. The mission is to converge a wide range of AI healthcare skills onto the patient, right at the point of care. Today's workflow is screen, then refer: a frontline worker spots something, and the patient travels to a specialist who may be hundreds of miles and many months away. That model dates to the 1890s, and it simply does not work in low- and middle-income countries. There are roughly 200,000 eye specialists on Earth and well over a billion people with diabetes, high blood pressure, and the slow diseases that take sight and life. The math does not work. SocialEyes is trying to fix the math.</p><p>The way in is the eye. It is the cheapest window we have into the rest of the body, the one place you can photograph blood vessels and nerve tissue directly, with no needle, using a camera that fits in a clinic. We lead with timely screening of the chronic diseases that increasingly dominate global health: diabetes, hypertension, and the conditions that take sight before anyone catches them. From the same image we can also pull biomarkers and risk factors, and in some cases a specialist-level read of certain retinal conditions. The retina is just one of the tissues the SocialEyes architecture can handle. By itself, signs of dozens of diseases, in the eye and systemically, can ultimately be detected.</p><p>That is what makes it worth two borrowed months. The work is grounded where it is needed most: low-resource clinics across low- and middle-income countries, anchored in Nepal. The first deployments are at clinics and other fixed sites, not yet in the hands of community health workers in the field. That comes later, with a dedicated device called MARVIN. For now: a camera, a trained model, and a clinic a long way from the nearest specialist. That is the product. The retinal AI is how it gets there.</p><p>The training data is grounded in public research datasets, which is a deliberate choice. It keeps the method something we can talk about openly, even when specific results stay in the lab. Which is convenient, because that is exactly what I am doing.</p><div><hr></div><h2>How I ended up in the room</h2><p>SocialEyes is Nick's mission. I build the AI for it, and have been part of the team since the start of the year.</p><p>They found me through the open work. The experiments on a DGX Spark on my desk, the writing on the blog, the <a href="https://rundatarun.io/p/im-justin-johnson-i-build-things">habit of building the real thing and showing it</a>. Nick reached out, I started helping, I dug the mission, and now we are doing some of the coolest work I have done.</p><p>That origin is the whole way I think about building. Build the real thing, share it in public, and the right people show up. A working system and an honest write-up change the conversation from "should we" to "how do we do this faster." This collaboration is what that looks like when it happens.</p><p>The four interns are the same story from the other side. Students who wanted to do real work, not a slideshow about AI. So they are doing real work. They have stood up their first shared code repository, they watch the GPU dashboards, they own a lane of the operations. Three are in college and one is about to start. Most people their age get coffee. These four are helping run a supercomputer.</p><div><hr></div><h2>The borrowed iron</h2><p>Here is the through-line for anyone who has followed the work. The small box came first.</p><p><a href="https://rundatarun.io/p/three-days-one-petaflop-and-an-ai">A DGX Spark</a>, NVIDIA's desktop AI machine, sits on my desk with 128GB of memory. It is a real research rig at a tiny fraction of data-center cost. Most of the hard questions, which tools to use, which model, how to make everything behave, got beaten out on that small box first, where a mistake costs minutes instead of metered GPU-hours. The little box is the cheap rig that de-risks the expensive one.</p><p>The borrowed node of eight H100s, running in the cloud, gets to stand on all of that. Eight H100s, about as much compute as you can put in a single machine, is enough hardware to train a foundation model from scratch and run dozens of experiments beside it. That is the thing the small box could never do. The lineage is the point: de-risk on the desktop, then deploy at scale in the cloud. And the clock is running. Two months, then the iron goes back.</p><blockquote><p><strong>The DGX Spark on the desk taught us the method for a year. The borrowed node is where we finally get to run it.</strong></p></blockquote><div><hr></div><h2>The pace</h2><p>What follows is the short version of a long few days. We brought the node up from a web console to a first running job, stood up a private reasoning model on the box so the team's tools ran on our own hardware, and split its storage into permanent and scratch on a machine that erases itself if you misfile a single thing. We gave the whole team secure access through one entry point, nobody holding the keys to wipe it, and staged terabytes of retinal imagery across four kinds of eye scan, CFP, IR, FAF, and OCTA, without losing a day to a careless mistake. Before spending a GPU-hour, we read the literature and locked the recipe. That alone caught a starting setting roughly sixteen times too aggressive, before it burned a wasted run. Then we trained. The first foundation-model run came back with an uncomfortable verdict: color photos alone topped out no better than an off-the-shelf vision model. So we did not argue with the data. We pivoted to a richer, multimodal approach on a second kind of scan that had been sitting unused on the box, and launched a fresh run inside the same week.</p><p>Around the training, we built an evaluation suite to score every model on the tasks that count, ran ablation sweeps overnight with the box chaining its own experiments while we slept, and red-teamed a result that looked too good, catching three stacked bugs before any of them shipped as a finding. I put <a href="https://ai.rundatarun.io/practical-applications/waking-the-research-engine">ARIA</a>, the autonomous research engine I built, on the node and watched it run its first experiment on its own, then re-contracted it for a world where compute is suddenly cheap and gave it a cheat-detector and a wall of baselines so nothing scores well by accident. All of it ran across three AI coding agents in three lanes at once, operations, modeling, and the research engine, kept from colliding by a shared board with clear ownership. We even survived a disaster: a sync tool deleted the shared workspace, and we had it all back the same afternoon from versioned history and a nightly backup, then locked the setting so it cannot repeat. And the whole time, four interns stood up their first shared code repository and ran a real lane of the operations, the GPU dashboards and a slice of the ops.</p><p>None of that pace is about working longer hours. It is the <a href="https://rundatarun.io/p/the-harness-is-the-moat">harness</a>. The stack of AI tools I have spent two years building and writing about: the coding agents, the <a href="https://rundatarun.io/p/she-already-built-it">autonomous research engine</a>, the skills that turn a one-line instruction into a finished job. The harness does the grunt work. The five of us do the deciding. Point it at a borrowed node and a mission worth the effort, and all of that is a few days, not a few quarters. It also means you can afford to be wrong fast and right next, the way we were when the first model fell flat and the second one took its place inside the same week.</p><blockquote><p><strong>The pace is not longer hours. It is the harness doing the grunt work while the five of us do the deciding.</strong></p></blockquote><div><hr></div><h2>AIXplore, the lab</h2><p>I have mentioned the other blog in passing before but never properly introduced it. <a href="https://ai.rundatarun.io">AIXplore</a> is the lab. Where Run Data Run tells the story, AIXplore shows the work: the engineering, in detail, for the people who do sit in the code. How you give a team access to a borrowed machine without handing out the keys to wipe it. How you stage terabytes of scans without losing a day. How you serve a model locally when the documentation lags the code by a month. Same work as here, one altitude down.</p><p>It got a full rewrite this week, too. While ARIA ran SocialEyes experiments on the node in the background, I rebuilt the entire platform. Every new post carries interactive widgets you can poke at, side quests that branch off into the deep cuts without derailing the main read, and, best of all, a reproduction prompt on every piece. Copy it, point your own Claude Code at it, and build the thing yourself. Whether you have shipped models for years or you are trying to start this weekend, there is a door in.</p><p>The whole SocialEyes build is going up there as it happens, in a series called <a href="https://ai.rundatarun.io/series/borrowed-iron">Borrowed Iron</a>: standing up the node, the access plumbing, the data staging, the autonomous engine, the training runs. If this post made you want the real detail, that is where it lives. And almost none of it is specific to retinas. The lessons transfer to anyone doing serious work on rented GPUs.</p><div><hr></div><h2>Come along</h2><p>I am on a break between jobs this summer, a real sabbatical, and this is what I am doing with it. I will get to a beach at some point. I will also <a href="https://builder-leader.com">publish a book</a> before the summer is out. And the rest of it goes to a borrowed supercomputer, a founder, and four interns, building healthcare for people a long way from a hospital. This is how a sabbatical should go. I figure I will sleep once I start the new job in August.</p><p>So that is the setup. A mission I believe in. A partner who backed it with serious hardware. An unlikely little crew, two of us and four interns, moving faster than the crew size has any right to. And a standing invitation to watch us make the most of two borrowed months in public.</p><p>New parts land as the work happens, here and on AIXplore both. Come along.</p><p><em>Justin writes Run Data Run for the people who lead the work, and AIXplore (ai.rundatarun.io) for the people who build it. The Borrowed Iron series runs on AIXplore as the SocialEyes build happens. If this was useful, the easiest way to support it is to subscribe and forward it to one person who would want it.</em></p>]]></content:encoded></item><item><title><![CDATA[The Harness Is the Moat]]></title><description><![CDATA[The model is the part of your AI strategy you can order off a price sheet. The thing you build around it is the part nobody can copy. I made mine public so you can see the shape of it.]]></description><link>https://rundatarun.io/p/the-harness-is-the-moat</link><guid isPermaLink="false">https://rundatarun.io/p/the-harness-is-the-moat</guid><dc:creator><![CDATA[Justin Johnson]]></dc:creator><pubDate>Thu, 25 Jun 2026 11:29:55 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/e6607538-817e-4148-939a-a6d92a8e283d_1376x768.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!uDli!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa7059c64-a3c8-49c3-98df-e7fab56dcb51_1376x768.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!uDli!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa7059c64-a3c8-49c3-98df-e7fab56dcb51_1376x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!uDli!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa7059c64-a3c8-49c3-98df-e7fab56dcb51_1376x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!uDli!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa7059c64-a3c8-49c3-98df-e7fab56dcb51_1376x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!uDli!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa7059c64-a3c8-49c3-98df-e7fab56dcb51_1376x768.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!uDli!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa7059c64-a3c8-49c3-98df-e7fab56dcb51_1376x768.jpeg" width="728" height="409.5" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a7059c64-a3c8-49c3-98df-e7fab56dcb51_1376x768.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:false,&quot;imageSize&quot;:&quot;normal&quot;,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:728,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;captionedImage&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!uDli!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa7059c64-a3c8-49c3-98df-e7fab56dcb51_1376x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!uDli!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa7059c64-a3c8-49c3-98df-e7fab56dcb51_1376x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!uDli!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa7059c64-a3c8-49c3-98df-e7fab56dcb51_1376x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!uDli!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa7059c64-a3c8-49c3-98df-e7fab56dcb51_1376x768.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Earlier this month I wrote that <a href="https://rundatarun.io/p/you-dont-have-to-write-the-code">two things held their price</a> while AI made almost everything else cheap. Anthropic had just measured one of them across four hundred thousand coding sessions: what the human brings to the work, the command of a domain that lets you direct the machine and catch it when it is confidently wrong. That number was the closest thing we have to proof of an argument I have been making for a while now.</p><p>But there were two things, not one. The second one Anthropic could not measure, because it does not live in their data. It lives in yours. It is the system you build around the model, and it is the part of your AI strategy that no competitor can copy.</p><p>Recently I made mine public so you can look at it. It is a project called Claudelicious, and before I open it up, I want to convince you why the category it belongs to is where your real advantage hides.</p><blockquote><p><strong>The model is the part of your AI strategy you can order off a price sheet. The part that compounds is the part you build around it.</strong></p></blockquote><div><hr></div><h2>The part everyone shops for is the cheap part</h2><p>Most of the AI strategy conversations I sit in still open with the same question. Which model. Which one is smartest this month, which one is cheapest, which one to standardize on. It is a reasonable question to ask once, and a strange one to keep losing sleep over. The models keep getting cheaper and they keep converging; the gap between the best one and the third-best one is a few weeks, not a few years, and the price of yesterday's best falls by half on a schedule you could set a watch to. <strong>Picking a model is a decision you will remake four times before this sentence feels dated.</strong> It is the commodity layer: you can buy it, you can switch it, and so can the company across the street.</p><p>Last month I wrote that <a href="https://rundatarun.io/p/the-model-is-no-longer-the-frontier">the model is no longer the frontier</a>. The first academic conference on agentic systems had just convened, and almost none of its papers touched the model itself. The interesting problems had all moved to the layer around the model: how it improves itself, whether you can tell when it is wrong, what it is allowed to do. That was the research community's verdict. If the frontier moved to the system around the model, then the durable advantage moved there too.</p><p>I have been circling this since last fall, when I started writing about <a href="https://rundatarun.io/p/the-quiet-week-claude-became-your">Claude as a colleague rather than a chatbot</a>. I put the explicit flag down this spring, in a piece called <a href="https://rundatarun.io/p/start-with-claude-code">Start With Claude Code</a>: a year into running this setup, the harness is the moat. This is the full case for it.</p><div><hr></div><h2>What a harness actually is</h2><p>A raw model, even a very good one, is a brilliant contractor with no memory. It shows up sharp, does excellent work for an hour, and then forgets everything the moment the session ends. It does not know your conventions. It does not remember the mistake it made last Tuesday. It cannot reach for the right tool unless you hand it over, every single time.</p><p>The harness is everything you build to fix that. It is the layer that turns a forgetful genius into a system that gets better at <em>your</em> work over time. You can hold all of mine in your head without an engineering degree:</p><ul><li><p><strong>Rules</strong> the model reads at the start of every session, so it already knows how I work before I ask for anything.</p></li><li><p><strong>Skills</strong> it can reach for, small packaged procedures for recurring jobs, so I describe the outcome and it picks the right one. I keep 48 of them.</p></li><li><p><strong>Memory</strong> it carries between sessions, in four tiers, so a lesson learned today is still there next week. Behind it sits roughly seventy thousand documents of accumulated context about my work.</p></li><li><p><strong>A learning loop</strong> that catches its own mistakes and writes them down, so the same error does not happen twice.</p></li><li><p><strong>Continuity</strong> that lets me close the laptop mid-thought and have the next session pick up exactly where this one stopped.</p></li><li><p><strong>Always-on agents</strong>, eight of them spread across a small mesh of machines, that keep working long-horizon jobs while I sleep.</p></li></ul><p>None of that is the model. All of it is the thing that decides whether the model is useful to me specifically. A model is a thing you rent for an hour. A harness is a thing you live inside.</p><blockquote><p><strong>A raw model is a brilliant contractor with no memory. The harness is everything that turns it into a colleague who remembers.</strong></p></blockquote><p>This is also why the choice of harness has a character to it, which is the argument I made in <a href="https://rundatarun.io/p/three-harnesses-three-characters">Three Harnesses, Three Characters</a>: the same underlying model behaves like a different colleague depending on the rig you wrap it in. You are not picking an engine. You are picking who shows up to work.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!j4O7!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1a74e7ec-eff3-45da-955f-fcde3b6050dc_1376x768.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!j4O7!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1a74e7ec-eff3-45da-955f-fcde3b6050dc_1376x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!j4O7!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1a74e7ec-eff3-45da-955f-fcde3b6050dc_1376x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!j4O7!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1a74e7ec-eff3-45da-955f-fcde3b6050dc_1376x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!j4O7!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1a74e7ec-eff3-45da-955f-fcde3b6050dc_1376x768.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!j4O7!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1a74e7ec-eff3-45da-955f-fcde3b6050dc_1376x768.jpeg" width="728" height="409.5" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/1a74e7ec-eff3-45da-955f-fcde3b6050dc_1376x768.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:false,&quot;imageSize&quot;:&quot;normal&quot;,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:728,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;captionedImage&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!j4O7!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1a74e7ec-eff3-45da-955f-fcde3b6050dc_1376x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!j4O7!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1a74e7ec-eff3-45da-955f-fcde3b6050dc_1376x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!j4O7!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1a74e7ec-eff3-45da-955f-fcde3b6050dc_1376x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!j4O7!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1a74e7ec-eff3-45da-955f-fcde3b6050dc_1376x768.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div><hr></div><h2>What makes a harness great</h2><p>Not every harness is worth building. A bad one is a pile of half-used scripts that slows you down and a maintenance burden you resent. A great one has three properties, and none of them is about how smart the model is.</p><p><strong>It is shaped to you.</strong> Every rule encodes a preference I formed by getting something wrong. Every skill is a procedure I refined over dozens of runs. Every memory is a piece of context about my work, my data, my judgment calls that took months to accumulate. A competitor can license the identical model tomorrow. They cannot license eighteen months of my context, because it does not exist anywhere to license.</p><p><strong>It learns.</strong> Because the harness has a learning loop, it does not sit still. It gets a little better every time it catches a mistake, and it does that whether I am watching or not. So the distance between a tuned harness and a fresh one does not hold steady. It grows. The organization that started building this a year ago is not a year ahead. It is a year ahead and pulling away.</p><p><strong>It stays small.</strong> This is the property people get wrong, and getting it wrong is how they waste a year. More scaffolding is not better. Every rule the model reads, every skill in its reach, is something it holds in its attention on every single turn, and attention is finite. My own skill library peaked at 134 packaged procedures. I cut it to 48, and the cut was the work, not the cleanup. <strong>Compounding does not mean accumulating. It means keeping the ones that work and ruthlessly cutting the ones that don't.</strong></p><p>Put those three together and you have something a competitor cannot buy, because it is personal, and cannot catch, because it keeps moving. That is the moat, and it is the reverse of the model question. The model is the layer where every company is equal.</p><blockquote><p><strong>The model is where everyone is equal. The harness is where the distance opens up.</strong></p></blockquote><div><hr></div><h2>How the harness crosses the gap</h2><p>Now widen the lens, because this is where the harness stops being a productivity story and becomes a strategy one.</p><p>There is a real divide in how people see AI right now, and it is not skeptics against believers. Andrej Karpathy named it this spring: two groups talking past each other. The split that matters is between the people who have built something with these systems and the people who have only read about them, used the free tier, or sat through a demo. Almost every boardroom argument about what AI can and cannot do is a symptom of that divide, not a real debate. The two sides are calibrated to different machines, different use cases, different tiers of the thing.</p><p>You cannot read your way across that gap. You cannot procure your way across it either. <strong>You cross it by building, and building does not mean writing code.</strong> It means directing a harness to produce real work inside a domain you already command. The harness is the thing you wear to do it. My book calls it the exoskeleton, which is the whole title: <a href="https://builder-leader.com">Builder-Leader: The AI Exoskeleton That Crosses the Gap</a>. The skills to operate one are not new. They are the same instincts that got you to senior in the first place: setting intent and the boundaries around it, designing the handoffs, judging the output against what good actually looks like. Pointed now at a system of agents instead of a team of people.</p><p>This is where those pieces and this one are the same argument seen from different sides. <strong>You don't have to write the code</strong> was the human half <em>measured</em>: domain command, not coding, is what predicts whether you get working results out of an agent, across four hundred thousand sessions. <a href="https://rundatarun.io/p/what-ai-didnt-reprice">What AI didn't reprice</a> was the human half <em>priced</em>: when output gets cheap, the judgment that knows which output is right gets dearer, and it is the one asset on your books that went up in value this year. The harness is the other half, the system that judgment directs. The two are useless apart. A great harness pointed by someone who does not know the field just produces wrong answers faster. Deep expertise with no harness leaves most of its own capability sitting on the table.</p><blockquote><p><strong>A leader who commands the domain and has built the harness holds a position the company across the street cannot copy. Not the model. The pairing.</strong></p></blockquote><p>That pairing is the whole point. One half is what you carry in your head. The other half is what you build around the model to put that head to work at machine speed. Together they are the thing no price sheet sells.</p><div><hr></div><h2>Why I made it public</h2><p>So I wrote mine down. Not as a tutorial, and not as a flex. As a cookbook of <em>why</em>: why each piece exists, what problem it solves, how the parts fit together, and where I drew the line. It is called Claudelicious, and it lives on <a href="https://github.com/BioInfo/claudelicious">GitHub</a>.</p><p>It is built around one idea: run the model as a system, not a chat box. And it is deliberately <strong>not a catalog.</strong> The community already maintains a good directory for finding a skill, a hook, a connector. Claudelicious is the other half, the part you cannot reverse-engineer from a screenshot: how those pieces fit together into a harness that remembers across sessions, corrects itself when you catch it, and keeps working while you sleep.</p><p>So it walks the six pieces I named earlier and shows the wiring under each. The rules that make the model know how you work before you ask. The skill library, and the discipline of keeping it small. The memory and the second brain behind it. The learning loop that makes a correction stick instead of recurring. The continuity that survives a closed laptop. The always-on agents and the small fleet they run on. <strong>Every chapter is principle first, then a worked example, then plain notes on what to copy and what to leave behind.</strong> There is a quickstart for getting the spine in place in your first week, and a longer version for reading the whole thing as one story.</p><p>It is also the working companion to the book. The book makes the case that the people who win with AI are not the ones who pick the best model, but the ones who build the system around it and lead from inside it. Claudelicious is that case made concrete: one such system, with the reasoning shown, that you can open and read.</p><p>I made it public for a specific reason. <strong>You cannot buy a harness.</strong> There is no SKU for the thing I have described, and there never will be, because the whole point is that it is yours. But you can look at the shape of someone else's, see which parts map to your work, and decide what is worth building and maintaining in your own organization. A worked example beats an abstract argument. So here is mine, with the reasoning shown.</p><p>Do not walk away with a shopping list. Walk away with the question underneath it.</p><blockquote><p><strong>Your AI edge is not on the model menu. It is the answer to a question no vendor will answer for you.</strong></p></blockquote><p>The question is this. A year from now, the model everyone runs will be cheaper and smarter than today's and roughly the same as your competitor's. So what will be different about how <em>your</em> organization uses it? What did you build around it that the next org did not, that learned from your work while theirs sat still, that no one can license because it is made of your own accumulated judgment?</p><p>If the answer is nothing, that is the gap to close this quarter. <strong>Not which model. What is the harness, who owns it, and is it learning.</strong></p><div><hr></div><h2>Where this goes next</h2><p>The harness is one of the two things AI did not make cheaper. The other is the thing you point it with, and that one did not just survive the markdown. It got more expensive. I made that case <a href="https://rundatarun.io/p/what-ai-didnt-reprice">earlier this week</a>: when output gets cheap, the scarce ability to judge it gets dearer, and the judgment of someone who knows your field is the one line on the books that AI repriced upward. The harness is the system. That judgment is what tells the system where to aim. Most teams are still budgeting for neither.</p><p>For now, sit with the pair. A leader with a domain in their head and a harness around the model is not waiting for a better model to arrive. They are already pulling away from the ones who are.</p><div><hr></div><h2>Sources</h2><ul><li><p><a href="https://rundatarun.io/p/you-dont-have-to-write-the-code">You Don't Have to Write the Code</a>. Anthropic's 400,000-session study, the population-scale measurement of the human half: domain command, not coding, predicts success.</p></li><li><p><a href="https://rundatarun.io/p/what-ai-didnt-reprice">What AI Didn't Reprice</a>. The pricing of that same human half: judgment is the one asset AI marked up while it marked almost everything else down.</p></li><li><p><a href="https://rundatarun.io/p/the-model-is-no-longer-the-frontier">The Model Is No Longer the Frontier</a>. The first agentic-systems conference, where the research frontier moved off the model and onto the layer around it.</p></li><li><p><a href="https://rundatarun.io/p/the-quiet-week-claude-became-your">The Quiet Week Claude Became Your Coworker</a>. Last fall, where the colleague-not-chatbot version of this argument started.</p></li><li><p><a href="https://rundatarun.io/p/start-with-claude-code">Start With Claude Code</a>. This spring, where I put the explicit flag down: the harness is the moat.</p></li><li><p><a href="https://rundatarun.io/p/three-harnesses-three-characters">Three Harnesses, Three Characters, One Working Week</a>. The same model behaves like a different colleague depending on the rig around it.</p></li><li><p><a href="https://github.com/BioInfo/claudelicious">Claudelicious</a>. The public cookbook: the why and the wiring behind a full harness.</p></li><li><p><a href="https://builder-leader.com">Builder-Leader: The AI Exoskeleton That Crosses the Gap</a>. The book the harness is the working companion to.</p></li></ul><div><hr></div><p><em>Run Data Run is free, no paywall. If this was useful, the easiest way to support it is to subscribe and forward it to one person on your team who'd want it. If it wasn't, tell me why. I'll make it better.</em></p>]]></content:encoded></item><item><title><![CDATA[Ep. 1: Two Groups]]></title><description><![CDATA[Same AI, same task, a different result. The difference isn't the model.]]></description><link>https://rundatarun.io/p/ep-1-two-groups</link><guid isPermaLink="false">https://rundatarun.io/p/ep-1-two-groups</guid><dc:creator><![CDATA[Justin Johnson]]></dc:creator><pubDate>Tue, 23 Jun 2026 12:06:14 GMT</pubDate><enclosure url="https://api.substack.com/feed/podcast/203232843/1a09d4a1a72d8840d995d28664a70485.mp3" length="0" type="audio/mpeg"/><content:encoded><![CDATA[<p>Two people pay for the same AI, run the same model, hand it the same task. One gets work the other can&#8217;t come close to. Episode one is about why, and why the gap that decides your next two years isn&#8217;t the one everyone is arguing about.</p><h3>In this episode</h3><ul><li><p>The two stories you heard this week, and why both are true</p></li><li><p>Why &#8220;the truth is somewhere in the middle&#8221; is the expensive wrong answer</p></li><li><p>The distribution gap: where the gains are actually clustering, and why code sprints while writing, search, and advice crawl</p></li><li><p>The split nobody covered: inside the paid tier, the people who use these systems versus the people who build the harness and drive them</p></li><li><p>Why this isn&#8217;t just another early-adopter curve that closes on its own</p></li><li><p>What &#8220;driving&#8221; actually looks like, and why most people never cross over</p></li><li><p>The one new muscle to add if your career was built on judgment, not code</p></li></ul><h3>Referenced in this episode</h3><ul><li><p>Andrej Karpathy&#8217;s April thread on the two groups talking past each other</p></li><li><p>Jason Lemkin (SaaStr): &#8220;how awed you are by AI is almost perfectly correlated with how much you actually use it to build&#8221;</p></li><li><p>Anthropic&#8217;s report on how agents actually get used (software engineering was nearly half of all agent activity)</p></li><li><p>The roughly one-million-conversation study on how people use ChatGPT (coding around four percent of messages)</p></li><li><p>Gary Marcus calling Claude Code the single biggest advance in AI since the large language model</p></li><li><p>2025 McKinsey, Gartner, and IT-leader surveys on stalled AI projects and agent sprawl</p></li></ul><h3>From the book</h3><p>This is chapter one of Builder Leader: The AI Exoskeleton That Crosses the Gap. Each episode takes one idea from one chapter and talks it through. Subscribe to get the next one.</p>]]></content:encoded></item><item><title><![CDATA[What AI Didn't Reprice]]></title><description><![CDATA[AI marked down almost every skill on your team this year. One it marked up. If you run a budget, that one number should change where you're spending.]]></description><link>https://rundatarun.io/p/what-ai-didnt-reprice</link><guid isPermaLink="false">https://rundatarun.io/p/what-ai-didnt-reprice</guid><dc:creator><![CDATA[Justin Johnson]]></dc:creator><pubDate>Mon, 22 Jun 2026 12:02:31 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/36179c56-9a7c-4efa-a798-73fa8ab3f7a3_1376x768.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!j94_!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F28c600a8-2160-46d1-8ddc-226fda99e591_1376x768.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!j94_!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F28c600a8-2160-46d1-8ddc-226fda99e591_1376x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!j94_!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F28c600a8-2160-46d1-8ddc-226fda99e591_1376x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!j94_!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F28c600a8-2160-46d1-8ddc-226fda99e591_1376x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!j94_!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F28c600a8-2160-46d1-8ddc-226fda99e591_1376x768.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!j94_!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F28c600a8-2160-46d1-8ddc-226fda99e591_1376x768.jpeg" width="728" height="409.5" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/28c600a8-2160-46d1-8ddc-226fda99e591_1376x768.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:false,&quot;imageSize&quot;:&quot;normal&quot;,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:728,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;captionedImage&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!j94_!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F28c600a8-2160-46d1-8ddc-226fda99e591_1376x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!j94_!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F28c600a8-2160-46d1-8ddc-226fda99e591_1376x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!j94_!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F28c600a8-2160-46d1-8ddc-226fda99e591_1376x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!j94_!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F28c600a8-2160-46d1-8ddc-226fda99e591_1376x768.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Think of the last two years of AI as one long, indiscriminate markdown sale. Skill after skill that used to be scarce got cheap, fast. Writing a clean function. Drafting a competent memo. Summarizing a dense report. Producing a first-pass analysis. A year ago each of those was something you hired for, waited on, or did yourself at the cost of an afternoon. Now they are a prompt and a few seconds.</p><p>The sale hit almost everything. It missed one thing, and the thing it missed didn't just hold its price. It went up.</p><p>Last week I <a href="https://rundatarun.io/p/you-dont-have-to-write-the-code">put four hundred thousand sessions behind that claim</a>: Anthropic's analysis of real coding work found that what predicted success wasn't your job title or your syntax, it was whether you understood the problem. That was the evidence. The implication: a few things got more valuable this year, not less, while everything around them got cheaper.</p><blockquote><p><strong>Almost everything got cheaper this year. The ability to tell right work from confident-looking wrong work got more expensive.</strong></p></blockquote><div><hr></div><h2>The half that got cheap was the doing</h2><p>Start with what fell in price.</p><p>What AI commoditized is the <em>execution</em> of a known task. Give a model a well-specified job inside a domain it has seen a million examples of, and it will do it about as well as a competent junior, for free, instantly, at three in the morning. That is most of what people mean when they say AI is changing knowledge work, and they are right about it.</p><p>I've made this argument from a few angles already. When <a href="https://rundatarun.io/p/the-model-is-no-longer-the-frontier">the model stopped being the frontier</a>, the interesting problems all moved to the layer around it. When <a href="https://rundatarun.io/p/apple-rented-its-brain">Apple rented its brain</a>, it conceded that the model itself is now the rentable, swappable part of the stack, even for the most vertically integrated company on earth. The engine is a commodity. You can buy it, switch it, and so can your competitor.</p><p>So if the doing got cheap, the question is what didn't, and why it climbed while everything around it fell.</p><div><hr></div><h2>The half that got expensive was the judging</h2><p>"AI does the work" treats the work as if it were one thing. It isn't. Every task has a doing half and a judging half, and AI only ate the first one.</p><p>The doing is producing the output. The judging is knowing whether the output is right. Not right in general, not right on a benchmark, but right <em>here</em>, for this patient, this contract, this market, this dataset with its ten-year-old quirk that everyone who has worked it knows about and no document anywhere records.</p><p>That second half is domain expertise, and it resists the markdown for a structural reason: it isn't in the training data. A model learns from what people wrote down. The deepest domain knowledge was never written down. It lives in the head of the person who has run the assay four thousand times, sat across from the regulator, watched the trade go wrong in 2019, and can look at a fluent, plausible, beautifully formatted answer and say, with a flat certainty, "no, that's not how this behaves." You cannot prompt your way to that, because the corpus the model trained on doesn't contain it.</p><p>This follows a basic rule about complements. When you flood a market with cheap supply of one thing, the price of its complement goes <em>up</em>. Cheap output makes the scarce ability to judge that output worth more, not less, because now there is a hundred times more output to judge and the same small number of people who can tell which of it is wrong. AI didn't just spare domain expertise. It bid the price up.</p><blockquote><p><strong>Cheap output doesn't lower the value of judgment. It raises it. There's a hundred times more to judge and the same few people who can.</strong></p></blockquote><div><hr></div><h2>Two voices, different worlds, one conclusion</h2><p>What turned this from a hunch into a post was watching it land from people who don't share a desk, a discipline, or a reason to agree.</p><p>In March, the Berkeley <em>California Management Review</em>, about as institutional as business thinking gets, <a href="https://cmr.berkeley.edu/2026/03/tacit-knowledge-is-your-next-competitive-moat/">told executives their next competitive moat is tacit knowledge</a>: not the data, not the models, but the judgment embedded in their people, and the systems that capture it before it walks out the door. Two months later, a working software engineer named Aaron Brethorst, with no reason to read a business-school journal, <a href="https://www.brethorsting.com/blog/2026/05/domain-expertise-has-always-been-the-real-moat/">wrote that domain expertise has always been the real moat</a>. His point: the framework knowledge, the syntax, the boilerplate, the part everyone used to grind years to acquire, was never the durable thing. AI making it free didn't destroy the moat. It drained the water and showed you where the moat actually was the whole time.</p><p>One came at it from the top of the org chart, thinking about strategy and knowledge graphs. The other came at it from the keyboard, watching a model write code he then had to check line by line. Different altitude, different vocabulary, different month. Same conclusion. Put their two arguments next to the four hundred thousand sessions from last week and you have a management journal, a working engineer, and a population-scale measurement all pointing at the same line through the work. When the same claim arrives from corners that don't read each other, the content is almost secondary. The convergence is the signal.</p><div><hr></div><h2>The one asset that appreciated</h2><p>Domain expertise is the one appreciating asset on the books. Everything else in your AI strategy is depreciating. The model you standardized on this quarter will be cheaper and roughly matched by your competitor's within months; you'll remake that decision four times before it matters.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!HTsN!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F853a514e-f0a1-4c7b-a106-b95e7a5ee047_1200x896.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!HTsN!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F853a514e-f0a1-4c7b-a106-b95e7a5ee047_1200x896.jpeg 424w, https://substackcdn.com/image/fetch/$s_!HTsN!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F853a514e-f0a1-4c7b-a106-b95e7a5ee047_1200x896.jpeg 848w, https://substackcdn.com/image/fetch/$s_!HTsN!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F853a514e-f0a1-4c7b-a106-b95e7a5ee047_1200x896.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!HTsN!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F853a514e-f0a1-4c7b-a106-b95e7a5ee047_1200x896.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!HTsN!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F853a514e-f0a1-4c7b-a106-b95e7a5ee047_1200x896.jpeg" width="728" height="409.5" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/853a514e-f0a1-4c7b-a106-b95e7a5ee047_1200x896.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:false,&quot;imageSize&quot;:&quot;normal&quot;,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:728,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;captionedImage&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!HTsN!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F853a514e-f0a1-4c7b-a106-b95e7a5ee047_1200x896.jpeg 424w, https://substackcdn.com/image/fetch/$s_!HTsN!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F853a514e-f0a1-4c7b-a106-b95e7a5ee047_1200x896.jpeg 848w, https://substackcdn.com/image/fetch/$s_!HTsN!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F853a514e-f0a1-4c7b-a106-b95e7a5ee047_1200x896.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!HTsN!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F853a514e-f0a1-4c7b-a106-b95e7a5ee047_1200x896.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Which changes what the work actually is. The advantage isn't holding the expertise. It's how fast you're turning it into something the machine can use before the person carrying it retires, quits, or simply forgets the quirk from 2019: into context, into the checks that catch a wrong answer before it ships, into the loop that lets one expert's judgment steer a hundred agents instead of one. That capture is an investment with a return, and most organizations aren't making it because they still have expertise filed under cost.</p><p>There's a second thing AI didn't reprice, and it pairs with this one: the system you build around the model, the harness that turns a forgetful chat box into something that compounds on your work. I put the flag down on that one in <a href="https://rundatarun.io/p/start-with-claude-code">Start With Claude Code</a>, and it's the subject of this Sunday's deep dive. A great harness pointed by someone who doesn't know the field just produces wrong answers faster. The judgment is what tells the system what good looks like. The two are a pair, and neither is on the model menu.</p><blockquote><p><strong>Everything in your AI strategy is depreciating except two things: the judgment of the people who know your field, and the system you build to put it to work.</strong></p></blockquote><div><hr></div><h2>The decision that compounds</h2><p>The decision that compounds isn't which model to pick. You remake that one on a schedule, and so does everyone else.</p><p>The companies that win the next few years won't be the ones with the best model. Everyone has roughly the same model. They'll be the ones who looked at the markdown sale, noticed the single thing whose price went the other way, and spent like they understood why.</p><div><hr></div><h2>Sources</h2><ul><li><p><a href="https://rundatarun.io/p/you-dont-have-to-write-the-code">You Don't Have to Write the Code</a> (Run Data Run). Anthropic's 400,000-session study: domain understanding, not job title or coding skill, predicted success. The evidence under this post's argument.</p></li><li><p>Berkeley <em>California Management Review</em>, <a href="https://cmr.berkeley.edu/2026/03/tacit-knowledge-is-your-next-competitive-moat/">Tacit Knowledge Is Your Next Competitive Moat</a> (March 2026). The institutional framing: the differentiator that lasts is the judgment embedded in your people, and the systems that capture it.</p></li><li><p>Aaron Brethorst, <a href="https://www.brethorsting.com/blog/2026/05/domain-expertise-has-always-been-the-real-moat/">Domain Expertise Has Always Been the Real Moat</a> (May 2026). The builder's framing: framework knowledge going free made the real moat visible.</p></li><li><p><a href="https://rundatarun.io/p/start-with-claude-code">Start With Claude Code</a> (Run Data Run). The other thing AI didn't reprice: the harness you build around the model.</p></li></ul>]]></content:encoded></item><item><title><![CDATA[You Don't Have to Write the Code]]></title><description><![CDATA[Anthropic watched 400,000 sessions with a coding agent and found that what predicts success isn't your job or your syntax. It's whether you understand the work. That changes who gets to build, and what your expertise is worth.]]></description><link>https://rundatarun.io/p/you-dont-have-to-write-the-code</link><guid isPermaLink="false">https://rundatarun.io/p/you-dont-have-to-write-the-code</guid><dc:creator><![CDATA[Justin Johnson]]></dc:creator><pubDate>Wed, 17 Jun 2026 10:07:01 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/a3b81ead-a9e7-47dc-8fb3-3052d58b3a01_1408x768.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!1mzG!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ec3ec3c-01e7-4551-b000-e8aca7896a73_1408x768.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!1mzG!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ec3ec3c-01e7-4551-b000-e8aca7896a73_1408x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!1mzG!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ec3ec3c-01e7-4551-b000-e8aca7896a73_1408x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!1mzG!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ec3ec3c-01e7-4551-b000-e8aca7896a73_1408x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!1mzG!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ec3ec3c-01e7-4551-b000-e8aca7896a73_1408x768.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!1mzG!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ec3ec3c-01e7-4551-b000-e8aca7896a73_1408x768.jpeg" width="728" height="409.5" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/3ec3ec3c-01e7-4551-b000-e8aca7896a73_1408x768.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:false,&quot;imageSize&quot;:&quot;normal&quot;,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:728,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;captionedImage&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!1mzG!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ec3ec3c-01e7-4551-b000-e8aca7896a73_1408x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!1mzG!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ec3ec3c-01e7-4551-b000-e8aca7896a73_1408x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!1mzG!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ec3ec3c-01e7-4551-b000-e8aca7896a73_1408x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!1mzG!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ec3ec3c-01e7-4551-b000-e8aca7896a73_1408x768.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>This week Anthropic published <a href="https://www.anthropic.com/research/claude-code-expertise">an analysis of about 400,000 real Claude Code sessions</a>, run between last October and this April. It set out to answer a plain question: when someone sits down with a coding agent, what actually predicts whether they succeed? There is a comfortable answer and a more useful one, and almost everyone is repeating the comfortable one.</p><p>The comfortable answer, the one the headlines took, is that anyone can build software now. That part is true, and it is the least interesting thing in the report.</p><p>The finding underneath is the one I'd hand to anyone, whether they run a team or just their own work. What predicted success was not your job title, and it was not whether you could write code. It was whether you understood the problem you were trying to solve. Anthropic's own framing: success comes from how well a person understands the work, not whether they're trained in coding.</p><p>That is not a story about AI flattening the gap between your best people and your average ones. It is the opposite. The thing you spent fifteen years getting good at just became the thing that decides whether the most expensive tool you're buying actually pays off. I have been making that case for a while. This is the first time anyone has put 400,000 sessions behind it.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!1dkG!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd63cd64c-afe6-4bda-a07a-c7dbd6a6a6e0_1408x768.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!1dkG!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd63cd64c-afe6-4bda-a07a-c7dbd6a6a6e0_1408x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!1dkG!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd63cd64c-afe6-4bda-a07a-c7dbd6a6a6e0_1408x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!1dkG!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd63cd64c-afe6-4bda-a07a-c7dbd6a6a6e0_1408x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!1dkG!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd63cd64c-afe6-4bda-a07a-c7dbd6a6a6e0_1408x768.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!1dkG!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd63cd64c-afe6-4bda-a07a-c7dbd6a6a6e0_1408x768.jpeg" width="728" height="409.5" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d63cd64c-afe6-4bda-a07a-c7dbd6a6a6e0_1408x768.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:false,&quot;imageSize&quot;:&quot;normal&quot;,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:728,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;captionedImage&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!1dkG!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd63cd64c-afe6-4bda-a07a-c7dbd6a6a6e0_1408x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!1dkG!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd63cd64c-afe6-4bda-a07a-c7dbd6a6a6e0_1408x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!1dkG!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd63cd64c-afe6-4bda-a07a-c7dbd6a6a6e0_1408x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!1dkG!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd63cd64c-afe6-4bda-a07a-c7dbd6a6a6e0_1408x768.jpeg 1456w" sizes="100vw"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div><hr></div><h2>What the tool actually divided up</h2><p>Start with the cleanest number, because it doubles as a mental model you can carry into a Monday meeting.</p><p>In a typical session, the human made about <strong>70 percent of the planning decisions</strong> and the agent made about <strong>80 percent of the execution decisions</strong>. People decided what to build. The agent decided how to build it. That split held across every kind of work they measured, from writing code to running systems to analyzing data.</p><p>So the labor didn't disappear. It separated. The agent took the part a lot of people thought was the moat, the ability to actually produce the syntax, and handed back the part that was always the hard part: knowing what to ask for, and whether the answer is any good.</p><p>Here is what makes me trust the number. A practitioner reached the same shape from the other side. Addy Osmani, writing in January, found that the developers succeeding with these tools "spend 70 percent of their time on problem definition and verification strategy, 30 percent on execution." Two independent measurements, one watching a population and one watching working engineers, landing on the same line through the work. The tool made the typing cheap. It did nothing for the knowing.</p><blockquote><p><strong>The tool made the typing cheap. It did nothing for the knowing.</strong></p></blockquote><p>This is the thing I've been calling <a href="https://rundatarun.io/p/the-atomic-unit-of-work-just-changed">a change in the atomic unit of work</a>. The unit a person owns moved up the stack, from producing the thing to directing it. Now there's telemetry under the claim.</p><div><hr></div><h2>Your job title barely moved the needle</h2><p>This is the result Anthropic led with, and it earns the lead.</p><p>They scored a verified success rate, which is stricter than it sounds: a session counts as a win only if the model judged it successful and there was a hard signal to back that up, a real commit, a passing test, a user saying yes. By that bar, on code-producing work, software occupations succeeded <strong>34 percent</strong> of the time and everyone else succeeded <strong>29 percent</strong>. A marketer and a staff engineer finished a coding task at almost the same rate.</p><p>The instinct is to read that as "engineers are finished." It is the wrong read, and the same study hands you the right one. What collapsed was the premium on being a programmer. What held was the premium on understanding the problem. Measure expertise correctly and it mattered enormously: novice sessions succeeded <strong>15 percent</strong> of the time, and people who actually knew their domain landed between <strong>28 and 33 percent</strong>. Roughly double.</p><p>The reconciliation is the part that clarifies it, because it is the bet I've been making. Expertise here is not your title or your r&#233;sum&#233;. It is task-specific. Anthropic's own example: a senior engineer asking their first question about the Rust language is a beginner at Rust. The skill that predicts whether you get working software out of an agent was never "I am an engineer." It is "I understand this particular problem well enough to direct the work and catch it when it goes wrong."</p><p>That is the whole argument for the builder-leader, measured now across a population instead of asserted from a single desk. You do not have to become an engineer. You do not have to write the code. You direct it, inside a domain you already command, and the command is the thing that pays.</p><blockquote><p><strong>You don't have to become an engineer. You don't have to write the code. You direct it, inside a domain you already command.</strong></p></blockquote><div><hr></div><h2>The expert tell is recovery</h2><p>If I had to keep one number from the whole report, it would be this one.</p><p>Experts didn't just prompt better on a good day, though they did do more with each instruction: about <strong>12 agent actions per prompt</strong> versus <strong>5</strong> for novices, and roughly five times the output. The real difference showed up when things went sideways. When a novice hit trouble, they walked away with nothing written about <strong>19 percent</strong> of the time. For everyone with more domain knowledge, that abandonment rate was <strong>5 to 7 percent</strong>.</p><p>Sit with what that means. The expert hit the same wall the novice did. Then they routed around it, reframed the problem, and caught the fluent, confident, completely wrong answer before it shipped. The novice hit the wall and quit.</p><p>So the value your senior people add in an agentic workflow is not that they prompt cleaner on a Tuesday. It is that they fail better on a Thursday. They have the judgment to know when the plausible output is wrong, and the agent does not have that judgment and cannot get it from a model update. The bottleneck moved from typing to deciding. Pratima Arora at Smartsheet put it plainly this spring: the hours haven't changed, but the density of work has.</p><p>The scarcest thing in the building is now the thing your best people already carry. That is good news, and most of the coverage will skip right past it.</p><blockquote><p><strong>The scarcest thing in the building is now the thing your best people already carry.</strong></p></blockquote><div><hr></div><h2>The work moved up the stack</h2><p>Two more numbers close the loop, and they are the ones that tell you where this is going.</p><p>Between October and April, the share of sessions spent fixing broken code fell from a third to under a fifth, <strong>33 percent down to 19</strong>. Over the same stretch, the estimated value of the work people brought rose about <strong>27 percent</strong>, with the biggest jump in building something new, up <strong>43 percent</strong>.</p><p>Put those together and the trajectory is clear. People are not using the agent to fix more bugs. They are using it to attempt harder, more valuable, more end-to-end work, and bringing more judgment to bear when they do. The tool got cheaper per task, so the tasks got more ambitious.</p><p>This is the oldest pattern in economics wearing new clothes. Make a thing cheaper to produce and you do not produce less of it. You produce far more, and you need more judgment to steer all of it. That is the shape of this moment, and it cuts against the fear that the work is shrinking. The work isn't shrinking. It changed: what we attempt got bigger, and how we get there moved from doing to directing.</p><div><hr></div><h2>Two things held their price, not one</h2><p>There is a second thread here, and it is the one that turns a study into a strategy.</p><p>Anthropic measured what the human brings to the session, and found that what the human brings, domain command, is decisive. That is one of two things this whole shift never made cheaper. The other is what you build around the model.</p><p>The model is the commodity, and the system you build around it is the moat: the rules and skills and memory that turn a smart, forgetful chat box into something that gets better at your specific work over time. Birgitta B&#246;ckeler, writing on Martin Fowler's site, named the gap the model itself can never close. A coding agent, she wrote, has "no social accountability, no aesthetic disgust at a 300-line function, no intuition that 'we don't do it that way here,' and no organisational memory." Those are not features you buy a newer version of. They come from the person and the system around the model, or they don't come at all.</p><p>So the picture is narrower and more useful than "expertise wins." Two things resisted the repricing. What you bring to the model, which is domain command, and what you build around it, which is the system that holds your organization's judgment. The model in the middle, the part everyone is still shopping for on price, is the cheap layer between them. Anthropic just measured one of those two human pieces at a scale none of us could reach alone. The other one you build.</p><p>And both are leadership skills, not engineering ones. Setting intent, designing the handoffs, evaluating output against what good actually looks like. Those are the same instincts that got your best people to senior in the first place, pointed now at a system of agents instead of a team of people. No new species of human required. That is the part I keep coming back to, and it is why I think the people who internalize this will out-build the ones still arguing about whether the juniors are coming for the seniors.</p><div><hr></div><h2>One honest note</h2><p>At the level of the whole economy, you cannot see this yet in the aggregate numbers. A large Danish study of AI chatbots across occupations found no significant effect on earnings or recorded hours. I take it seriously, and I read it as a statement about timing, not a refutation. Session behavior changes before payroll does. The 400,000 sessions are the leading edge; the wage data lags. Hold the whole argument a little more loosely for it, and don't over-rotate on any single quarter's story.</p><div><hr></div><h2>What I'd do with this Monday</h2><p>If you run a team, here is the part that converts. If you don't, read it as where to put your own time.</p><p>Put the agent in the hands of your domain experts directly, not only your engineers. The advantage is highest exactly where deep knowledge meets the tool. The clinical-operations lead who has watched the trial workflow break in a dozen specific ways will get more out of an agent than a generalist engineer who hasn't, because the agent can produce the code but it cannot supply the judgment about what the code is for.</p><p>Hire and promote for understanding the problem and recovering well, not for who writes the cleanest syntax. The data says occupation barely predicts success and recovery strongly does. That is now a measurable thing to select for.</p><p>And grow your next builder on purpose. The way across this gap was always building, six months of directing real work inside a real domain, not sitting through a demo. The leaders who win the next few years won't be the ones with the best coders. They'll be the ones who turned their domain experts into builders before anyone told them it was allowed.</p><p>The agent decides how. You still decide what, and you still decide who learns to decide what next. Four hundred thousand sessions just told you those are the two jobs worth keeping. They are not a smaller job than the one before. They're the better one.</p><div><hr></div><p>Related: <a href="https://rundatarun.io/p/the-specialist-is-now-you">The Specialist Is Now You</a> is what one expert plus an agent can now do alone, and <a href="https://rundatarun.io/p/the-atomic-unit-of-work-just-changed">The Atomic Unit of Work Just Changed</a> is where the unit a person owns moved up the stack.</p><p><em>The fuller argument for treating judgment, intent, and domain command as the durable skills is the Builder Leader field guide (<a href="http://builder-leader.com">builder-leader.com</a>).</em></p>]]></content:encoded></item><item><title><![CDATA[The Wrong Tool Problem in Genomic AI]]></title><description><![CDATA[A leaderboard tells you which model wins. It can't tell you that you picked the wrong kind of model.]]></description><link>https://rundatarun.io/p/the-wrong-tool-problem-in-genomic</link><guid isPermaLink="false">https://rundatarun.io/p/the-wrong-tool-problem-in-genomic</guid><dc:creator><![CDATA[Justin Johnson]]></dc:creator><pubDate>Sun, 14 Jun 2026 11:23:55 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/320403dd-e4b7-43f9-abc1-ba6e54ee4ac8_1376x768.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!jH4V!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fad9183b6-98f5-4763-b74e-bb73b4838d52_1376x768.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!jH4V!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fad9183b6-98f5-4763-b74e-bb73b4838d52_1376x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!jH4V!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fad9183b6-98f5-4763-b74e-bb73b4838d52_1376x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!jH4V!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fad9183b6-98f5-4763-b74e-bb73b4838d52_1376x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!jH4V!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fad9183b6-98f5-4763-b74e-bb73b4838d52_1376x768.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!jH4V!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fad9183b6-98f5-4763-b74e-bb73b4838d52_1376x768.jpeg" width="728" height="409.5" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ad9183b6-98f5-4763-b74e-bb73b4838d52_1376x768.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:false,&quot;imageSize&quot;:&quot;normal&quot;,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:728,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;captionedImage&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!jH4V!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fad9183b6-98f5-4763-b74e-bb73b4838d52_1376x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!jH4V!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fad9183b6-98f5-4763-b74e-bb73b4838d52_1376x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!jH4V!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fad9183b6-98f5-4763-b74e-bb73b4838d52_1376x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!jH4V!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fad9183b6-98f5-4763-b74e-bb73b4838d52_1376x768.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>Every Sunday I pick one paper or release that's worth your time, break it apart, and tell you why it matters. No hype. No summaries of summaries. Just the idea, explained.</em></p><div><hr></div><p>A researcher walks over with a sequence on her screen. She has a stretch of DNA, a promoter, the switch that turns a gene on, and one variant sitting inside it that she suspects changes how loudly that gene gets expressed. She wants to know if the variant matters. So she asks the question everyone asks now: which AI model should I use? BOLT-LMM or DNABERT?</p><p>It sounds like a ranking question. Pick the better of two tools. It is not a ranking question. <strong>It is a category error</strong>, and it's one I've watched smart teams make without noticing.</p><p>BOLT-LMM and DNABERT do not belong in the same column. One is a statistical method built to scan a whole population and find which genetic differences track with a trait across thousands of people. The other reads a single stretch of DNA and turns it into a numeric representation a computer can work with. Asking which is better for a promoter variant is like asking whether a microscope or a telephone is better for measuring temperature. The answer is neither, and the question itself is the problem.</p><blockquote><p><strong>The hardest mistake in genomic AI isn't picking the second-best model. It's picking the wrong kind of model and not knowing it.</strong></p></blockquote><div><hr></div><h2>The phrase that hides nine different things</h2><p><strong>"Genomic AI model" sounds like one category. It is at least nine.</strong></p><p>There are DNA language models, which read raw sequence and learn its grammar the way a text model learns the grammar of English. There are sequence-to-function predictors, which take a piece of DNA and predict what it does in a cell, how strongly it drives expression, whether a region is open or closed. There are variant-effect predictors, which score how much a specific mutation changes that function. There are statistical-genetics engines, which never look at sequence grammar at all and instead test, across a population, which genotypes associate with which traits. There are single-cell models that work on tables of gene activity per cell, generative models that design new sequences, and a few more.</p><p>Every one of these gets called a "genomic AI model." Every one of these shows up in benchmark papers ranked against the others. And every one of these answers a different question.</p><p>The researcher with the promoter variant needs a variant-effect predictor or a sequence-to-function model, something that reads her one sequence and tells her what the change does. BOLT-LMM, the statistical-genetics engine, is built for a completely different shape of problem: thousands of people, their genotypes, their traits, and the search for associations across the cohort. It has no way to read the grammar of her single promoter. DNABERT, the DNA language model, reads grammar beautifully but doesn't produce a calibrated answer to "does this variant matter" without more machinery bolted on top.</p><blockquote><p><strong>The two tools the researcher was choosing between were both wrong. The benchmark she'd consult to choose between them can't tell her that.</strong></p></blockquote><div><hr></div><h2>Why the leaderboard can't catch this</h2><p>Here is the uncomfortable part. <strong>The benchmarks the whole field relies on are structured to make this error invisible.</strong></p><p>A leaderboard ranks models within a class. It lines up five DNA language models and tells you which embeds sequence best. It lines up four variant-effect predictors and tells you which scores variants most accurately. That is useful, and the people building those benchmarks are doing careful work. But a leaderboard assumes you already chose the right column. It answers "which is best?" and stays silent on the question that comes before it: best at what, and is that even what I need?</p><p>The wrong-class error happens upstream of every benchmark. By the time you're reading a leaderboard, the category decision is already made, usually without anyone noticing a decision was made at all. You typed a model name into a search box, found a paper that ranked it highly, and never asked whether the thing being ranked was the right kind of thing for your task.</p><p>Nobody walks into this on purpose. The pull comes from how the field names things. Two models share the word "genomic," both have an impressive accuracy number in a recent paper, both have a HuggingFace page and a star count, and the surface presentation makes them look like rival products on the same shelf. A leaderboard reinforces the illusion by placing them in the same table. Nothing in that experience signals "these answer different questions." The signal you'd need is a layer that sits before the ranking and sorts tools by what they're for, and that layer mostly doesn't exist.</p><p>I'd put this error ahead of picking the second-best model in the right class, and it's far more costly. Pick the second-best variant-effect predictor and you lose a few points of accuracy on a task that was at least the right task. Pick a statistical-genetics engine for a single-sequence question and you don't lose accuracy, you get an answer that means nothing, dressed up to look like it means something. The numbers come back formatted, plotted, ready to drop into a slide, and there's nothing on the surface that tells you the foundation was wrong. Weeks of analysis can ride on top of a category mistake made in the first thirty seconds, and the failure surfaces late, if it surfaces at all.</p><div><hr></div><h2>Classify first, rank second</h2><p>The fix is almost embarrassingly simple to state and surprisingly hard to do without help: <strong>classify the computational object before you rank it.</strong></p><p>Before you ask which model is best, ask what kind of object your task actually needs. Do you have one sequence and want a functional readout? You need a sequence-to-function or variant-effect model. Do you have a cohort of people and want to find trait associations? You need a statistical-genetics engine. Do you have a table of cells and want to label cell types? A single-cell model. The class is determined by the shape of your data and the shape of your question, not by which model has the most citations this month.</p><p>Get the class right and the leaderboard becomes useful again, because now you're ranking within the right column. Get the class wrong and the leaderboard is worse than useless, because it gives you a confident ranking of tools that can't do your job.</p><p>The reason this is hard in practice is that the class isn't printed on the box. A model's documentation tells you its architecture and its scores; it rarely says, in plain terms, "this is for cohort association, not single-sequence scoring." You have to infer the class from the shape of the inputs it takes and the outputs it produces, and that inference is exactly the expertise a researcher new to a method doesn't have yet. The whole point of a triage layer is to do that inference for you and to fail loudly when the tool you named is the wrong kind of thing.</p><div><hr></div><h2>What this looks like as a tool</h2><p>I built a small open thing to make this concrete, because it's easier to argue for "classify first" when you can watch it fire.</p><p>It's called <a href="https://modelmap.vercel.app">ModelMap</a>, and it's a decision layer rather than a leaderboard. You describe your task in plain terms, what you have and what you want, and it triages the class first. When the tool you named doesn't match the class your task needs, it returns a "wrong-tool" verdict: here is the class your task needs, here is the class this tool actually belongs to, and here is why they don't match. Ask it whether DNABERT-2 is right for a population GWAS and it routes you to the statistical-genetics column and flags DNABERT-2 as the wrong class, with the real function of each spelled out.</p><p>Underneath are nine method classes and around thirty model cards, with deterministic rules handling the clear-cut category calls and a grounded language-model layer translating free-text questions into the controlled vocabulary. It's a research-use proof of concept, source-available, and you bring your own model key so there's nothing to abuse. I built it because I couldn't find anything that does. The closest existing tool, OmniGenBench, is a leaderboard, excellent at ranking within a class and silent on which class you need.</p><blockquote><p><strong>A leaderboard answers "which is best." The question that wrecks more projects is "best at what, and is that what I need."</strong></p></blockquote><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!U9rW!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74376ae2-fb40-4c02-88e3-c777eb3fb21e_1408x768.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!U9rW!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74376ae2-fb40-4c02-88e3-c777eb3fb21e_1408x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!U9rW!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74376ae2-fb40-4c02-88e3-c777eb3fb21e_1408x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!U9rW!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74376ae2-fb40-4c02-88e3-c777eb3fb21e_1408x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!U9rW!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74376ae2-fb40-4c02-88e3-c777eb3fb21e_1408x768.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!U9rW!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74376ae2-fb40-4c02-88e3-c777eb3fb21e_1408x768.jpeg" width="728" height="409.5" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/74376ae2-fb40-4c02-88e3-c777eb3fb21e_1408x768.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:false,&quot;imageSize&quot;:&quot;normal&quot;,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:728,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;captionedImage&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!U9rW!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74376ae2-fb40-4c02-88e3-c777eb3fb21e_1408x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!U9rW!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74376ae2-fb40-4c02-88e3-c777eb3fb21e_1408x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!U9rW!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74376ae2-fb40-4c02-88e3-c777eb3fb21e_1408x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!U9rW!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74376ae2-fb40-4c02-88e3-c777eb3fb21e_1408x768.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>ModelMap's whole point is that one verdict card. It's the moment it tells you the two things you were comparing were never comparable, and points you back one step to the decision you skipped.</p><p>If you build tools like this, the harder question is how you let a language model sit in front of that verdict without letting it invent a class or a license. I wrote that part up for the people who'd build it, over on AIXplore: <a href="https://ai.rundatarun.io/AI%20Systems%20%26%20Architecture/grounded-llm-triage-layer">Build an LLM Triage Layer That Can't Freelance</a>.</p><div><hr></div><h2>The question to carry</h2><p>You don't need ModelMap, and you don't need to remember the nine classes. <strong>You need one habit.</strong></p><p>The next time someone on your team asks which genomic AI model to use, or which large model, or which forecasting method, or which anything, resist the pull to answer the ranking question they asked. Ask the one underneath it first. What kind of object does this task actually need? Are we even in the right column?</p><p>Most of the time the ranking question is the easy part, and the field has built good tools for it. The category question is the one that decides whether the answer means anything, and almost nothing in the standard toolkit asks it for you. That's the gap worth closing, in genomics and well beyond it.</p><div><hr></div><p><em>Sunday Deep Dive is a weekly series on Run Data Run. Every Sunday I pick one paper, release, or technique worth understanding, break it apart, and tell you what it means for your work. Free every Sunday, no paywall. If it was useful, the easiest way to support it is to subscribe and forward it to one person on your team who'd want it. If it wasn't, tell me why. I'll make it better.</em></p>]]></content:encoded></item></channel></rss>