<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[Run Data Run: Around the Corner]]></title><description><![CDATA[Short reviews of ideas worth watching; opt-in, not part of the weekly email]]></description><link>https://rundatarun.io/s/around-the-corner</link><image><url>https://substackcdn.com/image/fetch/$s_!t_Ch!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffa36f5aa-74af-4492-b8d7-93b03f14a337_1280x1280.png</url><title>Run Data Run: Around the Corner</title><link>https://rundatarun.io/s/around-the-corner</link></image><generator>Substack</generator><lastBuildDate>Tue, 11 Aug 2026 04:44:57 GMT</lastBuildDate><atom:link href="https://rundatarun.io/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[Justin Johnson]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[rundatarun@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[rundatarun@substack.com]]></itunes:email><itunes:name><![CDATA[Justin Johnson]]></itunes:name></itunes:owner><itunes:author><![CDATA[Justin Johnson]]></itunes:author><googleplay:owner><![CDATA[rundatarun@substack.com]]></googleplay:owner><googleplay:email><![CDATA[rundatarun@substack.com]]></googleplay:email><googleplay:author><![CDATA[Justin Johnson]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[Nobody Saves Money on the Model]]></title><description><![CDATA[A team just swapped in a model that costs twice as much per token, and their bill went down. Here is why that is not a paradox, and what it means for anyone trying to make AI cheaper at scale.]]></description><link>https://rundatarun.io/p/nobody-saves-money-on-the-model</link><guid isPermaLink="false">https://rundatarun.io/p/nobody-saves-money-on-the-model</guid><dc:creator><![CDATA[Justin Johnson]]></dc:creator><pubDate>Tue, 21 Jul 2026 12:30:17 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/417116f3-2fb3-4906-b055-5124f68d1532_1376x768.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!cAWE!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F18713613-e64e-4119-802d-79f139aba297_1376x768.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!cAWE!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F18713613-e64e-4119-802d-79f139aba297_1376x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!cAWE!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F18713613-e64e-4119-802d-79f139aba297_1376x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!cAWE!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F18713613-e64e-4119-802d-79f139aba297_1376x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!cAWE!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F18713613-e64e-4119-802d-79f139aba297_1376x768.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!cAWE!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F18713613-e64e-4119-802d-79f139aba297_1376x768.jpeg" width="728" height="409.5" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/18713613-e64e-4119-802d-79f139aba297_1376x768.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:false,&quot;imageSize&quot;:&quot;normal&quot;,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:728,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;captionedImage&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!cAWE!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F18713613-e64e-4119-802d-79f139aba297_1376x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!cAWE!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F18713613-e64e-4119-802d-79f139aba297_1376x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!cAWE!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F18713613-e64e-4119-802d-79f139aba297_1376x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!cAWE!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F18713613-e64e-4119-802d-79f139aba297_1376x768.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Last week a team at <a href="https://cognition.com/blog/devin-fusion">Cognition</a>, the company behind the coding agent Devin, published a number that reads like a typo. They replaced Opus 4.8 with Fable 5. Fable 5 costs about twice as much per token. Their bill went down.</p><p>Not their quality. Their bill.</p><p>Here is the part of their table that matters:</p><p><strong>Fable 5, inside their new architecture: score 57.6, cost $3.00 per task.</strong></p><p><strong>Opus 4.8, on its own: score 48.8, cost $3.24 per task.</strong></p><p>The expensive model, wired up correctly, was better <em>and</em> cheaper than the cheap model on its own. Not a trade. Both columns at once.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!yuC2!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa12f42d2-cfed-4c78-a945-d040df9b1b90_1376x768.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!yuC2!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa12f42d2-cfed-4c78-a945-d040df9b1b90_1376x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!yuC2!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa12f42d2-cfed-4c78-a945-d040df9b1b90_1376x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!yuC2!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa12f42d2-cfed-4c78-a945-d040df9b1b90_1376x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!yuC2!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa12f42d2-cfed-4c78-a945-d040df9b1b90_1376x768.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!yuC2!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa12f42d2-cfed-4c78-a945-d040df9b1b90_1376x768.jpeg" width="728" height="409.5" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a12f42d2-cfed-4c78-a945-d040df9b1b90_1376x768.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:false,&quot;imageSize&quot;:&quot;normal&quot;,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:728,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;captionedImage&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!yuC2!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa12f42d2-cfed-4c78-a945-d040df9b1b90_1376x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!yuC2!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa12f42d2-cfed-4c78-a945-d040df9b1b90_1376x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!yuC2!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa12f42d2-cfed-4c78-a945-d040df9b1b90_1376x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!yuC2!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa12f42d2-cfed-4c78-a945-d040df9b1b90_1376x768.jpeg 1456w" sizes="100vw"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>If you have ever sat in a meeting where someone proposed saving money by dropping down a model tier, that result should stop you. It stopped me, because I had run the opposite experiment, and I had lost.</p><div><hr></div><h2>What they actually built</h2><p>Cognition split the work between two models instead of routing between them.</p><p>The expensive model is the driver. It plans, it interprets what the request actually means, it makes the judgment calls, and it does the final review. It takes very few actions itself and it reads only what it must.</p><p>The cheap model is the sidekick, and it is not a helper function. It is a full agent with its own tools that goes and does the mechanical work: fetching context, running the slow tests, carrying out the routine implementations once someone has decided what they are.</p><p>Both hold their own memory. Both run at the same time. And this is not a demo. Eighty-eight percent of Cognition's own internal code changes now run through that automated split.</p><p>The reason it works fits in one line. <strong>You pay for thinking once, and for doing it many times.</strong></p><div><hr></div><h2>The problem is that the opposite is also true</h2><p>Everyone in this field has met the other result. You move a workload to a cheaper model to save money, and it costs you more. The cheap model misunderstands the task, produces something confidently wrong, and a person spends an afternoon unpicking it. The invoice went down and the total cost went up.</p><p>That happens, and I am not going to argue with it, because I have the receipts.</p><p>So we have two findings that appear to be at war. Expensive models save money. Cheap models cost money. Both are observed, both are honest, and most of the advice you will read picks one and ignores the other.</p><div><hr></div><h2>I ran the losing experiment</h2><p>I built a cost router. The logic was the logic everyone reaches for: work out how hard the task is, send the easy ones to a cheap model, keep the expensive one for the hard ones.</p><p>It saved nothing. Not a little less than I hoped. Nothing.</p><p>I want to be precise about why, because the failure is more useful than a success would have been. <strong>I routed by task type. The axis that pays is ambiguity.</strong></p><p>Every individual call my router made was defensible. This one looks like a simple rename, send it to the cheap model. That one looks like architecture, keep it upstairs. And the total refused to move, because "simple rename" was a description of the <em>work</em>, not a description of <em>how much had already been decided</em>. Half the jobs I was labelling easy still had open questions inside them, and an open question handed to a cheap model is the single most expensive thing you can buy.</p><div><hr></div><h2>The law</h2><p>Once you see it, the war stops.</p><blockquote><p><strong>A cheap model is expensive on an open question and cheap on a closed one.</strong></p></blockquote><p>An open question is one where something still has to be decided. What are we actually building. What does done mean here. Which of these two readings of the request is the real one. Give that to a cheap model and it will not tell you the question is open. It will pick an answer, sound sure, and hand you something plausible that you now have to check line by line.</p><p>A closed question has had the deciding done. Rename this function in these three files, and the test that proves it is this one. There is no judgment left in the task, and a cheaper model does it for a fraction of the price with nothing at risk.</p><p>Which gives the driver model a job description nobody writes down. <strong>Its work is not "the hard parts." Its work is to turn open questions into closed ones.</strong> And that conversion has a name we already use for it. It is called a plan.</p><p>This also explains the finding that keeps embarrassing people who try to build a committee of cheap models and vote. A group at <a href="https://arxiv.org/abs/2502.00674">Princeton</a> tested that directly last year and found that running the single best model several times and combining its own answers beat mixing different models together, on every one of the thirteen mixed configurations they tried, using roughly half the forward passes. Their explanation is blunt: mixing models of different quality drags the average quality down. <strong>You cannot vote your way to judgment.</strong> Where the question is still open, quality dominates, and diversity is not the free lunch it looks like.</p><div><hr></div><h2>Cheap is not a property of the model</h2><p>Here is where I think the whole conversation is framed wrongly, including by the people getting the right answers.</p><p>Everyone calls this "big model, small model." My own setup says that is not the axis.</p><p>The models I push my mechanical work to are not small. They are large, capable models. They cost me nothing at the margin, because I bought them on flat monthly subscriptions instead of by the token. That is not a smaller brain. It is a different contract.</p><p>So there are two dials, and they are independent. <strong>Where does the marginal cost live, and where does the judgment live.</strong> A fine-tuned small model is cheap because it was narrowed. A large model on a flat rate is cheap because of how you bought it. Both are "the cheap one," and treating them as the same thing is how people end up sending an open question to a bargain and wondering why the quarter went sideways.</p><div><hr></div><h2>The sentence I had already written</h2><p>I found a note I made a few weeks ago, while building the thing that hands my grunt work off to those subscription models. I had written the law without recognising it:</p><blockquote><p><strong>A vague spec produces confident garbage, and the cheaper the model, the truer that is.</strong></p></blockquote><p>Read that again as a cost statement, because that is what it is. A cheap model is not a discount. It is a <strong>multiplier on the quality of your specification.</strong> Specify tightly and it multiplies your savings. Specify loosely and it multiplies your mess.</p><p>Which is why the discipline I run alongside it is not optional. The expensive model writes the entire work order: the task, the files, what done means, and the exact command that proves it. And when the cheap model comes back and says the job is finished, that claim is worth nothing. I read the difference it made, and I run the test myself. It has never once been the model's word that closed a task.</p><div><hr></div><h2>So the small-model story is not a cost play</h2><p>This is the part I keep chewing on, because it inverts something I believed.</p><p>Everyone is excited about training small models to do one narrow job extremely well, and the excitement is framed as a cost saving. It is not, or at least the saving is not where people are pointing.</p><p>If you can specify a task tightly enough to train a small model on it, you have <strong>proved the task is closed.</strong> All the ambiguity was removed by somebody, at some point, doing the expensive work of deciding. The training run is just collecting the winnings.</p><blockquote><p><strong>A small model is not a discount. It is a receipt for judgment already spent.</strong></p></blockquote><p>And the same is true of every cost reduction I have ever managed to make stick at scale. Every one of them was bought earlier, by an act of judgment that closed a question. The saving showed up in the invoice. It was created somewhere else entirely.</p><div><hr></div><h2>The boundary, in their words</h2><p>Buried in Cognition's own write-up, past the numbers, is the caveat I would have led with:</p><blockquote><p><strong>The sidekick fails when judgment is the deliverable.</strong></p></blockquote><p>Their example is a hard feature whose subtle intent got lost the moment it was handed down. The work came back correct and wrong at the same time.</p><p>Anyone who has run a team knows exactly where that line sits, and knows it is not about the seniority of the person you handed it to. You can delegate the work. You cannot delegate the judgment about what the work is for. The failure looks identical from the outside either way: something arrives, it is technically defensible, and it is not what the thing was for.</p><div><hr></div><h2>The honest caveat</h2><p>One thing about that Cognition result deserves saying out loud, because they said it themselves and nobody repeating the headline has.</p><p>Their best row was measured on Fable 5 during a stretch when the model was briefly pulled from sale, suspended in June under a US export-control order and restored at the start of July. Their own note adds two more caveats: those numbers were taken before the interruption, and that configuration was never tuned the way the others were. The model is back on sale now. The untuned-config caveat is not.</p><p>The pattern transfers. But the exact recipe is one vendor's best-case number on a setup they admit they never optimized, which makes it directional, not a benchmark. A result nobody has reproduced is a claim, not a finding, and that distinction is worth keeping close in a year when every week produces a new number.</p><div><hr></div><h2>What to do with this</h2><p>You do not save money by hiring cheaper people. You save money by putting your best judgment on the plan, and then making the execution mechanical enough that it does not need judgment. Every leader has run the cheaper-people experiment at some point. Most of us have the scar to show for it.</p><p>The unit of cost was never the hour. It was the outcome.</p><p>So the question to take into your next architecture review is not which model you are using, or what it costs per million tokens. Those are the numbers on the invoice, and the invoice is a lagging indicator of a decision somebody already made.</p><p><strong>Ask what you are paying per solved problem. Then ask who is doing the thinking.</strong></p><div><hr></div><p><em>Justin Johnson writes Run Data Run. His book on building with AI, Builder Leader, is at builder-leader.com.</em></p>]]></content:encoded></item><item><title><![CDATA[Workflows, Seven Weeks In]]></title><description><![CDATA[I called the economics of fan-out the day it shipped. Running it as a daily default since taught me the caveat I buried in a footnote is the actual problem.]]></description><link>https://rundatarun.io/p/workflows-seven-weeks-in</link><guid isPermaLink="false">https://rundatarun.io/p/workflows-seven-weeks-in</guid><dc:creator><![CDATA[Justin Johnson]]></dc:creator><pubDate>Mon, 20 Jul 2026 13:43:47 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/630b1b4e-c0d8-4f04-960f-307f64af8a9d_1664x928.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!D7J3!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F22b3a63f-8f6a-4b78-b63a-c82b4878e287_1664x928.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!D7J3!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F22b3a63f-8f6a-4b78-b63a-c82b4878e287_1664x928.jpeg 424w, https://substackcdn.com/image/fetch/$s_!D7J3!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F22b3a63f-8f6a-4b78-b63a-c82b4878e287_1664x928.jpeg 848w, https://substackcdn.com/image/fetch/$s_!D7J3!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F22b3a63f-8f6a-4b78-b63a-c82b4878e287_1664x928.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!D7J3!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F22b3a63f-8f6a-4b78-b63a-c82b4878e287_1664x928.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!D7J3!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F22b3a63f-8f6a-4b78-b63a-c82b4878e287_1664x928.jpeg" width="728" height="409.5" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/22b3a63f-8f6a-4b78-b63a-c82b4878e287_1664x928.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:false,&quot;imageSize&quot;:&quot;normal&quot;,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:728,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;captionedImage&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!D7J3!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F22b3a63f-8f6a-4b78-b63a-c82b4878e287_1664x928.jpeg 424w, https://substackcdn.com/image/fetch/$s_!D7J3!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F22b3a63f-8f6a-4b78-b63a-c82b4878e287_1664x928.jpeg 848w, https://substackcdn.com/image/fetch/$s_!D7J3!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F22b3a63f-8f6a-4b78-b63a-c82b4878e287_1664x928.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!D7J3!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F22b3a63f-8f6a-4b78-b63a-c82b4878e287_1664x928.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Seven weeks ago, the day Opus 4.8 and Dynamic Workflows shipped, I wrote that the unit of agentic work had moved from one model call to dozens of verified ones, and that the open question was no longer whether the model is good enough. It was who is verifying.</p><p>I have been running fan-out as a daily default since. The field taught me things the launch-day post could not.</p><div><hr></div><h2>The economics held</h2><p>The price math was the argument in May, and it played out the way the arithmetic said it would. Fifty parallel subagents on Fast at $10 in and $50 out stopped being a budget event. A review pass that spawns a skeptic per finding, a discovery sweep that runs five finders with different lenses and takes the union, a batch edit split across ten agents that each own a file. None of those feel like a stunt now. They feel like a Tuesday.</p><p>My own standing rule shifted to match. For anything that is a build, I hand the file-level work to subagents by default and keep the main context for the plan, the spec, and the diff review.</p><blockquote><p><strong>I do not decide to fan out. I decide not to, and only when there is a reason.</strong></p></blockquote><p>So the call on cost was right, and it was the easy call. The arithmetic was visible at launch.</p><div><hr></div><h2>The dispatch problem is still in my head</h2><p>The harder prediction was the one I was least sure of. I said the question of when a workflow is the right shape, and when a single careful pass is, was unsolved and lived in your head. Seven weeks later it still lives there.</p><p><code>ultracode</code>, the setting that makes Claude reach for a workflow on every task without being asked, still ships off. There is still no governance layer that says do not orchestrate this one. So the judgment of when to spend forty agents and when to spend one is mine, made fresh each task, and I get it wrong in both directions. I have fanned out a rename that a single pass would have finished cleaner and faster, and I have run one careful pass on a discovery job that wanted five blind finders and missed a third of the surface. Neither mistake announces itself. The over-orchestrated one costs more; the under-orchestrated one returns a confident, incomplete answer.</p><blockquote><p><strong>The tooling got cheaper. The taste did not.</strong></p></blockquote><div><hr></div><h2>The footnote became the fight</h2><p>The part I could not see in May is the one that has cost me the most.</p><p>I called the verifier the priced skill and warned the tooling for one was thin. What I did not know is that the harness puts a hard ceiling on the verifier's ability to do its job. Every subagent is capped at 8,000 output tokens per response, and the model's own thinking counts against that budget. Set effort to <code>xhigh</code>, which is the Claude Code default, and an adversarial verifier asked to reason hard about whether a finding is genuine will think its way past the ceiling and die before it emits a single word of verdict. The tool call comes back with zero output. No finding, no error, no crash. Silence that reads exactly like a clean pass.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!JMpf!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe520dfa9-46af-4bab-be00-69c3a40f29d7_1376x768.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!JMpf!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe520dfa9-46af-4bab-be00-69c3a40f29d7_1376x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!JMpf!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe520dfa9-46af-4bab-be00-69c3a40f29d7_1376x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!JMpf!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe520dfa9-46af-4bab-be00-69c3a40f29d7_1376x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!JMpf!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe520dfa9-46af-4bab-be00-69c3a40f29d7_1376x768.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!JMpf!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe520dfa9-46af-4bab-be00-69c3a40f29d7_1376x768.jpeg" width="728" height="409.5" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e520dfa9-46af-4bab-be00-69c3a40f29d7_1376x768.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:false,&quot;imageSize&quot;:&quot;normal&quot;,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:728,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;captionedImage&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!JMpf!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe520dfa9-46af-4bab-be00-69c3a40f29d7_1376x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!JMpf!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe520dfa9-46af-4bab-be00-69c3a40f29d7_1376x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!JMpf!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe520dfa9-46af-4bab-be00-69c3a40f29d7_1376x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!JMpf!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe520dfa9-46af-4bab-be00-69c3a40f29d7_1376x768.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>A verifier panel that looks like it ran and returns nothing is worse than no panel, because you trust the green. This is the failure I warned about in May, arriving through a door I did not know was there. The launch post said a fan-out that forwards fifty unverified findings is worse than the single pass it replaced.</p><blockquote><p><strong>A fan-out whose verifiers silently no-op forwards zero findings and tells you everything is fine.</strong></p></blockquote><p>The fix is mechanical once you know the ceiling exists. Bound what each agent writes, chunk large outputs across several small responses, and dial effort down for the producers so the verifiers have budget left to reason. But you have to know it is there. Most builders reaching for their first verifier panel do not, and the default settings hide it from them.</p><div><hr></div><h2>What I run now</h2><p>Effort is not one dial for the whole job. The producers, the agents writing code and editing files, run low, because their work is mechanical and the reasoning tax buys nothing. The judges and verifiers keep the high effort, and I split their task so each verdict fits inside one response instead of one giant pass that overflows. The verifier gets a real design, not a one-line prompt bolted onto the end of the workflow.</p><p>In May the verifier was a line item. Now it is the thing I spend design time on.</p><blockquote><p><strong>The verifier is the only part of the pipeline that fails without telling you.</strong></p></blockquote><div><hr></div><h2>The frame still holds</h2><p>The bridge between one model call and an agentic system is now a tool the model writes for itself, and the frameworks that were charging for that bridge have a lower ceiling on what they can charge. The vendor who ships the verifier primitive first still sets the pattern everyone copies.</p><p>I would add one clause to the closing line I wrote then. If the price of careful goes down, the price of casual goes up, and the default question stops being is the model good enough yet. It becomes who is verifying, and can your verifier afford to think.</p><div><hr></div><p><em>Around the Corner: short reviews of ideas worth watching. Opt-in section, not part of the weekly Run Data Run email. [Subscribe to the main list](https://rundatarun.io/subscribe) for longer essays.</em><a href="https://rundatarun.io/subscribe">Subscribe to the main list</a> for longer essays.*</p>]]></content:encoded></item><item><title><![CDATA[The skill that edits its own instructions]]></title><description><![CDATA[Self-editing skills, the ecosystem racing to build them, and the flywheel that makes any of it compound.]]></description><link>https://rundatarun.io/p/the-skill-that-edits-its-own-instructions</link><guid isPermaLink="false">https://rundatarun.io/p/the-skill-that-edits-its-own-instructions</guid><dc:creator><![CDATA[Justin Johnson]]></dc:creator><pubDate>Fri, 05 Jun 2026 01:39:26 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/f8c150f0-a43a-4b1f-b03a-d0c015174066_1376x768.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!WJFH!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb4375d91-9211-4aff-86f2-611dd25d7993_1376x768.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!WJFH!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb4375d91-9211-4aff-86f2-611dd25d7993_1376x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!WJFH!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb4375d91-9211-4aff-86f2-611dd25d7993_1376x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!WJFH!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb4375d91-9211-4aff-86f2-611dd25d7993_1376x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!WJFH!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb4375d91-9211-4aff-86f2-611dd25d7993_1376x768.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!WJFH!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb4375d91-9211-4aff-86f2-611dd25d7993_1376x768.jpeg" width="728" height="409.5" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/b4375d91-9211-4aff-86f2-611dd25d7993_1376x768.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:false,&quot;imageSize&quot;:&quot;normal&quot;,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:728,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;captionedImage&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!WJFH!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb4375d91-9211-4aff-86f2-611dd25d7993_1376x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!WJFH!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb4375d91-9211-4aff-86f2-611dd25d7993_1376x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!WJFH!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb4375d91-9211-4aff-86f2-611dd25d7993_1376x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!WJFH!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb4375d91-9211-4aff-86f2-611dd25d7993_1376x768.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>One of the routines I run a few times a day rewrote a line of its own instructions this week, and I let it.</p><p>The routine checks on a fleet of small agents I keep running, fixes what it knows how to fix, and flags the rest for me. The new part is the last step. Before it signs off, it reads back over the run and asks one question: what did this teach me that my instructions did not already cover? When the answer is real, it edits the checklist it works from, so the next run starts a little smarter.</p><p>It did not change the model. It changed its <em>skill</em>, the separate written set of instructions a tool loads when the task comes up. I thought this was a clever trick I had built. Then I went looking, and found half the field had built it too.</p><h2>Why skills are the unit</h2><p>A skill is less exotic than it sounds. It is a written set of instructions, plus any helper files, that an assistant pulls off the shelf when a task matches: a checklist for triaging your inbox, a runbook for closing the monthly books, the house style your team writes in. <a href="https://code.claude.com/docs/en/skills">Claude Code</a> keeps each one in its own folder and loads it only when it is relevant.</p><p>Here is why it reaches anyone who does not write code. The model is rented, and it gets swapped for a better one every few months. <strong>The skill is the part you own.</strong> It is where your specific know-how lives, the institutional memory that survives the upgrade. Most teams pour that knowledge into documents nobody opens twice. A skill is the same knowledge in a form the assistant actually uses, every time the work comes up. And it has crossed from a Claude Code feature to an <a href="https://www.agensi.io/learn/agent-skills-open-standard">open standard</a>: the same plain-text file now runs across Codex, Cursor, Copilot, and more than thirty other tools.</p><h2>My clever trick wasn't clever</h2><p>So I swept the last thirty days to see who else was doing this. The answer was humbling and useful: nearly everyone.</p><p>The pattern I thought I invented is written up as a recipe, a reflection step that runs after a skill is used, asks whether it helped, and proposes an edit to its own file. Anthropic's own skill builder does a sharper version, splitting your examples into train and test sets and keeping only the change that scores better on the held-out half. Microsoft's <a href="https://github.com/microsoft/SkillOpt">SkillOpt</a> tunes a skill's written instructions the way you would train a model, and its standout result is the one I care about: a skill tuned inside one tool kept almost all of its gain when they moved it to another. The know-how lived in the skill, not the software. It is already shipping inside products that crystallize skills as you work.</p><blockquote><p><strong>The thing I built to make my own setup compound turned out to be a pattern the whole field is converging on. My problem was not unique, which is exactly why the answer is worth keeping.</strong></p></blockquote><p>The scale is the real headline. One directory has now scraped <a href="https://skillsmp.com/">1.6 million of these skill files</a> off public GitHub, up from the roughly 790,000 a research team catalogued six months ago. Which is also the catch.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!CFBL!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F191d4fa2-79c1-4bb3-b91e-4898cc110173_1200x896.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!CFBL!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F191d4fa2-79c1-4bb3-b91e-4898cc110173_1200x896.jpeg 424w, https://substackcdn.com/image/fetch/$s_!CFBL!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F191d4fa2-79c1-4bb3-b91e-4898cc110173_1200x896.jpeg 848w, https://substackcdn.com/image/fetch/$s_!CFBL!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F191d4fa2-79c1-4bb3-b91e-4898cc110173_1200x896.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!CFBL!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F191d4fa2-79c1-4bb3-b91e-4898cc110173_1200x896.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!CFBL!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F191d4fa2-79c1-4bb3-b91e-4898cc110173_1200x896.jpeg" width="728" height="409.5" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/191d4fa2-79c1-4bb3-b91e-4898cc110173_1200x896.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:false,&quot;imageSize&quot;:&quot;normal&quot;,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:728,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;captionedImage&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!CFBL!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F191d4fa2-79c1-4bb3-b91e-4898cc110173_1200x896.jpeg 424w, https://substackcdn.com/image/fetch/$s_!CFBL!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F191d4fa2-79c1-4bb3-b91e-4898cc110173_1200x896.jpeg 848w, https://substackcdn.com/image/fetch/$s_!CFBL!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F191d4fa2-79c1-4bb3-b91e-4898cc110173_1200x896.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!CFBL!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F191d4fa2-79c1-4bb3-b91e-4898cc110173_1200x896.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>What's oversold</h2><p>Two things to hold back on, both visible in that same sweep.</p><p>The flood is real, and most of it is noise. The ecosystem went from empty to crowded in about six months, and the directories admit most skills barely trigger or quietly burn context. A skill that edits itself inside that flood does not automatically improve. It compounds whatever it already was, noise included, unless you gate the edits.</p><p>Write-once-run-everywhere is also sold harder than it ships. The open standard is real, but the formats are still converging, not interchangeable. A skill that sings in one tool can stumble in the next.</p><h2>Why this is worth watching</h2><p>I have <a href="https://rundatarun.io/p/delegation-not-automation-how-human">written before that the real skill is delegation, not automation</a>, knowing what to hand off and when to step back in. A skill that edits itself is the next turn of that screw, and the sweep handed me the guards that separate the durable version from the noise. <strong>Don't edit on a fluke:</strong> a problem has to recur before it earns a permanent change. <strong>Prove the new rule works:</strong> the added check has to catch the thing it was written for before it stays. <strong>Never re-add what you deleted.</strong> None of that is exotic. It is what you would want from a sharp junior teammate keeping the runbook current.</p><p>That loop is the real point, and it is bigger than skills. I write up something I think is clever, sweep the field to find the dozen people who already hit it, take the best of what they worked out, fold it back into my own setup, and share the result forward. The writing and the searching are not separate from the building. They are how the building compounds. This post is that loop turning once.</p><div><hr></div><p>The systems getting the press rewrite their own code against a scoreboard. The skills that will run your operations next year just keep better notes, in a file you can read, borrow the best ideas from everyone else, and get a little sharper every time they run.</p><p><em>Around the Corner: short reviews of ideas worth watching. Opt-in section, not part of the weekly Run Data Run email. [Subscribe to the main list](https://rundatarun.io/subscribe) for longer essays.</em><a href="https://rundatarun.io/subscribe">Subscribe to the main list</a> for longer essays.*</p>]]></content:encoded></item><item><title><![CDATA[Opus 4.8 and Workflows - One Careful Pass Is No Longer the Default]]></title><description><![CDATA[Anthropic shipped Opus 4.8 and Dynamic Workflows on the same day. Together they move the unit of agentic work from one model call to dozens of verified ones.]]></description><link>https://rundatarun.io/p/opus-48-and-workflows-one-careful</link><guid isPermaLink="false">https://rundatarun.io/p/opus-48-and-workflows-one-careful</guid><dc:creator><![CDATA[Justin Johnson]]></dc:creator><pubDate>Fri, 29 May 2026 09:14:52 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!So5X!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fad9345b0-786a-456b-b096-c368b2546540_1365x768.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!So5X!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fad9345b0-786a-456b-b096-c368b2546540_1365x768.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!So5X!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fad9345b0-786a-456b-b096-c368b2546540_1365x768.png 424w, https://substackcdn.com/image/fetch/$s_!So5X!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fad9345b0-786a-456b-b096-c368b2546540_1365x768.png 848w, https://substackcdn.com/image/fetch/$s_!So5X!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fad9345b0-786a-456b-b096-c368b2546540_1365x768.png 1272w, https://substackcdn.com/image/fetch/$s_!So5X!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fad9345b0-786a-456b-b096-c368b2546540_1365x768.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!So5X!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fad9345b0-786a-456b-b096-c368b2546540_1365x768.png" width="1365" height="768" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ad9345b0-786a-456b-b096-c368b2546540_1365x768.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:768,&quot;width&quot;:1365,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:190172,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://rundatarun.io/i/199713856?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fad9345b0-786a-456b-b096-c368b2546540_1365x768.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!So5X!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fad9345b0-786a-456b-b096-c368b2546540_1365x768.png 424w, https://substackcdn.com/image/fetch/$s_!So5X!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fad9345b0-786a-456b-b096-c368b2546540_1365x768.png 848w, https://substackcdn.com/image/fetch/$s_!So5X!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fad9345b0-786a-456b-b096-c368b2546540_1365x768.png 1272w, https://substackcdn.com/image/fetch/$s_!So5X!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fad9345b0-786a-456b-b096-c368b2546540_1365x768.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Anthropic shipped Opus 4.8 yesterday. The model bump alone is not the story.</p><p>The story is that two other things shipped beside it. A new Claude Code primitive called <a href="https://claude.com/blog/introducing-dynamic-workflows-in-claude-code">Dynamic Workflows</a>, which lets Claude write its own orchestration scripts and fan out to dozens of parallel subagents with adversarial verifiers built in. And a 3x cut to Opus Fast pricing, from $30/$150 per million tokens down to $10/$50, roughly 2.5x faster on top of it.</p><p>Those three changes are the same change. Anthropic just repositioned the unit of agentic work, and the pricing finally allows what the tooling implies.</p><h2><strong>What&#8217;s actually in 4.8</strong></h2><p>The model card is doing the usual benchmarks-go-up dance, but the prompting and runtime changes are where the day-to-day differences land.</p><p>The biggest is the move to <strong>adaptive reasoning only</strong>. <code>MAX_THINKING_TOKENS</code> is now ignored, and <code>CLAUDE_CODE_DISABLE_ADAPTIVE_THINKING</code> is gone. The model decides how hard to think. To pin behavior you use the new <code>/effort</code> slash command (or <code>--effort</code> flag) across <code>low / medium / high / xhigh / max</code>. Claude Code defaults to <code>xhigh</code> for coding work; claude.ai and Cowork default to <code>high</code>. For a one-off deep pass without raising the whole session&#8217;s effort, drop the literal word <code>ultrathink</code> into the prompt and that turn alone reasons harder.</p><p>Four new or refreshed slash commands round it out. <code>/ultrareview</code> runs a senior-engineer review pass over the diff. <code>/simplify</code> does a refinement pass on recently modified code. <code>/focus</code> hides intermediate work and shows only the final output. <code>/fewer-permission-prompts</code> scans the session and writes a safer allowlist into <code>settings.json</code> so the harness stops interrupting you on read-only bash and MCP calls.</p><p>The Max plan gets the 1M context window by default, and the <strong>Fast mode price cut</strong> to $10 in / $50 out per million tokens makes the difference between Fast and default Opus closer to a latency choice than a budget one. There is also a research-preview Auto mode behind Shift+Tab that auto-approves safe actions and pauses on risky ones, aimed at long-running tasks where you want to walk away.</p><p>None of that is revolutionary by itself. The change in posture is.</p><h2><strong>What Dynamic Workflows actually is</strong></h2><p>A Dynamic Workflow is a JavaScript orchestration script that the model writes for itself. The script does not call the model directly. It calls four primitives that the harness wires into the session: <code>agent()</code> spawns a subagent and returns its result, <code>parallel()</code> fans tasks out concurrently with a barrier, <code>pipeline()</code> streams items through multiple stages without barriers between them, and <code>phase()</code> groups subagent calls under a progress label.</p><p>Two activation modes ship with it. <strong>Explicit</strong> &#8212; you say &#8220;create a workflow to audit this codebase for X&#8221; and Claude designs and runs the script. <strong>Implicit</strong> &#8212; you flip the <code>ultracode</code> setting on and Claude evaluates every task as a workflow candidate, reaching for fan-out instead of single passes by default. Ultracode is off out of the box, and the docs are clear about why. It burns tokens fast.</p><p>The interesting part is what gets baked in. Schema validation through a structured-output tool means subagents return validated objects, not strings you have to parse. Workflows resume from a prior <code>runId</code>, so an edit to your script doesn&#8217;t re-run the agents that didn&#8217;t change. There is a 1,000-agent lifetime cap per workflow as a runaway-loop backstop. And the documented &#8220;quality patterns&#8221; &#8212; adversarial verify, judge panel, loop-until-dry, multi-modal sweep, completeness critic &#8212; show Anthropic&#8217;s own hand on what good fan-out looks like.</p><blockquote><p>The most telling pattern is the adversarial verify. Spawn three independent skeptics per finding, prompt each to refute it, kill the finding if a majority succeed. Anthropic is not selling fan-out as more answers. They are selling it as more checks.</p></blockquote><p>This is the part that changes how you build.</p><h2><strong>The price math is the whole story</strong></h2><p>The orchestration-first pattern has been technically possible for a year. LangGraph wired it up. CrewAI wired it up. The reason almost nobody runs it as a default is the bill. Fifty parallel subagents on Opus 4.7 Fast was $30 in and $150 out per million tokens. A serious review pass on a real pull request would burn through a tank of compute and produce something a senior engineer could have written by hand in less time.</p><p>Opus 4.8 Fast is $10 in and $50 out. The same fan-out is now a third of the cost and 2.5x faster.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!mkiP!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F144ef84f-b9b9-48d3-b330-9577da361782_1408x768.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!mkiP!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F144ef84f-b9b9-48d3-b330-9577da361782_1408x768.png 424w, https://substackcdn.com/image/fetch/$s_!mkiP!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F144ef84f-b9b9-48d3-b330-9577da361782_1408x768.png 848w, https://substackcdn.com/image/fetch/$s_!mkiP!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F144ef84f-b9b9-48d3-b330-9577da361782_1408x768.png 1272w, https://substackcdn.com/image/fetch/$s_!mkiP!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F144ef84f-b9b9-48d3-b330-9577da361782_1408x768.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!mkiP!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F144ef84f-b9b9-48d3-b330-9577da361782_1408x768.png" width="1408" height="768" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/144ef84f-b9b9-48d3-b330-9577da361782_1408x768.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:768,&quot;width&quot;:1408,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:481703,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://rundatarun.io/i/199713856?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F144ef84f-b9b9-48d3-b330-9577da361782_1408x768.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!mkiP!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F144ef84f-b9b9-48d3-b330-9577da361782_1408x768.png 424w, https://substackcdn.com/image/fetch/$s_!mkiP!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F144ef84f-b9b9-48d3-b330-9577da361782_1408x768.png 848w, https://substackcdn.com/image/fetch/$s_!mkiP!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F144ef84f-b9b9-48d3-b330-9577da361782_1408x768.png 1272w, https://substackcdn.com/image/fetch/$s_!mkiP!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F144ef84f-b9b9-48d3-b330-9577da361782_1408x768.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>That is the number that changes behavior. Iteration loops that used to be <em>run this once, pray it found the bug</em> become <em>run it three times with different angles and trust the intersection</em>. Discovery sweeps that used to ship as a single grep become five finders with different lenses, deduped at the union. Verifier panels that used to be cosplay become a default.</p><blockquote><p>Fast at $10/$50 makes the fifty-agent review pass look like a Tuesday, not a stunt.</p></blockquote><h2><strong>What this looks like in practice</strong></h2><p>I have been running a version of this pattern on my own homelab since February. Eight named agents probe their own state in parallel, a separate evaluator scores their last forty-eight hours of output against a scope file, safe idempotent fixes auto-apply, and only the human-decision items surface for me. The expensive part was never the orchestration. It was the cost of getting eight independent passes to cohere before I trusted the verifier.</p><p>Anthropic&#8217;s own <code>review-changes</code> example shows the same shape with the rough edges sanded off. Dimensions fan out: bugs, performance, security, reuse, tests. Each dimension yields findings. Each finding is handed to a panel of independent skeptics whose prompt is <em>try to refute this</em>. A finding survives only if a majority of skeptics fail to refute. It is the same trick a good engineering org runs at code review, ported into the model layer and budgeted in tokens instead of senior-engineer hours.</p><h2><strong>What&#8217;s oversold</strong></h2><p>Two honest caveats.</p><p>The new skill being priced is not the orchestration script. It is the verifier. A fan-out that finds fifty plausible bugs and forwards all fifty is worse than the single careful pass it replaced, because it shifts the verification burden onto a human who now has to triage noise instead of read code. The workflows post handwaves the verifier as a prompt. Most builders I know do not have a verifier discipline yet, and the tooling for one is thin.</p><p>And ultracode, the setting that makes Claude reach for a workflow on every task without being asked, ships off by default. Anthropic&#8217;s documentation flags why. Fan-out burns quota fast, and there is no governance layer that says <em>do not orchestrate this task</em>. The dispatch problem, deciding when a workflow is the right shape and when a single careful pass is, is unsolved. Right now it lives in your head.</p><h2><strong>Why this is worth watching anyway</strong></h2><p>Two reasons.</p><p>The agentic frameworks built between 2024 and now (LangGraph, CrewAI, Autogen) were filling the gap where the model vendor did not ship orchestration. That gap just closed. The bridge between <em>one model call</em> and <em>agentic system</em> is now a tool the model writes for itself, in JavaScript, with caching and resume baked in. Whatever those frameworks were going to charge for over the next eighteen months, the ceiling on that price just dropped.</p><p>And the orchestration-first pattern was already where AI engineering was heading. <a href="https://rundatarun.io/p/evals-are-the-new-bottleneck">Evals are the new bottleneck</a> because at some point the question stops being <em>is the model good enough</em> and starts being <em>how confident am I that this particular answer is right</em>. Workflows give you a vocabulary for spending compute on that second question instead of just the first. The vendor that ships the verifier primitive first sets the pattern everyone else copies.</p><div><hr></div><p>If the price of careful goes down, the price of casual goes up. The default question stops being <em>is the model good enough yet</em>. It becomes <em>who is verifying</em>.</p><div><hr></div><p><em>Around the Corner: short reviews of ideas worth watching. Opt-in section, not part of the weekly Run Data Run email. <a href="https://rundatarun.io/subscribe">Subscribe to the main list</a> for longer essays.</em></p>]]></content:encoded></item><item><title><![CDATA[The Number That Predicts When Your Agent Will Break]]></title><description><![CDATA[A new benchmark gives a name to the failure practitioners keep calling "complex reasoning."]]></description><link>https://rundatarun.io/p/the-number-that-predicts-when-your</link><guid isPermaLink="false">https://rundatarun.io/p/the-number-that-predicts-when-your</guid><dc:creator><![CDATA[Justin Johnson]]></dc:creator><pubDate>Fri, 22 May 2026 12:12:22 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!kIK2!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe184f0eb-b20d-46bc-80f5-172c90b35b41_1376x768.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!kIK2!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe184f0eb-b20d-46bc-80f5-172c90b35b41_1376x768.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!kIK2!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe184f0eb-b20d-46bc-80f5-172c90b35b41_1376x768.png 424w, https://substackcdn.com/image/fetch/$s_!kIK2!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe184f0eb-b20d-46bc-80f5-172c90b35b41_1376x768.png 848w, https://substackcdn.com/image/fetch/$s_!kIK2!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe184f0eb-b20d-46bc-80f5-172c90b35b41_1376x768.png 1272w, https://substackcdn.com/image/fetch/$s_!kIK2!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe184f0eb-b20d-46bc-80f5-172c90b35b41_1376x768.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!kIK2!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe184f0eb-b20d-46bc-80f5-172c90b35b41_1376x768.png" width="1376" height="768" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e184f0eb-b20d-46bc-80f5-172c90b35b41_1376x768.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:768,&quot;width&quot;:1376,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:527480,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://rundatarun.io/i/198834597?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe184f0eb-b20d-46bc-80f5-172c90b35b41_1376x768.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!kIK2!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe184f0eb-b20d-46bc-80f5-172c90b35b41_1376x768.png 424w, https://substackcdn.com/image/fetch/$s_!kIK2!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe184f0eb-b20d-46bc-80f5-172c90b35b41_1376x768.png 848w, https://substackcdn.com/image/fetch/$s_!kIK2!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe184f0eb-b20d-46bc-80f5-172c90b35b41_1376x768.png 1272w, https://substackcdn.com/image/fetch/$s_!kIK2!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe184f0eb-b20d-46bc-80f5-172c90b35b41_1376x768.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><p>A new paper asks a question that sounds simple and turns out to have teeth. When a frontier model fails at &#8220;reasoning,&#8221; what is it failing at?</p><p>Their answer is a number. They call it Relational Complexity, and it predicts model failure better than anything else they measured.</p><p>The paper, <a href="https://arxiv.org/abs/2604.12176">&#8220;Evaluating Relational Reasoning in LLMs with REL&#8221;</a> from Fesser, Ektefaie, Fang, Kakade, and Zitnik, borrows a construct from cognitive science. Relational Complexity (RC) is the minimum number of entities a system has to hold in mind and bind together at once to take a single reasoning step. &#8220;A is taller than B&#8221; is RC=2. &#8220;A is between B and C&#8221; is RC=3. The number climbs as the relations get wider.</p><p>The finding is clean and a little grim. As RC goes up, accuracy falls off a cliff, and nothing the authors tried pulled it back.</p><h2>What they measured</h2><p>The clever part is the benchmark itself. REL is a generative framework, not a fixed test set. It produces as many problems as you want at any RC level, across three domains the authors deliberately picked to look nothing alike: pattern-completion puzzles (Raven&#8217;s matrices), phylogenetic trees in biology, and molecular isomers in chemistry.</p><p>Why three unrelated domains? Because that lets them hold everything else constant. Same vocabulary, same input length, same task format, only the RC dial moving. Most reasoning benchmarks can&#8217;t separate &#8220;the task is harder&#8221; from &#8220;the task has more words&#8221; or &#8220;the task is in an unfamiliar domain.&#8221; REL can.</p><blockquote><p>Hold vocabulary, length, and format fixed, move only the relational complexity, and watch the accuracy curve bend. That is the whole experiment, and it is enough.</p></blockquote><p>They ran it against Claude Opus 4.5, Gemini 3 Pro, and GPT 5.2.</p><h2>The numbers</h2><p>At low complexity, the models look great. The pattern puzzles at RC=1-2 land around <strong>91% accuracy</strong> across all three models.</p><p>Then it collapses. Scale the matrices up, push RC to 6, and Claude and Gemini drop to roughly <strong>12%</strong>. The biology task tells the same story: phylogenetic homoplasy detection runs at 35% with four taxa and falls to <strong>1% at twenty-five taxa</strong>.</p><p>The authors then did the thing most benchmark papers skip. They ran a regression to check whether RC was actually the driver or just correlated with something else. With collinearity controls in place, <strong>RC explained 24 to 44% of the explainable variance</strong>. The next-strongest factor topped out at 17%.</p><blockquote><p><strong>It is not input length. It is not domain. It is the binding count</strong></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!w2Jg!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd8b5acff-d460-4cbd-88aa-bb28065276c3_1879x1100.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!w2Jg!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd8b5acff-d460-4cbd-88aa-bb28065276c3_1879x1100.png 424w, https://substackcdn.com/image/fetch/$s_!w2Jg!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd8b5acff-d460-4cbd-88aa-bb28065276c3_1879x1100.png 848w, https://substackcdn.com/image/fetch/$s_!w2Jg!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd8b5acff-d460-4cbd-88aa-bb28065276c3_1879x1100.png 1272w, https://substackcdn.com/image/fetch/$s_!w2Jg!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd8b5acff-d460-4cbd-88aa-bb28065276c3_1879x1100.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!w2Jg!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd8b5acff-d460-4cbd-88aa-bb28065276c3_1879x1100.png" width="1456" height="852" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d8b5acff-d460-4cbd-88aa-bb28065276c3_1879x1100.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:852,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:195013,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://rundatarun.io/i/198834597?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd8b5acff-d460-4cbd-88aa-bb28065276c3_1879x1100.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!w2Jg!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd8b5acff-d460-4cbd-88aa-bb28065276c3_1879x1100.png 424w, https://substackcdn.com/image/fetch/$s_!w2Jg!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd8b5acff-d460-4cbd-88aa-bb28065276c3_1879x1100.png 848w, https://substackcdn.com/image/fetch/$s_!w2Jg!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd8b5acff-d460-4cbd-88aa-bb28065276c3_1879x1100.png 1272w, https://substackcdn.com/image/fetch/$s_!w2Jg!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd8b5acff-d460-4cbd-88aa-bb28065276c3_1879x1100.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><strong>.</strong></p><p></p></blockquote><p>And the interventions did almost nothing. Extra test-time compute bought <strong>2 to 3%</strong>. In-context examples bought <strong>3 to 6%</strong>. Tool use, handing the chemistry model RDKit so it could compute instead of reason, produced a mean recall of <strong>0.094</strong> that got <em>worse</em> as the problem grew.</p><blockquote><p><strong>The gap is structural. You don&#8217;t prompt your way out of it.</strong></p></blockquote><h2>Why an agent builder should care</h2><p>I have written before that <a href="https://rundatarun.io/p/evals-are-the-new-bottleneck">evals are the new bottleneck</a> and that <a href="https://rundatarun.io/p/the-agent-archaeology-checklist-8">agent failures cluster around a small set of repeatable mistakes</a>. RC is the missing vocabulary for one whole class of those failures.</p><p>Think about what your agent does when it stalls on something that &#8220;should&#8221; be easy. A cross-document join where it has to reconcile three sources at once. A planning task with four interacting constraints. A loop where it has to hold the output of step two while reasoning about step five. Those are not long tasks or unfamiliar tasks. <strong>They are high-RC tasks.</strong> The model has to bind several interdependent things simultaneously, and that is the regime where frontier accuracy falls to a coin flip or worse.</p><blockquote><p>When a task needs three or more interdependent variables held in mind at the same time, the failure is not a smarter-model problem. It is a binding problem, and more compute does not fix it.</p></blockquote><p>This reframes the diagnostic. The next time an agent breaks on a task you expected it to handle, the useful question is not &#8220;is the model good enough yet.&#8221; It is &#8220;how many things does this step force the model to bind at once.&#8221; If the answer is four or more, you have your explanation, and the fix is architectural, decompose the binding into smaller steps with explicit intermediate state, rather than waiting for a better model.</p><h2>What I&#8217;d hold back on</h2><p>Two honest caveats, both the authors more or less own.</p><p>The tasks are stylized. Raven&#8217;s matrices and phylogenetic trees are clean lab instruments, and the jump from &#8220;RC in a synthetic tree&#8221; to &#8220;RC in your production workflow&#8221; is assumed, not proven. I would love to see RC mapped onto a naturalistic agent benchmark before treating the number as a planning constant.</p><p>And there is no human baseline. Every result frames frontier models as failing, but without a human RC-versus-accuracy curve we cannot tell whether people plateau at RC=5 or sail past it. That would settle whether this is &#8220;LLMs are uniquely bad at binding&#8221; or &#8220;binding is hard for everyone and LLMs are a bit worse.&#8221; Different stories, different implications.</p><div><hr></div><p>The contribution here is not the scary 12%. It is the ruler.</p><blockquote><p>For two years &#8220;complex reasoning&#8221; has been the phrase practitioners reach for when a model fails and they cannot say why. RC turns that shrug into a measurement.</p></blockquote><p>The generative code is <a href="https://github.com/ada-f/relational_reasoning">open on GitHub</a>, so you can instantiate REL-style probes against your own agent tasks instead of guessing.</p><p>Watch for the human baseline and the naturalistic mapping. If those land, RC stops being a benchmark curiosity and becomes a number you check before you ship an agent into a high-binding workflow.</p><p>The models are not getting dumber. We are just learning to name the shape of where they break.</p><div><hr></div><p><em>Around the Corner: short reviews of ideas worth watching. Opt-in section, not part of the weekly Run Data Run email. <a href="https://rundatarun.io/subscribe">Subscribe to the main list</a> for longer essays.</em></p>]]></content:encoded></item><item><title><![CDATA[Neural Networks Don't Think in Straight Lines]]></title><description><![CDATA[Goodfire's new work suggests most of our tools for understanding AI are pointed at the wrong shape.]]></description><link>https://rundatarun.io/p/neural-networks-dont-think-in-straight</link><guid isPermaLink="false">https://rundatarun.io/p/neural-networks-dont-think-in-straight</guid><dc:creator><![CDATA[Justin Johnson]]></dc:creator><pubDate>Tue, 12 May 2026 14:45:36 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!eehw!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbeba1ee6-a60d-43f2-b84a-d1ab50983e9c_1376x768.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!eehw!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbeba1ee6-a60d-43f2-b84a-d1ab50983e9c_1376x768.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!eehw!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbeba1ee6-a60d-43f2-b84a-d1ab50983e9c_1376x768.png 424w, https://substackcdn.com/image/fetch/$s_!eehw!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbeba1ee6-a60d-43f2-b84a-d1ab50983e9c_1376x768.png 848w, https://substackcdn.com/image/fetch/$s_!eehw!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbeba1ee6-a60d-43f2-b84a-d1ab50983e9c_1376x768.png 1272w, https://substackcdn.com/image/fetch/$s_!eehw!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbeba1ee6-a60d-43f2-b84a-d1ab50983e9c_1376x768.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!eehw!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbeba1ee6-a60d-43f2-b84a-d1ab50983e9c_1376x768.png" width="1376" height="768" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/beba1ee6-a60d-43f2-b84a-d1ab50983e9c_1376x768.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:768,&quot;width&quot;:1376,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:446462,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://rundatarun.io/i/197360370?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbeba1ee6-a60d-43f2-b84a-d1ab50983e9c_1376x768.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!eehw!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbeba1ee6-a60d-43f2-b84a-d1ab50983e9c_1376x768.png 424w, https://substackcdn.com/image/fetch/$s_!eehw!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbeba1ee6-a60d-43f2-b84a-d1ab50983e9c_1376x768.png 848w, https://substackcdn.com/image/fetch/$s_!eehw!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbeba1ee6-a60d-43f2-b84a-d1ab50983e9c_1376x768.png 1272w, https://substackcdn.com/image/fetch/$s_!eehw!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbeba1ee6-a60d-43f2-b84a-d1ab50983e9c_1376x768.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>A research team at <a href="https://www.goodfire.ai/">Goodfire</a> trained a tiny neural network to drive a virtual car up a hill. Then they looked inside the network to see how it represented the car&#8217;s position.</p><p>The answer was not where anyone expected.</p><p>Position didn&#8217;t live as a clean direction in the network&#8217;s internal space. It lived as a <strong>curve</strong>, threaded through the network&#8217;s neurons like a string. Every point on the string corresponded to a real-world position of the car.</p><p>When the team nudged the network along that curve, the car moved coherently. When they nudged it in a straight line across the curve, the way almost every modern interpretability tool does, the predictions broke. The car teleported. The simulation produced nonsense. The straight line wandered through regions of the network&#8217;s space that the model had never learned to handle.</p><p>Their new paper, <a href="https://www.goodfire.ai/research/the-world-inside-neural-networks">&#8220;The World Inside Neural Networks&#8221;</a>, argues this isn&#8217;t a quirk of one toy model. It&#8217;s how networks actually represent things.</p><h2><strong>The shape of the problem</strong></h2><p>Most of what we do to understand or steer large AI models <strong>assumes representations are straight</strong>. There&#8217;s a name for this assumption in the field, the <a href="https://transformer-circuits.pub/2024/scaling-monosemanticity/">linear representation hypothesis</a>: concepts inside a model live as directions in the network&#8217;s internal space, and you can adjust the model&#8217;s behavior by moving along those directions.</p><p>You see this assumption everywhere. It&#8217;s how Anthropic built <a href="https://www.anthropic.com/research/golden-gate-claude">&#8220;Golden Gate Claude&#8221;</a>, the version of its model that couldn&#8217;t stop talking about a bridge. It&#8217;s how researchers find &#8220;refusal directions&#8221; and &#8220;honesty vectors.&#8221; It&#8217;s how <a href="https://transformer-circuits.pub/2024/scaling-monosemanticity/">sparse autoencoders</a> (SAEs), the dominant tool for naming what&#8217;s inside a model, try to break activity into a clean list of concepts.</p><p>Add. Subtract. All of it assumes flat geometry.</p><blockquote><p>If the real structure is curved, every straight-line move is just an approximation along a tangent, and the further you push, the worse the approximation gets.</p></blockquote><p>That would explain a lot of unsolved noise in the field. Why steering tricks work in narrow zones and fall apart at the edges. Why killing one &#8220;feature&#8221; inside a model often breaks something unrelated. Why so many &#8220;we found the X concept&#8221; papers don&#8217;t reproduce cleanly when somebody else tries.</p><p>The field has been working with the wrong shape and getting partial credit for the effort.</p><h2><strong>What they actually showed</strong></h2><p>The Mountain Car experiment is the centerpiece. It&#8217;s small, but the intervention proved the geometry: walk along the curve, the model behaves; cut across it, the model breaks. That&#8217;s the difference between geometry as decoration and geometry as cause.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!W1rZ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0fcdadb1-c807-4657-9f2d-5662119e6938_1200x896.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!W1rZ!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0fcdadb1-c807-4657-9f2d-5662119e6938_1200x896.png 424w, https://substackcdn.com/image/fetch/$s_!W1rZ!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0fcdadb1-c807-4657-9f2d-5662119e6938_1200x896.png 848w, https://substackcdn.com/image/fetch/$s_!W1rZ!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0fcdadb1-c807-4657-9f2d-5662119e6938_1200x896.png 1272w, https://substackcdn.com/image/fetch/$s_!W1rZ!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0fcdadb1-c807-4657-9f2d-5662119e6938_1200x896.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!W1rZ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0fcdadb1-c807-4657-9f2d-5662119e6938_1200x896.png" width="1200" height="896" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/0fcdadb1-c807-4657-9f2d-5662119e6938_1200x896.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:896,&quot;width&quot;:1200,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:768503,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://rundatarun.io/i/197360370?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0fcdadb1-c807-4657-9f2d-5662119e6938_1200x896.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!W1rZ!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0fcdadb1-c807-4657-9f2d-5662119e6938_1200x896.png 424w, https://substackcdn.com/image/fetch/$s_!W1rZ!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0fcdadb1-c807-4657-9f2d-5662119e6938_1200x896.png 848w, https://substackcdn.com/image/fetch/$s_!W1rZ!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0fcdadb1-c807-4657-9f2d-5662119e6938_1200x896.png 1272w, https://substackcdn.com/image/fetch/$s_!W1rZ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0fcdadb1-c807-4657-9f2d-5662119e6938_1200x896.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The same lens shows up in their other work. Months and days form circles inside language models. Colors organize by hue and brightness. I <a href="https://rundatarun.io/p/the-specialist-is-now-you">walked through one of Goodfire&#8217;s biology pipelines</a> a few weeks back, where the same techniques surface features in a DNA model that look like splice sites and regulatory regions. The curved-geometry view is becoming their signature.</p><p>The harder claim, and the more important one, is what they say about sparse autoencoders. SAEs are the bet Anthropic, OpenAI, and DeepMind have all made on how to read large models. Goodfire argues SAEs <strong>break continuous structure into disconnected pieces</strong>. Their example: words ending in &#8220;-ore&#8221; form one smooth curve in the model&#8217;s internal space, and SAEs shatter that curve into a handful of unrelated features. The unity disappears.</p><blockquote><p>If that critique holds for big models, a meaningful slice of current AI safety research is studying artifacts of its own tools, not the model.</p></blockquote><h2><strong>What&#8217;s oversold</strong></h2><p>The framing, &#8220;the world inside neural networks,&#8221; does more work than the evidence supports. The paper smuggles in a big claim, that models contain a faithful map of reality, which is hard to disprove because nobody knows what would count against it.</p><p>What Goodfire actually showed is narrower and more useful. <strong>Representations are curved. The curves are causal. Tools that assume straightness are leaving capability on the table.</strong> That&#8217;s enough. The cosmic framing is marketing.</p><p>Two real gaps the paper doesn&#8217;t address:</p><ul><li><p><strong>Does it scale?</strong> They show the geometry is causal for one toy model. Does the same picture hold for a 70-billion-parameter language model? Open question.</p></li><li><p><strong>Is it the same picture across models?</strong> If different models trained on the same data find the same curves, the geometry is approximating something real about the world. If not, the curves are model artifacts and the philosophy crumbles.</p></li></ul><p>Both questions are answerable. Neither is in the paper.</p><h2><strong>Why this is worth watching anyway</strong></h2><p>Two reasons.</p><p>One, it reframes the tooling debate. The interpretability community has been arguing about which kind of feature dictionary to build. Goodfire is asking whether a dictionary is even the right object. A map of curves wants different math, different methods, different papers.</p><p>Two, the <strong>parallel with biology is getting hard to dismiss</strong>. <a href="https://www.quantamagazine.org/the-brain-maps-out-ideas-and-memories-like-spaces-20190114/">Grid cells, place cells, and head-direction cells</a> in mammalian brains encode space as exactly the kind of curved structure Goodfire is finding inside artificial networks. That work won <a href="https://www.nobelprize.org/prizes/medicine/2014/press-release/">the 2014 Nobel Prize in Physiology</a>. When evolved biology and trained silicon land on the same shape, the convergence is worth taking seriously.</p><div><hr></div><p>A year ago I wrote that <a href="https://rundatarun.io/p/the-race-to-understand-ais-black">interpretability was the race we couldn&#8217;t afford to lose</a>. Goodfire&#8217;s work is what running that race looks like when it goes well.</p><p>This is an important direction, half-marketed, and the next year of interpretability research will tell us whether the curved-geometry view replaces the feature-dictionary view or merges with it.</p><p>Watch the scaling question. Watch whether somebody bigger than Goodfire bets on this lens.</p><p>If they&#8217;re right, a lot of recent activation-steering work is about to age badly.</p><div><hr></div><p><em>Around the Corner: short reviews of ideas worth watching. Opt-in section, not part of the weekly Run Data Run email. <a href="https://rundatarun.io/subscribe">Subscribe to the main list</a> for longer essays.</em></p>]]></content:encoded></item></channel></rss>