Somebody else can hear the best idea in your own sentence before you do.
Episode 100 of Data Science Leaders, Domino Data Lab's show, is out on YouTube. Thirty minutes with Thomas Been, Domino's CMO. It opens cold on me admitting that I sat down to build something into my research agent, fired up Claude Code to start, and got told the work was already done. She'd shipped it in February. I found it in May.
They called the episode "Becoming the Builder-Conductor." I never said that word.
I said "conductor" once, in a clause, in the middle of an answer about something else. Three weeks later it was the title. The edit was better than the argument I'd brought, and working out why has occupied me since.
Why I said yes
The invitation came through the writing. Thomas had read Start With Claude Code and brought it up on a Domino team call before anyone contacted me. Somebody read the thing and wanted to argue with it. That's a reason to get on a call.
Guest-dominant, no interruptions, no gotcha. Thomas talks maybe a fifth of the time. Long answers get to finish. No slides, no self-introduction, one topic, thirty minutes.
He does the introduction himself, and he introduced me as having a degree in marine biology. It's molecular biology. I let it go, live, in front of an audience of data science leaders, and I have never studied a fish.
Which changes what preparing means. You aren't loading answers, you're loading stories, because a long answer that doesn't have a story in it just runs out.
I prepared nearly nine thousand words for a thirty-minute conversation. The best idea in the episode is a word that isn't in any of them.
The word I didn't write
What I came to argue is what I've been arguing all year. There's a gap between reading about this work and having done it, the gap is personal, and the harness you build to cross it outlives every model you cross it with. That's The Harness Is the Moat, and it's most of *Builder Leader*.
Describing where most technical leaders are sitting right now, I said they're using agents like a chatbot with a longer wait, and that the alternative is "much more of an orchestra."
One clause. Gone in two seconds. I moved straight on to the next point.
Domino built the episode around it, and the reason it works is a tension I hadn't noticed I was carrying. A builder has his hands on the thing. That's the whole book: cross the gap yourself, because you can't read your way across it. A conductor is the only person on that stage who makes no sound, and is held responsible for all of it. Nobody has ever applauded a conductor for playing well.
So which is the promotion?
A builder has his hands on the thing. A conductor never touches an instrument and owns the sound anyway.
I don't resolve that here, because I didn't resolve it there. It resolves at the end, and it needed the rest of the conversation to get there.
What happened while I slept
The metaphor is cheap without something underneath it. So:
I wanted to understand federated learning. Building is how I learn, so I opened a session to start building. Before I wrote a line, ARIA told me she'd already done it.
She'd scored the idea 9.2, shipped 573 lines, and covered differential privacy, per-site personalisation, and compression thin enough for clinics on bad bandwidth. All the parts I'd been looking forward to working out for myself. February, overnight, unattended. I found it three months later. I wrote that one up at the time in She Already Built It, and it remains the most humbling twenty minutes I've had with a computer.
She scored it 9.2, shipped 573 lines overnight, and waited three months for me to notice.
Pool ideas, score them, run the cheap version first, analyse the result, then critique it with a different model so she can't grade her own homework, and self-heal when something breaks. No human anywhere inside that cycle. I set direction and guardrails and sit on top of it.
She's moved house since I last wrote about her. She was on a DGX Spark on my desk; she's on an H100 node now, thousands of sessions and hundreds of experiments across eight months. The ARIA paper covers eighteen weeks of her: 19,364 commits, seven instances, four scientific domains. I still cannot make that number feel normal.
The work I care about most is the retinal imaging. Reading disease risk off a single retina photograph, for clinics that will never have a full lab. As I said to Thomas, the eye is "the one place that you can see the blood vessels and the nerves without cutting people open." Same kernel as everything else she does, pointed somewhere new. Not rewritten.
The question nobody prepped
Ten questions were worked in advance. The one that produced the best answer wasn't among them.
Thomas asked about ARIA's relationship to time. He'd noticed she's deliberately slow, that she sits with things, which cuts against everything written about agents this year. It appears nowhere in the nine thousand words.
What came out, and I didn't know it was the answer until I heard myself say it:
"She's wrong a lot, and she's really good at recognising when she's wrong, and that's more important than going off and getting things right a lot of the time. To give me an idea of where not to look is worth just as much to me."
That's the conductor argument, and I got to it by accident.
A conductor doesn't generate the notes. A conductor decides which of what the orchestra produces is the take. Generation got cheap. Selection didn't. A system that reliably knows when it's wrong has done half your selecting before you sit down.
The limit sits in the same breath, because it's the part people skip: knowing you're wrong is not the same as being right, and a model marking its own work is worth nothing. That's precisely why the critique stage runs on a different model from the one that did the work. Take that apart and the whole loop degrades into a machine that agrees with itself, confidently, forever.
Andrej Karpathy ran the same shape on a single box. A 630-line training script left running unattended, roughly seven hundred changes explored on its own across two days, about twenty of which stuck. I wrote about what decides whether a loop like that pays off or burns you while you sleep in The Loop Is Simpler Than It Sounds. The model is the commodity. The readable loop around it is the asset.
The delay was never the science
This pattern is older than ARIA.
The TCGA consortium spent five or six years sequencing and assembling human mutation data. We reprocessed the entire corpus in about 22 days, over the run-up to Christmas, with a novel algorithm. It surfaced roughly 17% novel variants that mapped straight back into a live drug portfolio.
Five years to a month is not a science result. It's an arithmetic result about everything that isn't science.
The delay is never the science. It's the machinery around it.
And none of that is a biology story. Jack Hanlon, who leads GenAI Media at Meta, said it sharper than I ever have: "The first time you see an engineer build something in 45 minutes that would have taken a week a year ago, but then see it not ship for another 6 weeks, you will be radicalized." Meta is a consumer company. Put the same build inside a patient-facing one and you add privacy impact assessment, security review, AI governance review, architecture review, procurement, legal redline and regulatory sign-off, each gate serial, each gate three weeks at the floor. Six of them in a row is half a year spent on a thing that was finished in the spring.
You can't read your way across
This was the part Domino pulled for their own feed, and they cut it in the wrong place.
Most technical leaders are still on the copilot side. Prompt, review, approve, repeat. Plenty of them believe they've moved past that, and are running agents like a chatbot with a longer wait. The gap between those two isn't tooling. It's whether anyone has personally operated the thing.
Senior sponsorship exists. Technical enthusiasm exists. The person in the middle who has done it with their own hands usually doesn't.
The quote Domino ran was "the crossing is personal because somebody's got to do it." True, and it's half a sentence. The half they cut is the half a leader needs:
The crossing is personal, because somebody's got to do it, and then somebody's got to teach somebody else how to do it.
That's how we did biology. It's how we did everything else. Nobody has explained why this one is different.
You can't read your way across, and you can't buy your way across by outsourcing it. Both routes feel like progress and neither produces anyone who can teach the next person.
The question nobody's asking
Thomas closed by asking what nobody is asking, and this is the one I'd want a team to take away.
AI has to eat the process around AI. I called it an ouroboros on the call, the snake that consumes itself, and I'll stand behind the image. We have pointed AI at the typing with real enthusiasm. We have not pointed it at the approvals, the reviews, the intake queues, or the committee that decides whether a model may be used at all.
In a regulated industry that process exists to protect patients, and it isn't a villain. But holding a technology you cannot deploy harms patients too, and that harm doesn't show up in anyone's risk register. A process that blocks the thing which would have helped them has stopped protecting patients and started protecting itself. Both of those are true at once, and pretending only the first one is true is how this argument usually gets lost.
So, concretely, four things:
Name the queue that's actually binding you. Not the model, not the budget. The queue. Almost nobody can name theirs, which is itself the finding.
Point the agents at the queue, not only at the code. The review packet, the evidence pack, the governance artefacts, the intake triage. That work is textual, repetitive, and enormously expensive in human weeks. It is exactly the shape of thing these systems are good at, and almost nobody is aiming there.
Stop buying point solutions and calling it a strategy. I said on the call that "we'll just buy Claude Cowork" is a band-aid, and the same goes for whatever the next one is called. Rebuilding how a company runs was never something you could buy.
Do the crossing yourself, then teach exactly one person. That's the whole adoption mechanism. It doesn't scale another way, and every org chart that assumes otherwise is buying software instead of capability.
The thing I'm starting on next is whether an entire company can run on agents. Finance, engineering, all of it. I'm sketching it now, and Thomas asked me back to report on how it went, which converts an idea into a deadline.
Which is where the conductor question finally resolves, and the answer is neither promotion nor demotion. The instrument got much bigger and the job moved. My own arc has gone from heavy steering of these systems to trying to stay out of their way so I'm not slowing them down, and I mean that as a description of work rather than of leisure. Thomas put it better than I did: it's like being on a boat, waking up to see where it went overnight, reading the weather, and setting the line.
Somebody still has to decide which of it is the take. That part hasn't moved at all.
I get to spend my mornings reading what a machine decided to try while I was asleep and then choosing what it means. It's awesome, and I've stopped looking for a graver word for it.
We made AI eat the keyboard. Now it has to eat the queue.
That last line is a post I already wrote, AI Ate the Keyboard. Now It Has to Eat the Queue, and I'm reusing it on purpose. It's the same argument, and I hadn't yet worked out that the queue is where the conductor actually stands.
Watch it
Episode 100 of *Data Science Leaders*, thirty minutes. Thomas gets more out of me on the biology-to-code route than I've written down anywhere, and there's a stretch on how ARIA handles time that isn't in this piece at all, because it deserves its own.
Thanks to Thomas Been and the Domino team, who ran the best-prepared interview I've done and then found a better title for it than I had.
The sources
*Data Science Leaders* episode 100, "Becoming the Builder-Conductor" - Domino Data Lab, released 2026-08-21, 30:16.
ARIA: Sustained Autonomous Research Agents in Biomedicine - Johnson and Bedworth, 2026. Zenodo. The 18-week deployment, 19,364 commits, seven instances, four domains.
The Harness Is the Moat - the argument I brought to the conversation.
Start With Claude Code - the post that produced the invitation.
She Already Built It - the federated learning discovery, written up at the time.
Inside ARIA: Teaching a Machine to Think Like a Scientist - how the loop is built.
The Loop Is Simpler Than It Sounds - the overnight loop pattern, and the three questions that decide whether it pays off.
The Year Nature Caught Up - the publishing-lag argument behind the TCGA section.
AI Ate the Keyboard. Now It Has to Eat the Queue - the process argument, and the closing line.
Three Harnesses, Three Characters, One Working Week - what running several of these at once looks like.
*Builder Leader* - the book the argument in this piece comes out of.
Run Data Run is free, no paywall. If it was useful, the easiest way to support it is to subscribe and forward it to one person on your team who'd want it. If it wasn't, tell me why. I'll make it better.



