This is a two-minute film about a genetic arrangement an AI agent spotted in a virus's DNA that the genome's own authors had never described. I started the conversation a little before nine on the morning of October 1, then went and did my actual job. The ninth version rendered at 4:23 that afternoon. I didn't write a frame of it.
The calendar says a day, because I was doing my actual job in between. By my rough count it was a couple of hours of computing time, about an hour of mine, and around $50 in tokens at API prices. That covered the finished phage film plus the storyboard, narration and first paintings for a second one. I pay a flat subscription, so the bill I actually saw was zero.
The science comes from a post I wrote last week. Anthropic pointed 949 agent sessions at public genome data. One of them, reading the raw letters of a jumbo phage, noticed a short sequence repeating beside a reverse transcriptase, an enzyme that copies RNA back into DNA. The phage's original genome report had described neither the repeats nor the gene beside them. When the team ran the same search ten more times, every rerun missed it.
I wanted that story told visually, for people who don't read preprints. So I asked Opus to make it into a film. It also made a second one, which I'll get to, because the second film turned out to be about the thing I spent my day doing.
What's actually inside one of these films
A film made this way is a software project with a soundtrack.
There's no video editor in the loop.
It starts as a storyboard: a table with one spoken line per row, and a description of what the viewer sees while that line plays. Opus writes it from the source material and a research pass on how good explainers handle the same subject. For the phage film that meant Drew Berry's molecular animations and the way 3Blue1Brown keeps one object on screen while the camera moves around it.
Then the voice goes first. A small text-to-speech model called Kokoro runs on my laptop, reads each line, and records the exact moment every word lands. Those timings drive everything after. When the narrator says "reverse transcriptase", an arrow draws itself on that word, because the animation is keyed to the timestamp.
Each scene is a small web page, animated on a single timeline and rendered one frame at a time. The renderer blends eight sub-frames into each frame for motion blur, then adds a little film grain. That's most of the difference between "website" and "film".
Code can't paint a virus, though. For the parts that need a real picture, Gemini generates painted backgrounds, which film people call plates. An open video model on my DGX turns a still plate into five seconds of slow motion, three and a half minutes per clip, for free.
Last comes a critic. A separate agent, told nothing about what was built or why, watches the render. It gets contact sheets, frames at every transition and a transcript of the audio, and writes up what a person would notice.
Opus builds the scenes in parallel, one agent per scene. The parent agent keeps the storyboard and the final cut.
That's a 67-second explainer of the pipeline, made by the pipeline, from the day before. It was the first one I approved. My entire review was "awesome!", which tells you how low the bar was before we tried for spectacular.
The skill rewrites itself between renders
The whole pipeline is a Claude Code skill: a set of instructions Opus loads when I say "make a video", plus a ledger of every ruling I've made and every trap it has fallen into. The ledger is the part that makes it fast. The skill, ledger included, is public in Claudelicious, my open cookbook for running Claude Code as a system: the skills, hooks, memory and learning loop behind everything in this post.
When I give a note, Opus does two things in the same session: it fixes the frame, and it writes the note into the skill as a dated rule. The next render already follows it. So does every film after it.
Nine versions fit into one afternoon because no mistake had to be made twice.
Here's what that looked like on the phage film.
The first draft was a slide deck that moved. Typed letters on navy, ticking counters, a corner label like a news channel. My note: "pretty basic." Within the hour the skill had a new first step: paint two looks as stills, I pick one, then build. The second film's look took me about thirty seconds.
A number landed in the wrong organism. The 8% figure came from a Staph phage, and a scene had drawn it inside E. coli. New rule: check the brief against the primary source before storyboarding.
A counter invented data. It ticked through 2 of 200 and 7 of 200 on its way to a value the source only reports at the end. New rule: print only the sourced number.
The ending just stopped. My note was "the ending is weak." Version nine lights the 14 real repeat rows, flares them on the last line and holds while the music resolves. New rule: every ending lands on its last word.
None of those were hard. Each needed somebody to look at the frame and say what was wrong, once.
The critic catches a lot, and it's still a guess
The critic ran three rounds on the phage film. It found real problems I'd have missed on a laptop speaker: a music bed so quiet it was effectively missing, diagrams shrunk to a sliver of the frame, the same caption style reused eight times until it read like a template. It flagged that the film's central number, one hit in eleven runs, never actually appeared on screen. Fixed by round two.
Round two also told a scene agent to make the reruns visibly miss the gene. The agent did exactly that, and round three flagged that the film now contradicted the science: the reruns reached the same genes and passed them by without reading the stretch of DNA beside them. Following the critic had produced a cleaner-looking film that was wrong.
So every critic note now gets checked against the fact sheet before it goes to a builder.
A critic's fix is a claim about the world, and I check it like any other finding an agent hands me.
The second film is the same lesson
The second film ties three posts together. The lab post is the first: the agents surfaced 17 candidate partner gene families, and only three held up as previously unreported. Telling those apart took experts, and then a lab bench. The second is OpenAI measuring 3.1 days of agent work for every day of human work, with over half of the long successful tasks still needing a person to step in. The third is about the record agents leave behind, which only becomes useful once someone has defined what a good result is.
It's painted like a watercolor this time: a night sky of star trails, a gold-panner's tray, an office tower lit floor by floor, a sealed envelope. Code draws everything that moves on top of the paintings.
That film's argument is that agents can do the work now, and the scarce part is someone who decided what good looks like before the work started. I spent a whole day living it.
Opus wrote the storyboards, generated the paintings, built the scenes, timed the voice, ran the critic and assembled the cuts. I read source material it had already summarized, picked a look from two options, and watched renders. The day's real decisions were mine and they fit in a few sentences: this is basic, this ending is weak, I like these backdrops, two paintings with a voice over them isn't an explainer.
That last one landed on day two. The first cut of the second film was two lovely paintings, a slow camera move and a narrator.
Every frame passed every check the pipeline has, and it was boring.
No critic would have flagged it, because nothing in it was broken. It just wasn't what I wanted, and only I knew that.
That's the bottleneck from all three posts, at the size of one person and one film. Three things helped.
Decide the look before the first frame. Picking between two painted directions took me thirty seconds and saved hours of building toward the wrong world. In a lab or a data team the equivalent is writing down what a good result looks like before the agents run, the way a clinical trial protocol names its endpoint before the first patient enrolls.
Give the critic the facts, not just the output. A critic that can only see the film can tell you it looks wrong. A critic with the source facts can tell you it's wrong.
Put the rule where the next run reads it, the same day. For me that's the skill's ledger. For a team it's the eval suite, the prompt or the checklist the agents run against, updated before the next run rather than in next month's retro.
I asked for a second opinion on this ending
The first version of this post ended on a count: 31 rulings, 42 traps, and a promise to publish the next number. Before calling it done, I ran it past a second-opinion skill I built. It hands a draft to two models from other labs, DeepSeek and GLM, and has them argue with my own take.
Both said the ending was bookkeeping. DeepSeek went further and pointed out that I'd written a rule into the video skill saying every ending lands on its last word, then ended my own post by just stopping. Fair. A little rude, but fair.
What they pitched was better than what I had. GLM said to keep the count, but as a starting line instead of a finish line. DeepSeek said to aim for better notes: fewer about pronunciation and invented counters, more about the story. Which is awesome, because that's the version of this I actually want.
Here's how I see it going. Every note I give teaches the skill a piece of the mechanics, and the mechanics stop needing me. What's left is the part I like: which story deserves two minutes of someone's attention, and what it should feel like to watch. Nobody else can give that note for me, and I wouldn't want them to.
What I'd make once my notes are only about the story is probably my next post. For now the ledger says 31 rulings, and the third film starts from all of them.



