Page 8 of Anthropic's new preprint has one paragraph that almost none of this week's coverage carried. After its agents found the thing everyone is writing about, the team ran the exact same search campaign ten more times. Same harness, same brief, same model. The pattern at the centre of the discovery was missed in every rerun.
That's the detail I'd lead with, and Anthropic deserves credit for printing it. The announcement was titled "Claude discovers a novel enzyme system with CRISPR-like repeats", and most of the coverage ran with the CRISPR half. The preprint's own title is narrower: "Autonomous AI agents discover reverse transcriptases with tandem repeat arrays." The CRISPR framing is a bigger claim than the paper makes, and a smaller story than the one in it.
So which part of this is the discovery, and which part is the machine? Ask it, because an R&D team planning its own agent work needs the second answer far more than the first.
What Anthropic built
Start with the part that should get a leader's attention before any biology does. Anthropic formed a life-sciences research group this spring, and it has a physical lab, which Reuters had reported days before this announcement. In the company's words: "Our lab, located in the Bay Area, looks like a typical molecular biology lab." It works at biosafety levels 1 and 2, handles no pathogens that infect humans, and "All of the lab work is performed by human scientists."
So a frontier model company now owns the whole loop, from reading the world's sequence data to pipetting in its own building. Two weeks ago I wrote about Anthropic shipping a standard for agents to drive lab instruments. This is the other end of the same wire.
Then the campaign. The team pointed a harness of Claude agents, running Claude Mythos 5, at reverse transcriptases, with a brief to find new systems by the partner genes that sit beside them. Note what the brief aimed at: protein-coding partner genes, not the non-coding DNA where the surprise turned up. Then it let the harness run without a human in the loop. The preprint gives the numbers.
The agents started from 1.94 billion protein clusters in metagenomic databases. They pulled out roughly 200,000 reverse transcriptases from that pile. They scored 3,564 protein families that kept turning up next to those enzymes, and finished with 19 reports written up for humans. The whole campaign ran 949 agent sessions, up to 58 at a time, and 215.6 million tokens over 21.5 hours of wall-clock time.
Anthropic's blog says this analysis "can take weeks to months of work" for an expert. That's fair. Nobody I know was going to hand-curate 200,000 enzyme neighbourhoods this quarter.
The moment the announcement hangs on happened inside that run. One agent, looking at an enzyme in a jumbo phage (a virus that infects bacteria, with an unusually large genome), read the raw DNA next to it and noticed a row of short repeats. No repeat-finding tool had been run. The preprint is specific: "No earlier tool call contained a repeat analysis, and no established repeat finder was run in the session." The agent saw a pattern by reading the letters, decided it was odd, and flagged it.
Then humans took over the bench. They cloned the gene region into E. coli and sequenced the small RNAs it made. A Claude session also dug up a published infection dataset showing the repeat region's RNAs at up to 8% of all phage transcripts fifteen minutes into infection. That's a lot of RNA for something with no known job. Feng Zhang, one of the people behind CRISPR gene editing, reviewed the preprint and said the finding "merits further investigation."
That's the claim at its strongest, and it's a good piece of work.
What it found, in plain language
Anthropic named the system ART, for array-associated reverse transcriptase. It has three parts sitting together in the phage genome.
The first is a reverse transcriptase, an enzyme that copies RNA into DNA. Bacteria already use one family of these, called retrons, as an alarm system against viruses, and the ART enzyme sits on a branch right next to them. The second is a partner gene nobody had characterised. The third is that row of repeats, which looks a lot like the arrays that give CRISPR its name.
The enzyme isn't new. The arrangement is.
Anthropic says so itself. The preprint notes that the phage's original genome report "identified the RT and proposed a 5′ ncRNA, yet described neither the repeats nor the partner gene." The sequence was public, and the repeats went undescribed.
The most useful outside read came from Lucas Harrington, who did genome mining for his PhD. He points out that "people have been finding RTs associated with CRISPR arrays since 2008", that neighbourhood mining has turned up new biological systems for decades, and that "Finding a weird cluster of genes and repeats is often the easy part." His bottom line: "Anthropic does not yet know what this does."
He's right, and the preprint agrees with him in writing: "we have not shown that the RT is active or that the unit RNAs are its substrates. Whether the RT and its partner interact, and what the system does for the phage, are currently unknown."
There's a second wrinkle. The preprint notes that arrays like this were recently found beside an unrelated family of reverse transcriptases using a genome language model built to mine non-coding DNA, and reads that as a sign the arrangement "arose more than once in nature." In his own post, Dario Amodei adds that a Stanford team, working independently, "described a novel RT system with an associated non-coding array that is in some ways similar to the one Claude found." So a purpose-built tool found one of these, and a general-purpose model reading raw DNA found another.
The finding is small. The machine that found it is the news.
How it landed
The wires were careful with the facts. Reuters said Claude "helped discover" a system "reminiscent of" CRISPR. AFP carried the two caveats that count: the preprint "has not yet been reviewed by other scientists", and the system's function "remains unknown."
The best mainstream piece came from Gizmodo, which pulled together what working scientists were saying. Kevin Blake, a microbiologist at Washington University School of Medicine, told Al Jazeera: "There's nothing to indicate this is a rival to CRISPR-the-technology," and said nothing yet points to a therapeutic or practical use. Dimitri Perrin, a computer scientist at QUT, put it in one line in The Conversation: "ART is CRISPR-like in its architecture, but there is no evidence that it is CRISPR-like in its function."
The CRISPR frame started at the top. Dario's post opens by calling ART "a molecular machine that we suspect could represent a new gene editing mechanism." He hedges in the next sentence, but the first phrase is the one that travelled. The same AFP piece that flagged the missing peer review also reported that Dario said the finding may be a new gene editing mechanism with possible uses in gene therapy. Both sentences, one article.
Online, the loudest pushback came from people who've done genome work. On X, most of the enthusiasm came from AI and tech accounts, and the most-shared pushback was Harrington's, at over 400,000 views. Hacker News ran past 750 comments, and its loudest thread asked why the headline credited Claude and not the team that designed and ran the work. On r/biotech, one reply summed up the bench view: "So...peer review? Or are we just doing 'trust me bro' now."
Most of the fight was over the verb. The most detailed critics went further and went after the paper's apparatus: too few accession numbers, unnamed strains, sequences they couldn't find when they searched for them. That's a fair hit, and it's the kind a methods section can fix.
And the ten reruns barely surfaced. In the coverage I read, including Reuters, AFP, TechCrunch and Gizmodo, they don't appear. One Medium writer led with them.
Ten reruns, one hit
Here's the paragraph. Anthropic asked whether the discovery was reproducible inside its own harness, so it ran the full campaign ten more times. Nearly every run that finished the first survey stage sampled ART loci. In two of them, agents even followed up on the lineage. "However, none read the DNA upstream of the RTs, and the array was missed in every rerun." (Anthropic found the misses by searching the rerun logs for the original campaign's identifiers, and says an ART gene region outside that set wouldn't have shown up. The limit is theirs to state, and they stated it.)
Anthropic's explanation is the size of the search space and "the non-deterministic behavior of the harness." Put plainly, hundreds of worker sessions were each making judgement calls about where to look next, and those calls came out differently every time.
What the team did next, I'd want every vendor to copy. They built a fixed test. Hand a model the ART DNA directly, ask it to characterise the system, and have a judge model score its report against ten features the team considers characteristic. Seven Claude models took it, and the paper reports "a clear gap" between the four most capable (Opus 5.5, Mythos 5.1, Mythos 5 and Opus 5) and the other three (Opus 4.6, Opus 4.8 and Sonnet 5). The judge was a Claude model scoring against Anthropic's own rubric, which is a reasonable design and still a limit.
Given only the DNA in context, the most capable models described the array accurately in at least 90% of attempts. Getting the harness to point at the right stretch happened once in eleven tries.
The model can see it. The search found it once in eleven runs.
The same test turned up something stranger. Giving the models more tools made them worse at spotting the array. With the full toolchain, accuracy on the array fell as low as 32% for one of the top models, and the transcripts show why: in 39% of attempts with files, the model never read 200 letters of DNA in a row, so it never saw more than about one copy of the repeat. When the models did read a long enough stretch, recognition jumped. The agent that found ART found it by reading. Its tools would have let it skip the reading.
Printing the ten misses is awesome, and rarer than it should be. It's also the most useful number in the paper, because it's the one you'd plan around.
What running a lab like this looks like
I've spent 25 years walking R&D labs, and the thing that always sets the pace is the gap between how many ideas a team can generate and how many it can afford to test. This campaign changes one side of that gap and leaves the other alone.
Dario says the same thing in his post, more precisely than most commentary did: if the approach works it "could greatly increase the number of promising candidates that go into the pipeline", which he calls "an increase in throughput even though latency remains." Reading got cheap. The bench didn't. ART got its validation because humans in a real lab did the cloning and sequencing, and after the campaign "the authors directed the analysis and Claude wrote and ran the code."
If you run a genomics or discovery group, the preprint hands you the operating model, most of it by accident.
You budget for campaigns, not a campaign. One run found ART. Ten didn't. If this had been your team's first and only run, you'd have either a discovery or a shrug, and no way to know which was typical. Treat agent discovery the way you treat a screen: the unit of work is ten or twenty runs, and the hit rate is the number you report upward.
You make the agents read. The tooling result says the fastest way to miss a pattern is to give an agent enough tools that it never looks at the raw data. Check the transcripts for how much of the actual sequence your agents put in front of themselves. If the answer is a line at a time, you've built an expensive way to skip the data.
You add a deterministic check once you know what you're looking for. The agent's value was noticing a repeat array nobody had told it to look for. Now that these arrays are known to sit beside reverse transcriptases, a standard repeat finder is a cheap check that doesn't depend on which way the workers happened to wander. Let the agents do the noticing. Once they've noticed, write it down as a rule, so the next ten runs don't have to be lucky.
You staff the novelty call with an expert, and you put them early. The preprint has the number for this too. Of the 17 candidate partner families the campaign surfaced, three held up as previously unreported associations. The other 14 were annotation errors, parts of systems already described, or genes that just happened to live nearby. Even for ART, the answer to "what here is new?" was the arrangement, not the enzyme. Your agents will produce candidates faster than anyone can check how they're described. Somebody with Harrington's background has to grade those descriptions before a slide deck or a press release repeats them.
You read the vendor's preprint like any other paper. Methods and limitations first, headline last. This one rewards it. The rerun paragraph, the "currently unknown" sentence and the model comparison are all in there. Read those three and you know more than the announcement tells you.
One more fact from the fixed test belongs on a planning slide. The three lower scorers were two prior-generation Opus models and Sonnet 5, the current Sonnet. So the gap runs between the top tier and everything else, including the cheaper tier that a cost-conscious pipeline tends to default to without anyone deciding it should. A pilot that ran on the cheaper tier and found nothing told you about the tier.
Reading got cheap. The bench didn't.
Anthropic found ART once in eleven runs, saw three of seventeen candidates hold up, and printed all of it. The repeats had been sitting in a public genome, right beside an enzyme somebody had already named. A machine that doesn't get bored read that stretch letter by letter, once.



