Agents do real work now. In LangChain's survey of more than 1,300 people building them, 57% had an agent in production. They draft the monthly report, answer the first round of customer email, reconcile invoices, and summarize the literature before the Monday meeting.
Most of those jobs end the same way. A person reads what the agent produced and decides what to do with it. A reviewer opens a summary drafted overnight, fixes one sentence that overstates a finding, and approves the rest.
That fix is the most useful thing to come out of the whole exchange, and in most companies it goes nowhere. The document gets better. The agent doesn't, and tomorrow it makes the same mistake for someone else.
There's a growing argument that those corrections, and everything else agents leave behind, could end up worth more than the work itself. I think the argument is right. I also think most companies can't use any of it yet, and drug development has a good way of explaining why.
No trial starts without a primary endpoint. Before the first patient enrolls, the protocol says what success means and how it will be measured.
Most agents went into production without one.
The exhaust
Satya Nadella gave the byproduct a name this summer. In an essay in July he wrote: "Models learn from 'exhaust,' the prompts people write, the tools agents use, and especially the corrections people make when the model is wrong."
The best example I have is the post you're reading.
It started with a tweet and a hunch. Agents went off and researched it. I read what came back and argued with half of it. More research. An outline. A draft that I told the agent read like a review journal article, so it got rewritten. Somewhere in the middle I decided I disliked the metaphor the whole piece had been built around, and it went too. Then another round of research, another rewrite, more guidance from me.
Every one of those steps left a record of a decision: what to keep, what to cut, which source was worth trusting, which argument held up. That's exhaust. It's a far more detailed record of how I think a piece through than anything I could write down if you asked me to.
So where do those decisions go?
Mostly, nowhere. The reason I dropped that metaphor lives in a chat window, and next month an agent may well hand me the same metaphor again.
Saving everything doesn't fix it either. My own agents send their model calls through a gateway that logged every one of them for five months. When I finally went to learn from that log, it held 469,930 rows, and the two columns meant to hold what was asked and what came back were empty in every single one.
I didn't have a record of decisions. I had plumbing.
MIT's NANDA group studied why most enterprise AI projects show no measurable return and ended up in the same place. Their report says: "The core barrier to scaling is not infrastructure, regulation, or talent. It is learning." Most of the systems they looked at "do not retain feedback."
The pitch for AI exhaust is that you stop throwing those corrections away. Capture them, learn from them, and your agents get better every week. Over time your company builds up something a competitor can't buy: a record of how your people make decisions. Nadella's worry was that this record leaks to whoever supplies your models, and that you should keep it.
In January I wrote about the context graph, the idea that agent work leaves behind a searchable record of how decisions get made. This is the harder half of that idea. A pile of records doesn't teach anything on its own. Three things have to be true first. The agent has to be doing a job you understand. You have to know what doing it well looks like. And you have to be able to connect what the agent did to what happened next.
Start with the job
Most agents never leave behind anything worth learning from, because they never get out of the pilot. The model is rarely the reason.
Every business process exists twice. There's the version in the procedure document, and there's the version people run. The document says invoices over a threshold need two approvals. The person who has done the job for eight years knows which three suppliers always send duplicates, checks those first, and waves through the rest.
An agent meets the second version on its first day. If nobody has written that version down, nobody can tell the agent about it, and nobody can say afterwards whether it did the job. So the pilot produces some good outputs and some strange ones, nobody can explain the difference, and the project stalls in a steering committee.
McKinsey tested 25 things companies do with AI to see which ones showed up in the bottom line. Redesigning the workflow had the biggest effect of any of them. Only 21% of companies using AI had fundamentally redesigned even some of their workflows.
Decide what good looks like
Once you know the job, you can say what doing it well means. In AI that written-down definition is called an eval, and I argued in May that evals are the new bottleneck. It's a set of checks that run every time the agent works. The summary cites a source for every number. The support reply resolves the issue without the customer writing back. The contract clause matches our fallback position. Some of those checks run as code, some need a person, and some use another model as the grader.
This is where coding agents got lucky, and why they took off first. Code grades itself. The tests pass or they fail. The developer accepts the suggestion or deletes it. Cursor feeds those accepts and rejects back into its autocomplete model and rolls out a new version every 1.5 to 2 hours. That's awesome, and it's only possible because every suggestion arrives with a free verdict.
A claims decision, a site-selection memo or a blog post comes with no verdict attached. Someone has to build one. That missing grader is most of the distance between coding agents and everything else a company does.
Most companies haven't built it yet. The same LangChain survey found 89% of respondents monitor what their agents are doing, and only about half evaluate whether the output was any good.
Watching tells you what happened. Grading tells you whether it should have.
You also won't get the definition right up front. The first time a team sits down to grade fifty agent outputs, they argue. One reviewer marks a summary down for dropping a caveat. Another thinks the caveat was noise. A third says the summary is fine for the internal meeting and wrong for the regulator. Shreya Shankar and her colleagues watched people grade AI output and named this criteria drift: "users need criteria to grade outputs, but grading outputs helps users define criteria."
I didn't know "reads like a review journal article" was one of my failure modes until I read one with my name on it. Now I know to look for it.
That argument over examples is the business process getting written down for the first time, by the people who run it. Understanding the process and writing the eval turn out to be the same piece of work.
Capture both halves
Now the exhaust starts to mean something, provided you capture two things rather than one.
The first is the process record: what the agent did and where a person stepped in. The second is the result: whether the thing the work was for actually happened. The customer didn't write back. The filing was accepted. The site enrolled on time.
Each half is only half a finding. The process record tells you what happened and not whether it was good. The result tells you it went well and not why. Joined together, they show you which step made the difference. A support reply the reviewer loved, followed by the same customer writing in again the next morning, tells you something neither record could tell you alone.
Drug development already runs a version of this. Big cardiovascular trials use a committee of clinicians to rule on every suspected event. Was it really a heart attack, a stroke, a bleed? Each ruling is made against a definition the trial wrote down before it started, so every one of them is a decision with the grading built in.
A team publishing in Circulation trained a model on one trial's committee rulings, then moved it to a second trial with a different drug, different patients and different definitions, adapting it with 20 of that trial's suspected events per endpoint. Among 13,885 suspected primary-endpoint events in the second trial, GPT-4o reading the records directly got 76.3% right. The purpose-built model got 86.4%. With humans taking the 30% of cases it was least sure about, the combination reached 95.6%.
The study is retrospective, and it doesn't separate how much of the gap came from those 20 examples and how much from everything else about the system. What it does show is a general model losing to a much smaller record of how one particular group of people rules, when every ruling came with its definition attached.
The hard part is that results come back at different speeds. Some arrive in seconds: the reviewer accepted the draft, edited it, or threw it out. Some take days: rework, escalations, how long the case stayed open. The ones you care about most can take months: the renewal, the regulator's response, the study readout.
So in practice you learn from the fast signals and keep checking them against the slow ones. Most of the time they point the same way. When they don't, when the drafts reviewers love are the ones that keep coming back from the regulator, the agent is getting better at the wrong thing, and only the joined record will show you.
This is also where Nadella's warning gets sharper. A pile of ungraded traces shows a competitor what you did.
A graded pile shows them what works.
Where the learning lands
Put it together and you get a loop. The reviewer's correction from this morning becomes a check the agent runs tomorrow, and next quarter's outcomes tell you whether that check was right.
Companies are already running it. Airbnb's customer-support team built one into its support AI, with the human agents' feedback on each suggested reply flowing straight into the next version. Retraining went from a months-long cycle to a matter of weeks.
Airbnb's loop feeds its lessons back into the model. That isn't the only place they can go, and where they get stored is the decision a leader actually owns. You can retrain the model itself, or you can change the harness around it.
The harness is everything that isn't the model: the instructions the agent follows, the examples it's shown, its tools and checks, and whatever it remembers from last time. The model is the part you rent. The harness is the part you build, which is why I called it the moat back in June.
That difference grows every quarter, because the models keep getting replaced. A better one arrives every few months, and a lesson trained into last year's model stays behind when you switch. A lesson written into the harness comes with you.
It's cheaper too. The researchers behind GEPA had an agent rewrite its own instructions from its mistakes, and it beat retraining the model by 6% on average while needing up to 35 times fewer attempts.
I've moved my own work between models several times this year. What carried over each time was everything around the model: the instructions, the examples, the checks. Anything that had only ever lived inside the old model stayed behind.
So when a correction turns into a new check, or an example of what good looks like, or a rule the agent follows next time, it compounds. A competitor can license the same model you use.
They can't license your harness or the record of every time your loop has turned.
The loop can fool you
There's one way this goes wrong that looks exactly like success.
Economists call it Goodhart's law: when a measure becomes a target, it stops being a good measure. Tune an agent against the same eval for a quarter and the eval starts measuring how well you tuned, and stops telling you much about the work.
I watched this happen on one of my own research agents this summer. It ran more than six hundred experiments in a couple of days with a green dashboard the whole time. Almost all of them were the same experiment. Every metric said it was working. What caught it was a person reading the results, which is the slow signal doing its job.
That's the protection. As long as results from the real world keep coming back and rewriting what counts as good, the eval can't drift too far from the job. The loop has to keep turning for the same reason it was worth building.
The endpoint
Go back to the trial for a moment. The endpoint isn't there to satisfy an auditor. It's there so that when the data comes back, everyone already agreed on what it would mean, and the result can change what happens next.
That's the whole argument for AI exhaust, and the whole problem with it. The records your agents leave behind are worth a great deal, but only after someone has written down the job, decided what good looks like, and gone to get the result. Then every correction becomes a lesson, every lesson lands in a harness you own, and the value of the exhaust is set by how fast that loop turns.
None of that work belongs to the agent. Somebody has to own the definition of success for each process an agent touches, and somebody has to go and collect the outcome when it arrives.
In most companies, nobody has been given either job.



