Read the field list before you read the pitch. PubMatic's governance announcement for its agentic stack says that every agent action is logged: what it did, when it did it, under what parameters it was running, and whether a human was in the loop. Four things, all of them real, all of them more than most vendors are currently offering. Now notice what a complete and honest record of those four things still cannot tell you.
It cannot tell you what the agent did not do. There is no field for the allocation that was available at the same moment and lost, no field for how close it came, no field for what it was measured against. The file is a full account of an event and contains no account of a choice.
I want to be careful with this, because the easy version of the argument is wrong and I nearly wrote it. Everyone who has spent a week near agentic systems in a marketing context has had the same uneasy feeling about audit logs, and most of the essays reaching for it end up demanding that the machine explain its reasoning. That demand cannot be satisfied, and asking for it is how you get dismissed by the people who build these things. The real gap is narrower, more boring, and entirely fixable.
What the industry calls this, and what it actually is
The trade framing is transparency. Buyers do not trust autonomous agents, so vendors are shipping visibility, and the visibility they ship has converged on one pattern: log the action, then attach a sentence explaining it. A natural-language rationale, generated after the fact, sitting in a column next to the event. It reads like an explanation. It is a caption.
The accurate description is duller and worse. Proof is a ratio, and an industry that celebrates evidence while forgetting what sits under the line is measuring effort rather than effect. These logs have a numerator with nothing beneath it. They record the thing that happened and not the population it was drawn from, which means every entry reads as inevitable, because the only option in the file is the one that won.
A record with one option in it is not a record of a decision. It is a record of an outcome, written by the thing that produced it.
Nobody is asking the machine to introspect
Here is where I have to concede the mechanics fully, because getting this wrong would be fatal to the argument.
Agents do not deliberate. There is no chamber in which options are weighed, no moment of hesitation, nothing that corresponds to a person sitting with a decision. There is a scoring function over an eligible set, and the top score is emitted. Anyone who has worked on bidding or ranking will tell you that talking about what the model considered and rejected is projecting a human process onto arithmetic, and they are right to say so.
But the thing I want logged was never the reasoning. A scoring function has a domain. To rank a set, the system has to build the set: the candidates exist, in memory, at the instant of the decision, with numbers attached to them. Then one of them is emitted and the rest are dropped.
That dropping is a default, not a law. The distinction that matters is between expensive to produce and cheap to produce and thrown away, and almost everything in this space belongs to the second category. Nobody is being asked to compute anything new. They are being asked to keep something they already made.
The size of the ask, stated honestly
The obvious counter is volume, and it is legitimate rather than evasive, so let me scope the claim before someone else scopes it for me.
Do not ask for candidate sets on auction-level bidding. A large advertiser participates in millions of auctions a day and a platform runs billions. Retaining per-auction candidate sets at that scale is a genuine infrastructure argument, not a dodge, and an essay that waved it away would deserve to be dismissed.
Agent-level decisions are a different order of magnitude. When an autonomous media-buying agent shifts budget between line items, changes a target, pauses a placement, or restructures a campaign, it does so tens of times a day per account, not millions. Those are precisely the decisions the industry is currently claiming to audit, and they are precisely the ones where keeping the candidate set costs a rounding error in storage. The gap between what is affordable and what is being logged is not technical. It is that nobody specified the field.
Three people who would push back
The platform engineer says I am anthropomorphising a softmax. Conceded above, and it is the objection that most improved this piece. What survives it is a request for the domain of a function, which is an ordinary engineering artifact.
The agency operations lead says they already log candidates, that it comes to tens of gigabytes a day, and that nobody has ever opened the file. Also true, and it is the more interesting objection, because it is about the difference between retention and legibility. A dump you never read is not a record. But that argues for a schema, sampling and a review cadence. It does not argue for discarding the data, and it certainly does not argue for a vendor discarding it on your behalf.
Then the pragmatist, who is the one worth answering at length: the agent hit its target, inside its guardrails, and the log proves it. What more do you want?
Compliance is not quality
This is the whole essay, so I will put it as plainly as I can.
An action log lets you audit compliance. Did the agent stay inside the limits you set. Did it act under an authorised identity. Was a human in the loop where your policy required one. These are real questions and the current logs answer them well.
A record carrying the set the choice came from lets you audit quality. Was the thing it picked better than the thing it passed over, and by how much, and has that margin been narrowing for three weeks. That is a different question and no amount of compliance evidence gets you to it.
Notice which of the two you can do anything with. If your record shows that the agent moved four hundred thousand dollars from one line item to another and stayed within policy, your entire relationship with that vendor is the sentence the agent behaved. There is nothing to negotiate; behaviour is binary and it passed. If your record shows the set it chose from and what each option scored, you can point at the one that came second and ask why it lost. Now you are having a conversation about performance with numbers in it, which is the only kind of conversation that ever moves a contract.
The pattern is one I keep running into from different directions. A number generated by the party being paid on it is a seller's counterfactual, and it can be perfectly accurate while remaining a quote. An audit log written by the party being audited is the same object in a different costume, and the fact that it is factually correct in every field is not the reassurance it appears to be. I made a version of this argument about whether you are permitted to hold evidence at all. This one sits underneath it: even where you are permitted, and even where you win custody, what you take custody of may not contain the thing you needed.
A seller can name the gap without filling it
There is a small, precise illustration of this available right now, and I include it as evidence rather than as the point.
Google's AI Max reporting adds a segment called search_term_match_source on the search term view, with three values: ADVERTISER_PROVIDED_KEYWORD for traffic your keywords earned, AI_MAX_BROAD_MATCH for traffic expanded out of an existing keyword, and AI_MAX_KEYWORDLESS for traffic matched from page and asset content with no keyword involved. That third value is a seller writing down, in its own developer documentation, that a category of its matching has no advertiser-supplied input behind it. The label is honest. It is also the whole of the disclosure: you learn that a keywordless match occurred, and nothing about what else the system could have matched instead. Naming a gap is not the same as filling it, and this is what naming looks like.
This site argues more than it claims, and an argument cannot be scored in retrospect. So here is one with a date on it.
On June 30, 2027, read the public product documentation for the major agentic media-buying offerings and check whether any of them exposes, in the advertiser-facing log, the candidate set and the per-candidate scores behind an agent-level budget or bid-target action. I claim none of them will. The logs will still carry the action, the timestamp, the parameters, the approval, and by then a longer and better-written natural-language rationale.
Two minutes to verify and it can go against me in public, which is the point. If a vendor ships the candidate set before then, I was too pessimistic about how the market responds to a question buyers have not yet learned to ask.
The instrument already exists and needs one row
The Counterfactual Audit was built for a narrower case: take a claim a seller generates about its own contribution, price it against your own denominator, and log whether the claim holds steady over time. An oscillating claim tells you, using the seller's own instrument, that the number is a quote.
Give it a row for the agent case. For each class of agent-level decision, record four things: whether the log contains the candidate set, whether it contains a score per candidate, what the margin was between the chosen option and the runner-up, and the date you last looked. Four fields. The margin column is the one that earns its keep, because a margin that collapses toward zero over successive weeks tells you the agent is choosing between options that no longer differ, which is the moment its autonomy stopped being worth what you are paying for it.
That is also the honest answer to the operations lead. You do not need forty gigabytes. You need a margin, sampled, on a cadence, with a name against it.
Who owns the shape of the record
Somebody in your organisation has to specify what the record contains, and they have to do it before procurement signs, because a field list is a contract term beforehand and a feature request afterwards. At most companies that role does not exist. Media has an owner. Tooling has an owner. The shape of the evidence has nobody, which means it has a default owner, and the default owner is whichever vendor wrote the schema first.
This is the same failure I described when a second seller enters the mix and two conversion definitions get added together in a board deck. In that case the definition of the unit was accepted by nobody in particular. Here it is the definition of the record. Both are finance decisions that arrive dressed as configuration, and both get made whether or not anyone decides to make them.
The vocabulary for this, decision classes, default owners, kill conditions and the ledgers underneath them, lives on the Decision Ownership hub. What this essay adds is small: the record of an agentic decision is itself a decision class, and it is currently unowned at almost every company buying one.
I am not certain how this resolves. It is possible that a serious buyer asks the question at scale next year and the schemas change quietly, in which case the claim above goes against me and I will take that. What I am confident about is much smaller. Before you sign an agentic media-buying contract, ask one question: does your log contain the candidate set and the scores for each agent action, or only the action that was taken? It costs nothing, the vendor can answer it today, and the answer is a fact about a schema rather than a matter of opinion. Ask it while it is still a term. Once the ink is dry, you are not negotiating, you are filing a request, and the definition you never signed off on has already become the one your numbers are made of.