Oversight has a list price now. Microsoft's Agent 365 went generally available on the first of May at fifteen dollars per user per month, and it sits next to the thirty-dollar Copilot seat that supplies the agents it watches. Salesforce meters the same idea three ways: two dollars a conversation, five hundred dollars per hundred thousand flex credits, or a per-user licence starting at a hundred and twenty-five. Agent revenue there ran around eight hundred million in a quarter, up from roughly five forty the quarter before. Whatever else is true, a market has formed, and it is the market for watching the thing you just bought, sold by the company that sold you the thing.
Meanwhile the obligation to watch it stopped being optional. The EU's new product liability directive has to be in national law by the ninth of December, and it treats software and AI systems as products in their own right, with strict liability attached: a claimant shows the defect and the damage, not that you were careless. California's AB 316 took effect on the first of January and removes the sentence that a defendant who developed, modified or used an AI system may not now say in court, which is that the AI did it on its own. Other defences survive. That one is gone.
Put those two together and the shape is clear enough to make out. The obligation to prove your agent behaved is landing on you. The instrument that discharges it is being sold to you by the party your agent works for.
The version of this argument that I do not want to make
The easy essay here writes itself and it is a bad essay. It says vendors are conflicted, the fox is guarding the henhouse, caveat emptor, and it ends with an exhortation to be vigilant. Everybody already believes that. Nobody changes a contract because of it, and articles making that case have been landing in inboxes for two years without moving a single procurement process I have seen.
The reason it does not move anything is that it is a claim about motive. Claims about motive can be answered with sincerity, and the people building these products are, in my experience, mostly sincere. A vendor can be entirely honest and the problem I am describing remains exactly as large, which is the tell that motive was never the variable.
So I want to make a claim about capability instead, and to get there I have to concede something substantial first.
Platforms run counterfactuals constantly
They do. This is not a field where nobody knows how to construct a comparison. Geo holdouts, conversion lift studies, ghost bidding, matched-market tests, incrementality experiments wired into the buying interface: the major platforms have been running these for years and in many cases run them better than the agencies asking for them. If the argument were that the industry lacks counterfactual machinery, an engineer would correct me in the first reply and would be right to.
The narrowing is this. Every one of those comparisons is authored by the seller. The seller picks the question, picks the arm, picks the window, picks when to stop. What no product on the market offers is the inverse: a comparison you specify, on a decision you name, at a time you choose, delivered raw.
And the reason is structural rather than sinister. The comparison a buyer would most want to commission is the one that tests whether being in the product beats not being in it. Nobody builds, unprompted, the machine whose most natural output is a reason to cancel. That is not corruption. It is what happens when a capability has no customer asking for it, and it is the same shape I keep meeting from different angles: a number produced by the party paid on it is a seller's counterfactual, and it can be accurate in every particular while still being a quote.
The missing thing is not the experiment. It is your authorship of the question.
But every industry buys assurance from interested parties
This is the strongest objection and it deserves a real answer, because the person making it is usually the most experienced person in the room. Financial audit is paid for by the audited company. It works anyway. Why should software be different?
It works anyway because of three things sitting around it, and it is worth naming them separately rather than gesturing at them as a bundle.
There is a licensing body that can end an auditor's ability to practise, which means the auditor has something at risk that is larger than the fee. There is a mandated standard, so the scope of the work is not renegotiated client by client and an auditor cannot quietly agree to look at less. And there is separation: the firm auditing the accounts does not also sell the company its accounting system, because that combination was tried and the consequences of trying it are the reason the rule exists.
Agent governance has none of the three. No licence. No standard for what an agent record must contain. And no separation whatsoever, since the same vendor sells the agent and the layer that observes it, a combination that in financial audit would be prohibited outright.
The conclusion I draw is not that software people are less trustworthy than accountants. It is duller and more useful. Every safeguard that makes financial audit credible was built after a failure, written into law by people cleaning up a wreck. Agent governance is at the stage financial audit occupied before any of that existed, and buying at that stage is not wrong, but it does mean the protections you are imagining around the transaction are not there yet. You are earlier than you feel.
What a seller's own baseline looks like in practice
Google's published figure for AI Max is that advertisers using the full feature suite see about seven percent more conversions or conversion value at a similar CPA or ROAS. I have no reason to doubt the number. What repays attention is what it is measured against, which Google states plainly: the full suite compared with using search term matching alone. Both arms are inside the product.
Through September, campaigns still running campaign-level broad match or the older automatically created assets were moved onto AI Max, with the window closing at the end of the month. The comparison a migrated advertiser would actually want is the obvious one, thirty days migrated against thirty days not, and it does not exist and will never be constructed, because the party able to construct it has no reason to and the party wanting it has no standing to ask. Nothing here is hidden. The baseline is published. It is simply the seller's baseline, and there is no second one.
I could delete the preceding two paragraphs and this essay would say the same thing, which is roughly the point. Google is not the villain of this piece. It is the convenient example because its documentation is unusually explicit about what it compared.
Three people who would push back
The vendor's architect says they run comparisons all the time, and has the experiment console open to prove it. Conceded above, and the concession improved this piece. What survives is narrower: the buyer cannot commission one.
The procurement lead says every industry buys assurance from interested parties. Answered above, and the answer took me somewhere I did not expect when I started writing, which is that the analogy holds beautifully right up until you list what makes it work.
Then the CFO, who is the one worth answering at length: you want me to fund a second evidence system to check the first one. Show me the line.
There is no line, because it is a clause
Here is the part that costs nothing, and it is the whole practical content of this essay.
While terms are still open, ask for one thing: the right to commission a comparison arm on a decision you name, at a frequency you name, with the result delivered raw. A held-out cohort, a suppressed geography, a parallel policy. Not better reporting. Not a dashboard. The right to ask a question the vendor did not choose.
Three specifics keep it from being satisfied with a slide. Name who at the vendor is obliged to run it. Name the maximum delay between request and delivery. And say in the text that the comparison may include a period of non-use of the product, because that is the arm everyone tiptoes around and it is the one that tells you what you are actually buying.
You are not funding a system. You are reserving a right, in a document being signed anyway, at a marginal cost of nothing. And the answer informs you either way: a vendor who agrees has confirmed the capability exists, and a vendor who declines has confirmed the same thing while telling you where it sits on the roadmap.
This site argues more than it claims, and an argument cannot be scored in retrospect. So here is one with a date on it.
On June 30, 2027, read the standard commercial terms of the major agentic marketing and media-buying platforms and check whether any of them grants the customer a buyer-commissioned comparison arm: the right to specify a holdout or alternative-policy cohort on a decision of the customer's choosing and receive the raw result. I claim none of them will. Vendor-designed experiment products and seller-reported lift studies do not count, because they are the thing that already exists.
That date already carries a separate claim of mine about what an agent log contains. The pairing is deliberate: one is about the contents of a record and one is about the rights attached to it, and I would rather be graded on both at once.
The instrument, and one more row
The Counterfactual Audit was built to price a seller's claim about its own contribution against a denominator you control, and to log whether the claim holds steady. It needs one more row for this case.
For each vendor relationship, record four things: whether your contract grants a buyer-commissioned arm, the last date you exercised it, the delay between asking and receiving, and whether the result arrived raw or summarised. The delay column is the one that earns its keep. A right that takes eleven weeks to exercise is not a right, it is a courtesy, and the number tells you which one you have long before any dispute makes it matter.
None of this works without something underneath it to compare against, which is the older point: proof is a ratio, and an industry celebrating evidence while forgetting what sits below the line will accept an arm it did not choose simply because an arm was offered.
Who owns this before procurement signs
Legal owns the terms. Procurement owns the price. Marketing owns the outcome. The clause that would let anyone check the outcome belongs to none of them, which means it belongs to the default owner, and the default owner is whoever drafted the contract. That is the vendor.
The vocabulary for this, decision classes and default owners and kill conditions and the ledgers underneath them, lives on the Decision Ownership hub. What this essay adds is one line: the right to commission a comparison is itself a decision class, it is currently owned by nobody on the buying side, and a decision class with no owner gets settled by whoever writes the document first.
I do not know how this resolves, and the honest version is that I have changed my mind once already while writing it. I started out thinking the problem was that vendors would not tell you the truth, and the more I looked the more it seemed they mostly do, about the questions they selected. What I am confident about is smaller and fits in a sentence. Before you sign, ask for the right to commission the comparison, name the person and the delay and the non-use arm, and get it in writing while it is still a term. Afterwards it is a feature request, and a question you never reserved the right to ask is a question the answer to which is permanently theirs.