Here is a design I admire. You see an ad inside an AI assistant's answer. If you click, you do not land on a website. A separate conversation opens, with an AI representative of the business, labelled as sponsored and kept apart from the assistant's own answers. The label is on. The wall between the two is, as far as anyone has described it, real. The assistant's independent answer is left alone.

It is a careful piece of engineering and it answers one question cleanly: is this an ad? It does not answer the question that comes straight after it, which is whether the thing in that conversation would have said the same if nobody had paid for the conversation. Those are different questions. The industry keeps building an answer to the first and letting it stand in for the second.

Two instruments that get mistaken for one

Disclosure tells you who is speaking. Verification tells you whether what they said is representative of what could have been said. A label is a disclosure instrument and it is a good one. It cannot do the second job, because the second job needs something outside the conversation to compare against, and a label is inside it.

I have run into this shape twice already on this site. An agent's audit log records every field correctly and leaves out the set the action was chosen from. A seller's counterfactual can be accurate in every particular and still be a quote from the party paid when you act on it. In each case the record is honest and the honesty is beside the point, because the thing you needed was never inside the record. This is a third instance, on a platform that is not Google, in a channel where my usual remedy does not work. More on that shortly.

The title is meant in the plain sense of both words. A label is what is printed on the outside of a thing. A ledger is the record you keep of what happened and what it was measured against. A product can carry an immaculate label and have no ledger at all, and a sponsored conversation that is clearly marked as an ad has told you its status, not its worth.

What you cannot decompose, you cannot manage

The failure I want to name is not that a number from a sponsored conversation might be fraudulent. Most will not be. The failure is that when it moves, you cannot tell why.

Say the conversion rate on those conversations rises by a fifth in a month. There are at least two stories. The product got more compelling to the people arriving, or the conversation got better at sounding compelling. They look identical in a dashboard and they call for opposite responses. If it is the first, you spend more. If it is the second, you are buying persuasion that may not travel with the customer once they leave the chat, and may stop working the moment another advertiser's agent persuades equally well or the format changes under you.

You can optimise a metric you cannot decompose. Turn a knob, watch the line, turn it again. What you cannot do is manage it, because management means knowing what a change in the number is a change in. That is the cost of ignoring this, and it is not a fake number. It is a real number whose cause you have outsourced.

You can optimise a number you cannot decompose. You cannot manage it.

Three people who would push back

The platform's product lead says: we already disclose. What more do you want, our ranking logic? No. Nobody is asking for the logic, and I would not expect to get it. The ask sits entirely on the advertiser's side of the table, which is why the remedy below needs nothing from the platform at all.

The pragmatist says: it converts, and I do not need to know why. That is a reasonable position for a quarter. It is a poor one for a budget line that is going to grow, because the first time the number turns down you will want to know which of the two stories you were living in, and you will not have kept the evidence that would tell you.

The third one is the objector I take most seriously, because most advertisers are this person. A small advertiser has no leverage with any platform. I cannot get an audit clause out of anyone, they say. What do I do on Monday? I wrote about contract clauses in the audit essay, and that remedy assumes an account big enough for someone to take your call. This essay is for the reader who does not have one. The answer has to be something you can do without asking anyone.

The worked example

Take OpenAI's Sponsored Agents, which is the format I have been describing. OpenAI's January 16 advertising principles say that ads "do not influence the answers ChatGPT gives you" and that ads will be "clearly labeled and separated from the organic answer." Its help page says Sponsored Agents are in limited alpha, available only to selected advertisers, with no early-access requests being taken. That page says nothing about how a sponsored conversation is distinguished from ChatGPT's ordinary responses, and nothing about the accuracy or independence of what is said inside one. Trade coverage attributes a line to the announcement, that these conversations are clearly labelled and kept separate from ChatGPT's independent answers. I could not find that sentence on the help page, so treat it as reported rather than confirmed. One trade write-up also notes that no pricing has been published and that it is not stated how an interaction will be counted.

Notice what the independence principle protects. It protects the answer from the ad. The sponsored conversation is the ad. Nothing in either document addresses whether that conversation's product claims are representative of what the same assistant would say unprompted, and, as far as I can find, there is no way for an outside party to test it, because no outside party has the conversation logs at scale. That absence is the evidence. It needs no disputed number from the seller, which is why I would rather build the argument on it.

A policy, not a negotiation

The policy is one sentence. Any number sourced from inside a sponsored conversation is signal until something outside the conversation corroborates it. The vocabulary is from the evidence-layer essay: the signal layer is what triggers a decision, the evidence layer is what scores whether the action worked, independently of the action. A sponsored-agent conversion is signal by construction, because it is the platform's channel counting the platform's conversation. Treating it as evidence is that essay's phrase for it, a signal-layer object wearing evidence-layer clothes. It is also a numerator with nothing under it, which is the problem the denominator essay is about.

In practice that means three small things. Tag the source on the line in your reporting, so nobody reads it as ordinary conversions. Keep it out of the decision-grade pages, the ones that move budget or reset targets, until something you hold corroborates it: orders in your own system, a holdout you control, returns and support contacts from the people who came in that way. And when it moves, ask the decomposition question before you act on it. This is the same instinct as asking, of any number a seller hands you, what would force it out if it were bad, applied to one channel. It also fits with counting two sellers' conversions as one total, which is a mistake the policy prevents.

None of this costs anything or needs a call. A measurement lead can adopt it this afternoon.

A claim you can check

This site argues more than it claims, and an argument cannot be scored later. So here is one with a date.

By December 31, 2027, no major conversational-commerce advertising surface will give an independent party a way to verify that a disclosed sponsored-agent conversation's specific product claims match what the same assistant would say to the same question unprompted. Disclosure that the conversation is sponsored will exist. Verification that its content is representative will not. Check it against public product documentation and terms on that date. This is a different dimension from the other dated claims on this site, which concern the contents of agent logs, the right to commission a comparison, and reporting. It is about representativeness inside a consumer-facing conversation. If an outside party can run that comparison by then, I am wrong.

I started out assuming this would be another essay about asking for something in a contract, and I am glad I could not write that version, because it would have been addressed to nobody who needed it. What I am sure of is smaller. A label tells you who is talking. It cannot tell you whether what they said would survive being compared with anything else, and a budget that grows on a number it cannot decompose is a budget that depends on nobody ever needing to know why the number moved. Tag the source, hold the line below the decision pages, and wait for something outside the conversation to agree with it.

Share LinkedIn Post