Instrument Agentic Ad Lab

Questions to Ask an Agentic Advertising Vendor, and the Answers That Disqualify

Seven diligence checks an agentic advertising vendor can fail on the wire, and the targeting dimensions a seller MUST honour.

21 min read 9 chapters

A tool, not a read. Bring your own buy and work through it.

A diligence question is only worth asking if the vendor can fail it, and almost none of the questions being published this year can be failed. The seven below can. Three of them — the A/B, the endpoint call, and the capabilities read in section four — converge on one output: a short list of targeting dimensions, drawn from the nine that get_adcp_capabilities puts a MUST behind, named in your contract with that public declaration as the acceptance test. The other four produce their own failable objects — an authorisation file, a replay window, an error code, an approval field — rather than feeding that list.

Three sales agents in the AdCP agent registry answered an anonymous call on 12 August 2026, out of a catalogue of 23 listings: 18 typed sales, 3 signals, 1 creative, 1 buying (the AdCP agent registry, measured). One of the three emptied its whole catalogue when I sent a structured filters.channels: ["audio"]. None of the three dropped a product when the brief told it, in plain English, to exclude the inventory it was about to return, and two of them asserted a match against "purple monkey dishwasher quantum bicycle". Three agents is not a rate and nothing below treats it as one. It is a wiring observation, and the wiring is lopsided in a way that three samples are enough to show: the structured half of the request did something on one agent, and the prose half did nothing anywhere.

Infillion’s 15 Questions to Ask Your DSP for the Agentic Era, the list most buyers are working from, sits behind an email form, with two of its questions visible:

“Is your platform built for AI-driven buying, or layered onto legacy infrastructure?”

“Do you own any of your supply, and does your system prioritize it?”

Read the first one as a vendor. You say “built for AI-driven buying”, and so does the vendor after you, and so does the one after that, so the question separates nobody. Gartner named the thing you are screening for, agent washing, and put the count of real builders in June 2025 at roughly 130 among the thousands claiming agentic AI. That number is Gartner’s own and isn’t independently audited, but its direction matches what buyers say in public: the highest-scoring reply under an r/programmatic thread on Yahoo’s six new agents is that “having an OpenAI API wrapper doesn’t make things agentic”.

Three checks are worth more than the four that follow them, in this order:

CheckWhat it costs youWhat it proves
Send the same product request twice, one constraint changedAn hour of an engineer’s time, or a screen-share in the meetingWhether a constraint you supply changes what comes back
Ask for the agent’s URL and call itAn email, then an hourWhether the agent exists anywhere outside the vendor’s own login
Open the publisher’s /.well-known/adagents.json in a browserThirty secondsWhether the publisher has authorised this vendor to sell it

The sections run in that order, then four more. Four of the seven need somebody who can call an endpoint, one runs in a browser, two are conversations with nothing to run, and check two has a browser half and an endpoint half.

Both stacks are in scope: AdCP, the Ad Context Protocol, whose agents each publish one machine endpoint, and AAMP, IAB Tech Lab’s Agentic Advertising Management Protocols, whose reference code you can read but whose agents you cannot dial.

One structured filter excluded anything. No English sentence did.

This is the highest-yield check here, and if you only get one thing into the meeting, get this one. Send a product discovery request, send it again with a constraint that should exclude most of the catalogue, compare the two sets. That is the whole test, and it is the one thing a demo cannot fake, because you supply the second input.

Five requests went to each of the three agents that answered anonymously on 12 August 2026, over MCP, with no credentials:

RequestCora AIEquativNo Fluff Advisory
US-only CPG brief, CTV and online video4 products, all Korean FAST CTV and news1 product, “High CTR APAC Video”5 products: 4 fixed, 1 generated (“US market entry & GTM operators”)
Plus “Exclude all non-US inventory. Exclude Korean-language…”same 4same 1same 5, generated product on the same topic
"purple monkey dishwasher quantum bicycle"same 4same 1same 5, generated product switches to “Clean room & data collaboration”
filters.countries: ["US"]same 4, and all four declare ["KR","US"], so returning them is defensiblesame 1same 5
filters.channels: ["audio"]0 productssame 1, self-described “Synthetic data for demo…”same 5

Two of the three asserted a match against the nonsense string. Cora returned filter_diagnostics: {total_candidates: 4, matched_candidates: 4, no_match_targeting: false, semantics: "approximate"} and a brief_relevance per product reading “Matched against buyer brief:” with the brief echoed back. Equativ returned brief_relevance: "Matches video campaign requirements for APAC market with CTR optimization segment" against a brief saying United States, and again against the nonsense.

A second run the same day went to Cora alone: sixteen calls, brief held byte-identical, varying only filters, to separate “filters no-op on a small catalogue” from “filters ignored”.

Varied fieldCora’s answer, from a catalogue of 4
channelsctv 3, display 1, olv 0
countriesUS 4, KR 4, GB 0
delivery_typeguaranteed 4, non_guaranteed 0
min_exposures999999999 returns 0
countries + channelsGB + ctv 0, US + ctv 3
start_date / end_date2030-01-01 to 2030-01-02 returns 4
budget_rangeUSD min 1 max 2 returns 4
is_fixed_pricetrue 4, false 4
standard_formats_onlytrue returns 4

That is the honoured-versus-decorative split, measured on a live agent: four fields exclude (channels, countries, delivery_type, min_exposures), only channels down to a partial set rather than all-or-nothing, and five accepted fields do nothing (start_date, end_date, budget_range, is_fixed_price, standard_formats_only).

Cora also says why it excluded. Its empty responses carry no_match_reason (“No configured CORA AI test products match the requested channel, country, inventory, product, pricing, or measurement filters”) and suggested_refinements (“Try channels ctv or display”, “Use countries KR or US”). An agent that can name the constraint it enforced is a different capability class from one returning a bare empty array, and no compliance verdict separates them.

Four products is a small catalogue, No Fluff’s is five, and at those sizes returning everything and matching everything look identical from outside. Prose was not entirely inert — No Fluff Advisory generates its fifth product per brief from a corpus of essays, and that topic did follow the brief — but it could change what a product slot was about and never take a product away. What survives those caveats is the shape of the wiring: the structured half of get_products is where the implementation is, thin as it is, and the natural-language brief that every deck opens with came back on two of three agents as an echo with a match assertion attached.

A real answer runs the A/B live and the sets differ, or names which fields are hard filters and which are hints. “The agent interprets the brief holistically” is the answer that ends it, and the follow-up that exposes it is to ask what changes in the response when you remove a requirement. When credentials do not arrive before the second meeting, have them screen-share both calls and read the product IDs off the screen with you. You still supply the second input, which is the only property that matters here.

Ask for the endpoint URL, then call it yourself

Ask before the meeting rather than after: the agent’s URL, the authentication scheme, and either a sandbox tenant or a test key. Vendors who have built the thing send that by email the same day. Any version of “the agent is available inside our platform” ends the conversation, because an agent that only exists behind the vendor’s own login is a feature of their UI, and the premise of both stacks is that somebody else’s software calls it.

Ask for the URL, the transport and one worked request, not “a URL ending in /mcp”. On 16 August 2026 six publisher tenants behind prebidselleragentgateway.optimera.nyc (WebMD, TheWrap, Bustle Digital Group, Carpenter Media Group, Trusted Media Brands, Mediatonik) returned HTTP 405 to a JSON-RPC initialize at their enrolled URL, and HTTP 200 with a clean handshake once /mcp was appended. Three other listings publish a path the instruction never describes: Equativ at /v1/discover, LoopMe at /mcp/seller, Content Ignite at a bare host. Their mcp_endpoint field is byte-identical to url, so there is no second field to fall back on, and a publisher enrolled that way looks dead to any buyer trusting the catalogue.

The browser half is thirty seconds: open agenticadvertising.org/api/registry/agents and search for the vendor. Enrolment runs through an AgenticAdvertising.org member profile with no self-registration, so an absent vendor is not necessarily an absent agent. The same API reported 68 further agents discovered by crawling adagents.json files, a set the AdCP agent registry does not publish and which may overlap the 23. The endpoint half needs an engineer. Every registry agent speaks MCP, so a call is an HTTP POST and a session header. An anonymous handshake plus get_products, across the 18 sales agents:

ResultCountAgents
Returned products to an anonymous caller3Cora AI, Equativ, No Fluff Advisory
Handshake succeeded, get_products did not3goTom Sandbox, Adzymic (SPH), Adzymic (Mediacorp)
Credentials required at the handshake6AdCP Test Agent, InMobi (prod and non-prod), LoopMe, Purrsonality Seller, Pubx
Typed sales, exposes no get_products tool3Advertible, Dstillery, vastlint
Transport failure before any protocol ran3BidMachine (503), Content Ignite (500), Rediads (522 from Cloudflare)

Rows two and three, the nine agents that demand credentials at one layer or the other, are the healthy ones. An agent that refuses an anonymous stranger is behaving correctly, and the protocol’s own public test agent is in that group: it answers 401 Missing bearer token, the HTTP code for absent credentials.

Two of row two are my fault. Both Adzymic listings failed with 1 validation error for call[get_products] / promoted_offering / Unexpected keyword argument, and promoted_offering appears in get-products-request at none of 2.5.3, 3.0.0, 3.1.0, 3.1.10 or 3.1.13. Read that row as one agent refusing credentials and two rejecting a request the specification allows, which leaves seven agents demanding credentials rather than nine. Then read the rejection as its own finding: get-products-request sets additionalProperties: true, so the field Adzymic hard-errored on is permitted, and Cora, Equativ and No Fluff — the three that returned products — were the conformant ones for ignoring it. That last point cuts into the measurements above, which came off the same harness: if promoted_offering rode along on those calls too, the filtering matrix was gathered with an unrecognised property on the wire. The schema says it should have been ignored, and nothing in the three responses suggests otherwise, but ignored-by-schema is not the same as measured clean, and only a rerun without the field settles whether an unrecognised property changed what came back.

The last two rows are the ones worth pausing on. Advertible answers unknown tool "get_products", Dstillery offers get_signals only, and vastlint’s 14 tools are VAST validation and content-standards management. Any of those rows can belong to a vendor worth buying from. Not being able to tell you which row they are in disqualifies them.

Vox’s authorisation file is missing two required fields

When a vendor says your inventory is discoverable by buying agents, or that they are authorised to sell somebody else’s, the assertion lives in a file at https://<domain>/.well-known/adagents.json, and anyone can open it in a browser. I opened four on the same day: straitstimes.com serves a pointer file naming sales-agent.adzymic.ai as the authoritative location, vox.com serves a full file naming https://salesagent.voxmedia.com as its authorised agent, and cnn.com and theguardian.com return 404. One line that sentence does not give you: cnn.com answers 302 to edition.cnn.com, and that host returns the 404, where the Guardian’s is direct. All four behaved the same way again on 16 August 2026.

Look for the vendor’s endpoint in authorized_agents[].url, then look at the entry itself, because Vox’s has a defect worth knowing about. It carries url, name and properties. The 3.1.13 schema allows six entry shapes; each requires authorization_type plus its own variant field, and each also allOf-references core/authorized-agent-base.json, whose required array is ["url", "authorized_for"]. The entry fails all six shapes twice, once on the missing discriminator and once in the shared base. The requirement has held across versions rather than lapsing: 2.5.3 required url, authorized_for and authorization_type on each of its four shapes, and the 3.1.x rewrite moved the first two into the base file rather than dropping them. A buyer agent validating that file against the published schema has grounds to reject the authorisation outright. Most clients don’t validate today, which is why nobody has caught it, and “most clients don’t validate today” is not a sentence you want load-bearing in a contract.

“Discovery is handled through our integration” ends the conversation. Discovery is a file on a domain the publisher controls, and what belongs inside it, including the multi-megabyte network files this spot check did not reach, sits on agent discovery files.

Nine dimensions carry a MUST. Twenty-eight filters carry nothing.

Those are the three worth most. The four that follow are narrower, and the first of them is where the sentence you can put in a contract comes from.

AdCP 3.1.13 puts campaign prose in a brief string with no schema and no validation, and puts everything a seller can be asked to filter on in a sibling filters object: 30 properties, 29 of them filters and one an ext escape hatch, with required_axe_integrations deprecated. Twenty-eight live filters — countries, regions, metros, postal_areas, geo_proximity, channels, format_ids, delivery_type, exclusivity, budget_range, signal_targeting, required_performance_standards and the rest — and not one obliges the seller to anything, which is what the Cora matrix shows in both directions. “We support AdCP” says nothing about how many of the 28 do anything.

The obligation sits in a different document. get_adcp_capabilities returns a targeting block under media_buy.execution holding exactly nine dimensions (geo_countries, geo_regions, geo_metros, geo_postal_areas, geo_proximity, age_restriction, language, keyword_targets, negative_keywords), unchanged across 3.1.10 through 3.1.13, and its description reads: “If declared true/supported, buyer can use these targeting parameters and seller MUST honor them.” That sentence attaches a normative obligation to a public declaration, over a list short enough to read in a meeting, and it is what a procurement lead can put in a contract. “We are AdCP conformant” is not.

Adzymic publishes the answer this question is looking for. Its live capabilities response declares three targeting dimensions true, geo_countries, geo_regions and device_platform, alongside content_standards: false. I would onboard the vendor who published three honoured fields over the one claiming twenty-eight in a deck, because only one of those two numbers can be checked.

One qualifier on the third: device_platform is not a property of that capabilities block at 3.1.10, 3.1.11, 3.1.12 or 3.1.13. It lives in core/targeting.json, a different 28-property document, per package rather than per request. Adzymic’s declaration therefore binds on two of the nine, and the third is a promise in a vocabulary the MUST does not reach.

A real answer names which fields cause a product to be excluded from the response and which are advisory. The disqualifying answer is “we are fully compliant with the spec”. A vendor who can’t split the list either hasn’t implemented filtering, or has implemented it without knowing which fields went live.

Run it yourself, with an engineer. One call to get_adcp_capabilities returns the declaration, and Adzymic’s is public.

The idempotency declaration, with the replay window in seconds

The same call carries the retry contract, and the reference code is the argument for asking. POST /api/v1/negotiations/messages on IAB Tech Lab’s seller-agent required an idempotency_key, never honoured it, and let a retried counter-offer burn a round from a negotiation’s max_rounds budget, so “enough retries can push a legitimate negotiation past max_rounds into an unintended REJECT” (issue 51). Issue 44 was the same class on quotes: the request arrives twice, the agent prices it twice. A required field that the handler never reads is worse than no field, because it buys the caller a guarantee that doesn’t exist. Both closed on 11 August 2026, to commits 47beb9b4 and 89d0eda3, filed and fixed by the same outside contributor, with the maintainer conceding on 44 that the key is “intentionally required - no silent auto-minting” and that /deals misses the payload-divergence case too.

Neither issue documents a duplicated create_media_buy landing against a live insertion order, so size the money exposure as theoretical for now and bounded by your daily cap: a retry doesn’t create budget, it creates a second commitment against the same one. That is still a commercial question rather than a technical annex, because the person who eats a double booking is not the person who wrote the retry loop.

AdCP’s declaration is an idempotency object next to adcp.supported_versions. The schema’s own language is unusually direct: clients MUST NOT assume a default, and a seller without the declaration “is non-compliant and should be treated as unsafe for retry-sensitive operations”. Its numbers are contractable too: replay_ttl_seconds runs from 3600 to 604800 with 86400 recommended, a divergent payload inside the window returns IDEMPOTENCY_CONFLICT and past it IDEMPOTENCY_EXPIRED, and in_flight_max_seconds is published so a buyer SDK sizes its retry budget against that bound rather than the far wider replay window. Cora declares {supported: true, replay_ttl_seconds: 86400}, the recommended value rather than an arbitrary one. goTom declares request signing supported and required for create_media_buy.

A real answer is the declaration itself, with the replay window in seconds. The disqualifying answer is a description of the retry policy with no idempotency object behind it, or a protocol version quoted from a slide with no supported_versions on the wire.

Run it yourself, with an engineer. Fetching the declaration takes one call, so ask them to make it while you are in the room.

The reference seller agent invents a price for a product that does not exist

Ask what the agent does when it hits something it cannot answer: a product missing from the catalogue, or a brief it cannot parse. The reference implementations answer that in public, and their answer isn’t reassuring. On 11 August 2026 the same contributor filed issue 57: POST /api/v1/deals/curated takes an arbitrary product_id with no existence check, and when the ID matches nothing it falls back to a hardcoded base_cpm = 12.0 and persists a confirmed, bookable deal at $13.20 total CPM, against real inventory the same instance prices at $45 base and a $35 floor. The docstring says the opposite: “A known-but-unpriced product (no base/floor CPM) is a 422 — never a fabricated price.” The reporter’s sharper point is the half a publisher has to price: that route carries no auth dependency at all, unlike every other deal-creation surface on the agent, so the caller minting the deal need not be anybody.

The guarantee is implemented for products that exist without pricing and skipped for products that do not exist, which is the failure mode to interrogate in both stacks: a confident, bookable number standing where an error belonged. Hold it to its evidence, though, which is thinner than the AdCP measurements above: this is a code read on a public repository, and nothing here dialled a deployed AAMP seller agent to see whether the route is reachable on anyone’s production tenant.

A real answer names the error code and shows you the response body. The disqualifying answer is “it falls back to a sensible default”.

Nothing to run here. Unless the vendor gives you write access to a sandbox, this one is a conversation, and the two issues above are the vocabulary for having it.

Ask which field records that a human was asked

Amy Porter of RPA told Digiday there is a “real risk they could also obscure critical decision-making if advertisers rely too much on AI”. She is right, and the remedy is narrower than the warning suggests. The obscurity nobody can fix sits in the model’s reasoning, which nobody was ever going to read. The obscurity you can fix is whether a protocol field records that a human was asked and what they said.

AdCP’s answer is thinner than the question deserves. Two wire fields record human review, both in the plan audit log: summary.statuses.human_reviewed, an integer described as a “supplementary count of checks that went through internal human review”, and summary.escalations[], whose items carry check_id, reason, a resolved_at timestamp and a free-text resolution the schema illustrates with “approved_by_human” and “rejected_by_human”. No identity appears anywhere in that object: no user_id, no approver, no reviewer. check_governance’s own response schema has no human-review field at all, and the audit-log schema is marked x-status: experimental. The wire records that somebody was asked and when it resolved, never who, and never in a value a machine can read. The state the buy sits in while it waits is named elsewhere, in AdCP’s submitted and input-required async states. AAMP’s buyer agent puts the switch upstream instead, in an ApprovalConfig with separate booleans for plan review, booking, creative and pacing adjustment.

So ask which specific operations require a second human, and which field enforces that, rather than which ones the vendor’s policy says require one. A role in the vendor’s UI is a setting their own staff can change on a Friday afternoon; the field is what the log shows on Monday. A real answer names the field and the state the buy sits in while it waits. The answer that ends it is “our platform has an approvals workflow” with no field named, because a workflow that leaves nothing on the wire leaves neither side a record on Monday. The autonomy ladder maps each rung to the mechanism behind it.

Nothing to run here either, until you have a live integration and can read your own logs.

The compliance badge is a tiebreak, not one of the seven

AdCP runs a hosted compliance programme with storyboards, tracks and a badge, 38 storyboards in the universal set. It is the most developed conformance apparatus in agentic advertising, and it doesn’t test the thing the A/B tests. Issue 2902, open since April and filed by an outside implementer, says why:

“An endpoint can accept start_date/end_date, return a correctly-shaped response, pass all compliance checks — and silently return the same data regardless of what dates were passed. The schema is valid. The contract is broken.”

Not a hypothetical: Cora returned the same four products to a flight window of 2030-01-01 to 2030-01-02, on that exact field pair, in the run tabulated above. Issue 6374, open since 11 August, is the reporting half, with verified exposed as a bare boolean, three unrelated causes behind a false and no field to tell them apart. A third, issue 5495, had the runner silently strip a spec-defined field an agent’s tool schema did not declare, log a console warning that never reaches the report, and score the step a pass: “A conformance tool reporting ‘pass’ for a capability the agent never exercised is a false positive.” That one closed as completed on 18 June 2026. The case rests on 2902’s assertion, not on my having opened the storyboard bodies to see whether any sends a filter and asserts an exclusion.

The AAMP side has no badge to weaken. Its conformance kit is a script the vendor runs in their own CI, emitting a gap_report.json, with no third-party register of who passed. So on that stack, ask for the gap report. “AAMP-compliant” is, today, entirely self-asserted.

Rank the badge accordingly: it is a tiebreak between two vendors who both passed the A/B, it is evidence of effort and I would weight it above a case study, and it is not a substitute for running the A/B yourself.

The other tiebreak in the same class is who fixes their stack, how fast, and which bugs they leave alone. Issues 44 and 51 went from filed to closed within days. Issue 57 was opened one minute after 51 closed and is still open with zero comments on 16 August 2026. Two retry bugs and one that mints money out of nothing, sorted the wrong way round.

Where these checks stop working

Three agents, one day, no credentials, and anonymous access shows only what these systems do for a stranger. The catalogue behind it holds what an AAO member enrolled at public visibility, and nothing else. Rows move fast, too: BidMachine went from the worst outcome class to the healthiest between 12 and 16 August 2026, while the six mis-enrolled Optimera tenants stayed broken across the same window. Rerun the checks rather than reading the rosters as a shortlist.

The checks verify mechanism. They say nothing about whether the buy performed better, which is the objection carrying the most weight in both threads I read. On value, from r/programmatic:

“so far there’s not a good value preposition [sic] that can’t be addressed by a bulk sheet or API.” — u/GreenFlyingSauce, r/programmatic

And on proof, from r/adops:

“I have yet to see any agentic experience that you can actually trust to execute, the ones that work don’t save enough time nor have any proof of improved performance because attributed metrics don’t mean incremental performance. AI has no absolute accountability either.” — u/Toast687, r/adops

Both are practitioner sentiment rather than measurement, and nothing on the wire answers either. Neither stack carries incrementality in any field.

A vendor who passes all seven checks has shown you the integration works. Whether it buys better is a separate claim, and none of these checks touch it. Named deployments sit in who is running agentic buys, where every headline figure came from the vendor that sold the campaign. Which is why the check that generalises is the endpoint call, and why a vendor who says plainly which parts are not built yet can still be the right pilot partner.

When one of them fails the A/B and you want them anyway, put the fix in the pilot contract rather than in the minutes. Name the dimensions from the nine your planning depends on, quote the schema’s own sentence back at them, and set a date by which get_adcp_capabilities declares each one supported. That declaration is the reporting obligation and the acceptance test in one object: public, machine-readable, and re-readable on the date without asking anybody’s permission.