Instrument Agentic Ad Lab

The Autonomy Ladder for Agentic Media Buying: What a Counterparty Can Prove

Five rungs from read-only reporting to unattended spend, and the AdCP and AAMP fields, credentials and contract clauses that decide which of them a counterparty can actually verify on a live buy.

25 min read 10 chapters

A tool, not a read. Bring your own buy and work through it.

Every vendor deck in this category lands on the same sentence: humans set the rules, agents operate within them. Nobody publishes the rules. Five rungs run from reading your campaign data at the bottom to moving your money at the top, and each one is held, or not held, by something specific: a field on the wire for some, a credential scope or a contract clause for the rest. Those fields invert the framing. Autonomy is not a dial the buyer sets. It is a property of whichever counterparty’s implementation is loosest, and the buyer usually can’t see it.

An ad ops lead put the question to r/adops directly:

How much trust do you put in these new AI agents to execute optimizations and shift budgets, or are you also just using them as reporting assistants? What’s a part of the process that is NEVER going to be taken by AI?

The top reply:

we automate the boring plumbing, data pulls, normalization, anomaly alerts, basic pacing dashboards, because that’s repeatable and easy to QA at scale. budget shifts and bid changes stay human since context like client politics, promo timing, and platform quirks still matter a lot.

The Ad Context Protocol says this about itself, in the schema for the actions a seller currently allows on a media buy: the mode and sla values “are advisory at the moment of emission; sellers MAY resolve to a different mode by the time the mutation arrives (state can change), in which case the request is rejected with ACTION_NOT_ALLOWED (reason: mode_mismatch).” So a rung is a claim about a counterparty, and a claim counts only as far as that counterparty publishes something you can check.

The five rungs

RungWhat the agent may doWhat breaks if it is wrongHow you undo itWhat the counterparty publishes
1. ReadPull delivery, forecasts, product catalogues; normalise; alertA wrong number in a client deckResend the deckMCP readOnlyHint, set by 3 of the 13 listed agents that answered tools/list on 12 August 2026
2. ProposeAssemble the exact buy and stop before the writeNothing, as long as the proposal stays off the wireDiscard the draftNothing. The nearest affordance, paused: true on create_media_buy, commits you and withholds impressions, and nothing states that a human cleared it
3. EnvelopeMutate a live buy inside tolerances the seller declaresOver-delivery on the wrong package at your own CPMupdate_media_buy mid-flight, if the mode still resolves your way; impressions already served are not recoverableavailable_actions[].mode, advisory when emitted and re-resolved when your mutation lands; budget.reallocation_threshold
4. Post-hocCommit spend; findings logged, not blockedA weekend of budget in the wrong placePause, then argue about the cancellation feecancellation_policy.notice_period, required on the policy. Governance mode arrives on the response, after the decision it describes
5. UnattendedReallocate freely up to the plan totalThe cap is the only thing between the agent and the whole budgetNothing below the capbudget.reallocation_unlimited, which you declare to your own governance agent and the seller never sees; floor_price and round caps, which are seller-agent constants rather than protocol fields

The fourth column matters more than the second. Between rungs 3 and 4 the undo stops being a keystroke and becomes a contract negotiation.

Two dials are in play in that table and only one of them is a ladder. Rungs 1, 2, 3 and 5 grade how much the agent may write. Rung 4 grades when a human looks, and post-hoc review attaches to any rung above the first: a plan can carry unlimited budget reallocation and still escalate every targeting decision, and conflating the two dials is the most common misreading of this layer. Write scope is the dial a seller advertises. Review timing is the dial a buyer cannot prove to anyone. So when rung 4 appears below, read it as that second dial rather than as a wider write scope than rung 3: what changes at 4 is not what the agent may touch but that nobody looks until the spend has landed.

Picking a rung

Three artifacts decide the moves up, and not one of them is a field you set on your own request. The move from 1 to 2 needs nothing from anyone else, which is why it is the only one you can make unilaterally. From 2 to 3 you need a tolerance document: the numeric edges of a seller’s conditional_self_serve envelope, which AdCP leaves out-of-band until issue #4425 lands. From 3 to 4 you need a cancellation policy, which carries notice_period and cancellation_fee, both required, the fee in one of four types and charged “when the notice period is not met” — that document is what prices the undo once the undo stops being a keystroke. From 4 to 5 you need a declared aggregation window: aggregation_window_days, the trailing period over which a governance agent sums committed spend before testing any dollar threshold. The first two are documents you obtain; the third is a declaration you read off the governance agent’s get_adcp_capabilities response.

Start at 1. Move up only when the next rung’s entry condition — a tolerance document, a cancellation policy that is not full_commitment, a declared aggregation window — is true of the specific buy in front of you.

Then adjust for the account. Duluth Trading hands bidding and creative iteration management to agents through its agency, and Ellie Uberto, its director of marketing, reports being comfortable with it. The risk profile of a bid adjustment on a retail account is not that of a pharmaceutical campaign at Bayer, whose digital media activation lead told Digiday’s Programmatic Marketing Summit in May 2026, “I want a person overseeing the bot.” That is Glenniss Richards, senior director of digital media activation.

Rung 1: read-only is where nearly everyone is

A wrong read costs you a correction. That’s the whole risk, and it is where every practitioner in that thread reported shipping. Own_Corner1016, posting for TeqBlaze: “Anything involving pacing, budgets, or major decisions - still remains fully human.” When a small-business owner posted an open-source plugin letting Claude reallocate budget through the Google Ads API, the first reply was one line: “The official Google ads mcp server is read only.” One line from one reply is the closest thing there is to a public statement of where a large platform put its own connector.

Weigh it accordingly: the r/adops thread scored 6 with nine comments, four of them from two accounts, and the Google line is a six-upvote reply from u/assassinofkings316 rather than anything Google published.

AdCP has no read-only mode for production data. account.sandbox looks like one and does something else, marking “a sandbox account — no real platform calls, no real spend,” a separate test account returning fake delivery. Rung 1 is pulling real delivery and alerting on real anomalies, so sandbox cannot express it.

What holds this rung is out-of-band. Your credential is scoped to discovery and reporting calls, and nothing in the request or the response says so. The one machine-readable hint is MCP’s readOnlyHint annotation, and on 12 August 2026 three of the thirteen listed agents that answered tools/list set it: goTom, No Fluff Advisory’s sales agent, vastlint. goTom annotates get_products true and create_media_buy false, which is the shape a client can act on. Four of the other ten expose create_media_buy with no annotation at all, so against those a client cannot tell a read from a write by inspection.

Rung 2: propose the diff, spend on a second call

A tool author in the same thread put up a proposal and asked practitioners to shoot it down:

i’d automate reads, normalization, anomaly checks and draft actions. for spend-changing writes, i’d let the agent prepare the exact before/after but require a second call with identical args. that’s the line i’m testing in adport; i’m the author and would love adops people to tell me where the approval gets too annoying.

The shape is right. Approving “shift 15% from CTV to display” is not the same act as approving the twelve-field object that shift turns into, and a second call with identical arguments makes the human’s approval refer to the object. Every buy with a spend-changing write in scope belongs here, whatever anyone’s maturity model says, and it costs you one extra call.

AdCP doesn’t give you this. create_media_buy has no dry_run; at 3.1.13 the flag ships on three sync operations, sync_catalogs (which is inventory housekeeping), sync_accounts and sync_creatives, none of which is the call that spends money. The nearest affordance is paused: true on create_media_buy, which commits you contractually and then withholds impressions.

Issue #6381 on the AdCP repository, filed 11 August 2026 by pkras, states the gap exactly:

Today the only way to honor this is inside the buyer’s own agent orchestration — simply not calling create_media_buy until a human clicks approve. That gate is invisible to the counterparty, can’t be verified by a heterogeneous partner, and leaves no shared record.

Rung 2 is real, widely intended, and lives entirely off the wire: your approval log and your seller’s booking log are disconnected silos.

The RFC narrows its own complaint. A signed compact JWS governance_context already rides on core/protocol-envelope.json carrying sub, plan_hash, authorized_commitment, policy_decisions and audit_log_pointer, and in 3.1 all sellers MUST verify it. What that receipt does not say is that a person cleared the spend, so the issue calls the remaining gap narrow and proposes a human_approval claim with an opaque approver_ref, minted by the governance agent because “Buyers MUST NOT construct, modify, or re-sign the token”. It is open, with one comment and a needs-wg-review label: a proposal, not a direction.

AAMP’s buyer agent ships rung 2 in code, and that reference repository is the only shipped evidence anyone has. At commit 0c3a1545 it holds three guards that complete in three different directions, each on a different rung: an approval timeout here on rung 2, a pacing hold on rung 4, a spend ceiling on rung 5. The first fails closed: wait_for_approval times out after 3600 seconds and returns approved=False, timed_out=True, so an unseen request becomes a rejection and the campaign stalls rather than errors, which means it stops without paging anyone. ApprovalConfig in models/campaign_brief.py configures that same guard rather than adding a second one, with four booleans: plan_review and booking default to True, creative and pacing_adjustment to False. Plan and spend gated, creative and pacing not, which is close to what the thread describes.

Those defaults answer the tool author’s question about where approval gets too annoying, and the answer is milder than the fear. Pacing adjustment, the high-frequency stage, is ungated, so approval volume scales with deals booked rather than with days on flight: a ten-line plan costs a couple of clicks at the front and nothing after until somebody changes the plan.

Check two things before citing the config as a standard. check_approval_required is called nowhere outside the approval module. flows/deal_booking_flow.py has its own hard-coded stop at AWAITING_APPROVAL plus an approve_all() method, so the booking path does not consult the config to decide it should pause. And flow_state.BookingState.campaign_brief is typed dict[str, Any] while flow_state.CampaignBrief, the nine-field class that flow describes, carries no approval_config at all; the gate reads its configuration out of a campaign-store row instead. No gate setting on the money path is enforceable by type. The repository has two brief classes and one gate config, and the config is attached to the class the booking flow does not use.

A comment in consolidate_recommendations records what that cost once. Under an or_ trigger, CrewAI 1.14 treated the five research methods as a racing group and fired the listener a single time, so a fast no-budget channel triggered consolidation while a funded channel was still working, pending_approvals stayed empty, and “approval silently approved nothing and jobs ‘completed’ unbooked”. An approval gate reporting success over an empty set is the failure rung 2 exists to prevent, and it happened in the IAB Tech Lab reference implementation.

Rung 3: the mode is on the wire, and the wire says it is advisory

This is the rung the market is selling: the agent operates freely inside a pre-agreed boundary and escalates outside it. AdCP models it properly, giving each action a seller allows on a buy a mode from a three-value enum. The column that matters is the third.

ModeWhat the seller does with your mutationCan you predict the outcome before you call?
self_serveHonours it synchronously, no approvalYes
conditional_self_serveAuto-approves inside declared tolerances, escalates outside themNo. The tolerances are not on this surface
requires_approvalPuts a human on its own side in the loop, asynchronous, resolved by poll or webhookYes, in the sense that the answer is always “wait”

The middle value is where programmatic guaranteed lives, and the mode description names FreeWheel, Magnite and GAM as the platforms “where small mutations clear automatically but large ones queue for human review.” The same description then concedes the problem:

Constraint metadata defining the tolerances is out of scope for v1 […] until #4425 lands, tolerances are declared out-of-band and buyers cannot statically predict which mutations will auto-approve from this surface alone.

An envelope whose edges you can’t read is a surprise with paperwork on it, and your agent finds the edge by trying, in production, on a live guaranteed buy. So the entry condition for this rung is a document rather than a field: the seller’s conditional_self_serve tolerances, in numbers, in writing. If nobody has sent you one, you are on rung 4 with extra steps.

Two other clauses decide who holds the lever. On the product template a seller may publish several modes[] for one action, and “SDKs that see multiple modes MUST NOT assume which one will fire”; the singular resolved mode on the buy is the one to read, and that is the value the schema calls advisory. Recovery from mode_mismatch “is a flow switch, not a retry against the same task”. Repricing is not modelled as a mode at all: a seller whose quote no longer covers your update returns REQUOTE_REQUIRED. A publisher can advertise self-serve-within-tolerance, escalate at will, reprice outside the envelope, and stay conformant throughout.

The third mode deserves more than the joke, because publisher-direct sellers will live on it. requires_approval is “human-in-the-loop, asynchronous, no proposal artifact,” and the call returns status: 'submitted' with a task_id covering “long-running execution (hours to days).” A seller can put a clock on that: core/sla-window.json carries response_max and completion_max as ISO 8601 durations, and response_max is defined mode by mode, including “queue ack for requires_approval”. Both are optional, absence “means no commitment, not zero commitment”, and enums/task-status.json has nine values and no expiry among them, so nothing terminates a request nobody opens. The clock that decides when a pending approval has gone stale has to live in your own orchestration, and somebody has to be paid to watch it.

The buyer-side half is better specified. A campaign governance plan pushed through sync_plans carries budget.reallocation_threshold, the “amount above which budget reallocations require human escalation”, movable “up to this threshold per change without asking a human” and settable “to 0 to require approval for every reallocation”. A numeric autonomy dial denominated in the plan’s currency, and the single field here I would most want an agency to fill in deliberately. Its mutually exclusive partner, reallocation_unlimited, is rung 5’s field. Eight of the 13 files under dist/schemas/3.1.13/governance/ carry "x-status": "experimental", and those eight are the request and response pair for each of the four callable governance tasks — sync_plans, check_governance, report_plan_outcome, get_plan_audit_logs. The five that carry no such marker are shared definition files nobody calls. So every wire path this dial travels is pre-release. Pin the release you build against, keep the governance calls behind an adapter you own, and price the rewrite: AdCP’s versioning permits field-level schema change between releases, and these are the files it is most likely to change.

AdCP reads “per change” literally. An agent under a reallocation_threshold of 5,000 moves 4,999 as often as it likes. The defence sits on a different surface: get_adcp_capabilities lets a governance agent declare an aggregation_window_days, and without it “a buyer can split a single large spend into many sub-threshold commits across plans / task surfaces / time and bypass every dollar-gated escalation,” with absence meaning buyers “MUST assume per-commit evaluation only.” Ask for the declared window before you pick a threshold, because against an agent that declares none the number you set caps one change and nothing else. Then ask what it is keyed on. Aggregation “is keyed on (buyer_agent, seller_agent, account_id)”, so one buyer running two buying agents, or two account ids, sits outside a window that was genuinely declared.

AAMP’s envelope is buyer-local by construction: DealPreferences.max_cpm, a FrequencyCap of max_impressions and period_hours, a PacingModel with four values (EVEN, FRONT_LOADED, BACK_LOADED, CUSTOM). None of them is asserted to the seller.

Rung 4: spend first, review after

AdCP’s governance mode enum is the cleanest description of this rung in either corpus. audit always returns approved, advisory returns findings but does not block (“Human reviews happen post-hoc”), enforce blocks. Escalation severity sits alongside it with three values, of which two gate behaviour: warning means “agent may proceed but human should review within a deadline,” critical means “agent must not proceed until human approves.” The third, info, is “logged for audit, no action required.”

Neither enum is a control a buyer can set. Across the 3.1.13 release, enums/governance-mode.json is referenced by check-governance-response.json and get-plan-audit-logs-response.json, enums/escalation-severity.json by those two plus report-plan-outcome-response.json, and by no request schema at all. The response says what the mode is for: letting “counterparties, regulators, and auditors distinguish whether a finding blocked execution (enforce) or was logged silently (audit).” You configure the mode inside your own governance agent, and the wire reports which one was in force once it no longer matters.

What makes rung 4 expensive is the contract underneath it, not the protocol. cancellation-policy.json requires both a notice_period and a cancellation_fee, and the fee is “applied when the notice period is not met” in one of four types, of which full_commitment is the worst: “buyer owes the full committed budget regardless of delivery.” Post-hoc review of a guaranteed buy means reviewing something you may have to pay for either way, so the rung is only defensible on a buy whose cancellation policy is something other than full_commitment. Price the miss as well as the type, because a cancellation inside the notice window costs nothing under any of the four and the number that decides which side you land on is your own detection latency. The exit is one-way regardless: update_media_buy.canceled reads “Cancellation is irreversible — canceled media buys cannot be reactivated”, and sellers “MAY reject with NOT_CANCELLABLE”.

Henry Webster of Kelly Scott Madison put the fear in the form every agency lead will recognise, at Digiday’s Programmatic Marketing Summit in May 2026: “Would it blow a quarter’s worth of budget in a weekend?” Webster is the agency’s SVP director of analytics and insight.

The nearest thing to an answer is thinner than it looks, and it is not in AdCP. IAB Tech Lab’s buyer agent declares a campaign state called PACING_HOLD, distinct from a manual PAUSED, with legal transitions ACTIVE → PACING_HOLD labelled “automated pacing deviation threshold” and back again on “deviation resolved, auto-resume.” Separately, pacing/engine.py measures deviation at 10% for a warning and 25% for critical in both directions, and on breach emits a PACING_DEVIATION_DETECTED event and returns a PacingAlert. Nothing joins the two. Outside models/state_machine.py and a status-string map in storage/campaign_store.py, PACING_HOLD appears nowhere in the source, no code performs the transition, and the pacing guide hedges to match by saying the state machine “can transition” the campaign. The industry’s strongest automated brake on this rung is an unwired enum sitting next to a detector that files a report, and 25% is a critical alert threshold rather than a hold trigger. Guard two of three: declared, legal, never performed.

Post-hoc review itself has no AAMP mechanism, with no governance mode and no escalation severity anywhere in the corpus. Rungs 4 and 5 are an AdCP-only conversation. What an AAMP buyer gets instead is DecisionRecord, an audit object carrying a DecisionActor whose kind discriminates human from machine, plus rationale, policy refs and a money_effect. It records who decided and why, after the fact, and it is never embedded on a Deal, Order or Quote, so it documents the spend without being able to interrupt any of it.

Rung 5: unattended, inside a cap

AdCP has a field for declaring this on purpose, and the file recommending it also warns against the alternative it recommends two lines earlier. reallocation_threshold says “Set equal to total for effectively unlimited reallocation.” The next property, reallocation_unlimited, says “Use this for deliberate full-autonomy declarations rather than setting reallocation_threshold: total (which silently tightens when total changes).” A buyer reading sync-plans-request.json top to bottom takes the advice the following paragraph withdraws. Someone thought about the failure mode where a budget cut quietly revokes an autonomy grant.

plan.human_review_required is orthogonal to the money dial and the buyer can’t switch it off. When a resolved policy carries requires_human_review the governance agent must set the flag: “A buyer cannot opt out of human review by omitting the flag.” It exists for GDPR Article 22 and EU AI Act Annex III verticals, which is why the money dial and the review dial move independently of each other.

The third AAMP guard fails open, and says so in its own docstring. booking/spend_ceiling.py is a deterministic cap called at both money-committing sites, before a Deal ID is minted and before booked lines are created, with “no LLM involvement, no configuration flag to disable it”. It exists because max_cpm previously “reached the selection LLM only as prose and the seller only as an advisory search filter, so a seller quoting far above the buyer’s ceiling would still be issued a Deal ID.” Supply no limit and the check is skipped with a warning logged, which the module calls “an explicit choice to preserve current demo behavior”. PacingConfig.max_reallocation_pct caps a single reallocation proposal at 30% of total budget. Both fit rung 5’s definition, neither is asserted to the seller, so no counterparty can confirm either exists.

On the sell side the cap is set by a grade somebody else issues. Any publisher deploying IAB Tech Lab’s seller agent unmodified has put price negotiation on rung 5 while their buy-side counterpart is still arguing about rung 2. Its negotiation engine accepts, counters or walks away with no human in the path, floored by floor_price and bounded by per-tier limits that run from three rounds and an 8% cumulative concession for a public buyer to six rounds and 20% for an advertiser-tier one. start_negotiation reads buyer_context.effective_tier, maps it through TIER_STRATEGY_MAP and looks up STRATEGY_LIMITS, and that tier derives from a trust status the AAMP agent registry issues. Four tiers, four concession budgets, graded by a third party.

What a live sales agent publishes, measured 16 August 2026

Schemas describe what a seller may implement. Cora AI, at sales-agent.coraai.org/mcp, is one of three AdCP sales agents that returned a product catalogue to an anonymous caller in the 12 August 2026 registry sweep, which makes its buying surface readable by a stranger. The method is four MCP calls, ending in a get_products carrying a US connected-TV brief.

That sweep produced a verdict about rung 1. The agent returned four products in a 25 KB payload, and the strings allowed_actions, available_actions, cancellation_policy and conditional_self_serve appear in it zero times. Two of those four are properties of core/product.json: allowed_actions and cancellation_policy, both optional on a schema that requires 7 of its 49 properties, so the agent is conformant. The other two could not have appeared in a product payload at all — available_actions[] resolves on the buy response, and conditional_self_serve is a value of the action-mode enum rather than a field — and their absence matters only because allowed_actions[], the product-level template where a seller advertises which actions and modes it will allow, is the thing that never arrived. The mode template that would place this agent on rung 3 and the cancellation terms that would price rung 4 are simply not published, which puts rung 1 at the ceiling of what a buyer can verify here.

Cora AI’s AAO Verified badges name governance and media-buy at 3.0, issued 25 June 2026, and the table below scores the agent against 3.1.13. Those are not the same yardstick, and the badge is the weaker one: it names a major version where AdCP’s own adcp_version is VERSION.RELEASE, as in 3.1.13, so a buyer reading the badge does not learn which release’s fields to expect. Score the absences against the 3.0 line instead and only one of them is excused. create_media_buy.paused arrives with 3.1; plan_id on create_media_buy, revision on update_media_buy and the five required fields are present in every 3.0 release from 3.0.1 to 3.0.18.

Its tool surface, re-probed on 16 August 2026, repeats the pattern. All 24 tools are published, none carries an MCP annotation of any kind, and three fields this ladder depends on are absent.

FieldIn AdCP 3.1.13On the live agent
create_media_buy.pausedPresentAbsent
create_media_buy.plan_idPresent, required when the account has governance agentsAbsent
update_media_buy.revisionPresent, optional, optimistic concurrencyAbsent
create_media_buy required fields50

Of the 24 tools it publishes, one carries a governance surface. sync_governance binds a governance agent to an account, so it is the wire path every control on rungs 3 through 5 arrives by. Here it takes accounts as an array of bare untyped objects with no inner structure published at all, and requires only accounts where 3.1.13 also requires idempotency_key, so the deduplication key that stops one retry firing two approval flows is optional.

Elsewhere in the same probe the agent publishes populated schemas: 20 properties on create_media_buy, 18 on get_products, 9 on update_media_buy with media_buy_id the only required one. Read the 20 carefully, because it is not the release’s 20. AdCP 3.1.13’s create_media_buy request carries 20 properties of its own, and adcpexplorer.com prints 22 for the release because it counts the envelope’s adcp_version and adcp_major_version alongside them. The agent publishes neither version field, and a dozen of its 20 — advertiser, advertiser_name, budget, currency, flight_start, flight_end, product_ids, measurement_terms, signal_targeting, signal_targeting_groups, reporting_dimensions, brief — are not in the 3.1.13 request schema at all. Matching totals, different surfaces, which is how a schema can be the same size as the spec while missing paused and plan_id. get_products does the same thing at 18 against 18, five fields on each side that the other does not have.

docs/protocol/calling-an-agent.mdx warns that some AdCP MCP servers publish no per-tool parameter schemas at all, where “every tool shows {type: 'object', properties: {}}”, and this is not one of those. The absences are chosen, and so is the empty required list, which is the worst line in the table. A create_media_buy that requires nothing accepts a call with no budget, no dates and no packages, and whatever happens next is decided somewhere the buyer cannot see.

The missing fields cost specific rungs. Booking paused in one call is gone, though update_media_buy does publish paused, so the move survives as create-then-pause: two calls with a live window between them and no revision on either, so nothing detects a second writer inside that window; sellers MUST reject on a revision mismatch, but only when the buyer sends one, and this schema has nowhere to put it. Without plan_id, nothing binds the buy to a governance plan, and that is how reallocation_threshold and human_review_required reach the seller at all.

Run the same check against whatever agent you are buying from: send initialize, then notifications/initialized, then tools/list, and read each inputSchema. Those schemas tell you which rung that seller can hold you to. Read the result as a ceiling, the way the table above reads it. No allowed_actions[] on the products and nothing declares a mode, so rung 1 is the top. No cancellation_policy and rung 4 has no price on it, so you cannot tell what the undo costs. No plan_id on create_media_buy and rungs 3 through 5 have no way to reach the seller at all, whatever your own governance agent is configured to do.

Why the dial belongs to the counterparty

Every mechanism above fails in the same direction, and what is left when they fail is whatever each counterparty happened to implement.

I trust the practitioners in that r/adops thread over the vendor framing, because agreeing the rules turns out to be the hard part and they are the only people describing it. TeqBlaze’s r/programmatic post offering free testing of its AdCP sales agent reports that “the challenge is getting teams to agree on the rules before the system runs. If goals or limits are vague, the agent simply scales that ambiguity faster.” A commenter answered: “most publisher teams can’t even agree on their floor prices lol. we spent 3 weeks just trying to get sales and ops to align on what a ‘good deal’ looked like.” No flair on that account, so read it as a practitioner claim rather than a publisher one.

AdCP’s own governance documentation is closer to those practitioners than its vendors are, and it has the same problem. Escalate when confidence is insufficient for the risk, it says, and check-governance-response.json carries confidence as a number from 0 to 1 on each finding, plus an uncertainty_reason “Present when confidence is below a governance-agent-defined threshold.” The string confidence occurs zero times in the request schema. You receive the score and the counterparty sets the bar.

There is a structural reason rungs 3 through 5 have no counterparty yet. sync_plans, check_governance, report_plan_outcome and get_plan_audit_logs are all Required on the governance agent role, while sync_governance on a sales agent is Conditional, “Required when the buyer uses campaign governance”. Every buyer-side control on this page lives on a governance agent, and the 23 listings in the public catalogue on 12 August 2026 were 18 sales, 3 signals, 1 creative and 1 buying. Zero governance agents. There is nothing in the catalogue to bind to.

So run the ladder against the counterparties you actually trade with, one at a time.

Eight questions for the vendor call

  1. Which rung does your default configuration sit on, and which field sets it? An answer that does not name a field is not an answer. (rung placement)
  2. If the agent proposes a spend-changing write, what exactly do I approve: a sentence, or the payload? (rung 2)
  3. Does my approval leave any record the seller can verify, or only a row in your database? Issue #6381 is why the honest answer is currently the second one. (rung 2)
  4. For conditional_self_serve actions, what are the tolerances, in numbers, in writing? (rung 3)
  5. On the inventory you would let the agent commit, what is the cancellation_policy: notice period, fee type? full_commitment means the undo costs the whole budget. (rung 4)
  6. What is your reallocation_threshold default, and if it is set equal to the plan total, what happens to it when the total is cut mid-flight? The schema’s own warning is that it silently tightens. (rung 5)
  7. What aggregation_window_days does your governance agent declare, and what is it keyed on? (rungs 3 and 5)
  8. Which AdCP release do you implement today, and what happens to my plan configuration when you upgrade to the next one? Eight of the 13 governance schema files at 3.1.13 are experimental, so the shapes on rungs 3 through 5 will move. (all rungs)

Question 4 decides which rung you are actually on, and the answer has to come back in numbers and in writing before the buy goes live.