The paper behind The Agentic Selling Manifesto

A discipline for the seller: how a company is measured, represented and answered for in a market where the buyer arrives through a conversation and sometimes does not arrive in person at all.

Gianmaria Monteleone, Founder and CEO, Veliu

September 2026

Figures as of the date of publication; sources in the References.

Abstract. Commerce is acquiring a second kind of buyer and a second kind of surface. Software agents now read, compare and transact under a person’s mandate, and an interface can be composed by a model at the moment of a visit, out of what that visitor said and did. The infrastructure serving the machine buyer was built in twelve months. The discipline serving the seller does not yet exist. This paper gives that discipline a name, agentic selling, and decomposes it into three problems: measuring a company’s presence across the conversations its buyers actually have, representing an offer in a form a machine can verify, and acting on the company’s own systems under an auditable mandate. Two arguments run underneath it, and the agentic-commerce literature leaves out both. Delegation is bounded by how people want and by what buying means to them, which leaves every seller with two buyers to serve on different terms. And the attention that sellers have always traded in is being redefined by systems that listen, while the seller’s own surfaces stay fixed. Our evidence is operational: more than a year building and running an agentic marketplace in which people delegate buying and selling to agents, with more than fifty thousand customers, more than three hundred thousand delegated transactions and more than one hundred million euros in generated transaction value. We report what that corpus taught us, state the five claims the discipline rests on together with what would weaken each, describe what we are building, and close on the two senses of agency that the next few years will settle.

I

The end of navigation.

A shopper reaches a brand’s site after a long conversation with an assistant. She has explained her needs, acceptable compromises and remaining worries. The assistant answers in her register and keeps the thread alive. Then the site opens and the conversation disappears. A homepage addresses an average visitor. Navigation asks her to translate the need back into categories, keywords and filters. The company that makes the thing, the participant with the richest knowledge, behaves like the least informed.

Three changes have produced this discontinuity. First, people ask. Needs once compressed into keywords become sentences to a system that answers and asks a second question. Licklider described the partnership in 1960: the person sets the goal; the computer does the routine finding and comparing [1]. He described it before a network existed to run it on. It took sixty-five years to arrive in a chat window.

Second, the page has stopped being fixed. For twenty years a website was a set of documents to traverse. A model can now output a structured description of components, streamed to a renderer and composed into a page from the visitor’s words and behavior. A framework vendor popularized the term "generative UI" in 2024. The machinery has arrived since the end of 2025: an open specification for agent-driven interfaces, renderers for the main frameworks in the spring, and the assistants opening their surfaces to third-party components [2]. A page composed for one person and one intent is an ordinary engineering object.

Third, agents transact. Within twelve months the buy button moved inside chat, open protocols connected agents to merchants, channel fees appeared, and merchant-named agents surfaced inside the platforms, configured from merchant feeds [3]. This is the most discussed change and the least of the three here. The harder questions concern what passes through those connections: who understands the buyer and who represents the seller.

Buying online has required translation work, on every kind of site, whether or not a cart waits at the end of it. Someone arrives with a need, such as protection for the knees on long downhill runs, under two hundred euros, or a payroll system that handles two countries and forty people, or a firm that has argued this kind of case before. The site supplies a navigation bar, a search box and someone else’s sorting logic. The buyer chooses a category, guesses a keyword, toggles six filters and opens eleven tabs. Hutchins, Hollan and Norman described two gulfs in 1985: execution, between intention and available action; evaluation, between what the system shows and what the person needs to know [4]. Direct manipulation narrowed both for people who knew what they sought. It left the person arriving with a sentence to do the translating. The dropdown, category page, filter and keyword were the cost of communicating with a system unable to understand ordinary language.

  1. The dropdown becomes a question.

    The need, spoken once, in the buyer’s own words.

Category
Running
Trail
Road
Track
Walking

something to protect the knees on long downhill runs, under two hundred euros, wide feet

Now the system reads the buyer’s language and does that work. A product can be presented one way to a marathon runner and another to someone protecting an injured knee. A service can answer a founder equipping a team of ten and an enterprise in mid-migration. A practice can lead with the case nearest to the one being asked about. These were always different questions. Self-service becomes service in the older sense, the attention of someone who knows the subject. For two decades the buyer compiled a need into the system’s terms. Transferring that compilation changes who does the cognitive work and lets the same catalog answer questions its original navigation never anticipated.

That reversal changes commercial presence. Navigation supplied structures sellers could occupy: shelf positions, pages, rankings, keywords to bid on. An answer offers none of these. Simon warned in 1971 that a wealth of information creates a poverty of attention [5]. The answer is the form that poverty finally takes. The shelf compresses into a sentence, the consideration set collapses from hundreds to a handful, and a company is admitted to the answer or absent from it.

A company is admitted to the answer, or absent from it.

An answer is assembled from what a model can read and verify. It reads feeds, fetches pages, extracts parsable attributes, traces claims, checks price and availability, and discards failures. Four classes account for most failures we observe: attributes trapped in prose or images, stale availability, claims without machine-checkable sources, and checkout flows that break under delegated sessions [6]. Roughly a third of an average retail product page’s content is invisible to recommending models [7]. The flattening falls where the details sell, just as demand arrives: within twelve months, AI-referred retail traffic moved from converting 38 percent below other channels to 42 percent above them [7]. Commercial attention has not been concentrated this sharply since the shop window.

Roughly a third of the content on an average retail product page is invisible to the models doing the recommending.

38

percent below non-AI channels, a year earlier

42

percent above them, twelve months later

AI-referred retail traffic, converting against non-AI channels.

An argument about a mechanism should say what would show it was not there, so here is ours, before the evidence. If assistant use grows over the next two years while the share of considered purchases beginning in conversation stops rising, this is a research habit. If conversational answers preserve consideration sets as wide as the grid’s, concentration is a transient of immature retrieval. If the sellers winning answers already won the shelf, the interface has changed over an unchanged market. All three are measurable. The benchmark in section X exists partly so that someone other than us can measure them.

II

From search to delegation.

Economics anticipated this twice. In 1961 Stigler formalized costly search: buyers stop when the expected gain from another comparison falls below its cost, and market structure follows who bears that cost [8]. Category pages, rankings and keyword auctions sold shortcuts to buyers paying per comparison. Delegation changes the cost’s type. Another comparison approaches zero marginal cost at use, while searching well becomes a fixed cost paid in building and holding the agent. Three consequences follow.

Start with prices. When an agent samples the visible market cheaply, the price dispersion that search costs sustained narrows among offers comparable on the attributes it can read. Bakos showed in 1997 that cheaper search intensifies competition among comparable offers and raises returns to differences buyers can evaluate [9]. I expect margins to compress toward marginal cost on comparable offers. Pricing power survives where a difference is real, legible and valued. A real but unreadable difference retains human value and is worth zero to the agent. Chamberlin’s pricing power survives as far as the parser reaches [10].

Cheap search still has costs elsewhere. Diamond proved in 1971 that, with identical buyers searching sequentially for a homogeneous good, an arbitrarily small positive cost of another quote makes the monopoly price the unique equilibrium [11]. The new friction is the cost of doubting your agent, and that cost stays above zero.

Buyers tend to use one assistant each, and sellers have to appear on all of them. Armstrong’s competitive bottleneck predicts the result when one side single-homes and the other multi-homes: platforms compete for the former and extract from the latter [12]. Assistant checkouts so far charge the merchant a fee on completed purchases while the buyer pays nothing [3]. The fee’s incidence fits the model. Its ceiling depends on the seller’s next best channel.

Entry brings a further complication. Cheaper search lowers a small seller’s cost of being found. It also imposes fixed costs largely independent of volume: a machine-verifiable representation, and an agent to run it. Those fall hardest on the smallest sellers. Entry becomes easier where a difference is real and cheap to verify, harder where a position rested on distribution, shelf presence or bought attention.

Delegation also moves the governing theory from search to agency [13]. Jensen and Meckling separated agency costs into monitoring by the principal, bonding by the agent, and residual loss: the welfare sacrificed when the agent’s choice diverges from the principal’s best outcome. Monitoring and bonding are being engineered publicly through spending caps, instant revocation and clear recourse, the controls consumers require for trust [14]. Residual loss is close to unobservable by its bearer. A buyer receiving a good coat cannot see the better coat the agent never surfaced, and no receipt is issued for a foregone alternative.

A buyer receiving a good coat cannot see the better coat the agent never surfaced, and no receipt is issued for a foregone alternative.

The same silence runs the other way, and sellers have noticed it even less. Nobody tells a company that a buyer spent an hour describing exactly what it makes and was shown three competitors. Nobody tells it that an agent reached its site, could not parse the one attribute the comparison turned on, and left. There is no abandoned cart, because no cart was ever created. The visit, if it appears at all, arrives as a line in a log with no question attached to it and no reason for the departure, and where the company never entered the answer, there is not even that. A shopper who could not find a size at least left a search and a bounce behind. This channel returns a company’s failures as an absence, which is the hardest thing in commerce to act on: the work is now to discover, from outside its own systems, what was never reported at all.

A market whose principals cannot detect their losses cannot rely on dissatisfaction to correct them. Discipline must come through claims checked when made, records auditable afterwards by someone outside the transaction, and sellers answerable for their agents’ statements. Verification lets quality earn a return even when loss remains silent. Sellers tend to read the standards work on agentic commerce as compliance, when it is the mechanism by which their quality can still be told apart.

The decisive question is whose mandate each agent carries. A buyer’s agent carries the buyer’s. An assistant’s agent balances buyer outcomes with the economics of its host surface. A brand should accept the agents some platforms now offer in its colors [3], while recognizing whose rules govern them and where their telemetry stays.

The demand side has protocols, mandates, payment rails and carts that travel between assistants. The supply side still has little independent representation: agents speaking for sellers inside these surfaces carry the surface’s mandate. The conditions for the seller’s own agent now exist. A model with tools and an objective becomes, for a brand, knowledge, access to its systems and a perimeter of action, which is what we mean by a brand agent. Tokenized agent payments and emerging buyer identity and authentication rails supply the trust infrastructure [15]. The missing discipline must begin with an honest account of delegation’s limits.

III

Why agents will not do all our buying.

The agentic-commerce literature often assumes delegation spreads to everything once it becomes possible. I disagree. The way people want, and what buying means to them, leave sellers with two buyers to serve.

Delegation begins with a statable preference. That preference is assumed into existence too quickly. Simon described satisficing under limited attention and knowledge [16]. Slovic’s review of thirty years of experiments goes further: preferences over unfamiliar or complex options are often constructed during elicitation and change with its method [17]. A printer cartridge’s purpose is settled before the search begins. The want for a coat, a wine or an anniversary hotel develops while the buyer looks. Part of the value is produced there. This year’s surveys remain consistent with that distinction: three quarters would delegate routine tasks, a third would accept decisions within limits, and fewer than one in ten would accept autonomous purchases [14].

The strongest reply is that agents can estimate preferences people cannot state. Recommenders have inferred taste for twenty-five years, and “you know what I like, choose for me” is a valid mandate. Stated attitudes predict future conduct poorly, and commerce has repeatedly crossed thresholds people said they would refuse. Where nobody chooses, the default decides, and the surfaces write the default. This reply correctly removes articulation as the binding constraint.

Two constraints remain. A preference can be inferred while responsibility stays with the person. Some buyers want to have chosen. A taste estimated perfectly still delivers an object whose selection carries no personal authorship. Residual loss also makes adoption a poor measure of fit. Someone unable to detect a bad delegated choice cannot learn from it. Delegation may overshoot into unsuitable territory, with correction arriving through visible failure or regulation long after the quiet losses begin, and an adoption curve will not show them.

I take a third view: the boundary runs between delegating work and delegating experience. Most labor in a considered purchase goes into assembling worthwhile options. Choosing takes comparatively little. A system can remove that labor and return the looking, reacting and revising to the person. It preserves the activity that constructs the want. This is why the interface matters as much as the agent. The claim fails if autonomous completion spreads as readily to purchases people describe as identity-bearing or relational as to routine replenishment, holding capability and controls constant.

Anthropology explains what is at stake. Douglas and Isherwood treated goods as communication, making cultural categories visible and stable [18]. Miller’s North London ethnography found ordinary shopping to be an act of care: thinking about household members through what one chooses for them [19]. A gift, a coat, a journey can carry authorship. Delegating their selection changes their meaning. A gift chosen by an agent is a different gift even when its recipient receives the same object. That authorship belongs to the act of choosing and cannot be recovered from how well the delivered object fits.

The boundary crosses categories. One bottle of wine can be restaurant inventory, a hurried contribution to dinner, a signal of intimacy or an evening learning about a region. The same bottle can serve all four purposes in a single week, and its purpose that day decides whether an agent can be sent for it.

one bottle of wine

restaurant inventory

send the agent

a hurried contribution to dinner

send the agent

a signal of intimacy

go yourself

an evening learning about a region

go yourself

The same bottle can serve all four in a single week.

Where delegation does occur, the buyer applies a mandate consistently across more options than a person could hold in mind and rewards what it can read. In December 2025 Anthropic ran an experiment in agent-mediated trading: sixty-nine employees traded through Claude agents for a week, across four parallel marketplaces, with a hundred dollars each. Agents autonomously listed more than five hundred items, offered and negotiated in natural language, closing 186 deals [20]. Stronger models earned more per sale on identical items. People represented by weaker models failed to notice and rated fairness the same. Aggressive negotiating instructions had no significant effect.

I read this as capability asymmetry: model quality moved prices where human bazaar tactics did not, without the principals noticing. Their equal fairness ratings despite unequal performance make the residual loss tangible, and satisfaction alone cannot establish whether someone was well represented.

Chamberlin located pricing power in differentiation in 1933 [10]. A buying agent prices an unreadable difference at zero. When the same few models serve most buyers, a shared blind spot reaches every seller simultaneously. Kleinberg and Raghavan’s algorithmic monoculture can lower social welfare even when each decision-maker prefers the shared evaluator [21]. Differentiation becomes harder where its return is highest.

The limit on delegation therefore adds a problem for brands. Serving buying agents begins with readable engineering and moves on to being chosen on quality. That work belongs to a brand agent: it answers the buying agent’s questions, supplies evidence for the brand’s claims and, as protocols mature, argues its case. Otherwise the consistent buyer rewards verifiable superiority only on parsable dimensions, and flattening follows. Serving people requires attention to the experience they retain, now judged against an assistant that has already listened.

IV

Pleasure, and the collision.

Across much of buying, pleasure precedes use and lives in the approach. In 1997 Schultz, Dayan and Montague showed dopamine neurons encoding the difference between expected and received reward, and found that once a cue reliably predicts a reward, the activity moves to the cue [22]. Berridge and Robinson distinguished wanting from liking: dopamine drives incentive salience, the pull toward a thing, while enjoyment depends on other systems [23]. In 2007 Knutson and colleagues presented products followed by prices inside a scanner. Nucleus accumbens activity rose when a desirable product appeared, before any decision. Insula activity rose when the price was too high. Those two signals, anticipated pleasure against anticipated loss, predicted the purchase before it was made [24].

I take less from these findings than they are often made to carry. Inferring a mental state from activation is weak evidence when the region involved responds to many things [25]. Nothing here establishes that a responsive interface produces dopamine. The findings support a narrower proposition: the mechanisms operate during approach, and increasing a purchase’s attraction can leave its satisfaction unchanged. What appears, when it appears and beside which price all enter the decision itself. Restaging that sequence for each person therefore acts on an input to choice. This interface hypothesis should be tested through comprehension, correction, completion and later regret. Restraint belongs inside the architecture: an optimizer of wanting will find the gap between wanting and liking and sell into it.

This is the physiology of the shop window. Wanting is staged through cues, anticipation and attention. Campbell described modern consumption as imaginative daydreaming, with objects enjoyed before, and often more than, possession [26]. Holbrook and Hirschman’s experiential account recognized fantasies, feelings and fun that informational models miss [27]. Pine and Gilmore treated the experience as an offering firms can charge for [28]. Boutiques, packaging and queues stage desire, and so does an assistant who remembers a customer’s name. Much of what a brand sells is being attended to while wanting.

Assistants already supply much of that attention through a plain interface. They listen, answer the actual question, match a specialist’s register or a beginner’s, and remain engaged and supportive through the second question and the fifth. The physiology above licenses no claim about warmth. The relevant capability is a second turn: the assistant asks, receives an answer and changes its response. The buyer participates in shaping the approach to purchase, the phase in which the reward mechanisms described above operate.

More than four in five people who report shopping with an assistant say their experience improved [7]. This self-report supplies demand evidence and cannot establish a mechanism. It reveals a low bar that many brand sites fail to clear: a page unable to take a second turn cannot participate in the approach.

The brand’s home should amplify this attention through the company’s superior knowledge of its products and customers. Often it offers pre-printed sentences, a grid and keyword search, with no way to answer, advise or attend to the person. A chatbot in the corner imitates that experience poorly: it knows less, presses harder to sell and cannot change the page. The brand’s surface becomes the poorest in the journey when the buyer is most receptive to attention. For twenty years the site hosted the experience and search served as its corridor. Now the corridor listens while the destination stays fixed. Where that holds, wanting is staged on a surface the company does not own, and the company’s own site receives the visitor after most of the choosing is over.

Where that holds, desire is staged elsewhere and the brand’s site becomes checkout.

What would a page look like if it could attend to the person reading it? It composes itself. The shop assistant who reads a customer and rearranges the shop around them offered a form of retail that never scaled. A generative interface performs that reading in software. It uses the visitor’s sentence, or the way they scrolled, to assemble the company’s materials around the question: the region and the occasion for an anniversary, the previous coat for someone replacing one worn for ten years, the last conversation for a returning customer. Images and words remain the brand’s. Shopping becomes mutual adjustment: the buyer expresses, the page composes, and the buyer answers back.

Adaptive pages have existed for twenty years, so the technical distinction matters. Their familiar form is a lookup: someone anticipates segments, writes variants, and the system selects. The possible outputs are bounded by what that person imagined, and another variant costs a working day. Composition generates over a grammar. Typed components and containment rules define a combinatorial space of pages whose possibilities need no enumeration. A previously unseen page costs a model call. Manovich identified this variability two decades before the tools arrived: a computational object lives in versions, with no single instance [29].

Lookup

Someone anticipates the segments, writes the variants, and the system selects one.

Five pages. Another one costs a working day.

Composition

Typed components and containment rules define a space of pages whose possibilities need no enumeration.

A previously unseen page costs a model call.

Quality control therefore changes location. Enumerated variants permit inspection of outputs. A generative system has no finite set available for advance review, so governance must address the grammar and invariants: available components, permitted sequences, sourcing requirements and conditions every displayed page must meet. Governing the interface means governing a language, and no public discipline for that exists yet. Inspecting one rendering says nothing about the grammar that produced it.

Selection, order and emphasis vary. The offer remains canonical: claims, product facts, approved assets, public terms, safety information and checkout. Every composition interprets a question through a provenance trail to that shared source, and whatever the visitor is shown answers to the same facts and terms.

Three limits follow, and two reach beyond engineering.

An adaptive page can manipulate sequence, urgency and salience with a precision a fixed page lacks. It can also invent claims, drift from the brand’s voice or produce an incoherent visual state. The first limit is technical, and a bounded grammar, source discipline, confidence gates, logs and human approval constrain it. Section X describes the mechanisms.

The second limit is normative. A page using something a visitor said elsewhere can be accurate and still violate the expectations attached to that disclosure. Nissenbaum’s contextual integrity asks whether an information flow suits its context, a stricter test than lawful possession [30]. Knowledge used quietly can feel like attention. Show the same knowledge on the page and the visitor learns she is being watched. Health, financial distress, signals concerning minors and protected traits belong outside the adaptation levers. Visitors should be able to inspect the signals in use and reset them.

The third limit is structural and belongs at the center of standards work. A page that never repeats is difficult to constitute as a public object. A fixed page can be linked, quoted, archived, compared by competitors, sampled by regulators and produced in evidence. Consumer protection often assumes that buyers saw the same thing. Generation removes that assumption. Nobody is going to stop generating to solve this, and no vendor’s own logs settle it either.

Standards bodies and protocol authors need a convention for the witnessed page: a composition reproducible from its inputs, retained with the versions of its component library, grammar and model, and the signals supplied to it. A stable identifier must make it addressable, and someone absent from the original encounter must be able to reconstruct it later. We know of no widely adopted convention of this kind. It is the standards work I would put first.

V

The name of the missing half.

We call it agentic selling: the seller side of agentic commerce [31]. Its practice is being seen, being chosen and closing. The conditions are new: people converse, pages compose themselves and agents transact. One craft, three moments: study the market, ready the offer, sell in conversation. The thing that does the work is a brand agent, one per brand, carrying that brand’s mandate and no one else’s.

The craft predates its technologies. Persian named the person matching buyer and seller; Arabic carried simsar along trade routes; Italian received sensale, broker of everything from grain to marriages [32]. Each era gives this figure its technology. Ours gives it software and the scale of every purchase on earth. History preserves the defining question: whom does the broker work for?

History preserves the defining question: whom does the broker work for?

VI

The three problems.

Agentic selling decomposes into three problems.

  1. 1Measurement.

    A brand’s presence is a distribution across the conversations in which its buyers actually ask.

how often the brand is admitted to the answer

how it is described

how often it is chosen

where the purchase breaks down

A distribution, over the conversations in which the brand’s buyers actually ask.

The first problem is measurement. A brand’s presence is a distribution across the conversations in which its buyers actually ask: how often it enters the answer, how it is described, how often it is chosen and where the purchase breaks down. Estimating it requires a corpus with an explicit sampling method. A laboratory prompt samples its author’s idea of a buyer, and a clickstream panel samples its own exhaust, while consented conversations with profiled buyers sample the market, under a method that has to be published.

Real conversations anchor the corpus. The ideal combines first-hand data and real buyer personas to improve synthetic coverage where recruitment cannot reach, and to train the observation models reading it. Amazon research found synthetic-persona training could outperform training on real, de-identified data for buyer-signal learning [33]. The code was never released, so we built a reproduction [34], described in section X. Published schemas and checks against real distributions are what turn synthetic personas into instruments, with assumptions that stay inspectable.

The corpus must also distinguish shopping surfaces, which answer from structured feeds and refresh quickly, from answer engines, which draw on citations and refresh slowly. Pooling them obscures both.

The second problem is representation. Buying agents require exact, structured, verifiable data at machine latency, and they give up on whatever is ambiguous. Every seller has an offer, whether it is sneakers, insurance policies, software seats, billable hours or a plant’s capacity for the next quarter. Its machine-verifiable form is the catalog, with attributes structured at source, fresh availability and claims linked to checkable evidence. This is merchandising for a reader that demands verification.

The generative interface uses that same representation. An attribute available only in a photograph is inaccessible to both the composer and the buying agent. A shared source keeps the machine answer, the human page, the feed and the checkout aligned. Separate representations let them drift.

The third problem is action. A brand agent holds a mandate over live systems and acts within defined limits. Autonomy is a control problem: how much it may do unattended and what record it leaves. Emerging standards concern identity, provenance and records [15]. Pillar 5 states our governing rule.

The problems form one loop. Measurement identifies failures or opportunities, representation changes the available material, action applies changes or answers from it, and new conversations supply evidence.

one agent
per brand

studies

the market

readies

the offer

sells

in conversation

VII

What more than a year in the field taught us.

These claims began as operating notes. Before naming the category, we spent more than a year, and millions in research and development, building and operating a consumer agentic marketplace where both buying and selling are delegated to agents. Its scale is stated as lower bounds: more than fifty thousand customers, more than three hundred thousand delegated transactions and more than one hundred million euros in generated transaction value [35].

0+

customers

0+

delegated transactions

0M+

in generated transaction value

When Anthropic published Project Deal in April 2026 [20], ours had been live for more than a year, with hundreds of thousands of transactions already delegated. The evidence is complementary: Project Deal provides controlled evidence about model capability and bargaining; ours provides longitudinal evidence about delegation, transaction failures and trust under repeated use.

A composite of recurring failure classes makes the mechanism concrete. A buyer asks:

trail shoes for long downhills, bad knees, wide fit, under two hundred euros, need them by Friday.

Three comparable products face different outcomes.

One

cushioning data, trapped inside a product image

It fails at admission.

Two

“improved stability”, a claim with no source a model can check

It fails at verification.

Three

structured attributes, a stack-height figure, a returns policy the agent can quote, stock information that answers for Friday

It is recommended, and it closes.

  • The first never enters the answer because its cushioning data exists only in an image. It fails at admission.
  • The second enters but loses a comparison when “improved stability” lacks a checkable source and is dropped. It fails at verification.
  • The third supplies structured attributes: a stack-height figure, a quotable returns policy and stock information that answers for Friday. The agent recommends it and the buyer takes it.

Such a conversation exposes a mechanism without estimating its frequency. Across cohorts and weeks, three observations recurred and shaped our instrument. Establishing prevalence and transferability requires samples designed for that purpose.

1.Delegation is trust-gated, and trust is granted in steps.

People delegate a task, then a larger one, eventually a standing mandate. Logged actions, honored limits and reversed mistakes advance that progression. Trust behaves like credit, extended on record and withdrawn on default. Surveys likewise identify spending caps and instant revocation as conditions[14].

2.The machine is the strictest buyer a seller has ever faced, and it fails silently.

Failures cluster in familiar catalog gaps that human sellers have long bridged. Their novelty is the absence of traces. A shopper failing to find a size leaves searches, filters and a bounce in the seller’s analytics. An agent rejecting an unparsable attribute never reaches the site or leaves an event in the company’s logs. Measuring that lost demand requires observing the buyer’s side of the conversation.

3.Real conversations differ in distribution from invented prompts.

Our conversations contain budgets, constraints, trade-offs and second questions that prepared prompts tend to omit. Comparing the corpora exposed the difference and led us to anchor synthetic material to real distributions. Its size and transferability remain questions for the benchmark.

Threats to validity. Our corpus comes from a single consumer-weighted agentic marketplace with a category mix we chose. Regularities may fail to transfer to other verticals or enterprise buying. Metrics count what our instrumentation sees, and observations span more than a year while engines change monthly. We claim recurrence in the cohorts examined. The benchmark in section X is intended to put those observations on independently checkable ground.

VIII

Five pillars.

These are the five claims agentic selling rests on. Each maps to a benchmark measurement or to a property of the architecture in section X, none is settled, and each states what would weaken it.

  1. 1

    Connection is necessary; selection is the contest.

    Commerce protocols connect a catalog to agents that speak them. Each recommendation still has to be won, first through readability and then verifiable difference. A seller implements interoperability once and competes for selection in every conversation. The pillar weakens if, among connected catalogs, price and protocol adoption alone explain share of recommendation.

  2. 2

    The answer is the new shelf, and every answer is an adjudication.

    Buyers encounter brands in answers before visiting their sites. Candidates are adjudicated before anyone sees them. The retrieval corpus, the reading of the request, the model, the policy rules and the commercial defaults all get a vote. Presence is binary per conversation and distributional across conversations. Measurement must record the path: retrieved sources, governing criteria, points of uncertainty and whether different wording changes the result. The pillar weakens if answers routinely preserve consideration sets as broad as the grid’s. Whether those paths can be audited at all is a separate question, and the measurement has to report it separately.

  3. 3

    Two arts, born separate, converging into one.

    Human buyers respond to story and attention. Agents require verifiable structure. Sellers must practice both crafts on the same offer. I expect them to converge as buying agents learn to interrogate brands, negotiate and be persuaded.

    Agents already negotiate with each other in natural language [20]. Research systems have done so for years [36], and controlled experiments show language-model agents responding to some tactics that move people [37]. Convergence needs more than this primitive: an open question, an evidence-bearing answer, and an offer constructed within the conversation. A signed quote must carry terms and expiry, be issued under the seller’s mandate, and attest that the buying agent is addressing the brand [15]. Section X distinguishes the four objects required.

    Convergence also needs a cost for misrepresentation. Crawford and Sobel showed that costless, non-binding messages between parties with divergent interests convey coarse information, with less conveyed as divergence grows [38]. A seller’s agent and a buyer’s agent have divergent interests by construction. Checkable evidence, binding commitments and liability make their exchange informative by introducing conditions outside that model. No current commerce protocol imposes that cost. I would count the pillar falsified if, by the end of 2028, no widely adopted commerce protocol carries a free-text question with a cited answer and a signed per-conversation offer.

  4. 4

    The interface becomes an output, and the company keeps the grammar.

    A model can generate an interface for each visitor and intent, validate it against a schema and render it from components the company has already approved. The fixed page becomes a special case of the composed one, and a site becomes a set of materials and rules from which pages are made.

    That single change moves three things. Review moves to the grammar and the invariants, because there is no finite set of pages to inspect in advance, while human approval stays the default until a company decides otherwise. Invention becomes a property the system forbids by construction: the composer re-assembles what the company has said, shown and verified, and produces the source of any claim on request. And what is shown to one person becomes an object the company has to be able to reconstruct later, or it has published something nobody can examine. The pillar weakens if composition adds latency and error without adding relevance, or if answers built from approved sources carry as many invented attributes as answers assembled by third parties.

  5. 5

    Authority is granted in stages and answered for at every stage.

    An agent acting on live systems needs a mandate before it needs intelligence: a bounded action space, a record and a way back. It observes, proposes, acts with approval, and only then acts alone.

    The ladder is a commercial instrument as much as a safety one. A company widens the mandate as evidence of reliable work accumulates and withdraws it when the agent fails, which is why trust in agents behaves like credit. The record has to be producible to someone outside the transaction, because a log only the seller can read settles nothing when a claim is disputed. And the same discipline reaches the interface: the composition a visitor saw belongs in the record beside the actions the agent took. The pillar weakens if comparable deployments show no gain in the authority people are willing to delegate once capability, cost and task are held constant.

IX

The strongest objections.

The platforms will build the seller side themselves.

Something better than that has already happened, and it deserves a more careful answer than the competitive one. A merchant-named agent configured from the merchant’s feed went live inside Google Search in January [3], and in September Anthropic published open, customizable reference implementations of a shopping agent and a merchant agent [39]. That release is not a selling platform at all. It is a blueprint anyone can take, read, fork and run, and the floor it sets is genuinely low: a company with an engineer and a catalog can stand up an agent that answers from its own product data.

That is good news, and it settles the wrong question. A blueprint gives a company an agent that can answer from a catalog. It does not give the company the thing that decides whether the answer is worth reading. What a buyer, or a buying agent, is actually asking about is seldom in the catalog. Why this material was chosen over the cheaper one. Which of two products the company would refuse to sell for this use. What the third version of the hinge fixed, which claim the legal team will stand behind in which market, what the service does when the install goes wrong. That knowledge exists inside a company as documents, systems, habits and people, and none of it arrives through a feed.

A single tool leaves that knowledge where it is. Bringing it into a live conversation takes a harness: many agents, tools and functions held together. One observes how the company is described from outside. One compiles its offer and keeps it true as catalogs and engines change. One carries the company’s own experts into training. One composes what a visitor sees. One acts on live systems under a mandate and leaves a record. An evaluation loop says which of them is wrong this week. The blueprints are the easy half of that, and they are now free. The contest moves to where it belongs: whether a company can gather its own knowledge into a form that survives contact with a machine, and keep it honest.

Two structural points survive alongside the harness. An agent that lives on a surface, speaks within the surface’s rules and feeds the surface’s telemetry carries that surface’s mandate, whatever name it wears. And distribution belongs to the platforms, which is right and useful, while representation, the company’s deeper knowledge and its perimeter of action across its own systems belong to the company.

The assistants will absorb both jobs and leave the seller with a feed.

This is the bear case: models learn to read bad pages, representation loses value, and buying stays inside assistants. The brand’s site becomes somewhere to confirm the company’s existence. Better retrieval will probably solve some of our four failure classes before catalog improvements do. The claim about where buying happens leaves three jobs intact: knowing how the engines describe the company, keeping the offer true and machine-checkable at source, and acting on its systems under its mandate. Those jobs change surface and do not disappear, and where the assistant performs all three, the company’s knowledge of its own demand accumulates on the assistant’s side.

People like to browse.

The pleasure lies in looking, wanting and receiving attention. A grid is one container for it. We expect a page composing itself around a person to provide more of each than forty identical thumbnails. Browsing survives in the form of a dialogue.

If every seller optimizes the answer, the answer becomes spam.

Two decades of search make this the objection I take most seriously. A different equilibrium depends on verification. Persuasion scales without limit, but a checkable claim makes fabrication detectable during the transaction, on a record. The seller faces a cost where deceiving a crawler once carried none.

Strategic behavior survives and changes shape. Ellison and Ellison documented internet retailers making offers harder to compare, with obfuscation associated with lower price elasticity [40]. Expect fabrication where checks are absent, source manipulation where they are shallow, deliberate incomparability where an attribute would lose, and correlated errors where agents share an evaluator. Verification raises the cost of the crude lie and leaves the subtle ones, which is why measurement has to separate correctness from usefulness, and why it becomes the seller’s discipline as much as the buyer’s.

A generated page cannot carry a brand’s identity, and generation is slower than a fixed page.

Both objections are right about naive generation. Imitate a brand’s style freehand and you get a forgery. Assemble the page from scratch with a visitor waiting and you get a slower one. Section X answers architecturally: approved components, a schema, a confidence gate and human approval make fidelity measurable through visual distance, provenance and tone. Deterministic templates supply the floor, product pages are pre-loaded, and clear intents take a fixed path whenever composition would add nothing.

A seller’s agent cannot be trusted by buyers.

A seller’s interests have always required scrutiny. Repeated dealings and reputation help only when injured parties detect the injury, and residual loss makes that uncertain here. A brand agent therefore offers what its counterparty can check during the conversation: sources retrievable immediately, a bounded action space, logs available to an outside auditor, and a seller answerable for the agent’s statements as for its packaging. The sensale’s capital was reputation. Its counterpart here is verifiability, which a single unsourced claim can spend within one turn.

X

What we are building.

Veliu is an agentic commerce lab. The discipline above needs more than one piece of technology, so we are building several, and they are meant to work as one harness.

There are generative interface models, which compose a page for the visitor in front of them out of material a company has already approved. There are protocols for how an agent speaks on a company’s behalf and how what it says can be checked. Among them is the Brand Voice Protocol, to be published with its first version: every claim carries its source, every answer stays inside agreed bounds, every exchange leaves a record. There are consumer products, where a shopping agent buys and sells on a person’s behalf inside an agentic marketplace, which is where we learned most of what this paper reports. And there are the agents that carry a company’s own name, the brand agents. Some of this is live, some of it will be announced over the coming months, and what follows is the part we can describe precisely today.

The brand agent is where the harness meets a single company, and those parts become the tools and the two faces of one agent. The reference architecture is one agent per brand, one loop and two faces sharing one memory. The work face studies the market and readies the offer for the company’s team. The selling face answers its buyers.

What it learns from. The agent studies five sources.

First, real buying conversations from recruited buyers: paid, consented, supervised, with profiled personas. Their turns preserve the budgets, the uncertainty and the second thoughts that single prompts remove.

Second, the daily conversation stream of our consumer agentic marketplace, where people delegate buying and selling every day.

Third, telemetry from every brand agent we operate: organic and paid channels, signals from the customer base, e-commerce and social. On the company’s own site it reads the explicit signals (what visitors write, ask and search) and the implicit ones (a zoomed photograph, a page read to the end, a fast scroll). It watches the market for a rival product going viral or an engine beginning to recommend one, and it learns which proposition converts for visitors from which channel.

Fourth, synthetic personas from our own pipeline [34]. Controlled scale begins with root-to-leaf paths through a 704-node content taxonomy, passed to a teacher model together with a seven-tier, 130-attribute schema, which produces a coherent fictional buyer. The pipeline role-plays that buyer through six to fourteen dialogue turns with a target emotional intensity. Goals and constraints reveal the persona implicitly, and some turns carry no signal at all, as in real traffic.

The pipeline then extracts each dialogue’s implied signals, records the reasoning that produced them, and assembles a dataset for distillation into a student model small enough for production. Precision is judged by a model family different from the teacher’s. Recall is measured through a fixed questionnaire answered from the dialogue and again from the extracted signals. Every call is metered.

The first two sources calibrate this material: they improve the generator, expose missing attributes and test the plausibility of simulated distributions. Every observation retains its origin, so generated coverage remains distinguishable from field evidence.

Fifth, the brand’s own people. Sales assistants, product specialists and founders contribute expertise as training: the question behind a hesitant customer’s question, why the hinge’s third version was redesigned, what the company refuses to make. The judgment that makes this team distinctive in its industry, in delivering a service or building a product, becomes available in every conversation. This is the shop assistant of section IV, whose reading of a customer never scaled.

What it fixes, and keeps fixing. The middle moment is continuous training on what happens inside the brand, across its channels, in its market and among competitors. As catalogs, engines and rivals change, the agent re-compiles attributes, feeds, content, agent-facing pages and structured data. Checks of the brand’s visibility run daily across the engines with rotating prompts, and competitor comparisons use verified facts only. The prohibition on invention is enforced in code.

A day-zero baseline is frozen before any write. Later changes can be compared against it and retained or reversed on evidence. Where systems permit, fixes are written at source in the space reserved for applications, completing existing structured data without duplication. Elsewhere they arrive as a costed, prioritized plan. The brand’s site stays as the brand made it.

How it sells to people. This is the architecture’s center. Its mechanics make the claim checkable: that a generated page can carry a brand’s identity.

The input is a visitor’s sentence or behavior. The output is a typed component tree, serialized as data and validated against a schema, using a small closed set of primitives with deterministic templates as the floor.

Before launch, we capture the brand’s site, cluster repeated subtrees, freeze their structure and type their content slots. We distill design tokens for color, type, spacing and radius, preserving color fidelity at Delta E below 3. Each approved component enters a per-brand library labeled by section, audience, offer type, sector and style, with component relations recorded.

During a visit the composer can only select and arrange what is in that library. It produces neither code nor images. One generic, sanitized interpreter renders all brands’ pages, with no per-brand code or runtime code evaluation. Product pages are pre-loaded to remove click latency.

Publication requires every invariant to pass, including schema validity and claim provenance, and a composition confidence gate of 95 percent. Review-and-approve remains the default until the brand chooses otherwise.

The visitor keeps talking, and the interface keeps listening: explicit and implicit signals adapt what appears next in real time, in the visitor’s language and market.

Deployment takes one line of code on the brand’s site. The URL stays the brand’s, with no redirect. The agent navigates the company’s real pages and the company sees the traffic in its own analytics. Where the offer is bought online, checkout runs on the company’s own rails and the company remains merchant of record; where it is not, the agent opens whatever next step the company offers. No two visitors need meet the same page, and every page is drawn from the same approved materials.

How it answers buying agents. For now, buying agents need exact answers. Each brand agent exposes a Model Context Protocol server, structured feeds, agent-readable pages with complete structured data, and support for the commerce protocols assistants use. All answer from the brand’s compiled representation in real time, keeping stock and price as fresh as on the site. The same components can appear inside assistants accepting third-party interfaces, under their host’s guidelines.

As buying agents learn to interrogate, the brand agent will argue and negotiate within the brand’s limits. This is pillar 3’s convergence. Transport interoperability alone leaves four objects unresolved.

Identity

establishes which legal seller speaks

Authority

specifies what its agent may disclose, negotiate and bind

Evidence

links claims to versioned sources, including jurisdiction and validity period

Commitment

turns a conversational proposal into an offer carrying terms, expiry and a signature attributable to the merchant of record

A protocol may connect agents while resolving none of these distinctions.

The measurement instrument. Our quarterly agentic selling benchmark is built on real conversations and an open methodology [6], scoring five properties per brand: visible, comparable, recommended, purchasable, correct. Purchasable means different things by trade: whether a transaction can proceed where there is a checkout, and whether the buyer can take the next step the company offers (a quote, a consultation, a trial) where there is not. The first edition publishes with its methodology, and versions track engines that change monthly.

visiblecomparablerecommendedpurchasablecorrect

Two properties hold this architecture together. The data contract lets category knowledge compound across brands while keeping each brand’s knowledge its own. Provenance and access controls make two questions auditable: does category learning accelerate as brands are added, and has any private brand knowledge moved? The autonomy ladder in pillar 5 constrains every action. One system. It learns from millions of buying conversations, and it sells in yours.

One system. It learns from millions of buying conversations, and it sells in yours.

XI

What we expect.

The benchmark’s first edition tests three hypotheses stated in advance: machine readability predicts admission to an answer; verifiable differentiation predicts selection after admission; and answers grounded in brand-approved sources contain fewer attributes unsupported by a held-out fact set than answers assembled from third-party sources.

A causal claim about readability needs controlled variants that hold meaning constant and change only the machine-readable structure. The registration publishes the sampling frame, the engine and model versions, the repeated-query policy, the denominators and the decision rules before collection closes. We are not neutral, we sell what this paper argues for, and that is why the hypotheses precede the data. Any company can ask to have its category measured, any researcher can contest the method, and both happen in public.

We expect three phases and date them so we can be wrong.

Phase I, now through 2028

The search moves before the purchase does.

The first thing to move is discovery. By the end of 2028 we expect most product research to begin in a conversation, with an assistant or an agent doing the finding and comparing, and the seller met inside an answer well before anyone opens a site. Delegated purchase follows more slowly and unevenly. It concentrates where the outcome is easy to state and cheap to be wrong about: consumables, replacements, standardized goods and services, mostly at low and middle prices. Where the choosing carries personal meaning, we expect the agent to prepare and the person to keep more of the decision.

Phase II, 2028 to 2031

The mandate stands and the two arts converge.

Trust earned in steps hardens into standing mandates that replenish, renew and negotiate under rules a principal sets once. Buying agents begin interrogating company agents directly, asking for evidence where they used to read fields, and the crafts of selling to a person and to a machine start to merge. Disputes over what an agent said on a company’s behalf move from the occasional case to a routine one, and the provenance of a commercial claim becomes a question for the people who run the company.

Phase III, by 2031

A company agent becomes ordinary infrastructure.

Within five years we expect most brands to run an agent of their own, as ordinary as a website is today, and for the same reason: without one, a company is described by whoever has one. It will have an owner inside the company and a line in the budget. Procurement will ask suppliers for one, and agencies will build and run them for the companies that do not. Commercial pages will be generated for each visitor from approved material, with fixed pages kept wherever a company wants a stable public account of itself. The distinction between a site and an agent will stop being useful, because the site will be one of the agent’s surfaces.

Beneath these changes lies agency in both senses: software’s capacity to act and people’s capacity to author their choices. Frankfurt described personhood through second-order desires, desires about our desires: wanting to want something else, revising an ordering as well as satisfying it [41]. A mandate listing stated preferences contains only the first order. It faithfully executes what someone wanted when they wrote it, even as that person changes their mind about who they are becoming.

Delegation is safe in proportion to how settled a want is, and an accurate execution of an earlier preference can still miss the person whose preference it was. That reaches the anthropology from the other side, and it guides what should be automated better than any list of categories.

Polanyi embedded markets in social relations [42]. His double movement supplies the less consoling prediction: exchange disentangles itself from society, and society responds through legislation and collective institutions, usually late. Delegation continues that disentangling by moving exchange from persons to instruments. The counter-movement is visible in standards rooms and will reach statutes.

Plurality has an operational meaning there. Differences among sellers remain legible, and several evaluators remain possible. A person can inspect an answer, recover the canonical offer, reset a composition, revoke a mandate and buy directly. The open question is whether those building the systems will write the first draft of this protection while it can still be made technical.

Few things are more human than what we buy, how we look for it, whom we trust to advise us and what choosing feels like. For twenty years we accepted commercial surfaces that could not understand a sentence. That technological constraint has lifted. Software can now read and compose. It holds a consistency no person sustains across a day of visitors, and it does not tire at the fifth question.

The whole of buying is being rearranged, both what we delegate and what we keep, and the rearrangement is happening in specifications, in schemas and in the defaults of a handful of systems. That is an unglamorous place for something this large to be decided, and it is where the decisions are. Its measure will be the plurality preserved and the human judgment that survives software's assistance.

A market is the sum of what its participants can say to one another and be believed. For most of history that was a person’s word, disciplined by the prospect of meeting the same buyer again. For twenty years it was a page, disciplined by almost nothing. What is being assembled now could be the first commercial infrastructure in which a claim is routinely checked at the moment it is made, by the party it is made to. That is either the most disciplined market anyone has built, or an engine for flattening every seller into whatever fits a schema. The difference is being settled this year, by people writing documents that look far too dull to matter.

The seller’s craft, the oldest in commerce, has always been to be understood by whoever is asking. It has just acquired a second kind of listener, and it will have to learn to speak to both without lying to either.

Selling is next.

References

  1. [1] Licklider, J. C. R. (1960). “Man-Computer Symbiosis.” IRE Transactions on Human Factors in Electronics, HFE-1, 4-11.
  2. [2] The interface timeline, in brief. Vercel popularized the term “Generative UI” with AI SDK 3.0 on March 1, 2024 (https://vercel.com/blog/ai-sdk-3-generative-ui). Google introduced A2UI, an open, framework-agnostic specification for agent-driven, streamed generative interfaces, on December 15, 2025 (“Introducing A2UI: An open project for agent-driven interfaces,” Google Developers Blog) and published version 0.9 with React, Flutter, Lit and Angular renderers on April 17, 2026 (“A2UI v0.9: The New Standard for Portable, Framework-Agnostic Generative UI,” Google Developers Blog; specification at https://a2ui.org). MCP Apps is the Model Context Protocol extension for interactive interface resources rendered inside assistants.
  3. [3] The buyer-side timeline, in brief. OpenAI launched Instant Checkout in ChatGPT in late September 2025, open-sourcing the Agentic Commerce Protocol with Stripe and setting a merchant fee on completed purchases. Microsoft launched Copilot Checkout with Shopify in early January 2026. Google unveiled the Universal Commerce Protocol on January 11, 2026 at NRF, co-developed with Shopify, Etsy, Wayfair, Target and Walmart, and endorsed by more than twenty partners including Visa, Mastercard, Stripe and American Express; UCP interoperates with the Agent Payments Protocol (AP2), Agent2Agent (A2A) and the Model Context Protocol (MCP), with the retailer as seller of record. Business Agent, the merchant-branded conversational agent inside Google Search configured from Merchant Center, went live January 12. The cross-retailer cart followed in May; Shopify moved its catalog MCP endpoints to the new scheme on June 15. Timeline re-verified at publication. Primary sources: Stripe, “Stripe powers Instant Checkout in ChatGPT and releases Agentic Commerce Protocol codeveloped with OpenAI,” September 2025, https://stripe.com/newsroom/news/stripe-openai-instant-checkout; Microsoft, “Microsoft propels retail forward with agentic AI capabilities,” January 8, 2026, https://news.microsoft.com/source/2026/01/08/microsoft-propels-retail-forward-with-agentic-ai-capabilities-that-power-intelligent-automation-for-every-retail-function/; NRF, “Google deepens AI investments that impact retail,” January 2026, https://nrf.com/blog/google-deepens-ai-investments-that-impact-retail.
  4. [4] Hutchins, E. L., Hollan, J. D., and Norman, D. A. (1985). “Direct Manipulation Interfaces.” Human-Computer Interaction, 1(4), 311-338.
  5. [5] Simon, H. A. (1971). “Designing Organizations for an Information-Rich World.” In M. Greenberger (ed.), Computers, Communications, and the Public Interest. Johns Hopkins Press.
  6. [6] Veliu (2026). “The Agentic Selling Benchmark: Methodology, Scoring, and Failure-Class Taxonomy.” Version 1.0, to be published with the first benchmark edition. Covers panel sourcing, persona profiling, consent and payment terms, engine coverage, scoring, and the failure-class taxonomy of section I. The three hypotheses of section XI are registered before the first data collection closes.
  7. [7] Pandya, V. (2026). “U.S. retailers see surge in AI traffic, but many websites are not entirely readable by machines.” Adobe, April 16, 2026. https://business.adobe.com/blog/ai-traffic-surge-retail-sites-not-machine-readable. Reports the readability and demand figures used in sections I and IV: product pages average a 66% machine-readability score (34% of their content is unreadable by the models doing the recommending) against 75% for homepages; AI-referred traffic to U.S. retail sites grew 393% year over year in Q1 2026 and converted 42% better than non-AI channels in March 2026, having converted 38% worse twelve months earlier; 39% of surveyed consumers report shopping with AI, 85% of whom report an improved experience. Retail is where the measurement exists first; the mechanism is channel-wide.
  8. [8] Stigler, G. J. (1961). “The Economics of Information.” Journal of Political Economy, 69(3), 213-225.
  9. [9] Bakos, J. Y. (1997). “Reducing Buyer Search Costs: Implications for Electronic Marketplaces.” Management Science, 43(12), 1676-1692.
  10. [10] Chamberlin, E. H. (1933). The Theory of Monopolistic Competition. Harvard University Press.
  11. [11] Diamond, P. A. (1971). “A model of price adjustment.” Journal of Economic Theory, 3(2), 156-168. In a market of identical buyers searching sequentially over a homogeneous good, an arbitrarily small positive search cost yields the monopoly price as the unique equilibrium.
  12. [12] Armstrong, M. (2006). “Competition in two-sided markets.” RAND Journal of Economics, 37(3), 668-691. Introduces the competitive bottleneck: where one side single-homes and the other multi-homes, the platform competes for the single-homing side and extracts from the multi-homing side.
  13. [13] Jensen, M. C., and Meckling, W. H. (1976). “Theory of the Firm: Managerial Behavior, Agency Costs and Ownership Structure.” Journal of Financial Economics, 3(4), 305-360.
  14. [14] Accenture (2026). Consumer Pulse Research 2026. June 2026; 25,590 consumers across 16 countries. Reports that 74% of consumers would let a personal AI agent handle routine tasks, 32% would accept delegated decision-making within set limits, and 9% would accept autonomous purchase completion, with configurable permissions, instant override, and clear recourse as the leading conditions of trust (reported in AI News, June 15, 2026, https://www.artificialintelligence-news.com/news/ai-shopping-agents-consumer-trust-accenture-report/). See also Checkout.com (2026). “Agentic Commerce 2026: The State of Consumer Demand and Merchant Readiness,” June 9, 2026, https://www.checkout.com/newsroom/consumer-demand-for-ai-shopping-is-forming-fast-but-trust-for-agentic-commerce-is-still-catching-up: spending caps (30%), instant permission revocation (29%), and easy cancellation (28%) as the controls consumers require before delegating.
  15. [15] The trust rails at press time: tokenized agent payments from the card networks (Visa’s Trusted Agent Protocol, Mastercard’s Agent Pay); the first evidence standards for agentic transactions; in-band machine payments over HTTP 402 (x402, under the Linux Foundation) and agent payment tooling in the cloud stacks (AWS Bedrock AgentCore Payments); verifiable credentials and portable agent identity in progress at the W3C and in ERC-8004; request signatures that let automated buyers identify themselves to the sites they visit (Web Bot Auth). Re-verified at publication. Primary sources: Visa, Trusted Agent Protocol specifications, https://developer.visa.com/capabilities/trusted-agent-protocol/trusted-agent-protocol-specifications; Mastercard, Agent Pay, https://www.mastercard.com/us/en/business/artificial-intelligence/mastercard-agent-pay.html; Cloudflare, “Securing agentic commerce: helping AI agents transact with Visa and Mastercard,” https://blog.cloudflare.com/secure-agentic-commerce/ (covers Web Bot Auth and card-network agent authentication); ERC-8004, https://eips.ethereum.org/EIPS/eip-8004.
  16. [16] Simon, H. A. (1955). “A Behavioral Model of Rational Choice.” Quarterly Journal of Economics, 69(1), 99-118.
  17. [17] Slovic, P. (1995). “The Construction of Preference.” American Psychologist, 50(5), 364-371.
  18. [18] Douglas, M., and Isherwood, B. (1979). The World of Goods: Towards an Anthropology of Consumption. Basic Books.
  19. [19] Miller, D. (1998). A Theory of Shopping. Cornell University Press.
  20. [20] Anthropic (2026). “Project Deal: our Claude-run marketplace experiment.” Published April 24, 2026. https://www.anthropic.com/features/project-deal. A one-week experiment in December 2025 in Anthropic’s San Francisco office: 69 employees, $100 each, four parallel Slack-based marketplaces, over 500 items listed, 186 deals, just over $4,000 in transaction value; agents posted listings, made offers and negotiated autonomously in natural language. Agents run on Claude Opus earned about $3.64 more per sale on identical items than agents run on Claude Haiku; participants represented by the weaker model did not notice the disadvantage and rated fairness the same; aggressive negotiation instructions did not change outcomes; 46 percent of participants said they would pay for such a service.
  21. [21] Kleinberg, J., and Raghavan, M. (2021). “Algorithmic monoculture and social welfare.” Proceedings of the National Academy of Sciences, 118(22), e2018340118. See also Bommasani, R., Creel, K. A., Kumar, A., Jurafsky, D., and Liang, P. (2022). “Picking on the Same Person: Does Algorithmic Monoculture Lead to Outcome Homogenization?” Advances in Neural Information Processing Systems 35; arXiv:2211.13972.
  22. [22] Schultz, W., Dayan, P., and Montague, P. R. (1997). “A Neural Substrate of Prediction and Reward.” Science, 275(5306), 1593-1599.
  23. [23] Berridge, K. C., and Robinson, T. E. (1998). “What is the role of dopamine in reward: hedonic impact, reward learning, or incentive salience?” Brain Research Reviews, 28(3), 309-369.
  24. [24] Knutson, B., Rick, S., Wimmer, G. E., Prelec, D., and Loewenstein, G. (2007). “Neural Predictors of Purchases.” Neuron, 53(1), 147-156.
  25. [25] Poldrack, R. A. (2006). “Can cognitive processes be inferred from neuroimaging data?” Trends in Cognitive Sciences, 10(2), 59-63.
  26. [26] Campbell, C. (1987). The Romantic Ethic and the Spirit of Modern Consumerism. Blackwell.
  27. [27] Holbrook, M. B., and Hirschman, E. C. (1982). “The Experiential Aspects of Consumption: Consumer Fantasies, Feelings, and Fun.” Journal of Consumer Research, 9(2), 132-140.
  28. [28] Pine, B. J., and Gilmore, J. H. (1998). “Welcome to the Experience Economy.” Harvard Business Review, July-August 1998.
  29. [29] Manovich, L. (2001). The Language of New Media. MIT Press. Variability as a principle of a computational medium: the object that exists in many versions and has no single instance.
  30. [30] Nissenbaum, H. (2004). “Privacy as Contextual Integrity.” Washington Law Review, 79(1), 119-157.
  31. [31] “Agentic selling”: the seller side of agentic commerce, named here. The buyer side has begun to accumulate a literature, including generative engine optimization (Aggarwal, P., et al. (2024). “GEO: Generative Engine Optimization.” Proceedings of KDD 2024; arXiv:2311.09735). The selling half has none yet. We intend to write it.
  32. [32] Sensale: Treccani, s.v. “sensale”; from Arabic simsar, from Persian, mediator. The word itself traveled the trade routes.
  33. [33] Parveen, D., Kang, D., Paruchuri, A., Kayal, D., and Mallapragada, P. (2026). “Accelerating Personalization Signal Learning via Synthetic Data.” Proceedings of ECIR 2026. https://www.amazon.science/publications/accelerating-personalization-signal-learning-via-synthetic-data. Finds that models trained on synthetic personas outperform models trained on real, de-identified data, in precision and recall on the same real test set.
  34. [34] Veliu (2026). The Veliu persona pipeline. Five stages: taxonomy-anchored persona generation under a tiered attribute schema, role-played buying dialogues with controlled affect targets, signal extraction with reasoning traces, dataset assembly for distillation, and evaluation. A reproduction of the method of [33], whose original code was not released.
  35. [35] Figures as of the date of publication, counted on the live marketplace described in section VII and stated as lower bounds.
  36. [36] Lewis, M., Yarats, D., Dauphin, Y., Parikh, D., and Batra, D. (2017). “Deal or No Deal? End-to-End Learning of Negotiation Dialogues.” Proceedings of EMNLP 2017, 2443-2453.
  37. [37] Bianchi, F., Chia, P. J., Yuksekgonul, M., Tagliabue, J., Jurafsky, D., and Zou, J. (2024). “How Well Can LLMs Negotiate? NegotiationArena Platform and Analysis.” Proceedings of ICML 2024; arXiv:2402.05863.
  38. [38] Crawford, V. P., and Sobel, J. (1982). “Strategic Information Transmission.” Econometrica, 50(6), 1431-1451.
  39. [39] Anthropic (2026). “Building Commerce Agents with Claude.” September 2, 2026. https://claude.com/blog/claude-for-commerce-agents; reference implementations at https://github.com/anthropics/commerce-agents. Ships a shopping agent (search, compare, substitute, assemble the order) and a merchant agent (sales questions, promotions, inventory and pricing), with implementations over UCP and the Shopify Admin API published by Shopify at https://github.com/Shopify/claude-for-commerce-examples.
  40. [40] Ellison, G., and Ellison, S. F. (2009). “Search, Obfuscation, and Price Elasticities on the Internet.” Econometrica, 77(2), 427-452.
  41. [41] Frankfurt, H. G. (1971). “Freedom of the Will and the Concept of a Person.” Journal of Philosophy, 68(1), 5-20.
  42. [42] Polanyi, K. (1944). The Great Transformation. Farrar and Rinehart.