Back to the Lab

Research Lab · Article · Agentic selling

Which Agentic Commerce Statistics Show Actual Adoption?

Agentic commerce statistics separated into transactions, deployments, surveys, readiness measures, and forecasts, with primary-source checks.

By Veliu Editorial Team13 min read
Which Agentic Commerce Statistics Show Actual Adoption?

In brief. Agentic commerce statistics measure consumer demand, merchant adoption, infrastructure readiness, transaction activity, and forecast scenarios. For merchants appearing in ChatGPT, Gemini, Perplexity, or Copilot, completed orders and production logs show observed behavior; surveys show reported attitudes; forecasts show possible scale. No single figure proves market-wide adoption.

  • Agentic commerce surveys, deployments, transactions, and forecasts answer different questions and should never share one adoption label.
  • Completed orders provide the strongest behavioral evidence when the denominator, timeframe, geography, cancellations, and refunds are disclosed.
  • Merchant enrollment proves participation at a stated program level, while active usage and order volume require separate records.
  • Forecasts describe possible future scale under stated assumptions and do not measure current market adoption.
  • Merchants can act now by measuring SKU-level consistency across product pages, feeds, commerce endpoints, and checkout responses.

Agentic commerce statistics measure consumer demand, merchant adoption, infrastructure readiness, transaction activity, and forecast scenarios. For merchants appearing in ChatGPT, Gemini, Perplexity, or Copilot, completed orders and production logs show observed behavior; surveys show reported attitudes; forecasts show possible scale. No single figure proves market-wide adoption.

A merchant can enable a product feed today and still have zero agent-completed orders. That gap matters for US brands and ecommerce operators reading adoption claims about ChatGPT, Gemini, Perplexity, and Copilot.

These systems can increasingly read current merchant feeds, indexed product pages, schema.org facts, and clean catalog records. If a size variant has no identifier or the feed says $89 while the page says $99, that product may fail an eligibility check or create trouble during comparison and checkout, depending on the platform.

This article separates observed behavior from surveys, deployments, transaction records, and forecasts. It also shows what each number can prove about a store.

What do the five evidence classes prove?

Agentic commerce statistics provide evidence about consumer demand, merchant adoption, infrastructure readiness, transaction activity, and future scenarios. Each class answers a different question. A survey can measure willingness, a protocol release can establish technical availability, and an order record can show completed behavior within a defined system.

That distinction prevents a common analytical error. A merchant announcement, a consumer saying they would delegate a purchase, and a completed delegated payment are three separate events with three separate denominators.

The unit changes the conclusion.

Six terms keep the numbers honest

Agentic commerce: software agents participate in product research, comparison, cart creation, checkout, or payment under user authority.

Buying agent: the buyer-side system acting on a shopper’s intent, such as finding a waterproof carry-on under $250 and comparing eligible products.

Agentic Selling: the seller-side discipline through which a brand studies the market, readies its store, and sells to people and buying agents.

Observed metric: a measured event or production deployment, such as 12,000 completed checkout sessions during a stated month.

Survey result: a reported attitude, intention, or claimed behavior from a defined sample using a disclosed question.

Forecast: a modeled future scenario built from assumptions about adoption, spending, coverage, and time.

Verdict: A credible statistic exposes its denominator

A credible evidence ledger preserves the unit, denominator, observation period, geography, and source. “Thousands of merchants” is weak by itself. “5,000 enrolled merchants shipping to the US as of November 2025,” when reported by the program operator, is a scoped program statistic.

Evidence classWhat it measuresStrongest acceptable sourceWhat it can proveMain limitation
Completed transactionsOrders, payment attempts, completion, refundsAudited merchant, processor, or platform recordsPurchases occurred within the disclosed scopeOne system cannot establish industry-wide adoption
Production deploymentsFeatures available to real merchants or shoppersOfficial launch record and product documentationA capability was live for the stated usersAvailability does not show active use
Catalog or protocol readinessEligible products, valid feeds, working endpointsValidation logs, protocol records, merchant systemsProducts or checkout paths passed defined checksReadiness does not show selection or demand
Behavioral telemetrySearches, product-card views, sessions, cart eventsFirst-party event logs with methodologyObserved use on one surface and periodLogs seldom explain motivation
Consumer or merchant surveysAttitudes, recall, intention, claimed useOriginal report with questionnaire and sampleWhat sampled respondents reportedSelf-report does not verify an order
ForecastsModeled future revenue or adoptionOriginal model owner with assumptions and rangeA possible scenario under stated inputsA scenario does not measure current adoption

Transaction records and production activity show behavior. Surveys show responses. Forecasts show scenarios.

Verdict: Demand percentages need the exact question

Demand evidence moves through distinct stages: awareness, current use, willingness to delegate, trust, purchase intent, and completed delegated purchase. Movement from one stage to the next requires separate measurement.

The appeal is obvious: one large percentage can make the market look settled. The evidence does not support that shortcut because question wording, sample composition, country, and fieldwork date change the meaning of every survey result.

A defensible demand table records the exact question before recording the percentage. This edition excludes survey percentages that cannot be checked against an original questionnaire and methodology file.

Evidence entryRequired scopeClassificationNamed limitation
Current use of a buying agentFieldwork date, geography, sample size, shopper definition, recall windowClaimed behaviorSelf-report cannot verify an order
Willingness to delegateFieldwork date, geography, sample size, task and approval conditionsStated intentionIntention may not become use
Trust in autonomous purchaseFieldwork date, category, price limit, approval rule, merchant contextAttitudeTrust varies by product risk
Completed delegated purchaseObservation window, geography, transaction denominator, completed-order definitionObserved behaviorPlatform scope limits generalization

This evidence control keeps unsupported data out of the analysis. Adding an unattributed “AI shopping” percentage would make the page look fuller while weakening the conclusion.

A useful survey also separates a recommendation from delegated action. Asking ChatGPT for running-shoe ideas is AI-assisted research. Authorizing a buying agent to choose a size, create a cart, and submit payment is delegated commerce.

Verdict: Merchant enrollment proves participation at a stated level

Merchant adoption has at least five levels: announcement, pilot, protocol support, production availability, and active usage. Count the level stated by the primary source.

Perplexity reported that its Merchant Program had expanded to more than 5,000 merchants in November 2025. Its program announcement establishes catalog participation at that date and within the stated scope of merchants shipping to the US. The announcement does not report orders for every enrolled merchant.

OpenAI’s current product-feed specification documents fields and eligibility rules for product ingestion. Because specifications change, this article does not freeze an attribute count. A feed that passes the applicable checks establishes technical readiness for that profile and version.

Google announced the Universal Commerce Protocol, or UCP, with Shopify and named partners in its NRF remarks. That source supports an announcement and its stated partner scope. Merchant-level activity requires separate usage records.

Access labels belong beside every deployment claim:

LabelMeaning
BroadGenerally available to the stated merchant or user population
LimitedRestricted by geography, account, category, or invitation
BetaLive testing with change risk and incomplete coverage
Partner-mediatedAvailable through a named commerce platform or provider
AnnouncedPublicly committed, with production use still unverified

A logo slide cannot supply the missing denominator.

Verdict: Transaction activity is strongest and scarcest

A completed order records behavior. Merchant-side checkout completion, delegated-payment success, repeat use, cancellations, and refunds reveal more than impressions or feature availability.

The denominator is decisive. A 92% payment success rate means little if it excludes authentication failures before the payment attempt. Likewise, 10,000 completed orders need a time window, geography, merchant count, and refund policy.

MetricRequired denominatorRequired scopeDisclosure status
Completed delegated ordersInitiated eligible checkout sessionsDates, geography, cancellations, later refundsRarely public
Checkout-session completionAll created agent checkout sessionsFixed period, expiry and cancellation rulesMerchant or platform record
Delegated-payment successAll submitted payment attemptsFixed period, retries and reversalsProcessor record
Repeat delegated purchaseBuyers with one completed delegated orderCohort window and refund eligibilityRarely public
Net transacted valueCompleted settled ordersCurrency, fixed period, refunds and chargebacksAudited disclosure preferred

No public protocol repository can fill this table for the whole industry. Repositories document capabilities and releases. Industry totals require transaction records from a defined population.

Verdict: Forecasts describe possible scale

McKinsey modeled a scenario in which the US business-to-consumer retail market could see up to $1 trillion in orchestrated revenue by 2030. Its original analysis defines the scenario and its boundaries. “Could,” “up to,” “US B2C,” “orchestrated,” and “2030” are essential qualifiers.

Orchestrated revenue may include spending influenced or coordinated by agents. The figure does not automatically equal autonomous checkout volume, payment value processed by one protocol, or incremental retail revenue.

Forecast evidenceObserved transaction evidence
Uses assumptions about future adoption and spendingRecords an event that already occurred
Has a future horizon and scenario rangeHas an observation window and denominator
Frames possible market scaleEstablishes activity in a defined system
Supports capital-planning scenariosSupports operational measurement

Forecasts belong in the scenario column of the ledger.

Buyer-side agents create seller-side work

Buying agents express and execute shopper intent. A brand agent covers the seller side of agentic commerce, where the store answers with current product facts and carries the order onto the brand’s commerce rails.

Who sells when AI buys
The brand sets offer, answer, checkout.
DimensionBuying agentBrand agentMeasurable evidence
MandateActs on a shopper’s stated intent and authoritySells for one brand inside its rules and voiceRecorded intent, query logs, approved selling rules
Product-data dependenceUses comparable price, stock, variant, delivery, and return factsMaintains and serves approved product factsAccuracy, completeness, and freshness checks
Transaction roleBuilds or approves a cart under buyer authorityAnswers product questions and supports merchant checkoutSession creation, cart validation, completion
Checkout pathSends an authorized action into an available commerce flowKeeps the sale connected to the brand’s commerce systemMerchant-of-record and order-system records

Buyer-side adoption estimates say little about whether a specific brand can answer, “Is the navy size 10 in stock, and can it arrive in Austin by Friday?” That readiness requires a SKU-level test.

Five measurable stages connect a question to an order

A useful measurement model follows the transaction path: question, retrieval, comparison, cart, and merchant checkout. Each stage creates a separate event that can be tested.

1. The question creates a retrieval task

“Find a carry-on under $250 that fits Delta limits” creates a request for dimensions, current price, stock, and suitability. Retrieval-augmented generation, or RAG, means an AI system can fetch external records or pages while producing an answer. Some commerce surfaces use retrieval because price and stock can change between model-training cycles.

Measure the query set, retrieval success, market, and test date. A test with 100 fixed prompts remains comparable over time only when the prompts and engine settings remain fixed.

2. Feeds and indexed pages supply product facts

Commerce surfaces can draw from combinations of merchant feeds, indexed pages, structured data, and platform catalog records. Access, freshness, and supported fields vary by platform and account, so each test should name the surface and date.

On the open web, schema.org Product, Offer, and AggregateRating markup can make facts explicit to machines. Core schema.org does not impose one universal validity profile. For Google product rich-result eligibility, price is required within an Offer, while properties such as priceCurrency and availability follow Google’s current required or recommended rules for the applicable experience. If the page says “In stock” and the feed says “Out of stock,” record a source divergence.

3. The engine compares eligible products

Eligibility can place a product into a candidate set. Selection depends on the engine, query context, available data, and proprietary systems.

OpenAI says organic shopping results can consider relevance, availability, price, quality, and whether the merchant is the maker or primary seller. Its shopping documentation states those factors, while their exact weights remain undisclosed.

4. Protocols can carry cart and authorization data

Commerce protocols describe different parts of product exchange, cart creation, checkout, and payment authorization. Their supported stages, access rules, and maturity vary by specification version and implementation.

The Model Context Protocol, or MCP, can expose resources and callable tools such as product search or inventory checks. The MCP specification does not itself define a universal commerce checkout or payment-authorization flow. The agentic commerce protocol map explains the current roles and gaps.

5. Merchant records establish the completed order

In one common reference flow, the merchant or its commerce provider returns the authoritative cart, handles an approved payment method, and records the order. Implementations vary, including who creates the cart and how payment authority is represented.

The practical measures are checkout creation, valid-cart rate, payment success, cancellation, fulfillment, and refund. This closes the earlier denominator problem: protocol support can establish readiness, while merchant order records establish completed purchases.

From found to bought
Protocol support shows readiness. Merchant order records show completed purchases.

Why does this matter when an agent does the buying?

ChatGPT, Gemini, Perplexity, Copilot, and AI Overviews can increasingly use structured merchant feeds, indexed product pages, and schema.org facts, with availability and mechanics varying by surface. Each engine controls its own retrieval, eligibility, and selection decisions.

A market forecast says little about whether an engine can read a brand’s current price, stock, size, shipping promise, return policy, and checkout path. For a merchant, the immediate question is whether a navy jacket in size 10 appears as the same variant across the page, feed, and checkout response.

That is why generative engine optimization for ecommerce starts with readable, consistent product facts. An article can earn a citation from prose. A product with current commercial facts is easier to compare and carry into checkout.

Seven questions expose a weak statistic

QuestionPassCautionFail
Who published it?Primary operator, regulator, or standards bodyNamed analyst with disclosed sourcesAnonymous aggregation
What exactly was counted?Defined event and unitBroad category with notes“Activity” without definition
What evidence class is it?Observed, reported, announced, or forecast is labeledClassification is inferableClasses are mixed
What are the dates?Publication and measurement datesPublication date onlyNo date
What are the geography and sample?Both disclosedOne disclosedNeither disclosed
What is the denominator?Full denominator and exclusionsPartial denominatorPercentage alone
Can it be traced?Original methodology and URLSecondary link naming the originalCircular citations

A statistic passes only when another analyst can reconstruct what the number means. Reproducing the result may require private records, but reproducing the definition should be possible.

Merchants can measure readiness before public totals exist

Field completeness: measure the percentage of SKUs with the properties required by each target feed or consumer profile. Segment the full catalog because perfect hero products can hide a weak long tail.

Surface divergence: compare the product page, feed, and commerce endpoint. Count every price, stock, and variant mismatch by SKU.

Propagation latency: time the interval from a stock or price change in the commerce system to its appearance on the page, feed, and sampled engine answers.

Answer accuracy: run a fixed prompt set by engine and compare stated price, availability, and specifications with the live source. Preserve screenshots, timestamps, surface, and market.

Checkout coverage: measure the percentage of active SKUs eligible for the relevant agent checkout implementation, then record checkout-session completion and payment success with their denominators.

MetricFormula or unitMinimum segmentation
Product-data completenessApplicable fields passed divided by active SKUsCategory, market, long tail
Page-feed-endpoint divergenceConflicting fields divided by checked fieldsField type and SKU
Propagation latencyMinutes from source change to observed updateSurface and change type
Answer accuracyCorrect sampled facts divided by tested factsEngine, prompt intent, market
Eligible-SKU coverageEligible active SKUs divided by active SKUsImplementation and market
Checkout completionCompleted orders divided by created sessionsDevice, market, merchant
Payment successSuccessful payments divided by submitted attemptsProvider and failure reason
Answer or citation shareAppearances within a fixed prompt setEngine, intent, date

Answer or citation share is an outcome proxy inside a fixed prompt set. It does not establish revenue or a position outside the measured prompts.

The verified timeline records availability

Infrastructure availability and usage are separate records. This compact timeline includes milestones supported by the cited primary pages and avoids unsupported release dates.

DateMilestoneWhat the source establishesScopePrimary source
November 2024Anthropic published MCPAn open protocol for connecting AI applications with tools and resourcesGlobal specificationMCP
November 2025Perplexity reported 5,000+ merchants in its programProgram enrollment and selected shopping capabilitiesMerchants shipping to the USPerplexity
January 2026Google announced UCP with Shopify and named partnersProtocol announcement and stated partner supportAvailability described in the remarksGoogle NRF remarks

A launch proves that infrastructure existed at its stated scope. Usage requires another row backed by activity data.

Five actions turn market claims into store measurements

  1. Label every number as transaction, deployment, readiness, telemetry, survey, or forecast evidence.
  2. Preserve denominators, measurement dates, geography, exclusions, and exact question wording.
  3. Prioritize transaction records and SKU-level readiness measures when allocating implementation work.
  4. Audit the agent view from outside the store, then compare current price, stock, variants, shipping, returns, and checkout across each relevant surface.
  5. Repeat the same fixed tests after every catalog or protocol change, so movement reflects the store and the documented method.

Veliu’s brand agent studies the market, readies the store, and sells to people and to the AI shopping agents that arrive. Catalog retrieval, normalization, feeds, and commerce endpoints help it answer with a current price and a valid variant. On the brand’s site, it can adapt the selling experience from permitted signals and approved components, answer shopping agents machine to machine, and return the questions customers asked. Checkout stays on the brand’s rails.

The sale stays on the brand’s rails
The brand agent asks, verifies, answers, then carries the sale to the brand’s checkout.

Author: Veliu Editorial Team

Editorial methodology: This evidence ledger prioritizes first-party product documentation, official protocol repositories, standards organizations, and original research methodology. Every numerical claim is classified by evidence type and checked for date, geography, denominator, production scope, and limitation. Vendor statements establish what the vendor documented within the cited scope. The timeline excludes milestones whose exact dates or release status could not be confirmed from the linked primary record.

The next operational step is concrete: export 100 active SKUs and count every price, stock, identifier, and variant mismatch across the page, feed, and checkout response.

The next paper, when it is written.

One email per paper, and nothing else.

Subscribe

More from the blogResearch Lab

The report