Back to the blog

What Is Generative UI, and Who Composes Your Storefront?

What Is Generative UI, and Who Composes Your Storefront?

Generative UI composes the interface per request. What Google proved, what stays deterministic, and how a brand agent runs the agentic storefront.

By Veliu Editorial Team11 min read

Key takeaways

Generative UI is an interface an AI model composes at the moment of use. Google measured an 82.8% preference for generated interfaces over text answers and shipped the approach inside Gemini and Search. For a brand the question is control: platforms already compose your products on their surfaces, while a brand agent can compose the brand's own selling surface from approved components, with checkout staying deterministic on the brand's rails.

  • Generative UI is an interface an AI model composes at the moment of use, from approved components and current catalog data.
  • Google Research measured an 82.8% preference for generated interfaces over text answers, with checkout, conversion, and generation speed left untested.
  • AI platforms already compose brand products inside their own generated surfaces, while Adobe measured AI-referred retail traffic converting 42% better in March 2026 and product pages scoring worst on machine readability.
  • A brand agent with a governed component grammar composes the brand's own selling surface, and commerce state stays deterministic on the brand's rails.

On November 18, 2025, Google changed what an answer looks like. Inside the Gemini app and Search AI Mode, the interface itself is now generated per query, an approach Google Research calls generative UI. Ask about a jacket and the reply can arrive as product cards, a size filter, and a comparison, composed on Google's surface from whatever product data the model can read.

Google measured the effect before shipping it. In pairwise human ratings, generated interfaces beat standard markdown answers in 82.8% of comparisons on the LMArena prompt set. People prefer a composed interface to a wall of text, by a wide margin.

For a brand or ecommerce operator, this raises a sharper question than the design press is asking. ChatGPT, Gemini, Perplexity, and Copilot are becoming shopping surfaces, and each of them can now render your products inside an interface you never designed. This article defines generative UI in plain terms, separates what the evidence proves from what it does not, and lays out what has to be true in your catalog before a generated interface can sell for your brand.

The storefront is becoming an output.

What is generative UI?

Generative UI is a user interface that an AI model composes at the moment of use: the model chooses which components to show, fills them with current data, and arranges the layout for this specific request, within limits set by whoever governs it. In a store, that means one shopper question can produce three product cards, a care comparison, and a size filter that did not exist as a page a second earlier.

The terminology is young. The CHI 2026 workshop on generative UI, whose organizers span Microsoft Research, UC San Diego, MBZUAI, Apple, and MIT CSAIL, concedes that the community has not yet agreed on shared definitions: it proposes only "working definitions", and the broadest of them covers any interface created by an AI model, whether during design or at the point of use.

In practice the term means three different things depending on who is speaking, and the difference decides what a merchant should build.

CampWhat generative UI means thereWho it servesWhere it runs
Developer tooling (Vercel AI SDK, CopilotKit, tambo)A model's tool call streams a pre-built component into the conversationApp developersYour application
UX research (Nielsen Norman Group)An interface generated in real time around one user's needs and contextDesignersAny product
Platforms (Google, since Gemini 3)The model generates the entire experience, including shopping viewsThe platform's usersThe platform's surface

Vercel's documentation defines it mechanically as "connecting the results of a tool call to a React component." Nielsen Norman Group defined it in March 2024 as "a user interface that is dynamically generated in real time by artificial intelligence to provide an experience customized to fit the user's needs and context." Google Research treats it as a model that "generates not only content but an entire user experience."

All three camps share one omission. None of them says a word about selling.

How is generative UI different from the personalization engine you already bought?

Merchandising software has adapted pages for a decade, so the fair question is what actually changed. The answer is the unit of work.

ApproachWhat changes per visitorWho decidesGranularity
Merchandising rules and A/B slotsWhich banner or block fills a fixed slotA rule a human wrote in advanceSegment
Adaptive UIOrder and visibility of predefined elementsPredefined triggersSession
Generative UIWhich components exist on the surface, with which data, in which arrangementA model choosing within an approved setSingle request

A rules engine asks which of five banners to show. A generative interface asks what the right selling surface for this request is, then assembles it. When a shopper types "a machine-washable coat for a 7-year-old, under $90, here by Friday", no pre-written rule owns that intersection of care, age, price, and delivery; a composed comparison of the eligible products does.

Adaptive personalization rearranges the furniture. Generative UI builds the room for the question.

The evidence: people prefer a composed interface

Google Research's evaluation, announced with the Gemini 3 rollout and released in full on arXiv, is the strongest public data on the question, and its fine print matters as much as its headline.

The headline is real. Generated interfaces were preferred over generated markdown in 82.8% of pairwise comparisons on prompts sampled from LMArena, and the margin grew on information-seeking queries. In the same study's ELO ranking, generative UI scored roughly 380 points above the top search-result website (1736 against 1352), while human-expert-built sites stayed on top at 1800.

Interfaces raters preferred, ELO score (LMArena set)

Pairwise human preferences converted to ELO; generation speed excluded from the evaluation.

CategoryELO
Human-expert site1,800 ELO
Generative UI1,736 ELO
Generated markdown1,438 ELO
Top search result1,352 ELO
Plain text1,174 ELO

Experts still win. Raters preferred the expert-built website over the generated interface in about half of the comparisons, and the generated one in roughly a third. The bar was not a brand's design team: the expert pages were commissioned from freelancers at 100 to 130 dollars per site, built in a few hours each.

The capability is new. One model generation earlier this simply did not work: the study reports output error rates of 29% to 60% for the previous small models, against 0% for the current ones, measured with the study's repair passes in place. Whatever a team concluded from a 2024 prototype is stale.

The clock was off. Live generation "can often take a minute or two", by the paper's own admission, and raters were shown pre-cached results, so the 82.8% was measured with latency removed. A shopper on a product page does not wait two minutes.

Preference is proven. Selling is not.

Why a generated interface cannot run your checkout

The same study is blunt about what it never tested. The paper contains no evaluation of checkout, live inventory, or purchase outcomes, and the raw model outputs needed a repair pipeline that fixed JavaScript errors, CSS errors, and hallucinated assets before any rater saw them.

Commerce already has case law for what happens when a probabilistic system speaks for the money side. In February 2024 a Canadian tribunal held Air Canada liable for a bereavement refund policy that its website chat tool had invented, and ordered the airline to honor it. An interface may compose the shelf; it cannot improvise the contract.

Veliu's research position on the component grammar of the agentic storefront draws the operating boundary from the design side: agents may compose presentation, while governed systems retain commerce state. A model can select an approved comparison component and explain why three coats fit the request. Price, stock, cart, consent, payment, and checkout stay behind deterministic validation, and the purchase completes in the store's own governed checkout.

That division is what the University of Bristol position paper presented at CHI 2026 demands of the discipline: designers must "retain accountability" when generation enters the interface, including for information architecture and regulatory requirements. For a storefront, the brand retains accountability for every claim a generated card makes.

How a generated interface sells: the component grammar

The popular belief says a truly generative storefront must write the whole experience from scratch at runtime. The evidence points the other way: Google needed detailed system instructions and post-processors to keep free-form generation coherent, and the developer camp converged on the same lesson. Tambo, one of the infrastructure vendors in the space, wrote the rule down: "Don't generate UI from scratch when you can provide well-designed components."

The working model is dynamic bounded composition. The brand approves a library of domain-agnostic components (a comparison grid, a size selector, a product carousel, a care panel), each carrying the brand's visual tokens and each allowed to claim only specific catalog fields. The agent chooses which components fit the request, which products and data fill them, and which approved layout arranges them. It composes in real time; it does not invent a cart implementation or unreviewed code per request.

Walk one request through it. A shopper asks for an in-stock, machine-washable coat in size 8. The agent maps the question to catalog constraints, retrieves the eligible records, and composes three product cards with a care comparison and a size filter. If the care field is missing on one coat, the interface says so or omits the claim; a governed enrichment process can recover the missing attribute later, and a human reviews it before it becomes catalog truth.

Where composition stops and commerce begins: 1. Shopper intent (Question plus permitted context); 2. Catalog constraints (Size, care, price, stock); 3. Approved components (Cards, filters, comparisons); 4. Shopper correction (Inspect and revise context); 5. Commerce rules (Cart, consent, payment, checkout)
Where composition stops and commerce begins

Two rules keep the composition trustworthy.

Signals stay under shopper control. A remembered size may narrow a shortlist only if the interface shows that memory, dates it, and offers an immediate correction. "I selected these in size 8; tell me if you need a different size" is the honest form.

Catalog truth sets the ceiling. A polished card cannot create a missing care instruction, repair a detached size variant, or resolve two conflicting prices. Every generated sentence must trace to a current, governed product record, which is the same canonical catalog work an agentic storefront depends on before any interface is generated at all.

Who composes your storefront if you do nothing?

The platforms are not waiting for merchants to decide. Each major AI surface already composes brand products inside its own generated experience, fed by the brand's structured data.

SurfaceWhat it composesThe seller's inputStatus
Google SearchBusiness Agent: shoppers chat with your brand in Search; AI Mode generates shopping viewsMerchant Center data plus your websiteRolling out in the US since January 2026, eligibility-gated
ChatGPTShopping results and agentic storefronts; purchases typically complete on the merchant's own storeA product feed per OpenAI's specificationRevamped around discovery in March 2026
Microsoft CopilotCopilot Checkout and on-site Brand AgentsMerchant program; Shopify merchants auto-enrolled with opt-outUS rollout since January 2026
PerplexityShopping hub with agent-led buyingFeed via its merchant program; PayPal merchants act as merchant of recordInstant Buy live in the US since November 2025

Note the pattern in the fine print: after OpenAI's March 2026 shopping revamp and Perplexity's merchant-of-record model, discovery happens on the platform's surface while the transaction lands back on the merchant's storefront. The store still closes the sale; the platform's generated interface decides which stores get the chance. Their agent works for the channel. Yours works for you.

The traffic behind this is measured. Adobe Analytics reported AI-referred traffic to US retail sites up 393% year over year in Q1 2026, and in March 2026 that traffic converted 42% better than non-AI traffic. The same report scored product detail pages worst of all retail page types on machine readability, at an average of 66%.

The pages that sell are the pages agents read worst.

This is the control question hiding inside a definition. When the only generative UI in your category runs on the platforms' side, your products appear however their model composes them. An agent that works for the brand rather than the channel is the counterweight: the same governed catalog facts serve human shoppers on your domain, your composed selling surface, and the external engines that read structured product data to answer buying questions.

What gets indexed when every session is different?

A composed surface raises a fair technical objection: crawlers cannot index an interface that exists once. The answer is that generative UI rides on top of the record layer; it never replaces it.

Canonical product pages, feeds, and structured data remain stable and deterministic, and that layer is what search crawlers and answer engines keep reading. Composition changes the presentation of those records per request; it does not change the records. A storefront built this way serves three readers at once, from one source of truth: people, the brand's own composed interface, and the buying side of agentic commerce that arrives machine to machine.

Where the brand agent fits

Veliu is the brand agent that does the selling—a standing mandate that goes beyond answering a chat session: it studies how the category gets shopped, readies the store so every product fact holds up, and sells to people and to the AI shopping agents that arrive, in the brand's voice. On the brand's own domain it composes the selling surface per visitor from permitted signals and approved components, and it answers shopping agents machine to machine from the same governed records; reading the public catalog from outside is what lets it fill every card with the current price and a valid variant. Checkout stays on the brand's rails, with the brand as merchant of record. Behind the agent sit 50K+ customers, 300K+ delegated transactions, over €100M transacted, and over a year of R&D on a live delegated-commerce marketplace.

What this means for your product

  1. Treat the interface as an output and the catalog as the input. Map every claim a generated card could display to a current catalog field, and disclose or omit whatever has no field behind it.
  1. Approve the component set before wiring any model. Decide which structures may appear on your storefront, which fields each may claim, and which layouts are acceptable, in writing, before generation touches a visitor.
  1. Keep remembered context correctable. Any stored fact that narrows a shortlist must be visible, dated, and reversible by the shopper on the spot.
  1. Keep commerce deterministic. Price, stock, cart, consent, payment, and checkout stay behind validation and explicit approval, on your rails, whatever the interface does upstream.
  1. Measure failures by field, then expand. Record which catalog gaps blocked which component states, fix the governed source, and only then add the next adaptive layout.

The concrete next step fits in one afternoon: run one journey (an in-stock, machine-washable coat in size 8) against 25 live SKUs, and log every component state whose displayed claim cannot be traced to a current catalog field.

Component stateClaim shownCatalog field behind itStatusAction
Product cardPrice and availabilityGoverned offer recordpass / fail / unknownFix source before layout
Care comparisonMachine-washableCare attributepass / fail / unknownEnrich, then human review
Size filterIn stock in size 8Child-variant stockpass / fail / unknownRepair variant links

Author: Veliu Editorial Team

Methodology note: Veliu Editorial Team reviewed the linked primary sources as available on August 3, 2026: Google Research's generative UI evaluation and blog post, Nielsen Norman Group's definition, the CHI 2026 workshop proposal, the University of Bristol position paper, Google Business Agent documentation, OpenAI commerce documentation and announcements, Microsoft's Copilot commerce announcement, PayPal's Perplexity partnership notes, and Adobe Analytics AI traffic reporting. The review compared published definitions, evaluation conditions, and stated limitations. It ran no new user study and no live checkout test, and no figure in this article measures Veliu's own performance.

Get the next piece when it ships

We send new notes on making your catalog readable and buyable by AI agents as they come out.

Subscribe to the newsletter