ChatGPT Now Answers With Interfaces. Generative UI Has Gone Mainstream

Soedja Editorial Team14 min read

Key takeaways

  • On 7 October 2026, OpenAI began rolling out GPT-6 with Intelligent UI, which lets ChatGPT answer with charts, forms, calculators and small apps instead of only text. With Google's Gemini (November 2025) and Anthropic's Claude (March 2026) already doing versions of this, all three of the largest AI assistants now generate interfaces on demand.
  • The strongest public evidence comes from Google's research paper: raters preferred generated interfaces over standard markdown answers 82.8% of the time, and over the top Google Search result 90% of the time. Pages built by human experts still won overall.
  • The capability arrived suddenly. Gemini 2.0 Flash-Lite produced broken interfaces on 60% of prompts; Gemini 2.5 and Gemini 3 produced none.
  • The industry is splitting into two architectures: models that write the interface as code, and models that assemble it from a trusted component library. The second is winning on speed, safety and brand control.
  • In our pull of npm data, monthly downloads of the MCP Apps SDK grew from about 890,000 in January 2026 to 16.6 million in September. The plumbing for generated interfaces is being adopted faster than the consumer features built on it.

The answer is turning into an interface

For three years, the dominant shape of an AI answer was a block of text with a few headings and bullet points. Developers called it the markdown wall. It was flexible and fast to produce, but it treated every question the same way, whether the person wanted a definition, a comparison of three laptops or a mortgage estimate.

That shape is now changing at the largest scale available. On 7 October 2026, OpenAI started rolling out GPT-6 to ChatGPT with a capability it calls Intelligent UI, describing it as a way "to answer quickly with fully interactive user interfaces." Paid plans received it first; Free and Go users followed the next day. OpenAI's examples include a roast-dinner timeline next to the recipe, a road-trip route on a map, a bill splitter, a retirement calculator and an explorable breakdown of a bicycle's drivetrain. The model decides on its own when an interface is better than prose. No toggle, no special prompt.

OpenAI says ChatGPT serves more than 1.2 billion people a week. That number is what makes this launch different from the experiments that preceded it. Google introduced generative UI in the Gemini app and in Search's AI Mode in November 2025. Anthropic gave every Claude plan inline interactive charts and diagrams in March 2026. Those were meaningful. But a default change inside ChatGPT is the point at which generated interfaces stop being a feature some people try and become a format most people see.

1.2 billion weekly ChatGPT users now receiving answers that can include generated interfaces

82.8% of head-to-head comparisons in which raters preferred a generated interface over a markdown answer

18.7× growth in monthly MCP Apps SDK downloads between January and September 2026

What "generative UI" actually means

The term gets used loosely, so it helps to separate it from things that look similar.

Generative UI is an interface created at runtime, for one specific request, by the model answering it. Ask how compound interest works and you may get a slider for the rate and a chart that redraws as you drag. Ask the same question with different numbers tomorrow and the interface may be built differently. Google's researchers put it simply: the model generates "not only content but an entire user experience."

It is not the same as AI design tools such as v0, Lovable or Figma's AI features, which help people produce interfaces at build time that are then shipped and reused. It is also different from classic personalisation, where a fixed interface swaps content based on who is looking. Generative UI changes the structure of the screen itself, and it does so per question rather than per user segment.

The closest older idea is what the Nielsen Norman Group called outcome-oriented design: instead of designing every screen, designers define goals and constraints, and the system assembles the right view. For years that was mostly theory. The models were not reliable enough to do it.

People prefer it. Experts still win

The most rigorous public evaluation so far is Google's paper Generative UI: LLMs are Effective UI Generators, by Yaniv Leviathan, Dani Valevski, Yossi Matias and colleagues. The team took prompts from LMArena and a second set of information-seeking questions, then had human raters compare five kinds of answer: a generated interface, a standard markdown answer, plain text, the top Google Search result for the query, and a website built specifically for that prompt by a paid human expert.

Generated interfaces beat every format except a human expert

Human ratings of five answer formats for the same prompts

Source: Leviathan et al., "Generative UI: LLMs are Effective UI Generators" (Google, arXiv 2604.09577, February 2026), Tables 1, 2 and 7. Generation time was excluded from ratings.

Three things stand out. First, the gap between generated interfaces and markdown is large, roughly 300 Elo points, which is a bigger jump than the one from plain text to markdown. Second, generated interfaces beat the top search result decisively, which is uncomfortable reading for anyone whose business depends on being that result. Third, the human experts remain ahead. The paper describes the generated pages as worse than expert-crafted ones but "at least comparable in 50% of cases."

The economics matter as much as the ranking. Google paid its contractors US$100 to US$130 per website, and each took three to five hours. The model produced a comparable page about half the time in a minute or two. For a single prompt, that is not a contest humans can win on cost.

A generated interface does not need to beat a designer. It only needs to beat a wall of text, and it already does, by a wide margin.

One caveat: the evaluation deliberately ignored generation time. Raters judged the finished result, not the wait. As we explain below, latency is precisely the problem the industry has spent the last year trying to solve.

The capability appeared almost overnight

The same paper ran the identical system on five generations of Gemini models. The results explain why generative UI went from demo to default so quickly.

Generative UI is an emergent ability: it switched on between model generations

The same generative UI system run on five Gemini models (LMArena prompts)

Source: Leviathan et al., "Generative UI: LLMs are Effective UI Generators" (Google, arXiv 2604.09577, February 2026), Table 3.

Gemini 2.0 Flash-Lite failed to produce a usable page for 60% of prompts. One generation later, failures dropped to zero, and quality kept rising. The authors call the ability emergent: it is not something that was gradually engineered into the product, it is something that became possible once the underlying models crossed a threshold of reliability in writing working front-end code.

That has a practical consequence. Any company building on top of these models gets the capability for free as models improve, which is why the infrastructure race described below has moved so fast.

Two ways to build an interface on the fly

Behind the consumer features sit two fundamentally different architectures, and the choice between them shapes speed, safety and how much control a brand keeps over what users see.

Write the code. The model writes HTML, CSS and JavaScript for each answer, and the app runs it in a sandbox. Google's research system works this way, as do Claude's inline visuals and the interactive apps that run inside chat windows through MCP Apps. It is maximally flexible: anything a web page can do, the answer can do. It is also slow, because the model must write every line, and risky, because the result is executable code that nobody reviewed. Google's own paper notes that generation "can often take a minute or two."

Compose the components. The model does not write code. It describes an interface using a fixed catalogue of pre-built, trusted components, and the app renders them natively. Google's open A2UI standard takes this route. Its launch post states the design goal as making UI "safe like data, but expressive like code," and warns that "running arbitrary code generated by an LLM may present a significant security risk." OpenAI's Intelligent UI also leans this way: the company says it built "a library of native, streamable components" and a compiler that renders the interface while the model is still generating it.

Write the codeCompose the components
What the model outputsHTML, CSS and JavaScriptA structured description that references known components
ExamplesGoogle's Generative UI research, Claude visuals, MCP AppsChatGPT Intelligent UI, A2UI, json-render, AG-UI front ends
SpeedSlow; whole page must be writtenFast; components stream in as generated
SecuritySandboxed execution of unreviewed codeOnly pre-approved components can render
Visual consistencyVaries answer to answerInherits the host's design system
ExpressivenessAnything a web page can doLimited to what the catalogue offers

In practice the two are converging. A third-party teardown of Intelligent UI by the generative UI company Thesys, which OpenAI has not confirmed, reports a catalogue of roughly 70 native components plus an escape hatch for self-contained HTML apps when the catalogue is not enough. That hybrid, components by default and code when necessary, looks likely to become the norm.

For brands, the component approach has an underrated implication. If an assistant renders answers from a catalogue, then a company's design system stops being an internal style guide and becomes something closer to an API: the set of building blocks an agent is allowed to use when it speaks on the brand's behalf.

The plumbing is being built in public

Consumer launches get the headlines, but most of the last year's work happened in open standards and developer kits. The timeline below separates the two.

From text box to interface in 31 months

Key generative UI launches. Tap an event for detail; filter by type.

    Sources: Vercel, Anthropic, OpenAI, Google Research, Google Developers Blog, Model Context Protocol blog, The New Stack.

    To measure how quickly developers are adopting that plumbing, we pulled monthly download counts from the public npm registry for the main JavaScript packages behind each standard, covering October 2025 to September 2026.

    Developers are adopting open generative UI standards fast

    Monthly npm downloads of the main package behind each standard. Toggle packages below.

    Source: Soedja analysis of the public npm registry downloads API (api.npmjs.org), monthly totals October 2025 to September 2026, pulled 10 October 2026. Packages: @modelcontextprotocol/ext-apps, @ag-ui/core, @json-render/core, @a2ui/web_core, @mcp-ui/client, @openai/chatkit-react. Downloads include automated installs. Growth view indexes each package to its first full month (= 1×). Hover or tap the chart for monthly values.

    The pattern is steep and consistent. The official MCP Apps SDK, which lets tools show interactive interfaces inside Claude, ChatGPT and other clients, went from about 890,000 downloads in its first full month to 16.6 million in September. Google's A2UI core library grew almost 14-fold in five months. AG-UI, the event protocol that connects agents to front ends, grew about 15-fold over the year. Combined, the four protocol-level packages in the chart were downloaded about 38 million times in September 2026, against roughly 590,000 a year earlier.

    Two counter-signals are worth noting. MCP-UI, the community project whose ideas fed into MCP Apps, has plateaued since mid-year as developers move to the official SDK. And OpenAI's own ChatKit React package peaked in June and has fallen about 65% since. Read together, the data suggests developers prefer open, cross-vendor standards over a single assistant's toolkit, even when that assistant is the largest.

    Download counts are a proxy, not a census. They include automated builds and dependency installs, and one popular framework adopting a package can move the numbers. The direction is still unambiguous. For context, Vercel's general-purpose AI SDK, which many of these projects build on, reached 106.5 million downloads in September, six times its level a year earlier.

    What still breaks

    Generated interfaces introduce failure modes that text answers do not have, and the early reviews have already found them.

    • Plausible but wrong visuals. A chart is more persuasive than a sentence. When a generated diagram encodes an error, readers are less likely to question it. Google's paper notes that outputs "occasionally contain inaccuracies"; OpenAI says "there's still work ahead to improve the model's design judgment."
    • Overuse. Not every question deserves a widget. Early user reactions to Intelligent UI include complaints that simple questions now come back with unnecessary visuals. Knowing when not to build an interface is a design skill the models are still learning.
    • Fragile state. The Thesys teardown reports that Intelligent UI interfaces reset on reload and that the model cannot see values a user changes in a widget. An interface the model cannot read back is a dead end for follow-up questions.
    • Accessibility. A hand-coded page can be audited for screen readers and keyboard use. A page generated fresh for every query cannot, unless it is built from components that are accessible by design. This is another argument for the component approach.
    • Measurement. An interface rendered inside a chat window has no URL, no page view and no analytics tag. Wired asked whether Intelligent UI is a real capability gain or presentation polish. For marketers, the more pressing question is how anyone will know what users saw.

    What this means for brands and builders in Indonesia

    The shift extends a trend we have tracked on this blog. In September, we reported that Google's AI Overviews now appear on most branded searches. Generative UI pushes the same dynamic further: the assistant no longer only summarises a website, it can rebuild the website's most useful function, such as a calculator, a comparison table or a booking form, inside the answer. When that happens, the reason to click through gets weaker.

    It also raises the stakes in the platform contest we described in Gemini vs ChatGPT in Indonesia. Google's advantage in Indonesia is distribution through Android. OpenAI's answer, as of this week, is a richer answer format for everyone, including free users.

    For Indonesian companies, three responses are practical now. The first is to make data machine-readable: clean product specifications, prices and availability that an assistant can turn into an accurate comparison. The second is to treat the design system as a product, documented as components with clear rules, because that is the form agents can consume. The third, for businesses with a genuinely useful tool, is to publish it as an app inside the assistants through MCP Apps, so the interaction happens on the brand's terms rather than as an assistant's improvised copy.

    ChangeImpactWhat to do about it
    ChatGPT, Gemini and Claude all generate interfacesAnswers replace simple web tools and comparison pagesAudit which of your pages an assistant could rebuild
    Generated UI beats the top search result 90% of the time in Google's testsFewer clicks for informational queriesShift content toward depth, original data and things assistants cannot synthesise
    Component-based rendering is becoming the normBrand consistency depends on cataloguesDocument your design system as reusable components
    MCP Apps adoption is acceleratingBrands can ship interactive tools inside assistantsPrototype one high-value tool as an MCP App
    Generated interfaces have no URL or page viewAttribution gets harderTrack assistant referrals and branded search as proxy signals

    Frequently asked questions

    What is generative UI?

    Generative UI is a user interface that an AI model creates in real time for a specific request, such as an interactive chart, a calculator, a form or a small app, instead of replying with text only. The layout is decided per question rather than designed in advance.

    What is ChatGPT Intelligent UI?

    Intelligent UI is OpenAI's generative UI capability in ChatGPT, launched with GPT-6 on 7 October 2026. It lets answers combine text, visuals and interactive elements such as buttons, sliders and charts. It rolled out to paid plans first and to Free and Go users from 8 October 2026.

    How is generative UI different from AI website builders?

    AI website builders and design tools help people create interfaces that are then published and reused. Generative UI creates an interface at the moment a question is asked, for that question only, inside the assistant's answer.

    Do people actually prefer generated interfaces?

    In Google's published evaluation, raters preferred generated interfaces over markdown answers 82.8% of the time and over plain text 97% of the time. Websites built by human experts were still rated higher overall, though generated pages were at least comparable in about half of cases.

    What are A2UI and MCP Apps?

    Both are open standards for letting AI agents show interfaces. A2UI, from Google, has agents describe an interface as structured data that the app renders with its own trusted components. MCP Apps, an official extension of the Model Context Protocol since January 2026, lets tools deliver interactive HTML interfaces that run sandboxed inside clients such as Claude and ChatGPT.

    Is generative UI safe?

    It depends on the architecture. Systems that run model-written code rely on sandboxing. Systems that render from a fixed component catalogue limit what can appear on screen, which is why Google describes A2UI as "safe like data, but expressive like code."

    Will generative UI reduce website traffic?

    For simple informational and utility queries, it likely will, because the assistant can deliver the answer and the tool in one place. Pages with original data, depth and experiences an assistant cannot reproduce are better positioned.

    Keep reading

    Soedja Editorial Team

    Editorial Team

    Share:

    Need help with this?

    Let us know what you need and we will get back to you.