kulikowski.me
← Writing

When a WebMCP tool returns an MCP App

8 min read#agentic-web #webmcp #mcp #mcp-appsview as markdown

Coauthored with

Liad Yosef · MCP Apps / Ora

WebMCP gives agents tools from a website. MCP Apps gives humans UI from an MCP server.

So what happens when a WebMCP tool returns an MCP App? 🤔

Picture an agent working with a website you already have open. Most steps are plain tool calls, but then it needs you to compare a few options, move a slider or check a preview.

Could the website open a small UI inside the agent's side panel for exactly that step? We built an experiment to find out.

Two directions

You probably know WebMCP already - a page registers tools, then an agent in the browser calls them. One detail matters here: WebMCP defines tools only. No MCP resources and no prompts.

MCP Apps goes in the other direction. An MCP tool points to a ui:// resource, and the host renders it as a sandboxed View: a chart, map, form or whatever the human needs for the next step. Think of it as the last mile of interaction: the agent does the heavy lifting, and the View gives the human the right controls for the decision that remains.

So which one should you use? For us, one question makes the choice much simpler:

Where is the human looking?

At the website

WebMCP

The agent operates the live page instead of replacing it.

At the agent

MCP Apps

A small UI travels into whichever agent the human uses.

Between the two

BothThis post

The page stays the main surface, one step moves into the agent.

The split case is the one we wanted to test. The website remains the main surface, but one step works better as a focused UI inside the agent's side panel.

Two ways to combine them

There are two bridges between WebMCP and MCP Apps, and they run in opposite directions:

Bridge 1 · MCP Apps draft

View → host

The View the host rendered exposes tools, so the agent can operate it. WebMCP-shaped, carried over the MCP Apps channel.

Bridge 2 · in neither spec

WebMCP tool → ViewThis post

A page tool returns UI, and the host opens it for the human beside the website.

1️⃣ A View exposes tools to the host. The MCP Apps draft spec proposes this as App-provided tools: the host can operate the View it rendered. The tools are shaped like WebMCP tools, but they travel over the MCP Apps channel, not WebMCP.

2️⃣ A WebMCP tool returns a View. When the agent calls the tool, the website gives the human an interface.

We tested the second one 👇

Demo - Form / Factor

We built Form / Factor, a home-gym equipment page. It recommends equipment based on your measurements, the space you have and your preferred balance between strength, variety and convenience.

The whole demo is on GitHub if you want to run it yourself: Kulikowski/webmcp-apps - a local page plus an unpacked Chrome extension. You'll need to enable chrome://flags/#enable-webmcp-testing.

The page's WebMCP tools in DevTools (left), the live website (middle) and the fit-controls View the agent opened in the side panel (right).

The human asks the side-panel agent for help choosing equipment. The agent is scripted, with no LLM - we're testing the bridge, not the model.

The agent calls gym_open_fit_sidecar, which returns text for the model, the page state as structuredContent, and the View itself as an embedded ui:// resource, plus a _meta.ui.resourceUri link telling the host which View to open for the human.

document.modelContext.registerTool({
  name: "gym_open_fit_sidecar",
  description: "Open interactive equipment-fit controls for the user. Returns an MCP Apps UI resource.",
  inputSchema: { type: "object", properties: {}, additionalProperties: false },
  annotations: { readOnlyHint: true },
  async execute() {
    return {
      content: [
        { type: "text", text: "Interactive equipment fit controls opened for the user." },
        // The View itself, as a standard MCP embedded resource
        { type: "resource", resource: {
          uri: "ui://form-factor/equipment-fit",
          mimeType: "text/html;profile=mcp-app",
          text: FIT_SIDECAR_HTML,
        } },
      ],
      structuredContent: snapshot(),
      _meta: {
        // MCP Apps puts this on the tool; registerTool() has no _meta
        ui: { resourceUri: "ui://form-factor/equipment-fit" },
        // Not standard: stands in for MCP Apps' visibility: ["app"]
        "webmcp-apps": {
          allowedPageTools: ["gym_update_profile", "gym_set_preferences"],
        },
      },
    };
  },
});

The extension renders that View in the side panel through a sandbox proxy - controls for measurements, room size and the training mix. As the human changes them, the View calls two more page tools, gym_update_profile and gym_set_preferences, and the recommendation updates on the live page 😎

The Form / Factor loop · Bridge 2

The page hands the agent a View, the human decides inside it, and their choices flow back to the page.

Website

owns the state

Agent

extension · MCP Apps host

View

sandboxed iframe

1gym_open_fit_sidecar
the agent calls a page tool
2text + ui:// resource
the result carries the View
3sandbox proxy
the host renders it in a sandbox
4ui/initialize
the View opens the handshake

Human moves the sliders and drags the training mix

5tools/call
the View asks the host
6gym_set_preferences
or gym_update_profile · allowlisted

Recommendation updates on the live page

WebMCPMCP AppsOur bridge - in neither spec

The agent finds the tool and opens the View, the human makes the decision, and the website remains the source of truth.

What's standard, and what isn't

Inside the side panel, the View follows the MCP Apps protocol. It's a real ui:// resource with the MCP Apps MIME type, and the extension renders it through a sandbox proxy, completes the ui/initialize handshake, then exchanges JSON-RPC tools/call requests and tool-result notifications over postMessage.

WebMCP has tools only, so three pieces of MCP Apps have to move:

  • The link moves to the result. MCP Apps puts _meta.ui.resourceUri on the tool definition, but WebMCP's registerTool() has no _meta, so the host can't know about the View in advance.
  • The View is embedded in the result. An MCP Apps host fetches it with resources/read. WebMCP has no resources, so the same resource contents ride in the result's content. MCP Apps deferred embedded resources; MCP-UI uses them.
  • App-only tools become an allowlist. WebMCP can't mark tools with MCP Apps' visibility: ["app"], so the page names them in allowedPageTools under its own webmcp-apps key in _meta.

Rendering the View at all is still a private convention between our page and our extension: WebMCP has no rule that a host should render UI found in a tool result. A client that doesn't know the convention may pass the whole result to the model as text, HTML included. This experiment shows that the two can meet, not that they already interoperate.

That also means our extension has to provide the guardrails on its own 🔑

  • Render only a known ui:// URI, and only when the result embeds that exact resource with the MCP Apps MIME type.
  • Talk only to the expected page origin, and render the View through a sandbox proxy - the double-iframe pattern of the MCP Apps reference host - under a deny-by-default CSP. The View gets an opaque origin, so it can't read the site's cookies or storage or reach the extension's privileges.
  • Keep an allowlist of page tools the View can call. gym_update_profile and gym_set_preferences behave like MCP Apps app-only tools, but WebMCP cannot mark them that way, so any other agent on the page still sees them.

Open questions

The experiment shows the pieces can fit together. It doesn't settle how they should. These are the questions we kept asking ourselves, and where we currently lean.

1. Should the tool return the UI, or point to it up front?

In MCP Apps, a tool points to its UI before anyone calls it - that ui:// pointer from earlier. The host can review and prefetch the UI.

Our gym_open_fit_sidecar does the opposite. Nothing tells the agent that UI is coming until the result arrives with the _meta.ui.resourceUri link and the View inside. WebMCP's registerTool() has no _meta, so there's nowhere to put the link up front.

We lean towards both: point to the UI when the tool is registered, so the agent can prepare for it, then include it in the result.

2. What if nobody can see the UI?

WebMCP also describes a headless browsing scenario: the agent works on its own, then pulls in a human when it needs one.

A tool that opens UI sounds especially useful here. But what if nobody is around to see it? 🙃

Two different things matter:

  • Presence: is a human looking at the page, looking at the agent, or not present at all?
  • Rendering: can the agent show an MCP App at all, or only text?

Presence decides whether and where to show UI. Rendering decides the format. If nobody is there, the tool should say "needs a human" instead of opening UI for an empty room.

The agent could simply tell the page what it supports, but that gives sites a new way to fingerprint agents. A safer option might be to return text plus optional UI, like email's multipart/alternative, and let the agent pick what it can show. That's what Form / Factor does: a text block for any agent, and the View for one that can render it.

3. Where should a confirmation happen, and who should confirm it?

Deleting a GitHub repository shows a confirmation dialog. If an agent calls a "delete repo" tool while you watch the page, do you need the same confirmation again in the side panel?

We lean towards the page UI - it's the dialog users already know and trust. The side panel is for steps the page has no good UI for, like our training-mix triangle.

Browser agents can also click the page. If the agent can press "Delete" itself, showing the dialog doesn't prove that a human approved the action. The same problem applies to a "confirm" button inside the View.

So the real question isn't only where the dialog lives, but who is allowed to answer it. We think a tool needs a way to say "this step is for the human": pause, show the UI, and let the browser verify that the answer came from a person, not the agent. WebMCP is already discussing something close to this: requestUserInteraction() / requestUserInput().

A page could also draw something that imitates the agent's own interface. At minimum, page-originated UI should show its origin, the way browsers label permission prompts. We don't know yet whether a click inside that View should approve the next tool call.

4. Can the agent operate the View it just opened?

"Can you fill in that form?" MCP Apps lets a View expose tools to the host - the App-provided tools from bridge 1️⃣.

They look like WebMCP tools, but they aren't. To the browser agent, the View is a sandboxed iframe, not a page.

Our preference is one authoring API with two transports. The View could use an API shaped like document.modelContext.registerTool, while the SDK carries those calls over the MCP Apps channel. Developers learn one API, and the protocols stay separate.

5. Who should move: WebMCP or MCP Apps?

Our bridge needed three substitutions because the two specs don't meet in the middle. There are two ways to close that gap.

WebMCP adds resources. Then the MCP Apps flow works almost unchanged: if tools can also carry _meta to link a ui:// resource, the host reads it before the call, so it can review and prefetch the UI. The cost is a new primitive for a standard that has stayed tools-only so far - though there's already an open proposal for resources.

MCP Apps accepts embedded resources. Then WebMCP stays tools-only, and a tool result can carry its View directly. MCP already has the content block for this, and MCP-UI shipped it. The cost is losing up-front review, unless tools can also point to their UI at registration.

The second is the smaller change for both specs. The first keeps MCP Apps' security model intact. We don't think the answer is obvious.

It touches a bigger question: how MCP-compatible should WebMCP actually be?

What we'd propose

Starting points for discussion, not positions of either group:

  • For WebMCP: let tool results carry UI next to text, let registerTool() carry _meta so tools can point to their UI up front, support app-only tools, and provide a standard way to hand a step to the human and wait.
  • For MCP Apps: define a web host profile for displaying Views supplied by web pages, labelling their origin and routing calls back to the page.
  • For both: share a model of presence and rendering.

One loop

Not every website should become an MCP App, and an MCP App shouldn't rebuild a full website inside chat.

But the two can meet for one useful moment: the agent works with the website, the website opens the right controls, the human decides, and the agent carries on 🤝

Today that takes a custom extension. If you have opinions on the questions above, bring them to the specs: WebMCP issues for WebMCP and ext-apps issues for MCP Apps - we'd love to hear them 💪

The demo code is in Kulikowski/webmcp-apps, and issues about the demo itself are welcome there.

More soon. ✨