Dev.to AI πŸ€– Ai πŸ‘ 0 πŸ“– 13 min read

Building Clay: an A2UI host where the model emits interface instead of prose

Technical notes from building an agent that renders live React widgets, keeps per-session state in a Durable Object, and refuses any UI it doesn't recognize. The problem with chat Every agent product today

Building Clay: an A2UI host where the model emits interface instead of prose

Technical notes from building an agent that renders live React widgets, keeps per-session state in a Durable Object, and refuses any UI it doesn't recognize.

The problem with chat

Every agent product today has the same output medium: a paragraph. Ask for a trip budget and you get a wall of text containing numbers you cannot manipulate. Ask what happens if the hotel costs more and you re-prompt, and hope it recomputes consistently.

The interesting alternative is to let the model emit interface β€” a declarative description of widgets bound to a data model β€” and have the host render it as real controls. That's the A2UI idea: agent-to-UI. The model never sends code. It sends JSON that names components from a catalog the host already owns.

Clay is a complete implementation of that loop:

prompt ──► model ──► A2UI message ──► React renderer ──► live widgets
                                              β”‚
                                     widget event
                                              β–Ό
                          agent recomputes dependents ──► patch

The whole app is three panes: a chat rail, a generated surface, and an inspector showing the exact messages, because in this design the protocol is the product.

Why "declarative, not code" is the actual security model

The tempting shortcut is to let the model emit HTML or JSX and inject it. Then your sandbox is your own diligence, and one bad <script> is a bad day.

With a component catalog, the model's vocabulary is ten nouns:

Column Row Text Slider Toggle Table BarChart Stat Badge Button

Anything else cannot render, by construction. There is no dangerouslySetInnerHTML anywhere in the client, no eval, no dynamic import. Props arrive as data and get spread into components the bundle already contains. The model can ask for a Slider; it cannot ask for a <script>.

That turns "is the output safe?" from a scanning problem into a type problem.

The contract

One JSON object per turn, four allowed top-level keys:

{
  "surfaceUpdate": {
    "surfaceId": "s1",
    "components": [
      { "id": "root",   "component": { "Column": { "children": ["hdr", "budget", "table", "chart"] } } },
      { "id": "hdr",    "component": { "Text":   { "text": "5-day trip plan", "variant": "h2" } } },
      { "id": "budget", "component": { "Slider": { "label": "Budget", "min": 10000, "max": 80000,
                                                   "step": 1000, "value": 40000, "bind": "budget" } } },
      { "id": "table",  "component": { "Table":  { "columns": [ … ], "rows": [ … ] } } },
      { "id": "chart",  "component": { "BarChart": { "series": [{ "label": "Day 1", "value": 7200 }],
                                                     "unit": "INR" } } }
    ]
  },
  "dataModelUpdate": { "surfaceId": "s1", "dataModel": { "budget": 40000, "days": [ … ] } },
  "text": "one short sentence"
}

Two details do most of the work.

The tree is flat and id-linked. Column.children holds ids, not nested objects. Flat lists are far easier for a language model to emit correctly than deeply nested JSX-shaped trees, and patching a node is a replace-by-id.

bind is a promise. A widget's value must equal the value stored at its dotted path in the data model. That single invariant is what makes the surface live rather than decorative: the slider isn't a picture of a number, it's a handle on the model.

Architecture

                β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                β”‚  BROWSER                                                 β”‚
                β”‚  App.tsx ─► useSurface() ─► SurfaceRenderer.tsx           β”‚
                β”‚              β”‚ state, phases,   resolve via registry.ts  β”‚
                β”‚              β”‚ optimistic bind   β–Ό                        β”‚
                β”‚              β”‚ writes            Slider Table BarChart …  β”‚
                β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                               β”‚ /api/*  (Vite dev proxy)
          β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
          β–Ό                                         β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”            β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ RUNTIME=workers  :8787   β”‚            β”‚ RUNTIME=node   :3001     β”‚
β”‚ agent/index.ts           β”‚            β”‚ server-node/index.ts     β”‚
β”‚  Worker fetch()          β”‚            β”‚  node:http               β”‚
β”‚   session id β†’ DO stub   β”‚            β”‚   session id β†’ Map       β”‚
β”‚  ClayAgent               β”‚            β”‚                          β”‚
β”‚  this.state / setState() β”‚            β”‚                          β”‚
β”‚  SQLite per session      β”‚            β”‚                          β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜            β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
             β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                               β–Ό
                  agent/turn.ts  (shared turn engine)
             build messages β†’ call model β†’ classify reply
                  β†’ apply to surface β†’ measure dependents
                               β–Ό
                  agent/inference.ts (plain fetch)
             POST {baseUrl}/chat/completions Β· stream:false
                    response_format:{type:"json_object"}

Both runtimes import the same turn engine, prompt and guardrail. The fallback reimplements exactly one thing: who owns state. That's the point of the split β€” RUNTIME=node proves the Durable Object isn't load-bearing for the protocol, only for persistence.

The Cloudflare Agents SDK, from the actual API surface

The SDK's state API is the piece that changes between releases, so the installed package's type declarations were the ground truth. What [email protected] gives you:

import { Agent } from 'agents';

type SessionState = {
  sessionId: string; surfaceId: string;
  components: A2UIComponent[]; dataModel: Record<string, unknown>;
  turns: TurnRecord[]; updatedAt: number;
};

export class ClayAgent extends Agent<ClayEnv, SessionState> {
  initialState: SessionState = { /* … */ };

  async onRequest(request: Request): Promise<Response> {
    const path = new URL(request.url).pathname;
    if (path === '/api/chat')     return this.handleChat(request);
    if (path === '/api/interact') return this.handleInteract(request);
    return this.json({ error: 'not_found', detail: path }, 404);
  }
}

Five things worth knowing before you write against it:

  1. Implement onRequest, don't override fetch. The base class's fetch() resolves dynamic sub-agent paths and then delegates to your hooks. Overriding it skips that machinery.
  2. setState() is a full replace, synchronous, returns void. Forgetting {...this.state} silently wipes fields. There is no merge on the object (only on workflow steps), so every write spreads.
  3. SQLite storage is enabled by the migration, not a binding flag. new_sqlite_classes in wrangler.jsonc β€” there is no durable_objects.bindings[].sqlite key:
   {
     "compatibility_flags": ["nodejs_compat"],
     "durable_objects": { "bindings": [{ "name": "CLAY_AGENT", "class_name": "ClayAgent" }] },
     "migrations": [{ "tag": "v1", "new_sqlite_classes": ["ClayAgent"] }]
   }

After a few sessions the evidence is on disk: one .sqlite file per session,
under .wrangler/state/v3/do/clay-agent-ClayAgent/.

  1. wrangler dev does not promote process env into env. .env has to be parsed and passed explicitly, which is what scripts/run-agent.mjs does:
   for (const [key, value] of Object.entries(bindings(envFile.values))) {
     args.push('--var', `${key}:${value}`);
   }

Same script pins XDG_{CONFIG,CACHE,DATA}_HOME into the project so local state never escapes into your home directory.

  1. There is no sessionId concept for plain agents. The Durable Object's name is your session id. You mint it, you put it in the URL or a header, and namespace.get(namespace.idFromName(sessionId)) routes to it.

The Worker entry owns the public paths and forwards the body once:

// a Request stream cannot be consumed twice β€” read text, parse, forward the same text
const text = await request.text();
const body = text ? JSON.parse(text) : {};
const sessionId = clientSessionId(request, body) ?? newSessionId();
const upstream = await stubFor(env, sessionId).fetch(
  new Request(url.toString(), { method: request.method, headers: request.headers, body: text })
);

(That line about streams is not a footnote. My first version called
request.clone().json() and then .text(), and every /api/chat returned a 500 with Body has already been used.)

Deliberately unused: the SDK's WebSocket transport, agents/react, AI SDK chat helpers, MCP. Updates ride on the HTTP response, which keeps the protocol inspectable β€” and inspectable is the whole demo.

The hard part isn't rendering, it's the guardrail taxonomy

Validating the model's reply with zod is easy. Deciding what to do about an invalid reply is where the design lives.

A small model fumbles JSON constantly. Observed failure modes, all real:

  • truncated output mid-object
  • {"key, ": "total"} β€” a stray comma inside a quoted key
  • a component entry with no id at all
  • two components in one component body
  • a Slider whose value sits outside its own min/max
  • bind: "trip.travellers" while the data model wrote travellers flat

The naive answer is "retry everything." But then a model that invents
HologramMap gets retried until it behaves, and your unsupported-component path never fires β€” which is one of the things the app is supposed to demonstrate.

So the split is by kind of failure, not by severity:

Kind Examples Response
emission defect truncated JSON, missing id, unrecognized key, out-of-range slider value, two components in one body re-ask, up to 3 attempts, feeding the parser's complaint and the model's own previous answer back
contract violation unknown component name, forbidden prop, injection pattern, duplicate ids, dangling children, two roots refuse, final β€” invalid_surface, never retried

The discriminator is the zod issue code:

const result = a2uiResponseSchema.safeParse(parsed);
if (!result.success) {
  const issue = result.error.issues[0];
  // A prop the schema does not know, a missing prop, or a component body that
  // matches no variant is the model fumbling JSON β€” an emission defect, so the
  // loop asks again. Anything reported by superRefine (duplicate id, dangling
  // child, two roots, a bind with no data behind it) is a contract violation
  // and stays final.
  const defectCodes = new Set(['unrecognized_keys', 'invalid_type', 'invalid_union']);
  return {
    kind: defectCodes.has(String(issue?.code)) ? 'malformed' : 'rejected',
    detail: issuePath(result.error),
  };
}

And the re-ask, which never touches the payload:

conversation = [
  ...conversation,
  { role: 'assistant', content: raw.slice(0, 12_000) },
  { role: 'user', content: `That reply was rejected by the A2UI parser: ${emission.detail}. ` +
    `Re-emit the entire A2UI JSON object. Every entry in "components" must be an object with ` +
    `BOTH "id" (a non-empty string) and "component" (exactly one catalogued name with its props). ` +
    `Do not truncate, do not elide, do not explain. JSON only.` },
];

The rule that makes this honest: the agent loop may ask again; it may not patch. A refusal returns the model's own text (raw, first 300 chars) and the exact zod path, and the previous surface stays on screen. The inspector shows what was refused and why. Fixing the model's JSON locally would be a repair, and a repaired payload is no longer evidence of anything.

The catalog itself is one object, shared/catalog.ts, read three ways:
serialized into the system prompt, wrapped in zod schemas server-side, and
imported as the client's render allowlist. They cannot drift.

Making a drag actually change other widgets

This is the acceptance test that separates a live surface from a picture of one. Move a slider and the table, chart and stat cards must show new arithmetic.

Sequence:

you drag Slider "stay" 2500 β†’ 6000
  β”œβ”€ client: setPath(dataModel, "stay", 6000)      optimistic, thumb tracks instantly
  β”œβ”€ phase β†’ "patching"
  β–Ό
POST /api/interact { surfaceId, componentId, value }
  β”œβ”€ server writes bind into the dataModel BEFORE asking
  β”‚    (the model's job is downstream arithmetic, not re-deriving what the slider says)
  β”œβ”€ prompt carries: the event, the interactive widgets, the full current dataModel
  β”œβ”€ model returns the FULL component list + dataModelUpdate
  β”œβ”€ recomputedIds(previous, next, touchedId)
  β”œβ”€ if only the touched widget moved β†’ ask once more, pointing at its own reply
  β–Ό
200 { recomputed: ["chart","statTotal","table","statRemain", …] }

Two decisions matter.

Seed the bound value server-side. Asking the model to both apply the event and recompute everything is how you get half-updated surfaces. Applying bind is protocol semantics, not guessing at UI; the model's job becomes purely the arithmetic downstream.

Measure the right numbers. "Did dependents change?" needs a definition that can't be gamed:

// shared/a2ui.ts β€” numeric leaves, ignoring the control's own range
export function numericFields(component: A2UIComponent): Record<string, number> {
  const flat = flattenProps(componentProps(component));
  const out: Record<string, number> = {};
  for (const [key, value] of Object.entries(flat)) {
    // min/max/step describe the control, not a computed output.
    if (/(^|\.)(min|max|step|maxValue)$/.test(key)) continue;
    if (typeof value === 'number' && Number.isFinite(value)) out[key] = value;
  }
  return out;
}

Without that exclusion, a slider drag "changes" the moved widget's own value and a build that recomputes nothing still passes. The comparison also excludes the touched component id entirely.

Finally, syncBoundValues() re-writes every interactive widget's value from the data model after a patch lands. That's not repair either β€” it's the bind invariant enforced, so a stale echo in the model's reply can't rewind what the user just dragged.

In a real browser, one drag of Hotel per night β‚Ή2,500 β†’ β‚Ή6,000 moved the trip total (β‚Ή28,000 β†’ β‚Ή38,500), the remaining budget (β‚Ή12,000 β†’ β‚Ή1,500), the per-day average (β‚Ή7,700), the day table, the chart and the status badge β€” recomputed: badgeStatus, dayChart, dayTable, note, statPerDay, statRemaining, statTotal. The client computed none of those numbers.

Inference: one boring code path

const response = await fetchImpl(`${settings.baseUrl}/chat/completions`, {
  method: 'POST',
  headers: { 'content-type': 'application/json', authorization: `Bearer ${settings.apiKey}` },
  body: JSON.stringify({
    model: settings.model, messages, stream: false,
    temperature: settings.temperature, max_tokens: settings.maxTokens,
    response_format: { type: 'json_object' },
  }),
  signal: AbortSignal.timeout(90_000),
});

Particle.ai, LM Studio, Ollama and Gemini's OpenAI-compatible endpoint all speak this. Presets are data in the client; Custom covers everything else. Switching providers is a Settings click, not an edit.

Three details that bit:

  • max_tokens needs headroom. Reasoning models spend output budget on hidden thinking before writing any JSON. 8192 truncated surfaces regularly; 32768 doesn't. A truncated surface is not a small failure β€” it's an unusable turn.
  • The key never goes to the browser. Settings stores it in localStorage and sends it to the server; nothing echoes it back, and /api/test redacts it from upstream error bodies before returning them.
  • Fence tolerance is parsing, not repair. Some models wrap JSON in

```json

. Extracting the object before validation is fine; the payload is still validated strictly.

The renderer, and one type trick that prevents drift

// client/src/renderer/registry.ts
export const registry: Record<ComponentName, ComponentType<WidgetProps>> = {
  Column, Row, Text, Slider, Toggle, Table, BarChart, Stat, Badge, Button,
};

export function resolveComponent(name: string): Resolved {
  const decision = decide(name);                       // pure allowlist check
  if (decision.kind === 'unsupported') return { kind: 'unsupported', name };
  const Widget = registry[name as ComponentName];
  return Widget ? { kind: 'widget', name, Widget } : { kind: 'unsupported', name };
}

Record<ComponentName, …> is exhaustive: add a name to the catalog and the client fails to compile until a widget exists for it. The allowlist can't silently lag the server's.

An off-catalog name renders a red card and is logged:

UNSUPPORTED COMPONENT: HologramMap
id map Β· refused by the client allowlist Β· logged to the inspector

The renderer walks the tree from the single unreferenced root, guards depth, and renders missing children as a placeholder rather than throwing. allowlist.ts is a separate pure module specifically so the verifier can execute the client's decision in Node β€” testing the real function instead of re-implementing it in the test.

Verification you can't satisfy by cheating

npm run verify runs 11 checks over the four that matter, against a live model. The interesting part is what each one refuses to accept:

  • Different prompts β†’ different trees. Compares sorted multisets of component types, not response codes. A 200 OK with a canned tree passes the naive version and fails this.
  • Same prompt twice β†’ different tree, identical schema. Proves generation rather than a template, while asserting every component is catalogued.
  • A drag changes β‰₯2 other components' numbers. Uses the numeric-leaf definition above; ignoring min/max/step and the touched id.
  • Bogus components are refused twice over β€” the client allowlist function refuses the name, the server guardrail refuses an off-catalog payload, and an engineered prompt tests the live path. Then /api/health proves nothing crashed.
  • No renderable response contains <script, onclick or javascript: β€” with the diagnostic raw/detail fields stripped, because those deliberately echo a refused payload into the inspector, where React renders it as escaped text in a <pre>.

And the honest caveat, stated in the README too: this suite is model-dependent and will occasionally fail. A run earlier today failed check 2 because the model put two components in one body during an interact turn. The verifier retries refused turns the way the app's retry button does, but if you weaken a guardrail to make the suite green you've deleted the thing you were testing.

npm run verify:screenshots drives the same assertions in headless Chrome over CDP β€” send a prompt, wait for the turn, drag a slider, read displayed numbers out of the DOM before and after, then fire the injection attempt at the running app and assert nothing markup-shaped reached [data-surface-root].

Things that broke, so you don't have to

Symptom Cause Fix
Body has already been used on every POST request.clone().json() then .text() in the Worker read text() once, parse it, forward the same string
ERR_UNSUPPORTED_TYPESCRIPT_SYNTAX: parameter property Node's strip-only mode rejects constructor(readonly code: string) declare fields explicitly in the class
Cannot find module '…/shared/catalog' type-stripping needs explicit .ts specifiers import '../../shared/catalog.ts' in server-side modules
npm ERESOLVE on install agents declares an optional peer on react@^19; the client is React 18 .npmrc with legacy-peer-deps=true, and declare the non-optional peers (@modelcontextprotocol/*, zod) explicitly
Verifier screenshots showed the wrong session CDP script attached to a Chrome already listening on the shared debug port β€” someone else's session allocate an exclusive free port per run
A drag "changed 12 numbers" that meant nothing DOM probe compared .tabular-nums by index, including the dragged widget's own readout tag each number with whether it lives inside a slider card and exclude those
Inspector tab click silently did nothing labels are lowercase in the DOM with CSS capitalize match case-insensitively
Interact turns refused ~30% of the time slider out-of-range and stray-key slips classified as contract violations reclassify emission defects as retryable

What I'd build next

  1. Per-component patches (patchComponents / removeComponents) so a drag sends three components instead of twenty-three.
  2. Streaming surfaces β€” assemble the tree as the model emits it.
  3. Server-checked arithmetic: let the model attach a formula to a derived component and have the host evaluate it, so a wrong total becomes a visible failure instead of a plausible-looking one.
  4. TextField / Select / DateRange β€” the catalog's biggest gap, and the four-file checklist makes adding one mechanical.
  5. Turn scrubber β€” /api/session/:id already returns the whole turn log, so replaying turn n reconstructs any earlier surface for free.

Try it

npm install && cp .env.example .env && npm run dev

Ask for a trip plan, drag a cost slider, and watch the inspector's diff tab list which component ids the agent decided to recompute. Then ask it to put a <script> in a Text prop and read the refusal. Both paths are the app.

Clay 1

Clay 2

Clay 3

Code & more: https://www.dailybuild.xyz/project/270-clay

πŸ“° Read the original article on Dev.to AI

Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes β€” full credit and traffic to the original publisher.