Building Clay: an A2UI host where the model emits interface instead of prose
Technical notes from building an agent that renders live React widgets, keeps per-session state in a Durable Object, and refuses any UI it doesn't recognize. The problem with chat Every agent product today
Technical notes from building an agent that renders live React widgets, keeps per-session state in a Durable Object, and refuses any UI it doesn't recognize.
The problem with chat
Every agent product today has the same output medium: a paragraph. Ask for a trip budget and you get a wall of text containing numbers you cannot manipulate. Ask what happens if the hotel costs more and you re-prompt, and hope it recomputes consistently.
The interesting alternative is to let the model emit interface β a declarative description of widgets bound to a data model β and have the host render it as real controls. That's the A2UI idea: agent-to-UI. The model never sends code. It sends JSON that names components from a catalog the host already owns.
Clay is a complete implementation of that loop:
prompt βββΊ model βββΊ A2UI message βββΊ React renderer βββΊ live widgets
β
widget event
βΌ
agent recomputes dependents βββΊ patch
The whole app is three panes: a chat rail, a generated surface, and an inspector showing the exact messages, because in this design the protocol is the product.
Why "declarative, not code" is the actual security model
The tempting shortcut is to let the model emit HTML or JSX and inject it. Then your sandbox is your own diligence, and one bad <script> is a bad day.
With a component catalog, the model's vocabulary is ten nouns:
Column Row Text Slider Toggle Table BarChart Stat Badge Button
Anything else cannot render, by construction. There is no dangerouslySetInnerHTML anywhere in the client, no eval, no dynamic import. Props arrive as data and get spread into components the bundle already contains. The model can ask for a Slider; it cannot ask for a <script>.
That turns "is the output safe?" from a scanning problem into a type problem.
The contract
One JSON object per turn, four allowed top-level keys:
{
"surfaceUpdate": {
"surfaceId": "s1",
"components": [
{ "id": "root", "component": { "Column": { "children": ["hdr", "budget", "table", "chart"] } } },
{ "id": "hdr", "component": { "Text": { "text": "5-day trip plan", "variant": "h2" } } },
{ "id": "budget", "component": { "Slider": { "label": "Budget", "min": 10000, "max": 80000,
"step": 1000, "value": 40000, "bind": "budget" } } },
{ "id": "table", "component": { "Table": { "columns": [ β¦ ], "rows": [ β¦ ] } } },
{ "id": "chart", "component": { "BarChart": { "series": [{ "label": "Day 1", "value": 7200 }],
"unit": "INR" } } }
]
},
"dataModelUpdate": { "surfaceId": "s1", "dataModel": { "budget": 40000, "days": [ β¦ ] } },
"text": "one short sentence"
}
Two details do most of the work.
The tree is flat and id-linked. Column.children holds ids, not nested objects. Flat lists are far easier for a language model to emit correctly than deeply nested JSX-shaped trees, and patching a node is a replace-by-id.
bind is a promise. A widget's value must equal the value stored at its dotted path in the data model. That single invariant is what makes the surface live rather than decorative: the slider isn't a picture of a number, it's a handle on the model.
Architecture
ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β BROWSER β
β App.tsx ββΊ useSurface() ββΊ SurfaceRenderer.tsx β
β β state, phases, resolve via registry.ts β
β β optimistic bind βΌ β
β β writes Slider Table BarChart β¦ β
ββββββββββββββββ¬ββββββββββββββββββββββββββββββββββββββββββββ
β /api/* (Vite dev proxy)
ββββββββββββββββββββββ΄βββββββββββββββββββββ
βΌ βΌ
ββββββββββββββββββββββββββββ ββββββββββββββββββββββββββββ
β RUNTIME=workers :8787 β β RUNTIME=node :3001 β
β agent/index.ts β β server-node/index.ts β
β Worker fetch() β β node:http β
β session id β DO stub β β session id β Map β
β ClayAgent β β β
β this.state / setState() β β β
β SQLite per session β β β
ββββββββββββββ¬ββββββββββββββ ββββββββββββββ¬ββββββββββββββ
βββββββββββββββββββ¬ββββββββββββββββββββββ
βΌ
agent/turn.ts (shared turn engine)
build messages β call model β classify reply
β apply to surface β measure dependents
βΌ
agent/inference.ts (plain fetch)
POST {baseUrl}/chat/completions Β· stream:false
response_format:{type:"json_object"}
Both runtimes import the same turn engine, prompt and guardrail. The fallback reimplements exactly one thing: who owns state. That's the point of the split β RUNTIME=node proves the Durable Object isn't load-bearing for the protocol, only for persistence.
The Cloudflare Agents SDK, from the actual API surface
The SDK's state API is the piece that changes between releases, so the installed package's type declarations were the ground truth. What [email protected] gives you:
import { Agent } from 'agents';
type SessionState = {
sessionId: string; surfaceId: string;
components: A2UIComponent[]; dataModel: Record<string, unknown>;
turns: TurnRecord[]; updatedAt: number;
};
export class ClayAgent extends Agent<ClayEnv, SessionState> {
initialState: SessionState = { /* β¦ */ };
async onRequest(request: Request): Promise<Response> {
const path = new URL(request.url).pathname;
if (path === '/api/chat') return this.handleChat(request);
if (path === '/api/interact') return this.handleInteract(request);
return this.json({ error: 'not_found', detail: path }, 404);
}
}
Five things worth knowing before you write against it:
-
Implement
onRequest, don't overridefetch. The base class'sfetch()resolves dynamic sub-agent paths and then delegates to your hooks. Overriding it skips that machinery. -
setState()is a full replace, synchronous, returnsvoid. Forgetting{...this.state}silently wipes fields. There is no merge on the object (only on workflow steps), so every write spreads. -
SQLite storage is enabled by the migration, not a binding flag.
new_sqlite_classesinwrangler.jsoncβ there is nodurable_objects.bindings[].sqlitekey:
{
"compatibility_flags": ["nodejs_compat"],
"durable_objects": { "bindings": [{ "name": "CLAY_AGENT", "class_name": "ClayAgent" }] },
"migrations": [{ "tag": "v1", "new_sqlite_classes": ["ClayAgent"] }]
}
After a few sessions the evidence is on disk: one .sqlite file per session,
under .wrangler/state/v3/do/clay-agent-ClayAgent/.
-
wrangler devdoes not promote process env intoenv..envhas to be parsed and passed explicitly, which is whatscripts/run-agent.mjsdoes:
for (const [key, value] of Object.entries(bindings(envFile.values))) {
args.push('--var', `${key}:${value}`);
}
Same script pins XDG_{CONFIG,CACHE,DATA}_HOME into the project so local state never escapes into your home directory.
-
There is no
sessionIdconcept for plain agents. The Durable Object's name is your session id. You mint it, you put it in the URL or a header, andnamespace.get(namespace.idFromName(sessionId))routes to it.
The Worker entry owns the public paths and forwards the body once:
// a Request stream cannot be consumed twice β read text, parse, forward the same text
const text = await request.text();
const body = text ? JSON.parse(text) : {};
const sessionId = clientSessionId(request, body) ?? newSessionId();
const upstream = await stubFor(env, sessionId).fetch(
new Request(url.toString(), { method: request.method, headers: request.headers, body: text })
);
(That line about streams is not a footnote. My first version called
request.clone().json() and then .text(), and every /api/chat returned a 500 with Body has already been used.)
Deliberately unused: the SDK's WebSocket transport, agents/react, AI SDK chat helpers, MCP. Updates ride on the HTTP response, which keeps the protocol inspectable β and inspectable is the whole demo.
The hard part isn't rendering, it's the guardrail taxonomy
Validating the model's reply with zod is easy. Deciding what to do about an invalid reply is where the design lives.
A small model fumbles JSON constantly. Observed failure modes, all real:
- truncated output mid-object
-
{"key, ": "total"}β a stray comma inside a quoted key - a component entry with no
idat all - two components in one
componentbody - a
Sliderwhosevaluesits outside its ownmin/max -
bind: "trip.travellers"while the data model wrotetravellersflat
The naive answer is "retry everything." But then a model that invents
HologramMap gets retried until it behaves, and your unsupported-component path never fires β which is one of the things the app is supposed to demonstrate.
So the split is by kind of failure, not by severity:
| Kind | Examples | Response |
|---|---|---|
| emission defect | truncated JSON, missing id, unrecognized key, out-of-range slider value, two components in one body |
re-ask, up to 3 attempts, feeding the parser's complaint and the model's own previous answer back |
| contract violation | unknown component name, forbidden prop, injection pattern, duplicate ids, dangling children, two roots |
refuse, final β invalid_surface, never retried |
The discriminator is the zod issue code:
const result = a2uiResponseSchema.safeParse(parsed);
if (!result.success) {
const issue = result.error.issues[0];
// A prop the schema does not know, a missing prop, or a component body that
// matches no variant is the model fumbling JSON β an emission defect, so the
// loop asks again. Anything reported by superRefine (duplicate id, dangling
// child, two roots, a bind with no data behind it) is a contract violation
// and stays final.
const defectCodes = new Set(['unrecognized_keys', 'invalid_type', 'invalid_union']);
return {
kind: defectCodes.has(String(issue?.code)) ? 'malformed' : 'rejected',
detail: issuePath(result.error),
};
}
And the re-ask, which never touches the payload:
conversation = [
...conversation,
{ role: 'assistant', content: raw.slice(0, 12_000) },
{ role: 'user', content: `That reply was rejected by the A2UI parser: ${emission.detail}. ` +
`Re-emit the entire A2UI JSON object. Every entry in "components" must be an object with ` +
`BOTH "id" (a non-empty string) and "component" (exactly one catalogued name with its props). ` +
`Do not truncate, do not elide, do not explain. JSON only.` },
];
The rule that makes this honest: the agent loop may ask again; it may not patch. A refusal returns the model's own text (raw, first 300 chars) and the exact zod path, and the previous surface stays on screen. The inspector shows what was refused and why. Fixing the model's JSON locally would be a repair, and a repaired payload is no longer evidence of anything.
The catalog itself is one object, shared/catalog.ts, read three ways:
serialized into the system prompt, wrapped in zod schemas server-side, and
imported as the client's render allowlist. They cannot drift.
Making a drag actually change other widgets
This is the acceptance test that separates a live surface from a picture of one. Move a slider and the table, chart and stat cards must show new arithmetic.
Sequence:
you drag Slider "stay" 2500 β 6000
ββ client: setPath(dataModel, "stay", 6000) optimistic, thumb tracks instantly
ββ phase β "patching"
βΌ
POST /api/interact { surfaceId, componentId, value }
ββ server writes bind into the dataModel BEFORE asking
β (the model's job is downstream arithmetic, not re-deriving what the slider says)
ββ prompt carries: the event, the interactive widgets, the full current dataModel
ββ model returns the FULL component list + dataModelUpdate
ββ recomputedIds(previous, next, touchedId)
ββ if only the touched widget moved β ask once more, pointing at its own reply
βΌ
200 { recomputed: ["chart","statTotal","table","statRemain", β¦] }
Two decisions matter.
Seed the bound value server-side. Asking the model to both apply the event and recompute everything is how you get half-updated surfaces. Applying bind is protocol semantics, not guessing at UI; the model's job becomes purely the arithmetic downstream.
Measure the right numbers. "Did dependents change?" needs a definition that can't be gamed:
// shared/a2ui.ts β numeric leaves, ignoring the control's own range
export function numericFields(component: A2UIComponent): Record<string, number> {
const flat = flattenProps(componentProps(component));
const out: Record<string, number> = {};
for (const [key, value] of Object.entries(flat)) {
// min/max/step describe the control, not a computed output.
if (/(^|\.)(min|max|step|maxValue)$/.test(key)) continue;
if (typeof value === 'number' && Number.isFinite(value)) out[key] = value;
}
return out;
}
Without that exclusion, a slider drag "changes" the moved widget's own value and a build that recomputes nothing still passes. The comparison also excludes the touched component id entirely.
Finally, syncBoundValues() re-writes every interactive widget's value from the data model after a patch lands. That's not repair either β it's the bind invariant enforced, so a stale echo in the model's reply can't rewind what the user just dragged.
In a real browser, one drag of Hotel per night βΉ2,500 β βΉ6,000 moved the trip total (βΉ28,000 β βΉ38,500), the remaining budget (βΉ12,000 β βΉ1,500), the per-day average (βΉ7,700), the day table, the chart and the status badge β recomputed: badgeStatus, dayChart, dayTable, note, statPerDay, statRemaining, statTotal. The client computed none of those numbers.
Inference: one boring code path
const response = await fetchImpl(`${settings.baseUrl}/chat/completions`, {
method: 'POST',
headers: { 'content-type': 'application/json', authorization: `Bearer ${settings.apiKey}` },
body: JSON.stringify({
model: settings.model, messages, stream: false,
temperature: settings.temperature, max_tokens: settings.maxTokens,
response_format: { type: 'json_object' },
}),
signal: AbortSignal.timeout(90_000),
});
Particle.ai, LM Studio, Ollama and Gemini's OpenAI-compatible endpoint all speak this. Presets are data in the client; Custom covers everything else. Switching providers is a Settings click, not an edit.
Three details that bit:
-
max_tokensneeds headroom. Reasoning models spend output budget on hidden thinking before writing any JSON. 8192 truncated surfaces regularly; 32768 doesn't. A truncated surface is not a small failure β it's an unusable turn. -
The key never goes to the browser. Settings stores it in
localStorageand sends it to the server; nothing echoes it back, and/api/testredacts it from upstream error bodies before returning them. - Fence tolerance is parsing, not repair. Some models wrap JSON in
```json
. Extracting the object before validation is fine; the payload is still validated strictly.
The renderer, and one type trick that prevents drift
// client/src/renderer/registry.ts
export const registry: Record<ComponentName, ComponentType<WidgetProps>> = {
Column, Row, Text, Slider, Toggle, Table, BarChart, Stat, Badge, Button,
};
export function resolveComponent(name: string): Resolved {
const decision = decide(name); // pure allowlist check
if (decision.kind === 'unsupported') return { kind: 'unsupported', name };
const Widget = registry[name as ComponentName];
return Widget ? { kind: 'widget', name, Widget } : { kind: 'unsupported', name };
}
Record<ComponentName, β¦> is exhaustive: add a name to the catalog and the client fails to compile until a widget exists for it. The allowlist can't silently lag the server's.
An off-catalog name renders a red card and is logged:
UNSUPPORTED COMPONENT: HologramMap
id map Β· refused by the client allowlist Β· logged to the inspector
The renderer walks the tree from the single unreferenced root, guards depth, and renders missing children as a placeholder rather than throwing. allowlist.ts is a separate pure module specifically so the verifier can execute the client's decision in Node β testing the real function instead of re-implementing it in the test.
Verification you can't satisfy by cheating
npm run verify runs 11 checks over the four that matter, against a live model. The interesting part is what each one refuses to accept:
-
Different prompts β different trees. Compares sorted multisets of component types, not response codes. A
200 OKwith a canned tree passes the naive version and fails this. - Same prompt twice β different tree, identical schema. Proves generation rather than a template, while asserting every component is catalogued.
-
A drag changes β₯2 other components' numbers. Uses the numeric-leaf
definition above; ignoring
min/max/stepand the touched id. -
Bogus components are refused twice over β the client allowlist function refuses the name, the server guardrail refuses an off-catalog payload, and an engineered prompt tests the live path. Then
/api/healthproves nothing crashed. -
No renderable response contains
<script,onclickorjavascript:β with the diagnosticraw/detailfields stripped, because those deliberately echo a refused payload into the inspector, where React renders it as escaped text in a<pre>.
And the honest caveat, stated in the README too: this suite is model-dependent and will occasionally fail. A run earlier today failed check 2 because the model put two components in one body during an interact turn. The verifier retries refused turns the way the app's retry button does, but if you weaken a guardrail to make the suite green you've deleted the thing you were testing.
npm run verify:screenshots drives the same assertions in headless Chrome over CDP β send a prompt, wait for the turn, drag a slider, read displayed numbers out of the DOM before and after, then fire the injection attempt at the running app and assert nothing markup-shaped reached [data-surface-root].
Things that broke, so you don't have to
| Symptom | Cause | Fix |
|---|---|---|
Body has already been used on every POST |
request.clone().json() then .text() in the Worker |
read text() once, parse it, forward the same string |
ERR_UNSUPPORTED_TYPESCRIPT_SYNTAX: parameter property |
Node's strip-only mode rejects constructor(readonly code: string)
|
declare fields explicitly in the class |
Cannot find module 'β¦/shared/catalog' |
type-stripping needs explicit .ts specifiers |
import '../../shared/catalog.ts' in server-side modules |
| npm ERESOLVE on install |
agents declares an optional peer on react@^19; the client is React 18 |
.npmrc with legacy-peer-deps=true, and declare the non-optional peers (@modelcontextprotocol/*, zod) explicitly |
| Verifier screenshots showed the wrong session | CDP script attached to a Chrome already listening on the shared debug port β someone else's session | allocate an exclusive free port per run |
| A drag "changed 12 numbers" that meant nothing | DOM probe compared .tabular-nums by index, including the dragged widget's own readout |
tag each number with whether it lives inside a slider card and exclude those |
| Inspector tab click silently did nothing | labels are lowercase in the DOM with CSS capitalize
|
match case-insensitively |
| Interact turns refused ~30% of the time | slider out-of-range and stray-key slips classified as contract violations | reclassify emission defects as retryable |
What I'd build next
-
Per-component patches (
patchComponents/removeComponents) so a drag sends three components instead of twenty-three. - Streaming surfaces β assemble the tree as the model emits it.
-
Server-checked arithmetic: let the model attach a
formulato a derived component and have the host evaluate it, so a wrong total becomes a visible failure instead of a plausible-looking one. -
TextField/Select/DateRangeβ the catalog's biggest gap, and the four-file checklist makes adding one mechanical. -
Turn scrubber β
/api/session/:idalready returns the whole turn log, so replaying turn n reconstructs any earlier surface for free.
Try it
npm install && cp .env.example .env && npm run dev
Ask for a trip plan, drag a cost slider, and watch the inspector's diff tab list which component ids the agent decided to recompute. Then ask it to put a <script> in a Text prop and read the refusal. Both paths are the app.
Code & more: https://www.dailybuild.xyz/project/270-clay
Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes β full credit and traffic to the original publisher.

