Dev.to AI šŸ¤– Ai šŸ‘ 0 šŸ“– 8 min read

A Coding Agent Added Jetpack Compose Support to My Android Renderer MCP. I Wouldn't Have Started This Task Myself.

Android UI Renderer MCP is a local MCP server that lets a coding agent render Android UI without an emulator or a device: it gets back a PNG plus a component tree it can query for bounds, text, and state. It already han

A Coding Agent Added Jetpack Compose Support to My Android Renderer MCP. I Wouldn't Have Started This Task Myself.

Android UI Renderer MCP is a local MCP server that lets a coding agent render Android UI without an emulator or a device: it gets back a PNG plus a component tree it can query for bounds, text, and state.

It already handled classic Android UI — XML/View, RecyclerView, Activity + Fragment, overlays, different screen configurations. For XML there's a clear entry point: load the layout resource, apply a fixture, measure, lay out the View hierarchy, draw it, serialize the tree.

XML layout → inflate → fixture with data → measure/layout → PNG + View tree

Jetpack Compose has no equivalent entry point. There's no file to inflate. There's a Kotlin function:

@Composable
fun BookItem(
    book: Book,
    selected: Boolean,
    onClick: () -> Unit
)

Rendering it means understanding its signature, building correct Kotlin values for every parameter (enums, nullable types, data classes, collections, callbacks), resolving the app's theme, and only then calling the composable inside an environment where Compose can actually draw. In a real app, a ViewModel, DI, navigation, and data loading show up right next to it.

That's not "support one more UI format." That's a separate research problem, and it's exactly the kind of problem I usually never got around to starting — not because it's unsolvable, but because the cost of finding out how to approach it could easily exceed the value of the feature itself.

This time I handed that research problem to a coding agent. What follows is what it built, where it broke, and what I had to check before trusting the result.

I rejected the agent's first proposal

The safe option was the obvious one: let every Compose project provide its own renderer entry point — a function that already knows how to build the needed state, wire dependencies, and call the composable, with the MCP just invoking that prepared adapter.

It would have worked. It also would have killed the point of the tool.

The renderer isn't meant to be another test framework bolted onto a project. It's meant to support a loop an agent runs on an existing project with no manual prep:

change UI → render → inspect PNG + structure → fix

If a developer has to hand-write an adapter for every screen first, agent autonomy ends exactly where the real UI begins. So instead of asking "how do we render a composable," I set a constraint on the result:

The MCP must adapt to Kotlin/Compose on its own. The project must not have to write a special renderer entry point.

I didn't know how to implement that. I specified a property the result had to have, and left the research and implementation to the agent.

What render_compose does

The MCP got a new tool. Instead of an XML layout, the agent passes the fully qualified name of a top-level composable plus named JSON arguments:

{
  "function": "io.github.example.ComposeBookItem",
  "arguments": {
    "selected": true,
    "book": {
      "title": "Designing Data-Intensive Applications",
      "status": "AVAILABLE"
    }
  },
  "widthPx": 1280,
  "heightPx": 800,
  "densityDpi": 240
}

From there the renderer:

  1. finds the Kotlin file with the target top-level @Composable;
  2. parses the function's parameters;
  3. checks for extra and missing arguments;
  4. builds typed Kotlin values from the JSON;
  5. generates an ordinary Kotlin call to the composable;
  6. substitutes no-op lambdas for missing callbacks;
  7. finds the project-level AppTheme, or uses an explicitly passed wrapper;
  8. creates a temporary Kotlin probe;
  9. runs it via Gradle/Robolectric;
  10. draws the ComposeView to a PNG.

The project's own sources aren't touched — the temporary probe lives in .android-ui-renderer. Nullable values, enums, data classes, mutable properties, List, and Set are all supported, so the agent never has to pre-build a renderForMcp() function inside the app.

To verify this wasn't just working on a toy composable, I asked the agent to clone an open-source Jetpack Compose sample project end to end. The repo now has two equivalent apps, sample/ (XML/View) and sample-compose/ (Compose), sharing the same book library and the same seven scenarios: a book row, a long title at fontScale = 1.3, a list, master-detail, full screen, a loading state, and a Russian dark theme.

XML/View Jetpack Compose
Full-screen Shelf demo screen rendered from XML/View Same Shelf demo screen rendered from Jetpack Compose

Same scenario, two different UI stacks, both captured on a 1920Ɨ1200 px canvas at 240 dpi. This isn't a pixel-perfect comparison between XML and Compose — it demonstrates that the renderer reproduces equivalent states on the same canvas through two different mechanisms.

The first PNG worked. The tool's contract still didn't.

The most useful failure showed up almost immediately. On a real Compose component, the agent got back a normal, correct-looking PNG. By screenshot-tool standards, that's "done."

But tree serialization broke right after: one of the Compose semantics nodes had no className, which the renderer's existing protocol expected on every tree node. The image existed; a full MCP response didn't.

For a plain screenshot tool, that's a rounding error. For an agent tool, it isn't — an agent doesn't just look at a PNG, it needs to ask the renderer concrete questions afterward:

where is this element? what are its bounds? what text does it have?
is it clickable? selected? checked? what's inside it?

XML answers that through the View tree. For Compose, the renderer now returns a semantics tree instead — bounds, text, contentDescription, enabled, clickable, selected, checked, and child nodes. The agent added a fallback for semantics nodes without a conventional className, rebuilt the renderer, and reran the same request. Here's a fragment of the resulting view-tree.json, with bounds and state the agent can act on next:

{
  "id": "semantics-33",
  "className": "androidx.compose.ui.semantics.SemanticsNode",
  "clickable": true,
  "bounds": { "left": 650, "top": 500, "right": 799, "bottom": 560 },
  "children": [{
    "id": "semantics-35",
    "text": "Reserve",
    "bounds": { "left": 686, "top": 515, "right": 763, "bottom": 545 }
  }]
}

Only after that fix did the tool's actual contract pass — image and structure, not just a picture an agent has to eyeball.

Then a second failure, unrelated to Compose itself

The renderer temporarily injects test dependencies via a Gradle init script. In the new Compose demo project, this hit a strict repositoriesMode: the project's Gradle settings forbade adding a Robolectric dependency that way.

So the UI rendering already worked, but the reproducible open demo didn't build. The agent changed sample-compose's repositories mode to PREFER_PROJECT and left a comment explaining why the renderer needs to add a temporary test dependency at all. This is a familiar shape in agentic work: "add Compose support" doesn't stop at the Compose API — the agent has to walk the whole path to a reproducible result, including Gradle policy that has nothing to do with Compose semantics.

Timeline: under two hours from constraint to working demo

Per the session log: 12:26, first note of the technical problem. 12:41, I rejected the mandatory project-side entry point. ~13:00, first real Compose render — PNG succeeds, semantics tree serialization fails. Then a second successful run, the Gradle policy fix, tests, and all seven scenarios. 14:05, the main commit lands in main: Compose support, tests, Compose demo. 14:09, a follow-up commit after I flagged that the two comparison images had been captured at different canvas sizes.

From stating the constraint to a fixed, public result: about 1 hour 43 minutes.

I'm not turning that into "a human would take N days." I don't know how long this would've taken me, because the honest answer is I probably wouldn't have started it — I'd have had to separately learn how to call an arbitrary composable programmatically, how Compose compiler transformations affect that, how to get a semantics tree, how to run all of it on Robolectric, and how to avoid turning the project into a pile of special test entry points. All of that is researchable. The cost of that research was the actual stop factor, and that's what changed: not the speed of writing code, but the economics of deciding whether an idea is worth trying.

What the agent did, and what I did

Looking only at the diff, my part looks small. The agent designed and implemented render_compose, built the typed Kotlin call generator, built the Compose probe and semantics tree, added unit tests for enums/data classes/mutable state/callbacks, verified against a real Compose component, created the open sample-compose project with seven reproducible scenarios, and fixed both failures above.

I didn't write that code. What I did do:

  • rejected the mandatory project-side adapter and set the actual constraint (MCP adapts, project stays a normal project);
  • required verification against a real, working component instead of a toy composable;
  • caught the canvas-size mismatch in the XML/Compose comparison after the first "success."

That's a real shift in where my time goes. The old chain was study the unfamiliar area → find an approach → implement → debug → test → demo → verify. The new one is state the required property → set constraints → let the agent run research/implementation/tests/demo/fixes → verify it's actually the result I need. The technical complexity didn't disappear — it's still in the code. I just stopped having to walk the whole research path myself before I even get to decide if the idea is worth it.

Limits, stated plainly

The current implementation is deliberately narrow. The renderer works with a top-level composable, not a production navigation flow — no DI, no real data loading, visual state passed explicitly as arguments. Callbacks are replaced with no-op lambdas, so there's no simulated user interaction. Kotlin signature parsing is a custom regex-based parser, not the Kotlin compiler API, so complex overloads, type aliases, and unusual annotations may need more work. Compose returns a semantics tree, not an Android View tree, so identifiers are synthetic. And I haven't verified pixel-perfect agreement with a physical device — I'm not claiming that.

For the task I actually had, this is enough: take a specific visual state, render real Compose UI, get a PNG plus a machine-readable structure, with no adapter pre-written inside the project.

The actual result for me isn't Compose support

The formal result is simple: the renderer now supports both XML/View and Jetpack Compose, with two equivalent sample apps and side-by-side output.

But the thing worth taking away is the shift in how I size up a feature. I used to ask "how long will it take me to even figure out how to do this?" — and that question quietly killed a lot of useful-but-optional ideas before they started. Increasingly the question I can ask instead is "can I state precisely what the result needs to be, and which constraints actually matter?" I didn't know how to implement Compose rendering, and wouldn't have started researching it on my own. I did know exactly what I wanted from the tool. That was enough for the agent to cover the part I didn't know, including the two failures it had to debug on the way to a working public demo.

I keep the full case study, including the detailed session timeline and the complete set of assets, on my site: https://hram.github.io/en/articles/android-ui-renderer-compose/

šŸ“° Read the original article on Dev.to AI

Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.