Dev.to AI 🤖 Ai 👁 0 📖 4 min read

Your agents need a desk, not another chat pane

Most agent demos still look the same. You type into a box. Something happens somewhere else. A wall of tool calls scrolls by. Then a summary appears and you are supposed to trust it. That model works for narrow tasks. I

Most agent demos still look the same. You type into a box. Something happens somewhere else. A wall of tool calls scrolls by. Then a summary appears and you are supposed to trust it.

That model works for narrow tasks. It falls apart the moment the agent has to use the same computer you use: browsers, sheets, tickets, messy UIs that were never designed for automation. If you cannot see the pointer, you are not collaborating. You are waiting.

I build product and engineering for an infinite canvas OS. The design bet we keep returning to is simple: agents and humans should share a spatial workspace. Same board. Separate cursors. Interruptible work.

Here is what that actually implies when you try to ship it.

Chat is a log. A desk is an address space.

Chat is great for intent. "Book the cabin for Oct 8–13 if it is under $250." Fine.

What chat is bad at is state you can point at. Developers already know this from debugging. A stack trace is not the same as an open debugger with the heap still live. Agents that only report back in text force you to reconstruct the world from narration.

A spatial workspace flips that. Windows sit next to each other. The agent's browser tab is a place, not a sentence. Your notes stay beside the sheet it is filling. When something looks wrong, you do not ask for another summary. You look.

That is closer to how humans already work together on a real desk. One person owns the left side of the table. Another owns the right. Nobody waits for a paste into Slack.

Agents that steal your mouse are not teammates

On a normal desktop there is one pointer. If an agent uses the computer through that pointer, two bad outcomes show up immediately.

Either the agent takes your mouse and you sit there while a spinner owns the session. Or the agent runs out of sight and you only see the after-action report.

Neither feels like pair programming. Pair programming works because you can watch the other person's hands.

So give the agent its own cursor. Label it. Color it. Keep yours. Now you can keep typing in a sheet while the agent works in a booking page next door. Collision becomes a layout problem, not a trust problem.

We ended up treating each agent seat as a mouse and keyboard inside the compositor. Add an agent, drop a new cursor into a new window. Up to sixteen on one canvas in our case. Helpers get attached to a lead, one task per window, so the board does not turn into a pile of anonymous activity.

Interruptibility is the feature people actually want

The first time someone watches an agent click the wrong listing, they do not ask for a better model. They ask for a stop button that works mid-click.

Pause. Take the window back. Type a correction without restarting the whole run.

Those controls sound small. They are the difference between "autonomy" and "supervision." Autonomy without interruption is just a long-running script with better prose.

A useful pattern: the lead agent writes a plan you can edit before it goes. Helpers execute pieces in view. A separate check reads a fresh screenshot before anything is marked done. "Done" stops meaning "the model said it finished."

If you are designing an agent OS or agentic OS layer, bake interruptibility in early. Retrofitting it onto a fire-and-forget runner is painful.

Shared canvas beats shared screen for multiplayer work

Developers already hate the remote debugging ritual. One person shares. Everyone else narrates. "Scroll up." "No, the other tab." "Wait, who has control?"

Screen share is a video of a desk. A shared canvas is the desk.

On a shared spatial workspace, every cursor works. A teammate can type into a window that still runs on your machine. Agents belonging to a teammate show up labeled as theirs. The room stays when the call ends.

That last part matters more than it sounds. Most collaborative sessions die with the Zoom. The board evaporates. Tomorrow you rebuild context from memory and a half-updated doc.

Saved layout is underrated infrastructure. Apps, positions, tabs, still there in the morning. Boring. Extremely useful.

What I would require before calling something "human + agent workspace"

Skip the marketing labels. Ask for these concrete properties:

  • Visible agent input (cursor, rings, typing), not only a transcript
  • Agent windows that land beside yours instead of covering them
  • Pause / steer / reclaim without restarting the job
  • Verification that is not the same process that did the work
  • A place for your own notes and apps to remain in view while agents run
  • Multiplayer that does not serialize everyone onto one pointer

You can build pieces of this with existing tools. Browser automation plus screen share plus a chat sidebar gets you part of the way. The failure mode is the same every time: the human loses the spatial picture, and the agent becomes a black box again.

We built Novastart around the opposite assumption — that the computer itself should be the canvas where people, apps, and agents stay in sight. Install as an app on Windows or Intel Mac to try it, or as an OS on a dedicated PC if you want the full speed.

Closing note for people shipping agent products

Stop asking whether users "trust AI." Ask whether they can see what it is doing, stop it, and keep working in the same space.

If your agent only lives in a pane, you have built a clever assistant. If it sits at the desk with its own hands, you have started building software people can actually co-operate with.

That is the bar I care about.

📰 Read the original article on Dev.to AI

Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.