Dev.to AI šŸ¤– Ai šŸ‘ 0 šŸ“– 4 min read

How i verify what my coding agent says it fixed

A few months ago I was building a piece of software with Claude Code and at some point I ended up with 72 things to check by hand, the agent had finished, the tests were green, and I had no idea which of those 72 things

A few months ago I was building a piece of software with Claude Code and at some point I ended up with 72 things to check by hand, the agent had finished, the tests were green, and I had no idea which of those 72 things I had actually tried and which ones I had only thought about trying. Some notes in a text file, some in my head, a couple in the terminal, I'm pretty sure a few slipped through.

That's when I realized the hard part of working with coding agents isn't writing code anymore, it's knowing if what they wrote really works.

How I work with the agent
My workflow is nothing special, but it matters for the rest of the story. First I write the plan together with the agent, features, architecture, the decisions I care about, then I ask it to split the work into phases small enough to review, and for every phase I explicitly ask it to write tests, not at the end but as part of the phase.

Before moving to the next phase I read the critical parts of the code and I test things by hand. The first parts the agent handles pretty well on its own, the manual testing is where things used to fall apart.

Why green tests aren't enough
Automated tests check what the agent expects to happen, but the agent wrote the code and very often the tests too, so they share the same blind spots.

Tests won't tell you that the page is blank on Safari but fine on Chrome, or that the error message shows up but the form still submits, or that "remember me" survives a browser restart but not a server restart. Sometimes it technically does what you asked, just not what you meant.

You only find these things by using the software, and when the list gets long you lose track, you test the same thing twice, you skip another one, and a week later the bug you were sure was fixed is back.

What I wanted
I didn't want another test framework, I wanted something much simpler. The agent knows what it touched, so it should be the one writing the list of what to verify, I should be able to go through that list while I'm testing, often on my phone and not in the terminal, and my results should go back to the agent in a way it can act on, without copy-pasting anything.

So I built it, it's called palmtop.

How palmtop works
You ask the agent to do something and to put the checks on palmtop:
implement phase 2 and put the checks on palmtop
The agent does the work, then writes a checklist of what to verify, it shows up live on your paired devices (phone, tablet or browser) and you get a notification.

You try each check on your project and mark it: āœ“ if it works, āœ• if it doesn't, with a short note on what happens, or ? if you need more info before you can say.

A real example from a login flow:

  • āœ“ Wrong password shows an error
  • āœ• /settings while logged out goes to /login → blank page on Safari
  • ? "Remember me" survives a restart → browser or server restart?

When you're done you tell the agent "I reviewed the checks on palmtop", or just "Done", and it reads your review. For every āœ• it fixes the problem and asks you to check that item again, for every ? it just explains, without changing the code. Re-checking matters more than it seems, if a fix broke something else you usually catch it in the next round.

You can also add your own checks from the phone when you notice something the agent didn't think about, it reads them as requests together with the review.

Where the data lives
This was important to me. palmtop runs a small daemon on your computer, next to the agent, checklists, reviews and notes stay in ~/.palmtopai, and your devices reach it through a relay that only forwards messages and doesn't store anything on disk.

It's a preview, so to be honest: messages go over TLS but they're not end-to-end encrypted yet, that's the next big thing on the roadmap. Until then, don't put secrets in your checklists.

Trying it
It works with Claude Code and Codex out of the box, and with any other agent that can run shell commands, you need macOS, Linux or WSL and Node.js 22+.

curl -fsSL https://palmtop.mtwa.it/install.sh | bash

The installer sets up your agents and then shows a QR code to pair your phone, it never uses sudo and you can read the script before running it.

What I learned so far
Just writing down the checks is half the value, having the list makes me test more carefully even before I mark anything. Short notes are enough, "blank page on Safari" gives the agent more to work with than a long explanation. And the ? is underrated, half of my doubts weren't bugs at all, just things I didn't understand about what the agent did.

palmtop is free, no account needed, and I'm building it on my own in my spare time. If you use coding agents and you ever lost track of what you already tested, I'd really like you to try it and tell me where it breaks, feedback from someone other than me is exactly what it needs right now.

Someone on Reddit already suggested attaching screenshots to failed checks, and it's on the list. What would you add?

šŸ‘‰ palmtop.mtwa.it

šŸ“° Read the original article on Dev.to AI

Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.