Dev.to Security 🔐 Cybersecurity 👁 0 📖 4 min read

Vibe coding risks I didn't see coming as a non-developer

I've done SEO since 2009 and I've never written a line of code by hand. This year I still shipped real SaaS products by vibe coding, with Claude Code doing the typing. The question I get asked most is some version of "o

I've done SEO since 2009 and I've never written a line of code by hand. This year I still shipped real SaaS products by vibe coding, with Claude Code doing the typing.

The question I get asked most is some version of "ok but can a non-developer really build like a developer?"

Honestly, no. I can ship things. What I can't do is read what I ship, and pretty much every risk below came from that gap.

Here are the ones that actually bit me, roughly in the order I ran into them.

1. Green tests are not proof

On one build all the end to end tests passed and I was ready to ship.

Then I asked two AI reviewers from different companies to look at the same code. They came back with 14 more issues, including concurrency bugs, a few silent failures and a credential that could leak. 7 of the 14 were only caught by one of the two.

That part stuck with me. With just one reviewer I'd have missed 3 or 4 real problems and had no clue.

Why did the tests pass? The AI wrote them with the same understanding it used to write the code. So they checked what it expected, and nobody checked if it expected the right thing.

2. The AI grading its own homework

Claude Code can review its own plan before it writes any code, and to be fair that step does catch real stuff.

But when I sent one plan to outside models they found things Claude's review hadn't, and two were serious. The daily cost cap never actually read the cost, so it looked like a safety net without being one. And on a fresh install the tool would have made a draft reply for every unread email in the inbox. Dozens of drafts nobody asked for, plus a surprise API bill.

Neither was a coding bug. The code would have done exactly what the plan said, the plan was just wrong.

3. The bug is often in the belief, not the code

A lot of the worst mistakes in my builds had no bad code in them. The code did the wrong thing perfectly.

One report said zero records had a certain field when ten did, because the query looked in the wrong place. Another time a field was on almost every record but a lot of them were empty. And once I gave users a setup command that couldn't run at all. It looked fine. Nobody had actually run it.

Reading the diff won't catch any of that. What catches it is opening a real record, running the real command, looking at the real screen.

4. More AIs agreeing doesn't make them right

After that first scare I figured the fix was easy: ask more models, trust whatever they agree on.

So I tested it. I pulled a random sample of 20 findings from AI reviews of 30 public apps and had two models from different companies check each one against the code. On the findings both checkers agreed on, about 44% weren't real.

Agreement didn't help much either. Findings raised by several models were real 5 times out of 10. Findings from just one model were real 4 times out of 5. Tiny sample, so treat it as a warning and not a stat. My takeaway: a second model is great at spotting what the first one missed, but two models agreeing doesn't make it true.

5. My own checker made the same mistake

This one humbled me.

Full disclosure, I now build a code review tool, so keep that bias in mind for this part. It has a step that fact checks claims before anything gets built on top of them.

One day that step decided a real npm package wasn't valid. It had looked at six web pages and not one was about that package. Then it threw away four good findings, one of them about passwords stored in plain text, and told the user none of them had passed the audit.

The whole tool is built on the idea that one model shouldn't check its own work. And there was a single checker marking its own homework, inside my own tool. I'd applied the rule to the models and forgot the plumbing around them.

Someone outside noticed. Usually that's how it goes.

So can a non-developer build like a developer?

Not like one, no. Developers read code. I'm not good at that and I've stopped pretending I am.

What I can do is check evidence. These days before I trust anything I run every command I publish, exactly how a user would paste it. I open one real record before I believe a count. I look at the actual screen, not just the test results. I read data back out of the database after it's saved. And I get a second opinion from a model that didn't write the code, and treat it as input, not orders.

It's slower than pure vibing. Still a lot faster than cleaning up after the version of me who shipped those bugs.

If you read code for a living I'd really like to know: what's the first thing you check when someone hands you AI written code?

📰 Read the original article on Dev.to Security

Originally published by Dev.to Security. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.