Dev.to Security πŸ” Cybersecurity πŸ‘ 0 πŸ“– 4 min read

I re-scanned 393 AI-built repos without using AI. 1 in 8 shipped a critical flaw.

In September I published what a language model found across 400 AI-built repositories. The most useful criticism was also the most obvious: a model's judgement is not reproducible, and you cannot audit it. Run it again a

In September I published what a language model found across 400 AI-built repositories. The most useful criticism was also the most obvious: a model's judgement is not reproducible, and you cannot audit it. Run it again and you might get a different answer. Ask "how exactly did you count that?" and the honest reply is "the model decided."

So I ran the whole thing again with deterministic rules. Same frozen corpus, no model, 26 rules, every finding traceable to a rule id and a line number.

The headline number went down. That turned out to be the most interesting part of the exercise.

The corpus

400 public repositories built with Lovable, Bolt and v0, selected by build artefact rather than README text β€” a repo qualifies because it contains .lovable, lovable-tagger or .bolt, not because someone wrote "built with Lovable" in the readme. One repo per owner, forks excluded, up to three files each, frozen on 5 September 2026 so the sample cannot drift underneath the results.

Of 1,185 files, 1,163 were readable across 393 repos. The missing 22 have been deleted from GitHub since the freeze. The whole run takes about 90 seconds and costs nothing, which matters more than it sounds β€” the first pass exhausted an API balance partway through.

What the rules found

309 of 1,163 files β€” 27% β€” had at least one finding. 58 files had something critical, spread across 49 repos. That is 1 in 8 repositories carrying at least one critical flaw, by rules alone.

Cookie set without Secure or SameSite       111
Access-Control-Allow-Origin: *               74
dangerouslySetInnerHTML from a variable      60
RLS policy written as using (true)           39
SQL built by template interpolation          14
innerHTML assigned from a variable            6
Google API key committed in source            3
eval(), weak token randomness, no RLS         7

The finding I did not expect

39 files contained a row-level security policy written as using (true).

This is worse than simply forgetting to enable RLS, because it is the signature of someone who tried. They read that Supabase tables need row-level security. They enabled it. Then they wrote a policy that returns true for every row β€” which passes every check, for every user, every time.

-- Enabled, and completely ineffective
alter table orders enable row level security;
create policy "read orders" on orders for select using (true);

-- What it needs to say
create policy "read own orders" on orders
  for select using (auth.uid() = user_id);

The dashboard shows RLS as enabled. The checkbox is ticked. The table is as open as it was before.

An AI coding tool will happily generate the first version when asked to "add RLS", because it is a literal answer to the question. Nothing in the app breaks, nothing warns you, and the anon key that makes the table readable ships in the frontend bundle by design.

The number that went down

The first study reported 59% of scanned repos with a critical issue. This pass says 12%. Both are true, and the gap is the point.

On the 115 files both passes read:

  • the model flagged something in 113
  • the rules flagged something in 32
  • they agreed on the verdict 30% of the time

Rules catch syntactic patterns β€” a cookie without a flag, a wildcard CORS header, SQL glued together with a template literal. They are blind to everything that needs intent to be read: an endpoint that never checks who is asking, user input reaching a sink three functions later, an error handler returning a stack trace to the browser.

So 12% is a floor, not an estimate. It is the share of repositories where pattern matching with no understanding of the code could prove something was wrong. The real figure is higher. I publish the floor because it is the number I can defend line by line.

This is the trade-off nobody mentions when they publish an LLM-generated security statistic: you get breadth you cannot reproduce, or rigour that systematically undercounts. You do not get both, and which one you want depends entirely on whether someone is going to argue with the number.

One more thing worth noticing

Only three secrets turned up in source files across 1,163 files.

That sounds like good news and is not. Neither pass opens .env files β€” and the earlier study found 18.5% of these repositories had committed one. Secrets in these projects are not pasted into components. They sit in the environment file that went in on the first push, which is also the one place a casual look at the code will never find them.

If you build this way, three checks

  1. Search your SQL for using (true). If one sits on a table holding customer data, this whole post is in your own repository.
  2. Run git log --all -- .env. Any output means the file is in your history. Rotating the keys is the only fix β€” deleting the file is not, because the history keeps the old blob.
  3. Check what a signed-out request can read, rather than trusting that RLS being "on" means it is doing anything. From outside, an empty table and a protected table look identical, so the only real proof is two accounts: sign in as the second and try to read the first one's rows.

Method, per-rule counts and caveats are written up in full here: https://www.vibesafe.info/blog/393-ai-built-repos-rules-scan

Happy to share the per-file and per-finding CSVs if anyone wants to check the working β€” every row carries the repository, file, rule id and line number. I am not publishing the repository names: these are real projects by real people, most of whom have no idea, and a list of vulnerable apps is just a target list.

πŸ“° Read the original article on Dev.to Security

Originally published by Dev.to Security. Aggregated on AIWithGhost for educational purposes β€” full credit and traffic to the original publisher.