Dev.to Security πŸ” Cybersecurity πŸ‘ 0 πŸ“– 4 min read

I leaked my MongoDB password, so I built a tool that rotates it with zero downtime

I pushed a .env file to a public GitHub repo. It held the MongoDB password and the JWT secrets of Collabify, a real-time Kanban app I run on Vercel. Deleting the file doesn't help: it stays in the git history, and bots

I leaked my MongoDB password, so I built a tool that rotates it with zero downtime

I pushed a .env file to a public GitHub repo. It held the MongoDB password and the JWT secrets of Collabify, a real-time Kanban app I run on Vercel.

Deleting the file doesn't help: it stays in the git history, and bots scrape new commits for credentials within minutes. The only real fix is to rotate: replace each secret so the leaked one stops working.

Rotating by hand is where people take their own app down. This post is about the order that avoids that, and the tool I built to follow it every time.

leakfix rotating secrets: the first deploy fails and everything is rolled back, the second run succeeds

The order that keeps the app up

1. prepare   create a NEW credential next to the leaked one (both work)
2. switch    put the new value in the deployment's environment variables
3. redeploy  one production redeploy for all changed variables
4. verify    call a health route that fails if the database is unreachable
5. revoke    delete the leaked credential

Two rules make it safe:

  • Until step 5, everything can be undone. If any step fails, the completed steps are undone in reverse: variables get their old values back, production is redeployed on the old config, and the new credential is deleted. The app keeps running on the old credential the whole time.
  • Once step 5 starts, never roll back. Rolling back would mean pointing production at the leaked credential again. If revoking fails, the tool says what to finish by hand instead.

I found the second rule the honest way: a test where the revoke step failed, and my first version "helpfully" rolled production back onto the leaked password.

For MongoDB Atlas, step 1 means a new database user with the same roles (app β†’ app-lf20261001), created through a service account. For a JWT secret, it's a new random value; there's nothing to revoke, because tokens signed with the old secret simply stop verifying (users sign in again).

Testing against fake clouds

You can't run hundreds of rotations against real Atlas and Vercel accounts. So the tests run against an in-memory imitation of each API leakfix calls: Atlas database users, Vercel env vars and deployments, Render services, and a health endpoint. Each fake can be told to fail: the next deploy errors, the health check returns 503, Atlas blocks the caller's IP.

That makes failure paths ordinary test cases:

it("rolls back and redeploys the old config when the health check fails", async () => {
  const { state, fetchImpl } = setup();
  state.healthy = false;
  const result = await execute(buildPlans(leaked(), cfg, creds, fetchImpl).plans[0]!);

  expect(result.rolledBack).toBe(true);
  expect(state.env.get("MONGODB_URI")!.value).toBe(LEAKED_URI);     // old value back
  expect(state.atlasUsers.has("collabify")).toBe(true);              // old user still works
});

There are 22 of these now.

Then reality

Green tests didn't mean it worked. The first real run against my production app failed twice:

  1. Atlas returned 403. My ISP had moved me from one IP to the next, and the service account only allowed the old one. leakfix stopped at step 1 and changed nothing. (Now leakfix init suggests allowing the /24 range.)
  2. Vercel's redeploy ended with git_info_fail. I had redeployed with only the old deployment's id; for a Git-connected project, Vercel needs the Git source (repo id and commit). leakfix had already switched three variables and created a new Atlas user. It put all three values back, deleted the new user, and production never noticed.

After fixing both, the third run rotated everything with zero downtime. My fake Vercel now fails with git_info_fail unless the request names the commit, so that bug can't come back.

I later added Render and verified it on a live service, including a rollback: the app's /health route prints a short hash of the current secret, so you can watch the new value arrive, and then watch the old one come back when the rotation is rolled back.

Try it

npx leakfix scan        # which secrets are committed?
npx leakfix init        # finds your Vercel/Render project and Atlas project, checks access
npx leakfix rotate      # shows the plan; add --yes to run it
npx leakfix fix-repo    # stops tracking .env, adds .gitignore entries and .env.example

It also runs as a GitHub Action that fails a pull request containing a secret:

- uses: safayet404/[email protected]

What it isn't: a replacement for GitHub secret scanning or GitGuardian (they find leaks better), or for Vault, Doppler or Infisical (if your secrets already live there, use their rotation). leakfix is for the common case in between: secrets in .env and platform environment variables, and an alert you need to act on now.

Today it covers MongoDB Atlas, JWT secrets, Vercel and Render. Railway, Stripe keys and a GitHub App are next. Adding a platform is one file, and the contributing guide walks through it.

Code: https://github.com/safayet404/leakfix

What does your stack look like? That decides what I build next.

πŸ“° Read the original article on Dev.to Security

Originally published by Dev.to Security. Aggregated on AIWithGhost for educational purposes β€” full credit and traffic to the original publisher.