My Stripe idempotency key was correct. It just stopped working after 24 hours.
I added a reward feature that pays out through Stripe's customer balance, and I did what every payments guide tells you to do: I passed an idempotency key on the write that moves money. A security review pass then found
I added a reward feature that pays out through Stripe's customer balance, and I did what every payments guide tells you to do: I passed an idempotency key on the write that moves money. A security review pass then found a double-payment path anyway, and the key was not wrong. The key was correct, unique, and deterministic. It simply had an expiry date that my retry design ignored.
This post is about the gap between "I used an idempotency key" and "this operation is idempotent." They are not the same claim, and the distance between them is a time window.
The stack is Next.js, Drizzle, Neon and the Stripe Node SDK, but the lesson has nothing to do with any of those.
The feature
When a referred user's paid subscription goes through, the referrer gets a credit on their Stripe customer balance. In Stripe terms that is a negative customer balance transaction, which Stripe automatically applies to the customer's next invoice.
The reward logic lives in one function, rewardReferralIfEligible, and it can be called from three places: a webhook for customer.subscription.updated, a webhook for invoice.paid, and a manual sweep script I run to catch anything the webhooks missed. Because the same function can be reached from several triggers, it has to be safe to call repeatedly. That requirement is where the story starts.
There is a referrals table with a status column: pending or rewarded. The function's job is to move a row from pending to rewarded exactly once and, as a side effect of that single transition, create exactly one balance transaction on the referrer's Stripe customer.
Layer one: make the database the gate
The first thing I wrote was a compare-and-swap. Before touching Stripe, claim the row:
const updated = await db
.update(referrals)
.set({ status: "rewarded", rewardedAt: new Date() })
.where(and(eq(referrals.id, referral.id), eq(referrals.status, "pending")))
.returning();
if (updated.length === 0) return;
The WHERE status = 'pending' clause is the important part. If two webhook deliveries race, both run this UPDATE, but only one of them gets a row back from .returning(). The loser sees an empty array and exits. Postgres serializes the two updates on the same row, so exactly one caller proceeds. This also means the "already processed" check and the "claim" are one atomic statement instead of a read followed by a write, which is the classic time-of-check-to-time-of-use hole.
So far so good. The winner now owns the row and goes on to talk to Stripe.
Layer two: the idempotency key
The winner then creates the balance transaction. This is the call that actually moves money, so this is where I put the key:
const tx = await stripe.customers.createBalanceTransaction(
referrerCustomerId,
{
amount: -REFERRAL_REWARD_CENTS,
currency: "usd",
description: `Referral reward — referred user converted to Pro Annual ${marker}`,
},
{ idempotencyKey: `referral-reward-${referral.id}` }
);
The key is derived from referral.id, which is a UUID assigned once when the referral row is created and never changes. That is the right shape for a key: the same logical operation always produces the same key, and different operations never collide. If the HTTP response is lost and the call is retried, Stripe recognizes the key and returns the original result instead of creating a second transaction.
I stopped here at first and considered the problem solved. Database CAS stops concurrent callers. The idempotency key stops duplicate writes to Stripe. Two layers.
The rollback that made it dangerous
There is a third piece of code, and it is the one I wrote for a good reason. The CAS flips the row to rewarded before the Stripe call. That ordering is deliberate: it is what makes concurrent callers safe. But it means that if the Stripe call then fails, the database says "rewarded" while no credit exists. That is a silent loss for the referrer, and the row is now permanently skipped because it is no longer pending.
So the error path reverts the row (trimmed to the relevant lines):
} catch (err) {
try {
await db
.update(referrals)
.set({ status: "pending", rewardedAt: null })
.where(eq(referrals.id, referral.id));
} catch (rollbackErr) {
console.error(
`[referral] CRITICAL: rollback to pending also failed for referral ${referral.id}. ` +
`Row is stuck as "rewarded" without an actual credit. Manual reconciliation required.`,
{ originalError: err, rollbackError: rollbackErr }
);
throw err;
}
throw err;
}
Rolling back to pending means the next trigger, whether it is a later webhook or my sweep script, will pick the row up again and retry. That is what I want for a clean failure like a network error before Stripe received anything.
Now put the two mechanisms next to each other and ask what happens in the ambiguous case.
An ambiguous failure is when the request reaches Stripe, Stripe creates the balance transaction, and the response never makes it back to me. A timeout, a dropped connection, a function that gets killed mid-flight. From my code's point of view the call threw. From Stripe's point of view the credit exists.
My catch block cannot tell these two cases apart. It sees an exception and rolls the row back to pending. Now the database says "not rewarded" and Stripe says "rewarded", and the row is waiting for a retry.
If the retry happens within 24 hours, the idempotency key saves me: Stripe sees the same key, returns the original transaction, and nothing is duplicated.
If the retry happens later, it does not.
The 24-hour window
Stripe's documentation says that keys can be removed from the system automatically once they are at least 24 hours old, and that Stripe generates a new request if a key is reused after the original was pruned. That is the whole mechanism. After the window, the same key means "new request," and a new request creates a new balance transaction.
I had read that sentence before. I had even written a comment in the code saying the key protects retries "within 24h." What I had not done is ask who controls how long a row can stay in pending after a rollback.
The answer was: not me, at least not tightly. A row reverted to pending waits for the next trigger. The webhook triggers are driven by Stripe's event delivery, which I do not control. The sweep script runs when I run it. If nothing touches that row for a day and a bit, and then the next trigger arrives, the function runs the whole sequence again with an expired key. The sequence passes every check I had written: the row is pending, the subscription is eligible, the CAS succeeds, and Stripe happily creates a second credit.
The reviewer's finding, stated plainly: a rollback-and-retry design plus a 24-hour key means the retry safety net has a hole that opens exactly when the system has been quiet the longest. And "quiet for a day" is not an exotic condition. It is what a weekend looks like.
This is the part worth sitting with. Every individual piece of the code was defensible. The CAS was correct. The key was correct. The rollback was correct. The bug lived in the interaction between the rollback's retry delay, which is unbounded, and the key's lifetime, which is bounded.
The fix: check Stripe's state, not your own memory
An idempotency key is a cache of "I already did this" held by someone else, with a TTL. When the question is "did this money already move," the authoritative answer is the ledger, not the cache. So before creating the balance transaction, I now look at what is actually on the referrer's customer balance:
const marker = `(referral ${referral.id})`;
const allTxs = await stripe.customers
.listBalanceTransactions(referrerCustomerId, { limit: 100 })
.autoPagingToArray({ limit: 1000 });
if (allTxs.length >= 1000) {
throw new Error(
`referrer ${referrerCustomerId} has 1000+ balance transactions; manual reconciliation required`
);
}
const existing = allTxs.find((tx) => (tx.description ?? "").includes(marker));
if (existing) {
console.warn(
`[referral] credit already present on Stripe for referral ${referral.id}; marked rewarded without a new credit`
);
await recordBalanceTransactionId(referral.id, existing.id);
return;
}
The marker is a string embedded in the description of every credit I create, containing the referral's own ID. If a credit with that marker already exists on the customer, the function does not create another one. It records the existing transaction's ID against the row and finishes. The state of the world is read from Stripe at the moment of the decision, so it does not matter how long the row sat in pending.
I kept the idempotency key. It is still the cheapest protection against the common case, a retry seconds after a flaky response, and it costs nothing. But it is now the first line of defense, not the only one. The ledger lookup is what is left when the key has expired.
Two things I want to be honest about, because this fix is not free of tradeoffs:
- Matching on a description string is a convention, not a constraint. It works because I control every writer of these credits. If someone adds a credit by hand in the dashboard with a different description, it is invisible to the check, and if someone edits a description, the check can miss it. I treat it as a second line of defense, not a proof.
- The scan is bounded. Listing every balance transaction for a customer is fine for a customer with a handful of them and not fine for one with thousands. I cap the scan at 1,000 entries, and if the cap is hit I throw instead of guessing, which sends the row to manual reconciliation. Failing closed seemed better than silently assuming "not found" on a list I did not finish reading.
A stronger anchor: record the transaction ID
The description marker is a search. A search can be wrong in ways a direct lookup cannot, so I added a column to the referrals table that stores the balance transaction ID returned by Stripe:
async function recordBalanceTransactionId(referralId: string, txId: string): Promise<void> {
try {
await db
.update(referrals)
.set({ stripeBalanceTransactionId: txId })
.where(eq(referrals.id, referralId));
} catch (err) {
console.error(
`[referral] CRITICAL: credit ${txId} granted but failed to record on referral ${referralId}. ` +
`Row will show as rewardedWithoutCredit in KPI; reconcile manually.`,
err
);
}
}
and the top of the function checks it before doing anything else:
if (referral.stripeBalanceTransactionId) {
await db
.update(referrals)
.set({ status: "rewarded", rewardedAt: referral.rewardedAt ?? new Date() })
.where(and(eq(referrals.id, referral.id), eq(referrals.status, "pending")));
return;
}
If the transaction ID is already recorded, the credit exists by definition, and the function just repairs the status without calling Stripe at all.
Notice the comment I left in recordBalanceTransactionId: it deliberately does not rethrow. The credit has already been granted at that point. If recording the ID failed and I threw, the outer catch would roll the row back to pending, and I would have built a brand-new route to the exact double-credit I had just closed. The error handling for "the money moved but my bookkeeping failed" has to be different from the error handling for "the money did not move." Mixing them is how a safety mechanism turns into a hazard.
The related trap I found one pass later
While the same function was under review, a second finding turned up that has the same flavor: an assumption about what a Stripe field means.
The reward was originally gated on the subscription's status being active. That sounds like "the customer paid." It is not. When a trial ends, Stripe moves the subscription to active before the payment for the first invoice has been collected. For a short window the subscription is active and nothing has been paid, which means a status check alone can hand out a reward for money that never arrived.
The gate now looks at the subscription's latest invoice:
const sub = await stripe.subscriptions.retrieve(stripeSubscriptionId, {
expand: ["latest_invoice"],
});
if (sub.status !== "active") return;
// ...
if (!latestInvoice || latestInvoice.status !== "paid" || (latestInvoice.amount_paid ?? 0) <= 0) {
return;
}
An invoice that is paid with amount_paid > 0 is evidence that money moved. A subscription status is a state machine label. I also added invoice.paid as a trigger for the function so that the payment itself, not a status change, wakes it up. Because the function re-derives eligibility from Stripe on every call and returns quietly when the answer is "not yet," it is safe to trigger from any event, and a trigger that arrives too early costs nothing.
What I would tell myself before writing this
An idempotency key is a time-limited promise from someone else. Before relying on one, write down its TTL next to the longest delay your own system can introduce between an attempt and its retry. If the second number can exceed the first, the key is not your idempotency story. Mine could, because a rolled-back row waits for an external trigger I do not control.
Rollback plus retry is where idempotency actually gets tested. Straight-line code with a key looks airtight. The risk concentrates wherever you convert an ambiguous failure into a "try again later." An ambiguous failure means you do not know whether the side effect happened, and "later" is unbounded unless you bound it.
Ask the ledger, not the cache. When the side effect lives in someone else's system, the safest idempotency check reads that system's current state at the moment of the decision. It costs an extra API call. It also stays correct no matter how long the row sat, how many times it was retried, or whether a key was pruned in between.
Order your layers by what they protect against. In this function the database CAS stops concurrent callers, the idempotency key stops fast duplicate retries, the ledger lookup stops slow ones, and the stored transaction ID makes the repair path cheap. Each one covers a failure the others do not. Removing any single layer reintroduces a specific bug, so it helps to be able to say which one.
Treat "the money moved but my bookkeeping failed" as its own case. It must not share an error path with "the money did not move." Rolling back after a successful side effect is how you manufacture the duplicate you were trying to prevent.
Status fields are not receipts. If a decision depends on "was it paid," find the object that records payment, an invoice or a charge, and check that. A status can be reached by paths you did not anticipate.
The uncomfortable part
I want to be clear about how this was found: not by a test and not by production. A review pass that was specifically told to attack the money path read the rollback code and the key's lifetime side by side and asked how long a row could stay pending. I had the 24-hour fact in a code comment, one screen away from the bug, and still did not connect the two.
My tests would not have caught it either, because reproducing it needs an ambiguous failure followed by a gap of more than a day, and nobody writes a test with a 25-hour sleep. If a bug only exists across a time window longer than your test suite's patience, the fix has to come from reasoning about the window, not from waiting for a failure to show it to you.
The fix is a handful of extra lines. The expensive part was learning to read idempotencyKey as a claim with an expiry date instead of as a property of the operation.
Originally published by Dev.to WebDev. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.