'\w Is ASCII-Only, and My Spam Filter Accused Four Strangers of Being a Bot Ring'
By:Ronny Cruz También disponible en español: [https://dev.to/candornetwork/w-solo-reconoce-ascii-y-mi-filtro-antispam-acuso-a-cuatro-desconocidos-de-ser-una-red-de-bots-4odl] I run a signup screening service for Fe
By:Ronny Cruz
También disponible en español: [https://dev.to/candornetwork/w-solo-reconoce-ascii-y-mi-filtro-antispam-acuso-a-cuatro-desconocidos-de-ser-una-red-de-bots-4odl]
I run a signup screening service for Fediverse instances. It scores account applications — the reason-for-joining text, the IP, the email, signup velocity — and returns pass, flag, or block. Instance admins were drowning after last month's spam wave, so the thing exists because people needed it in a hurry.
I red-teamed it against a labeled corpus. Over one day the catch rate went 37% → 60% → 74% → 76%, and at every step the legitimate applications came back clean. Thirty-two out of thirty-two, zero false positives, every run.
That number was the problem. Not because it was wrong, exactly, but because it was hiding one.
The first block
The suite that tests content signals in isolation had never produced a hard block — only flag, which routes to a human. That was intentional design: text alone shouldn't auto-reject anybody. So when a block appeared after a change to the duplicate-detection thresholds, I went looking for which case had earned it.
Four applications had been grouped as a coordinated ring. Weight 40, coordinated_duplicate_content, enough to stack into a hard rejection.
Three of them were unrelated Cyrillic-language applications. The fourth was a single period.
What was actually happening
Near-duplicate detection worked by normalizing the text and hashing it, so that "I want to join!" and "i want to join" land in the same bucket. The normalizer looked like this:
const normalized = content
.toLowerCase()
.replace(/\s+/g, ' ')
.replace(/[^\w\s]/g, '') // strip punctuation
.trim();
return crypto.createHash('sha256').update(normalized).digest('hex').slice(0, 12);
In JavaScript, \w is [A-Za-z0-9_]. ASCII only. It does not become Unicode-aware when you add the u flag — that's \p{L} and friends, and you have to ask for them by name.
So [^\w\s] doesn't mean "strip punctuation." It means strip everything that isn't an ASCII letter, digit, underscore, or whitespace.
Feed it Cyrillic and you get an empty string. Same for Chinese, Japanese, Korean, Arabic, Hebrew, Greek, Devanagari, Thai, Armenian, Georgian. Every one of them normalizes to "", and SHA-256 of the empty string is always the same value — e3b0c44298fc.... Every application written in any non-Latin script hashed to one identical bucket.
Two consequences, and the quieter one is worse.
The feature never worked. Near-duplicate detection was completely non-functional for anyone not writing in Latin script. A spam ring posting the same Russian text across forty accounts would slip through the dedup layer entirely, because all forty hashed to the same value as every legitimate Russian application on the instance. The signal was noise, so it carried no information at all. No error, no warning, no test failure. It just quietly did nothing for a large fraction of the world.
And then it did something. Once enough non-Latin applications accumulated in that one bucket, the count crossed the coordination threshold. At that point the system started reporting unrelated people as an organized ring — and by construction, that false positive could only ever land on people who don't write in Latin script. A Russian speaker, a Japanese speaker, and an Arabic speaker who had never heard of each other, grouped and rejected as coordinated spam.
My perfect false-positive record was measured against a corpus written in English.
It's not even binary. [^\w\s] strips accented Latin characters too — señor becomes seor, café becomes caf. Spanish, French, German, Polish, Portuguese, Vietnamese text all get partially mangled, which degrades match quality without erasing it. There's a gradient here running from "works fine" (English) through "works badly" (accented Latin) to "collides catastrophically" (everything else), and the gradient tracks distance from ASCII.
The fix
const normalized = content
.toLowerCase()
.replace(/[^\p{L}\p{N}\s]/gu, '') // all letters, all numbers, all scripts
.replace(/\s+/g, ' ')
.trim();
// An empty or near-empty normalization isn't evidence of anything.
// Bucketing those together manufactures coordination that isn't there.
if (normalized.length < 8) return null;
\p{L} is any Unicode letter, \p{N} any Unicode number, and /u makes the property escapes legal. Three Cyrillic texts now produce three distinct hashes. Japanese survives instead of being mangled into whatever ASCII fragments happened to be embedded in it.
The null return matters as much as the character class. A lone . normalizes to nothing, and "nothing" is not a fingerprint — grouping all the empty normalizations together is exactly how four unrelated people became a ring. When the hash is null, the dedup layer is skipped entirely. Short applications are still covered by a separate minimum-length signal; they just don't get fingerprinted.
The bug next door
While I was in there, the same function's storage had a second problem. Hash records carried a count and a set of account IDs, but no timestamp, and the cleanup routine only deleted records with a count below 2. So anything seen twice was kept for the life of the process.
Which means a common phrase — "i want to join this instance" — accumulated unrelated real humans indefinitely. Week one, three people. Week six, nine people. Eventually it crosses the coordination threshold and starts flagging legitimate applicants for the crime of writing an ordinary sentence.
My clean false-positive record was partly an artifact of testing against a freshly started process. The failure mode needed weeks of uptime to develop, and my test runs lasted seconds.
The fix was a rolling 24-hour window with per-event timestamps, capped at 200 events per hash. Coordination now means four in a day, not four since March. Bounding the counts also made it safe to lower the threshold, which caught real rings that had previously slipped under it.
Why the tests didn't catch any of this
Here's the part I find hardest to shrug off.
The test suite has three Cyrillic cases. They check that a Russian-language reason with a declared Spanish locale raises a mismatch, that the same text with a Russian locale does not, and that no locale means no assumption. Good tests. They pass.
They passed identically before and after the fix.
Cyrillic text was flowing through the engine the entire time, hashing to e3b0c44298fc, and no assertion ever looked at the hash. The tests weren't missing. They were adjacent — they covered non-Latin input for one feature while a different feature silently failed on the same input.
Non-Latin coverage in a suite doesn't mean non-Latin coverage of the thing you just changed. That's an uncomfortable lesson because it doesn't resolve into "write more tests." It resolves into "know which assertion covers which failure," which is harder.
The coda, which is arguably the real story
I fixed this in production by running a patch script directly against the server. Verified it there, moved on, felt good.
Five days later I went to write this article and grepped my own machine first, on the theory that a shared engine class tends to get copied around. Twelve copies of that engine on my laptop. Every single one of them still had [^\w\s].
Including the git repository. Including the tagged release.
The fix existed in exactly one place on Earth: an unversioned, unbacked-up file on one VPS. Any redeploy from source would have silently reverted it — reintroducing both the Unicode collision and the never-expiring hash store into production, with no test failure to announce it.
Patch-the-server workflows create orphan fixes. If you take one thing from this, take that one, because it's the one that generalizes furthest and I nearly published a triumphant bug writeup about code that was still broken everywhere but the box I typed it on.
Grep your own stuff
If you have a JavaScript pipeline that normalizes text before hashing, clustering, deduplicating, or fingerprinting, this is worth ten seconds:
grep -rn --include='*.js' --exclude-dir=node_modules -e '\[\^\\w' -e '\\W' .
[^\w\s], \W, and [^a-z0-9\s] are all the same bug wearing different clothes. The tell isn't the regex on its own — it's the regex feeding something that groups records together. A normalizer that mangles text is a cosmetic problem. A normalizer that mangles text into a shared bucket is an accusation engine.
And check whether the fix is anywhere except the machine you fixed it on.
Honest limits on all of this
The catch-rate numbers come from a hand-built corpus of 32 legitimate and 34 spam applications, which is small. Identical code produced 18, 21, and 18 catches across three runs, so treat any single figure as ±3. Four of the labeled spam cases are textually indistinguishable from real applications, so 100% content-only detection was never the ceiling.
I don't have field data on how often the coordination false positive fired against real users, because the service is young and the window it needed to develop is longer than its production life. The bug is confirmed; the blast radius in the wild is not measured. I'd rather say that than imply a number I don't have.
The fixes are verified behaviorally — a red-team harness on the host, plus simulations against real corpus text before shipping. They still have no unit coverage. Reverting that character class today would show a fully green suite, which is the same gap I just spent a section complaining about. It's on the list.
Originally published by Dev.to Security. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.