How a crypto library got my Tauri app flagged as a trojan (and how I found it by bisection)
My desktop app started getting flagged by Microsoft as Trojan:Win32/Wacatac.B!ml. The code hadn't learned any new tricks. The culprit turned out to be the encryption library I had just added, and I found it by bisecting
My desktop app started getting flagged by Microsoft as Trojan:Win32/Wacatac.B!ml. The code hadn't learned any new tricks. The culprit turned out to be the encryption library I had just added, and I found it by bisecting builds against the VirusTotal API.
This post covers how I tracked it down, how I replaced the library without breaking the files users already had on disk, and what I'd do differently.
The app
RAM β Roblox Account Manager is an open-source (GPL-3.0) Windows app built with Rust + Tauri 2 and React. It manages many Roblox accounts, launches several game clients at once and automates things like rejoining servers.
That last part matters. An app that stores session cookies, spawns processes, brings windows to the front and sends synthetic input already looks suspicious to a heuristic engine. There is no margin for anything else to look off.
The symptom
After one release, Windows started warning users about the .exe, and VirusTotal showed 1/75: Microsoft, Trojan:Win32/Wacatac.B!ml.
The !ml suffix is the important part: it's a machine-learning verdict, not a signature match. No known malware pattern was found. A model just decided the binary resembles malware.
First guess: wrong
My first instinct was to compare the last clean build with the flagged one. I dumped both with pefile:
- same 16 "heuristically heavy" imports (
SendInput,SetForegroundWindow,TerminateProcess,DuplicateHandle,RegSetValueExWβ¦) - same section entropy, same linker
- only four new imports, all boring (
FlushFileBuffers,OpenEventW,RemoveDirectoryW,MessageBoxW)
Conclusion at the time: "the behavior profile is identical, the model just changed its mind, nothing to fix." It sounded reasonable, and it was wrong. The import table isn't everything the model looks at.
Bisecting against VirusTotal
So I stopped guessing and bisected. The VirusTotal API is free (4 requests/minute, 500/day). A small script hashes the file, looks up an existing report and uploads it only if the hash is new:
const digest = new Bun.CryptoHasher("sha256").update(bytes).digest("hex");
const existing = await fetch(`${API}/files/${digest}`, { headers });
if (existing.ok) {
// already analyzed: read last_analysis_stats
} else {
// upload, then poll /analyses/{id} until "completed"
}
Then plain git bisect: build the release .exe at each step, scan it, mark it good (0/75) or bad (1/75). Each step takes a few minutes, mostly the Rust release build.
(A tip: my first attempt bisected by date and broke on merge commits. Use real git bisect.)
The first bad commit was the one that integrated encryption for the account file, using sodiumoxide, the Rust bindings to libsodium.
Why libsodium?
sodiumoxide statically links the libsodium C library into your binary. A blob of optimized native crypto primitives, sitting inside an app that also stores credentials and injects input, is apparently close enough to what ransomware and stealers ship for the model to fire.
Nothing is wrong with libsodium itself. It's the combination the classifier learned to dislike.
Replacing it without locking anyone out
Users already had encrypted files on disk, and some had a password on them. The replacement had to read those files byte for byte. I went with pure-Rust RustCrypto crates and matched every libsodium parameter:
argon2 = "0.6"
crypto_secretbox = "0.1" # XSalsa20-Poly1305
sha2 = "0.10" # SHA-512 of the password, as before
/// Same as libsodium's crypto_pwhash (argon2i13, OPSLIMIT/MEMLIMIT_MODERATE):
/// Argon2i v0x13, t_cost = 6, m_cost = 128 MiB, parallelism 1, 32-byte output.
pub fn derive_key(password_hash: &[u8], salt: &[u8]) -> Result<Key, CryptoError> {
let params = Params::new(131072, 6, 1, Some(32)).map_err(|_| CryptoError::InvalidData)?;
let argon = Argon2::new(Algorithm::Argon2i, Version::V0x13, params);
let mut key: Key = [0u8; 32];
argon.hash_password_into(password_hash, salt, &mut key)
.map_err(|_| CryptoError::InvalidPassword)?;
Ok(key)
}
crypto_secretbox from RustCrypto already produces libsodium's layout: the 16-byte Poly1305 MAC in front of the ciphertext.
What made me trust the swap was a test. Before removing libsodium, I encrypted a tiny payload with the old implementation, pasted the bytes into the test as hex, and the new code must decrypt them exactly:
#[test]
fn decrypts_a_libsodium_fixture() {
// Bytes produced by the OLD libsodium/sodiumoxide implementation.
// If this fails, existing users' files would be locked. Never relax it.
let fixture = hex_decode(concat!(
"526f626c6f78204163636f756e74204d616e6167657220637265617465",
// β¦the rest of the header, salt, nonce, MAC and ciphertext
));
let hash = hash_password("test-password");
let plain = decrypt(&fixture, &hash).expect("fixture must decrypt");
assert_eq!(plain, b"{\"account\":\"example\"}");
}
If any parameter is off by one β Argon2 version, memory cost, MAC position β this test fails. That's exactly the kind of bug that locks users out of their data.
Side effect: tests got 17x slower
Pure-Rust Argon2 in a debug build is painfully slow: 5.4 s per derivation, against 0.3 s optimized. libsodium never had this problem, because it's precompiled C.
Since cargo test runs in debug, the suite went from minutes to most of the CI time. The fix is one block in Cargo.toml: optimize dependencies even in dev builds, keep your own crate unoptimized so it still compiles fast.
[profile.dev.package."*"]
opt-level = 3
The Rust suite went from 230 s to 41 s, faster than it was with libsodium.
Results
| File | Before | After |
|---|---|---|
App .exe
|
1/75 (Wacatac.B!ml) |
0/75, Defender clean |
| MSI installer | 0/61, but required admin | 0/75, per-user, no admin |
| NSIS setup | 3/71 | 1/75 (one generic ML engine, on the packager) |
No code signing, and no removing the encryption.
What I'd tell anyone shipping an unsigned desktop app
-
!mlverdicts are a moving target. The same source built with different feature flags got different verdicts from the same model. You can't "win" once; you measure every release. -
VirusTotal's Microsoft engine and local Defender disagree. A build passed local
MpCmdRunand still got flagged on VirusTotal. Check both, because users see both. -
Bisect, don't theorize. My careful import-table analysis pointed at the wrong conclusion. A few rounds of
git bisectpointed at the right commit. -
Make your scanner fail loudly. My own scan script once reported "all clean" while the Defender step had crashed (a UTF-8 em dash in a BOM-less
.ps1breaks Windows PowerShell 5.1) and the VirusTotal step exited 0 on any verdict. A scanner that didn't run must never count as clean. That rule now has a test. - Native C dependencies are a heuristic risk. If a pure-Rust equivalent exists and you can prove format compatibility with a fixture, it's worth considering.
The durable fix is still code signing (there are free options for open-source projects, like SignPath Foundation). Until then, measuring every release and keeping the recommended download clean is what I can do.
The project is open source, and I'd love feedback on the Rust/Tauri side: github.com/luanmacea/roblox-account-manager
I wrote this post with the help of an AI assistant; the investigation, numbers and code are from the project's history.
Originally published by Dev.to Security. Aggregated on AIWithGhost for educational purposes β full credit and traffic to the original publisher.