Remove Watermark Without Uploading: How We Built an In-Browser AI Editor
Most online watermark removers ask you to hand over your photo first and ask questions later. We built ClearPix the other way around: the AI models come to your device, your files never leave it, and inference runs entir
Most online watermark removers ask you to hand over your photo first and ask questions later. We built ClearPix the other way around: the AI models come to your device, your files never leave it, and inference runs entirely in your browser tab. This post is the architecture overview of how that actually works in production β the ONNX Runtime Web + WebGPU inference stack, the model distribution pipeline (R2 CDN, multi-mirror fallback, SHA-256 self-healing cache), the places it broke on us, and how we keep the privacy claim structural instead of rhetorical. No marketing math; every number below comes from our own engineering logs.
Does a watermark remover upload your photos? Usually, yes
The upload question is not paranoia β it is the default architecture of the category. Nearly every "free online watermark remover" is a thin frontend over a server-side inference API. Your image goes up, gets processed on someone else's GPU, and comes back down. Whether the privacy policy says the file is deleted afterward, you have no way to verify it, and neither do we. The structural fact is that your media crossed the wire.
We kept running into this as a genuine user anxiety, and we shared it. Photos with watermarks are often personal β screenshots, ID documents, family pictures, client work. "Trust us, we delete it" is a weak foundation for a tool people use on sensitive files. So we asked the harder question: can the inference just happen on the user's machine?
The answer, in 2026, is yes β if you are willing to solve the distribution and reliability problems that server-side tools offload to a data center. That trade is the whole story of ClearPix.
The architecture in one paragraph
ClearPix is a static site. There is no processing backend. When you open a tool page, your browser downloads ONNX model weights (once, then cached forever), loads them into ONNX Runtime Web, and runs inference on WebGPU where available with a WASM fallback. Your image or video is decoded locally, processed locally, re-encoded locally, and never serialized into a network request. The only bytes that travel are the models coming down to you, and a small set of aggregate analytics events that contractually cannot contain your media.
That single paragraph hides about 90% of the engineering. Let's unpack it.
In-browser AI image processing: the runtime layer
We run all neural inference through ONNX Runtime Web (ort 1.30.x), loaded via dynamic import so its ~1 MB parsed weight stays out of the initial page bundle β Core Web Vitals matter even for tools. Session creation prefers the WebGPU execution provider and falls back to WASM:
export async function createSession(
modelBytes: Uint8Array,
providers: readonly string[] = ['webgpu', 'wasm'],
) {
const ort = await getOrt();
// ... try each execution provider in order, return first that works
}
WebGPU is worth real money here: when we enabled the WebGPU EP (ort 1.30+), our inpainting step went from 2100β3500 ms to roughly 400 ms per region in our benchmarks. But "session created" does not mean "inference works" β we have seen WebGPU sessions deadlock silently. So the session runs behind a watchdog, and a dead WebGPU session falls back to WASM rather than hanging the tab. The fallback chain is not an afterthought; every production stall we have ever shipped was closed by one.
Even the runtime binaries need their own distribution strategy. The jsep.wasm file (28 MB) is a hard dependency of the WebGPU EP β if it fails to load, ORT silently drops back to CPU inference and the user just experiences a mysteriously slow tool. We serve the ORT runtime from jsDelivr first (measured ~7Γ the throughput of our own R2 custom domain for these files) with a 5-second HEAD probe, falling back to our self-hosted R2 copy when the CDN is unreachable β which happens in some network environments.
A client-side watermark remover lives or dies by its model pipeline
Here is the uncomfortable truth of client-side AI: you are now a software distribution company. A server-side tool updates weights by deploying a container. We have to push tens of megabytes of model files to arbitrary browsers on arbitrary networks, keep them intact, and never strand a user.
Distribution. Models live in a Cloudflare R2 bucket behind a custom domain on our zone. This was a measured decision, not a default: direct r2.dev URLs gave us 75β110 KB/s, while the same-zone custom domain delivered 2.3 MB/s β a 20β30Γ difference that turns a model download from a coffee break into a pause. Models are also deliberately excluded from our deploy bundle; stripping them (plus internal bench assets) took our release zip from 27.4 MB down to 10 MB.
Permanent caching. Once downloaded, model bytes go into IndexedDB and are reused forever β repeat visits load from the device, and the tool works offline. Two details cost us real debugging time. First, big models are written in 32 MB chunks (above a 64 MB threshold), because a single structured clone of a 198 MB model pushed the renderer process into OOM. Second, the cache is self-healing: we verify the SHA-256 of a cached model on every hit (~100 ms for a 67 MB file), because a download interrupted mid-write would otherwise poison a cache that never expires β permanently bricking that user's tool with no visible error.
Mirrors and fallback specs. Each model spec carries an ordered mirror list, a pinned SHA-256, and an optional fallback spec with a different weight format (fp16 β fp32) and an independent cache key. The first mirror whose bytes hash-match wins; a stalled connection (no bytes for 30 seconds) is aborted and handed to the next mirror; if every mirror fails, the fallback spec takes over:
for (const url of spec.urls) {
try {
const bytes = await fetchWithProgress(url, onProgress);
if (hash(bytes) !== spec.sha256) throw new Error('integrity check failed');
await idbSet(spec.key, bytes); // best-effort
return bytes;
} catch (e) { lastErr = e; } // next mirror
}
if (spec.fallback) return loadModel(spec.fallback, onProgress);
We learned this the hard way: on 2026-09-26 a mirror-chain failure took model loading down in production. Since then the rule is R2 first, public mirrors as backup, fallback spec chain underneath β designed before launch, not patched after an incident.
Making the privacy claim structural
"Private photo editor, no upload" is easy to print on a landing page. We wanted it to be checkable from inside the app.
The Privacy Shield. Every tool page renders a live status bar: a "0 bytes uploaded" chip, per-model download state with an offline-ready badge, and β once an inference session exists β the active backend translated into plain language. It is driven by a status bus inside the model loader, so it reflects what the engine is actually doing, not what the copywriter hoped. You can open devtools and verify the claim yourself: the network tab shows model bytes coming in and nothing going out.
An event contract we publish. We do collect a small set of aggregate events to know whether the tools work β things like process_complete (duration, backend, device tier) and result_feedback (thumbs up/down with reason tags). The contract, stated in our privacy policy, is that events never include images, file names, or file contents. Including them would defeat the point of the site.
Analytics that can't leak media. We use Microsoft Clarity for session replays to fix confusing UI. Because your files never reach a server this is already low-risk, but we additionally mark every canvas, image, and video element that can display user media with Clarity's data-clarity-mask attribute, so replays are structurally incapable of containing your content even if a selector goes wrong.
What this architecture costs us
Honesty section, because dev.to deserves one:
- First-visit cost is real. Downloading models and compiling WebGPU shaders makes the first run slower than a server round-trip. We hide shader compilation and encoder init behind prefetch and warmup, and IndexedDB makes visit two effectively free β but visit one pays the toll.
- We gave up control of the hardware. A data center GPU is a known quantity; a user's browser is not. WebGPU flakiness, missing cross-origin isolation (which caps WASM threading), and mobile memory limits are all our problem now. The fallback chains exist because each of these has bitten us.
- Model choice is constrained by the browser. We pick architectures designed for the deployment environment over paper SOTA β swapping our upscaler to a compact architecture was a 10Γ per-frame win at equal-or-better quality, something no amount of tuning the heavyweight model would have produced. That is a full post on its own.
We think the trade is overwhelmingly worth it, but it is a trade.
Try it
Everything described above is live and free, no account:
- Remove a watermark from a photo β detection, inpainting, and export all in-tab
- Our free video watermark remover β streaming pipeline, up to two-minute clips
- Remove hardcoded subtitles from a video
- Upscale an image 4Γ β the compact SR model mentioned above (video upscaling lives on the same stack)
What's next in this series
This overview skipped the interesting fights. The rest of the series goes deep on each one:
- The video pipeline rewrite β how replacing ffmpeg.wasm with WebCodecs + MediaBunny took full video processing from 95 s to 9 s, and how a streaming single-pass pipeline got us to O(1) memory.
- Model selection and quantization in the browser β why fp16 was free bandwidth for our upscaler but destroyed our inpainting model (mask-gated architectures and precision sensitivity), and the decision framework we now apply before touching a weight.
- Video inpainting performance β motion-adaptive frame skipping lets 65% of frames avoid inference entirely, plus the temporal EMA that keeps the result flicker-free (and the UX change that accidentally switched the whole optimization off).
- Running a 198MB model in a browser tab β tiles, memory walls, the watchdog/fallback structure, and the optimizations we benchmarked and killed.
- Automatic watermark detection β a YOLO + OCR dual-engine detector and the glyph-level masks that raised precision without losing recall.
Part 1 of the ClearPix engineering series β how we build free, private, in-browser media tools at clearpix.org.
Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes β full credit and traffic to the original publisher.