Dev.to AI πŸ€– Ai πŸ‘ 0 πŸ“– 3 min read

How I built object-tracking GIF captions on serverless GPUs (and kept the bill bounded)

I built GifGadgets, a set of GIF tools that run in the browser. Most of it is boring client-side work. Two features are not: captions that follow an object through a GIF, and background removal or replacement. Both use s

I built GifGadgets, a set of GIF tools that run in the browser. Most of it is boring client-side work. Two features are not: captions that follow an object through a GIF, and background removal or replacement. Both use segmentation models (SAM 2 and 3.1) on a serverless GPU (Modal). This post covers how they're put together and the choices that keep them cheap to run and low-maintenance. I'm a solo developer with a day job, so I wanted to set it up so that I could deploy it and just kind of forget about it (Although it's hard to do the forgetting part)

What it does

  • Follow an Object: click something in a GIF, and a caption tracks it frame by frame.
  • Remove or swap the background: click the subject and get a mask across every frame, then drop in a new background or leave it transparent.
  • Everything else (resize, crop, reverse, convert, add text) runs in the browser, so files never leave the device.

Architecture

Browser β†’ Lambda API β†’ Modal.

  1. The browser asks the API for a presigned S3 upload URL and uploads the GIF straight to private storage. The API never touches the file bytes.
  2. A small CPU "broker" on Modal receives the job and starts a separate GPU function. The broker answers status polls, so a waiting browser never occupies a GPU.
  3. The GPU job reads the GIF from S3, runs the model and writes the result back to S3.
  4. The browser polls for status, downloads the result and does the rest locally.

Two details I like:

  • Masks, not finished images. The server returns a gzip-compressed, bit-packed mask per frame. All compositing (new background, text placement, encoding) happens in a web worker in the browser. Changing the replacement background doesn't need another GPU run, which is where most of the cost saving comes from.
  • Real progress. The GPU job writes its progress to a Modal Dict keyed by call ID, so the UI can show "Starting GPU…" and then "frame 12 of 80". If the write fails, the job doesn't, and the UI falls back to a plain "running".

Models

Tracking uses SAM 2.1 on an L4. Background removal uses SAM 3.1 on an H100. Each lives on its own image with its own pinned dependencies, so upgrading one can't break the other. Model weights are baked into the container image at build time, so a cold start never waits on a multi-gigabyte download.

Keeping the bill bounded

Serverless GPU scales to zero, which is great until something goes viral. The safeguards, in order of how much they matter:

  • Prepaid credit. The hard ceiling is the credit balance on the Modal. If it runs out, the feature degrades; it doesn't keep billing.
  • Per-IP quota enforced in DynamoDB, plus a WAF rate rule.
  • A kill switch. One environment variable disables the segmentation route without a redeploy.
  • Budget alerts on the AWS side, set to notify me early.
  • Precomputed demos. The sample GIFs use stored results, so people can try the feature without spending a GPU second.

None of these is a true global cap, and I say so in my own docs. The failure mode I accepted is that the AI features might run out of credit on a busy day, so UI says so

What went wrong

At first I tried running the SAM inferences on lambda but it was extremely slow and a poor user experience (for me when I was testing and using it). Processing the tracking frames would take over a minute. To cut the time, I went to Modal serverless functions to utilize the GPU, and that improved the speed. However, on gifs with more frames it would still run pretty slow, so I sampled to 10 frames per second for the object tracking, and still got acceptable results.

Things I'd do differently

Currently it has a src folder where the frontend is built. It uses jinja to build the html files which get shoved into a folder called frontend, which gets synced to s3 and served up through cloudfront.

I feel like the src and build can be more "DRY" but for now, the thing works and I'm just trying to leave it alone

Try it

The demo is at gifgadgets.com. The background feature in the demo is limited: [what the limit is]. If it's busy, the sample GIFs still work. I'd like feedback, especially on GIFs where tracking drifts or the mask is wrong.

πŸ“° Read the original article on Dev.to AI

Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes β€” full credit and traffic to the original publisher.