Dev.to WebDev 🛠 Dev 👁 0 📖 4 min read

10,000+ animated photos later: what our 68-point face pipeline gets wrong

We crossed 10,000 animated photos at Živá Fotka last month. Živá Fotka is an AI tool that turns a static photo into a short living video, and can also edit and colorize old or black-and-white photos so the result looks n

We crossed 10,000 animated photos at Živá Fotka last month. Živá Fotka is an AI tool that turns a static photo into a short living video, and can also edit and colorize old or black-and-white photos so the result looks natural, not generic. At Inithouse, a studio shipping a growing portfolio of products in parallel, we decided to look at the data and figure out where the pipeline actually breaks.

The short version: our 68-point facial landmark model handles front-facing portraits well. It handles everything else with varying degrees of gracelessness. Here is what 10,000+ runs taught us.

Processing time is not a bell curve

Average processing time sits at roughly 18 seconds per photo. But the distribution is heavily skewed. Most photos finish in 12 to 16 seconds. The tail stretches past 40 seconds, almost always because of one of three things: high-resolution scans above 4000px on the long edge, group photos with three or more detected faces, or severely damaged inputs where the model retries landmark detection multiple times.

Input category Median time (s) 90th percentile (s) Share of total
Standard portrait (front-facing, 1 face) 14 19 58%
Partial profile (15-45 degree turn) 16 24 14%
Scanned photo (detected by noise pattern) 21 38 12%
Group photo (2+ faces) 23 41 9%
Low resolution (under 500px face height) 15 22 7%

The remaining roughly 4% are edge cases: heavy occlusion, non-human subjects people uploaded to see what happens, and photos where landmark detection fails entirely and the pipeline falls back to a simpler motion model.

Where the 68-point model fails

The model places 68 landmarks on each detected face: jawline, eyebrows, nose bridge, lip contour, eye corners. When those points land accurately, the animation looks like a person shifting slightly, blinking, maybe turning their head a fraction. When they land wrong, the result looks like the face is melting.

Three categories account for most quality complaints:

Profile and three-quarter views. The 68-point model was trained predominantly on frontal and near-frontal faces. At around 45 degrees of head turn, landmarks on the far side of the face start collapsing into each other. The eye and eyebrow points on the occluded side bunch up, and the animation produces an asymmetric stretch. We see this in roughly 1 in 8 uploads.

Glasses, especially reflective lenses. The model sometimes maps eye-corner landmarks onto the frame edge instead of the actual eye. Thick frames make this worse. Reflective or tinted lenses can cause the model to miss the eye region entirely, which leads to a stiff, mask-like animation around the upper face. This affects around 6% of all inputs.

Scanned and damaged photos. Roughly 12% of inputs come from scanned physical prints. These bring their own problems: scanner-bed shadows along edges, visible paper grain, fold creases running across faces, and color casts from aging. The landmark detector handles mild damage well enough, but a crease running directly across the nose or mouth throws off the contour points. Our pre-processing step that attempts to detect and interpolate across creases catches about 70% of these cases. The remaining 30% produce visible artifacts in the animation.

What we changed

After looking at these numbers, we made three adjustments.

First, we added a confidence threshold for landmark placement. If the model reports low confidence on more than 12 of the 68 points, the pipeline switches to a region-based animation that moves the face as a whole rather than point-by-point. The result is subtler but avoids the melting-face problem.

Second, for scanned inputs we introduced a dedicated pre-processing branch. It runs noise-pattern detection to identify scanner artifacts, applies targeted denoising that preserves facial features, and attempts crease interpolation before landmarks are placed. Processing time for scans went up by roughly 4 seconds on average, but the share of scans producing visible artifacts dropped from about 30% to under 15%.

Third, we built a resolution normalization step. Photos above 4000px on the long edge get downscaled before processing, which cut 90th-percentile times for high-res inputs from 38 seconds to around 24 seconds without measurable quality loss in the output video.

The colorization pipeline adds its own failure modes

Živá Fotka also handles colorization of black-and-white photos. About 18% of scanned inputs are B&W, and many users run colorization before animation. The colorization model generally produces plausible results for skin tones and common clothing colors. It struggles with two things: unusual fabric patterns (military uniforms from specific eras, regional folk costumes) and backgrounds where the model has to guess whether something is green foliage or brown earth. These are not face-pipeline failures, but users experience them as part of the same product, so they show up in the same feedback channel.

Numbers we are watching next

We track a handful of metrics week over week. The ones we are paying closest attention to right now:

  • Share of outputs where the confidence threshold triggers the fallback animation (currently around 11%, target is under 8%)
  • Scan artifact rate after pre-processing (currently around 15%, down from 30%)
  • 90th-percentile processing time across all categories (currently 28 seconds, target is 22)

If you want to try it, the product is live at alivephoto.online (English) and zivafotka.cz (Czech). We are part of Inithouse, a studio building and running a growing portfolio of products in parallel.

📰 Read the original article on Dev.to WebDev

Originally published by Dev.to WebDev. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.