Dev.to WebDev 🛠 Dev 👁 0 📖 6 min read

Checking passport photos against government specs in the browser with MediaPipe

Government photo rules read like a spec sheet. The DS-160 (US visa) wants a square JPEG between 600×600 and 1200×1200 px, at most 240 KB, with the head (chin to top of hair) taking 50–69% of the image height and the eyes

Government photo rules read like a spec sheet. The DS-160 (US visa) wants a square JPEG between 600×600 and 1200×1200 px, at most 240 KB, with the head (chin to top of hair) taking 50–69% of the image height and the eyes 56–69% up from the bottom. The Chinese visa wants 354×472 px, 40–120 KB, and the crown 10–70 px below the top edge. India's OCI photo must have a light background that is not white.

A spec sheet can be turned into code, so I built a checker that runs entirely in the browser: the photo never leaves the device, and there's no server-side inference to pay for. This post walks through the pipeline and the parts that were harder than expected.

The stack

  • @mediapipe/tasks-vision with two models, served from our own origin:
    • FaceLandmarker (478 landmarks + blendshapes + a facial transformation matrix)
    • ImageSegmenter with the selfie segmentation model
  • Plain <canvas> for rotation, cropping, pixel statistics and JPEG encoding
  • No uploads and no third-party scripts
const fileset = await FilesetResolver.forVisionTasks(WASM_PATH); // self-hosted wasm
const face = await FaceLandmarker.createFromOptions(fileset, {
  baseOptions: { modelAssetPath: "/models/face_landmarker.task", delegate },
  runningMode: "IMAGE",
  numFaces: 2, // we need to know if there's more than one
  outputFaceBlendshapes: true,
  outputFacialTransformationMatrixes: true,
});

numFaces: 2 is deliberate. If you ask for one face you can't tell a clean portrait apart from a photo with someone standing behind you, and "no other people in the photo" is a rule almost everywhere.

Step 1: level the eyes before measuring anything

Phone photos are rarely level. Every measurement downstream (head height, eye line, crop box) assumes an upright head, so the first step is to rotate around the midpoint between the irises:

const roll = rollFromEyes(irisLeft, irisRight); // atan2 of the eye line, in degrees
if (Math.abs(roll) > 1) {
  source = rotate(source, roll, cx, cy);
  res = face.detect(source); // re-run landmarks on the levelled image
}

Is rotating allowed? For in-plane tilt, yes: the US State Department's own photo tool rotates. What you must never do is warp or retouch the face. After rotating, the canvas corners are blank, so the crop search is constrained to rectangles that stay inside real pixels.

Step 2: finding the top of the head (landmarks stop at the forehead)

This was the least obvious part. Every official spec measures head height from the chin to the crown (top of the hair), but the face mesh stops at the hairline. Using the forehead landmark underestimates head height by roughly 20%, and that is enough to fail a 50–69% window.

The fix is the segmentation mask. Scan down from the top of the image inside the face's column band and take the first row where several columns are "person":

export function findCrownY(mask, w, h, x0, x1, maxY, minColumns = 2) {
  for (let y = 0; y <= maxY; y++) {
    let hits = 0;
    for (let x = x0; x <= x1; x++) if (mask[y * w + x] === PERSON) hits++;
    if (hits >= minColumns) return y;
  }
}

Two details matter:

  • Only scan the middle 60% of the face width, so shoulders and raised hands at the edges don't count as "head".
  • Require a minimum run of columns (about 10% of the face width) so a few stray mask pixels don't pull the crown up.

If the mask finds nothing (a bald head against a white wall can blend in), fall back to forehead − 0.2 × (chin − forehead).

Model size matters for a page people open on their phone. I started with the 16 MB multiclass segmenter (hair, face skin, body, background) and replaced it with the ~250 KB person/background selfie model, which treats hair as "person". On my test portraits the crown positions agreed within about 1.5% of head height, which is well inside every spec's tolerance.

Step 3: solve for the crop, don't guess it

With the crown, chin, eye line and face center known, the crop is a small constraint problem. Each spec gives some of:

  • head height as a fraction of the image height (e.g. [0.50, 0.69])
  • eye line as a fraction from the bottom (e.g. [0.56, 0.69])
  • crown margin from the top edge (China: 10–70 px of 472)
  • output size and aspect ratio

Pick the scale that puts the head at the middle of its allowed range, place the eyes in the middle of theirs, center horizontally, then check that the rectangle still sits inside the image. If the person is too close to the camera, no rectangle fits, and we tell them to step back instead of silently producing a bad crop.

Some specs publish no head-size number at all (the UK digital visa photo, New Zealand passport). For those we use a clearly labelled "recommended framing" rather than inventing a rule.

Step 4: the checks

Everything after the crop is measured on the output-sized image, because that is what the government system will see.

Rule Signal
Eyes open blendshapes eyeBlinkLeft/Right < 0.5
Mouth closed jawOpen < 0.25, plus the lip gap relative to face height
Facing the camera yaw from the facial transformation matrix, warn at 5°, fail at 10°
Background mean and standard deviation of luma over non-person pixels
Lighting mean luma of the face box within 70–210
Sharpness variance of the Laplacian on the face box, resized to 256 px wide

One surprise: pitch from MediaPipe's transformation matrix reads 10–17° on perfectly frontal official sample portraits. Failing on pitch would reject good photos, so pitch only ever warns.

Background checks depend on the spec: "white" (US) needs a mean above 200 with low variance, while OCI needs a light background that is not white. The same pixels can be a pass for one document and a fail for another, which is why every check reads its thresholds from the spec object instead of hard-coding them.

Step 5: hitting the file-size window

Many portals have a minimum and maximum file size (China: 40–120 KB, Australia: 70 KB–3.5 MB) and some cap the compression ratio. canvas.toBlob(cb, "image/jpeg", q) is monotonic enough in q that a binary search works:

export async function fitJpeg(encode, { minBytes, maxBytes }, iterations = 8) {
  const best = await encode(1);
  if (best.size <= maxBytes) return best.size >= minBytes ? ok(best) : tooSmall(best);
  const worst = await encode(0.3);
  if (worst.size > maxBytes) return tooLarge(worst);
  let lo = 0.3, hi = 1, fit = worst;
  for (let i = 0; i < iterations; i++) {
    const q = (lo + hi) / 2, r = await encode(q);
    if (r.size <= maxBytes) { fit = r; lo = q; } else { hi = q; }
  }
  return fit.size >= minBytes ? ok(fit) : tooSmall(fit);
}

encode is injected so the search is unit-tested without a canvas. Eight iterations are enough to land within a few KB of the limit, and we take the largest file that fits, because a bigger file means less compression and better quality.

What I'd tell anyone building something similar

  1. Store the source for every number. Each spec in the code links to, and has an archived copy of, the government page it came from. Requirements disagree across pages (the Vietnam e-visa portal states both a 2 MB and a 10 MB limit) and you need to be able to say why you chose one.
  2. Test on official sample photos. Governments publish "acceptable" examples. If your checker fails them, your thresholds are wrong, and that is how the pitch issue showed up.
  3. Never retouch. It's tempting to "fix" backgrounds or smooth skin, but US, UK, Australian and New Zealand guidance explicitly rejects edited photos. Crop, level and resize only.

The spec data (18 documents, with sources) is open on GitHub: tensam/passport-photo-requirements. The checker itself is at pixtidy.com. The check is free; I'm the developer, so feedback on edge cases (glasses glare, very dark skin against a white wall, babies) is very welcome.

📰 Read the original article on Dev.to WebDev

Originally published by Dev.to WebDev. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.