CPU probes for AI tattoo generators — what we measured
By Lena Hart tattooprobes resamples a design to a print size on CPU, measures lines and gaps in millimetres, and can blur the raster before it scores the skeleton. A PNG from an AI tattoo generator goes in as an array.
By Lena Hart
tattooprobes resamples a design to a print size on CPU, measures lines and gaps in millimetres, and can blur the raster before it scores the skeleton. A PNG from an AI tattoo generator goes in as an array.
Why the probes stayed on CPU
probes.py and simulate.py import NumPy, OpenCV, and scikit-image. The locked environment is still a CPU torch build, because reward scoring and the CLIP labeler share the venv. requirements.lock.txt pins torch==2.14.1+cpu, open_clip_torch==3.3.0, transformers==4.46.3, and image-reward==1.5. The geometry pass ran without a GPU, without a paid scoring API, and without new annotators. Seed 0 is the simulator default, torch.manual_seed(0) in the classifier, and SEED = 0 in the fusion script.
Two functions
Excerpt from tattooprobes/probes.py:
NEEDLE_MM = 0.30 # approx. single-needle/3RL line (practitioner guidance; assumption, see paper)
GAP_MM = 0.50 # gaps narrower than this are assumed to close after healing (assumption)
WORK_PX_PER_MM = 10.0 # analysis resolution after physical rescaling
def to_gray(img):
a = np.asarray(img.convert('RGB') if hasattr(img, 'convert') else img)
return cv2.cvtColor(a, cv2.COLOR_RGB2GRAY) if a.ndim == 3 else a
def physical_resample(gray, size_cm):
"""Rescale so that the longest side spans size_cm at WORK_PX_PER_MM."""
h, w = gray.shape
target = int(round(size_cm * 10 * WORK_PX_PER_MM))
s = target / max(h, w)
interp = cv2.INTER_AREA if s < 1 else cv2.INTER_CUBIC
return cv2.resize(gray, (max(1, round(w * s)), max(1, round(h * s))), interpolation=interp)
size_cm * 10 converts centimetres to millimetres, then WORK_PX_PER_MM scales the long side. INTER_AREA handles downscales and INTER_CUBIC handles upscales. to_gray accepts a PIL image or an array.
ink_mask is Otsu, or 128 when the gray standard deviation is at most 1. Ink is gray < t. _skel_widths samples a distance transform on the skeleton, doubles the radius, and divides by 10 px/mm. line_w_p10_mm is the 10th percentile of that sample. frac_sub_needle is the share under 0.30 mm. frac_tight_gap is the same width share on background inside the ink box, counted where the width is under 0.50 mm. midtone_frac is the share of gray values strictly between 60 and 200. edge_density_per_mm2 is the Canny count at 50 and 150, divided by WORK_PX_PER_MM and by area in mm².
PROBE_SIGNS is fixed a priori. line_w_p10_mm is +1. frac_sub_needle, open_ends_per_cm, frac_tight_gap, midtone_frac, border_std, edge_density_per_mm2, specks_per_cm2, and edge_transition_mm are −1. The composite is the unweighted mean of signed robust z-scores, median and IQR over the pool. edge_transition_mm is post hoc, from after the first degradation run, and probes_preregistered omits it.
from tattooprobes.probes import compute_probes
from tattooprobes.simulate import retention
features = compute_probes(img, size_cm=5.0)
kept = retention(img, size_cm=5.0, spread_mm=0.30, seed=0, noise=0.0)
features is a float dict. kept holds skel_f1 and gap_survival. We call both at 3, 5, 10, and 15 cm. When spread_mm > 0, simulate_on_skin runs cv2.GaussianBlur(g, (0, 0), spread_mm * WORK_PX_PER_MM) and adds noise only if noise > 0. The runs below pass noise=0.0.
Excerpt from tattooprobes/simulate.py:
deg = simulate_on_skin(img, size_cm, spread_mm, seed, noise)
# v0.3 fix: threshold the simulated image with the SAME Otsu threshold as the reference, and
# compare skeletons with a tolerance expressed in mm (0.2 mm) instead of a fixed 5-px window,
# which previously made the match stricter at larger physical sizes (resize artifact).
from skimage.filters import threshold_otsu
t0 = threshold_otsu(ref) if ref.std() > 1 else 128
m1 = deg < t0
s0, s1 = skeletonize(m0), skeletonize(m1)
k = 2 * int(round(TOL_MM * WORK_PX_PER_MM)) + 1
tol = cv2.dilate(s1.astype(np.uint8), np.ones((k, k), np.uint8)) > 0
TOL_MM is 0.2. F1 is the harmonic mean of the two dilated skeleton overlaps. Gap survival is the share of reference background inside the ink box that is still background after the blur.
How the rows got onto disk
HPDv2 train, prompts matching "tattoo": 2,010 pairs, 327 prompts, 1,340 images, stream-extracted from a 31.7 GB tar in 23 minutes, about 60 MB kept on disk. ImageRewardDB train+val tattoo prompts: 370 rows. 47 images are 0-byte in the upstream zips, leaving 323 usable images, 809 within-prompt rank pairs, and 51 prompts. Pick-a-Pic v2 came from a liuhuohuo2 parquet mirror of 672 shards, metadata only: 2,608 tattoo rows, 1,687 labeled and different pairs, 185 prompts. Images for that pass were not fetched. The license is unverified. The original card is gone, and the mirrors say MIT. Drozdik/tattoo_v0 is CC-BY-NC-SA-4.0, research only: 120 designs sampled, 6 defect types at 2 levels, 1,560 images.
Labels are a keyword rule plus CLIP ViT-B/32 zero-shot (open_clip ViT-B-32, pretrained openai). HPDv2: design 334, on_skin 87, incidental 674, other 146, ambiguous 99. ImageRewardDB: design 48, on_skin 12, incidental 238, other 18, ambiguous 7. tattoo_subject is design plus on_skin, which is 463 pairs and 119 prompts on HPDv2, and 148 pairs and 9 prompts on ImageRewardDB. A spot-check of 24 random design images looked tattoo-like for ~21/24 on HPDv2 and 24/24 on ImageRewardDB.
Per-dimension rows below are the alignment after the reward re-score: 528 pairs and 125 prompts. The composite stays on 463 pairs and 119 prompts.
Agreement, holdout, and the spread ladder
Pairwise accuracy is the rate at which the scorer orders a pair as the human label did. On the HPDv2 tattoo-subject per-dimension table, n_pairs is 528 and n_prompts is 125. Holm p is 0.0065 on every row.
| Scorer | Accuracy A | A−0.5 | Cohen h | Holm p | sig |
|---|---|---|---|---|---|
| frac_tight_gap@5 | 0.335 | −0.165 | −0.34 | 0.0065 | True |
| edge_density_per_mm2@5 | 0.335 | −0.165 | −0.34 | 0.0065 | True |
| midtone_frac@5 | 0.369 | −0.131 | −0.26 | 0.0065 | True |
| line_w_p10_mm@5 | 0.428 | −0.072 | −0.14 | 0.0065 | True |
| CLIPScore | 0.682 | +0.182 | +0.37 | 0.0065 | True |
| PickScore | 0.723 | +0.223 | +0.46 | 0.0065 | True |
| ImageReward | 0.684 | +0.184 | +0.38 | 0.0065 | True |
| HPSv2.1 | 0.790 | +0.290 | +0.62 | 0.0065 | True |
probe_composite@5 is the other alignment: 0.305 [0.256, 0.353] on 463 pairs and 119 prompts.
Fusion fits LogisticRegression on winner-minus-loser features after a random side flip, with a StandardScaler and GroupKFold on prompt id. HPDv2 rows in results/fusion_lopo.csv are GroupKFold(prompt,k=5). A literal leave-one-prompt-out loop will not match that file. Preregistered probes, with edge_transition_mm excluded, score 0.708 [0.667, 0.750]. Rewards score 0.790 [0.748, 0.827]. Rewards plus probes score 0.767 [0.729, 0.803]. Probes alone beat chance in that fit only with inverted coefficient signs.
In the default simulator, noise stays at 0. At σ = 0, mean skeleton F1 is 1.000 at every tested size. At σ = 0.30 mm, mean skeleton F1 by print size is 3 cm 0.477, 5 cm 0.563, 10 cm 0.668, 15 cm 0.733. At 5 cm and σ = 0.30 mm, Spearman ρ is −0.824 for frac_tight_gap vs gap survival, −0.599 for edge density vs gap survival, and +0.507 for line_w_p10_mm vs skeleton F1.
What we would change in the next eval
Fusion at 0.708 [0.667, 0.750] tracks the human bit, and the probes clear chance in that fit only with inverted coefficient signs. The signed composite on 463 pairs is 0.305 [0.256, 0.353]. Log PROBE_SIGNS for the millimetre direction. Fit the human bit for this preference label. A training loss built from the signed mean follows the 0.305 direction. line_w_p10_mm is the only preregistered key with sign +1, so a fatter Otsu mask, whether from thickening or from blur, moves that coordinate the way the heuristic favors.
HPSv2.1 was trained on HPDv2, and ImageReward was trained on ImageRewardDB, so those agreements are in-distribution leakage.
ImageRewardDB tattoo-subject has 9 prompts. The Pick-a-Pic design slice has 11 prompts. Probes are not Holm-significant on either slice, so both stay out of the 528-pair table.
frac_tight_gap and gap survival both read background inside the ink box: widths already under 0.50 mm, and pixels still background at σ = 0.30 mm. ρ = −0.824 is partly circular by construction. Edge density vs gap survival is ρ = −0.599. line_w_p10_mm vs skeleton F1 is ρ = +0.507.
The 5 px window in the excerpt made the match stricter at larger physical sizes. TOL_MM = 0.2 keeps the band in millimetres. At σ = 0 the published mean skeleton F1 is 1.000 at every tested size.
The ink-spread model is an isotropic Gaussian, not validated on healed skin. tatany.app was not ranked.
Reproduce
git clone https://github.com/Stark-Will/tattoo-probes
The archive is doi:10.5281/zenodo.23292900. Install requirements.lock.txt, keep the CPU torch wheel, and leave the seeds at 0. Entry points: compute_probes(img, size_cm=5.0) and retention(img, size_cm, spread_mm, seed=0, noise=0.0). Read per-dimension rows from the 528-pair, 125-prompt alignment, and the composite from the 463-pair, 119-prompt alignment.
Disclosure: I work on tatany.app, an AI tattoo generator; the study does not rank it or any other product.
Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.


