Detection Boxes to Redrawn Glyphs: A Developer Pipeline for Text Inside Images
Ten ways to change the words baked into a finished image — and keep the layout, type, and background a designer already signed off. You do not translate an image by translating a string. You find boxes, read what is ins
Ten ways to change the words baked into a finished image — and keep the layout, type, and background a designer already signed off.
You do not translate an image by translating a string. You find boxes, read what is inside them, and put different glyphs back into the same pixels. A campaign poster leaves design as English. A menu leaves as Japanese. A product label leaves as German. A diagram leaves with Spanish callouts. What you hold is the export: PNG or JPEG, layers gone, type painted into the raster. The brief does not change. Ship the other language. Do not ship a different design.
A document translator can let the layout flex. An image cannot. Column width, drop shadow, and the texture under a headline are pixel counts. Erase the source word and you also erase the pixels it covered. Image localization is that constraint: meaning has to change, and the picture is not allowed to.
This piece follows the pipeline a developer would build, then the ten real ways to own each stage. Scripts, OCR, and APIs come before any phone-camera story. Fidelity, cost, and failure modes stay attached to the method that has them.
The Detection Box Is the Unit of Work
Treat every text run as a rectangle, then as a string, then as a hole you have to fill, then as ink you have to put back. If you build the loop in Python, the stages are boring on purpose. Boring is what you can test.
- Detect. Find the regions that contain glyphs. OpenCV edge or contour detection will give you boxes on high-contrast labels. A deep OCR model — EasyOCR, or Tesseract in a mode that returns word boxes — does better on photos. You want coordinates, not one blob of page text. A box that is too tight clips diacritics. A box that is too loose eats a logo or an arrowhead.
- Read. OCR the crop. Language hints matter. Spanish on a label is not the same pass as Chinese on a package. Keep the confidence score. A low-confidence box is a box you should not silently translate.
- Translate. Send the string, not the image, to a machine-translation API or to a person. Google Translate, DeepL, and Baidu Translate are the usual engines here. Lock tokens that must survive unchanged: brand names, SKUs, registered marks, and numbers whose meaning is not linguistic.
- Erase. Cut the glyphs out and reconstruct the background. On a flat fill this is a rectangle of the sampled color. On a photo, gradient, or texture you need an inpaint. A small script usually fills with a flat color or a neighborhood average. A diffusion inpainter is what you call when that average looks like a bandage.
- Redraw. Draw the translation into the box with Pillow or OpenCV. You choose the closest font you actually have a file for. You set weight, fill, and size. You measure the advance width. If the string does not fit, you shrink, wrap, or track. You do not pretend it fits.
- Write. Save a bitmap at the same resolution you loaded. Keep the original. Diff them. The diff should light up inside the old text boxes and almost nowhere else.
That is the whole product, whether a person runs it by hand or a job runs it overnight. Hosted tools hide steps 1 through 5 behind one request. Multimodal models hide them inside a prompt. Diffusion models ask you for the mask and then invent the pixels. A designer does the erase and the redraw with a stylus and a font menu. The phone in your pocket does a coarse version of the loop and stamps a reading overlay on top, which is a different product.
Atlas Cloud has described the same split: after OCR lifts the words, someone still has to place them back into the design. A pure raster edit merges the change onto a new layer and gives you less control over individual letters. Manual typesetting is slower and, when the font is right, the most precise. A script sits between those two. You get coordinates. You do not get the original text frame, the original leading, or the original OpenType features unless you reconstruct them.
Input is a bitmap. Output is a bitmap of the same size. The value every method either computes or fakes is a list of boxes with strings. If you cannot see those boxes, you cannot debug a single bad caption.
Four Constraints a Redraw Has to Survive
The layout is frozen. The words are not. Every method below is a trade against the same four limits.
Length is not preserved. A short English phrase often grows by 30% or more in German or French, and usually shrinks in Chinese or Japanese. The box does not grow with the grammar. You resize, re-wrap, or re-track. Assuming the translation fits unchanged is how buttons overflow.
Type has to be rebuilt. Family, weight, color, letter-spacing, and any shadow, stroke, or outline are part of the design. A right sentence in a default sans still reads as a patch. Without the identical font file, a full match is hard. Photoshop and Pillow share that limit: they can only render a face you can load.
The background under the glyphs is destroyed when you erase them. Solid color is easy. A photo, gradient, or texture is reconstruction. Feather the edge and check the seam at 1:1. A visible rectangle means the fill failed.
Non-text geometry is in scope. Arrows, callouts, numbering, and legends contain text or point at it. OCR returns strings. It does not return "this arrow belongs to this caption." A person, or a model that sees the whole figure, has to.
The right method depends on fidelity, how much text is in the frame, and whether the source file still exists.
Scripts Before Buttons: a Custom Pipeline and OCR Typesetting
Own the batch job before you rent a button. The phone-camera path is last on purpose.
Method 8: A computer-vision script you run yourself
How it works. This is the loop from the first section, as code. Detect regions with OpenCV edge or contour detection, or with a deep OCR model. OCR and translate each box. Cut the glyphs out and fill the background. Draw the new string. Write the file. There is no off-the-shelf GUI. The script is the product.
Tools. Python with Pillow, OpenCV, EasyOCR, and pytesseract. A Baidu or Google Translate API for the string. The same job exists in Adobe PixelSense and in C# or .NET image code.
Input and output. A normal bitmap in. A same-resolution bitmap out. Keep the box list beside the file so a bad label can be blamed on OCR, translation, or the draw.
Automation. High once the script is actually finished, and fully batchable. Thresholds, language packs, and font fallbacks are the unfinished part.
Fidelity. Medium, and it depends on those details. You can roughly honor layout by drawing into the original boxes. You usually cannot match fonts automatically — you pick the closest face by hand. The fill is commonly a flat color or a neighborhood average. That is less natural than a diffusion inpaint, and it fails on complex backgrounds and non-standard lettering.
Time and cost. Development cost is high. Runtime is on the order of seconds per image. Money can sit near zero next to a third-party service if the pieces are open source or already licensed. The invoice is the coding.
Pros and cons. Full control, batch, and a path that can stay off the public network. Design-detail fidelity trails a dedicated image model. You will tune parameters and correct misses.
Best for. An enterprise process requirement, or small-scale label replacement. The usual case is a merchant scripting OCR, translation, and write-back on product-image labels. Demanding, and flexible.
Method 2: OCR, then typeset like you still have a layout
How it works. Same front half as Method 8 — OCR for content and position — then a person or a layout app owns the back half. Translate by machine or by a translator. Re-render at the original positions and styles. The pipeline is recognition, translation, layout reconstruction, output image.
Tools. Tesseract, EasyOCR, and Adobe Acrobat OCR. Google Translate, DeepL, and Baidu Translate. Layout in Adobe InDesign or Photoshop, or in Pillow and Matplotlib when the figure is a chart.
Input and output. PNG or JPEG in, a new image out. A PDF or AI detour is worth taking: a text frame can change leading without you repainting a gradient.
Automation. Medium. OCR and translation script cleanly. Final typesetting usually does not, once source and target lengths diverge and someone must change size and line spacing without hitting a photo.
Fidelity. High when fonts and frames match. Expansion still breaks the original spacing, so you scale or re-wrap. A bad box on a busy figure hurts more than a weak synonym. Kerning and ligatures drift even in the right family. Leader lines and numbering are not handled. They stay manual.
Time and cost. OCR plus translation is seconds to minutes. Typesetting is minutes to tens of minutes. Money stays low on open-source OCR and an API you already pay for. The labor is the typesetting hour.
Pros and cons. Faster than a blank-layer redraw, and reviewable while the text is still text. You depend on OCR accuracy and font matching. One button will not publish this.
Best for. Chart annotations, technical manuals, and educational diagrams. PDF translators follow the same idea: parse structure, translate region by region, replace in place, preserve the layout. On a flat image you are running that algorithm without a text layer.
Call an API When You Want the Model to Repaint
Delegating detection or inpainting does not delegate the check on numbers and names.
Method 6: Multimodal models that take an image and an instruction
How it works. GPT-4V, a later vision-capable GPT-4-class model, ChatGPT image editing, or a DALL·E edit takes the image plus an instruction to translate the text and leave the design alone. The model is supposed to detect, understand, and generate together. What you can operate is: image and instruction in, recognition and translation, a new image, then a human proofread. Do not treat a follow-up question as a box list.
Tools. OpenAI ChatGPT with image features such as GPT-4V. The DALL·E 3 image-editing API when you call it. Google Imagen Editor, mostly research-stage rather than a production control. Inside a service, developer access is the OpenAI API.
Input and output. Image plus prompt in. A new image out, at the original resolution or a size you set. You do not get layers back. If the model regenerates the frame, inspect more than the caption.
Automation. High. One prompt can carry the job. Log the prompt, the model, and the image hash anyway. A prompt is a bad audit log.
Fidelity. Generally high on layout and style. OpenAI's examples include a coffee-machine instruction diagram translated into Spanish with the layout kept. Atlas Cloud has reported GPT-4V translating infographic text and leaving the rest of the figure intact. Models still miss strings, misspell them, and shift elements when the new length is far from the old one. Budget a fix. This is not fine typesetting. It is often good enough for marketing art, which is a different bar from a regulated label.
Time and cost. Medium. One image is often seconds to tens of seconds, plus an API or subscription cost and your review time. Quota is finite.
Pros and cons. Arrows, legends, and hierarchy can survive, because the model sees the figure rather than a bag of boxes. Fine control is limited. Expect spelling errors, leftover source language, or compositing artifacts.
Best for. Infographics, posters, and illustrated promotions. A usable prompt is concrete: translate all text into Chinese, and do not change anything else. When one line overflows, redo that region. Redoing the poster is how the background drifts.
Method 7: Diffusion inpainting, where you own the mask
How it works. Mask the text, then prompt a replacement that keeps the background. DALL·E 3 or Stable Diffusion inpainting fills the mask. The flow is detect, build a mask, prompt ("replace X with Y, keep the background unchanged"), fill, inspect. Uploading the original and a mask to an editing API is the same idea over HTTP.
Tools. The OpenAI DALL·E 3 editing API. Adobe Firefly Generative Fill. HuggingFace Diffusers, including runwayml/stable-diffusion-inpainting, when you run the weights. ControlNet on text detection, and LaMa for the inpaint, when the mask should follow glyphs instead of a hand-drawn blob.
Input and output. Original, text mask, and prompt in. An image out. The model is supposed to take language and style from the prompt. Check the glyphs. These models are not spellcheckers.
Automation. Medium-high. Generation is one call. The mask is not. A person draws it, or OpenCV or OCR emits it. Too small and old serifs remain. Too large and the model repaints a product that was fine.
Fidelity. High, a step under models built to edit text. Long strings break words. Style tracks the prompt. The strength is background repair: a fill can blend into a complex texture more naturally than a neighborhood average. The weakness is letterforms — slight blur, a bad shape, an extra detail. Tuned masks and prompts are the difference.
Time and cost. Medium. Typically ten to tens of seconds, or your local GPU. Cloud adds queue time. You pay in compute or an API fee.
Best for. Creative work where a slight style shift is acceptable: concept art, slogans in illustrations. The banner fixture is Stable Diffusion taking "Summer sale — 30% off everything" to "Sommer-Sale — 30 % auf alles," background held, lettering re-rendered. Short lines survive more often than paragraphs. Set a regulated number with a real font.
Method 3: Format-preserving document services
How it works. Document translators parse layout and fonts, machine-translate each block, and re-typeset. Steps: layout parsing (coordinates and fonts), translation by a cloud service or a local engine, re-rendering, file write. Graphic images are typically left untouched, which is right for a scanned brochure and wrong for a figure whose labels had to change.
Tools. PDFTranslator, routing through Gemini and ChatGPT engines, the Pangeanic ECO platform, Pairaphrase, and Smartcat. Many accept a PDF or an image and return a translated PDF or graphic in the original styling. Integrate them as a service unless you also have an API you are allowed to call.
Input and output. Scans, images, or PDFs — JPG, PNG, or PDF — in. A file in the same general format out, with layout, tables, and graphic positions essentially where they started. A table that reflows by a row is still a defect.
Automation. High, end to end, with a light human proofread and a heavier one where numbers live. Unattended publishing is a different claim.
Fidelity. High. Careful OCR plus a dedicated translation system or an LLM, aimed at coordinate-level mapping. Layout and fonts are preserved. Document-level fidelity in this class is described as reaching 90% or better. Remaining errors are mostly length, such as a German cell in an English-sized column. Unusual layouts and handwriting are weak. A strong financial report does not imply a strong whiteboard photo.
Time and cost. Seconds to minutes for most jobs. A 60-page PDF is a matter of minutes. Paid service, or purchased API quota. Bulk and long documents are where the bill shows up.
Pros and cons. Automation and fidelity, against fees. Weird layouts and handwriting are poorly supported. A flat PDF back means the next fix looks like Method 2.
Best for. Enterprise technical documents, product manuals, training material, and financial reports that must keep their layout. Under the hood it is still OCR, translation, and automatic typesetting. Debug it like a pipeline.
A Hosted Loop When You Do Not Want to Own the Mask
When the source file is gone and the image is ordinary marketing art, a tool built for this loop is the short path that still tries to keep the design.
Method 4: Online AI that detects, translates, and repaints
How it works. Upload. Auto-OCR the regions. Translate, with a chance to proofread. The model erases the source glyphs and repaints the translation in a similar face. Download. OCR, machine translation, and image completion are one flow.
Tools. ReWords AI, ImageGPT, EditTextImage, Codia, and VisualGPT. Useful interfaces let you edit a suggested line and lock a line that must not change. A lock is how a brand name survives.
Input and output. PNG or JPEG in. A same-resolution bitmap out, translation rendered close to the original style. Not a layered source file. Layers mean Method 10 or Method 1.
Automation. High for one asset: upload and generate. Batch becomes credits, an API, or a queue.
Fidelity. High overall. Fills often match font style and background texture well enough that the edit is not obvious. Dense text, complex lighting, and decorative fonts still flaw. OCR tracks source quality. Neural translation varies and needs a human review. A clean look is a draft, not a proof.
Time and cost. Low to medium. Seconds to a minute. Nothing to install. A free allowance is normal. ReWords includes a few free generations per image, then paid credits. Bulk means payment or credits.
Pros and cons. You replace text on the image without the design file. Some images need a retry or a tighter instruction. Copyright and sensitive content stay yours.
Best for. E-commerce images, ads, social graphics, menus, flyers, and comics. On packaging, "SUMMER SALE" becomes "夏季特卖" with the background unchanged. An EditTextImage-style line turns "Summer sale — 30% off everything" into "Sommer-Sale — 30 % auf alles" with the layout left in place. A dense plate with leader lines belongs back in Method 2 or Method 3.
On a service built for this, you translate text in an image with a short runbook. Upload and choose the image-text function — a "SUMMER SALE" banner is the fixture. Pick the language, correct the line-by-line suggestions, and lock brand names. Generate or apply: the original glyphs come out, and the translation is redrawn in the original color, a similar face, and the existing background. Download and review at 1:1. That is the default when there is no source file and the clock is short. Numbers and proper nouns still fail, for the same reason they fail in a script.
Source Files, a Designer, a Studio, and a Phone
These four are what you do when you still hold the design, when fidelity outranks speed, or when you only need to read a sign. Two of them are the highest-fidelity options. One is the cleanest path if the source exists. One is never the file you publish.
Method 10: Edit the source, not the export
How it works. If the bitmap came from vector or layout software, replace the string in the text frame. InDesign or Illustrator, with scripts or plugins, can pull text out by OCR or by conversion, translate it, and put it back. Another route is PDF, then OCR and translate, then import. Effects and layers already exist. You are not guessing them from a JPEG.
Tools. Adobe Acrobat for OCR and an export to Word. Illustrator find-and-replace. InDesign scripts. Lokalise, or Transifex's InDesign plugin, when the strings belong in a translation memory.
Input and output. An editable document or a PDF in. Text replaced in the original file out. The PNG is a build product.
Automation. Medium, and only as scriptable as the design app. Swapping strings in tagged frames is automation. Guessing which outlined text used to be which sentence is not.
Fidelity. Very high, because you edit the source and layers survive — if you actually have the file. A flattened export is Method 2, whatever someone named it.
Time and cost. Medium. Replacement still needs a person for overflow. Cost stays low when the licenses already exist.
Pros and cons. Precise typesetting from the real file. No file, no method.
Best for. Work where you hold the design source, such as an InDesign manual. The emergency variant is a rough OCR pass for reviewers, then a refine inside the frames before anyone prints.
Method 1: A designer redraws it
How it works. A person runs the pipeline. Extract the text by OCR or by typing it off. Translate and proofread. Find or recreate the face. In Photoshop, Illustrator, or a similar tool, erase, set the translation, and rebuild effects. Order: extract, translate and proof, match the font, erase, insert, adjust.
Tools. Adobe Photoshop, Illustrator, InDesign, plus GIMP and Figma. Tesseract can speed extraction. It should not speed approval.
Input and output. Any image in, including PNG or JPEG. A new image out, or better a layered PSD or AI file. A flatten-only delivery throws away the next edit.
Automation. Low. Only OCR reasonably scripts.
Fidelity. Highest from a person on the pixels. The designer controls font, position, and style. Closeness depends on skill. It is slow and sensitive to expansion. Without the identical font, a full match is hard.
Time and cost. High. Minutes to hours, plus a learning curve on the software, so the tool is not free before the labor. Wrong default at volume. Often right for one flagship piece.
Pros and cons. Layout and style can be held for demanding brand work, including arrows and callouts. Slow, expensive, and easy to miss an accent or a shadow left over from the old word.
Best for. Ad posters, packaging, UI prototypes, and print that must stand next to the original. Atlas Cloud's point sits here: after OCR lifts the words, placing them back is still a design act. A pure raster edit merges onto a new layer with less control over individual letters. Manual setting is slower and, in skilled hands, the most precise option.
Method 9: A translator plus a DTP studio
How it works. You hire the pipeline. A translator translates. A desktop-publishing designer replaces the text. Steps: identify and edit the source if it exists (PSD or InDesign), translate, adjust the layout, re-render in the same software. A bitmap-only job is Method 1 with a real translation process. An InDesign package is Method 10 with bilingual review.
Tools. The Adobe suite, InDesign, Illustrator. CAT tools such as Trados or MemoQ keep terminology stable. The CAT tool does not export the poster. The design tool does.
Input and output. Any image or source package in. A high-quality localized image, and usually an editable file, out. Ask for the editable file. A flatten-only delivery is the wrong artifact at this price.
Automation. Low. People intervene at every stage that matters.
Fidelity. Highest, shared with careful manual design and often stricter, because a translator handles cultural nuance and a DTP operator handles format. You still review against the source and the glossary.
Time and cost. Highest of the ten. Translation fees plus design time, usually far above any machine option. Correct for sensitive work. Wrong for a tile that expires on Thursday.
Pros and cons. The quality ceiling, and the bill. Time and money sit above every automated method.
Best for. Brand campaigns, major trade-show pieces, image content in government or legal documents, and publications that need a human proofread. A wrong numeral is an incident. Put this method next to Method 1, not next to a camera app.
Method 5: Phone camera apps, for reading and only for reading
How it works. Google Translate, Baidu Translate's camera, Microsoft Translator, or WeChat scan-translate takes a photo or an import, runs OCR, translates, and overlays the result or saves a snapshot. Capture, recognize, translate, mark the original, show it. Comprehension on the spot. Not a master.
Tools. Those four, and the cousins inside other apps. No project file. No font menu that matters.
Input and output. A phone photo or screenshot in. An overlay or extracted text out. Some apps save an annotated image. That image is a reading aid.
Automation. Fully automatic. You point the camera.
Fidelity. Low, the bottom of this set. Default face, coarse mask, original layout and style not preserved. OCR and translation can sit near professional engines. The render is still not a design file you can publish. People who photograph text with Baidu Translate and save it often get the app's own mark, or a patch, on the frame.
Time and cost. A few seconds, often real time, usually a free app. Right price for "what does this say." Wrong price to treat as production art.
Pros and cons. Fast, no expertise, enough for daily life. They cannot preserve the original design. A logo or watermark may land on the output. Reference reading only.
Best for. Foreign menus, street signs, and manuals while traveling. Not the ad, not the label, not the file you print. A camera-roll screenshot is a hint at best. Rebuild with one of the other nine.
Ten Methods Side by Side
Fidelity here means how closely the translated image matches the original design, not how fluent the sentence is. Phone apps sit at the bottom of that column because they help you read and they add a large mask. Manual design and a human DTP pass sit at the top because a person protects layout and type. Do not read "low cost" as "best." A cheap path that wrecks the file is expensive on the day you reprint.
| Method | Design fidelity | OCR / recognition | Multilingual | Technical difficulty | Cost | Speed | Best fit |
|---|---|---|---|---|---|---|---|
| 1. Pro design software (manual) | Very high | N/A (human) | Any | High (needs a designer) | Highest (labor) | Slow (manual) | High-spec brand / marketing assets |
| 2. OCR + typesetting | High | High (structured text) | Many | Medium-high (software/scripts) | Low (open source) | Medium | Technical docs, labeled diagrams, structured charts |
| 3. Document translation service | High | High (professional OCR) | Very many | Low (turnkey) | Medium-high (service fee) | Fast | Bulk PDF/scans, professional documents |
| 4. Online AI tools (ReWords and similar) | High | High (AI) | Many | Low (turnkey) | Low/medium (free allowance + paid) | Fast | E-commerce posters, comics, everyday multi-image localization |
| 5. Phone camera apps | Low | Medium-high (handheld) | Many | Very low (turnkey) | Low (free) | Real-time | Quick reading for travel / daily life |
| 6. Multimodal large-model editing | High | High (model) | Very many | Low (turnkey) | Medium-high (API fee) | Medium | Infographics, ad posters, UI screenshots needing fast reasoning |
| 7. Diffusion models (DALL·E and others) | Medium-high | Medium (needs a mask) | Many | Medium (prompt craft) | Medium (API/compute) | Medium | Creative design, where a style shift is acceptable |
| 8. Custom Python pipeline | Medium | Medium (tool-dependent) | Many | High (programming) | Low (open source) | Flexible (batch) | Automation environments, private or enterprise batch jobs |
| 9. Professional human + DTP | Very high | N/A (human) | Any | Very high (expert team) | Highest (service fee) | Very slow (manual) | High-end publications, strict text and format demands |
| 10. Direct source-file translation | Very high | High (from the source) | Many | Medium (design software) | Low (if the software already exists) | Medium | Projects that still have an editable source |
Phone apps are for comprehension and add a large mask, so fidelity is low. Professional design and a human DTP pass preserve layout and type, so they score at the top. Method 8 is medium fidelity: open-source cheap to run, expensive to build. Method 4 is fast, with a free allowance and a paid tier after it. Method 7 is medium-high because long text breaks. Do not read the cheapest row as the best row.
Three Tutorials With the Code in the Fence
The three tutorials are complete enough to run or to reject. The first calls an image edit API with a mask. The second is the hosted runbook. The third is Tesseract, a translation client, and Pillow. Hash comments stay inside the fences. They are notes to the interpreter, not headings.
Tutorial 1: GPT-4 Vision, or a masked image edit
Overview. This is the smallest program that asks a vision-capable OpenAI model, or a DALL·E-style edit endpoint, to change text inside an image. The fixture translates the English "Summer sale" into the Chinese "夏季特卖."
Prep. An OpenAI account and an API key. A sample image that contains the source words. A mask aligned to that image. The filenames below are orig.png and mask.png. The mask covers the "Summer sale" region and as little else as you can manage. Mask polarity differs by API version. Read the current field definition before you burn a batch.
Flow. Load the original. Call the GPT-4V path or the DALL·E editing endpoint. The model recognizes and translates. The model generates. You download and proofread. Skip the last verb and you have a demo, not a pipeline.
-
Align the files.
orig.pngis the source.mask.pngmatches its pixel size. The mask covers the source phrase and not the product, the logo, or a legal line you did not mean to regenerate. - Call the API. The snippet uses the OpenAI Python SDK shape that takes an image, a mask, a prompt, a count, and a size. The prompt tells the model to translate the masked words into Chinese and to leave the rest untouched.
import openai
openai.api_key = "YOUR_API_KEY"
# Ask the model to translate the text in the masked area into Chinese
# while leaving everything else unchanged.
response = openai.Image.create_edit(
image=open("orig.png", "rb"),
mask=open("mask.png", "rb"),
prompt="用中文翻译上面的文字,不改变其他内容。",
n=1,
size="1024x768"
)
# Save the returned image to disk
with open("output.png", "wb") as f:
f.write(response['data'][0]['image'])
The prompt asks for Chinese and forbids other changes. n=1 requests one image. size="1024x768" sets the size in this example. The first returned image is written to output.png.
-
Review
output.png. "Summer sale" should now read "夏季特卖," with a similar face and the same background. A second fixture for this class of model is Korean in and Vietnamese out: same picture, different glyphs. Compare files, not memories. - Save it as the deliverable only after that look. If position or weight is slightly off, nudge it in an editor instead of regenerating the whole frame.
Caveats. Proofread every string, including type you thought the mask excluded. A missing line means a tighter prompt or a smaller block. The API bills by image size and by generation count. If Image.create_edit is not the symbol in your installed SDK, keep the contract — orig.png, mask.png, that prompt, one image, an explicit size — and map it to the current client.
Tutorial 2: A hosted tool, exercised like a test
Overview. Use ReWords, or a similar site such as EditTextImage, to run Method 4 on one fixture. This walkthrough uses ReWords. No local install. The test is whether you can explain the output.
Prep. Open the site. Have one image whose source text you can read yourself. If you cannot read the source, you cannot grade the translation.
Flow. Upload. The system detects and OCRs. It shows a suggested translation. You edit and lock. It renders. You download and compare.
- Upload. Choose the function that translates text in the picture and upload the fixture, an ad that contains "SUMMER SALE." Learn the clicks on a short banner before you try a menu or a comic panel. Short text makes a bad read obvious.
- Pick the language and edit the strings. Detection should return every line. Select the target. Chinese is the example. Correct mistranslations in the suggestion list before you generate. Lock tokens that are not translations: brand names, model numbers, a price already in the target currency.
- Generate. The render removes the original glyphs and draws the translation back in the original color, on the original background, in a face that tries to match. Same erase-and-redraw contract as the Python tutorial. You inspect the result as if you had watched the draw call.
- Download and compare. In the example, one side still says "SUMMER SALE" and the ReWords result says "夏季特卖," with the package or banner background still in place. Flip them at 1:1. Look at edges, shadows, and any character that might have been invented.
How to judge the run. These products usually include a free allowance. ReWords includes several free generations per image. Past that you pay or you spend credits. That is enough to learn the tool and not enough to pretend a catalog is free. Quality on this class of banner is often high, and it is still a draft until someone checks numbers and proper nouns. The method earns its place when you cannot get the design source and you still need a localized image rather than a caption under the old picture.
If you need the same loop named as a product action, edit text in an image is the job the upload-and-generate flow is for. The acceptance strings are the spec: "SUMMER SALE" to "夏季特卖," and, on the EditTextImage-style banner, "Summer sale — 30% off everything" to "Sommer-Sale — 30 % auf alles." There is no screenshot in this article on purpose. Grade the strings and the seam, not a caption under a missing figure.
Tutorial 3: Python, Tesseract, and a redraw you can debug
Overview. A script reads text off an image, translates it, and paints the translation back. The stack is Tesseract, OpenCV or Pillow, and a translation client. The fixture is Spanish into English. The example output is translated_image.png.
Flow. Load the original. Detect and OCR with Tesseract. Call a translation engine — here, the Google Translate client in googletrans. Clear the original region with a draw. Paint the English with Pillow. Save. The first version uses one known rectangle so the draw call is visible. Production uses the boxes OCR returns.
- Install the binary and the libraries. The Python package does not include Tesseract. On a Debian-family machine:
# Install Tesseract (example for Ubuntu)
sudo apt-get install tesseract-ocr
# Install Python libraries
pip install pytesseract pillow googletrans==4.0.0-rc1
The pin on googletrans==4.0.0-rc1 matches the Translator class used below. On macOS the Python packages stay the same and the system package does not. Do not skip language data. lang='spa' fails in a boring way when the Spanish traineddata is missing. If you later swap in an official translation client, keep the OCR, the rectangle erase, and the measured draw. Replace only the client.
-
Run the script. The image is
image_with_text.png. Recognition is Spanish. The destination is English. The rectangle is a stand-in, not a detector.
from PIL import Image, ImageDraw, ImageFont
import pytesseract
from googletrans import Translator
# Load the original image
image = Image.open("image_with_text.png")
# Run OCR (recognize Spanish text)
text = pytesseract.image_to_string(image, lang='spa')
print("Recognized text:", text)
# Translate the text
translator = Translator()
result = translator.translate(text, src='es', dest='en')
print("Translation:", result.text)
# For this example the text position is known; in production you would
# use pytesseract.image_to_boxes to get a box per character.
x, y, w, h = 50, 50, 300, 50 # example coordinates
# Erase the original text on the image
draw = ImageDraw.Draw(image)
draw.rectangle(((x, y), (x + w, y + h)), fill="white") # cover with white
# Draw the translated text
font = ImageFont.truetype("arial.ttf", size=36)
draw.text((x, y), result.text, font=font, fill="black")
# Save the output
image.save("translated_image.png")
What the script actually did. Tesseract reads Spanish from the whole image and prints it.
googletranstranslatesestoenand prints that too. The rectanglex, y, w, h = 50, 50, 300, 50is a known stand-in: origin (50, 50), width 300, height 50. Pillow covers it with white, draws the English inarial.ttfat 36 pixels, and writestranslated_image.png. One fixed region, one flat fill, one face you did not match. That is the illustration, not a localizer.What production changes. Take boxes from
pytesseract.image_to_boxes, union them into a line, and erase the line rather than each character. Sample a nearby color instead of hard-coding white. Measure width with Pillow and wrap or shrink when the string no longer fits — German often will not, by 30% or more. Loop for multiple lines. Use OpenCV contours when the boxes are noisy. Layout fidelity still trails an inpainting model, so review one image before you batch. A useful split is JSON of boxes and strings, then a second process that renders, so a bad name is caught before it is painted across a catalog.
Pick From the Constraint You Actually Have
Choose with three facts: how faithful the file must be, what the work may cost, and whether anyone can still open the source. A preference for a vendor is not a fact.
- No source file, and the calendar is short. This is the common case. Use an online AI tool in the Method 4 family, one you can use to translate the text inside an image without rebuilding the design by hand. Upload, choose the language, correct and lock lines, generate, download, then proofread. The render itself is seconds to a minute, plus the review you owe the file.
- The image is text-heavy and structured. Charts, manuals, diagrams. Use OCR plus typesetting (Method 2). If the job is a pile of PDFs or scans, a format-preserving document service (Method 3) can own the reflow. A 60-page PDF in minutes is a Method 3 story. A labeled schematic with leader lines is a Method 2 story, because the arrows will not survive a blind reflow.
- The file is print-bound or brand-critical. Use manual design (Method 1) or a translator plus DTP (Method 9). These are the slow, expensive, very-high-fidelity rows. Pick them when a visible patch costs more than the labor.
- The work has to live inside a system you run. Use a multimodal API (Method 6) when a prompt-and-review loop is acceptable and the images may leave the network. Use a custom Python pipeline (Method 8) when you need batch control, a private network, or a box list you can store. Use diffusion inpainting (Method 7) when the background is the hard part and the text is short enough not to shatter.
- You only need to understand the picture. Use a phone camera app (Method 5). A few seconds. Not a brand review. Overlay, default font, possible watermark or patch. Reading only. Never the file you publish.
- The editable source is still available. Edit it directly (Method 10). This is the cleanest path that exists. The moment the file is a flattened export, you are on Method 2, even if someone named it like a source.
When two constraints conflict, fidelity wins on anything public and durable, and speed wins on anything internal and temporary. A slide for tomorrow can be a Method 6 draft. A package that will be printed cannot.
Five Habits That Keep the Layout From Slipping
- Plan for length before you draw. German and French commonly expand by 30% or more. Chinese and Japanese commonly contract. Resize, re-wrap, or re-track. Never assume the new string fits the old box. Measure the rendered width and fail on overflow instead of clipping glyphs.
- Match the type, not only the language. Reproduce weight, color, letter-spacing, and any stroke or shadow. A correct sentence in the wrong face still looks wrong. If you lack the identical font, record the substitution where a reviewer will see it.
- Rebuild the background until the seam disappears. Feather the edge and check it at 1:1. On a script, leave the tutorial's white rectangle behind once the background is not white. Texture that suddenly goes smooth is an erase that did not heal.
- Account for what is not a paragraph. Arrows, callouts, numbering, and legends carry meaning or point at it. A pipeline that returns only strings will skip them. Put them on the review list.
- Proofread the words, not only the pixels. Machine translation is fast and imperfect. Numbers, units, brand names, and proper nouns need a human check before you publish. A clean fill around a wrong price is worse than a visible seam around a right one.
Keep the source image, the box list, the glossary, and the output in one job record. Otherwise a retired claim walks back in through a cached render.
What You Should Ship
Match the method to the file, not to the demo that was easiest to click.
Labels you can batch are Method 8: detect, OCR, translate, erase, redraw, seconds per image, and a fill that may be flat color or a neighborhood average. Diagrams with leader lines are Method 2, seconds to minutes for the read and the translation, then minutes to tens of minutes of typesetting. A long paid document is Method 3: fidelity described as 90% or better, a 60-page PDF in minutes, handwriting and strange layouts left out. A poster, package, menu, or comic with no source is the hosted repaint, free generations on a fixture and credits after that. "SUMMER SALE" to "夏季特卖" on the same background is a pass. An InDesign or Illustrator file is Method 10. A flatten is Method 2. A flagship, a trade-show wall, or a legal artifact is Method 1 or Method 9 — highest fidelity, highest time, highest cost.
A multimodal call is the coffee-machine diagram in Spanish, or an infographic whose words move and whose geometry stays: seconds to tens of seconds, and a miss when a line is skipped, misspelled, or far longer than the box. Diffusion is ten to tens of seconds, stronger on texture, weaker on long words. A phone app is a few seconds of reading and not a master.
Detection boxes are how you know which pixels may change. Pick the owner of that loop whose failure you can see before the file ships, and proofread the words before you trust the redraw.
Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.