Combining Images into a PDF Does Not Make Them Searchable: A Practical Browser Export Workflow
You photograph ten receipts, combine them into a PDF, and send the result to a colleague. The file opens. Every receipt is visible. Then your colleague searches for a supplier name and finds nothing. Nothing necessaril
You photograph ten receipts, combine them into a PDF, and send the result to a colleague.
The file opens. Every receipt is visible. Then your colleague searches for a supplier name and finds nothing.
Nothing necessarily went wrong with the conversion. The PDF may simply contain images of text rather than actual text.
An image-to-PDF tool packages visual pages. Searchability, semantic structure, and accessibility are separate capabilities.
That distinction matters when you build exports for receipts, screenshots, catalogs, portfolios, and administrative documents. A successful download should not create expectations the file cannot meet.
1. Define the output before implementing the converter
Start by deciding what the recipient needs.
| Requirement | Image-only PDF can provide it? | Additional work to consider |
|---|---|---|
| View several images in one file | Yes | Page order and readable layout |
| Print a visual packet | Yes | Page size, margins, and effective resolution |
| Search text photographed on the pages | Not by embedding images alone | OCR and verified text output |
| Copy text from a receipt photo | Not by embedding images alone | A usable text layer |
| Navigate a structured document | Not automatically | Structure, tags, reading order, and testing |
W3C's PDF OCR technique addresses adding actual text to scanned documents. OCR can support search and reading, but recognized text still needs review, and OCR alone does not establish complete accessibility.
For an expense packet, images may be sufficient. For a document someone must read with assistive technology, a different authoring workflow may be necessary.
Make that distinction visible before export, not after the user has uploaded a finished document elsewhere.
2. Page order is part of the data
File-selection order, filename order, and the intended document order are not always the same.
Consider:
receipt-1.jpg
receipt-10.jpg
receipt-2.jpg
Plain lexical sorting places receipt-10 before receipt-2. Numeric-aware filename sorting can help, but it still cannot know the intended chronology when names are inconsistent.
For a proposed interface, I would show numbered previews before export and allow explicit reordering. Keep that order in a stable array instead of reconstructing it from whichever asynchronous conversion finishes first.
const documentPages = [
{ id: "receipt-a", file: firstFile },
{ id: "receipt-b", file: secondFile },
];
// Export this array in order; do not sort by completion time.
If one input fails, tell the user which page is missing. Do not quietly produce a nine-page document that looks like a complete ten-page packet.
3. Fit the image without changing its proportions
A landscape screenshot should not be stretched to fill a portrait page. A tall receipt should not be cropped just to avoid whitespace.
For a reference layout, fit the full image inside a page box while preserving its aspect ratio. PDF coordinates are commonly expressed in points; one point is 1/72 inch.
function fitImageToPage({
imageWidth,
imageHeight,
pageWidth,
pageHeight,
margin = 24,
}) {
for (const value of [imageWidth, imageHeight, pageWidth, pageHeight]) {
if (!Number.isFinite(value) || value <= 0) {
throw new RangeError("Dimensions must be positive finite numbers");
}
}
if (!Number.isFinite(margin) || margin < 0) {
throw new RangeError("Invalid margin");
}
const availableWidth = pageWidth - 2 * margin;
const availableHeight = pageHeight - 2 * margin;
if (availableWidth <= 0 || availableHeight <= 0) {
throw new RangeError("Margins leave no drawing area");
}
const scale = Math.min(
availableWidth / imageWidth,
availableHeight / imageHeight,
);
const width = imageWidth * scale;
const height = imageHeight * scale;
return {
width,
height,
x: (pageWidth - width) / 2,
y: (pageHeight - height) / 2,
};
}
This computes placement. It does not resample or improve the source image.
For unusually tall receipts, you might offer a custom-height page or a deliberate split across multiple pages. Explain the option. Cropping off a total is not a formatting improvement.
4. Embedding and re-encoding are different choices
Some PDF libraries can embed JPEG or PNG bytes directly. Others use a canvas-based path that decodes the source and creates a new encoded image first.
Re-encoding can change quality and file size. Direct embedding avoids that particular conversion step, but does not solve layout, orientation, or every compatibility issue.
The PDF-LIB API reference documents image embedding and document serialization. Here is a deliberately narrow reference example using that library as an injected dependency:
async function makeImagePDF(files, PDFDocument) {
const document = await PDFDocument.create();
for (const file of files) {
const bytes = await file.arrayBuffer();
let image;
if (file.type === "image/jpeg") {
image = await document.embedJpg(bytes);
} else if (file.type === "image/png") {
image = await document.embedPng(bytes);
} else {
throw new Error(`Unsupported input in this example: ${file.name}`);
}
const page = document.addPage([595.28, 841.89]);
const placement = fitImageToPage({
imageWidth: image.width,
imageHeight: image.height,
pageWidth: page.getWidth(),
pageHeight: page.getHeight(),
});
page.drawImage(image, placement);
}
return document.save();
}
The calling application supplies PDFDocument from its installed pdf-lib dependency and uses the layout helper above. This is not BatchSet's implementation.
The example supports only JPEG and PNG, trusts file MIME labels as routing hints, and aborts on the first error. It does not normalize orientation, accept WebP, run OCR, add document tags, or apply production resource limits. Real decoding must validate bytes rather than treating a MIME label as proof of the file's format.
WebP input requires an appropriate conversion path before using these embedding methods. Do not assume every image format accepted by the browser is directly embeddable by your PDF library.
5. Judge resolution at the final placed size
A 1200-pixel-wide image printed six inches wide has an effective resolution of 200 pixels per inch:
1200 pixels / 6 inches = 200 pixels per inch
That is an illustrative calculation, not a guarantee of readable receipt text. Source sharpness, camera focus, contrast, compression, and the actual printer still matter.
Changing a DPI metadata field does not manufacture missing detail. Neither does placing a blurry photo into a higher-resolution canvas.
Inspect the smallest relevant text at the intended print size. For a screenshot packet, that may be an error message. For a receipt, it may be a tax identifier or total. If those details fail, a visually pleasant thumbnail is not enough.
6. A PDF wrapper is not a compression guarantee
The PDF must store image content and document structures. It can be larger than the source images, depending on representation and overhead.
My suggested export report includes:
- Input image count and actual output page count.
- Total input encoded bytes.
- Final PDF bytes.
- Any images resized or re-encoded.
- A list of failures or omissions.
If size reduction is required, make it a separate option and compare the actual output. Some screenshots and small images become larger when re-encoded.
Avoid blanket βcompress every page to JPEGβ behavior. Text screenshots, line art, transparent graphics, and photographs can have different needs. Review representative pages before applying one quality setting to the whole packet.
7. Browser-based does not mean unlimited
Even processing one page at a time, the PDF document object may retain embedded resources until serialization. Generating the final byte array can introduce another memory peak.
For large packets, expose a realistic batch limit, provide progress, and support cancellation where the chosen library permits it. If an operation has no abort interface, explain that cancellation may stop the next page rather than instantly interrupt the current one.
Keep originals available until the user has inspected the final file. Do not call a packet βcompleteβ before export and verification finish.
Local processing can keep source files away from a conversion server, but it does not prevent the user from later uploading the result to an email or document service. Privacy language should describe the actual conversion stage.
8. Verify what the recipient will receive
For a release or an important packet, my checklist is:
- Open the exported PDF in more than one viewer.
- Confirm page count and intended order.
- Check orientation, margins, and uncropped edges.
- Zoom into small text and review a print proof if printing matters.
- Search a known word to determine whether a text layer exists.
- Compare final bytes with any receiving system's size limit.
- Use an appropriate accessibility workflow if that is a requirement.
Searchability testing should include checking OCR accuracy where OCR was applied. Finding one word is not evidence that every page's text is correct.
Combine the pages with clear expectations
BatchSet's Image to PDF tool combines JPG, PNG, and WebP images into a PDF in the browser, with page-size and ordering options.
Use it when you need a visual packet in one file. Do not assume image conversion adds OCR, semantic tags, or an accessibility guarantee. Verify the exported pages, and choose a text-aware document workflow when the recipient needs more than pictures of words.
Originally published by Dev.to WebDev. Aggregated on AIWithGhost for educational purposes β full credit and traffic to the original publisher.