PDF is a terrible file format and every alternative is worse
I've spent the last two months building a PDF editor that runs in the browser. Merge, split, rotate, sign, watermark, add text. The kind of thing people do every day without thinking about what's underneath. Six weeks i
I've spent the last two months building a PDF editor that runs in the browser. Merge, split, rotate, sign, watermark, add text. The kind of thing people do every day without thinking about what's underneath.
Six weeks in, I understood something about PDFs that I wish I'd known before I started. The format isn't a document format. It's a page description format. And the difference matters for everything you try to do with it.
What a PDF actually is
A PDF isn't a document with paragraphs and headings and a reading order. It's a list of instructions for a printer.
When you open a PDF, what you're looking at is a set of drawing commands. "Move to this X, Y coordinate. Set the font to Helvetica. Draw the glyph at index 47. Advance 12 units. Draw the glyph at index 82." It's a script. The file tells a rendering engine how to paint pixels on a page, in the order they should be painted, at exact coordinates.
That's it. Everything else we think a document has, we're inferring.
There is no concept of a paragraph in a PDF. There is no concept of a sentence, a word, a heading, a line break, or a link between text and what it represents. Text isn't stored as text in the traditional sense. It's stored as positioned glyphs with a font reference. If you wanted to extract the sentence "hello world" from a PDF, you don't read it from a string. You find every glyph, get its position, sort by position, group by proximity, and hope the author didn't do anything weird with kerning or line spacing.
This is why PDF text extraction is so unreliable. This is why copying from a PDF sometimes gives you letters in the wrong order. This is why two-column layouts break every parser. The information is technically there, but not in the form you want.
The heritage problem
PDF comes from PostScript, which comes from the 1980s, which comes from a world where the goal was to describe how a printer should physically lay ink on paper. The whole format is built around that purpose. Long before "documents" and "content" and "accessibility" were things anyone cared about.
PostScript was designed in 1982 at Adobe, mostly to solve one problem: how do you tell a printer to render a page precisely, with the same output every time, regardless of the hardware? The answer was a stack-based programming language that draws things. PDF is that language, plus a container format and some compression.
Every quirk in PDF comes from this inheritance. Text isn't text, it's glyph positioning. Fonts aren't references, they're embedded subsets. Images aren't images, they're streams of pixels with specific encoding. Layout isn't layout, it's coordinates.
And here's the thing. It works. It works incredibly well. Twenty-five years later, PDF is the only format that reliably renders the same on every device, every printer, every screen. That's exactly what it was built for, and it still does that job better than anything else.
What this means when you try to edit
You cannot "edit" a PDF in any real sense. What you can do is draw new things on top of the existing drawing instructions.
When a tool says it can edit the text in a PDF, it's lying in a specific way. It's locating the glyphs of the text you want to change, removing those glyphs from the instruction list, and drawing new glyphs in the same coordinates. The result looks correct. The underlying file has no idea what happened. If someone opens it in a different renderer with different font handling, the new text may sit slightly off from where the old text was.
For anything else in the file, it's the same story. Move an image and you're really copying the image's drawing command to a new position. Delete a page and you're removing that page's instructions from the tree. Rotate a page and you're changing the viewport orientation.
There is no semantic model. There is no "this is a heading" or "this is a signature block." Everything is coordinates. Every edit is a manipulation of the instruction stream.
Once you understand this, three things become obvious:
Text reflow is impossible. When you delete text from a Word document, the rest of the paragraph moves to fill the gap. When you delete text from a PDF, you create a hole. The other text doesn't know it exists. It sits at its coordinates, exactly where it was drawn.
Editing breaks layout in subtle ways. Change a font size and the text may overlap what's next to it. The PDF doesn't know to reposition anything. It just draws what you told it to draw.
Text search is a heuristic. When you search a PDF, the renderer is guessing where words start and end based on proximity and spacing. It's right most of the time. When it's wrong, you get the mess we've all seen.
The alternatives
Here's where it gets frustrating. Every alternative to PDF has a fatal flaw that's worse than PDF's flaws.
DOCX. The obvious one. Real document structure, real text, real editability. But: proprietary format owned by Microsoft, changes behavior between Word versions, renders differently in LibreOffice, and requires a specific application to open. People send PDFs because PDFs look the same everywhere. DOCX files look slightly different on every machine.
ODT. The open-source answer. Same structural benefits, same open standard problem: it's a standard that almost nobody uses outside of LibreOffice. Compatibility with Word is imperfect. Nobody wants to receive a .odt file.
HTML. Can carry real structure, real text, real semantics. But: it doesn't print correctly, it doesn't paginate, it renders differently in every browser, and you can't email someone an HTML file and expect them to open it as a document. HTML is a delivery mechanism, not a document format.
PostScript. The ancestor of PDF. Same underlying model, worse container format, more verbose, less common. Strictly worse than PDF for everything.
Markdown and LaTeX. Both produce documents from structured source. Markdown is too limited for anything with real layout. LaTeX is powerful but the barrier to entry is enormous and the output is still PDF at the end.
Plain text. No layout, no fonts, no images. Perfect for programs, useless for contracts.
Every single one of these fails at the specific thing PDF succeeds at: rendering identically everywhere. That's the whole reason PDF won in the first place.
Why PDF won anyway
Because the actual problem people needed solved wasn't editability. It was "I need this document to look the same on your screen as it does on mine, and if you print it, it should print the same way."
Contracts, invoices, tickets, tax forms, boarding passes, university transcripts, medical records. None of these need to be editable by the recipient. They need to be final. They need to render correctly on any device, print identically on any printer, and preserve their visual appearance for years.
PDF does that. It does it better than any other format that has ever existed. That's why every alternative fails. They all solve a different problem and hope you didn't notice.
The tradeoff is that a PDF isn't really a document. It's a picture of a document, with some text embedded in the pixels for accessibility. When you accept that, everything else makes sense.
What I do about it
I built the PDF editor to work with the format instead of against it. The tools do things PDFs are good at:
- Merge (combine instruction streams)
- Split (divide instruction streams)
- Rotate (change viewport)
- Reorder (rearrange page tree)
- Delete pages (remove from tree)
- Compress (recompress object streams)
- Watermark (draw over the existing content)
- Sign (draw an image at specified coordinates)
- Add text (draw new text at specified coordinates)
None of these are "editing the document." They're all operating on the file as what it actually is: a container of instructions.
That's the honest version of what a PDF editor does. It doesn't edit the document because there is no document. It edits the instruction stream. The result looks correct because the instructions produce the right output. But if you look under the hood, it's just drawings on top of drawings.
The open question
Is there a format that could replace PDF if someone actually tried?
Something with the universality of PDF plus real document structure. Fixed layout for print, live text for editing, standard fonts, deterministic rendering.
I don't think it exists. Every attempt I've seen either sacrifices the deterministic rendering for editability (Word) or sacrifices editability for deterministic rendering (PDF). Nobody's built the one that does both, and I'm not sure anyone can.
But I'd like to be wrong. If you've seen a format that threads the needle, tell me what it is.
check out my pdf editor here. [SORO pdf editor]
Originally published by Dev.to WebDev. Aggregated on AIWithGhost for educational purposes β full credit and traffic to the original publisher.