What Building a Browser-Based PDF Editor Taught Us About PDF Coordinates, State and Export
Building a PDF editor looks deceptively simple at first. Render a PDF in the browser, place some controls around it, let the user add text or images, and then save the result. That sounds reasonable until the first obj
Building a PDF editor looks deceptively simple at first.
Render a PDF in the browser, place some controls around it, let the user add text or images, and then save the result.
That sounds reasonable until the first object has to survive zooming, page resizing and export.
While developing the Edit PDF tool for SwiftVecto, we found that the visible editing interface was only a small part of the problem.
The harder work was underneath it.
How do you make sure an object placed at a particular position in the browser appears at exactly the same position in the exported PDF?
How do you represent edits without constantly rewriting the original document?
What happens when pages have different dimensions?
What does "redact" actually mean?
And when should an operation happen in the browser rather than on the server?
These turned out to be some of the most interesting engineering problems in the project.
1. The screen is not the PDF
One of the first problems is coordinate systems.
Browser interfaces normally work from the top-left:
0,0 ─────────────→ X
│
│
↓
Y
PDF coordinate systems traditionally work from the bottom-left:
Y
↑
│
│
0,0 ─────────────→ X
Already, an object's Y position cannot simply be copied from the browser into the PDF.
But that is only the beginning.
Now introduce:
- zoom
- page rotation
- crop boxes
- media boxes
- different page dimensions
- device pixel ratios
- responsive editor layouts
A text object that appears at (300, 420) on the screen does not necessarily belong at (300, 420) in the PDF.
This becomes especially obvious when zooming.
Suppose the user positions an object while viewing a page at 80% zoom. If its stored coordinates are based on screen pixels, changing the zoom can effectively change the document.
That is the wrong model.
Zoom should change the view, not the underlying document state.
A much safer architecture is to treat coordinate conversion as its own subsystem:
PDF coordinates
↕
Editor coordinates
↕
Screen coordinates
Features then use the same transformation layer rather than implementing their own coordinate calculations.
That becomes increasingly important as the editor grows.
2. Don't rewrite the PDF after every edit
Another important decision was how to represent the document while it is being edited.
The tempting approach is:
User adds text
↓
Modify PDF
User adds image
↓
Modify PDF again
User moves object
↓
Modify PDF again
That quickly becomes difficult to manage.
Instead, we found it much more useful to think of the editor as maintaining its own document state.
Conceptually:
Original PDF
│
├── Original pages
│
└── Editor state
├── Page operations
├── Text objects
├── Images
├── Drawings
├── Annotations
├── Links
├── Forms
└── Redactions
An editor object can then have a structure along these lines:
{
id: 'obj_82',
type: 'text',
pageId: 'page_1',
bounds: {
x: 140,
y: 215,
width: 180,
height: 32
},
content: 'SwiftVecto',
style: {
fontFamily: 'Helvetica',
fontSize: 16,
bold: false,
italic: false,
color: '#111827',
opacity: 1
}
}
The important part is not the exact schema.
It is the separation between the source document and the editing state.
That separation makes features such as moving objects, undo/redo, deleting objects and changing properties considerably easier to reason about.
3. Stable page identity matters
Page management introduces another interesting problem.
Imagine a five-page document.
The user adds something to page 3.
Then they move page 3 to the beginning of the document.
Is the object attached to page 3, or to the page that is now page 1?
If objects are associated only with page numbers, reordering becomes dangerous.
A better model gives each page a stable internal identity.
For example:
{
id: 'page_a81f',
sourcePage: 3,
rotation: 0,
deleted: false,
objects: []
}
The visible page number can change.
The page identity does not.
That distinction makes page reordering much easier to manage because objects belong to a page entity rather than to a temporary position in an array.
4. Added content and existing content are different problems
Adding new text to a PDF and editing text that already exists in the PDF may look like the same operation to a user.
Technically, they are very different.
Adding new text can often be represented as a new editor object:
Page
+ text object
Existing PDF text may instead be represented internally as positioned glyphs, text runs, font references and drawing instructions.
There may not be anything resembling the editable paragraph you would find in a word processor.
This is one reason PDF editing becomes significantly more complicated once an editor moves beyond adding objects and starts modifying the document's existing content.
The UI may make both operations look similar.
The document model underneath them is not.
5. Redaction is not a black rectangle
This was one of the most important distinctions in the editor.
Visually covering information is not necessarily redacting it.
If sensitive text remains inside the PDF and a black rectangle is simply drawn over it, the underlying information may still exist.
Depending on how the PDF is constructed, it might still be extractable, searchable or recoverable.
A real redaction workflow therefore needs to treat redaction differently from ordinary drawing.
Conceptually:
Select sensitive region
↓
Identify affected content
↓
Remove or irreversibly alter that content
↓
Apply the visible redaction appearance
↓
Generate and validate the resulting PDF
That is fundamentally different from:
Draw black rectangle
The two may look identical on screen.
From a document-security perspective, they are not remotely equivalent.
6. OCR is another separate layer
OCR creates a similar distinction.
A scanned PDF may contain pages that are effectively just images.
A human can read the text.
The PDF cannot necessarily search or select it.
OCR can identify characters in those images, but an editor then has to decide what to do with that information.
Possible outcomes include:
- extracted text
- searchable text layers
- selectable text
- positioning data
- confidence information
So "run OCR" is not really a single UI operation.
It is a document-processing pipeline.
That affects where the feature belongs architecturally.
7. Not everything belongs in the browser
We wanted the editor to feel immediate, so interactive operations naturally belong close to the user.
Things such as:
- rendering
- selection
- dragging
- resizing
- drawing
- annotations
- signatures
- page thumbnails
- reordering
- zoom
- undo and redo
are well suited to browser-side interaction.
But some operations have different requirements.
For example:
- OCR
- true redaction
- document sanitisation
- complex content removal
- PDF repair or normalisation
- final output validation
may require heavier or more controlled processing.
For SwiftVecto, this led to a hybrid model.
The browser owns the interactive editing experience.
The processing layer handles operations where heavier processing or output correctness becomes more important.
This also means privacy claims need to be precise.
It would be misleading to say that every PDF operation always happens locally if some features legitimately require server-side processing.
8. Export is where everything has to agree
An editor can look perfect in the browser and still produce a bad PDF.
Export is where all the architectural decisions finally meet.
Conceptually, our pipeline looks more like:
Original PDF
+
Editor state
↓
Apply page operations
↓
Apply content changes
↓
Apply annotations
↓
Apply images, text and signatures
↓
Apply forms
↓
Apply redactions
↓
Validate output
↓
Final PDF
If the coordinate model is wrong, export exposes it.
If page identity is wrong, export exposes it.
If zoom has leaked into document coordinates, export exposes it.
If reordered pages have lost their objects, export exposes it.
This is why treating export as an afterthought is risky.
The editor and export pipeline need to share the same understanding of the document.
The biggest lesson
The biggest lesson from building the editor has been that a PDF editor is not primarily a toolbar.
The toolbar is simply the visible interface to a much larger document system.
Behind buttons such as:
Text
Image
Sign
Draw
Shapes
Annotate
Link
Forms
Redact
Crop
OCR
Pages
are several different engineering problems involving document state, coordinate transformation, page identity, rendering, processing and export.
Once those concerns are separated properly, adding features becomes much more manageable.
Without those boundaries, every new feature risks becoming another special case.
We wrote a more detailed breakdown of the architecture and the problems we encountered in the original SwiftVecto guide:
https://swiftvecto.com/guides/building-a-browser-based-pdf-editor
The browser-based editor discussed here is also available to try on SwiftVecto:
Originally published by Dev.to WebDev. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.