SwiftVecto SwiftVecto SwiftVecto

Search Results

No matching tools found

Try searching with a different keyword or browse one of our tool categories.

↑ ↓ Navigate Enter Open Esc Close
PDF 11 min read

Building a Browser-Based PDF Editor: The Problems You Don't See in the UI

A practical look at the engineering behind browser-based PDF editing, including coordinate systems, editor state, page operations, OCR, redaction and reliable export.

By SwiftVecto Published

A browser-based PDF editor can look deceptively simple. Open a document, click a toolbar, place some text or an image, move it into position and save the result. From the user's point of view, that is exactly how it should feel.

The engineering underneath is different. A PDF is not an HTML page and the browser is not editing the document directly. The editor has to reconcile the PDF's own page geometry and content with a visual workspace that can zoom, scroll, drag, resize and reorder objects. Then it has to turn that interactive state back into a reliable PDF.

While building SwiftVecto's Edit PDF workspace, several problems repeatedly proved more important than the visible toolbar. These are some of the problems behind the interface.

Rendering a PDF is not the same as editing it

The first distinction is fundamental. A PDF renderer can show a page and expose useful information about its text and geometry, but that does not automatically make the document editable.

An editor needs another layer above the rendered document. That layer has to know which page is active, which objects the user has added, where those objects are positioned, how they are styled, whether pages have moved, and what should happen when the user presses Undo, Redo or Save PDF.

Original PDF
    +
Editor state
    ↓
Visual editing workspace
    ↓
Export pipeline
    ↓
Edited PDF

That separation became one of the most useful architectural decisions in the editor. User actions change editor state first. The source PDF does not need to be destructively rewritten every time an object moves by a few pixels.

The coordinate-system problem

Positioning is one of the easiest parts of PDF editing to underestimate. Browser coordinates normally start at the top-left and increase down the screen. Traditional PDF coordinates are based on a different page coordinate system. Even before editing begins, those two spaces need to be translated correctly.

Browser view

0,0 ─────────→ X
 │
 │
 ↓
 Y

PDF page

 Y
 ↑
 │
 │
0,0 ─────────→ X

Now add zoom, page rotation, crop boxes, different page dimensions and display scaling. A text object that appears correctly at 80% zoom still has to land in exactly the same place in the exported document. The same applies to signatures, drawings, links, form fields, redactions and images.

The practical lesson is that individual editing features should not invent their own position calculations. A shared transformation layer should translate between PDF coordinates, editor coordinates and screen coordinates. Otherwise small inconsistencies accumulate until the exported file no longer matches what the user saw.

Zoom changes the view, not the document

This sounds obvious, but it has consequences throughout an editor. When the user zooms from 80% to 120%, an object becomes larger on screen. Its logical position on the PDF page should not change.

Dragging and resizing therefore cannot simply store whatever screen pixels happen to be visible at the time. The editor needs a stable representation independent of the current viewport. That stable state is what allows an object to survive zoom changes, page navigation and final export without drifting.

Page operations affect the whole document model

Page management is more than changing thumbnails. If page 3 becomes page 1, every object associated with that page still needs to belong to the correct page. If a blank page is inserted, the export process must create it. If pages are deleted or reordered, the final document has to reflect the same structure the user sees in the editor.

This is why stable internal page identifiers are useful. The visible page number can change while the editor continues to know which logical page owns a particular text object, signature, annotation or form field.

Added text is easier than existing text

Adding a new text box is relatively controlled. The editor owns the text, font choice, position and styling from the beginning. Editing text that already exists inside a PDF is a different problem.

Existing PDF text may use embedded or subsetted fonts, transformed glyphs and positioning information that does not map neatly to an ordinary browser text field. A replacement also has to deal with the original content rather than merely painting new text on top of it.

That is why a serious editor should be careful about promising perfect reproduction of every original font. The visual result matters, but so does being explicit about what the document format makes possible.

Redaction cannot just look correct

Redaction is the clearest example of the difference between appearance and document state. Drawing a black rectangle over an account number can make the number invisible while leaving the original information underneath. Depending on how the PDF was constructed, that hidden content may still be selectable, extractable or recoverable.

Visual concealment and redaction are not the same operation. A redaction workflow has to remove or sanitise the underlying content represented by the marked region, not simply cover it.

This is one of the operations where SwiftVecto deliberately crosses the browser/server boundary. Interactive marking belongs naturally in the browser, while authoritative document processing and validation can require the server-side PDF pipeline.

OCR solves a different problem

A scanned PDF may contain pages that are effectively images. There may be little or no useful text layer for the editor to work with. OCR can analyse those page images and return recognised text together with positional information.

Scanned page
    ↓
OCR
    ↓
Recognised text + coordinates
    ↓
Editor interaction

OCR does not magically turn every scan into a perfect original document. Recognition quality, language, image quality and layout all matter. What it does provide is another source of structured information that the editor can use when the PDF itself does not contain usable text.

Forms and links must remain real PDF objects

Some editor features are easy to fake visually. Blue underlined text can look like a link, but an exported PDF needs an actual link annotation if it is supposed to be clickable. Likewise, a rectangle labelled "Signature" is not automatically a PDF form field.

The export stage therefore needs to understand the intent of editor objects. Text, images, annotations, links and form fields are not interchangeable just because they can all be drawn on a canvas.

The browser/server boundary should follow the work

Keeping every operation in the browser sounds attractive, but it is not automatically the best architecture. Interactive work benefits from immediate browser feedback. More demanding document transformations benefit from a controlled processing pipeline.

Browser responsibilities

  • Rendering and navigation
  • Selection, dragging and resizing
  • Drawing and annotations
  • Page thumbnails and reordering
  • Zoom and editor state
  • Undo and redo

Processing responsibilities

  • OCR when required
  • True redaction
  • Complex content removal
  • Document normalisation
  • Final PDF generation and validation

This also affects privacy language. It would be inaccurate to claim that every Edit PDF operation always remains on the device if some operations legitimately use server-side processing. Privacy claims should describe the real processing path rather than the most attractive marketing version of it.

Export is where the editor has to prove itself

An editor can look excellent on screen and still fail if the saved PDF is wrong. Export is where page state, coordinates, object types and document operations have to converge.

Source PDF + editor state
    ↓
Apply page operations
    ↓
Apply content changes and annotations
    ↓
Apply images, text, signatures, links and forms
    ↓
Apply redactions where required
    ↓
Validate and generate final PDF

The important test is not simply "does Save PDF produce a file?" It is whether the exported document faithfully represents what the user created, across different page sizes, zoom levels and combinations of editing features.

What the toolbar does not show

Features such as text, images, signatures, drawing, shapes, annotations, links, forms, redaction, cropping, OCR and page management are the visible product. The less visible work is what allows those features to coexist without each one becoming a separate mini-editor.

A stable document model, shared coordinate transformations, explicit browser/server responsibilities and a reliable export pipeline are what turn a collection of toolbar buttons into an editing system.

That has been one of the main lessons from developing Edit PDF for SwiftVecto: in document software, what looks correct in the interface is only half of the job. The saved document is the other half.