Skip to content
Genom

Open a PDF in your browser

Free, no sign-up, and the file never leaves your computer.

Samplesample.pdf
Loading the viewer…
Also opensDrop any of them on the viewer above — the format is recognised from the bytes.

Drop a PDF above. It renders in the page, and the text comes back in reading order rather than in the order the glyphs happen to be drawn.

Nothing is uploaded. The same viewer that shows a Word file shows this one, through the same API.

How the format actually works

A PDF does not contain sentences

A PDF contains instructions to place glyphs at coordinates. There are no words, no lines and no paragraphs — those are things a reader infers from where the marks landed. Recovering text means clustering glyphs into lines by their baselines, working out where the spaces are from the gaps, and deciding which column comes first. This is why copying out of a PDF so often produces text in the wrong order, and why extraction quality varies so much between tools that all claim to support PDF.

What is supported

Everything below is parsed from the file, not approximated

Read from the bytes the format actually stores, into a model that does not know which generation produced it.

The file and its objects

Cross-reference tables and cross-reference streams, incremental updates, object streams, and the document catalogue down to the page tree. Encrypted documents where the scheme allows it.

Filters

Flate with all its predictors, LZW, ASCIIHex, ASCII85 and RunLength for data; and for images DCT (JPEG), JPX (JPEG 2000), CCITT Group 3 and 4 fax, and JBIG2 — the last three being the ones scanned documents actually use and most readers skip.

Fonts

Embedded Type 1, CFF and TrueType programs are parsed and rebuilt into something the browser will load, including Type 1 and Type 2 charstring interpretation. Encodings, CMaps and the standard fourteen fonts with their real metrics, so text spaced by a font that is not in the file still lands where it should.

Page content

The content stream interpreted as it is written: graphics state, transformation matrices, paths and clipping, colour spaces including Indexed, Separation and ICC-based, functions of all four types, and shading — axial, radial and mesh.

Text recovery

Glyphs clustered into lines by baseline, spaces recovered from the gaps rather than from characters that are usually absent, and reading order worked out across columns. This is what makes copying out of a document produce prose instead of a jumble.

Structure worth keeping

Table regions recovered from ruling lines and alignment, so a table in a PDF comes out of extraction as a table rather than as rows of loose words.
Being straight with you

What this does not do

Every renderer has a list like this. Most do not publish it.

  • A scanned page is an image with no text in it: it displays, but there is nothing to select and no OCR here.
  • Interactive forms are drawn in their current appearance and cannot be filled in.
  • Annotations other than links are not rendered.
  • Tagged-PDF structure trees are not used; reading order is inferred from geometry even when the file states it.
  • Some encryption schemes are detected but not opened.

If a file of yours renders wrong, send it to admin@genom.dev. It joins the corpus everything is measured against, which is what makes a fix stay fixed.

Questions

Things people ask

Is my PDF uploaded anywhere?
No. The file is opened by JavaScript running in this tab and is never sent anywhere — there is no upload, no server-side conversion and no copy kept. You can confirm it: open your browser’s network panel, then open a file. Nothing leaves. This also means the page works with no connection at all once it has loaded.
Can I select and copy text?
Yes. Text is recovered into reading order and rendered as real text, so selection, search and copy behave as you would expect rather than following the drawing order.
Does it handle scanned PDFs?
A scanned page is an image with no text in it, so the page displays but there is nothing to select. There is no OCR here.
Why offer a PDF viewer at all?
So that an application handling documents has one API for every format it might be handed, rather than one library for PDF and a different one for everything else.
Using it in your own code

The viewer, and the parser on its own

The parser produces a model that knows nothing about CSS or the DOM, so the same call runs in Node for extraction, search and conversion. PDF and its other generation are one package.

import { DocumentViewer } from '@genomdev/react';

<DocumentViewer file={file} fit="width" />;