@genomdev/pdf
PDF — parser, model, content adapter and viewer
120 exported symbols across 2 entry points
@genomdev/pdf— 114 exports@genomdev/pdf/view— 6 exports
@genomdev/pdf
Classes
A dictionary. Keys are stored without the slash. Values are stored *raw* — a `Ref` stays a `Ref` — and resolving one needs the cross-reference table, which is why lookup that follows references lives on `XRef` and not here. That separation is what lets a dictionary be parsed out of a content stream, where there is no document to resolve against at all.
A PDF name object: `/Type`, `/Font`, `/Contents`.
Raised when the byte offset given for an object holds something else.
A stream: a dictionary and a run of bytes whose meaning the dictionary gives. The bytes are kept encoded. Decoding needs the filter chain, the document's decryption key and, for an image, a decision about whether to decode at all — a JPEG is better handed to the platform than unpacked here — so it is a method on the document rather than a property of the object.
A PDF string: bytes, with the two text decodings available on request. `(Hello)` and `<48656C6C6F>` are the same object; the difference between literal and hexadecimal syntax is spelling. What the bytes *mean* depends entirely on where the string was found.
An indirect reference: `12 0 R`.
Functions
Decodes to one bit per pixel, rows padded to a byte. The output is what a PDF image of `/BitsPerComponent 1` expects, in the sense the image's own `/Decode` array will be applied to: a 0 bit is black unless `/BlackIs1` said otherwise.
A stream's bytes with every non-image filter undone. Failure of one filter is not failure of the document: a stream that will not decompress yields what came out of it, and a page missing one image is a better answer than a document that will not open. The exception is the empty result, which is returned as such and lets the caller tell "nothing here" from "something unreadable".
The matrix from PDF user space to the page as displayed. Flips the y axis, moves the origin to the crop box's corner, and applies the page's `/Rotate`. Derived rather than tabulated so that the four rotations cannot drift apart: the centre of the box is a fixed point of the rotation, which is what the offsets below say.
The dictionary of a stream, or the dictionary itself.
The size of a page as it is displayed: cropped, rotated, in points.
The first element of `/ID`, which the key derivation mixes in.
Rebuilds a font for the browser, or reports that it cannot be. Returns `undefined` for a font with no usable program — a standard-14 face the file did not embed, a Type 3 font (whose glyphs are content streams and are drawn rather than typeset), a program in a format nothing here reads. The caller then falls back to a substitute, which is what the viewer does anyway when this is not called at all.
Runs a page's content streams and returns everything they draw.
Содержимое PDF — структура, восстановленная из геометрии страницы. Для того, кто знает, какой у него файл. Принимает и уже открытый документ: вьюверу незачем разбирать файл второй раз, чтобы поискать в нём.
Every filter named on a stream, in application order, with its parameters.
What a glyph name says, according to Adobe's list.
The identity CMap, which is what `/Identity-H` and `/Identity-V` name. Two bytes per code, code equals CID. Also the honest fallback for a predefined CJK CMap we do not carry: the *splitting* is right, so the glyphs and their advances land in the right places even where the characters cannot be named.
The same image as a PNG, for consumers that want one file rather than pixels.
A decoded JPEG as RGBA. The colour transform is decided by the component count and the Adobe marker: three components are YCbCr unless the marker says otherwise, four are YCCK if it says 2 and CMYK if it says anything else. What the marker does **not** decide is whether the samples are inverted, and that is the trap. Photoshop writes CMYK JPEGs with every value flipped, and the PDF that embeds one says so with `/Decode [1 0 1 0 1 0 1 0]` — so the file's own statement is the one to follow. A reader that inverts whenever it sees an Adobe marker gets those files right and turns every *other* CMYK JPEG into a photographic negative, which is a striking way to be wrong.
Resolves a colour space object: a name, or an array with its parameters.
Builds a font from its dictionary.
Builds a callable from a function object, or an array of them.
`a` then `b`: the matrix that applies `a` first.
A glyph name to text, including the four conventions that are not the list. `uni0041`, `u1F600`, `g23`/`cid42` (no text at all — an index into the font), and the `name.alt` suffix a subsetter adds. Between them these cover most of what the glyph list misses in real files.
`/Contents` is one stream or an array of them, joined by a newline.
Parses a CMap program — a ToUnicode stream, or an embedded encoding CMap. `asEncoding` is for the second of those. An encoding CMap is supposed to map its codes with `cidchar`/`cidrange`, but a working minority of producers write `bfchar`/`bfrange` instead — the syntax is the same and the destination, a two-byte string, is the CID written as hex. Read only as text, such a font maps every code to CID 0 and the page comes out blank.
Reads a font program's table directory, and the tables worth reading. Takes the bytes of a `.ttf`, `.otf` or `.ttc`, and the first font of a collection where it is one. Nothing is returned for a file whose sfnt version is not one of the four in use — which is the check that keeps a `.pfb`, a bitmap font or a truncated download from being read as if the numbers at its head meant offsets. Every bounds check here is load-bearing rather than defensive: a font embedded in a document is arbitrary bytes from a stranger, and half of them have been through a subsetter.
The same order, with the tables left standing. A table is the one thing on a page that the recursive cut below gets exactly wrong: its columns are, to the cut, columns, so a table is read downwards — every number of the first column, then every number of the second — which loses the only thing a table means. So the tables are found first (`tables.ts`) and each is carried through the ordering as a single unit, neither split nor reordered, and comes out with its rows intact.
Puts the lines of a page into reading order. A recursive cut: find the widest vertical gap that no line crosses, split there, and order the halves left to right; if there is no such gap, split on the widest horizontal one and order top to bottom. This is the classical XY-cut, and it is here because the alternative — sorting by y — reads a two-column paper as alternating sentences from both columns, which is the single most visible failure of naive PDF extraction.
Rebuilds a TrueType or OpenType program with a new character map. The original tables are kept byte for byte — the outlines, the hinting, the metrics — so nothing about how the glyphs look can change here.
The rotation the matrix applies, in degrees, for text that is not upright.
Runs one charstring and returns the outline it draws.
How much the matrix scales lengths, along each axis. Used for the font size a glyph is really drawn at and for deciding whether two glyphs are on one line. Taken as the length of the transformed unit vectors, which is right for rotation and shear as well as for plain scaling — `m[0]` alone is the answer only for an upright matrix, and text set on its side has `m[0]` of zero.
`ABCDEF+Helvetica` is Helvetica with six letters saying it is a subset.
Splits a page's text marks into lines. Two glyphs are on the same line when their baselines agree to within a third of the type size *and* they overlap or nearly touch along the writing direction. The second half matters: a two-column page has two lines at every baseline, and a rule that looks only at y joins them into one sentence that reads across the gutter.
Reading a PDF as content. Every other walker in this package reads a structure the producer wrote down: a paragraph is a `<w:p>`, a cell is a `<c>`, a slide is a slide. A PDF has no such thing. It has glyphs at coordinates, and a paragraph is something this file *decides* — from where the lines are, how far apart, how they line up and what size they are set in. Two rules keep that honest. **The address is the file's order, not the reading order.** A block's locator names the index of the text run it starts at, counted in the order the page's content stream drew them, because that is a fact about the file that no change to these heuristics can move. The reading order decides what comes out first; the addresses stay where they are. It is the same separation the presentation walker makes, and for the same reason — a tuned heuristic must not invalidate every address anybody has stored. **Nothing is inferred that the geometry does not support.** A heading is a line set larger than the page's body size, with space above it, on its own — not a line that "looks like a title". A list item begins with a bullet or a number *and* is indented past the lines around it. Where the evidence is thin the block stays a paragraph, because a paragraph is never wrong, and a wrong heading reorganises somebody's whole retrieval index.
Builds a whole OpenType file around a bare CFF program. `/FontFile3` is a CFF with no sfnt wrapper at all — the format PDF prefers for a Type 1 font and what Distiller writes for almost everything. A browser cannot load one; wrapped in the eight tables below, the same bytes are an ordinary OTF. The metrics come from the PDF rather than from the font, which is not a compromise but the correct source: the widths the document lays out with are the widths in its font dictionary, and a disagreement between those and the font's own is resolved in the document's favour by every reader.
Builds a CFF from glyphs already interpreted into paths. Glyph 0 must be `.notdef` and is added if the caller did not supply one; the order of the rest is the order given, and that order is what a caller's character map will index by.
Interfaces
A rectangle as PDF states it, normalised so that x0 < x1 and y0 < y1.
CCITT Group 3 and Group 4, the fax codecs. This is what a black-and-white scanned page is made of, and it is the reason a scanned document without it draws as nothing at all rather than as something imperfect: there is no partial reading of a fax stream. The idea is older than the file format and simpler than it looks. A row of a bilevel image is a sequence of runs of white and black; Group 3 writes those run lengths with a Huffman code (one table for white, one for black, because the lengths are distributed quite differently); Group 4 writes each row as the *difference* from the row above it, which for a page of text is almost nothing at all. Everything below is those two things plus the housekeeping. `K` in the parameters says which: negative is Group 4, zero is Group 3 one dimension, positive is Group 3 with a mode bit at the start of every row.
A glyph selected by a character: what the new `cmap` will say.
An image ready to be handed to a browser or written out.
One character code, decoded.
Type 1 charstrings, interpreted into outlines. The conversion to CFF could be done operator by operator — Type 1 and Type 2 charstrings are cousins — and it is not done that way here. Type 1's oddities are not in its drawing commands but in the *escape hatch* it grew: flex and hint replacement are implemented by calling PostScript subroutines through `callothersubr`, with arguments passed on a separate stack and results fetched back with `pop`. A translation that tried to preserve those would have to preserve the protocol. So the charstring is *run* instead, and what comes out is a path: absolute points in glyph space, plus the advance width. Emitting a Type 2 charstring from a path is then arithmetic (`cff-writer.ts`). What is lost by going through a path is the hinting, which a screen at any reasonable size and a renderer with its own greyscale antialiasing do not miss.
Decodes an embedded JBIG2 image to one bit per pixel, rows padded to a byte. A 1 bit is **black** here, which is JBIG2's convention and the opposite of what a PDF image of `/BitsPerComponent 1` means by it — the caller inverts.
One glyph, placed.
One show operation: the glyphs of a single `Tj`, `TJ`, `'` or `"`. Kept as the file wrote it rather than split into words or lines. What a line is, is a question about geometry that the extractor answers with rules of its own; a run is what the *producer* considered one thing, and that information is worth keeping because it is often the only clue that two glyphs a millimetre apart belong to the same word.
One font program, read. Everything a caller can learn about a font without rasterising it: how many glyphs it has, how big its em is, which character reaches which glyph, what the glyphs are called — and the raw bytes of every table, so that a question this does not answer can still be asked of the file rather than of a second copy of it.
What a Type 1 program says about itself.
Type aliases
An embedded program, parsed as far as this package reads it.
Every mark may carry a `clip`: the rectangle the page had in force when it was drawn, in device units, and present only when it actually cuts the mark. It is a *box* and not a path, deliberately. The interpreter intersects the bounding boxes of the clipping paths it meets, so what a mark carries is a superset of the region the file describes — never smaller, so nothing that should be visible is ever hidden by it, and for the overwhelming majority of clips (`re W n`, which is what every producer writes to keep a picture inside its frame) it is exact. Without it a picture drawn twice its frame's size covers the text beside it, and a page that draws its own previous revision behind a clip shows both. Absent means "nothing was cut": either there was no clip, or the clip contained the mark entirely, in which case saying so would cost a `clipPath` per mark for no visible difference.
The 2×3 affine matrix everything in a PDF is positioned by. `[a b c d e f]` stands for the matrix that takes (x, y) to (a·x + c·y + e, b·x + d·y + f). Written as a tuple rather than an object because it is created constantly — one per glyph in a long run — and because that is the order the file states it in, so no reader of this code has to hold a translation in their head.
Values
CMYK to RGB by the naive formula. `255 · (1 − c) · (1 − k)` is not what a printer does and is what every screen reader does, including Acrobat's own preview. The alternative needs the output profile, which the file does not carry and the browser will not apply.
The codecs whose output is an image, not a byte stream.
PDF, as one module. The only format here that recognises itself outright: `%PDF-` at the start of the file and nothing else can claim it, so the registry never has to load a second module to find out. It is also the one format that shares nothing with the others — no ZIP, no OPC, no DrawingML — which is why an application that shows PDFs and nothing else carries neither `@genomdev/office-core` nor a line of it.
@genomdev/pdf/view
Classes
Functions
Interfaces
Values
The viewer's stylesheet. Small on purpose: a PDF carries its own appearance completely, so anything this adds is about the *reader* — the paper's shadow, the gaps between sheets, the selection colour — and never about the page.
PDF with its renderer attached. What `@genomdev/pdf/view` exports and what an application passes to the viewer. The universal half is the same object — the renderer is a field on it, not a second registration — so a format either brings a way to draw itself or it does not, and nothing has to pair them up by hand.