@genomdev/core
Genom core: byte sources, ZIP, streaming XML, format detection, the plugin registry, addressing and the content model
390 exported symbols across 6 entry points
@genomdev/core— 200 exports@genomdev/core/content— 94 exports@genomdev/core/dom— 58 exports@genomdev/core/dom/highlight— 20 exports@genomdev/core/viewer— 15 exports@genomdev/core/testing— 3 exports
@genomdev/core
Classes
A source backed by a browser `Blob`/`File`, read lazily in chunks.
A synchronous cursor-based reader over a `Uint8Array`. Binary formats (ZIP headers, OLE2/CFB, PDF xref tables, TIFF IFDs) are read as sequences of fixed-width fields, and tracking the offset by hand in every parser is a reliable way to introduce bugs. The reader does it for us and bounds-checks every step.
The operation was aborted through an AbortSignal.
The file was identified, but its contents violate the format specification.
The document is encrypted and no usable password was supplied.
Base class for every error raised by Genom. A common ancestor lets consumers distinguish "the file failed to open" from a genuine bug in their own code with a single `instanceof` check.
A source backed by HTTP range requests. Lets a document be opened from a URL without downloading it in full: a few kilobytes of tail data are enough to list the contents of a 200 MB file. If the server does not support ranges, the source downloads the file once and serves subsequent reads from memory.
A source backed by a buffer already held in memory.
A format feature that has not been implemented yet. A dedicated type lets the viewer show "this part of the document is not supported yet" instead of a generic read failure.
A read ran past the end of the byte source.
RC4 as a keystream generator: the cipher is symmetric, so one direction.
The file format was not recognised, or no plugin is registered for it.
Markup the reader walked past that nothing has declared it may walk past. Only ever raised in strict mode, which no viewer turns on: a reader that stops at the first unknown attribute is useless against real files, where every generator writes something nobody has seen. What it is for is the opposite situation — a development run over a corpus, where an element the parser silently ignores is indistinguishable from one it handles, and a gap therefore survives for as long as nobody happens to look at the right page. Strict mode makes the ignoring explicit: everything the parser passes over must be named in the registry of markup we have decided draws nothing, with the reason. Anything else stops the parse and names itself.
A streaming (pull) XML parser. This is the foundation of the whole OOXML layer and the reason large documents stay fast. Building a full node tree for a 40 MB `document.xml` costs several hundred megabytes and a second of allocation before any useful work starts. A pull parser lets the consumer walk the file once and build only the domain model it actually needs, allocating nothing per element it chooses to skip. Design notes that matter for performance: - Element and namespace names are interned. A document contains millions of `w:t`/`w:r`/`w:p` tags but only a few dozen distinct names, so interning turns name comparison into pointer comparison and removes almost all string allocation. - Attributes are parsed lazily. Most elements are visited without their attributes ever being read, so they are only materialised on demand. - Text is decoded lazily. Whitespace-only text between tags is extremely common and is skipped without ever becoming a JavaScript string. The parser deliberately supports no DTD or external entities: office files never use them, and processing them is a well-known vulnerability class (XXE).
Random-access ZIP archive reader. Works on top of a {@link ByteSource} rather than an in-memory buffer: listing the contents of a 100 MB xlsx only requires reading a few kilobytes of central directory at the end of the file. Individual entries are inflated on demand, which is essential for large workbooks where most of the weight sits in one or two sheets that may never be opened.
Functions
Reads `hhea` and `hmtx` into a table of advances. `hmtx` is not a plain array: it holds `numberOfHMetrics` pairs of advance and left side bearing, and then *bearings alone* for every glyph after that, all of which share the last advance. That tail is how a font of ten thousand fixed-width glyphs costs four bytes each instead of eight, and a reader that misses it returns nothing for most of a CJK font. Nothing is returned where the font does not say: a CFF-based OpenType file still carries `hhea` and `hmtx`, but a file missing either has no horizontal metrics to give and a caller must measure instead of guessing.
The advances of a family, or nothing where none were written.
AES-CBC. By default the first sixteen bytes of `data` are the initialisation vector, which is how PDF stores it. Padding is PKCS#7 and is stripped unless asked otherwise — a padding byte out of range is treated as no padding rather than as an error, because a stream whose last block is damaged is still a stream.
AES-CBC encryption with an explicit IV and no padding.
AES-ECB, one block at a time and no padding.
Attribute value. Unprefixed attributes are looked up with an empty namespace.
Boolean attribute in the OOXML sense. In the `ST_OnOff` schema `1`, `true` and `on` mean true; the absence of the attribute on a flag element (such as `<w:b/>`) also means true.
Numeric attribute value; `undefined` when absent or not a number.
The offsets at which a line may end, for a caller that wants a list. The end of the text is not among them: every line ends there and no caller has to be told so.
Whether this runtime can deflate at all; see the note above.
First direct child with the given name.
Direct child elements; text nodes are dropped.
All direct children with the given name.
The classical constants, in the units of an em of `unitsPerEm`.
Orders two locators the way the document reads. Needed wherever a range has to be normalised — a user selects backwards about half the time — and wherever highlights are merged. Flows are ordered by their kind and index so that a body address always precedes a footnote's, which is arbitrary but stable, and stable is the property that matters.
The parsed form of a language's patterns, parsed once per module.
Eight hex digits identifying the bytes a locator was made against. FNV-1a over the whole buffer, which is not a cryptographic hash and is not meant to be. The question it answers is "is this the same file as the one the address came from", where the alternative to a wrong answer is a wrong highlight, not a security breach. Sixty-four bits folded to thirty-two make an accidental collision a once-in-four-billion event between two files a user has open at the same moment, and a real hash would cost a megabyte of bookkeeping and a dependency for it.
Picks a readable text colour for a given background. Needed wherever a format specifies only a fill: a table header with a dark shade and default black text would be unreadable. The 0.5 threshold is on WCAG relative luminance.
The CRC-32 of a run of bytes, as an unsigned 32-bit number.
The cut of a family a run wants. A missing cut falls back to the regular rather than to nothing: a family that files one file for all four is written down once, and the browser synthesises the other three from it — which changes the shape of the letters and not their advances, so the regular's widths are the right answer for all four.
Turns the encoded ranges into something that can be asked about a character. Kept per table rather than per call: a document asks this for every character of every line, several times over as lines are tried and abandoned, and decoding a face on each of those would cost more than the layout.
CP1251: Cyrillic text in legacy Microsoft Office files.
Expands predefined entities and numeric character references.
Latin-1 (ISO-8859-1): used for legacy ZIP entry names and PDF strings.
UTF-16LE: the native string encoding of OLE2/CFB and many Windows structures.
Decodes an XML part, honouring the byte order mark it may start with. Nearly every OOXML part is UTF-8, and this exists for the ones that are not: XML permits UTF-16, a writer occasionally uses it, and Excel opens such a file without comment. Decoding those bytes as UTF-8 yields a string of NULs with no root element — a whole workbook lost to two bytes at the front.
Zlib-wrapped DEFLATE: the same stream with two bytes of header and a checksum.
All descendants at any depth with the given name.
Determines the file format from its contents, name and MIME type. Signature, name and MIME type only — this is the cheap pass. For a ZIP the result is `probable` and says so: telling docx, xlsx, pptx and odt apart means reading `[Content_Types].xml`, which `probe()` does in the next step. Keeping the two apart is what lets the expensive one be skipped for the formats that announce themselves in their first five bytes.
Encodes 8-bit RGBA pixels, top row first, as a PNG.
Escapes an attribute value: text, plus the quote that delimits it.
Escapes text content. `&` and `<` are required; `>` is escaped because Word escapes it, and a comparison against Word's own output should not turn on that. Characters XML cannot express at all — the C0 controls other than tab, newline and return — are dropped rather than written, because writing them produces a part no parser will read back, this one included. Tested by scanning rather than by a regular expression: the fast path runs over every character of every run of a document, and a scan that stops at the first character needing work is both quicker and legible.
First descendant at any depth with the given name.
Looks up a format by extension. Accepts `docx`, `.docx` or a whole file name.
Builds a locator string. The inverse of {@link parseLocator} for every input it accepts; the pair is covered by a round-trip test rather than by inspection, because the grammar is small enough to be exhaustively generated and too fiddly to eyeball.
The shape each code point takes, in the string's own order. `undefined` where the character has no shape to take — a space, a digit, a vowel mark — so that a caller can tell "the character as written" from "the letter, isolated", which reach the same glyph but are not the same question. The rule reads a letter's two neighbours *through* the transparent marks between: a beh followed by a fatha and then a reh is medial, because as far as the join is concerned the fatha is not there.
The key most recently passed to {@link setLicenseKey}, or `null`. Returned verbatim and unparsed. Nothing in the library calls this yet; it exists so that an application can confirm its own bootstrap ran, and so that the eventual check has somewhere to read from.
Whether this project has written the advances of a family down.
Whether a language tag has patterns at all; asked before anything is loaded.
The hyphenation of one document: its languages, resolved. Asked for every `w:lang` the document states — Word hyphenates each run by its own language, and a Ukrainian paper whose body is English is hyphenated as English. A tag nobody wrote patterns for is simply absent, and the word comes out whole.
The places a line may end inside this word, counted in characters kept. `hyphenationPointsOf('Silbentrennung', german)` is `[3, 6, 9]` — `Sil-`, `Silben-`, `Silbentren-` — and the caller decides which of them the column has room for. Nothing is returned for a word the minima leave no room in, so a caller can ask of every word and pay a map lookup for the short ones.
Zlib-wrapped DEFLATE: the same compression with two bytes of header. Needed because the two are not interchangeable and nothing says which is which. A ZIP entry is raw; the picture inside a Word document says "DEFLATE" in its header and is a zlib stream — a fact discoverable only by looking at the bytes, where `78 01` at the front is the giveaway. Feeding one to the other's decoder fails immediately rather than producing rubbish, which is the one mercy of it.
Raw DEFLATE with no wrapper, decompressed synchronously.
DEFLATE with or without its zlib wrapper, whichever the bytes turn out to be. Deliberately not two functions. Files lie about this constantly — a PDF stream marked `/FlateDecode` is supposed to be zlib-wrapped and a good number are raw, and the header check is two bytes — so the tolerant reading is the only one worth having at a call site that has just been handed a document.
True when the string is shaped like a locator. Does not validate the path.
Interprets an `ST_OnOff` value.
Whether the bytes begin with a zlib header rather than raw DEFLATE.
The joining type of one code point. The two zero-width controls and the bidirectional marks live outside the Arabic block and are answered before it is searched: `ZWJ` joins where no letter does, `ZWNJ` refuses to, and the left-to-right and right-to-left marks are invisible to the join in the way a vowel mark is. A document that sets Persian writes them by the hundred, and reading one as a letter would break every word it stands in.
Whether a string holds a letter whose shape depends on its neighbours.
Converts a {@link Length} to a CSS string, or `undefined` when it is `auto`.
Converts a {@link Length} to pixels; percentages need a reference size.
The break action after every UTF-16 code unit of `text`. The last unit is always {@link BREAK_MANDATORY}: LB3 says a line ends at the end of the text, and a caller that wants to join two runs must therefore ask about them together rather than concatenate two answers.
The Line_Break class of a code point, before LB1 resolves it.
The flow a locator addresses, without parsing the rest of it.
The path steps of a locator.
The same locator with a different character offset.
The name at an index of the standard Macintosh ordering, if it has one. Exported for the other reader of this list: an unembedded composite font whose CIDs are glyph indices into a font nobody has. The core Microsoft faces are all built in this order, so the ordering is the only statement there is about what those indices meant — see `unicodeOf`.
Applies one rule to what the probe found out.
The best any of a module's rules could do on this file.
Reads the constants out of a `MATH` table. Nothing is returned for a font without one, which is nearly every font: a caller wanting to set an equation in such a face should fall back to {@link classicalMathConstants}, not refuse.
MD5, sixteen bytes out.
Normalises text for quote matching and records where every character came from. The map back is the whole point, and the reason this is not three chained `replace` calls: a match found in normalised space has to become a range in the document, and every transformation that changes a length has to be accounted for as it happens. Unicode normalisation is applied per cluster rather than to the finished string. Running NFC over the assembled result would silently shorten the text out from under the offsets just recorded — the map would be right for every document without a diacritic and wrong for every document with one, which is the worst failure mode on offer. Per *character* would be no better in the other direction: NFC composes a base and its combining mark into one character, and a character examined alone has nothing to compose with, so decomposed text would stay decomposed and a quote stored precomposed would never match it. So a base character and the marks that follow it are normalised together, and every character that comes out is mapped to where the base came from.
Parses `ST_HexColor`: six or eight hex digits, with or without a leading hash. The special value `auto` means "the application picks the colour", so it returns `undefined` and lets the caller apply its own contextual rule (usually black text on a light background).
Parses a locator, or throws. Throwing rather than returning `undefined` because a malformed locator is a programming error on the caller's side — locators are produced by this library, not typed by hand — and a silent `undefined` here surfaces three layers away as a highlight that does not appear.
Reads a font program's table directory, and the tables worth reading. Takes the bytes of a `.ttf`, `.otf` or `.ttc`, and the first font of a collection where it is one. Nothing is returned for a file whose sfnt version is not one of the four in use — which is the check that keeps a `.pfb`, a bitmap font or a truncated download from being read as if the numbers at its head meant offsets. Every bounds check here is load-bearing rather than defensive: a font embedded in a document is arbitrary bytes from a stranger, and half of them have been through a subsetter.
Parses XML into a node tree. Built on top of {@link XmlPullParser} so there is exactly one lexer in the codebase. The tree form is the right tool for the small configuration parts of an OOXML package — `styles.xml`, `numbering.xml`, `theme1.xml`, `.rels` — which are read in full, revisited repeatedly, and small enough that the convenience of random access outweighs the allocation cost. For `document.xml`, worksheets and slides, use the pull parser directly: those parts are read once, sequentially, and can be several tens of megabytes. The options are the pull parser's own — in practice the namespace canonicalisation rule, which a tree parse needs for exactly the reason a streaming one does. Without it a Strict package reads into a tree whose every element is in a namespace no constant names, and every lookup against it silently returns nothing.
The patterns of a language tag, or nothing where none are written down.
The string with every letter that has a presentation form replaced by it, and lam–alef by its ligature; everything else as it was. `covers` is asked of each replacement, and a face that has not got one keeps the letter as written — a measurement in the wrong shape is still nearer than one of a missing glyph.
Formats a value as a CSS point string.
Formats a value as a CSS pixel string.
One buffer through RC4 with a fresh key schedule.
WCAG 2.1 relative luminance, 0..1.
Records the licence key for this process or page. Call it once, anywhere before the first viewer is created — the key is not a secret and belongs in your source, committed, alongside the rest of your bootstrap. Calling it again replaces the previous value; calling it with an empty string, `null` or `undefined` clears it. Nothing observable happens as a result. The library renders identically with a key, without one, and with a key that is complete nonsense.
Reads the substitutions of the joining features out of a font's `GSUB`. Nothing is returned for a font without the table, which is most Latin ones and every bitmap font; an empty map for a font that has it and states none of these features. Both mean the same to a caller — measure what `cmap` reaches — and are distinguished only so that "asked and there is none" can be cached.
All text in the subtree, concatenated in document order.
Throws {@link CancelledError} if the signal has already been aborted.
Normalises any supported input into a {@link ByteSource}. Strings are URLs.
Strips the trailing NUL padding of a fixed-length string field.
Depth-first traversal of the subtree, including the element itself.
Interfaces
The cuts of a family: regular, bold, italic, bold italic. A cut is absent when it resolved to the same file as one already written — a family filing one file for all four asks the browser to synthesise the other three, and synthesis does not change an advance.
One cut of a family, encoded. `em` is the design grid the rest is stated in; `a`, `d` and `g` are the ascent, the descent (positive) and the line gap as fractions of it; `n` is the advance of a code point the face does not cover, in font units. The advances themselves are ranges, encoded the way `line-break-classes.ts` encodes Unicode's: `s` holds the first code point of each range as a base-36 delta from the range before it, `l` how many code points the range covers, `v` the advance in font units — all three in the same order. The length is not redundant with the next range's start, and the first version of this file left it out and was wrong. A range ends where its run of *covered* code points ends, which is not where the next one begins: the gap between them is code points the face has no glyph for at all. Without `l` every gap read as "covered, at the width of the character before it", and a face asked for a character it does not have would answer with a width instead of admitting it cannot draw it — which is the signal Word uses to substitute.
How wide each glyph of a font is, in units of the em. The whole of a string's width, for a face that is neither kerned nor ligated — which is Word's default, and which the renderer's own stylesheet already asks the browser for with `font-variant-ligatures: none`. Word kerns only above the point size `w:kern` names, and the corpus almost never names one; `w:ligatures` is rarer still. So a run's width is the sum of its characters' advances, and the advances can be read once and reused. That is what makes computing a layout cheaper than measuring one. `canvas.measureText` costs a call per distinct string; this costs a table per font and an addition per character, and it answers on a machine with no browser and no such font installed. ## What it is not Not shaping. Arabic, the Indic scripts, Thai and Hebrew reorder and substitute glyphs before any width is decided, and a sum of per-character advances is simply wrong for them. A caller that may meet such text has to ask something that shapes; this is for the scripts where a character is a glyph.
A random-access source of bytes. This is the central abstraction of the project: parsers never touch `File`, `Blob` or the network directly. That makes it possible to read a ZIP central directory at the end of a file, or a PDF xref table, without pulling the whole document into memory, and to run the same parser in a browser, in Node, or on top of HTTP range requests.
Result of format detection.
Metadata common to every format.
What a decoded table can answer. Deliberately the same shape as the useful half of `AdvanceTable` in `font-file.ts`: whoever consumes widths should not have to know which of the two sources answered, and the two are interchangeable for everything except the kerned width.
Type metrics for the families a document is most likely to ask for and a browser least likely to have, as fractions of the em. Generated by `node tools/fonts/generate.mjs` from the fonts themselves — `head` for the design grid, `OS/2`/`hhea` for the vertical extents, and the average advance over a sample of English text. Do not edit by hand. What they are for: a viewer cannot draw a font it has not got, and the substitute the browser picks is not the one the authoring application picked. A third of the presentation corpus's line mass is set in a family the browser cannot resolve — Aptos above all, because Office keeps its cloud fonts in a private directory it never registers — and where a font is substituted the lines wrap **later** than PowerPoint's on 472 slides against 280. The stand-in is narrower, so a line that should have broken carries one more word, and from there neither side has the same lines at all. With these an `@font-face` can make a font the machine *does* have take the missing one's measurements: `size-adjust` for the width of the letters, `ascent-override` and `descent-override` for the height of the line.
Format description: its name, how to open it, how to recognise it.
A format: everything one file type needs, in one value. `@genomdev/docx` exports one of these for Word, covering both generations. Nothing else has to know that `.doc` exists.
An opened document: the contract shared by every parser. Deliberately narrow — it only carries what is meaningful for any format. Everything else (docx sections, xlsx sheets, pptx slides) lives in subtypes inside the format packages. The viewer works against this interface so that it can still show a title and a page count for a format whose renderer is not registered.
What the engine holds: the languages it may hyphenate, and the answer.
A language's hyphenation patterns, as the generated modules write them. The strings are the pattern file's own words, space separated: a pattern is letters with digits between them (`a1bc2d`), and an exception is a word with hyphens at the points it may break (`as-so-ciate`). Kept as text rather than as a parsed structure because that is what a module can hold without paying for it at load time — see `compileHyphenation`, which parses once, lazily.
The same, parsed, which is what the algorithm walks.
Kerning pairs, in units of the em.
A format that has not been loaded yet. The descriptor carries the recognition rules, so the registry can decide whether this is the module it needs before paying for it. `load` resolves to the module itself, which in a bundler is a chunk of its own.
A length as stored in the file, together with the unit it was stored in. Keeping the unit lets the renderer decide how to emit it: some measurements are better expressed in `pt` so the browser can round them itself, others must be pixels because they take part in layout arithmetic.
Exact: a range between two locators. Valid while the file's bytes are.
One step down the tree. `kind` is a single letter so the whole path stays short — a locator is stored per chunk, and a corpus of a million chunks pays for every character. The letters are mnemonic rather than clever: `b` block, `t` table, `w` row (`r` was taken), `c` cell, `p` paragraph, `r` run, `s` shape, `i` inline object.
The constants, in font units. Named as the specification names them, because every one of them is quoted by that name in the rules that use it and a friendlier name would only need translating back. Divide by the em to use them.
A resolved range in the document model.
Text prepared for matching, with a way back to the original offsets. Quote matching has to ignore differences no reader would call a difference: a non-breaking space against a space, a soft hyphen left over from justification, the zero-width joiners a copy-paste through a word processor leaves behind, and the two spellings Unicode allows for any accented letter. Normalising all of that changes the offsets, so the map back is built at the same time — without it a match in normalised space cannot be turned into a range in the document.
Options shared by every parser.
An entry copied from another archive, exactly as that archive stored it. Every header field the source stated is carried, not only the content. The flags, the two version words, the DOS timestamp and the extra field mean nothing to a reader of an OOXML package — and they are bytes of the local header, so an entry rewritten without them is an entry that differs from the one it was copied from. "Open it and save it gives you your file back" is a claim about bytes, and it is only true if the copy is one.
What resolving a file produced, and how sure the registry is.
A resolution, and how much to trust it. `resolvedBy` is not decoration. An application that watches it can see its documents drifting — the day the exact selectors stop matching and everything falls through to quotes is the day somebody started editing the corpus — and it can see that before a user reports a highlight in the wrong place.
Colour arithmetic, in the form every format needs it. A colour as a number, as CSS, as HSL, and as the luminance that decides whether text on it should be black or white. Nothing here belongs to a particular format — the transforms DrawingML defines on top of this (`tint`, `shade`, `lumMod`) live in `@genomdev/office-core`, because they are elements of a schema rather than facts about colour.
Native: the address an Excel user already knows.
Native: a slide, and optionally one shape on it.
The substitutions one feature asks for, by the glyph they replace.
Portable: offsets into extracted text. For everyone who chunked the text with something else. LangChain and its cousins keep a start and an end and nothing else, and this is the selector that lets those chunks come back. `profile` is a hash of the extraction options that produced the offsets. Markdown with GFM tables and markdown with HTML tables are different strings, and without the profile the difference would show up as a highlight that is forty characters off rather than as an error.
Robust: the text with enough of its neighbours to be unambiguous. `prefix` and `suffix` are what separate the fourteenth "Total" in a workbook from the fifteenth. Thirty-two characters each is the figure the annotation community converged on: enough to disambiguate ordinary prose, short enough that storing it per chunk is free next to the chunk itself.
One font program, read. Everything a caller can learn about a font without rasterising it: how many glyphs it has, how big its em is, which character reaches which glyph, what the glyphs are called — and the raw bytes of every table, so that a question this does not answer can still be asked of the file rather than of a second copy of it.
An XML attribute with its namespace resolved.
Kind of the event the pull parser is currently positioned on. Numeric rather than string-valued, so the hot dispatch loop in the document parser compares integers. A plain enum rather than a `const enum` because the values cross package boundaries, which ambient const enums cannot do under `verbatimModuleSyntax`.
XML, written. The counterpart of `pull.ts`, and deliberately as small. It knows nothing about namespaces beyond the fact that a name may carry a prefix: OOXML declares its prefixes once on the root element and uses them everywhere below, so a writer that resolved namespaces per element would be solving a problem the format does not have — and would break the one thing that makes lossless saving possible, which is splicing a run of the *original* characters back into the output. Those characters carry the document's own prefixes. Declare the document's own prefixes on the root and they are valid wherever they land; invent new ones and every preserved fragment is broken. Two properties everything else rests on: - **It appends and never revisits.** The output is a list of strings joined once at the end, so writing a forty-megabyte part is linear rather than quadratic. - **It writes no whitespace of its own.** Word writes `document.xml` as one line, and so does this. Indentation inside `w:t` is content — a space between two runs is a space in the document — and a pretty-printer that cannot tell the difference welds words together or pulls them apart.
Options controlling how the archive caches inflated entries.
One entry of the ZIP central directory.
Type aliases
Everything Genom can turn into a {@link ByteSource}.
Container kind: what can be determined from the first bytes of a file. This intermediate layer exists because many formats share one signature: docx, xlsx, pptx, odt and epub all start with `PK\x03\x04`. The core detector identifies the container, and the package that knows how to read that container narrows it down to a specific format.
How to recognise a format without running its code. Rules are tried in order and the first match wins. A rule that names bytes or a content type answers outright; a rule that names only a container narrows the field and leaves the decision to `canOpen`.
Broad document category; determines which viewer applies.
Either an already-loaded format or a promise of one.
Identifier of a concrete file format. A string literal union rather than an enum: the values are part of the public API, get serialised to JSON, and are used as registry keys.
The shape a letter takes, named as OpenType names its features.
What a character does to the letters beside it.
The opaque wire form. Parse it with {@link parseLocator}.
The independent content streams of a document. A document is not one sequence of blocks. Headers, footers, footnotes, endnotes, comments and speaker notes are separate flows that interleave with the body only once it is laid out, and addressing them as if they were part of the body would make every address after the first footnote wrong.
How well a rule matched: enough to decide, or only enough to shortlist.
Every substitution a font states, by the feature that asks for it.
Values
A line **must** end here: a hard break, or the end of the text.
A line may end here.
No break may be taken here.
Every family a table is written for, in the case the file names them.
Catalogue of known formats. It also lists formats that have no parser yet: the catalogue answers "what is this file", not "can we open it". That lets the viewer say "this is a PowerPoint 97-2003 presentation, support is planned" instead of a bare "unknown format".
Whether a code point is East Asian by width: `ea ∈ {F, W, H}`.
Whether a code point is final punctuation, `\p{Pf}`.
Whether a code point is initial punctuation, `\p{Pi}`.
Whether a code point is `Extended_Pictographic`.
Whether a code point is reserved for an emoji Unicode has not assigned.
The Line_Break classes, in the order `VALUES` indexes them.
A hyphenation that never breaks a word, for a document that asks for none.
CSS pixels per inch — 96, as is conventional on the web.
Enums
@genomdev/core/content
Classes
An extracted document: the tree, and every way of writing it down. The serialisers are methods rather than free functions for one reason: the addresses. A caller who has the tree and reaches for a `render` from somewhere else gets text with no way back into the document, which is the thing this package exists to prevent. Everything that produces text from here can also produce the map beside it.
The finished map: a sorted list of segments and two ways to search it. Sorted and searched by bisection rather than kept in a hash: the queries are range queries, a thousand-page document produces a few hundred thousand segments, and an index that answers "which segment contains offset 41 322" has to be ordered anyway.
Builds output and its map at the same time. One pass, not two. Every serialiser writes through this, so the map costs a counter and an array push per fragment rather than a second traversal — and, more to the point, it cannot fall out of step with the text, because there is no second traversal to fall out of step with.
Functions
The slow pass, run separately. Handlers are the extension point the whole design was arranged around: OCR for a scanned figure, a vision model for a chart nobody stored the data for, a describer for a photograph, a LaTeX converter for an equation. Every one of them is a network round trip, and every one is optional. So they run here, over the finished tree, and not inside the parse. The parse stays synchronous and fast; this pass can be skipped entirely, retried, rate-limited, or cancelled halfway without leaving a half-parsed document behind. A hook inside the parser would have made the whole pipeline asynchronous to accommodate the OCR of one logo.
Applies transforms in order.
Turns a run of numbers into the label a list marker shows. Presentations number their bullets by scheme rather than by value: the file says `arabicPeriod` and the position in the list is everything else. The schemes are named the same way across DrawingML, so this is shared.
An address inside a run, restated as an offset into its block. The translation the viewer needs and cannot do for itself. Extraction addresses a character as "run three of block twelve, seven characters in", because that is what survives a relayout. The page has no runs to speak of — the renderer emits one span per *formatting*, not per source run, and merges neighbours that look alike — so the only thing both halves can count is characters from the start of the block. The extracted document knows the text of every run in order, so the sum is exact. Doing it here rather than in the viewer also means the two never have to agree about what a run is.
Both ends of a resolved range, as offsets into their blocks.
The text of a block and everything under it.
Builds the content document from a document that is already open. The hash is computed from the bytes where they are to hand and omitted where they are not: an address without one still resolves, it simply cannot warn that it came from a different file.
The children of a block, whatever it calls them.
Drops rows and cells that hold nothing, which a generous used range invents.
Drops the running head and foot after their first appearance. A header repeated on every page of an eighty-page report contributes eighty identical fragments, and identical fragments are the worst thing that can happen to a nearest-neighbour search: they crowd out the answer with copies of the company name. Keeping the first occurrence keeps the information. Matched on normalised text so that a header carrying a page number — which is most of them — still counts as the same header.
Removes a table of contents. A generated contents page is a list of every heading in the document with a page number after it, which means it matches almost every query about the document and answers none of them. Recognised by shape — many short lines, most ending in a number, most of whose text appears again as a heading later — rather than by the field that produced it, because half of them are typed by hand.
The whole path for one format: bytes or an open document, into content. What `extractDocx` and its three neighbours call. A document that arrives open passes through and is not closed — it belongs to the caller; one opened here is closed here.
Plain search over a document, returning anchors rather than offsets. What the viewer's find box is built on. Everything it returns is a selector set, so a hit found here highlights through exactly the same path as a chunk retrieved from a vector database — one mechanism, not two.
Whether a paragraph is a heading, and how deep. `auto` believes the style first and the outline level second. Not type size: a document whose author never used a heading style genuinely has no headings, and guessing them from how large the text is produces a table of contents made of pull quotes and drop caps. Wrong structure is worse than none — chunking follows headings, so an invented one splits a passage in half.
The inline content of a block, when it has any.
The text of a run of inlines, with nothing added.
Tells an open document from something a document has yet to be opened from.
Rejoins words a line break split with a hyphen. `manage-\nment` is one word that no search will find and no tokenizer will recognise. It comes from justified text in a two-column layout, which is most academic and legal PDF-adjacent material. Conservative on purpose: only a lowercase letter, a hyphen at the end of a line, and a lowercase letter after it. `well-\nknown` is a real hyphen and joining it would be an error, so a word that appears elsewhere in the document with its hyphen intact is left alone.
A short hash of everything that changes the output. This is what a `TextPosition` selector carries, and the reason it can be trusted. Offsets into markdown with GFM tables and offsets into markdown with HTML tables are different numbers for the same passage; without a fingerprint of the settings that produced them, the difference arrives as a highlight forty characters off rather than as an error anybody can act on. Handlers are folded in as *whether they were present*, not as what they are. A function has no stable identity to hash, and "an image handler ran" is the part that changes the text.
The output a profile was measured against, when it says.
The fingerprint a `TextPosition` carries. The options alone are not enough, and finding that out was expensive: character 89 of the markdown and character 89 of the plain text of the same document with the same settings are different characters in different paragraphs. A selector that recorded only the options resolved against whichever output the reader happened to build, and produced a confident answer pointing somewhere else entirely.
Reads a media part out of a document, where the format has one. Duck-typed rather than a field of the interface: not every format has media, and the image handler needs a way to reach the bytes where they exist.
The anchor triple for a range: exact, portable, robust. All three, always, and in this order. They cost a few hundred bytes together and they fail at different times — the exact one the moment somebody saves the file, the portable one the moment the extraction options change, and the robust one only when the words themselves go. Storing one of them is a decision to lose the anchor on a day nobody will connect to the cause.
The style rule that claims this paragraph, if any.
Every block in the tree, in reading order, parents before children. A generator because callers stop early far more often than they finish — finding the first heading, locating a block by its address — and building the flat list first would undo the point of walking at all.
Every inline in a tree of inlines, in reading order.
Interfaces
Something written outside the flow, attached to a place in it. The `anchor` is where the marker sits, so a reader following a footnote knows which sentence it belongs to and a highlight can land on that sentence rather than on the note.
A hard line break inside a paragraph.
A chart, as what it plots. Every other extractor in this space leaves a hole where a chart was, because a chart looks like a picture and pictures are hard. It is not a picture: it is a table with a title, stored as a table, and reading it out is the single largest thing this package does that the alternatives do not.
Chunking that knows what a document is. The state of the art in JavaScript is to take the text, split it every 512 characters, and hope. That destroys exactly the things retrieval depends on: a table loses its header three rows in, a heading is separated from the section it names, a sentence is cut in half, and every chunk arrives at the index with no idea where it came from. This walks the tree instead. Headings become breadcrumbs and stay with their content, a table row is never split, a table too big for one chunk is split by rows with its header repeated in each, and every chunk carries the selectors that find it again in the document it came from — which is the part nobody else has, and the reason a retrieval hit can be highlighted rather than merely quoted.
The tree, as data. A stable, versioned schema, and that is the point of it existing at all: the tree is a public contract the moment somebody stores one, and a contract with no version number is a contract nobody can migrate. `schema` is bumped when the shape changes in a way a reader would notice. `JSON.stringify(doc)` on the class would produce private fields, method-less blocks with `undefined` scattered through them, and no version — which is exactly the sort of accident that becomes a file format.
SmartArt, as the nesting it draws.
What to do with the things that are not words. The defaults are chosen for the commonest use, which is feeding a language model, and they are not the conservative choices. `charts: 'data'` reads a chart out as the numbers it plots, because a chart is a table and leaving a hole where one was is how every other extractor in this space loses the content of a quarterly report. `headers: 'first'` drops the running head repeated on every page, because eighty repetitions of a company name is what poisons a retrieval index.
Semantic HTML: the structure, not the appearance. Not a renderer, and the distinction is the whole design. `@genomdev/docx/view` reproduces what a document looks like — fonts, page boxes, measured line breaks. This produces what it *is*: headings that are headings, tables that are tables, a `<figure>` around a picture and its caption. There is no CSS and no colour, because the consumer is a language model, a search index, or a page that has its own stylesheet and does not want this one. Every element carries its address in `data-loc`, which is what makes an HTML extraction round-trip: a click in the rendered output can be turned back into a place in the document.
An image that flows with the text rather than standing on its own.
Markdown, written properly. The bar is not "produces something a renderer accepts". Every library in this space clears that. The bar is that a person reading the output can tell what the document said, and a language model reading it does not have to guess — which means the table alignment survives, the nested list stays nested, a pipe inside a cell does not end the column, a paragraph starting with `1.` does not silently become a list, and a footnote is a footnote rather than a number floating in the middle of a sentence. Every one of those is a bug this had at some point.
What a run of text is wearing. A short list on purpose. Everything a word processor can do to a character is not what a reader means by emphasis, and carrying all of it would produce markdown full of `<span style>`. These are the marks that survive being written down as text.
A reference to something written elsewhere. Footnotes, endnotes and comments are separate flows, and putting their text where the marker is would corrupt both the reading order and every offset after it. What goes inline is the marker; the text is an annotation.
Everything filled in.
Turning a stored anchor back into a place in a document. The selectors are tried in order of precision and the first that succeeds wins. Which one that was is reported, and it is not decoration: an application that watches `resolvedBy` can see its corpus drifting — the day the exact addresses stop matching and everything falls through to quotes is the day somebody started editing the documents — and can see it before a user reports a highlight in the wrong place.
A container for the format's own idea of a page. The three formats disagree about what a page is and only one of them is right in the way a reader means. A slide is a page. A worksheet is a page in the sense that matters (it is the unit you navigate to) and not in the sense that prints. A Word section is neither: pages inside it fall wherever the fonts on the machine put them, which is why extraction refuses to number them — see `pageHint`.
A document's text with the addresses it maps to, built once and reused. Resolving one quote means scanning the whole text of the document. Resolving a hundred citations from a chat answer means doing it a hundred times, so the normalised form and its map back are built once and kept.
Plain text: the output with nothing added. The one that has to be genuinely plain. It is what goes into a search index, a diff, a `grep`, and every pipeline whose next stage is not a markdown parser — and every character of markup in it is a false hit waiting to happen. So the only characters here that are not from the document are the separators between blocks, and a caller who wants none of those can say so.
What a walk returns: the block tree, and the annotations found on the way.
Type aliases
What a block is.
Decides the heading level of a paragraph, or that it is not one.
What a handler is given, and what it may say back. Handlers run over the finished tree rather than during the parse. That is a deliberate trade: the parse stays synchronous and fast, and everything slow is a separate pass that can be skipped, retried, rate-limited or cancelled without touching it. A hook inside the parser would have made the whole pipeline asynchronous to accommodate the OCR of one logo. Returning nothing leaves the block alone. Returning a block replaces it — which is how a caller turns an image into a table, or a table into prose.
Which serialisation a set of offsets was measured against.
Where every character of the output came from. Addresses are not written into the text. The text stays text — no markers, no spans, nothing that would land in an embedding, in a prompt, or in the bill for the tokens. And nothing that would survive anyway: the caller is going to chunk this with their own splitter, and a splitter throws away whatever markup it does not understand, taking the addresses with it. So the addresses travel beside the text, in this. Markdown is what makes it interesting. Markdown *adds* characters — `## `, `**`, `|`, a row of dashes — that exist in no document, and a map that did not know the difference would put every highlight out by the width of the syntax in front of it. Hence three categories: source characters from a text node of the document derived text computed from the model: a list number, a footnote marker syntax markdown the serialiser invented The same distinction the renderer makes with `data-synthetic` when it counts characters in the DOM, running the other way. One idea, both halves of the system.
A pass over the tree. Returning `undefined` for a block drops it. The three shipped with the package are the three problems every retrieval pipeline hits in its first week, and finding out about them a week in is expensive.
One format's walk: its model into the block tree.
@genomdev/core/dom
Classes
Base implementation of {@link DocumentView}. Handles what every renderer needs identically: the root element, the resize subscription, zoom, and correct resource cleanup. Subclasses only implement {@link renderContent}.
Accumulates CSS rules and installs them as a single stylesheet. Renderers must emit shared CSS classes rather than inline `style` attributes. The difference is not cosmetic: a document with 50 000 runs produces 50 000 inline style attributes, each of which the browser parses separately and none of which can be shared. Routing the same formatting through a handful of generated classes cuts both the DOM size and the style recalculation cost by an order of magnitude. The builder also deduplicates: identical declaration blocks collapse onto one class, which is exactly what happens in real documents where a few dozen distinct formatting combinations cover the entire text.
Functions
Listens for the zoom gestures on an element; the returned function stops. Registered with `passive: false` throughout, because every one of these has to be able to prevent the browser's own answer to the same gesture — and a listener registered passively cannot.
Removes every child of a node.
Picks a readable text colour for a given background. Needed wherever a format specifies only a fill: a table header with a dark shade and default black text would be unreadable. The 0.5 threshold is on WCAG relative luminance.
Creates an SVG element; needed for VML shapes and drawing fallbacks.
Builds a complete CSS rule from a selector and a style map.
Escapes a string so it can be used as a CSS class name. Word style identifiers may contain spaces, dots and non-ASCII characters; all of them have to be neutralised before they become part of a selector.
Escapes a string for safe insertion into HTML markup.
Whether a module brought a renderer with it.
Adds a stylesheet to the document exactly once. A page may host several viewers while the renderer stylesheet is shared; the key prevents a duplicate `<style>` on every mount.
Declares stand-ins for whichever of `families` the browser cannot resolve. Idempotent and additive: the rules accumulate in one stylesheet per document, so a viewer may call this once per slide without re-measuring what it has already answered. A family the table does not know is left alone — a wrong correction is worse than none, and the table is only as wide as the fonts whose metrics have actually been read.
Converts a {@link Length} to a CSS string, or `undefined` when it is `auto`.
Converts a {@link Length} to pixels; percentages need a reference size.
Subscribes to container size changes. Returns an unsubscribe function. When `ResizeObserver` is unavailable (older environments, server rendering) no subscription is created and the caller is responsible for triggering re-layout itself.
Parses `ST_HexColor`: six or eight hex digits, with or without a leading hash. The special value `auto` means "the application picks the colour", so it returns `undefined` and lets the caller apply its own contextual rule (usually black text on a light background).
Formats a value as a CSS point string.
Formats a value as a CSS pixel string.
WCAG 2.1 relative luminance, 0..1.
Re-scales a frame already in the document. The whole reason the model is worth having: changing the zoom is two style writes and no layout at all. Under the multiply-everything arrangement it was a full re-render — for Word, a re-pagination of the entire document on every step of the zoom control.
A box that draws at natural size and occupies the scaled one.
The element that actually scrolls above this one. A viewer is mounted into a container the application styled, and whether that container scrolls or one of its ancestors does is the application's choice rather than the viewer's. Anything that has to keep a point still while the document changes size — which is every zoom — needs the answer, and guessing wrong means the scroll offsets are written to an element that has none.
Serialises a style map into a CSS rule body. Used by the stylesheet generator, which emits real CSS rules instead of inline styles: one rule shared by ten thousand paragraphs is dramatically cheaper for the browser than ten thousand inline `style` attributes.
Interfaces
A live view of a document mounted into the DOM.
Thin DOM helpers. Renderers create thousands of elements per document, and calling `document.createElement` followed by one-by-one style assignment is the single biggest source of noise in that kind of code.
Type metrics for the families a document is most likely to ask for and a browser least likely to have, as fractions of the em. Generated by `node tools/fonts/generate.mjs` from the fonts themselves — `head` for the design grid, `OS/2`/`hhea` for the vertical extents, and the average advance over a sample of English text. Do not edit by hand. What they are for: a viewer cannot draw a font it has not got, and the substitute the browser picks is not the one the authoring application picked. A third of the presentation corpus's line mass is set in a family the browser cannot resolve — Aptos above all, because Office keeps its cloud fonts in a private directory it never registers — and where a font is substituted the lines wrap **later** than PowerPoint's on 472 slides against 280. The stand-in is narrower, so a line that should have broken carries one more word, and from there neither side has the same lines at all. With these an `@font-face` can make a font the machine *does* have take the missing one's measurements: `size-adjust` for the width of the letters, `ascent-override` and `descent-override` for the height of the line.
A font as far as measurement is concerned.
A length as stored in the file, together with the unit it was stored in. Keeping the unit lets the renderer decide how to emit it: some measurements are better expressed in `pt` so the browser can round them itself, others must be pixels because they take part in layout arithmetic.
Colour arithmetic, in the form every format needs it. A colour as a number, as CSS, as HSL, and as the luminance that decides whether text on it should be black or white. Nothing here belongs to a particular format — the transforms DrawingML defines on top of this (`tint`, `shade`, `lumMod`) live in `@genomdev/office-core`, because they are elements of a schema rather than facts about colour.
Mounts a parsed document into an element.
A format module that can also draw. What `@genomdev/docx/view` exports.
What a gesture asks for.
Type aliases
Changes the element a renderer just produced. Called *after* the default rendering, which is what makes it a stable contract: replacing a renderer breaks whenever its internals move, while adding a class or an attribute to a finished element does not. Returning an element replaces the original for callers who really do need that.
Values
What a zoom is allowed to be. Beyond this the browser stops being useful.
CSS pixels per inch — 96, as is conventional on the web.
@genomdev/core/dom/highlight
Classes
Paints ranges through the Custom Highlight API. One registry entry per highlight name rather than per range: the API takes a set of ranges, and a document with four hundred search hits should be four hundred ranges in one entry, not four hundred entries. The rules that colour them are injected once per name.
Paints ranges as a layer of positioned rectangles. The layer is per container rather than per document, and positioned against it, so a highlight moves with the page it is on and disappears with it. That is what makes this survive virtualisation without any bookkeeping: the rectangles belong to the same element the content does.
Functions
The block part of an address, without its character offset.
A DOM range from a pair of block offsets. Returns `undefined` when either end is not on the page. Half a range is not a usable answer: a highlight drawn from the start of a mounted block to the end of the document is worse than no highlight, because it looks deliberate.
The element a block was rendered into. Returns `undefined` when the block is not in the DOM, which under virtualisation is the normal case rather than an error: only the pages near the viewport exist. The registry deals with that by re-applying highlights when a page mounts — see `registry.ts`.
Every element a block was rendered into: more than one when a page split it.
The character offset an address carries, or zero.
A DOM position from a character offset inside a block. Walks the block's text in document order, skipping anything marked synthetic, until the count reaches the offset. An offset past the end of the block clamps to its last position rather than failing: a highlight whose end is one character beyond the text — which happens whenever a range ends on a paragraph boundary — should still be a highlight. `prefer` decides what to do on a boundary, where two DOM positions describe the same character index: the end of one text node and the start of the next. They are equivalent to the DOM and not to the eye. A range that *starts* at the end of a node produces a zero-width rectangle at the end of that line whenever the text wraps there — a stray tick in the margin — and a range that *ends* at the start of the next node produces the same thing at the start of the following line. So a start prefers the later position and an end the earlier, which is also what a browser does with its own selection.
True when the host can paint ranges without touching the DOM.
The text nodes of a block that came from the document. A `TreeWalker` rather than `textContent`, because the synthetic ones have to be left out and `textContent` cannot leave anything out. `FILTER_REJECT` on the element prunes the whole subtree, which is what a marker is.
The document text of a block, as the page holds it.
Interfaces
Highlights, as state rather than as an operation. This is the design decision virtualisation forces, and it is not a detail. Only the pages near the viewport exist in the DOM — that is what lets a thousand-page document open as quickly as a five-page one — so "highlight this passage" cannot mean "find it and paint it". The passage is usually not there yet. So a highlight is a registered intention. The registry keeps it, paints what is currently on the page, and repaints when the page changes: a scroll, a zoom, a mount. Add a highlight on page four hundred of an unopened document and nothing visible happens until page four hundred arrives, at which point it is already there. The other thing that falls out of this: a highlight is a *list of fragments*, never a rectangle. A paragraph split between page three and page four gives two groups of rectangles on two pages, and an API that promised one would have to be rewritten the first time somebody highlighted a long quotation.
Putting a highlight on the page without touching the page. The obvious implementation — wrap the range in a `<mark>` — is the wrong one here, and not marginally. This viewer paginates by measuring the content it actually rendered, so any element inserted into the flow changes where the pages break. A highlight would move the text it is highlighting, and a search across a long document would repaginate it on every hit. So nothing is inserted. Two techniques, in order of preference: The CSS Custom Highlight API takes `Range` objects and paints them from the highlight registry with no DOM involvement at all — designed for exactly this and available in every current engine. It is the primary path. An absolutely positioned layer of rectangles is the fallback, and is also the only option for the things a text highlight cannot express: a box around a picture or a merged cell, a highlight that has to be clickable, one that carries a label. The rectangles come from `Range.getClientRects()` rather than from arithmetic, because the browser has already solved bidirectional text, ligatures and line wrapping, and doing it again by hand would get all three wrong.
Values
The default look, for a host that would rather not write any. Injected by the caller rather than by the registry: a viewer embedded in an application has a design system, and a package that put a stylesheet in the head without being asked would be fighting it. The colours are stated as custom properties so overriding one does not mean copying the rule. `::highlight()` accepts a short list of properties — colour, background, decoration and shadow — and nothing that affects layout, which is the point of it: a highlight cannot move the text it marks.
Finding a place in the page from an address in the document. The last link of the chain, and the one where the two halves of the system have to agree exactly: Selector[] → ModelRange → **DOM Range** → rectangles → a highlight The renderer stamps `data-loc` on block-level elements only — a document with 200 000 runs would otherwise carry 200 000 attributes, which is the weight this viewer avoids by emitting shared CSS classes instead of inline styles. So a character inside a block is found by walking the block's text nodes and counting, which is cheap and is only ever done for the handful of blocks a highlight actually touches. The counting has one rule, and it is the whole reason this works: text the renderer computed rather than read is skipped. List numbers are put on the page as ordinary text (deliberately — so they are correct under virtualisation, correct in print, selectable and searchable), and so are footnote markers and tab leaders. All of it is real on the page and belongs to no run in the file. Counting it would put every highlight after the first numbered list a few characters late.
@genomdev/core/viewer
Classes
A minimal typed emitter. Handlers are copied before dispatch: a handler that unsubscribes itself — which is what "once" looks like, and what a React effect does on unmount — must not shorten the list being walked.
Functions
Builds the content of a document already open, using its own walk. The walk comes from the format module, which is why this needs no dispatcher and no knowledge of any format: the module that opened the document also knows how to read it out. That is the whole reason a format is one object.
Builds the search index for a document already on screen. The bytes are wanted but not required. With them the addresses carry a hash and a stored anchor from a different file is refused rather than resolved to whatever sits at that path; without them everything still works and simply cannot warn.
Creates a viewer mounted into the given element.
Interfaces
Search and highlighting, over one mechanism rather than two. The temptation is to write a find box that walks the DOM looking for a string, which every viewer has, and which is wrong here for three reasons: it cannot see the pages that are not mounted, it finds the list numbers and page numbers the renderer computed, and what it produces is a DOM node rather than an address — so a hit cannot be stored, sent anywhere, or found again after a relayout. Instead the search runs over the extracted text, which the extractor can produce in the browser because it depends on nothing that is not there. A hit is a range of that text, which is an address, which is a highlight. The same path a chunk retrieved from a vector database takes. So `find` and `showChunk` are the same operation with different inputs, and neither had to be built twice.
An extension that gets the viewer itself. Returns its own teardown.
Type aliases
@genomdev/core/testing
Functions
Interfaces
A ZIP archive builder for tests. Simpler and more reliable than keeping binary fixtures in the repository: a test states the compression method, the archive comment and the set of parts itself, and therefore exercises exactly the reader branch it was written for.