@genomdev/genom
Read and show Word, Excel, PowerPoint and PDF files — in TypeScript, in a browser and on a server
74 exported symbols across 5 entry points
@genomdev/genom— 54 exports@genomdev/genom/lazy— 5 exports@genomdev/genom/formats— 4 exports@genomdev/genom/viewer— 6 exports@genomdev/genom/node— 5 exports
@genomdev/genom
Classes
An extracted document: the tree, and every way of writing it down. The serialisers are methods rather than free functions for one reason: the addresses. A caller who has the tree and reaches for a `render` from somewhere else gets text with no way back into the document, which is the thing this package exists to prevent. Everything that produces text from here can also produce the map beside it.
The file was identified, but its contents violate the format specification.
The document is encrypted and no usable password was supplied.
Base class for every error raised by Genom. A common ancestor lets consumers distinguish "the file failed to open" from a genuine bug in their own code with a single `instanceof` check.
The file format was not recognised, or no plugin is registered for it.
Functions
Eight hex digits identifying the bytes a locator was made against. FNV-1a over the whole buffer, which is not a cryptographic hash and is not meant to be. The question it answers is "is this the same file as the one the address came from", where the alternative to a wrong answer is a wrong highlight, not a security breach. Sixty-four bits folded to thirty-two make an accidental collision a once-in-four-billion event between two files a user has open at the same moment, and a real hash would cost a megabyte of bookkeeping and a dependency for it.
Determines the file format from its contents, name and MIME type. Signature, name and MIME type only — this is the cheap pass. For a ZIP the result is `probable` and says so: telling docx, xlsx, pptx and odt apart means reading `[Content_Types].xml`, which `probe()` does in the next step. Keeping the two apart is what lets the expensive one be skipped for the formats that announce themselves in their first five bytes.
Read a document's content. Takes anything a byte source can be made of — a path is not one of them, that is `@genomdev/genom/node` — or an already-open document, which is the path a viewer takes when it wants to search what is on screen without parsing it twice.
The whole document as markdown, in one call.
The whole text, in one call. The shortest path in and the one most people take first, so it does the conventional thing without being asked: no section labels, no markers, no markup.
Plain search over a document, returning anchors rather than offsets. What the viewer's find box is built on. Everything it returns is a selector set, so a hit found here highlights through exactly the same path as a chunk retrieved from a vector database — one mechanism, not two.
Builds a locator string. The inverse of {@link parseLocator} for every input it accepts; the pair is covered by a round-trip test rather than by inspection, because the grammar is small enough to be exhaustively generated and too fiddly to eyeball.
Read the content of a document that is already open. A hash is computed from the bytes when they are to hand and omitted otherwise: an address with no hash still resolves, it simply cannot warn that it came from a different file.
The key most recently passed to {@link setLicenseKey}, or `null`. Returned verbatim and unparsed. Nothing in the library calls this yet; it exists so that an application can confirm its own bootstrap ran, and so that the eventual check has somewhere to read from.
Parses a locator, or throws. Throwing rather than returning `undefined` because a malformed locator is a programming error on the caller's side — locators are produced by this library, not typed by hand — and a silent `undefined` here surfaces three layers away as a highlight that does not appear.
Records the licence key for this process or page. Call it once, anywhere before the first viewer is created — the key is not a secret and belongs in your source, committed, alongside the rest of your bootstrap. Calling it again replaces the previous value; calling it with an empty string, `null` or `undefined` clears it. Nothing observable happens as a result. The library renders identically with a key, without one, and with a key that is complete nonsense.
The blocks, one at a time. For a caller who means to process a book without holding it. The walk itself is not incremental — a document has to be parsed before its structure is known — so this is a convenience over the finished tree rather than a different pipeline, and it says so rather than pretending otherwise.
Normalises any supported input into a {@link ByteSource}. Strings are URLs.
Interfaces
A random-access source of bytes. This is the central abstraction of the project: parsers never touch `File`, `Blob` or the network directly. That makes it possible to read a ZIP central directory at the end of a file, or a PDF xref table, without pulling the whole document into memory, and to run the same parser in a browser, in Node, or on top of HTTP range requests.
Chunking that knows what a document is. The state of the art in JavaScript is to take the text, split it every 512 characters, and hope. That destroys exactly the things retrieval depends on: a table loses its header three rows in, a heading is separated from the section it names, a sentence is cut in half, and every chunk arrives at the index with no idea where it came from. This walks the tree instead. Headings become breadcrumbs and stay with their content, a table row is never split, a table too big for one chunk is split by rows with its header repeated in each, and every chunk carries the selectors that find it again in the document it came from — which is the part nobody else has, and the reason a retrieval hit can be highlighted rather than merely quoted.
Result of format detection.
Options, plus the formats allowed to claim the file.
Metadata common to every format.
What to do with the things that are not words. The defaults are chosen for the commonest use, which is feeding a language model, and they are not the conservative choices. `charts: 'data'` reads a chart out as the numbers it plots, because a chart is a table and leaving a hole where one was is how every other extractor in this space loses the content of a quarterly report. `headers: 'first'` drops the running head repeated on every page, because eighty repetitions of a company name is what poisons a retrieval index.
A format: everything one file type needs, in one value. `@genomdev/docx` exports one of these for Word, covering both generations. Nothing else has to know that `.doc` exists.
An opened document: the contract shared by every parser. Deliberately narrow — it only carries what is meaningful for any format. Everything else (docx sections, xlsx sheets, pptx slides) lives in subtypes inside the format packages. The viewer works against this interface so that it can still show a title and a page count for a format whose renderer is not registered.
Semantic HTML: the structure, not the appearance. Not a renderer, and the distinction is the whole design. `@genomdev/docx/view` reproduces what a document looks like — fonts, page boxes, measured line breaks. This produces what it *is*: headings that are headings, tables that are tables, a `<figure>` around a picture and its caption. There is no CSS and no colour, because the consumer is a language model, a search index, or a page that has its own stylesheet and does not want this one. Every element carries its address in `data-loc`, which is what makes an HTML extraction round-trip: a click in the rendered output can be turned back into a place in the document.
A format that has not been loaded yet. The descriptor carries the recognition rules, so the registry can decide whether this is the module it needs before paying for it. `load` resolves to the module itself, which in a bundler is a chunk of its own.
Markdown, written properly. The bar is not "produces something a renderer accepts". Every library in this space clears that. The bar is that a person reading the output can tell what the document said, and a language model reading it does not have to guess — which means the table alignment survives, the nested list stays nested, a pipe inside a cell does not end the column, a paragraph starting with `1.` does not silently become a list, and a footnote is a footnote rather than a number floating in the middle of a sentence. Every one of those is a bug this had at some point.
What a run of text is wearing. A short list on purpose. Everything a word processor can do to a character is not what a reader means by emphasis, and carrying all of it would produce markdown full of `<span style>`. These are the marks that survive being written down as text.
Options shared by every parser.
A resolution, and how much to trust it. `resolvedBy` is not decoration. An application that watches it can see its documents drifting — the day the exact selectors stop matching and everything falls through to quotes is the day somebody started editing the corpus — and it can see that before a user reports a highlight in the wrong place.
Plain text: the output with nothing added. The one that has to be genuinely plain. It is what goes into a search index, a diff, a `grep`, and every pipeline whose next stage is not a markdown parser — and every character of markup in it is a false hit waiting to happen. So the only characters here that are not from the document are the separators between blocks, and a caller who wants none of those can say so.
Type aliases
What a block is.
Everything Genom can turn into a {@link ByteSource}.
Broad document category; determines which viewer applies.
Either an already-loaded format or a promise of one.
Identifier of a concrete file format. A string literal union rather than an enum: the values are part of the public API, get serialised to JSON, and are used as registry keys.
The opaque wire form. Parse it with {@link parseLocator}.
Values
Word documents, both generations, as one module. Recognition is stated as data so the registry can pick this module out of six without loading any of them: an OPC package announces its main part's content type, and that is the only thing that tells the three OOXML formats apart. The compound file is the case rules cannot settle — a `.doc`, an `.xls` and a `.ppt` share one signature — so it shortlists itself and `canOpen` reads the directory to be sure. One module rather than two because the binary reader produces the same model: nothing above here has to know which generation it was handed.
PDF, as one module. The only format here that recognises itself outright: `%PDF-` at the start of the file and nothing else can claim it, so the registry never has to load a second module to find out. It is also the one format that shares nothing with the others — no ZIP, no OPC, no DrawingML — which is why an application that shows PDFs and nothing else carries neither `@genomdev/office-core` nor a line of it.
PowerPoint presentations, both generations, as one module. Recognition is stated as data so the registry can pick this module out of six without loading any of them: an OPC package announces its main part's content type, and that is the only thing that tells the three OOXML formats apart. The compound file is the case rules cannot settle — a `.doc`, an `.xls` and a `.ppt` share one signature — so it shortlists itself and `canOpen` reads the directory to be sure. One module rather than two because the binary reader produces the same model: nothing above here has to know which generation it was handed.
Excel workbooks, both generations, as one module. Recognition is stated as data so the registry can pick this module out of six without loading any of them: an OPC package announces its main part's content type, and that is the only thing that tells the three OOXML formats apart. The compound file is the case rules cannot settle — a `.doc`, an `.xls` and a `.ppt` share one signature — so it shortlists itself and `canOpen` reads the directory to be sure. One module rather than two because the binary reader produces the same model: nothing above here has to know which generation it was handed.
@genomdev/genom/lazy
Values
Everything, lazily. Convenient and honest about its cost: naming all four means a bundler emits all four chunks. It still loads one — but an application that will never see a presentation should list the three it will.
Word, both generations, with its renderer.
PDF, with its renderer.
PowerPoint, both generations, with its renderer.
Excel, both generations, with its renderer.
@genomdev/genom/formats
Values
Word documents, both generations, as one module. Recognition is stated as data so the registry can pick this module out of six without loading any of them: an OPC package announces its main part's content type, and that is the only thing that tells the three OOXML formats apart. The compound file is the case rules cannot settle — a `.doc`, an `.xls` and a `.ppt` share one signature — so it shortlists itself and `canOpen` reads the directory to be sure. One module rather than two because the binary reader produces the same model: nothing above here has to know which generation it was handed.
PDF, as one module. The only format here that recognises itself outright: `%PDF-` at the start of the file and nothing else can claim it, so the registry never has to load a second module to find out. It is also the one format that shares nothing with the others — no ZIP, no OPC, no DrawingML — which is why an application that shows PDFs and nothing else carries neither `@genomdev/office-core` nor a line of it.
PowerPoint presentations, both generations, as one module. Recognition is stated as data so the registry can pick this module out of six without loading any of them: an OPC package announces its main part's content type, and that is the only thing that tells the three OOXML formats apart. The compound file is the case rules cannot settle — a `.doc`, an `.xls` and a `.ppt` share one signature — so it shortlists itself and `canOpen` reads the directory to be sure. One module rather than two because the binary reader produces the same model: nothing above here has to know which generation it was handed.
Excel workbooks, both generations, as one module. Recognition is stated as data so the registry can pick this module out of six without loading any of them: an OPC package announces its main part's content type, and that is the only thing that tells the three OOXML formats apart. The compound file is the case rules cannot settle — a `.doc`, an `.xls` and a `.ppt` share one signature — so it shortlists itself and `canOpen` reads the directory to be sure. One module rather than two because the binary reader produces the same model: nothing above here has to know which generation it was handed.
@genomdev/genom/viewer
Classes
Functions
Creates a viewer that can open any format Genom supports.
Interfaces
An extension that gets the viewer itself. Returns its own teardown.
Type aliases
The viewer: bytes in, a document on the screen, and a way to reach it after. The one place that knows the whole path — recognise, load the module, open, mount. The React, Vue and Angular wrappers are adapters over this class, so their behaviour is identical by construction rather than by convention. Three things here are the API rather than the implementation. **Formats are values.** `formats: [docx, pdf]` — eager modules or lazy descriptors, mixed freely. Nothing is registered globally, nothing is discovered, and an application that shows PDFs ships no Word parser. **Options are addressed by format.** `options: { xlsx: { formulaBar: false } }` rather than one flat bag whose keys collide the moment two formats want the same word. **Everything reports through one bus.** `viewer.on('page:change', …)`, and the format-specific events carry their format in the name. Extension has two levels, and the cheap one comes first: `decorate` changes an element a renderer already produced, which survives the renderer being rewritten; `plugins` get the viewer itself, for what nobody anticipated.
@genomdev/genom/node
Functions
Extracts many files with a fixed number in flight. Errors are collected rather than thrown: a corpus of ten thousand documents always contains a few that no parser can open, and a batch that stops at the first one is a batch nobody can run. The caller gets a result per input and decides what a failure means.
Extracts a document from a path.
The hash a file's addresses will carry, without parsing it.
Reads a file into a byte source, keeping its name for format detection.