Skip to content
Genom
API reference

@genomdev/genom

Read and show Word, Excel, PowerPoint and PDF files — in TypeScript, in a browser and on a server

74 exported symbols across 5 entry points

@genomdev/genom

Classes

ContentDocumentfrom @genomdev/core
class ContentDocument

An extracted document: the tree, and every way of writing it down. The serialisers are methods rather than free functions for one reason: the addresses. A caller who has the tree and reaches for a `render` from somewhere else gets text with no way back into the document, which is the thing this package exists to prevent. Everything that produces text from here can also produce the map beside it.

format
FormatId
hash
string
metadata
DocumentMetadata
blocks
readonly Block[]
annotations
readonly Annotation[]
profile
string
walk
() => Generator<Block>
Every block in reading order, parents before children.
sections
readonly Block[]
The top-level sections: slides, sheets, or the parts of a document.
toPlainText
{ (options?: TextOptions & { withMap?: false; }): string; (options: TextOptions & { withMap: true; }): { text: string; map: OffsetMap; }; }
toPlainText
{ (options?: TextOptions & { withMap?: false; }): string; (options: TextOptions & { withMap: true; }): { text: string; map: OffsetMap; }; }
toMarkdown
{ (options?: MarkdownOptions & { withMap?: false; }): string; (options: MarkdownOptions & { withMap: true; }): { markdown: string; map: OffsetMap; }; }
toMarkdown
{ (options?: MarkdownOptions & { withMap?: false; }): string; (options: MarkdownOptions & { withMap: true; }): { markdown: string; map: OffsetMap; }; }
toHtml
{ (options?: HtmlOptions & { withMap?: false; }): string; (options: HtmlOptions & { withMap: true; }): { html: string; map: OffsetMap; }; }
toHtml
{ (options?: HtmlOptions & { withMap?: false; }): string; (options: HtmlOptions & { withMap: true; }): { html: string; map: OffsetMap; }; }
toJSON
() => ContentDocumentJson
chunks
(options?: ChunkOptions) => Chunk[]
Structure-aware chunks, each carrying the selectors to find it again.
selectorsFor
(text: string, map: OffsetMap, start: number, end: number, output?: OutputKind) => Selector[]
Selectors for a range of one of this document's text outputs. The bridge for everybody who chunked the text themselves. Give it the offsets a splitter reported, the map that produced them and which output they came from, and it gives back the anchor triple: exact, portable and robust.
CorruptFileErrorfrom @genomdev/core
class CorruptFileError extends GenomError

The file was identified, but its contents violate the format specification.

offset
number | undefined
Byte offset where the violation was found, when known.
EncryptedFileErrorfrom @genomdev/core
class EncryptedFileError extends GenomError

The document is encrypted and no usable password was supplied.

GenomErrorfrom @genomdev/core
class GenomError extends Error

Base class for every error raised by Genom. A common ancestor lets consumers distinguish "the file failed to open" from a genuine bug in their own code with a single `instanceof` check.

code
string
Stable machine-readable code; unaffected by message wording changes.
UnsupportedFormatErrorfrom @genomdev/core
class UnsupportedFormatError extends GenomError

The file format was not recognised, or no plugin is registered for it.

detectedFormat
string | undefined

Functions

chunkfrom @genomdev/core
function chunk(document: ContentDocument, options?: ChunkOptions): Chunk[]
contentHashfrom @genomdev/core
function contentHash(bytes: Uint8Array): string

Eight hex digits identifying the bytes a locator was made against. FNV-1a over the whole buffer, which is not a cryptographic hash and is not meant to be. The question it answers is "is this the same file as the one the address came from", where the alternative to a wrong answer is a wrong highlight, not a security breach. Sixty-four bits folded to thirty-two make an accidental collision a once-in-four-billion event between two files a user has open at the same moment, and a real hash would cost a megabyte of bookkeeping and a dependency for it.

describeFormatfrom @genomdev/core
function describeFormat(id: FormatId): FormatDescriptor | undefined
detectFormatfrom @genomdev/core
function detectFormat(source: ByteSource): Promise<DetectionResult>

Determines the file format from its contents, name and MIME type. Signature, name and MIME type only — this is the cheap pass. For a ZIP the result is `probable` and says so: telling docx, xlsx, pptx and odt apart means reading `[Content_Types].xml`, which `probe()` does in the next step. Keeping the two apart is what lets the expensive one be skipped for the formats that announce themselves in their first five bytes.

extract
function extract(input: ByteSourceInput | GenomDocument, options?: DispatchOptions): Promise<ContentDocument>

Read a document's content. Takes anything a byte source can be made of — a path is not one of them, that is `@genomdev/genom/node` — or an already-open document, which is the path a viewer takes when it wants to search what is on screen without parsing it twice.

extractMarkdown
function extractMarkdown(input: ByteSourceInput | GenomDocument, options?: DispatchOptions): Promise<string>

The whole document as markdown, in one call.

extractText
function extractText(input: ByteSourceInput | GenomDocument, options?: DispatchOptions): Promise<string>

The whole text, in one call. The shortest path in and the one most people take first, so it does the conventional thing without being asked: no section labels, no markers, no markup.

findTextfrom @genomdev/core
function findText(document: ContentDocument, query: string, options?: { caseSensitive?: boolean; wholeWord?: boolean; limit?: number; }): { range: { start: Locator; end: Locator; }; text: string; }[]

Plain search over a document, returning anchors rather than offsets. What the viewer's find box is built on. Everything it returns is a selector set, so a hit found here highlights through exactly the same path as a chunk retrieved from a vector database — one mechanism, not two.

formatLocatorfrom @genomdev/core
function formatLocator(parts: { version?: number; hash?: string; flow: LocatorFlow; steps?: readonly LocatorStep[]; offset?: number | undefined; cell?: string | undefined; }): Locator

Builds a locator string. The inverse of {@link parseLocator} for every input it accepts; the pair is covered by a round-trip test rather than by inspection, because the grammar is small enough to be exhaustively generated and too fiddly to eyeball.

fromDocument
function fromDocument(document: GenomDocument, bytes: Uint8Array | undefined, options?: DispatchOptions): Promise<ContentDocument>

Read the content of a document that is already open. A hash is computed from the bytes when they are to hand and omitted otherwise: an address with no hash still resolves, it simply cannot warn that it came from a different file.

getLicenseKeyfrom @genomdev/core
function getLicenseKey(): string | null

The key most recently passed to {@link setLicenseKey}, or `null`. Returned verbatim and unparsed. Nothing in the library calls this yet; it exists so that an application can confirm its own bootstrap ran, and so that the eventual check has somewhere to read from.

parseLocatorfrom @genomdev/core
function parseLocator(locator: Locator): LocatorParts

Parses a locator, or throws. Throwing rather than returning `undefined` because a malformed locator is a programming error on the caller's side — locators are produced by this library, not typed by hand — and a silent `undefined` here surfaces three layers away as a highlight that does not appear.

resolveSelectorsfrom @genomdev/core
function resolveSelectors(document: ContentDocument, selectors: readonly Selector[], options?: ResolveOptions): ResolvedRange | undefined
setLicenseKeyfrom @genomdev/core
function setLicenseKey(key: string | null | undefined): void

Records the licence key for this process or page. Call it once, anywhere before the first viewer is created — the key is not a secret and belongs in your source, committed, alongside the rest of your bootstrap. Calling it again replaces the previous value; calling it with an empty string, `null` or `undefined` clears it. Nothing observable happens as a result. The library renders identically with a key, without one, and with a key that is complete nonsense.

streamBlocks
function streamBlocks(input: ByteSourceInput | GenomDocument, options?: DispatchOptions): AsyncGenerator<Block>

The blocks, one at a time. For a caller who means to process a book without holding it. The walk itself is not incremental — a document has to be parsed before its structure is known — so this is a convenience over the finished tree rather than a different pipeline, and it says so rather than pretending otherwise.

toByteSourcefrom @genomdev/core
function toByteSource(input: ByteSourceInput): Promise<ByteSource>

Normalises any supported input into a {@link ByteSource}. Strings are URLs.

toHtmlfrom @genomdev/core
function toHtml(document: ContentDocument, options?: HtmlOptions): { text: string; map: OffsetMap; }
toJsonfrom @genomdev/core
function toJson(document: ContentDocument): ContentDocumentJson
toMarkdownfrom @genomdev/core
function toMarkdown(document: ContentDocument, options?: MarkdownOptions): { text: string; map: OffsetMap; }
toPlainTextfrom @genomdev/core
function toPlainText(document: ContentDocument, options?: TextOptions): { text: string; map: OffsetMap; }

Interfaces

ByteSourcefrom @genomdev/core
interface ByteSource

A random-access source of bytes. This is the central abstraction of the project: parsers never touch `File`, `Blob` or the network directly. That makes it possible to read a ZIP central directory at the end of a file, or a PDF xref table, without pulling the whole document into memory, and to run the same parser in a browser, in Node, or on top of HTTP range requests.

byteLength
number
Total size of the source in bytes.
name?
string | undefined
File name when known. Used as a hint for format detection.
mimeType?
string | undefined
MIME type when the source reports one.
slice
(start: number, end?: number) => Promise<Uint8Array>
Reads the `[start, end)` range. Implementations must return exactly the requested number of bytes or throw {@link OutOfBoundsError }. Short reads are not allowed, otherwise every parser would have to re-check the length after each call.
dispose?
(() => void) | undefined
Releases retained resources such as caches or network connections.
Chunkfrom @genomdev/core
interface Chunk
text
string
The chunk as stored and embedded, contextualised if that was asked for.
contextLength
number
How much of `text` is the prepended heading path.
selectors
readonly Selector[]
Where to find it again: exact, portable and robust, in that order.
breadcrumbs
readonly string[]
The heading path above it.
section
string | undefined
`Slide 5`, `Sheet Budget`, or the section's label.
tokens
number
start
number
Offsets into the serialised output this came from.
end
number
ChunkOptionsfrom @genomdev/core
interface ChunkOptions

Chunking that knows what a document is. The state of the art in JavaScript is to take the text, split it every 512 characters, and hope. That destroys exactly the things retrieval depends on: a table loses its header three rows in, a heading is separated from the section it names, a sentence is cut in half, and every chunk arrives at the index with no idea where it came from. This walks the tree instead. Headings become breadcrumbs and stay with their content, a table row is never split, a table too big for one chunk is split by rows with its header repeated in each, and every chunk carries the selectors that find it again in the document it came from — which is the part nobody else has, and the reason a retrieval hit can be highlighted rather than merely quoted.

maxTokens?
number | undefined
Target size. Default 512.
overlap?
number | undefined
Overlap between neighbours, in tokens. Default 64.
tokenCounter?
((text: string) => number) | undefined
How to count. Default: characters over four. Pluggable because the right answer depends on a model this package has never heard of, and a default that shipped a tokenizer would be a megabyte of tables for something the caller can do in one line.
contextualize?
boolean | undefined
Prepend the heading path to each chunk's text. Default true. "Q3 Results › Risks › Currency exposure" in front of a paragraph that says "the position was closed in October" is the difference between a chunk that retrieves and one that does not. The context is marked in `contextLength` so a caller that wants the bare text can strip it.
tables?
"whole" | "rows" | undefined
`whole` keeps a table together; `rows` splits large ones. Default `rows`.
source?
"markdown" | "text" | undefined
Which output the chunk text comes from. Default `markdown`.
DetectionResultfrom @genomdev/core
interface DetectionResult

Result of format detection.

format
FormatId
container
ContainerKind
confidence
"certain" | "probable" | "guess"
How much the detector can be trusted. `certain` — an unambiguous signature matched; `probable` — the container was identified and the concrete format was inferred from the extension or MIME type and still needs confirmation from the contents; `guess` — no signature matched and the decision rests on the file name alone.
reason
string
What led to the decision — useful when debugging third-party files.
DispatchOptions
interface DispatchOptions extends ExtractOptions

Options, plus the formats allowed to claim the file.

formats?
readonly FormatEntry[] | undefined
Which formats may claim the file. Defaults to all four, eagerly. Pass descriptors from `@genomdev/genom/lazy` to fetch the one that matches instead of shipping all of them.
DocumentMetadatafrom @genomdev/core
interface DocumentMetadata

Metadata common to every format.

title?
string | undefined
author?
string | undefined
subject?
string | undefined
keywords?
readonly string[] | undefined
description?
string | undefined
producer?
string | undefined
The application that produced the file.
createdAt?
Date | undefined
modifiedAt?
Date | undefined
language?
string | undefined
Language of the main content, BCP 47.
custom?
Readonly<Record<string, string | number | boolean | Date>> | undefined
Format-specific fields that do not fit the common schema.
ExtractOptionsfrom @genomdev/core
interface ExtractOptions

What to do with the things that are not words. The defaults are chosen for the commonest use, which is feeding a language model, and they are not the conservative choices. `charts: 'data'` reads a chart out as the numbers it plots, because a chart is a table and leaving a hole where one was is how every other extractor in this space loses the content of a quarterly report. `headers: 'first'` drops the running head repeated on every page, because eighty repetitions of a company name is what poisons a retrieval index.

signal?
AbortSignal | undefined
concurrency?
number | undefined
How many handler calls may be in flight at once. Default 4. The handlers are the slow part — an OCR round trip is a second where everything else in this package is a microsecond — so the number that matters is this one, and it belongs to the caller who knows what is on the other end of it.
images?
"omit" | "alt" | "reference" | "dataUri" | undefined
`omit` drops images; `alt` keeps the block with its alt text and no bytes; `reference` adds the part name so the caller can fetch them; `dataUri` embeds them. Default `alt`.
charts?
"omit" | "title" | "data" | undefined
`data` reads the plotted values out. Default `data`.
diagrams?
"omit" | "nodes" | undefined
Default `nodes`: SmartArt comes out as the nesting it draws.
headers?
"omit" | "first" | "all" | undefined
Running heads and feet. Default `first`. `first` keeps one copy per section, which is where the information is; `all` keeps every repetition, which is what a fidelity-minded caller wants and a retrieval index does not.
comments?
"omit" | "annotate" | undefined
Default `annotate`: comments come out as annotations, not inline text.
notes?
"omit" | "annotate" | "inline" | undefined
Footnotes and endnotes. Default `annotate`.
revisions?
"final" | "original" | "markup" | undefined
Which side of a tracked change to take. Default `final`. `final` is the document as its last author left it; `original` is what it said before; `markup` keeps both, marked.
hiddenText?
boolean | undefined
Hidden text (`w:vanish`). Default false, which is what hiding it meant.
hiddenSheets?
boolean | undefined
Hidden worksheets. Default false: they hold the lookup tables.
formulas?
"value" | "formula" | "both" | undefined
Default `value`: the number as the sheet displays it, not the formula.
emptyCells?
boolean | undefined
Empty rows and columns inside a sheet's used range. Default false. Excel's idea of the used range is generous, and a sheet whose author once typed in `ZZ4000` reports four thousand rows of nothing.
headings?
"auto" | "style" | "outline" | "none" | HeadingRule | undefined
How a paragraph becomes a heading. Default `auto`. `auto` believes the style first, then `w:outlineLevel`, and does not guess from type size — a document whose author never used a heading style has no headings, and inventing them from font size produces a table of contents full of pull quotes.
styleMap?
readonly StyleRule[] | undefined
Map a named style onto an output element, the way mammoth does. The escape hatch for the house style nobody outside the company has heard of: `Zitat` is a blockquote, `Code Block` is code, `Untertitel` is a level two heading. Rules are tried in order and the first match wins.
onImage?
ImageHandler | undefined
onTable?
TableHandler | undefined
onChart?
ChartHandler | undefined
onMath?
MathHandler | undefined
transforms?
readonly Transform[] | undefined
Applied to the tree in order, after the walk and before serialisation.
FormatModulefrom @genomdev/core
interface FormatModule<TDocument extends GenomDocument = GenomDocument>

A format: everything one file type needs, in one value. `@genomdev/docx` exports one of these for Word, covering both generations. Nothing else has to know that `.doc` exists.

id
string
Stable identifier, used for options and events: `docx`, `pdf`.
formats
readonly FormatId[]
Every format id this module opens.
detect?
readonly DetectRule[] | undefined
How to recognise the format from the file.
open
(source: ByteSource, options?: OpenOptions) => Promise<TDocument>
canOpen?
((source: ByteSource, detection: DetectionResult) => Promise<boolean>) | undefined
Confirms that the source really is this format. Only needed where a rule cannot decide — an OLE2 file is a `.doc`, an `.xls` or a `.ppt`, and telling them apart means reading the directory.
walk?
((document: TDocument, hash: string, options: ResolvedOptions) => Promise<WalkResult>) | undefined
Turns the parsed model into the content tree. Declared as a method rather than as a property on purpose: TypeScript checks method parameters bivariantly, and without that a `FormatModule<DocxDocument>` could not be put in a list beside a `FormatModule<PdfDocument>` — which is the only thing a registry ever does with them.
GenomDocumentfrom @genomdev/core
interface GenomDocument

An opened document: the contract shared by every parser. Deliberately narrow — it only carries what is meaningful for any format. Everything else (docx sections, xlsx sheets, pptx slides) lives in subtypes inside the format packages. The viewer works against this interface so that it can still show a title and a page count for a format whose renderer is not registered.

format
FormatId
kind
DocumentKind
metadata
DocumentMetadata
pageCount?
number | undefined
Number of pages/sheets/slides, when the format exposes it cheaply. For docx this is `undefined` until the document has been laid out: splitting a text flow into pages depends on fonts, hyphenation and the printable area.
extractText
() => Promise<string>
Extracts the whole document text for search, indexing and previews. A dedicated method because this is the one operation every format needs in the same way and which requires no rendering.
dispose
() => void
Releases retained resources. Calling it twice is safe.
HtmlOptionsfrom @genomdev/core
interface HtmlOptions

Semantic HTML: the structure, not the appearance. Not a renderer, and the distinction is the whole design. `@genomdev/docx/view` reproduces what a document looks like — fonts, page boxes, measured line breaks. This produces what it *is*: headings that are headings, tables that are tables, a `<figure>` around a picture and its caption. There is no CSS and no colour, because the consumer is a language model, a search index, or a page that has its own stylesheet and does not want this one. Every element carries its address in `data-loc`, which is what makes an HTML extraction round-trip: a click in the rendered output can be turned back into a place in the document.

locators?
boolean | undefined
Write `data-loc` on every element that has an address. Default true.
figures?
boolean | undefined
`<figure>`/`<figcaption>` around images. Default true.
sections?
boolean | undefined
`<section>` per slide, sheet or document section. Default true.
wrap?
boolean | undefined
Wrap the output in `<article>`. Default false.
footnotes?
boolean | undefined
Footnotes as a `<section class="footnotes">` at the end. Default true.
withMap?
boolean | undefined
LazyFormatfrom @genomdev/core
interface LazyFormat

A format that has not been loaded yet. The descriptor carries the recognition rules, so the registry can decide whether this is the module it needs before paying for it. `load` resolves to the module itself, which in a bundler is a chunk of its own.

id
string
formats
readonly FormatId[]
detect?
readonly DetectRule[] | undefined
load
() => Promise<FormatModule>
MarkdownOptionsfrom @genomdev/core
interface MarkdownOptions

Markdown, written properly. The bar is not "produces something a renderer accepts". Every library in this space clears that. The bar is that a person reading the output can tell what the document said, and a language model reading it does not have to guess — which means the table alignment survives, the nested list stays nested, a pipe inside a cell does not end the column, a paragraph starting with `1.` does not silently become a list, and a footnote is a footnote rather than a number floating in the middle of a sentence. Every one of those is a bug this had at some point.

tables?
"html" | "gfm" | "list" | undefined
`gfm` for pipe tables, `html` for a `<table>`, `list` for one line a row. `html` is not a cop-out: a pipe table cannot express a row span, and a merged cell rendered as a pipe table is silently wrong in a way nobody notices. `list` is for tables so wide that either of the others is unreadable — a workbook of forty columns, most often.
tableColumnLimit?
number | undefined
Widest table, in columns, still worth a pipe table. Default 12.
images?
"text" | "omit" | "alt" | "link" | undefined
`alt` writes `![alt]()`; `link` writes a real path; `omit` drops them; `text` writes just the alt text with no image syntax at all.
footnotes?
boolean | undefined
`[^1]` footnotes plus a section at the end. Default true.
comments?
boolean | undefined
Comments as footnotes too, marked with the author. Default false.
frontMatter?
boolean | undefined
A YAML block of the document metadata at the top. Default false.
sections?
"none" | "heading" | "rule" | undefined
`---` between sections, and a heading naming each. Default `heading`.
bullet?
"-" | "*" | "+" | undefined
`*` or `-`. Default `-`.
nativeMarkers?
boolean | undefined
Keep the document's own list markers instead of `1.` and `-`. On by default, and it is the right default: Word's `%1.%2` patterns produce `2.3.1`, which three levels of markdown nesting cannot express, and a numbered list that restarts partway down is invisible to a counter. The marker is `derived` text and anchors to the item.
math?
boolean | undefined
`$…$` for maths. Default true where a LaTeX form is available.
marks?
boolean | undefined
Emphasis, strikethrough and the rest. Default true.
withMap?
boolean | undefined
Marksfrom @genomdev/core
interface Marks

What a run of text is wearing. A short list on purpose. Everything a word processor can do to a character is not what a reader means by emphasis, and carrying all of it would produce markdown full of `<span style>`. These are the marks that survive being written down as text.

bold?
boolean | undefined
italic?
boolean | undefined
strike?
boolean | undefined
code?
boolean | undefined
superscript?
boolean | undefined
subscript?
boolean | undefined
underline?
boolean | undefined
highlight?
string | undefined
A named or hex highlight colour, when the author marked the text.
insertion?
boolean | undefined
Set on text a revision marked as inserted or deleted.
deletion?
boolean | undefined
OpenOptionsfrom @genomdev/core
interface OpenOptions

Options shared by every parser.

signal?
AbortSignal | undefined
Aborts parsing of large files.
tolerant?
boolean | undefined
Keep parsing when the file locally violates the specification. Defaults to `true`: real files produced by office suites break the standard routinely, and failing the whole document where a single paragraph could be dropped is a bad trade.
password?
string | undefined
Password for encrypted documents.
onProgress?
((fraction: number) => void) | undefined
Progress callback, 0..1.
ResolvedRangefrom @genomdev/core
interface ResolvedRange

A resolution, and how much to trust it. `resolvedBy` is not decoration. An application that watches it can see its documents drifting — the day the exact selectors stop matching and everything falls through to quotes is the day somebody started editing the corpus — and it can see that before a user reports a highlight in the wrong place.

range
ModelRange
resolvedBy
"locator" | "position" | "quote" | "native"
confidence
number
1 for an exact match, lower for a fuzzy one.
TextOptionsfrom @genomdev/core
interface TextOptions

Plain text: the output with nothing added. The one that has to be genuinely plain. It is what goes into a search index, a diff, a `grep`, and every pipeline whose next stage is not a markdown parser — and every character of markup in it is a false hit waiting to happen. So the only characters here that are not from the document are the separators between blocks, and a caller who wants none of those can say so.

separator?
string | undefined
Between blocks. Default `\n`.
sectionBreaks?
boolean | undefined
An extra newline between sections. Default true.
sectionLabels?
boolean | undefined
Name each section (`Slide 5`) before its content. Default true.
tables?
"tab" | "align" | "lines" | undefined
How a table row is written. Default `tab`. `tab` keeps the columns machine-readable — a row is still a record — and `align` pads them so a person can read the table in a terminal. `align` costs a second pass over the table to measure it.
markers?
boolean | undefined
Include list markers. Default true.
annotations?
boolean | undefined
Append footnote and comment text at the end. Default true.
withMap?
boolean | undefined

Type aliases

Blockfrom @genomdev/core
type Block = SectionBlock | HeadingBlock | ParagraphBlock | ListBlock | ListItemBlock | TableBlock | TableRowBlock | TableCellBlock | ImageBlock | ChartBlock | DiagramBlock | MathBlock | BlockquoteBlock | CodeBlock | BreakBlock
BlockTypefrom @genomdev/core
type BlockType = 'section' | 'heading' | 'paragraph' | 'list' | 'listItem' | 'table' | 'tableRow' | 'tableCell' | 'image' /** A chart, carried as the data it plots. */ | 'chart' /** SmartArt, carried as the nesting it draws. */ | 'diagram' | 'math' | 'blockquote' | 'code' /** An explicit page or column break the author put there. */ | 'break'

What a block is.

ByteSourceInputfrom @genomdev/core
type ByteSourceInput = ByteSource | Blob | ArrayBuffer | Uint8Array | string

Everything Genom can turn into a {@link ByteSource}.

DocumentKindfrom @genomdev/core
type DocumentKind = 'text-document' /** A grid of cells: xlsx, ods, csv. */ | 'spreadsheet' /** A sequence of slides: pptx, odp. */ | 'presentation' /** Fixed page layout: pdf. */ | 'paged-document' /** A raster or vector image. */ | 'image'

Broad document category; determines which viewer applies.

FormatEntryfrom @genomdev/core
type FormatEntry = FormatModule | LazyFormat

Either an already-loaded format or a promise of one.

FormatIdfrom @genomdev/core
type FormatId = 'docx' | 'xlsx' | 'pptx' | 'doc' | 'xls' | 'ppt' | 'pdf' | 'odt' | 'ods' | 'odp' | 'txt' | 'md' | 'csv' | 'json' | 'xml' | 'html' | 'png' | 'jpeg' | 'gif' | 'webp' | 'bmp' | 'tiff' | 'svg' | 'unknown'

Identifier of a concrete file format. A string literal union rather than an enum: the values are part of the public API, get serialised to JSON, and are used as registry keys.

Inlinefrom @genomdev/core
type Inline = TextInline | LinkInline | BreakInline | ImageInline | ReferenceInline | MathInline
Locatorfrom @genomdev/core
type Locator = string

The opaque wire form. Parse it with {@link parseLocator}.

Selectorfrom @genomdev/core
type Selector = LocatorSelector | TextPositionSelector | TextQuoteSelector | SheetSelector | SlideSelector

Values

docxfrom @genomdev/docx
docx: FormatModule<DocxDocument>

Word documents, both generations, as one module. Recognition is stated as data so the registry can pick this module out of six without loading any of them: an OPC package announces its main part's content type, and that is the only thing that tells the three OOXML formats apart. The compound file is the case rules cannot settle — a `.doc`, an `.xls` and a `.ppt` share one signature — so it shortlists itself and `canOpen` reads the directory to be sure. One module rather than two because the binary reader produces the same model: nothing above here has to know which generation it was handed.

pdffrom @genomdev/pdf
pdf: FormatModule<PdfDocument>

PDF, as one module. The only format here that recognises itself outright: `%PDF-` at the start of the file and nothing else can claim it, so the registry never has to load a second module to find out. It is also the one format that shares nothing with the others — no ZIP, no OPC, no DrawingML — which is why an application that shows PDFs and nothing else carries neither `@genomdev/office-core` nor a line of it.

pptxfrom @genomdev/pptx
pptx: FormatModule<PptxDocument>

PowerPoint presentations, both generations, as one module. Recognition is stated as data so the registry can pick this module out of six without loading any of them: an OPC package announces its main part's content type, and that is the only thing that tells the three OOXML formats apart. The compound file is the case rules cannot settle — a `.doc`, an `.xls` and a `.ppt` share one signature — so it shortlists itself and `canOpen` reads the directory to be sure. One module rather than two because the binary reader produces the same model: nothing above here has to know which generation it was handed.

xlsxfrom @genomdev/xlsx
xlsx: FormatModule<XlsxDocument>

Excel workbooks, both generations, as one module. Recognition is stated as data so the registry can pick this module out of six without loading any of them: an OPC package announces its main part's content type, and that is the only thing that tells the three OOXML formats apart. The compound file is the case rules cannot settle — a `.doc`, an `.xls` and a `.ppt` share one signature — so it shortlists itself and `canOpen` reads the directory to be sure. One module rather than two because the binary reader produces the same model: nothing above here has to know which generation it was handed.

@genomdev/genom/lazy

Values

all
all: readonly LazyFormat[]

Everything, lazily. Convenient and honest about its cost: naming all four means a bundler emits all four chunks. It still loads one — but an application that will never see a presentation should list the three it will.

docx
docx: LazyFormat

Word, both generations, with its renderer.

pdf
pdf: LazyFormat

PDF, with its renderer.

pptx
pptx: LazyFormat

PowerPoint, both generations, with its renderer.

xlsx
xlsx: LazyFormat

Excel, both generations, with its renderer.

@genomdev/genom/formats

Values

docxfrom @genomdev/docx
docx: FormatModule<DocxDocument>

Word documents, both generations, as one module. Recognition is stated as data so the registry can pick this module out of six without loading any of them: an OPC package announces its main part's content type, and that is the only thing that tells the three OOXML formats apart. The compound file is the case rules cannot settle — a `.doc`, an `.xls` and a `.ppt` share one signature — so it shortlists itself and `canOpen` reads the directory to be sure. One module rather than two because the binary reader produces the same model: nothing above here has to know which generation it was handed.

pdffrom @genomdev/pdf
pdf: FormatModule<PdfDocument>

PDF, as one module. The only format here that recognises itself outright: `%PDF-` at the start of the file and nothing else can claim it, so the registry never has to load a second module to find out. It is also the one format that shares nothing with the others — no ZIP, no OPC, no DrawingML — which is why an application that shows PDFs and nothing else carries neither `@genomdev/office-core` nor a line of it.

pptxfrom @genomdev/pptx
pptx: FormatModule<PptxDocument>

PowerPoint presentations, both generations, as one module. Recognition is stated as data so the registry can pick this module out of six without loading any of them: an OPC package announces its main part's content type, and that is the only thing that tells the three OOXML formats apart. The compound file is the case rules cannot settle — a `.doc`, an `.xls` and a `.ppt` share one signature — so it shortlists itself and `canOpen` reads the directory to be sure. One module rather than two because the binary reader produces the same model: nothing above here has to know which generation it was handed.

xlsxfrom @genomdev/xlsx
xlsx: FormatModule<XlsxDocument>

Excel workbooks, both generations, as one module. Recognition is stated as data so the registry can pick this module out of six without loading any of them: an OPC package announces its main part's content type, and that is the only thing that tells the three OOXML formats apart. The compound file is the case rules cannot settle — a `.doc`, an `.xls` and a `.ppt` share one signature — so it shortlists itself and `canOpen` reads the directory to be sure. One module rather than two because the binary reader produces the same model: nothing above here has to know which generation it was handed.

@genomdev/genom/viewer

Classes

Viewerfrom @genomdev/core
class Viewer
state
ViewerState
container
HTMLElement
view
DocumentView | undefined
The live view, for reaching format-specific APIs such as page navigation.
registry
FormatRegistry
content
(options?: ExtractOptions) => Promise<ContentDocument>
The open document as content: blocks, markdown, chunks, addresses. Built with the walk the format module brought, so this needs no dispatcher and no second parse — and it is the same tree extraction produces on a server, which is why an address minted there resolves here. Memoised: the tree does not change while the document is open.
search
(options?: ExtractOptions) => Promise<DocumentSearch>
Search over the extracted text, with hits highlighted in the document.
on
<K extends keyof ViewerEventMap & string>(event: K, handler: Handler<ViewerEventMap[K]>) => () => void
Subscribes to a named event. Returns an unsubscribe function.
once
<K extends keyof ViewerEventMap & string>(event: K, handler: Handler<ViewerEventMap[K]>) => () => void
emit
<K extends keyof ViewerEventMap & string>(event: K, payload: ViewerEventMap[K]) => void
Emits an event. Renderers and plugins report through here.
subscribe
(listener: (state: ViewerState) => void) => () => void
Subscribes to state as a whole. Kept beside the event bus because a UI framework wants a store rather than a stream: `useSyncExternalStore` needs exactly this shape.
open
(input: ByteSourceInput | GenomDocument, options?: ViewerOptions) => Promise<void>
Opens a file and mounts its view. Calling it again before the previous call settles is correct: the older result is discarded. That is the normal case when a reader picks files in quick succession.
refresh
() => void
Re-renders the current view.
zoom
number
The scale the document is drawn at, and one otherwise. A renderer that cannot scale reports nothing, and a viewer over one reads as being at its natural size — which it is.
setZoom
(zoom: number) => void
Changes the scale the document is drawn at. Here rather than only on the renderer because this is the object an application holds: the React, Vue and Angular wrappers all pass their options in at mount and none of them could reach a zoom control without it, so the only way to change the scale through them was to open the file again. Nothing is laid out again — see `dom/scale.ts` for why a zoom is a transform.
pageCount
number
How many pages, sheets or slides the open document came to.
currentPage
number
The one the reader is looking at, counted from nought.
goToPage
(index: number, behavior?: ScrollBehavior) => void
Scrolls to a page, a sheet or a slide. Here for the same reason as {@link setZoom}: this is the object an application holds, and until it had this the renderers' navigation could not be reached from the React, Vue or Angular wrappers at all.
close
() => void
Closes the document and clears the container.
destroy
() => void
Releases every resource. The instance must not be used afterwards.

Functions

createViewer
function createViewer(container: HTMLElement, options?: ViewerOptions): Viewer

Creates a viewer that can open any format Genom supports.

Interfaces

ViewerOptionsfrom @genomdev/core
interface ViewerOptions extends OpenOptions
formats?
readonly FormatEntry[] | undefined
The formats this viewer can open, eager or lazy. Lazy descriptors are the interesting case: they carry the rules that recognise a file, so the right module is fetched and the others never are.
registry?
FormatRegistry | undefined
A registry built by hand, for an application that shares one.
zoom?
number | undefined
fit?
"none" | "width" | "page" | undefined
initialPage?
number | undefined
options?
Readonly<Record<string, Record<string, unknown>>> | undefined
Renderer options, addressed by format id: `{ xlsx: { formulaBar: false } }`.
decorate?
Readonly<Record<string, Decorator>> | undefined
Decorators, keyed `format:kind` — `xlsx:cell`, `docx:hyperlink`.
plugins?
readonly ViewerPlugin[] | undefined
ViewerPluginfrom @genomdev/core
interface ViewerPlugin

An extension that gets the viewer itself. Returns its own teardown.

name
string
setup
(viewer: Viewer) => (() => void) | void
ViewerStatefrom @genomdev/core
interface ViewerState
status
ViewerStatus
document
GenomDocument | undefined
detection
DetectionResult | undefined
error
Error | undefined
progress
number
Parse and layout progress, 0..1.

Type aliases

ViewerStatusfrom @genomdev/core
type ViewerStatus = 'idle' | 'loading' | 'ready' | 'error'

The viewer: bytes in, a document on the screen, and a way to reach it after. The one place that knows the whole path — recognise, load the module, open, mount. The React, Vue and Angular wrappers are adapters over this class, so their behaviour is identical by construction rather than by convention. Three things here are the API rather than the implementation. **Formats are values.** `formats: [docx, pdf]` — eager modules or lazy descriptors, mixed freely. Nothing is registered globally, nothing is discovered, and an application that shows PDFs ships no Word parser. **Options are addressed by format.** `options: { xlsx: { formulaBar: false } }` rather than one flat bag whose keys collide the moment two formats want the same word. **Everything reports through one bus.** `viewer.on('page:change', …)`, and the format-specific events carry their format in the name. Extension has two levels, and the cheap one comes first: `decorate` changes an element a renderer already produced, which survives the renderer being rewritten; `plugins` get the viewer itself, for what nobody anticipated.

@genomdev/genom/node

Functions

extractAll
function extractAll(paths: readonly string[], options?: DispatchOptions & { concurrency?: number; }): Promise<BatchResult[]>

Extracts many files with a fixed number in flight. Errors are collected rather than thrown: a corpus of ten thousand documents always contains a few that no parser can open, and a batch that stops at the first one is a batch nobody can run. The caller gets a result per input and decides what a failure means.

extractFile
function extractFile(path: string, options?: DispatchOptions): Promise<ContentDocument>

Extracts a document from a path.

hashFile
function hashFile(path: string): Promise<string>

The hash a file's addresses will carry, without parsing it.

openFile
function openFile(path: string): Promise<ByteSource>

Reads a file into a byte source, keeping its name for format detection.

Interfaces

BatchResult
interface BatchResult
path
string
document?
ContentDocument | undefined
error?
Error | undefined