Read text and markdown
Every document as a tree of blocks, and four ways to write that tree out.
Extraction produces one thing — a ContentDocument — and the outputs are views of it. Markdown for a language model or a diff, plain text for search, HTML for display, JSON for storage. Tables stay tables in all four, and a chart comes out as the numbers it plots.
import { extract } from '@genomdev/genom';
const doc = await extract(bytes);
doc.blocks; // the tree: headings, tables, charts as data, SmartArt as nesting
doc.toMarkdown(); // GFM, with the options that make it readable
doc.toPlainText(); // nothing added
doc.toHtml(); // semantic, with addresses on the elements
doc.toJson(); // the tree, serialisable
doc.metadata; // title, author, dates, languageWhen you know the format
extract from the umbrella considers every format, which means every parser is in the module graph. When the format is known, ask its package directly — it brings nothing else with it.
import { extractDocx } from '@genomdev/docx';
const doc = await extractDocx(bytes);Options that change what comes out
| Option | Values | What it does |
|---|---|---|
images | 'omit' | 'alt' | 'reference' | 'dataUri' | Whether an image becomes nothing, its alt text, a reference, or the bytes inline. |
charts | 'omit' | 'title' | 'data' | A chart as nothing, as its title, or as the series it plots. |
revisions | 'accepted' | 'markup' | Tracked changes as the author left them, or with insertions and deletions marked. |
transforms | Transform[] | Clean-ups over the finished tree: repeated headers, hyphenation, empty rows. |
onImage onChart onTable onMath | handlers | Run over the finished tree — OCR, a vision model, a LaTeX converter. |
signal | AbortSignal | Cancels a long parse at the next block boundary. |
Filling in what the file does not say
Handlers run after the parse rather than inside it, so the parse stays fast and everything slow is optional, cancellable and rate-limited. A scanned figure gets OCR; a chart nobody stored the data for gets a vision model; an equation gets a converter.
const doc = await extract(bytes, {
concurrency: 4,
onImage: async (image, { readMedia }) => ({
alt: await describe(await readMedia(image.source!.partName)),
}),
});From a terminal
npx @genomdev/genom md report.docx
npx @genomdev/genom text *.docx --out corpus.txt
npx @genomdev/genom json deck.pptx --charts data