Skip to content
Genom
Tasks

Read text and markdown

Every document as a tree of blocks, and four ways to write that tree out.

Extraction produces one thing — a ContentDocument — and the outputs are views of it. Markdown for a language model or a diff, plain text for search, HTML for display, JSON for storage. Tables stay tables in all four, and a chart comes out as the numbers it plots.

Running here, in this pageNothing is uploaded
Loading the example…
import { extract } from '@genomdev/genom';

const doc = await extract(bytes);

doc.blocks;          // the tree: headings, tables, charts as data, SmartArt as nesting
doc.toMarkdown();    // GFM, with the options that make it readable
doc.toPlainText();   // nothing added
doc.toHtml();        // semantic, with addresses on the elements
doc.toJson();        // the tree, serialisable
doc.metadata;        // title, author, dates, language

When you know the format

extract from the umbrella considers every format, which means every parser is in the module graph. When the format is known, ask its package directly — it brings nothing else with it.

docx only
import { extractDocx } from '@genomdev/docx';

const doc = await extractDocx(bytes);

Options that change what comes out

OptionValuesWhat it does
images'omit' | 'alt' | 'reference' | 'dataUri'Whether an image becomes nothing, its alt text, a reference, or the bytes inline.
charts'omit' | 'title' | 'data'A chart as nothing, as its title, or as the series it plots.
revisions'accepted' | 'markup'Tracked changes as the author left them, or with insertions and deletions marked.
transformsTransform[]Clean-ups over the finished tree: repeated headers, hyphenation, empty rows.
onImage onChart onTable onMathhandlersRun over the finished tree — OCR, a vision model, a LaTeX converter.
signalAbortSignalCancels a long parse at the next block boundary.

Filling in what the file does not say

Handlers run after the parse rather than inside it, so the parse stays fast and everything slow is optional, cancellable and rate-limited. A scanned figure gets OCR; a chart nobody stored the data for gets a vision model; an equation gets a converter.

handlers.ts
const doc = await extract(bytes, {
  concurrency: 4,
  onImage: async (image, { readMedia }) => ({
    alt: await describe(await readMedia(image.source!.partName)),
  }),
});

From a terminal

@genomdev/genom
npx @genomdev/genom md report.docx
npx @genomdev/genom text *.docx --out corpus.txt
npx @genomdev/genom json deck.pptx --charts data