The content model
One vocabulary for four formats, and an address for every character in it.
Extraction produces a tree of blocks. A block is a section, a heading, a paragraph, a list, a table, an image, a chart, a diagram, an equation, a quote, a code block or a break; inside a block are inlines, which carry marks and links. That vocabulary is the same whether the file was a document, a workbook, a deck or a PDF — which is what makes markdown out of a spreadsheet mean something.
The half that knows a format lives in that format’s package. @genomdev/docx can turn its model into this tree, and nothing else can; the tree itself, the serialisers and the chunker know no format at all.
import { walk, inlineText } from '@genomdev/core/content';
for (const block of doc.walk()) {
if (block.type === 'heading') console.log('#'.repeat(block.level), inlineText(block));
if (block.type === 'table') console.log(block.rows.length, 'rows');
if (block.type === 'chart') console.log(block.series);
}Charts and SmartArt are data
A chart in a document is a part that stores the numbers it plots. Extracting it as the word "chart" throws away the only thing about it that a language model or an index can use, so it comes out as its series, its categories and its title. SmartArt comes out as the nesting it represents.
Four outputs of one tree
Every output can be produced with an offset map ({ withMap: true }), which is what turns a character position in the output back into a position in the document.
| Output | For | Keeps |
|---|---|---|
toMarkdown() | models, diffs, storage | headings, tables, lists, links, emphasis |
toPlainText() | search, indexing | reading order and nothing else |
toHtml() | display | structure plus data-loc addresses |
toJson() | pipelines | the tree as it is |