Skip to content
Genom
Concepts

The content model

One vocabulary for four formats, and an address for every character in it.

Extraction produces a tree of blocks. A block is a section, a heading, a paragraph, a list, a table, an image, a chart, a diagram, an equation, a quote, a code block or a break; inside a block are inlines, which carry marks and links. That vocabulary is the same whether the file was a document, a workbook, a deck or a PDF — which is what makes markdown out of a spreadsheet mean something.

The half that knows a format lives in that format’s package. @genomdev/docx can turn its model into this tree, and nothing else can; the tree itself, the serialisers and the chunker know no format at all.

walking the tree
import { walk, inlineText } from '@genomdev/core/content';

for (const block of doc.walk()) {
  if (block.type === 'heading') console.log('#'.repeat(block.level), inlineText(block));
  if (block.type === 'table') console.log(block.rows.length, 'rows');
  if (block.type === 'chart') console.log(block.series);
}

Charts and SmartArt are data

A chart in a document is a part that stores the numbers it plots. Extracting it as the word "chart" throws away the only thing about it that a language model or an index can use, so it comes out as its series, its categories and its title. SmartArt comes out as the nesting it represents.

Four outputs of one tree

Every output can be produced with an offset map ({ withMap: true }), which is what turns a character position in the output back into a position in the document.

OutputForKeeps
toMarkdown()models, diffs, storageheadings, tables, lists, links, emphasis
toPlainText()search, indexingreading order and nothing else
toHtml()displaystructure plus data-loc addresses
toJson()pipelinesthe tree as it is