Skip to content
Genom
Tasks

Parse and inspect

When blocks are not enough: the model as the format actually describes it.

Extraction flattens four formats into one vocabulary, which is what makes it useful and what makes it lossy. When you need the six-layer style cascade, the sheet’s frozen panes or the slide’s layout chain, open the document and work with its own model.

Nothing here needs a DOM, so all of it runs on a server.

import { openDocx } from '@genomdev/docx';
import { MemoryByteSource } from '@genomdev/core';

const document = await openDocx(new MemoryByteSource(bytes));

document.body;        // the block tree as WordprocessingML describes it
document.styles;      // the full style table, resolved on demand
document.numbering;   // list definitions with their overrides
document.theme;       // colours and font slots
await document.extractText();

Both generations, one model

A .doc is not converted to a .docx and then read: the binary reader produces the same model directly. So the code above does not change, and document.format is the only place the difference survives.

openDocx and its neighbours read the modern format. The format module — docx, xlsx, pptx — decides from the container which reader to use, which is what the viewer and extract go through.

either generation
import { docx } from '@genomdev/docx';

// Chooses the reader from the container, not from the extension:
// a .docx that is really a .doc opens rather than failing.
const document = await docx.open(source);

Formulas

The Excel formula language — lexer, parser, evaluator and function library — is its own entry point. A formula can be parsed and evaluated without opening a workbook, which is what conditional formatting rules, data validations and defined names need.

formula.ts
import { parseFormula, Evaluator } from '@genomdev/office-core/formula';

const ast = parseFormula('SUM(A1:A10) * 1.2');