Skip to content
Genom

Extract a Word document as Markdown, text or blocks

The document on one side, what a pipeline reads on the other.

.docx.docOne package, every generation — @genomdev/docx

Loading the example…

The demonstration is not that Markdown comes out — every library in this space can show that. It is the loop: click a word in the extraction and it lights up in the document; click the document and the extraction scrolls to the line it came from.

That join is an address per block, and it is the same address a retrieval system can store and bring back months later. Tables survive as tables, and the list numbers are the document’s own rather than the renderer’s.

import { extractMarkdown } from '@genomdev/genom';

// Headings stay headings, tables stay tables, and the document's own
// list numbers come out as text rather than as CSS.
const markdown = await extractMarkdown(file); // handbook.docx

Put it in your own application

What is above is the published package running in this page, with nothing behind it. The documentation covers the same ground with the API in full.