Every document. One parse. No server.
Word, Excel, PowerPoint and PDF (the .doc, .xls and .ppt of the nineties included), rendered in the browser and read out as Markdown, text and structured blocks. One TypeScript library, no server, no dependencies. Every passage keeps an address, so what your RAG pipeline retrieves can be scrolled to and highlighted in the document it came from.
- Zero dependencies
- TypeScript, no WebAssembly, no service to run
- Nothing uploaded
- Parsed and rendered in the tab. The file never leaves it
- Markdown, text, blocks
- Extracted in JavaScript, headings and tables intact
- Built for RAG
- Every chunk has an address: scroll to it, highlight it
Four families, both generations of each
One package per family, covering the modern file and the one from the nineties that people still email. The old formats are read directly, not converted first.
Everything it does, running — one page per product, with the code beside each pane.
The document never leaves the machine it is on
Uploading a document to be converted is a disclosure. For a law firm it touches privilege, for a clinic it is PHI, for a bank it is a processor to name in a filing. Doing the work in the tab removes the question instead of answering it.
The file never leaves
No upload, no temporary copy on somebody else’s disk, nothing to delete afterwards. Open the network tab while you try the demo above — it stays empty.
There is no service to run
Nothing to deploy, scale, patch or pay for per document. It is a dependency in your package.json, and it works the same on a laptop with no network at all.
It is your page, not an iframe
The document is DOM in your document. Style it, put things next to it, scroll it with the rest of the page, and read the text out of it.
Put it on the screen, or read it with code
Both come out of the same parse, so what you show and what you index cannot drift apart. Every chunk carries an address that survives a vector database and comes back able to scroll to the passage it names.
Show it to somebody
A viewer that draws the document in the page — selectable, searchable, printable, and scrollable through a thousand pages without waiting for them.
How to embed it →import { DocumentViewer } from '@genomdev/react';
import { all } from '@genomdev/genom/lazy';
export function DocumentPane({ file }: { file: File }) {
// Only the parser for this file's format is fetched.
return <DocumentViewer file={file} formats={all} fit="width" />;
}Read it with code
The same parse handed back as text, as Markdown, or as a tree of blocks with headings, tables and lists intact — for search, for an index, for a model.
How to extract →import { extractMarkdown } from '@genomdev/genom';
// Headings stay headings, tables stay tables,
// lists keep their numbers.
const markdown = await extractMarkdown(file);Office grades it, not us
A renderer nobody can check is a renderer you have to take on faith. This one is scored against Microsoft’s own output, on documents nobody here wrote.
- 01
Run the whole corpus
A corpus of third-party documents — Common Crawl derivatives, the test suites of pandoc, LibreOffice, POI and others — is rendered on every change. None were written by us and none were picked for being easy.
- 02
Let Office score it
Word exports each document to PDF, the pages are read back, and ours are compared against them page by page. PowerPoint does the same for decks. There is no human judgement in the number and no way for it to flatter us.
- 03
Keep what wins, revert what loses
One hypothesis per iteration, because two cannot be told apart afterwards. A change that improves the corpus survives. A change that fixes one document and costs three is reverted, however clever it was.
What it is not
- This is not pixel-perfect Word, and it never will be. The browser lays out text, and its line breaking is not Word’s.
- What is achievable — and what is measured — is that the content is all there, in the right order, correctly styled, with nothing missing, overlapping or garbled.
- If your requirement is a legally identical reproduction for print or signature, convert to PDF on a server. That is a different product and we will say so.
- If your requirement is that a person can read the document, search it, select from it and trust what they see, that is this one.
Try it on your ugliest document
If something renders wrong, send it. It joins the set everything is measured against, which means it stays fixed.