Open an old Word 97-2003 file
Free, no sign-up, and the file never leaves your computer.
A .doc file is not a .docx. It has no XML, no parts and no element names — it is a run of characters with formatting scattered through the file in 512-byte pages and a table at the front saying where everything is.
This reads that structure directly. Nothing is converted to .docx on the way, and nothing is sent anywhere.
The piece table, and why old documents come out scrambled elsewhere
Word never rewrote a .doc file when you edited it — it appended your change to the end and updated an index called the piece table saying what order to read the fragments in. A reader that ignores the piece table gets the text of the document in the order it was typed over its whole lifetime, not the order it reads in. That is why so many tools produce a .doc that is almost right and subtly scrambled. Text boxes are a second trap: a shape’s words live in a story of their own, joined to it by an index hidden in the high half of a property, and 41 documents in our corpus were entirely blank until that was handled.
Everything below is parsed from the file, not approximated
Read from the bytes the format actually stores, into a model that does not know which generation produced it.
The container and the text
Formatting
Structure
Tables
Pictures and shapes
Text boxes
Shapes, groups and embedded objects
Word 6.0 and Word 95
What this does not do
Every renderer has a list like this. Most do not publish it.
- Shapes nested more than one group deep are not placed.
- A geometry the shared shape library does not know — a curved connector, a star — is drawn as the rectangle it occupies.
- A picture whose location points past the end of the stream that should hold it is dropped: about one in twelve of the corpus.
- Word 6.0 and Word 95 documents are read without character or paragraph formatting: their property opcodes are a byte wide, and reading them with the newer decoder would invent formatting rather than omit it.
- Encrypted documents are detected but cannot be opened.
If a file of yours renders wrong, send it to admin@genom.dev. It joins the corpus everything is measured against, which is what makes a fix stay fixed.
Things people ask
- Is my document uploaded anywhere?
- No. The file is opened by JavaScript running in this tab and is never sent anywhere — there is no upload, no server-side conversion and no copy kept. You can confirm it: open your browser’s network panel, then open a file. Nothing leaves. This also means the page works with no connection at all once it has loaded.
- Is the file converted to .docx first?
- No. The binary format is parsed straight into the same content model the modern format produces. Nothing round-trips through another format, so nothing is lost to a conversion step.
- What about Word 6.0 and Word 95 files?
- They open — the text, the paragraphs and the tables come out. Character and paragraph formatting does not: those versions encode formatting opcodes a byte wide, and reading them with the newer decoder would invent formatting rather than omit it.
- Why can other browser viewers not do this?
- Almost none read the binary format at all; they either refuse the file or ask a server to convert it. Reading it in a browser means implementing the OLE2 compound file, the piece table and the formatted disk pages, which is a large amount of work for a format everyone hoped would go away.
The viewer, and the parser on its own
The parser produces a model that knows nothing about CSS or the DOM, so the same call runs in Node for extraction, search and conversion. Word 97-2003 and its other generation are one package.
import { DocumentViewer } from '@genomdev/react';
<DocumentViewer file={file} fit="width" />;