Skip to content
Genom

Open an old Word 97-2003 file

Free, no sign-up, and the file never leaves your computer.

Samplesample.doc
Loading the viewer…
Also opensDrop any of them on the viewer above — the format is recognised from the bytes.

A .doc file is not a .docx. It has no XML, no parts and no element names — it is a run of characters with formatting scattered through the file in 512-byte pages and a table at the front saying where everything is.

This reads that structure directly. Nothing is converted to .docx on the way, and nothing is sent anywhere.

How the format actually works

The piece table, and why old documents come out scrambled elsewhere

Word never rewrote a .doc file when you edited it — it appended your change to the end and updated an index called the piece table saying what order to read the fragments in. A reader that ignores the piece table gets the text of the document in the order it was typed over its whole lifetime, not the order it reads in. That is why so many tools produce a .doc that is almost right and subtly scrambled. Text boxes are a second trap: a shape’s words live in a story of their own, joined to it by an index hidden in the high half of a property, and 41 documents in our corpus were entirely blank until that was handled.

What is supported

Everything below is parsed from the file, not approximated

Read from the bytes the format actually stores, into a model that does not know which generation produced it.

The container and the text

The OLE2 compound file, the file information block, and the piece table — which is what puts the text back in reading order after Word has spent an afternoon appending edits to the end of the file. Eight-bit text is decoded in the code page the document was written in, taken from the font, the metadata and the language in that order.

Formatting

Character and paragraph properties out of the formatted disk pages: fonts, sizes, colours, the toggles and their invert-what-you-inherited form, underlines, highlights, borders, shading, indents, spacing, alignment and tab stops. The stylesheet is resolved parent-first and named the way the modern format names styles, so a heading from 1998 picks up the same built-in defaults as one from last week.

Structure

Sections with page setup and columns, headers and footers for first, odd and even pages, footnotes, endnotes and comments with their authors, fields with their instructions and cached results, and list definitions with their overrides.

Tables

Rows reconstructed from the marks in the text and the row definition that follows them, cell boundaries turned back into widths, horizontal and vertical merges, cell borders and alignment, and tables nested inside cells.

Pictures and shapes

The Office Art drawing layer — the same one .xls and .ppt use — with its picture store and the shapes that name it. PNG and JPEG pass through, device-independent bitmaps are wrapped as PNG, and EMF and WMF are translated into SVG by performing the drawing calls they record, because that is all a metafile is. Inline pictures keep their scaling and cropping; anchored ones keep their rectangle, their wrapping and their rotation.

Text boxes

A shape’s words live in a story of their own, joined to it by an index hidden in the high half of a property. A poster, a form or a flow chart says everything it says inside one, and 41 documents of the corpus were blank without this.

Shapes, groups and embedded objects

A shape with neither picture nor words keeps its geometry, its fill and its outline, which is what the rules, frames and panels of a report are. A group is anchored once and holds its members in a coordinate space of its own, so they are placed by mapping that space onto the rectangle the anchor gives. And an object embedded twenty years ago keeps the preview picture it was pasted with, which is the only thing a reader can show for it.

Word 6.0 and Word 95

The generation before this one has a shorter header, bin tables of half the width and no table stream at all. Those it reads; the text, the paragraphs and the tables come out, which is the difference between a document being readable and not.
Being straight with you

What this does not do

Every renderer has a list like this. Most do not publish it.

  • Shapes nested more than one group deep are not placed.
  • A geometry the shared shape library does not know — a curved connector, a star — is drawn as the rectangle it occupies.
  • A picture whose location points past the end of the stream that should hold it is dropped: about one in twelve of the corpus.
  • Word 6.0 and Word 95 documents are read without character or paragraph formatting: their property opcodes are a byte wide, and reading them with the newer decoder would invent formatting rather than omit it.
  • Encrypted documents are detected but cannot be opened.

If a file of yours renders wrong, send it to admin@genom.dev. It joins the corpus everything is measured against, which is what makes a fix stay fixed.

Questions

Things people ask

Is my document uploaded anywhere?
No. The file is opened by JavaScript running in this tab and is never sent anywhere — there is no upload, no server-side conversion and no copy kept. You can confirm it: open your browser’s network panel, then open a file. Nothing leaves. This also means the page works with no connection at all once it has loaded.
Is the file converted to .docx first?
No. The binary format is parsed straight into the same content model the modern format produces. Nothing round-trips through another format, so nothing is lost to a conversion step.
What about Word 6.0 and Word 95 files?
They open — the text, the paragraphs and the tables come out. Character and paragraph formatting does not: those versions encode formatting opcodes a byte wide, and reading them with the newer decoder would invent formatting rather than omit it.
Why can other browser viewers not do this?
Almost none read the binary format at all; they either refuse the file or ask a server to convert it. Reading it in a browser means implementing the OLE2 compound file, the piece table and the formatted disk pages, which is a large amount of work for a format everyone hoped would go away.
Using it in your own code

The viewer, and the parser on its own

The parser produces a model that knows nothing about CSS or the DOM, so the same call runs in Node for extraction, search and conversion. Word 97-2003 and its other generation are one package.

import { DocumentViewer } from '@genomdev/react';

<DocumentViewer file={file} fit="width" />;