Examples
Every one of these runs in the page you are reading, on a real file, with the code that produced it underneath. Nothing is uploaded and there is no server behind any of it — the parsers are the published packages, fetched when the file is opened.
How this is arranged
By product, not by feature. A format family gets a group, and inside it are the things that family can actually be made to do — which are not the same from one to the next. A workbook evaluates formulas and a deck does not; a PDF has text to recover and a spreadsheet has values to read. Grouping the other way would force every format into the same short list of verbs, and the first format that did not fit would break it.
Each example is its own page, so the ones worth finding can be found: recovering the text of a PDF is a different problem from turning a Word document into Markdown, and somebody is looking for exactly one of them.
Everything here
- Render a documentPages broken where Word breaks them — measured, not guessed.
- Extract text and MarkdownThe document on one side, what a pipeline reads on the other.
- Search and highlightHits painted on pages that are not in the DOM.
- Render a workbookBoth axes virtualised, and the number formats the file asks for.
- Read a workbook as dataSheets, rows and cells with their computed values.
- Search across sheetsFind a value in a workbook that is mostly not rendered.
- Render a deckMaster → layout → slide, with shapes, charts and SmartArt drawn.
- Pull out the slide textTitles, bullets and speaker notes, in slide order.
- Render a PDFPages as the file describes them, with text you can select.
- Recover the textGlyphs at coordinates, clustered back into sentences.