Skip to content
Genom

Recover a PDF’s text in reading order

Glyphs at coordinates, clustered back into sentences.

.pdfOne package, every generation — @genomdev/pdf

Loading the example…

Copying out of a PDF so often produces jumbled text because the file records the order the glyphs were drawn in, which is not the order they are read in. Recovering the text means clustering glyphs into lines by their baselines, working out where the spaces are from the gaps, and deciding which column comes first.

Table regions are recovered from ruling lines and alignment as well, so a table in a PDF comes out of extraction as a table rather than as rows of loose words.

import { extractMarkdown } from '@genomdev/genom';

// Headings stay headings, tables stay tables, and the document's own
// list numbers come out as text rather than as CSS.
const markdown = await extractMarkdown(file); // report.pdf

Put it in your own application

What is above is the published package running in this page, with nothing behind it. The documentation covers the same ground with the API in full.