Recover a PDF’s text in reading order
Glyphs at coordinates, clustered back into sentences.
.pdfOne package, every generation — @genomdev/pdf
Loading the example…
Copying out of a PDF so often produces jumbled text because the file records the order the glyphs were drawn in, which is not the order they are read in. Recovering the text means clustering glyphs into lines by their baselines, working out where the spaces are from the gaps, and deciding which column comes first.
Table regions are recovered from ruling lines and alignment as well, so a table in a PDF comes out of extraction as a table rather than as rows of loose words.
import { extractMarkdown } from '@genomdev/genom';
// Headings stay headings, tables stay tables, and the document's own
// list numbers come out as text rather than as CSS.
const markdown = await extractMarkdown(file); // report.pdf