Quickstart
Four working programs — read, show, search and chunk — with nothing left out.
Read a file on a server
A path, and the document as markdown. Tables survive, headings stay headings, and a chart comes out as the numbers it plots rather than as the word "chart".
import { extractFile } from '@genomdev/genom/node';
const doc = await extractFile('report.docx');
console.log(doc.toMarkdown());
console.log(doc.blocks.length, 'blocks');
console.log(doc.metadata.title, doc.metadata.author);Show a file in a browser
The formats are a prop, not a built-in list. Passing descriptors from @genomdev/genom/lazy means the bundler emits a chunk per format and the reader downloads the one they opened.
import { DocumentViewer } from '@genomdev/react';
import { docx, pdf } from '@genomdev/genom/lazy';
export function Viewer({ file }: { file: File | undefined }) {
return (
<DocumentViewer
file={file}
formats={[docx, pdf]}
fit="width"
onPageChange={({ index, total }) => console.log(index + 1, 'of', total)}
/>
);
}Search what is on screen
The search runs over the extracted text rather than over the DOM, which is why it finds passages on pages that have not been mounted yet and never matches a page number the renderer computed.
const search = await viewer.search();
search.find('quarterly revenue'); // every hit, highlighted
search.next(); // scrolls to the next one
// A passage stored months ago, shown in the document it came from.
search.show(chunk.selectors);Chunk for retrieval
Each chunk carries the heading path above it and the selectors that find it again. Store the selectors next to the embedding; they are what turns a retrieved passage back into a highlight.
import { extractFile } from '@genomdev/genom/node';
const doc = await extractFile('handbook.docx');
for (const chunk of doc.chunks({ maxTokens: 512, overlap: 64 })) {
await index.add({
text: chunk.text,
selectors: chunk.selectors,
});
}