Skip to content
Genom
Start

Quickstart

Four working programs — read, show, search and chunk — with nothing left out.

Read a file on a server

A path, and the document as markdown. Tables survive, headings stay headings, and a chart comes out as the numbers it plots rather than as the word "chart".

read.ts
import { extractFile } from '@genomdev/genom/node';

const doc = await extractFile('report.docx');

console.log(doc.toMarkdown());
console.log(doc.blocks.length, 'blocks');
console.log(doc.metadata.title, doc.metadata.author);

Show a file in a browser

The formats are a prop, not a built-in list. Passing descriptors from @genomdev/genom/lazy means the bundler emits a chunk per format and the reader downloads the one they opened.

Viewer.tsx
import { DocumentViewer } from '@genomdev/react';
import { docx, pdf } from '@genomdev/genom/lazy';

export function Viewer({ file }: { file: File | undefined }) {
  return (
    <DocumentViewer
      file={file}
      formats={[docx, pdf]}
      fit="width"
      onPageChange={({ index, total }) => console.log(index + 1, 'of', total)}
    />
  );
}

The search runs over the extracted text rather than over the DOM, which is why it finds passages on pages that have not been mounted yet and never matches a page number the renderer computed.

search.ts
const search = await viewer.search();

search.find('quarterly revenue'); // every hit, highlighted
search.next();                    // scrolls to the next one

// A passage stored months ago, shown in the document it came from.
search.show(chunk.selectors);

Chunk for retrieval

Each chunk carries the heading path above it and the selectors that find it again. Store the selectors next to the embedding; they are what turns a retrieved passage back into a highlight.

index.ts
import { extractFile } from '@genomdev/genom/node';

const doc = await extractFile('handbook.docx');

for (const chunk of doc.chunks({ maxTokens: 512, overlap: 64 })) {
  await index.add({
    text: chunk.text,
    selectors: chunk.selectors,
  });
}