Skip to content
Genom
Tasks

Chunk for retrieval

Chunks that carry their context and know where they came from.

A chunker that splits on character count produces passages that begin mid-sentence and belong to nothing. This one splits on the structure the document already has — sections, headings, table rows — and prefixes each chunk with the heading path above it, so a passage retrieved on its own still says what it is about.

index.ts
import { extractFile } from '@genomdev/genom/node';

const doc = await extractFile('handbook.docx');

for (const chunk of doc.chunks({ maxTokens: 512, overlap: 64 })) {
  await index.add({
    text: chunk.text,          // contextualised with its heading path
    selectors: chunk.selectors, // how to find it again
    meta: { source: 'handbook.docx', hash: doc.hash },
  });
}

Chunk options

OptionDefaultNotes
maxTokens512Counted with your own counter if you pass one.
overlap64Capped at half of maxTokens.
contextualizetruePrefix each chunk with the headings above it.
tables'rows'How a table is split: by rows, or kept whole.
tokenCounterlength / 4Pass your model’s tokeniser for exact sizes.

Closing the loop

The selectors are the reason to use this rather than a text splitter. Store them beside the embedding, and a retrieved passage can be shown in the document it came from instead of quoted beside it.

Each address carries a hash of the file’s bytes, so an anchor arriving from a different file is detected rather than silently resolved to whatever sits at that path. That is how a corpus tells you it has drifted before a user reports a highlight in the wrong place.

Answer.tsx
const viewer = useViewer({ formats: all });

async function reveal(passage: RetrievedPassage) {
  await viewer.open(await fetchFile(passage.source));
  const search = await viewer.viewer!.search();
  search.show(passage.selectors);
}

A whole corpus

npx @genomdev/genom chunks corpus/**/*.docx --max-tokens 512 > chunks.jsonl