Chunk for retrieval
Chunks that carry their context and know where they came from.
A chunker that splits on character count produces passages that begin mid-sentence and belong to nothing. This one splits on the structure the document already has — sections, headings, table rows — and prefixes each chunk with the heading path above it, so a passage retrieved on its own still says what it is about.
import { extractFile } from '@genomdev/genom/node';
const doc = await extractFile('handbook.docx');
for (const chunk of doc.chunks({ maxTokens: 512, overlap: 64 })) {
await index.add({
text: chunk.text, // contextualised with its heading path
selectors: chunk.selectors, // how to find it again
meta: { source: 'handbook.docx', hash: doc.hash },
});
}Chunk options
| Option | Default | Notes |
|---|---|---|
maxTokens | 512 | Counted with your own counter if you pass one. |
overlap | 64 | Capped at half of maxTokens. |
contextualize | true | Prefix each chunk with the headings above it. |
tables | 'rows' | How a table is split: by rows, or kept whole. |
tokenCounter | length / 4 | Pass your model’s tokeniser for exact sizes. |
Closing the loop
The selectors are the reason to use this rather than a text splitter. Store them beside the embedding, and a retrieved passage can be shown in the document it came from instead of quoted beside it.
Each address carries a hash of the file’s bytes, so an anchor arriving from a different file is detected rather than silently resolved to whatever sits at that path. That is how a corpus tells you it has drifted before a user reports a highlight in the wrong place.
const viewer = useViewer({ formats: all });
async function reveal(passage: RetrievedPassage) {
await viewer.open(await fetchFile(passage.source));
const search = await viewer.viewer!.search();
search.show(passage.selectors);
}A whole corpus
npx @genomdev/genom chunks corpus/**/*.docx --max-tokens 512 > chunks.jsonl