Skip to content
Genom
Concepts

Addressing

What joins extraction to the viewer: an address that survives another process, another machine and a relayout.

Extraction produces text. Something downstream finds a passage in it and wants to see that passage in the document. Between the two there has to be an address, and coordinates cannot be one — a Word document is a flow of paragraphs and has no pages until it is laid out. So an address names a node of the model.

a locator
genom:1:9f3ac21b:body/t2/w4/c1/b0/r0+12
└ scheme ┘ │ hash  └ flow ┘└──── steps ────┘└ offset
           version

Why there is a hash in it

The eight hex digits are a hash of the file’s bytes, and they are the reason an address can be trusted. An anchor arriving from a different file is *detected* rather than silently resolved to whatever node happens to sit at that path — which is the difference between a corpus that tells you it has drifted and one where a user reports a highlight in the wrong place.

Three addresses, not one

An anchor follows the W3C Web Annotation model: the exact locator, character offsets into a named serialisation, and the quoted text with its neighbours. They fail at different times — the first when somebody saves the file, the second when the extraction options change, the third only when the words themselves go — and the resolver reports which one matched.

resolving
import { resolveSelectors } from '@genomdev/core/content';

const resolved = resolveSelectors(doc, stored.selectors);

resolved?.match;  // 'locator' | 'position' | 'quote'
resolved?.range;  // where it landed

Text the renderer invented

List numbers, footnote markers and tab leaders are put on the page as ordinary text, but nobody typed them. Anything counting characters to find a position has to skip them, so the renderer marks them data-synthetic and the serialisers file them as derived rather than source.

Getting this wrong is silent: the address lands a few characters off, only in documents that have a numbered list.