Addressing
What joins extraction to the viewer: an address that survives another process, another machine and a relayout.
Extraction produces text. Something downstream finds a passage in it and wants to see that passage in the document. Between the two there has to be an address, and coordinates cannot be one — a Word document is a flow of paragraphs and has no pages until it is laid out. So an address names a node of the model.
genom:1:9f3ac21b:body/t2/w4/c1/b0/r0+12
└ scheme ┘ │ hash └ flow ┘└──── steps ────┘└ offset
versionWhy there is a hash in it
The eight hex digits are a hash of the file’s bytes, and they are the reason an address can be trusted. An anchor arriving from a different file is *detected* rather than silently resolved to whatever node happens to sit at that path — which is the difference between a corpus that tells you it has drifted and one where a user reports a highlight in the wrong place.
Three addresses, not one
An anchor follows the W3C Web Annotation model: the exact locator, character offsets into a named serialisation, and the quoted text with its neighbours. They fail at different times — the first when somebody saves the file, the second when the extraction options change, the third only when the words themselves go — and the resolver reports which one matched.
import { resolveSelectors } from '@genomdev/core/content';
const resolved = resolveSelectors(doc, stored.selectors);
resolved?.match; // 'locator' | 'position' | 'quote'
resolved?.range; // where it landedText the renderer invented
List numbers, footnote markers and tab leaders are put on the page as ordinary text, but nobody typed them. Anything counting characters to find a position has to skip them, so the renderer marks them data-synthetic and the serialisers file them as derived rather than source.
Getting this wrong is silent: the address lands a few characters off, only in documents that have a numbered list.