Handwritten text recognition for unseen historical Danish & Swedish court hands (1778–2021)
A pipeline of domain-adapted line segmentation and trained recognisers, studied under font- and layout-distribution shift, with honest measured evaluation of where systems fail.
The general vision-language models do not just score lower — they fail silently, producing fluent output that is wrong with no signal. A domain-adapted segmenter and trained recognisers, with a period-aware lexicon, close most of the gap.
The segmenters feed trained line recognisers (CTC specialists and a vision-language model) with period-aware lexicon decoding. Full results are in an in-preparation benchmark paper co-authored with Andy Stauder (Transkribus / READ-COOP).