Recovering the scriptio inferior (undertext) from palimpsests remains a significant challenge in digital heritage and historical document analysis due to the severe degradation of the scraped ink and complex interference from the visible scriptio superior (overtext). A prime example is the collection at the Biblioteca Capitolare of Verona; despite previous digitization efforts, the undertext in these manuscripts often remains illegible to Optical Character Recognition (OCR) models and, in severe cases, even to expert paleographers. Traditional enhancement techniques frequently fail to disentangle these overlapping signals effectively. This paper presents a novel approach to this problem using Blind Image Decomposition (BID). We modify the standard BID architecture to incorporate asymmetric decoders, allowing for specialized processing of the distinct text layers. To address the erosion of the undertext, we introduce a feature consistency loss to ensure robust latent representations, alongside a connectivity-preserving loss designed to enforce structural continuity in faint, fragmented strokes. Furthermore, to bridge the domain gap between training and inference, we employ a masked pretraining strategy on synthetic mixtures that leverages the spatial priors of the undertext script. Experimental results show superior separation performance compared to the baseline BID method. Finally, qualitative evaluation on the degraded Codex Veronensis XL (38) from the Biblioteca Capitolare of Verona confirms the model’s efficacy as a tool for paleographic analysis. Code and data available at: https://github.com/bhrnprhmd-has/BID-palimpsest
Blind Image Decomposition for Recovering Overlapping Text Layers on Palimpsests
Leontaridis, Panagiotis;Scattolin, Paolo;
2026-01-01
Abstract
Recovering the scriptio inferior (undertext) from palimpsests remains a significant challenge in digital heritage and historical document analysis due to the severe degradation of the scraped ink and complex interference from the visible scriptio superior (overtext). A prime example is the collection at the Biblioteca Capitolare of Verona; despite previous digitization efforts, the undertext in these manuscripts often remains illegible to Optical Character Recognition (OCR) models and, in severe cases, even to expert paleographers. Traditional enhancement techniques frequently fail to disentangle these overlapping signals effectively. This paper presents a novel approach to this problem using Blind Image Decomposition (BID). We modify the standard BID architecture to incorporate asymmetric decoders, allowing for specialized processing of the distinct text layers. To address the erosion of the undertext, we introduce a feature consistency loss to ensure robust latent representations, alongside a connectivity-preserving loss designed to enforce structural continuity in faint, fragmented strokes. Furthermore, to bridge the domain gap between training and inference, we employ a masked pretraining strategy on synthetic mixtures that leverages the spatial priors of the undertext script. Experimental results show superior separation performance compared to the baseline BID method. Finally, qualitative evaluation on the degraded Codex Veronensis XL (38) from the Biblioteca Capitolare of Verona confirms the model’s efficacy as a tool for paleographic analysis. Code and data available at: https://github.com/bhrnprhmd-has/BID-palimpsestI documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.



