Abstract
Pixel-precise Document Layout Analysis (DLA) is a critical step in processing historical and handwritten manuscripts that face challenges such as limited annotations and severe class imbalance. Existing few-shot methods typically rely on handcrafted data augmentation and standard semantic segmentation backbones. Moreover, the absence of pre-trained weights further restricts generalization across heterogeneous scripts. In this work, we propose AF-CANet, an augmentation-free, curriculum-guided attention-based U-Net framework for few-shot DLA. Unlike prior methods that rely heavily on handcrafted augmentation or large-scale pretraining, AF-CANet introduces an Efficient Dilated Channel Attention module to capture rich multiscale global semantic features and a cross-dataset curriculum learning strategy to enhance generalisation under limited supervision. Our model outperforms state-of-the-art methods on the DIVA-HisDB benchmark, achieving F 1 and IoU gains of 8.96% and 8.8%, respectively, without external augmentation. Comprehensive ablation studies validate each component’s contribution.
| Original language | English |
|---|---|
| Pages (from-to) | 34-41 |
| Number of pages | 8 |
| Journal | Pattern Recognition Letters |
| Volume | 206 |
| DOIs | |
| State | Published - Aug 2026 |
Keywords
- Document layout analysis
- Few-shot learning
- Handwritten recognition
- Semantic segmentation
Fingerprint
Dive into the research topics of 'AF-CANet: Augmentation-free, curriculum-guided attention UNet framework for few-shot document layout analysis'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver