Abstract
Traditional forensic document analysis methods have focused on feature-classification paradigm where a machine learning based classifier is used to learn discrimination among multiple writers. However, usage of such techniques is restricted to availability of a large labeled dataset which is not always feasible. In this paper, we propose a Cotraining based approach that overcomes this limitation by exploiting independence between multiple views (features) of data. Two learners are initially trained on different views of a smaller labeled training data and their initial hypothesis is used to predict labels on larger unlabeled dataset. Confident predictions from each learner are used to add such data points back to the training data with predicted label as the ground truth label, thereby effectively increasing the size of labeled dataset and improving the overall classification performance. We conduct experiments on publicly available IAM dataset and illustrate the efficacy of proposed approach.
| Original language | English |
|---|---|
| Pages (from-to) | 36-40 |
| Number of pages | 5 |
| Journal | CEUR Workshop Proceedings |
| Volume | 768 |
| State | Published - 2011 |
| Event | 1st International Workshop on Automated Forensic Handwriting Analysis, AFHA 2011 - A Satellite Workshop of ICDAR 2011 - Beijing, China Duration: Sep 17 2011 → Sep 18 2011 |
Keywords
- Classifier
- Co-training
- Labeled and unlabeled data
- Views
- Writer identification
Fingerprint
Dive into the research topics of 'A co-training based framework for writer identification in offline handwriting'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver