TY - GEN
T1 - No-reference document image quality assessment based on high order image statistics
AU - Xu, Jingtao
AU - Ye, Peng
AU - Li, Qiaohong
AU - Liu, Yong
AU - Doermann, David
N1 - Publisher Copyright:
© 2016 IEEE.
PY - 2016/8/3
Y1 - 2016/8/3
N2 - Document image quality assessment (DIQA) aims to predict the visual quality of degraded document images. Although the definition of 'visual quality' can change based on the specific applications, in this paper, we use OCR accuracy as a metric for quality and develop a novel no-reference DIQA method based on high order image statistics for OCR accuracy prediction. The proposed method consists of three steps. First, normalized local image patches are extracted with regular grid and a comprehensive document image codebook is constructed by K-means clustering. Second, local features are softly assigned to several nearest codewords, and the direct differences between high order statistics of local features and codewords are calculated as global quality aware features. Finally, support vector regression (SVR) is utilized to learn the mapping between extracted image features and OCR accuracies. Experimental results on two document image databases show that the proposed method can accurately predict OCR accuracy and outperforms previous algorithms.
AB - Document image quality assessment (DIQA) aims to predict the visual quality of degraded document images. Although the definition of 'visual quality' can change based on the specific applications, in this paper, we use OCR accuracy as a metric for quality and develop a novel no-reference DIQA method based on high order image statistics for OCR accuracy prediction. The proposed method consists of three steps. First, normalized local image patches are extracted with regular grid and a comprehensive document image codebook is constructed by K-means clustering. Second, local features are softly assigned to several nearest codewords, and the direct differences between high order statistics of local features and codewords are calculated as global quality aware features. Finally, support vector regression (SVR) is utilized to learn the mapping between extracted image features and OCR accuracies. Experimental results on two document image databases show that the proposed method can accurately predict OCR accuracy and outperforms previous algorithms.
KW - Document image quality assessment
KW - High order statistics
KW - No-reference
KW - OCR accuracy
UR - https://www.scopus.com/pages/publications/85006736191
U2 - 10.1109/ICIP.2016.7532968
DO - 10.1109/ICIP.2016.7532968
M3 - Conference contribution
AN - SCOPUS:85006736191
T3 - Proceedings - International Conference on Image Processing, ICIP
SP - 3289
EP - 3293
BT - 2016 IEEE International Conference on Image Processing, ICIP 2016 - Proceedings
PB - IEEE Computer Society
T2 - 23rd IEEE International Conference on Image Processing, ICIP 2016
Y2 - 25 September 2016 through 28 September 2016
ER -