TY - GEN
T1 - Document ranking by layout relevance
AU - May, Huang
AU - DeMenthon, Daniel
AU - Doermann, David
AU - Golebiowski, Lynn
PY - 2005
Y1 - 2005
N2 - This paper describes the development of a new document ranking system based on layout similarity. The user has a need represented by a set of "wanted" documents, and the system ranks documents in the collection according to this need. Rather than performing complete document analysis, the system extracts text lines, and models layouts as relationships between pairs of these lines. This paper explores three novel feature sets to support scoring in large document collections. First, pairs of lines are used to form quadrilaterals, which are represented by their turning functions. A non-Euclidean distance is used to measure similarity. Second, the quadrilaterals are represented by 5D Euclidean vectors, and third, each line is represented by a 5D Euclidean vector. We compare the classification performance and computation speed of these three feature sets using a large database of diverse documents including forms, academic papers and handwritten pages in English and Arabic. The approach using quadrilaterals and turning functions produces slightly better results, but the approach using vectors to represent text lines is much faster for large document databases.
AB - This paper describes the development of a new document ranking system based on layout similarity. The user has a need represented by a set of "wanted" documents, and the system ranks documents in the collection according to this need. Rather than performing complete document analysis, the system extracts text lines, and models layouts as relationships between pairs of these lines. This paper explores three novel feature sets to support scoring in large document collections. First, pairs of lines are used to form quadrilaterals, which are represented by their turning functions. A non-Euclidean distance is used to measure similarity. Second, the quadrilaterals are represented by 5D Euclidean vectors, and third, each line is represented by a 5D Euclidean vector. We compare the classification performance and computation speed of these three feature sets using a large database of diverse documents including forms, academic papers and handwritten pages in English and Arabic. The approach using quadrilaterals and turning functions produces slightly better results, but the approach using vectors to represent text lines is much faster for large document databases.
UR - https://www.scopus.com/pages/publications/33947431007
U2 - 10.1109/ICDAR.2005.92
DO - 10.1109/ICDAR.2005.92
M3 - Conference contribution
AN - SCOPUS:33947431007
SN - 0769524206
SN - 9780769524207
T3 - Proceedings of the International Conference on Document Analysis and Recognition, ICDAR
SP - 362
EP - 366
BT - Proceedings of the Eighth International Conference on Document Analysis and Recognition
T2 - 8th International Conference on Document Analysis and Recognition
Y2 - 31 August 2005 through 1 September 2005
ER -