TY - GEN
T1 - Approximate SPARQL for error tolerant queries on the DBpedia knowledge base
AU - Tauer, Gregory
AU - Rudnicki, Ronald
AU - Sudit, Moises
PY - 2013
Y1 - 2013
N2 - The Resource Description Framework (RDF), a language for describing resources, is being used more commonly in information fusion systems. SPARQL is a standard query language that enables knowledge extraction from data encoded in RDF. A SPARQL query is, in essence, an exact subgraph matching problem. Unfortunately, many of the techniques that produce data in RDF (such as manual data entry, social network analysis, natural language processing, etc.) make annotation mistakes, which result in dirty RDF data. SPARQL performs suboptimally on RDF data containing errors since, as an exact graph matching tool, it is not designed to cope with noisy data. To improve knowledge extraction under these conditions, we propose an extension to SPARQL that permits approximate graph matches. This allows queries to cope with errors in the RDF graph, both on the attribute level (such as misspelled names) as well as on the structural level (missing or extra edges). We use the TruST heuristic algorithm to solve the underlying approximate graph matching problem and demonstrate the benefit it brings to answering questions on the DBpedia knowledge base.
AB - The Resource Description Framework (RDF), a language for describing resources, is being used more commonly in information fusion systems. SPARQL is a standard query language that enables knowledge extraction from data encoded in RDF. A SPARQL query is, in essence, an exact subgraph matching problem. Unfortunately, many of the techniques that produce data in RDF (such as manual data entry, social network analysis, natural language processing, etc.) make annotation mistakes, which result in dirty RDF data. SPARQL performs suboptimally on RDF data containing errors since, as an exact graph matching tool, it is not designed to cope with noisy data. To improve knowledge extraction under these conditions, we propose an extension to SPARQL that permits approximate graph matches. This allows queries to cope with errors in the RDF graph, both on the attribute level (such as misspelled names) as well as on the structural level (missing or extra edges). We use the TruST heuristic algorithm to solve the underlying approximate graph matching problem and demonstrate the benefit it brings to answering questions on the DBpedia knowledge base.
UR - https://www.scopus.com/pages/publications/84890826070
M3 - Conference contribution
AN - SCOPUS:84890826070
SN - 9786058631113
T3 - Proceedings of the 16th International Conference on Information Fusion, FUSION 2013
SP - 850
EP - 856
BT - Proceedings of the 16th International Conference on Information Fusion, FUSION 2013
T2 - 16th International Conference of Information Fusion, FUSION 2013
Y2 - 9 July 2013 through 12 July 2013
ER -