DocumentCode
3762072
Title
The XMLization of a dependency treebank in CoNLL format for evaluating linguistic queries using XQuery
Author
Ahmad Pouramini;Amine Naseri
Author_Institution
Department of Computer Engineering, Sirjan University of Technology, Sirjan, Iran
fYear
2015
Firstpage
1058
Lastpage
1062
Abstract
Treebanks are essential resources for both data-driven approaches to natural language processing (NLP) and empirical linguistic researches. Developing these resources is time- and cost-consuming and requires specialized expertise. Therefore, they should be designed to be reused for different purposes. Currently, there are several dependency treebanks for some languages which are annotated in CoNLL format. For some languages, such as Persian, they are the few available linguistic resources. These treebanks are more suitable for the input of data-driven parsers, and querying linguistic data in them is not easy. In recent years, XML has been widely used for formatting treebanks, and there are various tools available for querying and annotating a linguistic croups in this format. In this paper, we present a tool for converting a dependency treebank in CoNLL format to an appropriate XML format. We designed the XML scheme to be particularly suitable for writing linguistic queries in XQuery syntax.
Keywords
"Decision support systems","XML","Natural language processing"
Publisher
ieee
Conference_Titel
Knowledge-Based Engineering and Innovation (KBEI), 2015 2nd International Conference on
Type
conf
DOI
10.1109/KBEI.2015.7436191
Filename
7436191
Link To Document