CIVILICA We Respect the Science
(ناشر تخصصی کنفرانسهای کشور / شماره مجوز انتشارات از وزارت فرهنگ و ارشاد اسلامی: ۸۹۷۱)

The XMLization of a Dependency Treebank in CoNLL Format for Evaluating Linguistic Queries using Xquery

عنوان مقاله: The XMLization of a Dependency Treebank in CoNLL Format for Evaluating Linguistic Queries using Xquery
شناسه ملی مقاله: KBEI02_280
منتشر شده در دومین کنفرانس بین المللی مهندسی دانش بنیان و نوآوری در سال 1394
مشخصات نویسندگان مقاله:

Ahmad Pouramini - Department of Computer Engineering Sirjan University of Technology, Sirjan, Iran
Amine Naseri - Department of Computer Engineering Sirjan University of Technology, Sirjan, Iran

خلاصه مقاله:
Treebanks are essential resources for both data-driven approaches to natural language processing (NLP) and empirical linguistic researches. Developing these resources is time- and cost-consuming and requires specialized expertise. Therefore, they should be designed to be reused for different purposes. Currently, there are several dependency treebanks for some languages which are annotated in CoNLL format. For some languages, such as Persian, they are the few available linguistic resources. These treebanks are more suitable for the input of data-driven parsers, and querying linguistic data in them is not easy. In recent years, XML has been widely used for formatting treebanks, and there are various tools available for querying and annotating a linguistic croups in this format. In this paper, we present a tool for converting a dependency treebank in CoNLL format to an appropriate XML format. We designed the XML scheme to be particularly suitable for writing linguistic queries in XQuery syntax.

کلمات کلیدی:
component; Natural language processing; treebanks; depenency structure

صفحه اختصاصی مقاله و دریافت فایل کامل: https://civilica.com/doc/553330/