From Universal Dependencies to error tagging: Using dependency parsing to annotate errors in Greek learner sentences
Date
Authors
Journal Title
Journal ISSN
Volume Title
Publisher
Abstract
This master’s thesis concerns the development of tools for automatic annotation of learner data with a focus on Greek. Firstly, I developed a tool for automatic error annotation, on an attempt to extract mainly morphological errors from the morphosyntactic annotation of parallel learner and standard data in the Universal Dependencies (UD) framework. The human evaluation of the tool proves the difficulty of this task, though also the tool’s suitability for a restricted domain of application. In addition, different models for syntactic parsing were trained, using different percentages of UDannotated standard and learner data. The primary aim of this second task was choosing the most fitting training data for in-domain testing, as well as measure the effect that the particular nature of learner data, when occurring in the training set, can have on model performance. The quantitative and qualitative evaluation indicate a significant improvement of performance on the learner test setting when learner data have been utilized in the training process, and the best performing model across domains was the one trained on both standard and learner data. Moreover, the addition of learner data to the training did not seem to have a negative effect on model performance overall. Lastly, a combination of the aforementioned morphosyntactic and error annotation methods was attempted, in order to annotate novel learner data; the reported results show a compatibility of the employed methods, nevertheless a limitation of performance, and thus space for improvement. The importance of this work lies in the enhancement of computational methods and automated tools used for morphosyntactic parsing and error tagging, as well as in the promotion of linguistic research on language learning, boosting a re-evaluation of learning and teaching methods