Extending corpus annotation of Nepali:advances in tokenisation and lemmatisation

Hardie, Andrew and Lohani, Ram and Yadava, Yogendra (2011) Extending corpus annotation of Nepali:advances in tokenisation and lemmatisation. Himalayan Linguistics, 10 (1). 151–165. ISSN 1544-7502

[img]
Preview
PDF
HLJ1001G.pdf - Published Version
Available under License Creative Commons Attribution-NonCommercial-NoDerivs.

Download (488kB)

Abstract

The Nepali National Corpus (NNC) was, in the process of its creation, annotated with part-of-speech (POS) tags. This paper describes the extension of automated text and corpus annotation in Nepali from POS tags to lemmatisation, enabling a more complex set of corpus-based searches and analyses. This work also addresses certain practical compromises embodied in the initial tagging of the NNC. First, some particular aspects of Nepali morphology – in particular the complexity of the agglutinative verbal inflection system – necessitated improvements to the underlying tokenisation of the text before lemmatisation could be satisfactorily implemented. In practical terms, both the tokenisation and lemmatisation procedures require linguistic knowledge resources to operate successfully: a set of rules describing the default case, and a lexicon containing a list of individual exceptions: words whose form suggests a particular rule should apply to them, but where that rule in fact does not apply. These resources, particularly the lexicons of irregularities, were created by a strongly data-driven process working from analyses of the NNC itself. This approach to tokenisation and lemmatisation, and associated linguistic knowledge resources, may be illustrative and of use to researchers looking at other languages of the Himalayan region, most especially those that have similar morphological behaviour to Nepali.

Item Type:
Journal Article
Journal or Publication Title:
Himalayan Linguistics
Subjects:
ID Code:
62712
Deposited By:
Deposited On:
05 Mar 2013 14:18
Refereed?:
Yes
Published?:
Published
Last Modified:
26 Nov 2020 02:07