Lancaster EPrints

Writing the Vernacular: Transcribing and Tagging the Newcastle Electronic Corpus of Tyneside English

Beal, J. and Corrigan, K. and Smith, N. and Rayson, P. (2007) Writing the Vernacular: Transcribing and Tagging the Newcastle Electronic Corpus of Tyneside English. Studies in Variation, Contacts and Change in English, 1. ISSN 1797-4453

Full text not available from this repository.

Abstract

The Newcastle Electronic Corpus of Tyneside English (NECTE) presented a number of problems not encountered by those producing corpora of standard varieties. The primary material consisted of audio recordings which needed to be orthographically transcribed and grammatically tagged. Preston (1985), (2000), Macaulay (1991), Kirk (1997), Cameron (2001) and Beal (2005) all note that representing vernacular Englishes orthographically, e.g. by using "eye dialect", can be problematic on various levels. Apart from unwelcome associations with negative political, racial or social connotations, there are theoretical objections to devising non-standard spellings which represent certain groups of vernacular speakers, thus making their speech appear more differentiated from mainstream colloquial varieties than is warranted. In the first half of this paper, we outline the principles and methods adopted in devising an Orthographic Transcription Protocol (OTP) for such a vernacular corpus, and the challenges faced by the NECTE team in practice. Protocols for grammatical tagging have likewise been devised with standard varieties in mind. In the second half, we relate how existing part-of-speech (POS)-tagging software (CLAWS4, cf. Garside & Smith 1997; and Template Tagger, cf. Fligelstone et al. 1997) had to be adapted to take account of the non-standard grammar of Tyneside English.

Item Type: Article
Journal or Publication Title: Studies in Variation, Contacts and Change in English
Uncontrolled Keywords: cs_eprint_id ; 1518 cs_uid ; 355
Subjects: Q Science > QA Mathematics > QA75 Electronic computers. Computer science
Departments: Faculty of Science and Technology > School of Computing & Communications
ID Code: 13005
Deposited By: ep_importer_comp
Deposited On: 21 Jun 2008 21:50
Refereed?: Yes
Published?: Published
Last Modified: 17 Sep 2013 08:16
Identification Number:
URI: http://eprints.lancs.ac.uk/id/eprint/13005

Actions (login required)

View Item