University of Oulu

Coats, Steven (2019) A corpus of regional American language from YouTube. In: Navarretta, C., et al. (eds.) Proceedings of the Digital Humanities in the Nordic Countries 4th Conference Copenhagen, Denmark, March 5-8, 2019, 2364, pp. 79-91. http://ceur-ws.org/Vol-2364/7_paper.pdf

A corpus of regional American language from YouTube

Saved in:
Author: Coats, Steven1
Organizations: 1English Philology, University of Oulu, 90014 Oulu, Finland
Format: article
Version: published version
Access: open
Online Access: PDF Full Text (PDF, 1.3 MB)
Persistent link: http://urn.fi/urn:nbn:fi-fe2019102534827
Language: English
Published: RWTH Aachen University, 2019
Publish Date: 2019-10-25
Description:

Abstract

Recent years have seen an increase in the number of corpora of regional language variation for English, allowing new types of aggregate analysis to be conducted. While the creation of a corpus from written language material is relatively straightforward, transcribing speech is time-consuming, and thus there are no large corpora of transcribed American speech with broad geographic coverage. This paper describes the creation of a new corpus of regional American English from the automatically generated captions of videos from YouTube channels with a local American focus — mainly channels of regional and local government entities or civic organizations. The corpus, which consists of transcripts of over 29,267 hours of spoken language, will enable the analysis of regional patterns of lexical, morphosyntactic, and other types of variation in spoken American English. Exploratory analysis and mapping of the corpus data indicates regional variation in spoken language is evident.

see all

Series: CEUR workshop proceedings
ISSN: 1613-0073
ISSN-E: 1613-0073
ISSN-L: 1613-0073
Issue: 2364
Pages: 79 - 91
Article number: 7
Host publication: Proceedings of the Digital Humanities in the Nordic Countries 4th Conference Copenhagen, Denmark, March 5-8, 2019
Host publication editor: Navarretta, Costanza
Agirrezabal, Manex
Maegaard, Bente
Conference: Digital Humanities in the Nordic Countries
Type of Publication: A4 Article in conference proceedings
Field of Science: 6121 Languages
113 Computer and information sciences
112 Statistics and probability
1171 Geosciences
Subjects:
Copyright information: © by the paper’s authors. Copying permitted for private and academic purposes.