The 385+ million word Corpus of Contemporary American English (1990–2008+)
Abstract
The Corpus of Contemporary American English ( COCA ), which was released online in early 2008, is the first large and diverse corpus of American English. In this paper, we first discuss the design of the corpus — which contains more than 385 million words from 1990–2008 (20 million words each year), balanced between spoken, fiction, popular magazines, newspapers, and academic journals. We also discuss the unique relational databases architecture, which allows for a wide range of queries that are not available (or are quite difficult) with other architectures and interfaces. To conclude, we consider insights from the corpus on a number of cases of genre-based variation and recent linguistic variation, including an extended analysis of phrasal verbs in contemporary American English.
Journal: International Journal of Corpus Linguistics
Publisher: John Benjamins Publishing Company
Citations are the number of DOI-registered works in Crossref that cite this paper; references are how many works it cites. Full text is on the publisher site via the DOI link.