| Asia Pacific Journal of Corpus Research Vol. 7, No. 1, pp. 1-5 |
| Abbreviation: APJCR |
| e-ISSN: 2733-8096 |
| Publication date: 31 August 2026 |
| Received: 13 June 2026 / Received in Revised Form: 13 July 2026 / Accepted: 15 August 2026 |
| DOI: https://doi.org/10.22925/apjcr.2026.7.1.1 |
Revisiting the Brown Corpus: Compilation Principles, Comparative Research, and AntConc-based Pedagogy |
| Chae Kwan Jung (Incheon National University), SOUTH KOREA |
| Copyright 2026 APJCR This is an open access article distributed under the terms of the Creative Commons Attribution License 4.0, which permits unrestricted, distribution, and reproduction in any medium, provided the original work is properly cited. |
Abstract |
| This article reexamines the Brown Corpus as a landmark resource in corpus linguistics and a practical basis for comparative research and data-driven language education. Drawing on the official corpus manual, publications, and the Brown family of comparable corpora, it reviews the corpus’s compilation principles, sampling frame, genre composition, versions, annotation, and technological production. It also synthesizes five representative studies of sentence complexity, modal change, genitive alternation, relative clause variation, and word-frequency norms. The review shows that the corpus’s enduring value lies less in its one-million-word size than in its transparent sampling design and replicability across regional and diachronic corpora, including LOB, Frown, FLOB, and AmE06. At the same time, its 1961 publication date, restriction to edited prose, unequal genre sizes, excerpt-based structure, and version-dependent tokenization limit generalization and require documentation of analytical conditions. Three AntConc activities are consequently proposed to connect corpus design with hands-on inquiry: analyzing the meanings and genre distribution of must, comparing American and British English through Brown and LOB keywords, and distinguishing the near-synonyms begin and start. The article argues that Brown remains valuable when treated not as a model of present-day English, but as a controlled historical benchmark for reproducible comparison and corpus literacy. |
Keywords |
| Brown Corpus, Brown Family Corpora, Comparative Corpus Linguistics, AntConc, Data-Driven Learning |
References |
| 권혁승, 정채관. (2012). 코퍼스 언어학 입문. 서울: 한국문화사. / Kwon, H. S., & Jung, C. K. (2012). Corpus Linguistics Introduction. Seoul: Hankook Publishing House.
권혁승, 정채관, 김재훈. (2018). 코퍼스 언어학 기초. 서울: 한국문화사. / Kwon, H. S., Jung, C. K., & Kim, J. H. (2018). Corpus Linguistics Basic. Seoul: Hankook Publishing House. 정채관. (2026). 쉽게 배우는 코퍼스 영어학 입문. 인천: 국립인천대학교출판부. / Jung, C. K. (2026). An Easy Introduction to English Corpus Linguistics. Incheon: Incheon National University Press. Anthony, L. (2026). AntConc (Version 4.4.2) [Computer software]. Tokyo, Japan: Waseda University. Retrieved July 20, 2026, from https://www.laurenceanthony.net/software/AntConc Brysbaert, M., & New, B. (2009). Moving beyond Kučera and Francis: A critical evaluation of current word frequency norms and the introduction of a new and improved word frequency measure for American English. Behavior Research Methods, 41(4), 977-990. Francis, W. N., & Kučera, H. (1961). Brown corpus of present day American English [Data set]. Oxford Text Archive, University of Oxford. Retrieved June 2, 2026, from https://hdl.handle.net/ 20.500.14106/0402 Francis, W. N., & Kučera, H. (1979). Manual of information to accompany A standard corpus of present-day edited American English, for use with digital computers (Rev. and amplified ed.). Providence, RI: Department of Linguistics, Brown University. (Original work published 1964). Retrieved June 2, 2026, from http://icame.uib.no/brown/bcm.html Hinrichs, L., & Szmrecsanyi, B. (2007). Recent changes in the function and frequency of Standard English genitive constructions: A multivariate analysis of tagged corpora. English Language & Linguistics, 11(3), 437-474. Hinrichs, L., Szmrecsanyi, B., & Bohmann, A. (2015). Which-hunting and the Standard English relative clause. Language, 91(4), 806-836. Kučera, H. (1980). Computational analysis of predicational structures in English. In COLING 1980 volume 1: The 8th International Conference on Computational Linguistics (pp. 32-37). Retrieved June 2, 2026, from https://aclanthology.org/C80-1006/ Kučera, H., & Francis, W. N. (1967). Computational analysis of present-day American English. Providence, RI: Brown University Press. Leech, G. (2004). Recent grammatical change in English: Data, description, theory. In K. Aijmer and B. Altenberg (Eds.), Advances in corpus linguistics: Papers from the 23rd International Conference on English Language Research on Computerised Corpora (ICAME 23) (pp. 61-81). Amsterdam: Rodopi. |
The Author |
| Chae Kwan Jung is an Associate Professor in the Department of English Language and Literature at Incheon National University, South Korea. He holds a BEng (Hons) in Manufacturing Engineering and Japanese from the University of Birmingham, UK, and an MSc in Engineering Business Management and an EdD in Applied Linguistics and English Language Teaching from the University of Warwick, UK. He is the founder of the Institute for Corpus Research (ICR) and the founding editor of the Asia Pacific Journal of Corpus Research (APJCR). His research interests include corpus linguistics, English language education, English for Specific Purposes (ESP), language assessment, and curriculum design and development. His current interests include data-driven learning, AI-assisted language pedagogy and translation, and the application of ESP-informed genre- and corpus-based approaches to language teaching and learning. |
The Author’s Address |
| First and Corresponding Author Chae Kwan Jung Professor Department of English Language & Literature Incheon National University 119 Academy-ro, Yeonsu-gu, Incheon, 22012, SOUTH KOREA E-mail: ckjung@inu.ac.kr |
