The corpus
Every number and every quoted example in the book comes from these articles. They are open-access papers published under a Creative Commons licence, which is what makes it possible to reproduce and redistribute them here.
Provenance and licence
The articles are drawn from the Elsevier OA CC-BY Corpus, a collection of open-access research papers released for text and data mining:
Kershaw, D., & Koeling, R. (2020). Elsevier OA CC-BY Corpus (Version 2) [Data set]. Mendeley Data. https://doi.org/10.17632/zm33cdndxs.2
All articles carry a CC BY 4.0 licence, which permits redistribution and adaptation with attribution. Each article's authors, title, journal and DOI are listed below, and every example quoted on this site or in the book carries its source. Where a corpus sentence has been simplified for teaching purposes, that is stated at the point of use.
What is in it
Three entries in the original selection were dropped because text extraction had left them with too little body text to analyse; they are listed with reasons in the book manifest. Short Economics Letters papers were separated into their own corpus because they are a different genre — see Unit 14.
Download
The tagged corpus, the definitions, the analysis programs and the results are available from the workspace page. Redistribution is welcome under the same CC BY terms; please keep the attribution file with the text.