Methods and definitions

A statistic about hedging means nothing without a definition of hedging. This page carries both halves: what was counted (the word lists, below, generated from the same files the analysis reads) and how it was counted (the rules, which matter just as much).

General conventions

Body text
All section text except references, abstracts and journal boilerplate (funding, conflicts of interest, acknowledgements, ethics statements). This is the denominator for rates unless a measure says otherwise.
Rates
Occurrences per 10,000 words. A rate of 10/10k is roughly once per 1,000 words — about once per printed page.
Tokenizer
One definition throughout: a token starts with a letter or digit and may continue with letters, digits, apostrophes, hyphens, periods, % or $. So U.S., z-score, 23.6% are one token each.
Matching
Case-insensitive, on whole words, with no stemming: only the listed surface forms count. Some families are patterns rather than plain words, because the bare word has too many neutral uses in economics — see the notes on each list.
Sections
Classified from the article's headings. A heading matching no functional category is a thematic (or content) heading. Combined headings such as “Data and Results” count once, under a fixed priority order. Subsections are tracked separately: where a journal sets top-level headings in mixed case, a paper's real architecture lives one level down.
Mathematics
Equations do not survive text extraction, so theory papers' word counts are understated relative to print.

Counting rules that change the numbers

Negated boosters count as hedges
A booster preceded within about two words by no, not, never, without, hardly, scarcely, cannot, lack(s) or little is counted as a hedge, not a booster: “no clear evidence” weakens a claim. In this corpus that reclassification affects several hundred occurrences, and without it the hedge-to-boost ratio is badly wrong.
Frames, not bare words
Where a word's plain form is dominated by neutral uses, only the epistemic frame counts. Likely counts in “it is likely that” or “unlikely to be”, but not in “more likely to quit” (a probability claim about the world, not a hedge). Evidence counts only as “strong / clear / compelling evidence”. Appear counts as “appears to”, not “appears in Table 3”.
Sequential connectives are a lower bound
Only sentence-initial capitalised uses count — First, … Second, … — so that “the first column” is excluded. Mid-sentence organising uses (“we first estimate”) are missed; the figure understates rather than overstates.
Reporting-verb stance uses a two-ring window
The verb is taken from the same sentence as the citation, within 15 tokens, stopping if the writer's own voice takes over (we, our) and skipping nominal frames (“the Card (1990) study”). An anaphoric follow-up sentence (“He finds that…”) is tallied separately rather than mixed in.
Citations are reported two ways
Counting every reference inside a parenthetical cluster gives one figure; counting each cluster once gives another. Both are reported, because the choice moves the integral share by about ten points.
Diagnostic ranges come from the corpus
Where the book gives a target range, it is the middle half of published articles — quartiles computed at run time, not a figure chosen by the author.
Judgment is separated from counting
Two classifications could not be automated honestly: the structural family of each paper and the niche strategy of each introduction. Both were coded by hand from corpus evidence, one primary category each, with the criteria written down; the programs read the finished coding sheets. Presence of a rhetorical move is detected by pattern, but move segmentation — which sentence belongs to which move — is not attempted.
What the corpus cannot show
Tables, figures and reference lists do not survive extraction, so claims about them in the book are presented as professional advice rather than corpus findings. The same applies to the “common mistakes” sections: a corpus of published papers cannot display the errors that reviewers removed.

The word lists

Generated from definitions/*.yaml — the same files the analysis reads.

Loading definitions…