MOC Statistics and Inference
These notes share a theme, and it is worth stating rather than leaving implicit. Most of the statistical trouble I encounter in climate work is not a failure of technique. It is the conversion of a continuous quantity of evidence into a discrete verdict — significant or not, model selected or rejected — after which the verdict circulates and the evidence behind it does not.
Amrhein and colleagues make the case directly, and it is the paper I would read first. Everything else here is either the machinery that verdict-making runs on or an alternative to it.
The machinery
Statistical inference is the general problem: what a sample licenses about a population. Hypothesis testing is the dominant framework for doing it, test statistics are what it computes, and the hypothesis testing notes work through the mechanics.
Knowing the mechanics is what makes the critique usable rather than merely agreeable. In my experience the impulse to classify results comes from not knowing what a test actually does — a p-value is easier to treat as a verdict if you are unsure what it measures. The argument for reporting continuous evidence is much easier to hold once the machinery is clear.
Alternatives to verdicts
Information criteria compare models on a continuous scale rather than accepting or rejecting them, which is a genuine improvement — though nothing prevents them being used as thresholds too, and they frequently are. Probabilistic model selection takes the framing further by treating model choice as an inference problem in its own right.
Where it lands in the climate work
The connection is not incidental. Counterfactual climate data and attribution work produce changes in likelihood, not yes-or-no answers about whether an event was caused by warming, and the pressure to convert them into the latter is exactly the pressure this map is about. Applied climatology faces the same problem from the decision side, where a single number is demanded and is nearly always wrong in a way that matters.