Climate Data

Climate data covers the records used to characterise the state and variability of the climate system. The useful division is by how the numbers came to exist, because that determines what can legitimately be asked of them — and it is a division that gets flattened as soon as everything is on the same grid in the same netCDF file.

Three kinds, three sets of limits

Observations are measurements: station records, radiosondes, buoys, satellite retrievals. They are the only category with a direct link to reality, and they are sparse, unevenly distributed, and inhomogeneous in time. Instruments change, stations move, and the record’s own history is a source of spurious trends. Satellite retrievals are a partial exception in coverage and a complication in kind — a retrieval is a physical quantity inferred through a model, so it sits somewhere between observation and product.

Model output is a physically consistent evolution of a system that is not the real one. Its internal consistency is its strength: every variable is compatible with every other, and closed budgets can be computed. Its weakness is that consistency with itself implies nothing about correspondence with reality.

Reanalysis blends the two, assimilating observations into a model to produce a gridded, complete, physically consistent estimate. It is the most immediately convenient category and the most easily over-trusted. ERA5 is not an observational dataset. Where observations are dense it is strongly constrained; where they are sparse — the upper ocean, polar regions, the pre-satellite era, much of the tropics before 1979 — it is closer to a model run with occasional corrections. The same file gives no indication of which regime a given grid point is in, and that is the single most common way reanalysis gets misused.

The comparison problem

Most climate analysis compares across these categories, and the comparison is rarely like-for-like. A station measures a point; a grid cell represents an area average. These are different quantities, and the difference is not noise — area averages have systematically less variance and weaker extremes than point measurements. Comparing modelled grid-cell precipitation against a gauge and finding the model underestimates extremes may reflect nothing more than the scale mismatch.

The same applies in time. Daily means computed from hourly data and daily means computed from a max-min pair are different quantities that share a name.

I try to be explicit about what each number represents before comparing, and to state the mismatch rather than absorb it into the result. This is unglamorous and it is where a large fraction of apparent model errors turn out to live.

See also: ERA5 for the reanalysis I use most, datasets for the packaging unit, metadata for what makes any of it interpretable, and counterfactual climate data for a fourth category with its own rules.