33 terms
The glossary
B
- Base rate
How common something is in the population before any test or evidence.
The prior prevalence. If one person in a thousand has a condition, that is the base rate, and it governs what any test result actually tells you.
Ignoring it is the single most consequential statistical error in ordinary life. A test that is 99% accurate for a condition affecting one in ten thousand produces about a hundred false positives for every true one — not because the test is bad, but because there is so much more healthy population to be wrong about.
Probability storiesSeen in odds grid chartsSee also: false positive, probability, expected value
C
- CAGR
The constant annual growth rate that would produce the observed change.
Compound annual growth rate smooths a start-to-end change into one yearly figure. It is useful for comparing growth across different periods.
It is entirely determined by the first and last values, so it is blind to everything between them. A series that crashed and recovered and one that grew steadily can report the same CAGR.
Finance educationSeen in slope chartsSee also: compounding, log scale, nominal vs real
- Causation
One thing actually bringing about another.
Establishing it needs more than data that moves together: either an experiment that intervenes, or a careful argument ruling out the alternatives.
The useful habit is to name the third variable that would explain both. If you can, the causal claim is not yet earned; if you genuinely cannot, it becomes more plausible.
Statistical literacySee also: correlation, confounder
- Cohort
A group defined by when it started, tracked over time.
Cohort analysis compares like with like by following everyone who joined in the same period, rather than mixing new and old together in a monthly total.
An overall retention number can hold steady while every individual cohort gets worse — growth keeps refilling the top. Only the cohort view shows it.
Business and startupsSeen in matrix heatmap chartsSee also: survivorship bias, weighted average
- Compounding
Growth applied to a total that already includes previous growth.
Because each period grows the base for the next, small rate differences separate enormously over time. Intuition is linear and consistently underestimates this.
It runs in reverse too, and asymmetrically: a 50% loss needs a 100% gain to get back to level, which is why avoiding large drawdowns matters more than catching large rises.
Finance educationSeen in line chartsSee also: cagr, log scale, nominal vs real
- Confidence interval
A range of values compatible with the data, at a stated confidence level.
A 95% interval is produced by a method that captures the true value 95% of the time across repeated samples. The interval is the finding; the point estimate in the middle is a convenience.
Wide intervals are informative — they say the data does not settle the question. Reporting only the midpoint of a very wide interval turns uncertainty into a fact.
Statistical literacySee also: margin of error, statistical significance, sample size
- Confounder
A third factor driving both things you are comparing.
Ice cream sales correlate with drownings because both rise with temperature. Temperature is the confounder, and once you account for it the relationship disappears.
Most spurious findings are confounders that nobody looked for. Wealth, age and population size confound an enormous share of published comparisons.
Statistical literacySee also: causation, correlation, simpson's paradox
- Correlation
Two measures that move together, without any claim about why.
A correlation coefficient runs from −1 to 1 and describes how tightly two variables track each other in a straight line. It says nothing about direction of influence and nothing about non-linear relationships.
A correlation of zero means no linear relationship, not no relationship. A perfect U-shape can score near zero.
Statistical literacySeen in scatter chartsSee also: causation, confounder, outlier
D
- Distribution
The full shape of the data: which values occur and how often.
Every summary statistic is a compression of the distribution, and every compression throws something away. The distribution is what was there before the throwing away.
Most arguments about a statistic dissolve once the distribution is on screen. A mean and a median that disagree, a long tail, a second peak, a floor at zero — all of it is visible at once and none of it survives into a single number.
Statistical literacySeen in histogram chartsSee also: mode, skew, standard deviation
E
- Expected value
The average outcome if you repeated a gamble indefinitely.
Multiply each outcome by its probability and add them up. It is the right way to compare bets, prices and risks that repeat.
It is the wrong tool for one-off decisions with ruinous downsides. A bet with positive expected value that can bankrupt you on the first try is still a bad bet, because you do not get to run it indefinitely.
Probability storiesSee also: probability, base rate
F
- False positive
A test saying yes when the answer is no.
Every test trades false positives against false negatives; you cannot reduce both without a better test. Where the dial is set is a judgement about which error costs more.
The headline accuracy figure hides this entirely. What matters is how many of the positive results are real, and that depends on the base rate as much as on the test.
Probability storiesSeen in odds grid chartsSee also: base rate, probability
I
- Index
A series rebased so one period equals 100, for comparing shapes rather than levels.
Rebasing lets series of wildly different sizes be compared on growth. Everything starts at 100 and the lines show relative change from there.
The base period is a choice and it is doing a lot of work. Shift the start into a slump and everything afterwards looks like a boom.
Data visualization lessonsSeen in line chartsSee also: nominal vs real, log scale, cagr
L
- Log scale
An axis where equal distances mean equal multiples, not equal amounts.
On a log axis the gap from 1 to 10 is the same as from 10 to 100. It is the correct choice for anything growing by percentages, and it turns exponential growth into a straight line whose slope is the growth rate.
It also flattens dramatic curves, which can be honest clarification or a way to make a steep rise look tame. Log axes should always be labelled as such.
Data visualization lessonsSeen in line chartsSee also: index, cagr, compounding
M
- Margin of error
How far a sample figure is likely to sit from the truth.
Usually quoted as plus-or-minus some percentage points at 95% confidence, and usually covering sampling error only — not bad question wording, not non-response, not a skewed sample.
Two figures whose margins overlap have not been shown to differ. A three-point lead with a four-point margin is a tie being reported as a result.
Statistical literacySee also: confidence interval, sample size, selection bias
- Mean
The total divided by the count — the arithmetic average.
Add every value, divide by how many there are. It is the average most people mean by “average”, and it is the right summary when the data is roughly symmetrical.
Its weakness is that every value pulls on it in proportion to its size, so a handful of extreme cases can drag it somewhere no actual case sits. Average wealth in a room changes by billions when one person walks in; average height does not.
Statistical literacySee also: median, skew, outlier
- Median
The middle value: half the cases sit above it, half below.
Line every value up in order and take the one in the middle. Unlike the mean, it does not care how extreme the extremes are — only how many of them there are.
That makes it the honest summary for anything skewed: incomes, house prices, wait times, company sizes. When a report gives you a mean for one of those and no median, the omission is usually doing work.
Statistical literacySee also: mean, percentile, skew
- Mode
The most common value in a set.
The value that appears most often. It is the only one of the three averages that works on categories rather than numbers — the most common job title has a mode but no mean.
A distribution with two modes is usually two populations that have been added together, and splitting them is almost always more informative than summarising them.
Statistical literacySeen in histogram chartsSee also: mean, median, distribution
N
- Nominal vs real
Figures before and after adjusting for inflation.
Nominal figures are in the money of their day; real figures are restated in one year's money so they can be compared. Any money series spanning more than a couple of years needs the real version.
“Highest ever” claims about revenue, salaries or box office are usually nominal, and usually stop being records once adjusted.
Finance educationSeen in line chartsSee also: index, cagr, compounding
O
- Outlier
A value far from the rest — sometimes an error, sometimes the whole point.
An outlier is only defined relative to an expectation, so calling something an outlier is a claim about what you expected, not a property of the data.
Deleting them is a decision that needs a reason. A mis-keyed figure should go; a genuinely extreme case usually should not, because in skewed domains the extremes carry most of the total.
Statistical literacySeen in scatter chartsSee also: skew, mean, distribution
P
- p-value
The probability of seeing data this extreme if there were no real effect.
It is not the probability that the hypothesis is true, and it is not the probability the result was a fluke. It answers a narrower question than almost everyone reading it assumes.
Because a threshold of 0.05 passes one in twenty null results, testing twenty things and reporting the one that passed is a reliable method for finding nothing at all.
Statistical literacySee also: statistical significance, confidence interval
- Per capita
A total divided by population, so places of different sizes can be compared.
Totals rank countries roughly by how many people live in them. Per-capita figures rank them by intensity, and the two orders are often close to reversed.
Which one is right depends on the question. Total emissions matter to the atmosphere; per-capita emissions matter to an argument about responsibility.
Public interestSeen in world map chartsSee also: weighted average, index
- Percentile
The value below which a given percentage of cases fall.
The 90th percentile is the value that 90% of cases come in under. Percentiles describe position rather than distance, which makes them robust to extremes and easy to act on.
They are the right tool for anything skewed or long-tailed: response times, salaries, delivery delays. “The 95th percentile doubled” is a specific, checkable claim in a way that “the average rose slightly” is not.
Statistical literacySee also: median, distribution, outlier
- Probability
How often something happens across many chances, expressed from 0 to 1.
A probability is a long-run frequency, not a schedule. One-in-a-hundred odds do not mean once every hundred attempts; they mean that across very many attempts the rate settles near one in a hundred.
Small probabilities become near-certainties once the number of chances is large enough, which is why extremely unlikely things happen to somebody every day.
Probability storiesSeen in odds grid chartsSee also: base rate, expected value, sample size
R
- Regression to the mean
Extreme results tend to be followed by more ordinary ones.
When a result mixes skill and luck, an extreme showing usually had unusual luck in it. The luck does not repeat, so the next result lands closer to average — with no change in underlying ability.
This manufactures the illusion that criticism works and praise backfires: you intervene after the worst performances, and they would have improved anyway.
Sports analyticsSeen in slope chartsSee also: sample size, outlier, survivorship bias
S
- Sample size
How many cases a figure is based on.
Precision improves with the square root of the sample, so quadrupling the sample halves the error. Small samples do not just give vaguer answers — they give wilder ones.
Small-sample extremes explain a lot of apparent geography. The regions with both the highest and the lowest rates of a rare condition are usually the least populated ones.
Statistical literacySee also: margin of error, confidence interval, regression to the mean
- Seasonal adjustment
Removing the regular annual pattern so the underlying trend is visible.
Retail always rises in December and unemployment always rises when school ends. Adjustment strips out the part of a movement that happens every year, leaving what is actually new.
It means adjusted figures are estimates produced by a model, and they get revised. Comparing an adjusted figure with an unadjusted one is a category error.
Everyday dataSeen in line chartsSee also: index, nominal vs real
- Selection bias
A sample that was not drawn fairly from the population it claims to describe.
Anyone who answered the survey chose to answer it, and that choice usually correlates with the thing being measured. Online polls measure who was online and motivated.
No sample size fixes it. A biased sample of a million is more confidently wrong than a biased sample of a hundred.
Statistical literacySee also: survivorship bias, sample size, margin of error
- Simpson's paradox
A trend that holds in every subgroup and reverses when they are combined.
It happens when the groups differ both in their rates and in their sizes. Aggregate the data and the group sizes, not the rates, decide the answer.
It is not a curiosity. Any comparison of overall rates between two populations with different compositions is exposed to it — which is most comparisons between countries, hospitals or schools.
Statistical literacySee also: confounder, weighted average, base rate
- Skew
An asymmetric distribution, with one tail longer than the other.
Right-skewed data has a long tail of high values — income, city populations, book sales. Left-skewed is the mirror image and is rarer in practice.
The practical consequence is simple: in a right-skewed distribution the mean sits above the median, most cases sit below the mean, and quoting the mean overstates the typical case.
Statistical literacySeen in histogram chartsSee also: mean, median, outlier
- Standard deviation
A measure of how spread out values are around the mean.
Roughly, the typical distance between a value and the average. Two datasets can share a mean and be nothing alike; the standard deviation is the first number that tells them apart.
It assumes a reasonably symmetrical distribution to be interpretable. On heavily skewed data it is still computable and mostly useless — percentiles say more.
Statistical literacySee also: distribution, percentile, mean
- Statistical significance
A result unlikely to have arisen by chance alone, under a stated threshold.
It is a claim about chance, not importance. A trivial difference becomes significant with a large enough sample, and a large difference stays non-significant with a small one.
The right follow-up question is always the effect size: significant how much? A statistically significant improvement of 0.2% is a real effect nobody should act on.
Statistical literacySee also: p-value, confidence interval, sample size
- Survivorship bias
Studying only what made it through, and concluding something about everything.
The startups that failed do not write retrospectives, and the funds that closed leave the index. What remains is a sample selected on the outcome you are trying to explain.
The test is to ask where the missing cases went, and whether they would have had the trait you are crediting for success.
Business and startupsSee also: selection bias, sample size
W
- Weighted average
An average where some cases count more than others.
Averaging the rates of several groups without weighting by group size treats a village and a city as equals. That is occasionally what you want and usually a mistake.
Unweighted averages of group rates are one of the most common routes into Simpson's paradox.
Statistical literacySee also: mean, simpson's paradox, per capita
Missing a term you ran into? Tell us and it gets added.