Learn With Examples · Probability & Statistics
Nobody can tell you how many minutes your food delivery will take. But a good app can tell you something better: how likely each possible answer is. That complete picture of “what could happen and how often” is a probability distribution, and it sits underneath almost every forecast, insurance premium and quality check you’ll ever meet.
Open a food delivery app and it won’t say “your dinner arrives at 8:14 p.m.” It says “25 to 35 minutes.” That small range is doing something clever. The app knows perfectly well that the real answer might be 22 minutes, or 31, or, on a bad night with a rainstorm and a missing rider, 52. It can’t know which. What it can know, from millions of past deliveries, is how often each outcome happens, and it squeezes that knowledge into the range it shows you.
That is the whole idea of a probability distribution, and it’s far less intimidating than the name suggests. It is simply a complete list of the things that could happen, together with how likely each one is. A single probability answers “how likely is this one outcome?” A distribution answers the bigger question: “across everything that might happen, where does the likelihood pile up, and where does it thin out?”
I’ve spent a lot of years explaining this to people who came in convinced it was a topic for mathematicians. What usually changes their mind is realising they already use distributions constantly, without the vocabulary. Every time you pad a journey because “traffic could be bad”, or keep an umbrella because “it might rain”, you’re reasoning about a spread of outcomes rather than a single prediction. This article puts words and numbers to that instinct.
A probability distribution describes every possible outcome and its likelihood
Take any uncertain quantity: the total of two dice, the number of customers arriving this hour, the height of the next person through the door. The distribution of that quantity tells you which values it can take, and how probable each value (or range of values) is.
Add up all those probabilities and you always get exactly 1, or 100%, because something has to happen. That single fact is what makes a distribution a distribution.
Start with something you can count: two dice
The easiest way to see a distribution is to build one. Roll two fair dice and add them. What totals are possible, and how likely is each?
There are 36 equally likely ways two dice can land (6 faces on the first, times 6 on the second). Count how many of those 36 give each total, and you have the whole distribution:
Look at what the picture tells you that no single number could. A total of 7 isn’t just “possible”, it’s the most likely outcome, and six times as likely as a 2. The totals bunch up in the middle and thin out toward the ends. That shape, a pile-up in the centre with rarer extremes, is why board games built around two dice feel the way they do, and why casinos can price the game at a profit.
Two rules every distribution obeys
No negative probabilities
An outcome is either impossible (0) or has some chance of happening. There’s no such thing as a −5% chance.
Everything sums to 100%
The probabilities of every possible outcome add up to exactly 1, because one of them must occur.
Those two rules are also a handy error-check. Suppose a weather service claims a 30% chance of rain, a 50% chance of cloud without rain, and a 30% chance of clear skies. That adds to 110%, so at least one number is wrong, and you can spot it without knowing any meteorology.
Discrete or continuous?
Distributions come in two families, and the difference is one of the most important ideas in the subject.
Discrete
- Outcomes you can count: 0, 1, 2, 3…
- Examples: dice totals, goals in a match, defective items, calls per hour
- Each individual value has its own probability
- Drawn as separate bars
Continuous
- Outcomes you measure: any value on a scale
- Examples: height, waiting time, temperature, delivery time
- A single exact value has probability zero
- Drawn as a smooth curve; probability is the area under it
The continuous case surprises people, so it’s worth slowing down. What’s the probability that a randomly chosen adult is exactly 170.0000000… centimetres tall, with infinite precision? Zero. There are infinitely many possible heights, and the chance of hitting one precise value shrinks to nothing. What does make sense is a range: the probability of being between 169 and 171 cm. That’s why a continuous distribution is drawn as a curve, and why the probability of a range is the area beneath the curve over that range, not the height of the curve at a point.
A reassuring shortcut. You never need to calculate that area yourself. Software and printed tables do it. What matters is the reading: taller curve means “values here are more concentrated”, and area means probability. The height of the curve is a density, not a probability, which is why it’s fine for the curve to exceed 1 on very narrow distributions.
The two numbers that summarise a distribution
A full distribution is the complete story. Often you want a headline. Two numbers carry most of it: where the distribution is centred, and how spread out it is.
The mean (expected value): the centre of gravity
The expected value is the long-run average: what you’d get if you repeated the experiment a huge number of times. You calculate it by multiplying each outcome by its probability and adding up. For a single fair die: 1×⅙ + 2×⅙ + … + 6×⅙ = 3.5. For two dice it’s exactly 7, right at the peak of the bars above.
Expected value is where distributions start paying rent in real life, because it lets you judge a gamble before you take it. Here’s a scratch card that costs ₹50:
| Prize | Probability | Prize × probability |
|---|---|---|
| ₹0 | 90.0% | ₹0.00 |
| ₹100 | 8.0% | ₹8.00 |
| ₹500 | 1.9% | ₹9.50 |
| ₹10,000 | 0.1% | ₹10.00 |
| Total | 100% | ₹27.50 |
The expected prize is ₹27.50 on a ₹50 ticket, an expected loss of ₹22.50 per card. Nobody loses exactly that amount on any one card (you win 0, 100, 500 or 10,000), but across many cards the average outcome converges on it. Lotteries, insurance and casinos all run on this arithmetic: any single result is random, the average is not.
The standard deviation: how spread out
Two distributions can share the same average and behave completely differently. A delivery service that always takes 30 minutes and one that takes anywhere from 10 to 50 both average 30. The second is far less predictable, and the number that captures that is the standard deviation: roughly, the typical distance of an outcome from the average.
This is why the average alone is a dangerous summary. Someone told “the average commute is 40 minutes” will be on time about half the time. Someone told “usually 35 to 50, occasionally 70” can plan properly. The spread is the difference between a number and an honest forecast.
Five distributions you’ll meet everywhere
Thousands of distributions exist, but a handful cover most of everyday life. Each answers a different kind of question. Tap through them: every one uses a real scenario and exact calculated numbers.
Five distributions you will meet everywhere
tap oneUniform: Rolling a fair die discrete
Every outcome is equally likely, so every bar is the same height.
Each face has probability 1/6. The average is (1+2+3+4+5+6)/6 = 3.5, a value the die can never actually show, which is a useful reminder that the average of a distribution needn’t be a possible outcome. Anything picked at random from a fair list follows this shape: a raffle draw, a shuffled playlist, a randomly assigned seat.
Binomial: Defective items on a production line discrete
The count of “yes” outcomes across a fixed number of independent tries, each with the same chance.
A factory makes phone chargers with a 5% defect rate and tests a box of 20. The chance the box is perfect is 0.95 to the power 20, which is 35.8%: only about one box in three, even though each charger is 95% reliable. Three or more defective units turn up about 7.5% of the time. The same distribution covers free throws made, ad clicks from a fixed number of viewers, or patients responding to a treatment.
Poisson: Calls arriving at a help desk discrete
The count of events in a fixed stretch of time when they arrive independently at a steady average rate.
A help desk averages 4 calls an hour. A completely silent hour has probability e to the power minus 4, which is 1.8%. Exactly 4 calls, the average, is the single most likely count and still only 19.5%. And 8 or more, double the norm, happens in about 5.1% of hours, roughly one hour in twenty, which is why staffing to the average alone leaves you swamped. Buses at a stop, typos per page and goals per football match behave the same way.
Normal: Adult heights continuous
The bell curve: values cluster around an average, with symmetric, quickly thinning tails.
With a mean of 170 cm and a standard deviation of 7 cm, about 68% of adults fall within one standard deviation (163 to 177 cm) and 95% within two (156 to 184 cm). Only about 2.3% are taller than 184 cm. Measurement errors, exam marks and blood pressure readings often follow this shape too, because each is the sum of many small independent influences.
Exponential: Waiting time for a bus continuous
How long until the next event, when events arrive at a steady average rate. Short waits are common; very long ones are rare.
If a bus comes every 5 minutes on average and arrivals are random, the chance you wait more than 10 minutes is e to the power minus 2, which is 13.5%. The curve is tallest at zero and falls away steadily. It has a strange “memoryless” property: having already waited 10 minutes tells you nothing about how much longer you will wait. It pairs naturally with the Poisson: Poisson counts arrivals, exponential times the gaps between them.
Every figure is calculated from the standard formula for that distribution, using typical realistic parameters.
The useful skill isn’t memorising formulas. It’s recognising which question you’re asking. “How many out of a fixed number succeed?” points to binomial. “How many events in a stretch of time?” points to Poisson. “How long until the next one?” points to exponential. “How is a measurement with lots of small influences spread?” points to normal. Match the question to the shape and half the work is done.
Counting successes: the binomial in action
Flip a fair coin 10 times. How many heads? You might expect exactly 5 every time, but you’ll get 5 only about a quarter of the time. Here’s the full distribution:
Two lessons hide in that chart. First, randomness is lumpier than intuition expects: a run of 7 or 8 heads in 10 flips is perfectly ordinary (together about 16% of the time), which is why people so often see “patterns” in pure chance. Second, the distribution is symmetric and bell-shaped even though each individual flip is nothing like a bell curve. That’s a clue to the next, and most famous, distribution.
The bell curve: the normal distribution
Take the number of heads with 10 flips, then 100, then 1,000, and the bars get finer while the outline settles toward the same smooth symmetric hump. Add up enough small independent influences of almost any kind, and the total tends toward this shape. That result is called the central limit theorem, and it’s the reason the normal distribution turns up everywhere from exam scores to measurement errors to blood pressure.
That chart contains the most useful rule of thumb in practical statistics, usually called the 68-95-99.7 rule. For anything that follows a bell curve, roughly 68% of values fall within one standard deviation of the mean, 95% within two, and 99.7% within three. It lets you judge how surprising a value is at a glance. A height of 190 cm is nearly three standard deviations above average, so it’s genuinely rare. A height of 175 cm is unremarkable.
About 68%
Roughly two out of every three values. This is “normal”, the ordinary range.
About 95%
Nineteen out of twenty. Outside this range is unusual enough to notice.
About 99.7%
All but three in a thousand. Beyond this is rare enough to investigate as a possible error.
Not everything is a bell curve. Incomes, city populations, and the sizes of insurance claims are lopsided, with a long tail of very large values, and treating them as normal leads to serious underestimates of extreme events. The average income can sit far above what most people actually earn. Before applying the 68-95-99.7 rule, check that the data really is roughly symmetric.
Where distributions quietly run the world
Pricing risk
Insurers model the distribution of claims. The premium covers the average claim plus a margin for the spread. Get the tail wrong and the company fails.
Spotting faults
Factories track a measurement’s distribution. A value more than 3 standard deviations from target signals that the machine, not chance, has changed.
Is the difference real?
Websites compare two designs by asking how likely the observed gap would be if nothing had changed. That’s a question about a distribution.
Calls, queues, checkouts
The Poisson distribution tells a help desk how many calls to expect, and how often it will be swamped by twice the average.
“70% chance of rain”
A forecast is a distribution over outcomes, summarised into one number. “Seven times in ten, on days like this, it rains.”
Reference ranges
A “normal” blood result is usually the middle 95% of a healthy population’s distribution, so about 1 in 20 healthy people fall outside it by definition.
Six mistakes people make
Treating the average as the most likely outcome
They coincide for a bell curve, but not in general. The average roll of one die is 3.5, an impossible result. The average household income is well above the most common one. Mean, median and mode are three different questions about the same distribution.
Ignoring the spread
Two options with the same average can carry very different risk. Always ask “average, and how variable?” before comparing anything: investments, delivery times, exam results, blood pressure.
Expecting streaks to “even out”
After five heads in a row, the next flip is still 50/50. The distribution of future flips has no memory. Over many flips the proportion settles toward half, but not because tails are “due”. This mix-up is called the gambler’s fallacy.
Assuming everything is normal
The bell curve is common but not universal. Extreme events in finance, insurance and natural disasters follow heavier-tailed shapes, in which very large outcomes are far more likely than a normal curve predicts.
Reading a continuous curve’s height as a probability
The height is a density. Probability is the area under the curve over a range. The probability of any single exact value on a continuous scale is zero.
Forgetting probabilities must total 100%
If the numbers in any claimed distribution don’t add to 1, something’s wrong. It’s the quickest sanity check there is, and it catches errors in reports, forecasts and even published statistics.
Check yourself
Five questions. Open each to check. The correct option is marked.
1. A distribution has probabilities 0.2, 0.3, 0.1 and x for its four outcomes. What is x?
- 0.3
- 0.4
- 0.5
- 0.6
All probabilities must sum to 1. 0.2 + 0.3 + 0.1 = 0.6, so x = 1 − 0.6 = 0.4.
2. What is the probability of rolling a total of 7 with two fair dice?
- 1/12
- 1/36
- 1/6
- 7/36
Six of the 36 combinations give 7 (1+6, 2+5, 3+4, 4+3, 5+2, 6+1). 6/36 = 1/6, about 16.7%.
3. For a continuous distribution, what is the probability of exactly one specific value?
- The height of the curve at that point
- Zero
- 1 divided by the number of values
- Impossible to say
There are infinitely many possible values, so any single exact value has probability zero. Probabilities belong to ranges, as areas under the curve.
4. What is the expected value of one roll of a fair die?
- 3
- 4
- 3.5
- 6
(1+2+3+4+5+6) ÷ 6 = 3.5. It isn’t a possible roll, but it’s the long-run average.
5. Heights have mean 170 cm and standard deviation 7 cm. About 95% of adults fall between which values?
- 163 and 177 cm
- 156 and 184 cm
- 149 and 191 cm
- 170 and 184 cm
95% lies within two standard deviations: 170 ± 14 gives 156 to 184 cm.
Frequently asked questions
What is a probability distribution in simple terms?
It is a complete description of every possible outcome of an uncertain event and how likely each is. The probabilities always add up to 1 (100%). It shows not just what could happen but where the likelihood is concentrated.
What is the difference between discrete and continuous distributions?
Discrete distributions cover countable outcomes, such as dice totals or goals scored, and give each value its own probability. Continuous distributions cover measurements on a scale, such as height or time, and give probability only to ranges, as the area under a curve.
What is the most common probability distribution?
The normal (bell curve) distribution. It appears wherever a result is the sum of many small independent influences, which is why heights, measurement errors, exam scores and many biological readings roughly follow it.
What does the standard deviation tell you?
It measures how spread out a distribution is: roughly the typical distance of a value from the average. A small standard deviation means outcomes cluster tightly; a large one means they vary widely.
What is the difference between probability and a probability distribution?
A probability is a single number for one outcome, such as a 1/6 chance of rolling a three. A probability distribution is the full set of outcomes together with their probabilities, showing how likelihood is spread across everything that could happen.
How are probability distributions used in real life?
In insurance pricing, quality control, medical reference ranges, weather forecasts, staffing for call centres, A/B testing of websites, and any decision where the outcome is uncertain and you need to weigh how likely different results are.
The takeaway
A probability distribution is the full map of an uncertain outcome: everything that could happen, and how likely each possibility is. It obeys two simple rules (no negative probabilities, and the total is 100%), it comes in a discrete form (bars for countable results) and a continuous form (curves, with probability as area), and it can be summarised by a centre, the expected value, and a spread, the standard deviation.
The practical habit it builds is a good one: stop asking “what will happen?” and start asking “what’s the range of things that could happen, and how likely is each?” That’s the difference between a single guess that’s usually wrong and a forecast you can plan around, whether you’re timing a journey, pricing a risk or judging whether a result is a fluke.
Try it on something in your own week. Note how long your commute actually takes, every day for two weeks. Plot the results as a little bar chart. You’ll have built a real distribution, and you’ll almost certainly see it’s lumpier and wider than “about 35 minutes” ever suggested.
probability distributionnormal distributionexpected valuestandard deviationstatistics basicsbinomial
