Every test you have ever trusted can be wrong in exactly two ways. It can shout when nothing is happening, or it can stay silent when something is. Statisticians call these Type I and Type II errors, and once you see them clearly, you will start noticing them in smoke alarms, spam folders, courtrooms and hospital screenings.
A story I tell every new analyst
Years ago I worked with a team that monitored server performance. They had an alert that fired whenever response time crossed a threshold. In the first month it fired forty times, and thirty-nine were nothing. By the second month, people stopped reading the messages. In the third month the one alert that mattered arrived, and it sat unread for two hours while a real outage grew.
That team had made both classic mistakes, one after the other. First they set the alarm too jumpy, which produced false alarms. Then, by training everyone to ignore it, they made real signals get missed. If you understand why that happened, you already understand the core of this article.
Here is the promise: by the end, you will be able to name each error, explain the trade-off between them, use the words alpha, beta and power without flinching, and decide which error matters more in a given situation. No advanced maths required. When the numbers do appear, I have computed them so you can follow along.
The idea of a null hypothesis in plain words
Before we can talk about errors, we need one setup idea. Almost every statistical test begins with a boring default assumption called the null hypothesis, written H0. It says “nothing is going on.” The new drug does nothing. The new web page converts no better than the old one. The coin is fair. The defendant is innocent.
Then you look at evidence and ask a single question: is this evidence surprising enough, under the assumption that nothing is going on, that I should stop believing it? If yes, you reject the null. If not, you fail to reject it, which is different from proving it true. It just means you did not find enough evidence.
The whole topic in two lines.
Type I error: you reject the null when it was actually true. A false alarm, or false positive.
Type II error: you fail to reject the null when it was actually false. A missed signal, or false negative.
The four possible outcomes
Any test ends in one of four situations, depending on what is true in reality and what the test says. Two are correct decisions and two are errors. The grid below is worth memorising.
Look at how symmetrical the grid is, and how differently the errors feel. A false alarm is loud and visible: someone complains, someone investigates, someone wastes an afternoon. A missed signal is silent by nature. Nobody notices what did not happen. That imbalance in visibility is why people routinely underweight Type II errors.
The memory trick that actually works
Students confuse the two constantly, and honestly so did I in the beginning. The trick that finally fixed it for me is the boy who cried wolf.
- The first time, he shouts “Wolf!” and there is no wolf. That is a false alarm: Type I. Type one, the first mistake in the story.
- The second time, a real wolf shows up and the villagers ignore him. That is a missed signal: Type II. Type two, the second mistake.
The order in the fable matches the numbering. If you remember nothing else from this article, remember the wolf.
Alpha and beta: the two probabilities behind the errors
Each error has a probability with a Greek letter attached.
- α (alpha) is the probability of a Type I error. It is also called the significance level, and you choose it before you look at the data. The most common choice is 0.05, meaning you accept a 5% false-alarm rate when nothing is really happening.
- β (beta) is the probability of a Type II error, the chance you miss a real effect.
- Power = 1 − β is the probability you catch a real effect. Researchers often aim for at least 80%.
Think of alpha as how easily you are fooled by noise, and power as how well you can hear a real signal. A good study keeps the first low and the second high.
Seeing the trade-off with two overlapping curves
Here is the picture I draw on whiteboards. There are two bell curves. The left one shows what your test statistic looks like when nothing is going on. The right one shows what it looks like when there is a real effect. You choose a cut-off: any result to the right of the line triggers the alarm.
The red tail is the false-alarm zone. Even when nothing is going on, a small share of results land beyond the line, and with this cut-off that share is 5.0%. The blue tail is the missed-signal zone. When the effect is real, some results still fall short of the line, and here that is 19.6%.
Now imagine sliding the cut-off. Move it to the left, and you catch more real effects but raise more false alarms. Move it to the right, and false alarms shrink while you miss more real effects. You cannot push both down by sliding the line. This is the fundamental trade-off of hypothesis testing. Here are actual numbers for four cut-off positions on this same pair of curves.
| Cut-off (z) | False alarms (α) | Missed (β) | Power | Character |
|---|---|---|---|---|
| 1 | 15.9% | 6.7% | 93.3% | Very sensitive, cries wolf often |
| 1.645 | 5.0% | 19.6% | 80.4% | The classic 5% setting |
| 2.326 | 1.0% | 43.1% | 56.9% | Stricter, 1% false alarms |
| 3 | 0.1% | 69.1% | 30.9% | Very strict, misses more real effects |
Read down the columns. As the cut-off rises, the false alarms fall from a wild figure to a tiny one, but the missed signals climb steadily. There is no free lunch. The only question is which mistake you can better afford.
Which error is worse? It depends on the stakes
This is the part that separates textbook knowledge from professional judgment. There is no universal answer; the answer depends on what each mistake costs. Tap through five situations below.
Courtroom
The null hypothesis is “the defendant is innocent.” The jury either convicts (rejects H0) or acquits (keeps H0).
What it means. A Type I error here is convicting an innocent person, and the legal system deliberately builds walls against it: presumption of innocence, proof beyond reasonable doubt, unanimous juries. The price is that some guilty people go free (Type II). Societies choose which error hurts more, and courts choose to tolerate more Type II to avoid Type I.
Smoke alarm
The null hypothesis is “there is no fire.” The alarm sounds (rejects H0) or stays silent.
What it means. A false alarm (Type I) is annoying: burnt toast, a wasted evacuation. A missed fire (Type II) can be fatal. So alarm designers accept many false alarms to make sure real fires are almost never missed. The cut-off is set to be jumpy on purpose.
Medical screening
A screening test for a condition affecting 2% of people, with 90% sensitivity and 95% specificity, is used on 10,000 people.
What it means. Of 670 positive results, only 180 are real, about 27%, while 20 real cases slip through. Screening programs accept a fair number of false alarms because a follow-up test can clear them, whereas a missed case might not be found in time.
Spam filter
The null hypothesis is “this email is legitimate.” The filter flags it as spam (rejects H0) or delivers it.
What it means. A Type I error is a real email, maybe a job offer or an invoice, landing in spam. A Type II error is a junk email in the inbox. Most people forgive the second far more easily than the first, so filters are tuned to let some junk through rather than to lose real messages.
A/B test
You test a new checkout page. H0 says it does no better than the old one. Suppose you collect 100 orders and need 5% false-alarm protection.
What it means. The rule is to declare the new page better only if you see at least 59 wins in 100 comparisons when a fair coin would give 50. That keeps false alarms at 4.4%. If the new page truly wins 60% of the time, this test detects it only 62% of the time, so 38% of the time you would miss a real improvement. More data raises power without raising false alarms.
Notice the pattern. Courts guard hard against Type I errors (convicting the innocent). Smoke alarms guard hard against Type II errors (missing a fire). Spam filters lean cautious about Type I errors (losing real mail). Screening programs accept lots of false alarms to avoid misses. In every case the threshold is a moral and practical choice wearing the costume of a number.
A closer look at medical screening
Medical tests give the clearest illustration of why false alarms are common even when a test is good. Suppose a condition affects 2% of a population. A screening test correctly flags 90% of people who have it (sensitivity) and correctly clears 95% of people who do not (specificity). Sounds excellent. Now screen 10,000 people.
- 200 people have the condition. The test catches 180 and misses 20 (Type II errors).
- 9,800 people are healthy. The test wrongly flags 490 of them (Type I errors) and correctly clears 9310.
That gives 670 positive results, of which only 180 are real. If you get a positive result, the chance you actually have the condition is about 27%. This surprises almost everyone. It does not mean the test is bad. It means the condition is rare, so even a small false-alarm rate produces a big pile of false alarms in absolute terms. This is why doctors follow up screening with a second, more specific test, and why a single positive result is a reason to investigate rather than a diagnosis.
How to reduce both errors: more data
If moving the cut-off only swaps one error for another, how do you get better on both? The answer is to sharpen the picture itself. A larger sample makes the two bell curves narrower, so they overlap less. With less overlap, you can hold false alarms at 5% and still catch far more real effects.
Here is a concrete illustration. Suppose a real effect exists that is about 0.3 standard deviations in size, which is modest. With a one-sided test at α = 5%, power grows with sample size like this.
| Sample size | Power | Chance of missing it (β) |
|---|---|---|
| 10 | 24.3% | 75.7% |
| 25 | 44.2% | 55.8% |
| 50 | 68.3% | 31.7% |
| 75 | 83.0% | 17.0% |
| 100 | 91.2% | 8.8% |
| 150 | 97.9% | 2.1% |
| 200 | 99.5% | 0.5% |
A study with 10 subjects has a small chance of detecting this effect at all, and its “no significant difference” result would be nearly meaningless. It tells you the study was too small, not that the effect is absent. This is one of the most common misreadings of research: treating a failure to find something as proof that nothing is there.
Other ways to raise power include reducing noise in the measurements, using a more precise instrument, comparing matched pairs instead of independent groups, and looking for larger effects. Statisticians call planning this before a study “a power analysis,” and it is one of the cheapest ways to avoid wasting a research budget.
A worked example: the A/B test
Let me show you the numbers in a familiar business setting. You redesign your checkout page and compare it against the old one across 100 head-to-head comparisons. Under the null hypothesis, the new page is no better, so each comparison is like a fair coin flip and you would expect about 50 wins.
You decide in advance that you want no more than a 5% false-alarm rate. Working through the exact binomial probabilities, that means you declare the new page better only if it wins at least 59 of 100. Under a fair coin, the chance of hitting that bar by luck is 4.4%. That is your effective Type I error rate.
Now suppose the new page really is better and wins 60% of the time. How often would this test detect that? Only about 62% of the time. The other 38% of the time you would return “no significant difference” and shelve a page that really was better. That is a Type II error, and it is expensive because nobody ever finds out about the money that was left on the table.
The fix is not to loosen alpha. It is to collect more comparisons. If you had 400 comparisons instead of 100, the same 60% effect would be detected far more often while the false-alarm rate stays at 5%.
The multiple-testing trap: false alarms pile up
Here is a hazard that catches even experienced people. An alpha of 5% sounds low. But it applies to each test separately. If you run many tests on data where nothing is truly going on, the chance of at least one false alarm grows fast.
| Tests run | Chance of at least one false alarm |
|---|---|
| 1 | 5.0% |
| 5 | 22.6% |
| 10 | 40.1% |
| 20 | 64.2% |
| 50 | 92.3% |
Run 20 tests and you have a 64% chance of at least one “significant” result by pure luck. Run 50 and it is 92%. This is the mathematical heart of “p-hacking”: slicing the data many ways until something looks significant. It is also why a marketing dashboard with forty metrics will always show a few “winners.” The remedy is to decide your hypotheses in advance, limit the number of comparisons, or adjust the threshold using methods such as the Bonferroni correction, which divides alpha by the number of tests.
What the p-value has to do with all this
The p-value is the probability of seeing data at least as extreme as yours, assuming the null hypothesis is true. You compare it with your pre-chosen alpha: if p is below alpha, you reject the null. The p-value itself is not the probability that the null is true, and it is not the probability that you made an error on this particular test. The false-alarm rate belongs to the procedure across many uses, and alpha sets it.
A quick way to hold both ideas: alpha is the standard you set before the race; the p-value is the time you actually ran.
Real-world places these errors hide
A false alarm stops a traveller for a bag check. A missed threat is catastrophic. Systems lean toward false alarms.
Blocking a genuine card purchase annoys a customer (Type I). Missing a real theft costs money (Type II).
Approving a useless drug is a Type I error. Rejecting a useful one is Type II, and patients lose out.
Rejecting a good candidate is a Type II error. Hiring a poor one is Type I, if “good enough” is the null.
One more subtlety worth knowing: which error is called “Type I” depends on how you phrase the null hypothesis. If you flip the null and alternative, the labels flip. In practice, the null is the cautious default, and you should always state it out loud before you decide what each error means.
Six mistakes I keep seeing (tap to open)
1. Reading “not significant” as “no effect”
A non-significant result may simply reflect a small sample or noisy data. It means you did not find enough evidence, not that the effect is zero. Check the power.
2. Treating 5% as a law of nature
Alpha of 0.05 is a convention. In particle physics the standard is far stricter, and in early-stage screening a looser threshold can be sensible. Pick it based on the cost of each error.
3. Optimising one error and forgetting the other
Making a test extremely strict drives Type I errors near zero while Type II errors soar. A perfectly cautious test can be perfectly useless.
4. Running many tests and reporting the winners
Without a correction, false alarms accumulate. Decide the questions first, or adjust your threshold.
5. Mixing up the p-value with the error rate
The p-value belongs to the data you observed. Alpha is the false-alarm rate you are designing into the procedure.
6. Ignoring base rates
When the thing you are hunting is rare, even a low false-alarm rate produces more false alarms than real detections. The screening example shows how.
Quick quiz: test yourself
Tap each question to reveal the answer and its reasoning.
A pregnancy test says “not pregnant” but the person is pregnant. What kind of error is this?
- Type I
- Type II
- Both
- Neither
The null is “not pregnant.” Keeping it when it is false is a missed signal, a Type II error.
A fire alarm sounds because of burnt toast. Which error?
- Type I (false alarm)
- Type II (missed signal)
- Correct decision
- Power
Rejecting a true null (“no fire”) is a Type I error.
Which change lowers the chance of a Type II error without raising the Type I rate?
- Lowering the significance level
- Using a stricter cut-off
- Collecting a larger sample
- Ignoring the data
A larger sample sharpens the picture, raising power (1 − β) while α stays fixed.
If β = 0.20, what is the power of the test?
- 0.05
- 0.20
- 0.95
- 0.80
Power = 1 − β = 0.80.
You run 20 independent tests at α = 0.05 when nothing is real. About how likely is at least one false alarm?
- 5%
- About 64%
- 20%
- 100%
1 − 0.9520 = 64%. False alarms accumulate quickly across many tests.
Frequently asked questions
What is the difference between Type I and Type II errors?
A Type I error is a false positive: you reject a null hypothesis that is actually true. A Type II error is a false negative: you fail to reject a null hypothesis that is actually false.
How do I remember which is which?
Think of the boy who cried wolf. The first time he cried wolf with no wolf, that was a Type I error, a false alarm. The second time, when the wolf was real and nobody believed him, that was a Type II error, a missed signal.
What is the significance level (alpha)?
Alpha is the probability of a Type I error that you are willing to accept, commonly 5%. It is set before you look at the data.
What is statistical power?
Power is the probability of detecting a real effect, equal to 1 minus the probability of a Type II error. Researchers often aim for 80% power or more.
Can I reduce both errors at once?
Yes, by collecting more data or by reducing noise in your measurements. For a fixed sample, lowering one error raises the other.
Is the p-value the probability of a Type I error?
Not exactly. The p-value is the probability of seeing data at least this extreme if the null were true. Alpha, chosen in advance, is the false-alarm rate the procedure is designed to have.
The takeaway
A Type I error is a false alarm, a Type II error is a missed signal, and every real-world test trades one against the other. Decide which mistake costs more, set the threshold accordingly, and collect enough data that you are not forced to choose between two bad options.
Next time somebody proudly announces a “statistically significant” result, or a “no significant difference,” ask two questions: how many false alarms could this procedure produce, and how likely was it to catch a real effect in the first place?
