Bertrand’s Box Paradox: Why “It’s Obviously 50/50” Is Wrong

Bertrand's box paradox
Probability puzzles · Bayes · Everyday reasoning

Three boxes. Six coins. You reach into one at random and pull out gold. What is the chance the other coin in that box is gold as well? Almost everyone says one half. Almost everyone is wrong. The correct answer is two thirds, and the reason it surprises us says a lot about how human intuition handles evidence.

Reading time: about 15 minutesLevel: beginner friendlyTopic: conditional probability

The puzzle, exactly as it is usually told

The French mathematician Joseph Bertrand published this puzzle in 1889, and more than a century later it still catches smart people out. I have used it in workshops for years, with analysts, teachers and engineers, and the pattern is remarkably stable: about four in five people answer 50% within seconds and defend it with real conviction.

Here is the setup. There are three identical boxes, each with two drawers. Inside them:

GGBox 1: two goldSSBox 2: two silverGSBox 3: one of each
One box holds two gold coins, one holds two silver coins, and one holds a gold and a silver.

You choose a box at random, open one drawer at random, and see a gold coin. What is the probability that the other drawer in the same box also holds gold?

The answer in one line. It is 2/3, or about 66.7%. The tempting answer of 1/2 is wrong because the coin you saw is more likely to have come from the gold-gold box than from the mixed box.

Before I explain, I want you to feel why 50/50 is so seductive, because if you understand the temptation you will spot the same mistake in a medical report or a court case.

The argument that feels airtight

Here is the reasoning almost everybody uses. “I drew a gold coin, so this cannot be the silver-silver box. That leaves two boxes: gold-gold and gold-silver. They were equally likely at the start, so it is a coin toss whether I am holding the gold-gold box or the mixed one. Half the time the other coin is gold.”

Every sentence in that paragraph sounds fine. The first step is correct: silver-silver is out. The second step is where it quietly goes wrong. The two remaining boxes were equally likely before you saw a coin. But you did not just learn that the box is not silver-silver. You learned something more specific: a gold coin came out. And the two boxes are not equally good at producing gold coins.

  • The gold-gold box can only ever give you gold.
  • The gold-silver box gives you gold only half the time.

So the observation of gold is stronger evidence for the box that always produces it. That is the whole secret. Evidence does not just eliminate possibilities; it reweights the ones that survive.

The clean way to solve it: count coins, not boxes

When I teach this, I tell people to stop thinking about boxes altogether. Boxes are a distraction. The thing that is randomly chosen, in effect, is one of six coins, each equally likely to be the coin you touch. Label them and list their situations.

from GGGother side:Gfrom GGGother side:Gfrom SSSother side:Sfrom SSSother side:Sfrom GSGother side:Sfrom GSSother side:GOutlined = the drawn coin is gold. Three such cases; in two of them the other coin is gold too.(each of the six coins is equally likely to be the one you pull)
All six coins, the box each comes from, and what is on the other side of the drawer.

Now apply the evidence. You saw gold, so throw away every case where you drew silver. Three outlined cases remain:

  • Gold coin A from the gold-gold box. The other coin is gold.
  • Gold coin B from the gold-gold box. The other coin is gold.
  • The gold coin from the mixed box. The other coin is silver.

Three equally likely possibilities, and in two of them the other coin is gold. That gives 2/3. No formulas, no jargon, just counting the things that really are equally likely.

The reason I love this puzzle as a teaching tool is that the correct method is a habit rather than a trick: list the equally likely basic outcomes, cross out the ones that contradict what you saw, and count what is left. Once you adopt that habit, entire families of confusing probability questions become easy.

Do not trust me: run the experiment

Whenever a probability answer feels wrong, do not argue about it, simulate it. I wrote a short program that repeats the experiment: pick a box at random, open a random drawer, keep the run only if the coin is gold, and record whether the other coin was also gold. Here are the results, at increasing numbers of gold-first draws.

Gold-first drawsOther coin also goldShare
10550.0%
1006262.0%
1,00066666.6%
10,0006,64066.4%
100,00066,68166.7%
50% (the tempting answer)66.7% (the true answer)50.0%1062.0%10066.6%1,00066.4%10,00066.7%100,000Share of gold-first draws where the other coin was also gold
The share settles near 66.7% as the number of draws grows, and stays well clear of 50%.

At 10 draws the number can bounce around, because small samples are noisy. At 1,000 draws it sits close to two thirds, and by 100,000 it is unmistakable. This is the law of large numbers doing what it does. If you want to convince a stubborn colleague, a simulation is more persuasive than any argument, because it removes the feeling that you are just trying to trick them with words.

You can even do a physical version at home. Put two gold and one silver stickers on a few index cards, or use red and white cards, and repeat it about fifty times. You will see about two thirds emerge, not exactly, but clearly.

The same idea, five different costumes

Bertrand’s puzzle is not really about coins. It is about how new information changes probabilities when the information arrives through a process that favours some cases over others. Once you have the pattern, you will start noticing it everywhere. Tap through the tabs below. Each one is a real, well-known version of the same reasoning, with the numbers worked out.

Three cards

A hat holds three cards: one red on both sides, one white on both sides, one red on one side and white on the other. You draw one, look at one face at random, and it is red. What is the chance the other face is red?

Cards3
Red faces3
Red-red faces2
Answer2/3

Working it out. Count the red faces, not the cards. There are three red faces in the hat. Two of them belong to the double-red card and one belongs to the mixed card. Given that you are looking at a red face, there is a 2 in 3 chance that the back is red as well. Same structure as the boxes, just with paper and ink.

Monty Hall

You pick one of three doors. The host, who knows where the car is, opens a different door showing a goat and offers a switch. Should you switch?

Doors3
Stay wins1/3
Switch wins2/3
Edge2×

Working it out. Your first pick was right one time in three, and that does not change because the host opened a door. Since the host never opens the car door, the remaining 2/3 of probability piles onto the other closed door. It is Bertrand’s logic in a game-show costume: the host’s reveal is new information, but it is not information that treats all cases equally.

Two children

A family has two children. You learn that at least one is a boy. What is the chance both are boys?

Families4
At least 1 boy3
Two boys1
Answer1/3

Working it out. List the four equally likely families: BB, BG, GB, GG. “At least one boy” removes GG and leaves three, of which only BB has two boys. So the answer is 1/3, not 1/2. But if you are told the older child is a boy, only BB and BG remain and the answer really is 1/2. Small changes in wording change which cases survive, which is the point of the whole paradox.

Medical test

A screening test is 90% sensitive with a 9% false-positive rate for a condition that 1% of people have. You test positive. What is the chance you are actually sick?

Screened10,000
Sick100
Positives981
Truly sick9.2%

Working it out. Of 100 sick people, 90 test positive. Of 9,900 healthy people, 891 test positive by error. So of the 981 positives, only 90 are genuinely sick, about 9.2%. Most people, including many doctors in classic studies, guess something near 90%. It is the same trap: reading the reliability of the test as the probability of the condition.

Spam filter

An inbox gets 1,000 emails. 20 are phishing. A filter flags 95% of phishing and wrongly flags 5% of the rest. One email is flagged. How likely is it phishing?

Emails1,000
Phishing20
Flagged68
Real phish27.9%

Working it out. 19 phishing emails are flagged, and 5% of the 980 legitimate ones, which is 49, are flagged by mistake. That makes 68 flagged emails in total, and only 19 are phishing, roughly 28%. The filter is very good and still wrong about most of its flags because real phishing is rare. Whenever the thing you are hunting is rare, expect this.

Notice how each tab has the same three moves: list the equally likely cases, remove the ones contradicted by the evidence, and count what remains. The medical and spam examples add a fourth idea, that rare things stay rare even after a positive signal, which brings us to Bayes.

The Bayes view: updating your beliefs

Statisticians phrase all of this in terms of Bayes’ theorem. Do not let the name scare you. It is a rule for updating a belief when you get new evidence. Start with what you believed before (the prior), ask how likely the evidence is under each possibility (the likelihood), and rescale.

Posterior ∝ Prior × Likelihoodyour updated belief is proportional to what you thought before, times how well each option explains what you saw

For the boxes:

  • Prior: each box is chosen with probability 1/3.
  • Likelihood of drawing gold: gold-gold gives 1, gold-silver gives 1/2, silver-silver gives 0.
  • Multiply: 1/3 × 1 = 1/3 for gold-gold, 1/3 × 1/2 = 1/6 for gold-silver, 0 for silver-silver.
  • Rescale so the total is 1: gold-gold gets (1/3) / (1/2) = 2/3, gold-silver gets 1/3.

The two thirds falls right out. The gold-gold box started at one third and rose to two thirds because it was better at explaining the evidence. Same answer, different language. Some people find the coin-counting version more natural, others the Bayes version. I suggest learning both, because when the numbers get bigger, Bayes becomes the safer bookkeeping.

Where this goes wrong in real life

Medical screening

Consider a test for a condition that affects 1% of people. The test catches 90% of true cases and wrongly flags 9% of healthy people. You test positive. Intuition shouts that you are 90% likely to be sick. Let us count in a population of 10,000.

981 people test positive out of 10,000 screened90 truly sick (9.2%)891 healthy but flagged (90.8%)Assumes 1% prevalence, 90% sensitivity, 9% false-positive rate
Most positive results in a rare-condition screening come from healthy people.

Of the 100 sick people, 90 test positive. Of the 9,900 healthy people, 891 also test positive. So among 981 positives, only 90 are truly sick: about 9.2%. The test is not bad. The condition is just rare, so false alarms outnumber true detections. Doctors in famous studies have made this exact error, which is why good clinics follow a positive screening with a confirmatory test rather than reacting to the first result.

Fraud alerts and spam filters

Banks send you a “suspicious transaction” text. Most of the time it is a false alarm, because fraud is rare among millions of transactions. This does not mean the fraud system is broken. It means its precision is limited by the base rate, and it is a deliberate trade: a few annoying alerts for a lot of caught fraud.

Evidence in court

Lawyers call the mistaken version the prosecutor’s fallacy: taking “the chance of this evidence if the person were innocent is one in a million” and hearing it as “the chance the person is innocent is one in a million.” Those are two different conditional probabilities. Confusing them has contributed to real miscarriages of justice, and it is the same logical slip as reading “gold came out of the gold-gold box” as if it were “the box is gold-gold.”

Hiring and screening

A company designs a screening test that 95% of great candidates pass. Then it assumes anyone who passes is 95% likely to be great. If only a small percentage of applicants are great, most passers are not. Again the rate at which the thing occurs in the population is the piece people forget.

Why our brains keep choosing 50/50

Psychologists have a few explanations, and I find them all helpful when I am teaching this.

  • We count the visible options, not the weights. After the silver box is excluded, two boxes are visible, so we say half and half. The weights are invisible unless you deliberately look for them.
  • We confuse the box with the coin. The question is about boxes in our heads, but the random selection was really about coins.
  • We love symmetry. Two options often feel symmetrical even when they are not. The mind reaches for 50/50 when it is unsure.
  • We ignore how evidence was generated. A gold coin is more likely to appear from a box with more gold in it. Whenever a signal is easier to get from one hypothesis than another, seeing the signal shifts the odds.

Once you know these four traps, you can build a checklist. It is what I use before I trust any probability I have just worked out in my head.

The four-step checklist.

1. Write down the equally likely basic outcomes (not the tidy groups).
2. Remove the ones that contradict the evidence.
3. Count or weight what is left.
4. Ask whether the rate of the thing in the general population changes the answer.

Variations that test your understanding

A good way to make sure you really understand a puzzle is to change it slightly and predict what happens. Try these.

What if you are told only that the box is not silver-silver?

Then the two remaining boxes really are equally likely, and the chance that the box is gold-gold is exactly 1/2. The difference between this and the original is that you did not see a coin. The coin observation is what carries the extra weight. This variation is a fantastic reminder that how you learned something can matter as much as what you learned.

What if a coin is chosen at random from the gold coins?

Suppose someone gathers all three gold coins and hands you one at random, then asks whether its box-mate is gold. You would get the same 2/3, because you are effectively picking among the same three cases.

What if there are four boxes?

Add a second gold-silver box. Now there are four gold coins in total, two in the gold-gold box and two in the mixed boxes. Draw gold, and the chance the other coin is gold is 2 out of 4, or 1/2. The count changes, so the answer changes. This shows the method is flexible: nothing magical about two thirds, it is simply the result of the count.

What if the coin-drawing is not random?

If somebody peeked and deliberately showed you a gold coin whenever they could, the mechanism changes and the probabilities can change again. This is the same subtlety that makes the Monty Hall problem sensitive to the host’s rules. Always ask how did this evidence reach me?

A short history and why the name “paradox” is fair

Strictly speaking it is not a paradox in the sense of a contradiction; the mathematics is completely consistent. It is called a paradox because the correct answer clashes with a strong intuition. Bertrand himself used it to make a wider point: probability problems need to be stated carefully, and small changes in the description can change the answer. In his book he chose it partly to show that naive symmetry arguments can mislead.

The same family includes the Monty Hall problem, the two-child problem, and the Sleeping Beauty problem. If you enjoy Bertrand’s boxes, those are the natural next puzzles. They share a structure: some information is revealed, and the trap is to treat the remaining cases as equally likely when they are not.

Five mistakes people make with conditional probability

1. Treating the remaining cases as equally likely

After eliminating cases, do not assume what is left has equal weights. Check how likely each surviving case was to produce the evidence.

2. Mixing up P(A given B) with P(B given A)

The chance of a positive test given illness is not the chance of illness given a positive test. The two can differ enormously, especially when the condition is rare.

3. Ignoring the base rate

If the thing you are detecting is uncommon, even an accurate test will produce more false alarms than true finds. Always ask how common the condition is.

4. Forgetting how the evidence was produced

The same fact can mean different things depending on whether it was revealed at random or by someone who knew the answer. Monty Hall lives entirely in that difference.

5. Trusting intuition over a quick count

When two methods disagree with your gut, do the count or a simulation. It takes five minutes and it prevents embarrassing conclusions in a report.

Quick quiz: test yourself

Tap a question to reveal the answer and its reasoning.

In Bertrand’s box problem, you draw a gold coin. What is the chance the other coin in the same box is gold?
  1. 1/3
  2. 1/2
  3. 2/3
  4. 3/4

Three gold coins could be the one you drew. Two sit in the all-gold box. So 2 out of 3.

Why is the answer 1/2 tempting?
  1. Because the math is unfair
  2. Because it seems only two boxes remain and they look equally likely
  3. Because coins are random
  4. Because silver is impossible

After seeing gold, you can rule out the silver-silver box, which leaves two boxes. But the boxes are not equally likely to have produced a gold coin: the all-gold box has twice the chances.

Which change would make the answer exactly 1/2?
  1. Learning the box you picked is not all-silver, without seeing a coin
  2. Drawing two coins
  3. Adding a fourth box of silver
  4. Painting the coins

If all you know is that the box is not silver-silver, then the two remaining boxes are equally likely. It is the coin evidence that tilts the odds.

A test is 99% accurate but the disease affects 1 in 1,000 people. A positive result means the chance of disease is closest to:
  1. 99%
  2. About 9%
  3. 50%
  4. 1%

Per 100,000 people, 100 are sick and 99 test positive. Of the 99,900 healthy, about 999 test positive by mistake. So 99 of 1,098 positives are sick, about 9%. Base rates matter.

What is the general principle behind Bertrand’s paradox?
  1. Probabilities never change
  2. Always choose 50%
  3. Small samples are useless
  4. Condition on the evidence in terms of equally likely basic outcomes

List outcomes that are truly equally likely (here, the six coins), remove those that contradict the evidence, and count what remains.

Frequently asked questions

What is Bertrand’s box paradox?

A probability puzzle by Joseph Bertrand from 1889. Three boxes contain two gold, two silver and one of each. You pick a box at random, draw one coin, and it is gold. The chance that the other coin is also gold is 2/3, not the tempting 1/2.

Why is the answer 2/3 and not 1/2?

Because three gold coins could have been drawn, and two of them belong to the gold-gold box. Each coin was equally likely to be pulled, so the coin, not the box, is the right thing to count.

Is Bertrand’s box the same as the Monty Hall problem?

They share the same logic. Both involve information that arrives through a process that is not neutral, and both reward counting equally likely basic outcomes. Monty Hall adds a host who knows where the prize is.

Who was Joseph Bertrand?

A French mathematician (1822–1900) who published the puzzle in his book on probability in 1889. He was also known for the Bertrand paradox about random chords in a circle, which is a different puzzle.

How does this connect to Bayes’ theorem?

Bayes’ theorem updates a probability after new evidence. Here the prior chance of each box is 1/3, but a gold coin is twice as likely to come from the gold-gold box, so the posterior chance of that box becomes 2/3.

Where does this matter in real life?

Medical screening, spam filters, fraud alerts, court evidence and any situation where a positive signal is interpreted without considering how common the underlying thing is.

The takeaway

Bertrand’s box paradox is small enough to fit in your pocket and big enough to change how you read a headline. Evidence does not merely eliminate options; it reshapes how much each remaining option deserves to be believed. Count the equally likely basic outcomes, remove what the evidence rules out, and only then divide.

Next time someone says “it has to be fifty-fifty, there are only two possibilities,” you can smile and ask the veteran’s question: are the two possibilities really equally likely?

Bertrand’s box paradoxconditional probabilityBayes’ theoremprobability puzzlesMonty Hallbase rate fallacy

Comments

No comments yet. Why don’t you start the discussion?

Leave a Reply

Your email address will not be published. Required fields are marked *