The Bayesian Trap
Source: The Bayesian Trap, Veritasium, 10:36, uploaded 2017-04-05, category Other / Unclear, playlist index 1328.
Veritasium begins with a rare-disease test. A person feels slightly ill, receives a positive result for a disease that affects 0.1% of the population, and hears that the test identifies 99% of people who have the disease whilst falsely identifying 1% of people who do not. The tempting answer is a 99% chance of being ill. Bayes’ theorem gives 9%.
A positive result in its population
Bayes’ theorem asks for the probability that a hypothesis is true after an event occurs. In the medical example, the hypothesis is that the patient has the disease and the event is a positive test. The calculation starts with the prior probability, which describes how likely the disease was before the test, and multiplies it by the probability of a positive result when the disease is present. That product is divided by the total probability of a positive result, which includes true positives and false positives.
The prior is often the hardest part to set. In this case the disease’s frequency in the population gives a reasonable starting point: one person in a thousand. The test correctly identifies that one person. Among the remaining 999 people, its 1% false-positive rate produces about ten more positive results. A randomly selected positive result therefore comes from a group of eleven people in which one has the disease. The chance is 1 in 11, or about 9%.
The arithmetic makes the result easier to see than the formula. A test can be accurate for people who have a disease and still produce many more false positives than true positives when the disease itself is rare.
Evidence as repeated updating
Bayes first considered a thought experiment in which he sits with his back to a square table whilst an assistant throws a ball onto it. Bayes cannot see the ball’s position. The assistant throws another ball and reports whether it landed to the left, right, in front of, or behind the first. Each report narrows the possible position of the hidden ball. More throws make the estimate more accurate, although certainty never arrives.
The example describes a way of knowing. Bayes does not need to claim that reality is undetermined. He treats knowledge as an estimate that changes when new evidence arrives. The video says that Bayes abandoned the work for more than a decade, regarded it as unworthy of publication, and never submitted it to the Royal Society. After his death, relatives asked his friend Richard Price to examine the papers. Price found the theorem and prepared it for publication, then gave the same idea a different setting. A man emerges from a cave and sees the sun rise for the first time. One sunrise leaves open the possibility of a one-off event. Each sunrise after that raises his confidence that the sun behaves this way every day.
That repeated use matters in the disease example. A second positive test from another laboratory changes the first result into the prior for the next calculation. Assuming the tests are independent and keep the same rates, two positive results raise the chance of disease from 9% to about 91%. Two results from separate laboratories are unlikely to be chance alone. The updated probability still differs from the test’s reported 99% accuracy because the prevalence and the false-positive rate remain part of the calculation.
Bayesian filtering applies the same update to email. The words in a message provide evidence for or against the hypothesis that it is spam. The method estimates the probability of spam given those words, rather than treating one word as a verdict.
Priors and immovable beliefs
Bayes’ theorem updates a prior. It cannot choose the prior for us. That leaves room for people to begin with radically different degrees of certainty about the same claim. Someone who assigns a 100% prior probability will keep a 100% posterior probability after any evidence that the model can process. Someone who assigns 0% will remain at 0%. Nate Silver makes this point in The Signal and the Noise, which the video cites when it argues that debates between those two positions cannot change either person’s mind.
The point concerns the structure of the calculation rather than a licence to treat every belief as equally good. A prior can come from population data, past experience, a model, or a guess. The video says that some priors are no better than guesses. The result inherits that starting point, then changes as evidence accumulates.
Repeated outcomes and the trap
Veritasium’s concern goes beyond the usual claim that probability feels counterintuitive. People may become too comfortable with the Bayesian pattern. Repeated rejection, failure, or low pay can make a person update towards the belief that this is simply how the world works. The pattern starts to look like a property of nature rather than the result of a set of conditions.
The video quotes Nelson Mandela: “Everything is impossible until it’s done.” The Bayesian reading is precise. An event with no observed instances receives a prior close to zero, so it appears impossible until new evidence exists. The missing part of the calculation is that actions help determine the outcome. If a person treats a repeated result as fixed and keeps behaving in the same way, the behaviour can help produce the same result. The belief becomes a self-fulfilling prophecy.
The video description makes this qualification explicit. Repeated events may arise from an inevitable condition, or they may follow a sequence of steps that depends on what people do. Escaping the Bayesian trap therefore requires experiments that change the steps and produce evidence for a different outcome. The source leaves the experiment open rather than prescribing one. Its claim is that a long run of the same result can describe a stable arrangement of behaviour without proving that the result had to occur.
Limits
The mathematical explanations, historical account, figures, and quotations in this note follow Veritasium’s English captions and description. I have not independently checked the historical story about Bayes and Richard Price, the attribution to Pierre-Simon Laplace, the book’s account of Bayes’ work, or the numerical medical example against their underlying sources. The diagnostic calculation is an illustration of conditional probability, not medical advice. Its 9% and 91% results depend on the stated prevalence, test rates, and independence assumption.
The description names Nate Silver’s The Signal and the Noise and Sharon Bertsch McGrayne’s The Theory That Would Not Die: How Bayes’ Rule Cracked the Enigma Code, Hunted Down Russian Submarines, and Emerged Triumphant from Two Centuries of Controversy as useful references. The sponsorship segment and its call to Audible are omitted from the note.
Further reading / references
- Nate Silver, The Signal and the Noise, named in the video and description.
- Sharon Bertsch McGrayne, The Theory That Would Not Die: How Bayes’ Rule Cracked the Enigma Code, Hunted Down Russian Submarines, and Emerged Triumphant from Two Centuries of Controversy, named in the video and description.