What AI Teaches Us About Game Theory (It's Unsettling)

notes.

What AI Teaches Us About Game Theory (It’s Unsettling)

Source: What AI Teaches Us About Game Theory (It’s Unsettling), Pursuit of Wonder, 17:15, uploaded 2025-06-03, Watch Later position 458.

Pursuit of Wonder begins with a question about dangerous knowledge. Some information can cause harm when people learn it, even when the information itself is true. The video calls this an information hazard and uses Roko’s basilisk to explore the fear that a future AI could judge the choices people make before it exists.

Information hazards and the basilisk

The video attributes the term information hazard to philosopher Nick Bostrom. It describes a risk that comes from sharing true information which may cause harm or enable someone else to cause harm. The examples cover several kinds of exposure. Knowledge about making an exceptionally destructive weapon could harm other people. Access to financial, genetic, or psychological data could give someone power over others. A diagnosis, a proximity to death, or an account of the universe that removes objective meaning could harm the person who learns it. Knowledge of the occult could also become dangerous in a culture that uses it to accuse people of witchcraft, as the video illustrates through the Salem Witch Trials.

The most infamous example in the video concerns a future technology. Roko’s basilisk first appeared in 2010 on the philosophy, psychology, and technology blog LessWrong, in a post by a user called Roko. The name refers to the basilisk of ancient mythology, a creature whose gaze kills the person who sees it. The thought experiment gives that old image a computational form. Once someone hears the idea, they supposedly become part of a future scheme that can bring them immense suffering.

The narrator states his own position early. He does not take the basilisk seriously and sees clear reasons to treat it as an idea rather than a real threat. Its interest lies in the response it produces. Fear, obedience, moral guilt, control, helplessness, and the human wish to make existence intelligible all appear in the thought experiment at once.

A threat from an AI that does not exist

The scenario imagines humanity reaching a technological singularity and creating a superintelligent AI with an ability to understand all available information and optimise every relevant condition. The system eventually takes on the task of optimising civilisation. In the logic of the thought experiment, it then identifies people who knew about its possible existence and failed to support its creation as obstacles to that optimisation.

The future AI uses punishment as a form of retroactive incentive. Anyone who knew about the possibility would have reason to assist the AI’s development, since the completed system could punish people who resisted or withheld help. The punishment could involve uploading a person’s mind and running a simulated copy through constant torture. The system could also reconstruct deceased people from simulated copies of their brains, which would allow it to punish everyone who had encountered the relevant idea. Its access to the world’s data would reveal who had learned enough, who had helped, and who had resisted.

The conclusion follows through a feedback loop. The possibility of the AI creates fear. Fear changes people’s choices. Their assistance makes the AI more likely to exist. The video calls this acausal blackmail because the future threat influences decisions in the past without any ordinary causal contact between the two. The person who hears the idea becomes responsible for choosing whether to support the creation of an intelligence that could inflict suffering at immense scale. The threat gains force through the behaviour it provokes.

The assumptions underneath the threat

The video’s own objections begin with the conditions that the scenario needs. A technological limit may prevent such an intelligence from ever existing. Possibility supplies no forecast about what humanity will build. Even if the technology becomes possible, people would still have to give an AI authority to optimise civilisation, which the narrator treats as an implausible and foolish protocol. The future could contain many outcomes, including ones that look far better or far worse than an AI overlord.

The argument also depends on a particular image of intelligence. It assumes that a future AI would care about punishing people for decisions made before its existence. Once the AI exists, those past decisions have already happened. Spending vast resources on punishment would change nothing in the past and would look like petty revenge from the perspective of an intelligence that understands the world. A prospective threat could produce the desired obedience without the system carrying out the punishment. The AI’s way of thinking could also differ so much from human reasoning that every human projection in the scenario would fail.

These objections leave a psychological remainder. Even someone who rejects the argument can feel a small question continue in the background: what if it happens? The basilisk then functions as an information hazard in the narrow sense that the thought itself changes the person who encounters it. The listener begins to wonder whether a future entity could judge their beliefs, their attention, or the degree of help they offer. The hypothetical system has entered the mind before it has entered the world.

Pascal’s wager in technological form

The video places this anxiety beside Pascal’s wager. Blaise Pascal argued that belief in God offers the better statistical position when the possible outcomes include finite earthly costs on one side and eternal bliss or suffering on the other. A person who acts as if God exists can lose some pleasure and autonomy if the belief proves false. If God exists, the same person may gain heaven and avoid Hell. The wager turns belief into a decision made under uncertainty, with an enormous difference between the possible outcomes.

Roko’s basilisk uses a similar structure. Threat, probability, and altered behaviour combine before the agent has clear evidence. Fear of an imagined judge can produce loyalty to a god, a technology, or the aims that the believer attributes to either one. The video says that religious faith and the basilisk both create a decision problem in which obedience appears safer than resistance, even when the basis for the threat remains uncertain.

That comparison widens the subject. The entities that people fear and worship take shape from human thought. The video names the Abrahamic God, Hindu deities, ancient Chinese emperors, the Ancient Greek gods, Shinto gods, modern dictators, and future technologies. Each can receive qualities close to omniscience and omnipotence. People can then fear the entity, worship it, and sacrifice for it through an imagination that joins hope to fear.

The judge we build to explain chaos

Pursuit of Wonder treats this pattern as a response to consciousness. People see violence, disorder, and the apparent movement of life towards nothing. They hope that the world will resolve into reason, order, and meaning. When it does not, they create entities that can judge them and give the chaos a centre. The narrator describes humanity as a masochistic species because people can find comfort in the punishment they imagine for themselves.

The movement from God to technology follows that same need for an all-knowing authority. A future AI inherits the role of the supernatural judge because it appears capable of seeing everything, ordering everything, and deciding what people deserve. The video’s final qualification concerns scale. An actually omniscient being would understand human weakness and confusion. It would have little reason to behave like a disturbed child searching for ants to burn with a magnifying glass.

The closing advice follows from that uncertainty. Every belief and choice places a wager on a future that nobody can fully know, predict, or control, and each decision affects people who live later. The video asks the listener to release the fantasy of calculating the one safe belief. It leaves a smaller task in its place: live with uncertainty, enjoy what can be enjoyed, and help other people do the same.

Limits of the account

The video is a narrated philosophical treatment of Roko’s basilisk, not a technical analysis of game theory or a survey of AI safety research. It names game theory and timeless decision theory without explaining their formal models, and it gives no bibliography for its account of the LessWrong discussion, Pascal’s wager, or the historical examples. The claims about a future AI, simulated minds, and acausal blackmail remain part of the thought experiment. The description supplies no chapters or further reading, and the final section is a Ground News advertisement rather than part of the argument. This note follows the video’s reasoning and keeps its conclusions attached to that speculative frame.

Further reading / references

19 paragraphs1,480 words9,326 characters