Probability in the Quantum World
Probability in the Quantum World
Quantum mechanics does not predict what one atom will do, and this is often mistaken for a limitation. It is not: the theory predicts the frequencies to ten significant figures. This lesson sharpens the probability tools well enough to say that precisely — and then shows the one place where the quantum rules break away from the classical ones entirely.
Learning Objectives
After this lesson you will be able to:
- Compute probabilities by enumeration over equally likely outcomes, and combine them with the compound probability rules for "and" and "or", stating the conditions each rule requires.
- Recognize when events are dependent, and use the complement rule to turn a hard enumeration into an easy one.
- Quantify how far a finite run of shots should scatter from the predicted probability, using the binomial standard error, and decide how many shots an experiment needs.
- Distinguish classical randomness (ignorance about which member of an ensemble you hold) from quantum randomness (indeterminacy in an identically prepared state).
- Apply the Born rule to compute outcome probabilities from probability amplitudes, and explain why amplitudes, not probabilities, are what add over alternatives.
- Compute expectation values and uncertainties for spin measurements, and show that the classical projection law survives as the average of a two-valued outcome.
Intuition
The previous lesson left us with a theory that predicts probabilities and nothing more. That provokes a reasonable objection: if all we can say is "half the time this, half the time that," how can quantum mechanics be the most precisely tested theory in science?
The answer has two halves, and both are worth having straight before we do any calculating.
First, "random" does not mean "unpredictable". A fair coin is random, yet nobody doubts that a million flips will produce close to half a million heads, or that we can say how close. Randomness that obeys a known distribution is a strong form of knowledge, not the absence of it. The law is a prediction with no free parameters, and it is confirmed or refuted by counting.
Second, and more important: the quantities physicists measure to ten digits are usually not probabilities. They are energy differences, frequencies, and -factors — numbers that come out of the theory sharply and can be measured sharply. Probabilities are the expensive things to measure, because their error bars shrink only as . The randomness is real and irreducible; the precision is real too; they live in different parts of the theory.
So: probability first, for its own sake and done properly, and then the discovery that quantum probability follows a rule the classical calculus does not contain.
Theory
Counting equally likely outcomes
For a finite set of outcomes that are all equally likely, probability reduces to counting:
The whole content of the definition sits in the phrase equally likely, which is an assumption about the physical setup, not something the arithmetic can supply. It holds for a fair coin, a fair die, and a well-shuffled deck. It fails for a weighted die, and it fails in a way beginners find easy to miss: the outcomes of rolling two dice are equally likely as ordered pairs — all 36 of them — but the 36 pairs do not map onto 11 equally likely totals. There is one way to total 2 and six ways to total 7.
Example. Roll two dice; what is the probability the total is 7? The successful ordered pairs are — six of them, out of 36, so . A casino that pays 4 chips plus your stake on this bet is short-changing you: in 36 plays you expect 6 wins returning chips against 36 staked, a loss of 6. A fair payout would be 5 plus the stake, and the missing chip is the house edge [Fre, §1.3].
Compound probabilities: "and" multiplies, "or" adds
Enumeration becomes unmanageable quickly, so we combine simple probabilities instead. The rules of thumb are that "and" means multiply and "or" means add — but each carries a condition, and knowing the conditions is what separates using the rules from guessing with them.
For "or", addition requires the events to be mutually exclusive (they cannot both happen). In general
and the subtraction vanishes exactly when the overlap is empty. For "and", multiplication requires independence. In general
which reduces to precisely when , i.e. when learning tells you nothing about .
Example: the probability of flipping two heads is , agreeing with the enumeration . The probability of or is , since those two outcomes cannot both occur.
When independence fails
Here is the failure mode to keep in mind. Roll one die and ask for the probability that the result is even and greater than three. Enumeration: the successes are 4 and 6, so . But "even" has probability and "greater than three" has probability , and . The multiplication rule failed because the two statements are dependent — they constrain the same die. Knowing the roll is even raises the chance it exceeds three from to , and indeed . ✓
Change the question to "even on the first die and greater than three on the second," and the events are independent because they involve physically separate systems; now is correct. Independence is a physical claim about the setup, not a property of the sentences. This distinction returns with force in Term 1, where two spins that have interacted are exactly the case where it fails.
The complement trick
Since an event either happens or it does not, . So
and it is often far easier to enumerate the failures than the successes.
The classic demonstration: in a room of six people, what is the probability that at least two share a birth month? Enumerating the successes means tracking every way two, three, or more people can coincide — a mess. Enumerating the failure is one line: all six months distinct. The first person may have any month; the second must avoid one of twelve; the third must avoid two; and so on:
so . Nearly four times in five — far higher than intuition suggests, because the number of pairs among six people is 15, not 6. The same computation with 365 days and 30 people gives , the celebrated birthday problem.
How many shots? Fluctuations and the standard error
Every quantum prediction in this course is a probability, and every test of one is a finite run of shots, so we need to know how far the data should scatter.
If each of independent shots succeeds with probability , the number of successes is binomial, with
The absolute fluctuation grows with — more shots means noisier counts. But the estimated probability has standard error
which shrinks as . Both statements are true simultaneously and neither is a paradox: the counts spread out, the fractions converge. For a fair coin (), , so the interval is a two-standard-deviation band and catches about 95% of runs.
The practical consequence is the tax on probabilistic measurements: one more digit of precision costs a hundred times more shots. Getting to near needs . This is why the sharp numbers in quantum mechanics are frequencies and energy differences rather than probabilities.
Two kinds of randomness
We can now say precisely what changed between Lesson 1 and Lesson 2.
The classical Stern-Gerlach experiment was random, but only because the source was. Each atom left the oven with a definite moment pointing in a definite direction, followed a definite trajectory, and landed at a definite computable spot. Had we known the initial orientation, we could have predicted the landing point exactly. The distribution described classical randomness: our ignorance of which atom we had, an ensemble of definite-but-unknown cases.
The quantum experiment is not like that. In Experiment 3 every atom entering the second analyzer was prepared identically, by the same port of the same analyzer, and verified repeatable by Experiment 1. There is no residual ignorance to appeal to. Yet the outcomes still split 50–50. The randomness is a property of the state, not of our knowledge of it.
Both kinds appear in real experiments, often together, and the arithmetic for combining them is just compound probability — the rotating analyzer below is exactly that. But they are conceptually different, and much confusion about quantum mechanics comes from treating the second as though it were the first.
Amplitudes, not probabilities, are the fundamental objects
Here is the break with classical probability. In quantum mechanics the state does not carry probabilities; it carries complex probability amplitudes, and probability is what you get by squaring one at the very end. For a system in state and an outcome corresponding to , the amplitude is the overlap and the Born rule gives
Why does this matter? Because of what happens when an outcome can be reached by more than one route. Classically, the probability of reaching a destination by either of two mutually exclusive paths is the sum of the two path probabilities. Quantum mechanically, if the two routes are not distinguishable — if nothing in the world records which one was taken — you must add the amplitudes first and square afterwards:
The last term, the interference term, has no classical counterpart. It can be positive, negative, or zero, and it is the reason quantum probability is not simply probability theory with a quantum source of noise. When the paths are distinguishable — when a measurement records which was taken — the interference term disappears and the classical sum is restored.
Interference in the Stern-Gerlach interferometer
The cleanest demonstration uses hardware we already have. Prepare atoms in , send them into an SG analyzer, and consider two ways of finishing the experiment before a final SG measurement.
flowchart LR
P["prepare<br/>+z"] --> SPLIT["SG x<br/>splits the beam"]
SPLIT -->|"+x branch"| R{"recombine<br/>or block?"}
SPLIT -->|"-x branch"| R
R -->|"block the -x branch"| M1["SG z: 50 / 50"]
R -->|"recombine coherently"| M2["SG z: 100% plus"]Block one branch. The surviving atoms are in , so the final measurement gives . Half the atoms come out despite having been certified at the start — Experiment 4 of the previous lesson.
Recombine both branches coherently, with nothing recording which one each atom took. Now sum the amplitudes over the two routes:
so : every atom comes out , exactly as if the analyzer had never been there. (It is the completeness relation doing the work.) Had we added probabilities instead, we would have obtained — the wrong answer, and precisely the answer the blocked experiment gives.
So the difference between and is not a difference in the apparatus geometry: both setups split the beam identically. It is the difference between an alternative that is merely unobserved and one that is eliminated or recorded. That is the sharpest available statement of what measurement does, and Lesson 4 makes it a postulate.
Expectation values, and where the classical law went
Assign the value to a outcome and to a outcome. The expectation value is the probability-weighted average:
For an atom prepared along and measured along , substitute the Born-rule probabilities:
Look at what happened. Every individual measurement returns and never anything between — there is no atom whose deflection corresponds to . Yet the average follows the classical projection law of Lesson 1 exactly. The continuous classical answer did not vanish; it was demoted from a description of each atom to a description of the ensemble. This is the correspondence principle in one line, and it is why the classical picture worked so well for macroscopic magnets, which are averages over such outcomes.
The spread tells the rest of the story. Since the outcome is always , we have regardless of state, so
The uncertainty vanishes at and — the two cases where the outcome is certain — and is maximal at , where the atom is prepared along an axis perpendicular to the one measured. A sharp value on one axis is bought with maximal spread on the perpendicular one, which is the uncertainty principle arriving unannounced.
Compounding both kinds of randomness: the rotating analyzer
Finally, a calculation that uses every tool in this lesson at once [Fre, §1.4]. Prepare atoms in . Send them into a second analyzer that is rotating, snapping instantaneously between three stops — at , at , at — so that each arriving atom finds it at one of the three, with probability each.
Express it in words first, then convert: the atom exits if it is at and exits , or at and exits , or at and exits . "And" multiplies, "or" adds:
Half, as the three-fold symmetry of the stops suggests. Now break the mechanism so it only ever stops at or , each with probability :
larger than before, since we removed a stop that was unfavourable. Note the structure of these calculations: the 's and 's are classical probabilities describing our ignorance of where the machine happened to be, while the factors are quantum probabilities describing an indeterminacy no amount of extra knowledge would remove. Both enter the same compound-probability formula. Keeping track of which is which is what Term 1 formalizes as the difference between a statistical mixture and a superposition (1.5 Density Matrices).
Hands-on (Python)
import numpy as np
rng = np.random.default_rng(2)
# --- 1. Error bars on the cos^2(theta/2) law ----------------------------
print("theta N=100 N=10000 exact")
for deg in (0, 45, 90, 135, 180):
p = np.cos(np.deg2rad(deg) / 2) ** 2
cells = []
for N in (100, 10_000):
k = rng.binomial(N, p)
se = np.sqrt(p * (1 - p) / N)
cells.append(f"{k / N:.3f} +/- {se:.3f}")
print(f"{deg:5d} {cells[0]:18s} {cells[1]:18s} {p:.4f}")
# The N=100 estimates scatter by ~0.05; the N=10000 estimates by ~0.005.
# One more digit of precision costs a hundred times more shots.
# --- 2. Interference: recombine vs block --------------------------------
inv = 1 / np.sqrt(2)
plus_z, minus_z = np.array([1, 0], complex), np.array([0, 1], complex)
plus_x, minus_x = inv * (plus_z + minus_z), inv * (plus_z - minus_z)
# amplitudes for the two routes +z -> (+x or -x) -> +z
A_plus = np.vdot(plus_z, plus_x) * np.vdot(plus_x, plus_z)
A_minus = np.vdot(plus_z, minus_x) * np.vdot(minus_x, plus_z)
print(f"\ncoherent (add amplitudes, then square): {abs(A_plus + A_minus)**2:.3f}")
print(f"blocked/recorded (add probabilities): "
f"{abs(A_plus)**2 + abs(A_minus)**2:.3f}")
# 1.000 vs 0.500 -- the whole difference between an unobserved alternative
# and an eliminated one.
# --- 3. Rotating analyzer ------------------------------------------------
for name, weights in (("all three stops", [1/3, 1/3, 1/3]),
("broken (A, B only)", [1/2, 1/2, 0.0])):
stops = np.deg2rad([0, 120, 240])
p_plus = sum(w * np.cos(t / 2) ** 2 for w, t in zip(weights, stops))
print(f"{name:20s} P(+) = {p_plus:.4f}")Exercises
E1 (easy). Two fair dice are rolled. What is the probability that the total is 4 or 11? Solve by enumeration and check with the compound rules.
Solution
There are equally likely ordered pairs. Total 4: — three ways. Total 11: — two ways. The two totals are mutually exclusive, so we add: . Note that "4" and "11" are not equally likely outcomes; only the 36 ordered pairs are, which is why enumeration must be done at the level of pairs.
E2 (easy). Ten people are in a room. Estimate the probability that at least two share a birth month, using the complement rule. Why does the answer surprise most people?
Solution
requires ten distinct months out of twelve:
so — a near certainty. Intuition anchors on "ten people versus twelve months, so it's close," but the relevant count is the number of pairs, , each with a chance of matching. Once the number of pairs exceeds the number of categories, collisions become overwhelmingly likely.
E3 (medium). An atom is prepared in and is measured. Compute and , evaluate both at , and explain why reproduces the classical projection law even though no single measurement ever does.
Solution
, and since every outcome squares to , and .
At the mean is , a value no individual measurement can return — the outcomes are only , in the ratio . The classical law survives as a statement about the ensemble average, not about any atom; a macroscopic magnet is an average over such outcomes, which is why the classical law appeared exact for two centuries.
E4 (medium). In the Stern-Gerlach interferometer, suppose an experimenter inserts a detector that records which branch each atom takes, but does not block either branch — and then discards the record without reading it. What does the final SG measurement give: or ? Justify your answer in terms of amplitudes.
Solution
. What kills the interference term is that the two alternatives have become distinguishable in principle — the branch information now exists somewhere in the world, correlated with the atom. Whether a human reads it is irrelevant; the amplitudes for the two routes now belong to states that differ in the detector as well as in the atom, so they are orthogonal and cannot interfere: , leaving .
This is worth stating starkly because it is the single most common misconception about measurement: it is correlation with the environment, not human awareness, that converts a superposition into a mixture. Discarding the record does not restore coherence — you would have to erase it in a way that destroys the correlation.
E5 (hard). Return to the rotating analyzer with stops , , , always fed by a first analyzer oriented vertically. Compute for: (a) all three stops equally likely, but the incoming beam taken from the exit of the first analyzer; (b) a broken mechanism that stops only at or , each with probability , with a beam; (c) a mechanism that spends of its time at and each at and , with a beam. Comment on how (c) compares with the unbiased case.
Solution
For a beam the quantum factor is : at , at , at . For a beam the factors are the complements, : , , .
(a) . Same as the beam — unsurprising, since the three stops are symmetrically placed and the two input states are complementary.
(b) . Removing the favourable stop drops the yield from to .
(c) $P(+) = \tfrac23(1) + \tfrac16\left(\tfrac14\right) + \tfrac16\left(\tfrac14\right) = \tfrac23 + \tfrac1{12} = \tfrac34$.
Weighting the analyzer toward the aligned stop raises from to , and in the limit of always stopping at it would reach 1 — recovering Experiment 1. The general structure is : a classical average (the , reflecting ignorance of the machine's position) over quantum probabilities (the , reflecting genuine indeterminacy). No measurement of alone can separate the two contributions — which is exactly the ambiguity that the density matrix formalism is built to handle.
Checkpoint
- State the "and"/"or" rules for combining probabilities together with the condition each one requires. Give an example where the "and" rule fails.
- In a room of six people, why is it easier to compute the probability that no two share a birth month?
- Why do the absolute fluctuations in a count grow with while the error on the estimated probability shrinks? How many shots are needed to pin a probability near to ?
- What exactly is the difference between the randomness in the classical Stern-Gerlach experiment and the randomness in Experiment 3 of the previous lesson?
- In the Stern-Gerlach interferometer, what is when the branch is blocked, and when both branches are recombined coherently? Which arithmetic rule differs between the two cases?
Answers
- "Or" adds when the events are mutually exclusive (generally ); "and" multiplies when the events are independent (generally ). The "and" rule fails for "even and greater than three on one die": the truth is , not , because the two statements constrain the same roll.
- Because "no two share" is a single chain of constraints — each new person must avoid the months already used — while "at least two share" requires enumerating every pattern of coincidence. Take the complement: .
- The count has , which grows like ; the fraction divides by , giving . For at : (one standard error), or for a guarantee.
- Classically the randomness was ignorance of which definite initial orientation the atom had; in principle each trajectory was computable. Quantum mechanically the atoms entering the second analyzer are identically and completely prepared, so no ignorance remains, and the outcome is still random — indeterminacy in the state itself.
- Blocked: . Recombined coherently: . Blocking (or recording) makes the alternatives distinguishable, so probabilities add; leaving them indistinguishable means amplitudes add and are squared afterwards, and the interference term makes up the difference.
Further Reading
- [Fre] J. K. Freericks, Quantum Mechanics Done Right, §1.3 — the probability interlude this lesson follows, with an extended set of dice, birthday, and Penney's-game problems.
- [NC] Nielsen & Chuang, §2.2.3 — the Born rule and measurement postulate stated axiomatically.
- [Sak] Sakurai & Napolitano, §1.4 — expectation values, dispersion, and the route to the uncertainty relation.
- [Bil] Billingsley, Probability and Measure, Chs. 1 and 27 — the rigorous versions of the counting definition, independence, and the central limit theorem invoked here.
- 0.2.1 Probability Spaces and 0.2.2 Random Variables & Expectation — the program's own treatment of the classical machinery.
- 1.3.2 Expectation & Uncertainty — where and become operator statements.
← Prev: Quantum Stern-Gerlach Experiment · Up: Resources · Next: The Conundrum of Projections →