6.3 Every Converse Fails
The implications above run one way only. Each of the following examples settles one converse, and together they show the four modes are genuinely distinct rather than notational variants.
Example 6.3.1 (Distribution does not imply probability). Let \(X\sim N(0,1)\) and put \(X_n = X\) for every \(n\), with \(Y=-X\). By symmetry \(Y\sim N(0,1)\) too, so \(F_{X_n} = F_Y\) for every \(n\) and trivially \(X_n\underset {D}{\longrightarrow }Y\). But \[\left |X_n - Y\right | = \left |X - (-X)\right | = 2\left |X\right | ,\] which does not become small: \(P\left (\left |X_n-Y\right |\geq 1\right ) = P\left (\left |X\right |\geq \tfrac 12\right )\approx 0.617\) for every \(n\). So there is no convergence in probability.
Note. The example shows what convergence in distribution does and does not say. It is a statement about the laws of the variables, not about the variables themselves, and two variables can have identical laws while being as far apart as \(X\) and \(-X\). This is exactly why Slutsky’s theorem insists that one of its limits be a constant.
Example 6.3.2 (Probability does not imply quadratic mean). Let \[X_n = \begin {cases} n, & \text {with probability } \dfrac 1n,\\[4pt] 0, & \text {with probability } 1-\dfrac 1n.\end {cases}\] For any \(0<\varepsilon <1\), \(P\left (\left |X_n\right |\geq \varepsilon \right )=\frac 1n\rightarrow 0\), so \(X_n\underset {P}{\longrightarrow }0\). But \[E\left (X_n^{2}\right ) = n^{2}\cdot \frac 1n = n \longrightarrow \infty ,\] so there is no convergence in quadratic mean. The rare values are rare enough to vanish in probability and large enough to dominate the second moment.
Example 6.3.3 (Almost sure does not imply quadratic mean). Take \(X_n = n\) with probability \(\dfrac {1}{n^{2}}\) and \(0\) otherwise, independently across \(n\). Since \(\sum _n \frac {1}{n^{2}}<\infty \), the Borel–Cantelli lemma gives \(P\left (X_n\neq 0 \text { infinitely often}\right )=0\), so \(X_n\underset {a.s.}{\longrightarrow }0\). Yet \[E\left (X_n^{2}\right ) = n^{2}\cdot \frac {1}{n^{2}} = 1\] for every \(n\), which does not tend to zero.
Example 6.3.4 (Probability does not imply almost sure). Partition \([0,1]\) into blocks: one interval of length \(1\), then two of length \(\tfrac 12\), then four of length \(\tfrac 14\), and so on. Let \(X_n\) be the indicator of the \(n\)-th interval in this list, with \(U\) uniform on \([0,1]\). The lengths tend to zero, so \(P\left (X_n=1\right )\rightarrow 0\) and \(X_n\underset {P}{\longrightarrow }0\). But every point of \([0,1]\) lies in infinitely many of the intervals, so for every outcome the sequence \(X_n\) takes the value \(1\) infinitely often and never settles. There is no almost sure convergence anywhere.
Remark. Collecting the four examples, the picture is \[\begin {array}{ccc} \text {almost sure} & \Longrightarrow & \\ & \searrow & \\ & & \text {in probability} \ \Longrightarrow \ \text {in distribution}\\ & \nearrow & \\ \text {quadratic mean} & \Longrightarrow & \end {array}\] with no arrow reversible, and with almost sure and quadratic mean convergence neither implying the other in either direction. The only reversal permitted anywhere is the last one when the limit is a constant.
That exception is not a curiosity. It is what allows the weak law of large numbers to be proved by Chebyshev’s inequality — which delivers convergence in quadratic mean — and it is what makes the consistency of an estimator, a statement about a constant limit, checkable by whichever of the two is easier.
Questions on this section
Stuck on something here? Ask below and it stays attached to this topic.