1.3 Sufficiency
A sufficient statistic for a parameter \(\theta \) is a statistic that, in a certain sense, captures all the information about \(\theta \) contained in the sample. Any additional information in the sample besides the value of the sufficient statistic does not contain any more information about \(\theta \).
Theorem 1.3.1 (Sufficiency Principle). If \(T(X)\) is a sufficient statistic for the model \(\, \{f_{\theta }(x)\, :\, \theta \in \Omega \}\,\) then any inference about \(\theta \) should depend on the sample only through the value of \(T(X)\) i.e if \(x\) and \(y\) are two sample points such that \(T(x) = T(y)\), then the inference about \(\theta \) should be the same whether \(X = x\,\) or \(\, Y = y\,\) is used.
Definition 1.3.2. A statistic is sufficient for the model \(\, \{f_{\theta }(x)\, :\, \theta \in \Omega \}\,\) if the distribution of the data \(\, X_1\, , \, X_2\, , \, \cdots \cdots \, , X_n\,\) given \(T = t\) does not depend on the unknown parameter \(\theta \).
Example 1.3.3. Let \(\, X_1\, , \, X_2\, , \, \cdots \cdots \, , \, X_n\,\) be a random sample from the \(POI(\theta )\) distribution. Show that \(\, T = \sum ^n_i X_i\, \)is sufficient for this model.
Solution. \(\displaystyle {f_{\theta }(x) = \frac {e^{-\theta }\, \theta ^x}{x!}}\hspace {1cm} x = 0\, , \, 1\, , \, 2\, ,\, \cdots \cdots \cdots \) \begin {align*} f_{\theta }(x_1\, , \, x_2\, , \, \cdots \cdots \, , \, x_n) \, & = \, \prod ^n_{i = 1}\, \frac {e^{-\theta }\, \theta ^{x_i}}{x_i!}\, = \, \frac {e^{n \theta }\, \theta ^{\sum x_i}}{\prod ^n_{i= 1}\, x_i!}\,\hspace {1.4cm} x_i = 0\, , \, 1\, , \, 2\, , \, \cdots \end {align*}
\[T \, = \, \displaystyle {\sum ^n_{i =1}X_i}\, \thicksim \, POI(n\theta )\]
\[f_{\theta }(t) = \, \frac {e^{-n\,\theta }\, \left (n\, \theta \right )^t}{t!}\hspace {1.5cm} t = 0\, , \, 1\, ,\, 2\, , \, \cdots \cdots \cdots \]
\begin {align*} f\left (x/T = t\right ) & = \, P\left (X_1 = x_1\, , \, \cdots \cdots \, X_n = x_n\big /\, T = t\right )\\\\ & = \, \frac {P\left (X_1 = x_1\, , \, \cdots \cdots \, , \, X_n = n\, , \, \sum ^n_{i = 1}X_i = t\right )}{P_{\theta }\left (T = t\right )}\\\\ & = \, \frac {e^{-n\,\theta }\, \theta ^{\sum x_i}}{\prod \,x_i!}\, \times \, \frac {t!}{e^{-n\,\theta }\, \left (n\theta \right )^t}\hspace {1cm}\text {if}\hspace {0.5cm} \sum ^n_{i = 1} x_i = t \end {align*}
\begin {align*} & = \, \frac {\theta ^{\sum x_i}}{\prod ^n_{i = 1}\, x_i!}\, \times \, \frac {t!}{n^t\, \theta ^t}\hspace {1.8cm}\text {if}\hspace {0.5cm} \sum ^n_{i = 1} x_i = t\\\\ & = \, \frac {t!}{\prod ^n_{i = 1}\, x_i!}\, \left (\frac {1}{n}\right )^t\\\\ & = \, \frac {t!}{x_1!\, x_2!\, \cdots \cdots \, x_n!}\, \left (\frac {1}{n}\right )^t\hspace {0.3cm}, \hspace {1.5cm}\text {if}\hspace {0.2cm} \sum ^n_{i = 1} x_i = t \end {align*}
\[\text {i.e}\hspace {1.5cm} X\big /\, T = t\, \thicksim \, MULT\left (t\, , \, \frac {1}{n}\, , \, \cdots \cdots \, \frac {1}{n}\right )\]
Since \(f\left (x\, /\, T = t\right )\) does not depend on \(\theta , {T = \sum ^n_{i = 1}X_i}\) is a sufficient statistics for this model. □
Example 1.3.4. Let \(\, X_1\, , \, X_2\, , \, \cdots \cdots \, , \, X_n\,\) be a random sample from \(\, X\thicksim N(\theta \, , \, \sigma ^2)\,\) where \(\sigma ^2\) is known. Show that \(\, T = \sum ^n_{i = 1} X_i \, \) is sufficient for \(\, \theta \).
Solution. \(\displaystyle {f_{\theta }(x) = \frac {e^{-1/2\sigma ^2\left (x - \theta \right )^2}}{\sqrt {2\, \pi \, \sigma ^2}}}\hspace {1cm},\hspace {1.5cm} -\infty < x < \infty \)
\begin {align*} f_{\theta }(x_1\, , \, \cdots \cdots \, , \, x_n) \, & = \, \prod ^n_{i = 1}\, \frac {e^{-1/2\sigma ^2\left (x_i - \theta \right )}}{\sqrt {2\, \pi \, \sigma ^2}} = \frac {e^{-1/2\sigma ^2\, \sum ^n_{i = 1}\left (x_i - \theta \right )^2}}{\left (2\pi \, \sigma ^2\right )^{n/2}} \end {align*}
\[ T = \sum ^n_{i = 1} X_i \, \thicksim \, N\left (n\theta \, , \, n\sigma ^2\right )\]
\[f_{\theta }(t) = \, \frac {e^{\frac {-1}{2n\sigma ^2}\, (t - n\theta )^2}}{\sqrt {2\pi n\sigma ^2}}\]
\begin {align*} f(x_1\, , \, x_2\, , \, \cdots \cdots \, , x_n/\, T = t) \, & = \, \frac {f_{\theta }(x_1\, , \, x_2\, , \, \cdots \cdots \, , \, x_n)}{f_{\theta }(t)}\hspace {1.5cm}\text {if}\hspace {0.5cm} \sum x_i = t \end {align*}
\begin {align*} & = \, \frac {\frac {e^{\frac {-1}{2\sigma ^2}\sum (x_i - \theta )^2}}{\left (2\pi \, \sigma ^2\right )^{n/2}}}{\dfrac {e^{\frac {-1}{2n\sigma ^2}(t -n\theta )^2}}{\sqrt {2\pi n\sigma ^2}}}\\\\ & = \, \frac {\left (2\pi n\sigma ^2\right )^{1/2}\, e^{\frac {-1}{2\sigma ^2}\left (\sum x_i^2 - 2\theta \sum x_i + n\theta ^2\right )}}{\left (2\pi \, \sigma ^2\right )^{n/2}\, e^{\frac {-1}{2n\sigma ^2}\left (t^2 - 2nt\theta + n^2\theta ^2\right )}}\\\\ & = \, \frac {\sqrt {n}}{\left (\sqrt {2\pi \, \sigma ^2}\right )^{n - 1}}\, e^{\frac {-1}{2\sigma ^2}\left (\sum x_i^2 - t^2\right )}\hspace {1.5cm}\text {if}\hspace {0.5cm} \sum x_i = t\\\\ & = \, \frac {\sqrt {n}}{\left (\sqrt {2\pi \, \sigma ^2}\right )^{n - 1}}\, e^{\frac {-1}{2n\sigma ^2}\, n\left (\sum x_i^2 - \frac {t^2}{n}\right )}\hspace {1.5cm} \text {if}\hspace {0.5cm} \sum x_i = t \end {align*}
which does not depend on \(\theta \). Therefore \(\, {T = \sum ^n_{i = 1} X_i}\,\, \) is sufficient for \(\theta \). □
Problem 1.3.1. Let \(\, X_1\, , \, X_2\, , \, \cdots \cdots \, , \, X_n\,\) be a random sample from \(X\thicksim BIN(1,\theta )\,\) and let \(\, T = \sum ^n_{i = 1} X_i\)
- (a).
- Show that \(T\) is a sufficient statistic for this model.
- (b).
- Explain how you would generate data with the same distribution as the original data using only the value of the sufficient statistic.
- (c).
- Let \(\, U = \begin {cases} U\left (X_1\right ) = 1\, & \text {if}\hspace {0.3cm} X_1 = 1\\ 0 & \text {otherwise}\\ \end {cases} \)
Find \(\, E_{\theta }\left (U\right )\, \) and \(\, E\left (U \mid T = t\right )\)
Show solution
Solution.
(a)
The joint probability function is \[f_{\theta }(x_1,\dots ,x_n) = \prod _{i=1}^{n}\theta ^{x_i}(1-\theta )^{1-x_i} = \theta ^{t}(1-\theta )^{n-t},\qquad t=\sum _{i=1}^{n}x_i ,\] which depends on the data only through \(t\). Taking \(g(t,\theta )=\theta ^{t}(1-\theta )^{n-t}\) and \(h(x)=1\), the factorisation criterion gives sufficiency of \(T\).
(b)
Given \(T=t\), every arrangement of \(t\) ones among \(n\) positions has the same conditional probability, namely \(1/\binom {n}{t}\): the conditional probability of a particular \(x\) with \(\sum x_i=t\) is \[\frac {\theta ^{t}(1-\theta )^{n-t}}{\binom {n}{t}\theta ^{t}(1-\theta )^{n-t}} = \frac {1}{\binom {n}{t}},\] free of \(\theta \). So it is enough to choose a subset of \(\{1,\dots ,n\}\) of size \(t\) uniformly at random and put a one in those positions and a zero elsewhere. The result has exactly the distribution of the original sample, and \(\theta \) was never needed. That is what sufficiency means operationally: \(T\) carries all the information about \(\theta \), and the rest of the data can be regenerated by randomisation alone.
(c)
Here \(U=X_1\), so \[E_{\theta }(U) = P_{\theta }\left (X_1=1\right ) = \theta .\] For the conditional expectation, \(U\) is an indicator, so \[E\left (U \mid T=t\right ) = P\left (X_1=1 \mid T=t\right ) = \frac {P\left (X_1=1,\ \sum _{i\geq 2}X_i = t-1\right )}{P(T=t)} = \frac {\theta \,\binom {n-1}{t-1}\theta ^{t-1}(1-\theta )^{n-t}} {\binom {n}{t}\theta ^{t}(1-\theta )^{n-t}} = \frac {\binom {n-1}{t-1}}{\binom {n}{t}} = \frac {t}{n}.\] So \(E\left (U\mid T\right ) = T/n = \overline {X}\).
Note. This is the Rao–Blackwell theorem in miniature. \(U=X_1\) is unbiased for \(\theta \) but throws away \(n-1\) observations; conditioning it on the sufficient statistic returns \(\overline {X}\), which is unbiased, has variance \(\theta (1-\theta )/n\) instead of \(\theta (1-\theta )\), and is the UMVUE.
To use the above definition, we must guess a statistic \(T(X)\) to be sufficient, however, the next theorem allows us to find a sufficient statistic by simple inspection of the pdf.
1.3.1 The Factorisation Criterion
Theorem 1.3.5. Suppose \(X\) has probability (density) function \(\, \{f_{\theta }(x)\, :\, \theta \in \Omega \}\,\) and \(\, T(X)\,\) is a statistic. Then \(\, T(X)\,\) is a sufficient statistic for \(\, \{f_{\theta }(x)\, :\, \theta \in \Omega \}\,\) if and only if there exist two non negative functions \(g(\cdot )\) and \(h(\cdot )\) such that \(\, f_{\theta }(x) = g\left (T(x);\, \theta \right )\, h(x)\,\) for all \(\, x \, , \, \theta \in \Omega \).
Proof. “Discrete Case”
Suppose \(T\) is a sufficient statistic for the model, \begin {align*} f_{\theta }(x)\, & = \, P\left (X_1 = x_1\, ,\, X_2 = x_2\, , \, \cdots \cdots \cdots \, , X_n = x_n\right )\\\\ & = \, \underbrace {P\left (X_1 = x_1\, , \, X_2 = x_2\, , \, \cdots \cdots X_n = x_n\, , \, T = t\right )\,}_{h(x)}\, \underbrace {P(T = t)}_{g(x)} \end {align*}
\(h(x)\) is a function of \(x\) and does not depend \(\theta \) by definition and \(g(x)\) depends only \(\theta \) and \(x\).
“Only if”
Suppose that \(\, f_{\theta }(x) = g\left (T(x);\, \theta \right )\, h(x)\,\) for all \(x\, , \, \theta \in \Omega \). We need to show that \(f(x/ T = t)\) does not depend on \(\theta \). \begin {align*} P_{\theta }\left (T = t\right ) \, & = \, \sum _{x:\, T(x) = t} \, g\left (T(x);\, \theta \right )\, h(x)\, = \, g\left (T(x);\, \theta \right )\, \sum _{x:\, T(x) = t}\, h(x) \end {align*}
\begin {align*} \therefore \hspace {0.5cm} f\left (x/\, T = t\right ) \, & = \, \frac {P_{\theta }(X = x)}{P(T = t)}\hspace {1.5cm} \text {if}\hspace {0.5cm} T(x) = t\\ & = \, \frac {g\left (T(x); \theta \right )\, h(x)}{g\left (T(x);\, \theta \right )\, \sum _{x:\, T(x) = t} \, h(x)} \hspace {1.5cm} \text {if}\hspace {0.5cm} T(x) = t\\ & = \, \frac {h(x)}{\sum _{x:\, T(x) = t}\, h(x)} \end {align*}
which does not depend on \(\, \theta \).
\(\therefore \hspace {0.3cm}T(X)\,\) is a sufficient statistic. □
Example 1.3.6. Let \(\, X_1\, , \, X_2\, , \, \cdots \cdots \, , \, X_n\,\) be a random sample from the \(N(\mu , \sigma ^2)\) distribution. Show that \(\, T = \left ( \overline {X}\, , \, S^2\right )\, \) is a sufficient statistic for this model.
Solution. \(\displaystyle {f_{\theta }(x) = \, \frac {1}{\sqrt {2\pi \sigma ^2\,}}\, e^{\frac {-1}{2\sigma ^2}(x -\mu )^2}}\hspace {1cm}, \hspace {1cm}\theta = \left (\mu , \sigma ^2\right )\)
\begin {align*} f_{\theta }(x_1\, , \, x_2\, , \, \cdots \cdots \, , \, x_n) & = \, \prod ^n_{i = 1}\, \frac {1}{\sqrt {2\pi \sigma ^2\,}}\, e^{\frac {-1}{2\sigma ^2}(x_i - \mu )^2}\\\\ & = \, \frac {1}{\left (2\pi \sigma ^2\right )^{n/2}}\,e^{\frac {-1}{2\sigma ^2}\sum ^n_{i=1}(x_i -\mu )^2}\\\\ & = \, \frac {1}{\left (2\pi \sigma ^2\right )^{n/2}}\, e^{\frac {-1}{2\sigma ^2}\left (\sum ^n_{i=1}x_i^2 - 2\mu \sum x_i + n\mu ^2\right )}\\\\ & = \, \underbrace {\dfrac {1}{\left (2\pi \right )^{n/2}}}_{h(x)}\,\, \underbrace {\dfrac {1}{\left (\sigma ^2\right )^{n/2}}\, e^{\frac {-1}{2}\, \left (\sum _i x_i^2 - 2\mu \sum x_i + n\mu ^2\right )}}_{g\left (T(x);\, \theta \right )} \end {align*}
i.e \(\, {T_1(X) = \left (\sum ^n_{i = 1}X_i\, , \, \sum ^n_{i = 1} X_i^2\right )}\,\,\) is a sufficient statistic for this model.
\begin {align*} f_{\theta }(x_1\,, \, x_2\, , \, \cdots \cdots \, , \, x_n) & = \frac {1}{\left (2\pi \sigma ^2\right )^{n/2}}\, e^{\frac {-1}{2\sigma ^2}\sum ^n_{i = 1}\left (x_i - \overline {x} + \overline {x} - \mu \right )^2}\\\\ & = \, \frac {1}{\left (2\pi \sigma ^2\right )^{n/2}}\, e^{\frac {-1}{2\sigma ^2}\sum ^n_{i = 1}\left [(x_i - \overline {x})^2 + \underbrace {2(x_i - \overline {x})(\overline {x} - \mu )}_0 + (\overline {x} - \mu )^2\right ]}\\\\ & = \, \frac {1}{\left (2\pi \sigma ^2\right )^{n/2}}\, e^{\frac {-1}{2\sigma ^2}\left (\sum ^n_{i = 1} (x_i - \overline {x})^2 \, + \, n(\overline {x} - \mu )^2\right )}\\\\ & = \, \frac {1}{\left (2\pi \sigma ^2\right )^{n/2}}\, e^{\frac {-1}{2\sigma ^2}\left ((n - 1)S^2 \, + \, n(\overline {x} - \mu )^2\right )}\\\\ & = \, \underbrace {\frac {1}{\left (2\pi \right )^{n/2}}}_{h(x)}\,\, \underbrace {\frac {1}{\left (\sigma ^2\right )^{n/2}}\, e^{\frac {-1}{2\sigma ^2}\left ((n - 1)S^2\, + \, n(\overline {x} - \mu )^2\right )}}_{g\left (T(x);\, \theta \right )} \end {align*}
i.e \(\, \, T(x) = \left (\overline {X}\, , \, S^2\right )\, \) is a sufficient statistic. □
Note. Sufficient statistics are not unique.
Example 1.3.7. Let \(\, X_1\, , \, X_2\, , \, \cdots \cdots \, , \, X_n\,\) be a random sample from the \(\, WEI(1,\theta )\,\) distribution.
Find a sufficient statistic for this model.
Solution. \(\displaystyle {f_{\theta }(x) = \theta \, x^{\theta - 1}\, e^{-x^{\theta }}}\) \begin {align*} f_{\theta }(x_1\, , \, x_2\, , \, \cdots \cdots \, , \, x_n) \, & = \, \theta ^n\, \left (\prod \, x_i\right )^{\theta - 1}\, e^{-\sum ^n_{i = 1} x_i^{\theta }}\\\\ & = \, \underbrace {\theta ^n\, \left (\prod ^n_{i = 1}\, x_{(i)}\right )^{\theta }\, e^{-\sum ^n_{i = 1} x_i^{\theta }}}_{g\left (T(x);\, \theta \right )}\,\, \underbrace {\left (\prod ^n_{i = 1}x_i\right )^{-1}}_{h(x)} \end {align*}
\(T(X) = \left (X_{(1)}\, , \, X_{(2)}\, , \, \cdots \cdots \, , \, X_{(n)}\right )\,\) is a sufficient statistic for this model. □
Example 1.3.8. Let \(\, X_1\, , \, X_2\, , \, \cdots \cdots \, , \, X_n\,\) be a random sample from the \(UNIF(0,\theta )\) distribution. Find a sufficient statistics for \(\theta \).
Solution. \(\displaystyle {f_{\theta }(x) = \frac {1}{\theta }}\hspace {1cm}, \hspace {1cm} 0 < x < \theta \)
\begin {align*} f_{\theta }(x_1\,, \, x_2\, , \, \cdots \cdot \, , \, x_n ) \, & = \, \frac {1}{\theta ^n}\hspace {1.5cm}, \hspace {0.5cm} 0 < x_1\, \cdots \cdots \, x_n < \theta \\\\ & = \, \frac {1}{\theta ^n}\hspace {1.5cm} , \hspace {0.5cm} 0 < x_{(1)} < x_{(n)} < \theta \\\\ & = \, \frac {1}{\theta ^n}\hspace {1.5cm}, \hspace {0.5cm} x_{(1)} > 0\hspace {0.5cm} \text {and}\hspace {0.5cm} x_{(n)} < \theta \\\\ & = \, \frac {1}{\theta ^n}\, I\left (x_{(1)} > 0\right )\, I\left (x_{(n)} < \theta \right )\\\\ & = \, \underbrace {\frac {1}{\theta ^n}\, I\left (x_{(n)} < \theta \right )}_{g\left (T(x);\, \theta \right )}\, \, \underbrace {I\left (x_{(1)} > 0\right )}_{h(x)} \end {align*}
\(\therefore \, \, T(X) = X_{(n)}\,\) is sufficient statistic for \(\theta \). □
Note. \[I(x\in A) = \begin {cases} 1 & \text {if}\hspace {0.3cm} x\in A\\\\ 0 & \text {otherwise}\\ \end {cases}\]
Problem 1.3.2. Let \(\, X_1\, , \, X_2\, , \, \cdots \cdots \, , \, X_n\,\) be a random sample from \(X\thicksim N(\mu , \sigma ^2)\,\) with \(\mu \) known. Find a sufficient statistic for \(\sigma ^2\).
Show solution
Solution. With \(\mu \) known the joint density is \[f_{\sigma ^{2}}(x) = \left (2\pi \sigma ^{2}\right )^{-n/2} \exp \left \{-\frac {1}{2\sigma ^{2}}\sum _{i=1}^{n}\left (x_i-\mu \right )^{2}\right \}.\] The data enter only through \(\sum _{i=1}^{n}(x_i-\mu )^{2}\), so with \[g(t,\sigma ^{2}) = \left (2\pi \sigma ^{2}\right )^{-n/2}e^{-t/2\sigma ^{2}}, \qquad h(x)=1,\] the factorisation criterion gives that \[T(X) = \sum _{i=1}^{n}\left (X_i-\mu \right )^{2}\] is sufficient for \(\sigma ^{2}\).
Note. Note that \(T\) is not \( (n-1)S^{2}\). Knowing \(\mu \) means the deviations are taken about the true mean rather than about \(\overline {X}\), and \(T/\sigma ^{2}\sim \chi ^{2}_{(n)}\) rather than \(\chi ^{2}_{(n-1)}\) — one degree of freedom is recovered, which is exactly the information that would otherwise have been spent estimating \(\mu \).
Problem 1.3.3. Let \(\, X_1\, , \, X_2\, , \, \cdots \cdots \, , \, X_n\,\) be a random sample from the \(EXP(1,\theta )\) distribution. Find a sufficient statistic for this model.
Show solution
Solution. The EXP\((1,\theta )\) density is \(f_{\theta }(x)=e^{-(x-\theta )}\) for \(x>\theta \) and zero otherwise, so the joint density is \[f_{\theta }(x) = \prod _{i=1}^{n} e^{-(x_i-\theta )}\,I\left (x_i>\theta \right ) = \underbrace {e^{n\theta }\,I\left (x_{(1)}>\theta \right )}_{g\left (x_{(1)},\,\theta \right )}\ \cdot \ \underbrace {e^{-\sum _{i=1}^{n}x_i}}_{h(x)} .\] Every \(x_i\) exceeds \(\theta \) exactly when the smallest of them does, which is what turns \(n\) indicator functions into one. By the factorisation criterion \(T(X)=X_{(1)}\) is sufficient.
Remark. The parameter here is a threshold rather than a location in the usual sense, and the sufficient statistic is an order statistic rather than a sum. Whenever the support depends on the parameter, the indicator function carries the information and an extreme order statistic is what the factorisation produces.
Corollary 1.3.9. Let \(\, X_1\, , \, X_2\, , \, \cdots \cdots \, , \, X_n\,\) be a random sample from the model \(\, \{f_{\theta }(x)\, :\, \theta \in \Omega \}\). If \(T(X)\) is a sufficient statistic for the model then so is any one-to-one function of \(T(X)\).
Proof. Let \(\, T^*(X) = r \left (T(x)\right )\, \) where \(r\) is a one-to-one function with inverse \(r^{-1}\). Then \begin {align*} f_{\theta }(x) & = \, g\left (T(x), \theta \right )\, h(x)\\\\ & = \, g\left (r^{-1}\left (T^*(x)\right )\right )\, h(x)\\\\ & = \, g^*\left (T^*(x)\right )\, h(x) \end {align*}
where \(\, g^*\left (T(x), \theta \right ) = \, g\left [r^{-1}\, \left (T^*(x)\right )\right ]\)
\(\therefore \,\, \) by the factorisation Theorem \(T^*(x)\) is a sufficient statistic. □
Questions on this section
Stuck on something here? Ask below and it stays attached to this topic.