2.3 The Information Inequality

Suppose we consider estimating a parameter \(\tau (\theta )\), where \(\theta \) is a scalar, using an unbiased estimator \(T(X)\). A lower bound for the variance of \(T(X)\) can be found by the information inequality.

Theorem 2.3.1 (Information Inequality). Suppose \(T(X)\) is an unbiased estimator of the parameter \(\tau (\theta )\) in a regular statistical model \(\{f_{\theta }(x):\theta \in \Omega \}\). Then \[\var \left (T\right ) \geq \frac {\left [\tau '(\theta )\right ]^2}{J(\theta )}.\] Equality holds in and only if \(f_{\theta }(x)\) is regular exponential family with natural sufficient statistic \(T(X)\).

Note.

1.
If equality holds then \(T(X)\) is called an efficient estimator of \(\tau (\theta )\).
2.
The number \(\frac {\left [\tau '(\theta )\right ]^2}{J(\theta )}\) is called the Cramer- Rao lower bound (C.R.L.B).
3.
The ratio of the C.R.L.B to the variance of an unbiased estimator is called the efficiency of the estimator.

Proof. \(E(T) = \tau (\theta )\)

\[\int _A T(X)\, f_{\theta }(x)\, dx = \tau (\theta )\]

\[\frac {\partial }{\partial \theta }\int _A T(X)\, f_{\theta }(x)\, dx = \tau '(\theta )\]

\[\int _A T(X)\, \frac {\partial }{\partial \theta }f_{\theta }(x)\, dx = \tau '(\theta )\]

\[\int _A T(X)\, \frac {\frac {\partial }{\partial \theta }f_{\theta }(x)}{f_{\theta }(x)}\, \cdot \, f_{\theta }(x)\, dx = \tau '(\theta )\]

\[\int _A T(X)\, S(\theta , X)\, f_{\theta }(x) \, dx = \tau '(\theta )\] \[E\left [T(X)\, S(\theta ,X)\right ] = \tau '(\theta )\] \[\cov \left [T(X),S(\theta ,X)\right ] = E\left [T(X)\, S(\theta , X)\right ] - E\left [T(X)\right ]\, \underbrace {E\left [S(\theta ,X)\right ]}_0\] Therefore \(\, \cov \left [T(X), S(\theta , X)\right ] = \tau '(\theta )\).

By the Cauchy-Schwartz inequality \(\, \left [\cov (X,Y)\right ]^2 \leq \var (X)\, \var (Y).\) \[-1 \, \leq \, \frac {\cov (X,Y)}{\sqrt {\var (X)\, \var (Y)}}\, \leq 1.\] Therefore \[\left (\cov \left [T(X),S(\theta ,X)\right ]\right )^2 \leq \var (T(X))\, \cdot \, \var \left [S(\theta ,X)\right ]\hspace {0.3cm}\cdots \hspace {0.3cm} (2.1)\] \begin {align*} \left [\tau '(\theta )\right ]^2 & \geq \var (T(x))\, J(\theta )\\ \var (T(X))\, & \geq \, \frac {\left [\tau '(\theta )\right ]^2}{J(\theta )} \end {align*}

Equality in \((2.1)\) holds iff \(T\) and \(S(\theta ,X)\) are linear functions of each other. i.e \[S(\theta , x) = A_1(\theta ) \, T(x) + A_2(\theta )\] \[\int S(\theta , x)\, d\theta = \int \left (A_1(\theta )\, T(X) + A_2(\theta )\right )\, d\theta \] \[\int \frac {\partial }{\partial \theta }\log \, f_{\theta }(x)\, dx = C_1(\theta )\, T(X) + C_2(\theta ) + \underbrace {C_3(X)}_{\text {constant}}\] \[\log \, f_{\theta }(x) = C_1(\theta ) \, T(x) + C_2(\theta ) + C_3(x)\] \begin {align*} f_{\theta }(x) & = e^{C_1(\theta ) \, T(x) + C_2(\theta ) + C_3(x)} = \underbrace {e^{C_2(\theta )}}_{C(\theta )} \, \exp \Big \{\underbrace {C_1(\theta )}_{q(\theta )}\, \underbrace {T(x)}_{T(X)}\Big \}\, \underbrace {e^{C_3(x)}}_{h(x)} \end {align*}

which is an exponential family with natural sufficient statistic \(T(X)\). □

Note. If \(\tau (\theta ) = \theta \), then \[\var (T) \geq \frac {1}{J(\theta )}.\]

Example 2.3.2. Let \(X_1, \, X_2, \, \cdots \, , \, X_n\) be a random sample from the \(POI(\theta )\) distribution. Show that the \(UMVUE\) of \(\theta \) achieves the Cramer-Rao lower bound.

Solution. \(\tau (\theta ) = \theta \, \implies \, \tau '(\theta ) = 1\).
\(T(X) = \sum ^n_{i = 1}X_i\,\) is a complete sufficient statistic. Now \(E\left (\overline {X}\right ) = \theta \), therefore \(\overline {X}\) is the \(U.M.V.U.E\) of \(\theta \). \[L(\theta ) = \frac {e^{-\theta }\, \theta ^x}{x!}.\]

\[\mathcal {L}(\theta ) = -\theta + x\, \log \theta - \log x!.\]

\[S(\theta ) = -1 + \frac {x}{\theta }.\]

\[I(\theta ) = -\left (-\frac {x}{\theta ^2}\right ) = \frac {x}{\theta ^2}.\]

\[J_1(\theta ) = E\left [I(\theta ,X)\right ] = E\left (\frac {X}{\theta ^2}\right ) = \frac {\theta }{\theta ^2} = \frac {1}{\theta }.\] \[\therefore \hspace {0.3cm} J(\theta ) = \frac {n}{\theta }.\]

\[C.R.L.B = \frac {1}{J(\theta )} = \frac {\theta }{n}.\]

\[\var \left (\overline {X}\right ) = \frac {\var (X)}{n} = \frac {\theta }{n} = CRLB.\] Therefore the \(UMVUE\) attains the \(CRLB\).

OR
\(X\thicksim POI(\theta )\) is the \(REF\) distribution with natural sufficient statistic \(\sum ^n_{i = 1} X_i\). \(\overline {X}\) is also a natural sufficient statistic. Therefore \(\var (\overline {X}) = CRLB\). □

Example 2.3.3. Suppose \(X_1, \, \cdots \, , \ X_n\) is a random sample from the distribution with pdf \[f_{\theta }(x) = \theta \, x^{\theta - 1}\hspace {0.3cm} \, 0 < x < 1, \hspace {0.3cm} \theta > 0.\]

(a)
Show that the variance of the \(UMVUE\) of \(\theta \) does not achieve the \(CRLB\). What is the efficiency of the \(UMVUE\).
(b)
Find the \(CRLB\) for variances of unbiased estimators of \(\, \tau (\theta ) = \theta (1 - \theta )\).

Solution.

(a)
\(f_{\theta }(x) = \theta \, x^{\theta - 1}\, = \underbrace {\theta }_{C(\theta }\, \exp \Big \{\underbrace {\theta }_{q(\theta )}\, \underbrace {x}_{T(X)}\Big \}\, \underbrace {\frac {1}{x}}_{h(x)}.\)
Let \(\eta = \theta \) \[f_{\eta }(x) = \eta \, \exp \left \{\eta \, \log x\right \}\, \frac {1}{x}.\] Therefore \(f_{\theta }(x)\) is a \(REF\) distribution and hence \(\sum ^n_{i = 1}\log x_i\,\) is a compete sufficient statistic. \[-\log X_i \thicksim EXP\left (\frac {1}{\theta }\right )\]

\[T = - \sum ^n_{i = 1} \log X_i \thicksim GAM\left (n, \frac {1}{\theta }\right ).\] by problem 1.2.4 \[E\left (X^p\right ) = \frac {\beta ^p\, \Gamma \left (\alpha + p\right )}{\Gamma (\alpha )}\]

\[E\left (T^{-1}\right ) = \frac {\left (\frac {1}{\theta }\right )^{-1}\, \Gamma (n - 1)}{\Gamma (n)} = \frac {\theta \, \Gamma (n - 1)}{(n - 1)\, \Gamma (n - 1)} = \frac {\theta }{n - 1}.\]

\[\therefore \, \, E\left [\frac {n - 1}{T} \right ] = \theta .\] Therefore \(\,\frac {n - 1}{T}\,\) is the \(UMVUE\) of \(\theta \).

Now \[L_1(\theta ) = \theta \, x^{\theta - 1}\]

\[\mathcal {L}_1(\theta ) = \log \theta + (\theta - 1)\, \log x\]

\[S_1(\theta ) = \frac {1}{\theta } + \log x\]

\[I_1(\theta ) = -\left (-\frac {1}{\theta ^2}\right ) = \frac {1}{\theta ^2}\]

\[J_1(\theta ) = \frac {1}{\theta ^2}\]

\[\therefore \hspace {0.3cm} J(\theta ) = n\, J_1(\theta ) = \frac {n}{\theta ^2}.\]

\[CRLB = \frac {\left [\tau '(\theta )\right ]^2}{J(\theta )} = \frac {1}{J(\theta )} = \frac {\theta ^2}{n}.\]

Now \[\var \left [\frac {n - 1}{T}\right ] = (n - 1)^2\, \var \left (\frac {1}{T}\right )\]

\[\var \left (\frac {1}{T}\right ) = E\left (\frac {1}{T}\right )^2 - \left (\left (\frac {1}{T}\right )\right )^2.\]

\[E\left (\frac {1}{T}\right )^2 = E\left (T^{-2}\right ) = \frac {\left (\frac {1}{\theta }\right )^{-2}\, \Gamma (n - 2)}{\Gamma (n)} = \frac {\theta ^2\, \Gamma (n - 2)}{(n - 1)\, (n - 2)\, \Gamma (n - 2)} = \frac {\theta ^2}{(n - 1)\, (n - 2)}.\]

\[\var \left (\frac {1}{T}\right ) = \frac {\theta ^2}{(n - 1)(n - 2)} - \frac {\theta ^2}{(n - 1)^2}.\] \begin {align*} \therefore \,\, \var \left (\frac {n - 1}{T}\right ) & = (n - 1)^2\, \theta ^2\, \left [\frac {1}{(n - 1)(n - 2)} - \frac {1}{(n - 1)^2}\right ] = \left (\frac {n - 1}{n - 2} - 1\right )\, \theta ^2\\ & = \left (\frac {n - 1 - n + 2}{n - 2}\right )\, \theta ^2\\ & = \frac {\theta ^2}{n - 2} > \frac {\theta ^2}{n}. \end {align*}

Therefore the variance of the \(UMVUE\) of \(\theta \) does not achieve the \(CRLB\).

OR
\(T = -\sum ^n_{i = 1} \log X_i\) is a natural sufficient statistic. But \(g(T) = \frac {n - 1}{T}\) is not a natural sufficient statistic. Therefore \(g(T) = \frac {n - 1}{T}\) does not achieve the \(CRLB\). \[\textit {efficiency} = \frac {CRLB}{\var [g(T)]} = \frac {\theta ^2/n}{\theta ^2/n - 2} = \frac {n - 2}{n}.\] i.e \(\, \frac {n - 2}{n}\longrightarrow 1, \, g(T)\) is an asymptotically efficient estimator.

(b)
\(\tau (\theta ) = \theta (1 - \theta ) = \theta - \theta ^2\) \[CRLB = \frac {\left [\tau '(\theta )\right ]^2}{J(\theta )} = \frac {\left (1 - 2\theta \right )^2}{n/\theta ^2} = \frac {\theta ^2\, (1 - 2\theta )^2}{n}.\]

Problem 2.3.1. For each of the following, determine whether the variance of the \(UMVUE\) of \(\theta \) based on a random sample \(X_1, \, \cdots \, , \, X_n\) achieve the Cramer Rao Lower bound and find the efficiency of the \(UMVUE\). For each case find the \(CRLB\) for variances of unbiased estimators of \(\theta ^2\).

  • \(N(\theta ,\mu )\)
  • \(BIN(1,\theta )\)
  • \(N(0,\theta ^2)\)
  • \(N(0,\theta )\)

Show solution

Solution. Throughout, \(T=\sum _{i=1}^{n}X_i^{2}\) where it appears, and the bound for a target \(\tau (\theta )\) is \(\left [\tau '(\theta )\right ]^{2}/I_n(\theta )\).

(a) \(N(\theta ,\mu )\), variance \(\mu \) known

\(I_n(\theta )=n/\mu \); the UMVUE is \(\overline {X}\) with variance \(\mu /n\). The bound is \(\mu /n\), so it is attained and the efficiency is \(1\). For \(\tau =\theta ^{2}\) the bound is \(\left (2\theta \right )^{2}\mu /n = 4\theta ^{2}\mu /n\).

(b) BIN\((1,\theta )\)

\(I_n(\theta )=n/\left [\theta (1-\theta )\right ]\); the UMVUE is \(\overline {X}\) with variance \(\theta (1-\theta )/n\), equal to the bound. Efficiency \(1\). For \(\tau =\theta ^{2}\) the bound is \(4\theta ^{2}\cdot \theta (1-\theta )/n = 4\theta ^{3}(1-\theta )/n\).

(c) \(N\left (0,\theta ^{2}\right )\)

Here \(\theta \) is the standard deviation. Per observation \(-E\left [\partial ^{2}\log f/\partial \theta ^{2}\right ] = 2/\theta ^{2}\), so \(I_n(\theta )=2n/\theta ^{2}\) and the bound for \(\theta \) is \(\theta ^{2}/(2n)\). The UMVUE is \(c_n\sqrt {T}\) with \(c_n = \Gamma (n/2)\big /\left [\sqrt 2\,\Gamma \left ((n+1)/2\right )\right ]\), and its variance is \(\theta ^{2}\left (nc_n^{2}-1\right )\), which is strictly larger. The bound is not attained; the efficiency \(1\big /\left [2n\left (nc_n^{2}-1\right )\right ]\) equals \(0.915\) at \(n=2\), \(0.977\) at \(n=10\) and \(0.998\) at \(n=100\) — so it rises to one only in the limit. For \(\tau =\theta ^{2}\) the bound is \(\left (2\theta \right )^{2}\theta ^{2}/(2n) = 2\theta ^{4}/n\), and this one is attained, by \(T/n\).

(d) \(N(0,\theta )\)

Now \(\theta \) is the variance. Per observation the information is \(1/\left (2\theta ^{2}\right )\), so \(I_n(\theta )=n/\left (2\theta ^{2}\right )\) and the bound for \(\theta \) is \(2\theta ^{2}/n\). The UMVUE is \(T/n\) with variance \(2\theta ^{2}/n\): attained, efficiency \(1\). For \(\tau =\theta ^{2}\) the bound is \(4\theta ^{2}\cdot 2\theta ^{2}/n = 8\theta ^{4}/n\).

Remark. Parts (c) and (d) are the same model written with different parameters, and they answer the question differently. The bound is attained exactly when the score is an affine function of the estimator, which for an exponential family happens for the mean of the natural sufficient statistic and for nothing else. That statistic is \(\sum X_i^{2}\), whose mean is \(n\theta ^{2}\) in (c) and \(n\theta \) in (d) — so the bound is attained for \(\theta ^{2}\) in (c) and for \(\theta \) in (d). Attainability is a property of the parameterisation, not of the model.

Theorem 2.3.4 (The Multiparameter Case). If \(\theta = (\theta _1, \, \cdots \, , \, \theta _k)^t\) is a vector, \(\tau (\theta )\) is a function of \(\theta \) and \(T(X)\) is any unbiased estimator of \(\tau (\theta )\) in a regular model, then the information inequality becomes \[\var _{\theta }(T) \geq D(\theta )^t\, J^{-1}(\theta )\, D(\theta )\] for all \(\theta \in \Omega \) where \[D(\theta ) = \left (\frac {\partial \tau }{\partial \theta _1}\, , \, \cdots \, , \, \frac {\partial \tau }{\partial \theta _k}\right )^t.\]

Proof. Let \(S(\theta ,X)\) be the score vector, so that \(E_{\theta }\left [S\right ]=0\) and \(\var _{\theta }(S)=J(\theta )\). Differentiating the unbiasedness identity \(E_{\theta }\left [T(X)\right ]=\tau (\theta )\) with respect to \(\theta _j\) and interchanging derivative and integral, which regularity permits, \[\frac {\partial \tau }{\partial \theta _j} = \int T(x)\,\frac {\partial }{\partial \theta _j} f_{\theta }(x)\,dx = \int T(x)\, S_j(\theta ,x)\, f_{\theta }(x)\,dx = \cov _{\theta }\!\left (T, S_j\right ),\] the last equality because \(E_{\theta }\left [S_j\right ]=0\). Collecting the \(k\) components, \(\cov _{\theta }(T,S) = D(\theta )\).

Now take any fixed vector \(c\in \mathbb {R}^{k}\) and apply the Cauchy–Schwarz inequality to the scalar random variables \(T\) and \(c^{t}S\): \[\left [\cov _{\theta }\!\left (T,\, c^{t}S\right )\right ]^{2} \;\leq \; \var _{\theta }(T)\ \var _{\theta }\!\left (c^{t}S\right ),\] that is \[\left (c^{t}D\right )^{2} \;\leq \; \var _{\theta }(T)\ c^{t}J(\theta )\,c .\] Choose \(c = J^{-1}(\theta )\,D(\theta )\), which is legitimate because \(J(\theta )\) is positive definite in a regular model. Then \(c^{t}D = D^{t}J^{-1}D\) and \(c^{t}Jc = D^{t}J^{-1}JJ^{-1}D = D^{t}J^{-1}D\), so the inequality reads \[\left (D^{t}J^{-1}D\right )^{2} \;\leq \; \var _{\theta }(T)\ D^{t}J^{-1}D .\] Since \(J^{-1}\) is positive definite, \(D^{t}J^{-1}D>0\) unless \(D=0\), and dividing through gives \[\var _{\theta }(T) \;\geq \; D(\theta )^{t}\,J^{-1}(\theta )\,D(\theta ).\] If \(D=0\) the inequality is trivial. □

Remark. The scalar case is \(k=1\), where \(D=\tau '(\theta )\) and the bound collapses to \(\left [\tau '(\theta )\right ]^{2}/J(\theta )\). Equality holds exactly when \(T\) is an affine function of the score, which is why the bound is attained in exponential families and generally is not attained outside them.

Questions on this section

Stuck on something here? Ask below and it stays attached to this topic.