2.2 Properties of the Score and Information

Consider a family of probability density functions \(\{f_{\theta }(x):\, \theta \in \Omega \}\). Let \(A = \{x:\, f_{\theta }(x) > 0\}\). Then \(\int _Af_{\theta }(x) \, dx = 1\) and therefore \[\int _A\frac {\partial }{\partial \theta }\, f_{\theta }(x)\, dx = \frac {\partial }{\partial \theta }\int _Af_{\theta }(x)\, dx = 0\] provided that the integral can be interchanged with the derivative. Models that permit this interchange and calculations of the Fisher information are called regular models.

Theorem 2.2.1. If \(X = (X_1, \, \cdots \, , \, X_n)\) is a random sample from a regular model \(\{f_{\theta }(x):\, \theta \in \Omega \}\), then

(i)
\(E_{\theta }\left [S(\theta , X)\right ] = 0\)
(ii)
\(\var _{\theta }\left [S(\theta , X)\right ] = E_{\theta }\left [\left (S(\theta ,X)\right )^2\right ] = J(\theta )\).

Solution.

(i)
\(E_{\theta }\left [S(\theta ,X)\right ] = E_{\theta }\left [\displaystyle {\sum ^n_{i=1}}\dfrac {\partial }{\partial \theta } \, \log \, f_{\theta }(X_i)\right ] = \displaystyle {\sum ^n_{i= 1} E_{\theta } \left [\frac {\partial }{\partial \theta } \, \log \, f_{\theta }(X_i)\right ]}\). Now \begin {align*} E\left [\frac {\partial }{\partial \theta }\, \log \, f_{\theta }(X_i)\right ] & = \int _A\frac {\partial }{\partial \theta }\, \log \, f_{\theta }(x)\, f_{\theta }(x)\, dx\\ & = \int _A\frac {\frac {\partial }{\partial \theta }\, f_{\theta }(x)}{f_{\theta (x)}}\, \cdot \, f_{\theta }(x)\, dx\\ & = \int _A\frac {\partial }{\partial \theta } f_{\theta }(x)\, dx\\ & = \frac {\partial }{\partial \theta }\, \int _Af_{\theta }(x)\, dx\\ & = \frac {\partial }{\partial \theta }\, A(1)\\ & = 0. \end {align*}
(ii)
\(\begin {aligned}[t] \var \left [S(\theta ,X)\right ] = E\left [\left (S(\theta , X)\right )^2\right ] - \left (E\left [S(\theta ,X)\right ]\right )^2 & = E\left [\left (S(\theta ,X)\right )^2\right ] - 0^2\\ & = E\left [\left (S(\theta ,X)\right )^2\right ]\hspace {1cm}\cdots \cdots \hspace {0.5cm} (*) \end {aligned}\)

\(E\left [\left (S(\theta ,X)\right )^2\right ] = E\left \{\left [\sum ^n_{i = 1}S_1(\theta ,X_i)\right ]^2\right \}\,\) but

\begin {align*} S_1(\theta , X_i) = & \frac {\partial }{\partial \theta }\, \log \, f_{\theta }(X_i)\\ = & E\left \{\sum ^n_{i = 1}\left (S_1(\theta ,X_i)\right )^2 \, + \, \sum _{i\neq j}S_1(\theta ,X_i)\, S_1(\theta ,X_j)\right \}\\ = & \sum ^n_{i = 1} E\left (S_1(\theta ,X_i)\right )^2 \, + \, \sum _{i\neq j}\underbrace {E\left (S_1(\theta ,X_i)\right )}_0\,\underbrace {E\left (S_1(\theta ,X_j)\right )}_0\\ & \text {Since}\, X_i \, \text {and}\, X_j\, \text {are independent}\\ & = \sum ^n_{i = 1}E\left [\frac {\frac {\partial }{\partial \theta }\, f_{\theta }(X_i)}{f_{\theta }(X_i)}\right ]^2. \end {align*}

Recall Example 2.1.1 \[I(\theta ,X) = \sum ^n_{i = 1} \left [\left (\frac {\frac {\partial }{\partial \theta }\, f_{\theta }(X_i)}{f_{\theta }(X_i)}\right )^2 \, - \, \frac {\frac {\partial ^2}{\partial \theta ^2}\, f_{\theta }(X_i)}{f_{\theta }(X_i)}\right ]\] The Fisher information \[J(\theta ,X) = \sum ^n_{i = 1} E\left [\left (\frac {\frac {\partial }{\partial \theta }\, f_{\theta }(X_i)}{f_{\theta }(X_i)}\right )^2 \right ]\, - \,\sum ^n_{i = 1} E\left [ \frac {\frac {\partial ^2}{\partial \theta ^2}\, f_{\theta }(X_i)}{f_{\theta }(X_i)}\right ]\] but \begin {align*} E\left (\frac {\frac {\partial ^2}{\partial \theta ^2}\, f_{\theta }(X_i)}{f_{\theta }(X_i)}\right ) & = \int _A\frac {\frac {\partial ^2}{\partial \theta ^2}\, f_{\theta }(x)}{f_{\theta }(x)}\, \cdot \, f_{\theta }(x)\, dx\\ & = \int _A \frac {\partial ^2}{\partial \theta ^2}\, f_{\theta }(x)\, dx\\ & = \frac {\partial ^2}{\partial \theta ^2}\, \int _Af_{\theta }(x)\, dx\\ & = \frac {\partial ^2}{\partial \theta ^2}\, (1) = 0. \end {align*}

\[\therefore \,\,\, J(\theta ) = \sum ^n_{i = 1} E\left [\left (\frac {\frac {\partial }{\partial \theta }\, f_{\theta }(X_i)}{f_{\theta }(X_i)}\right )^2\right ].\]

\[\therefore \,\,\, \var _{\theta }\left (S(\theta ,X)\right ) = E_{\theta }\left [\left (S(\theta ,X)\right )^2\right ] = J(\theta ).\] □

Example 2.2.2. Let \(X = (X_1, \, \cdots \, , \, X_n)\) be a random sample from the \(POI(\theta )\) distribution, then \[L(\theta ) = \prod ^n_{i = 1}\frac {e^{-\theta }\, \theta ^{x_i}}{x_i!} = \frac {e^{-n\theta }\, \theta ^{\sum ^n_{i = 1} x_i}}{\prod ^n_{i = 1} x_i!}.\]

\[\mathcal {L}(\theta ) = \log \, L(\theta ) = -n\theta + \sum ^n_{i = 1} x_i\, \log \theta - \log \prod ^n_{i = 1} x_i!.\]

\[S(\theta ) = \frac {\partial \mathcal {L}}{\partial \theta } = -n + \frac {\sum ^n_{i = 1} x_i}{\theta }.\]

\[I(\theta ) = \frac {-\partial ^2\mathcal {L}}{\partial \theta ^2} = -\left (\frac {-\sum ^n_{i = 1} x_i}{\theta ^2}\right ) = \frac {\sum ^n_{i = 1} x_i}{\theta ^2}.\]

\begin {align*} J(\theta ) & = E\left [I(\theta , X)\right ] = E\left [\frac {1}{\theta ^2} \, \sum ^n_{i = 1} X_i\right ]\\ & = \frac {1}{\theta ^2} \, \sum ^n_{i = 1} E(X_i)\\ & = \frac {1}{\theta ^2}\, \cdot \, n\theta \\ & = \frac {n}{\theta }. \end {align*}

Now \begin {align*} E\left [S(\theta ,X)\right ] & = E\left [-n + \frac {1}{\theta }\, \sum ^n_{i = 1} X_i\right ] = -n + \frac {1}{\theta }\, \sum ^n_{i = 1} E(X_i)\\ & = -n + \frac {1}{\theta }\cdot n\theta \\ & =0. \end {align*}

\begin {align*} \var \left [S(\theta ,X)\right ] & = \var \left [-n + \frac {1}{\theta }\, \sum ^n_{i = 1} X_i\right ] = \frac {1}{\theta ^2}\, \sum ^n_{i = 1} \var (X_i)\\ & = \frac {1}{\theta ^2}\, \cdot \, n\theta \\ & = \frac {n}{\theta }\\ & = J(\theta ). \end {align*}

2.2.1 The Multiparameter Score and Information

Theorem 2.2.3 (The Multiparameter Case). Suppose \(\theta = (\theta _1, \, \cdots \, , \, \theta _k)^t\), then the score function \(S(\theta )\) is a \(k\)-dimensional vector whose \(i^{\text {th}}\) component is the derivative of \(\mathcal {L}(\theta )\) with respect to the \(i^{\text {th}}\) component of \(\theta \) \[S(\theta ) = \left [\frac {\partial \mathcal {L}(\theta )}{\partial \theta _1}\, , \, \cdots \, , \, \frac {\partial \mathcal {L}(\theta )}{\partial \theta _k}\right ]^t.\] The information function \(I(\theta )\) is a \(k\times k\) matrix whose \((i,j)^{\text {th}}\) element is minus the derivative of \(\mathcal {L}(\theta )\) with respect to the \(i^{\text {th}}\) and \(j^{\text {th}}\) components of \(\theta \), \[I(\theta ) = \left [I_{ij}(\theta )\right ]_{k\times k} = \left [\frac {-\partial ^2\mathcal {L}(\theta )}{\partial \theta _i\, \partial \theta _j}\right ]_{k\times k} \hspace {0.2cm} , \hspace {0.5cm} i,j = 1, \, 2, \, \cdots \, , \, k\] The Fisher information is a \(k\times k\) matrix whose elements are componentwise expected of the information matrix \[J_{ij}(\theta ) = E\left [I_{ij}(\theta ,X)\right ]\hspace {0.2cm} , \hspace {0.3cm} i,j = 1, \, 2, \, \cdots \, , \, k\] For a regular family of distributions \(E\left [S(\theta , X)\right ] = (0, \, \cdots \, , \, 0)^t\) and \[J(\theta ) = \var \left [S(\theta ,X)\right ] = E\left [S(\theta , X)\, \left [S(\theta ,X)\right ]^t\right ]\]

Note. The invariance property of the M.L estimator also holds in the multiparameter case.

Example 2.2.4. Let \(X_1, \, \cdots \, , \, X_n\) be a random sample from the \(N(\mu , \sigma ^2)\) distribution. Find the ML estimator of \(\theta = (\mu , \sigma ^2)^t\) the score function and the Fisher information.

Solution. Let \(\, \theta _1 = \mu , \, \, \theta _2 = \sigma ^2, \, \, \theta = (\theta _1, \theta _2)^{t}\) \[L(\theta ) = \prod ^n_{i = 1}\frac {1}{\sqrt {2\pi \theta ^2}}\exp \left \{-\frac {1}{2\theta _2}(x_i - \theta _1)^2\right \} = \frac {1}{\left (2\pi \theta _2\right )^{n/2}}\, \exp \left \{-\frac {1}{2\theta _2}\, \sum ^n_{i = 1} (x_i - \theta _1)^2\right \}\]

\[\mathcal {L}(\theta ) = \frac {-1}{2\theta _2} \, \sum ^n_{i = 1} (x_i - \theta )^2 - \frac {n}{2}\, \log \, 2\pi \theta _2\]

\[S(\theta ) = \begin {pmatrix} \dfrac {\partial \mathcal {L}}{\partial \theta _1}\\\\ \dfrac {\partial \mathcal {L}}{\partial \theta _2}\\ \end {pmatrix}\]

\[\frac {\partial \mathcal {L}}{\partial \theta _1} = \frac {-1}{2\theta _2}\sum ^n_{i = 1} 2(x_i - \theta _1)\, (-1) = \frac {1}{2\theta _2}\sum ^n_{i = 1} 2(x_i - \theta _1)\]

\[\frac {\partial \mathcal {L}}{\partial \theta _2} = \frac {1}{2\theta ^2_2} \sum ^n_{i = 1}(x_i - \theta _1)^2 - \frac {n}{2}\cdot \frac {2\pi }{2\pi \,\theta _2} = \frac {1}{2\theta ^2_2}\sum ^n_{i = 1} (x_i - \theta _1)^2 - \frac {n}{2\theta _2}.\]

\[S(\theta ) = \begin {pmatrix} \dfrac {1}{\theta _2}\sum ^n_{i = 1}(x_i - \theta _1)\\\\ \dfrac {1}{2\theta _2^2}\sum ^n_{i = 1} (x_i - \theta _1)^2 - \dfrac {n}{2\theta _2}\\ \end {pmatrix}.\]

\(S(\theta ) = 0\)

\[\frac {1}{\theta ^2}\, \sum ^n_{i = 1} (x_i - \theta _1) = 0\]

\[\widehat {\theta }_1 = \frac {1}{n}\sum ^n_{i =1} X_i = \overline {X}.\]

\[\frac {1}{2\theta ^2_2} \sum ^n_{i = 1} (X_i - \theta _1)^2 - \frac {n}{2\theta _2} = 0.\]

\[\frac {1}{n}\sum ^n_{i = 1} (X_i - \overline {X})^2 = \widehat {\theta }_2.\]

\(\therefore \, \widehat {\theta } = \left (\overline {X}\, \, , \, \, \frac {1}{n}\sum ^n_{i = 1}(X_i - \overline {X})^2\right )^t\,\) is the ML estimator of \(\theta = (\mu , \sigma ^2)^t.\)

\[I(\theta ) = \begin {pmatrix} -\dfrac {\partial ^2\mathcal {L}}{\partial \theta _1^2} & - \dfrac {\partial ^2\mathcal {L}}{\partial \theta _1\, \partial \theta _2}\\\\ - \dfrac {\partial ^2\mathcal {L}}{\partial \theta _2\, \partial \theta _1} & -\dfrac {\partial ^2\mathcal {L}}{\partial \theta _2^2}\\ \end {pmatrix}.\]

\[-\frac {\partial ^2\mathcal {L}}{\partial \theta _1^2} = -\frac {1}{\theta _2}\, \sum ^n_{i = 1}(-1) = \frac {n}{\theta _2}.\]

\[-\frac {\partial ^2\mathcal {L}}{\partial \theta _1\, \partial \theta _2} = -\left (-\frac {1}{\theta ^2_2}\sum ^n_{i = 1}(x_i - \theta )\right ) = \frac {1}{\theta ^2_2}\sum ^n_{i = 1} (x_i - \theta _1).\]

\[-\frac {\partial ^2\mathcal {L}}{\partial \theta ^2_2} = - \left (-\frac {1}{\theta ^3_2} \sum ^n_{i = 1} (x_i - \theta _1)^2 + \frac {n}{2\theta ^2_2}\right ) = \frac {1}{\theta ^3_2} \sum ^n_{i = 1} (x_i - \theta _1)^2 - \frac {n}{2\theta ^2_2}.\]

\[I(\theta ) = \begin {pmatrix} \dfrac {n}{\theta _2} && \dfrac {1}{\theta ^2_2}\sum ^n_{i = 1}(X_i - \theta _1)\\\\ \dfrac {1}{\theta ^2_2}\sum ^n_{i = 1}(X_i - \theta _1) && \dfrac {1}{\theta ^3_2}\sum ^n_{i = 1} (X_i - \theta _1)^2 - \dfrac {n}{2\theta ^2_2}\\ \end {pmatrix}.\]

\[J(\theta ) = E\left [I(\theta )\right ] = \begin {pmatrix} \dfrac {n}{\theta _2} && \dfrac {1}{\theta ^2_2}\sum ^n_{i = 1}(E(X_i) - \theta _1)\\\\ \dfrac {1}{\theta ^2_2}\sum ^n_{i = 1}(E(X_i) - \theta _1) && \dfrac {1}{\theta ^3_2}\sum ^n_{i = 1} E(X_i - \theta _1)^2 - \dfrac {n}{2\theta ^2_2}\\ \end {pmatrix}.\]

\[\sum ^n_{i = 1} \left (E(X_i) - \theta _1\right ) = \sum ^n_{i = 1} (\theta _1 - \theta _1) = 0\] \begin {align*} \frac {1}{\theta _2^3}\sum ^n_{i = 1} E\left (X_i - \theta _1\right )^2 - \frac {n}{2\theta ^2_2} & = \sum ^n_{i = 1} \var (X_i) - \frac {n}{2\theta ^2_2} = \frac {1}{\theta ^3_2}\, n\theta _2 - \frac {n}{2\theta _2^2}\\ & = \frac {n}{\theta ^2_2} - \frac {n}{2\theta _2^2}\\ & = \frac {n}{2\theta _2^2}. \end {align*}

\[\implies \hspace {0.5cm} J(\theta ) = \begin {pmatrix} \dfrac {n}{\theta _2^2} & 0\\\\ 0 & \dfrac {n}{2\theta _2^2}\\ \end {pmatrix} = \begin {pmatrix} \dfrac {n}{\sigma ^2} & 0 \\\\ 0 & \dfrac {n}{2\sigma ^4}\\ \end {pmatrix}.\] i.e \(\hspace {0.3cm} \left (\overline {X}\, , \, \frac {1}{n}\sum ^n_{i = 1}(X_i - \overline {X})^2\right )\) \[\left (\overline {X}\, ,\, \dfrac {n - 1}{n}\, S^2\right ).\] □

Problem 2.2.1. Let \((X_1, \, X_2, \, X_3) \thicksim MULT(n, \, \theta _1, \, \theta _2, \, \theta _3)\). Find the M.L estimator of \(\theta = (\theta _1, \theta _2)^t\), the score function and the Fisher information. Hint: use \[f_{\theta }(x_1, x_2, x_3) = \frac {n!}{x_1!\, x_2!\, (n - x_1 - x_2)!}\, \theta _1^{x_1}\, \theta _2^{x_2}\, \left (1 - \theta _1 - \theta _2\right )^{n - x_1 - x_2}\]

Show solution

Solution. Write \(\theta _3 = 1-\theta _1-\theta _2\) and \(x_3 = n-x_1-x_2\). Then \[\log L = \text {const} + x_1\log \theta _1 + x_2\log \theta _2 + x_3\log \left (1-\theta _1-\theta _2\right ).\]

Score

\[S_j(\theta ,X) = \frac {X_j}{\theta _j} - \frac {X_3}{1-\theta _1-\theta _2}, \qquad j=1,2 .\]

Maximum likelihood estimator

Setting both components to zero gives \(X_1/\theta _1 = X_2/\theta _2 = X_3/\theta _3\); with \(\theta _1+\theta _2+\theta _3=1\) and \(X_1+X_2+X_3=n\) the common ratio is \(n\), so \[\widehat {\theta }_1 = \frac {X_1}{n},\qquad \widehat {\theta }_2 = \frac {X_2}{n} .\]

Fisher information

Differentiating again, and using \(E(X_j)=n\theta _j\), \[-\frac {\partial ^{2}\log L}{\partial \theta _j^{2}} = \frac {X_j}{\theta _j^{2}} + \frac {X_3}{\theta _3^{2}}, \qquad -\frac {\partial ^{2}\log L}{\partial \theta _1\partial \theta _2} = \frac {X_3}{\theta _3^{2}},\] so \[J(\theta ) = n\begin {pmatrix} \dfrac {1}{\theta _1}+\dfrac {1}{\theta _3} & \dfrac {1}{\theta _3}\\[8pt] \dfrac {1}{\theta _3} & \dfrac {1}{\theta _2}+\dfrac {1}{\theta _3} \end {pmatrix}.\] The off-diagonal entries are not zero: the two estimators are negatively correlated, as they must be, because a larger \(\widehat {\theta }_1\) leaves less probability for \(\theta _2\).

2.2.2 The Likelihood Equation and its Solution

Theorem 2.2.5 (Newton’s Method). Suppose that the M.L estimate is determined by the likelihood equation \(S(\theta ) = 0\).
It frequently happens that an analytic solution for \(\theta \) cannot be obtained. If we begin with an approximate value for the parameter, \(\theta ^{(0)}\), we may update that value as follows: \[\theta ^{(i + 1} = \theta ^{(i)} + \frac {S\left (\theta ^{(i)}\right )}{I\left (\theta ^{(i)}\right )}\hspace {0.3cm} , \hspace {0.5cm} i = 0, \, 1, \, \cdots \] and provided convergence of \(\theta ^{(i)}\), \(i\longrightarrow \infty \) obtains, it converges to the score equation above.
In the multiparameter case, where \(S(\theta )\) is a vector and \(I(\theta )\) is a matrix, then Newton’s method becomes \[\theta ^{(i + 1)} = \theta ^{(i)} + \left [I\left (\theta ^{(i)}\right )\right ]^{-1}\, S\left (\theta ^{(i)}\right )\, , \hspace {0.2cm} i = 0, \, 1, \, \cdots \] This algorithm is also called the Newton-Raphson method. A similar algorithm is obtained \(I(\theta )\) replaced by the Fisher information \(J(\theta )\).

Note. This is based on Taylor series: i.e \[f(x) = f(x_0) + \frac {(x - x_0)}{1!}\, f'(x_0) + \frac {(x - x_0)^2}{2!} \, f''(x_0) + \cdots \cdots \]

\[f(x) \approx f(x_0) + (x - x_0)\, f'(x_0)\] We solve the equation \(f(x) = 0\) \[f(x_0) + (x - x_0) \, f'(x_0) = 0\] \[\frac {f(x_0)}{f'(x_0)} + x -x_0 = 0\] \[x = x_0 - \frac {f(x_0)}{f'(x_0)}\] i.e \[\theta ^{(i + 1)} = \theta ^{(i)} + \frac {S\left (\theta ^{(i)}\right )}{I\left (\theta ^{(i)}\right )}.\]

Example 2.2.6. Suppose \(X_1, \, \cdots \, , \, X_n\) is a random sample from the distribution with pdf \[f_{\theta }(x) = \frac {\alpha \, \beta }{\left (1 + \beta \, x\right )^{\alpha + 1}}\, , \hspace {0.3cm} x > 0;\, \, \alpha , \, \beta >0.\] Explain how you find the M.L estimate of \(\theta = (\alpha , \beta )^t\). Find the Fisher information.

Solution. \[L(\theta ) = \prod ^n_{i = 1} \theta \, x_i^{\theta - 1}\, e^{-x_i^{\theta }} = \theta ^n\, \prod ^n_{i = 1} x_i^{\theta - 1}\, e^{-\sum ^n_{i = 1} x_i^{\theta }}\]

\[\mathcal {L}(\theta ) = n\, \log \theta + (\theta - 1)\, \log \prod ^n_{i = 1} x_i - \sum ^n_{i =1} x_i^{\theta }.\]

\[S(\theta ) = \frac {n}{\theta } + \sum ^n_{i = 1}\log x_i - \sum ^n_{i = 1} x_i^{\theta }\, \log x_i.\] \(S(\theta ) = 0\,\,\) cannot be solved explicitly

\[I(\theta ) = -\frac {\partial ^2\mathcal {L}}{\partial \theta ^2} = -\left (-\frac {n}{\theta ^2} - \sum ^n_{i = 1} x^{\theta }_i \, \left (\log x_i\right )^2\right ) = \frac {n}{\theta ^2} + \sum ^n_{i = 1} x_i^{\theta }\, \left (\log x_i\right )^2.\] Solve for \(\theta \) iteratively using \[\theta ^{(i + 1)} = \theta ^{(i)} + \frac {S\left (\theta ^{(i)}\right )}{I\left (\theta ^{(i)}\right )}\hspace {0.2cm} , \hspace {0.6cm} i = 0, \, 1, \, 2, \, \cdots \] iterate until convergence to obtain \(\widehat {\theta }\). □

Example 2.2.7. Suppose \(X_1, \, \cdots \, , \, X_n\) is a random sample from the \(GAM(\alpha , \beta )\) distribution. Explain how you would obtain the \(M.L\) estimator of \(\theta = (\alpha ,\beta )^t\). Find the Fisher information.

Solution. \[L(\theta ) = \prod ^n_{i = 1}\frac {x_i^{\alpha - 1}\, e^{-\frac {x_i}{\beta }}}{\beta ^{\alpha }\, \Gamma (\alpha )} = \frac {\left (\prod ^n_{i = 1} x_i\right )^{\alpha - 1}\, e^{-\frac {1}{\beta }\sum ^n_{i = 1} x_i}}{\beta ^{n\alpha }\, \left (\Gamma (\alpha )\right )^n}.\]

\[\mathcal {L}(\theta ) = (\alpha - 1)\, \log \prod ^n_{i = 1} x_i - \frac {1}{\beta } \sum ^n_{i = 1} x_i - n\alpha \, \log \beta - n\, \log \Gamma (\alpha ).\]

\[S(\theta ) = \begin {pmatrix} \dfrac {\partial \mathcal {L}}{\partial \alpha }\\\\ \dfrac {\partial \mathcal {L}}{\partial \beta }\\ \end {pmatrix}\hspace {2cm} \begin {matrix} \dfrac {\partial \mathcal {L}}{\partial \alpha } = \sum ^n_{i = 1} \log x_i - n\, \log \beta - n\, \Psi (\alpha )\\\\ \dfrac {\partial \mathcal {L}}{\partial \beta } = \dfrac {1}{\beta ^2}\, \sum ^n_{i = 1} x_i - \dfrac {n\alpha }{\beta }\\ \end {matrix}\] i.e \(\, \Psi (z) = \frac {d}{dz}\, \log \Gamma (z) = \) diagamme function.

\( S(\theta ) = \begin {pmatrix} 0\\ 0\\ \end {pmatrix}\,\) cannot be solved explicitly.

\[I(\theta )= \begin {pmatrix} -\dfrac {\partial ^2\mathcal {L}}{\partial \alpha ^2} & -\dfrac {\partial ^2\mathcal {L}}{\partial \alpha \, \partial \beta }\\\\ -\dfrac {\partial ^2\mathcal {L}}{\partial \beta \, \partial \alpha } & -\dfrac {\partial ^2\mathcal {L}}{\partial \beta ^2}\\ \end {pmatrix}\]

\(-\dfrac {\partial ^2\mathcal {L}}{\partial \alpha ^2} = n\, \Psi '(\alpha )\),    \(-\dfrac {\partial ^2\mathcal {L}}{\partial \alpha \, \partial \beta } = -\dfrac {\partial ^2\mathcal {L}}{\partial \beta \, \partial \alpha } = \dfrac {n}{\beta }\) and

\(-\dfrac {\partial ^2\mathcal {L}}{\partial \beta ^2} = -\left (-\dfrac {2}{\beta ^3} \sum ^n_{i = 1} x_i + \dfrac {n\alpha }{\beta ^2}\right ) = \dfrac {2}{\beta ^3}\, \sum ^n_{i = 1} x_i - \dfrac {n\alpha }{\beta ^2}\).

\[I(\theta ) = \begin {pmatrix} n\, \Psi '(\alpha ) & \dfrac {n}{\beta }\\\\ \dfrac {n}{\beta } & \dfrac {2}{\beta ^3}\sum ^n_{i = 1} x_i - \dfrac {n\alpha }{\beta ^2}\\ \end {pmatrix}.\]

Solve for \(\theta \) iteratively using \[\theta ^{(i + 1)} = \theta ^{(i)} + \left [I\left (\theta ^{(i)}\right )\right ]^{-1}\, S\left (\theta ^{(i)}\right )\] iterate until convergence to obtain \(\widehat {\theta } = \begin {pmatrix} \widehat {\alpha }\\ \widehat {\beta }\\ \end {pmatrix}.\)

\begin {align*} J(\theta ) = E\left [I(\theta ,X)\right ] & = \begin {pmatrix} n\, \Psi '(\alpha ) & \dfrac {n}{\beta }\\\\ \dfrac {n}{\beta } & \dfrac {2}{\beta ^3}\sum ^n_{i = 1} E(X_i) - \dfrac {n\alpha }{\beta ^2}\\ \end {pmatrix} = \begin {pmatrix} n\, \Psi '(\alpha ) & \dfrac {n}{\beta }\\\\ \dfrac {n}{\beta } & \dfrac {2}{\beta ^3}\, n\alpha \beta - \dfrac {n\alpha }{\beta ^2}\\ \end {pmatrix}\\\\ & =\begin {pmatrix} n\, \Psi '(\alpha ) & \dfrac {n}{\beta }\\\\ \dfrac {n}{\beta } & \dfrac {n\alpha }{\beta ^2}\\ \end {pmatrix} = n\,\begin {pmatrix} \Psi '(\alpha ) & \dfrac {1}{\beta }\\\\ \dfrac {1}{\beta } & \dfrac {\alpha }{\beta ^2}\\ \end {pmatrix}. \end {align*} □

Problem 2.2.2. Suppose \(X_1, \, \cdots \, , \, X_n\) is a random sample from the distribution with pdf \[f_{\theta }(x) = \frac {\alpha \beta }{\left (1 + \beta x\right )^{\alpha + 1}}\hspace {0.4cm}, \hspace {0.4cm} x > 0;\hspace {0.4cm} \alpha , \, \beta > 0\] Explain how you find the \(M.L\) estimate of \(\theta = (\alpha , \beta )^t\). Find the Fisher information.

Show solution

Solution.

Finding the estimate

\[\log L = n\log \alpha + n\log \beta - (\alpha +1)\sum _{i=1}^{n}\log \left (1+\beta x_i\right ).\] Differentiating, \[\frac {\partial \log L}{\partial \alpha } = \frac {n}{\alpha } - \sum _{i=1}^{n}\log \left (1+\beta x_i\right ), \qquad \frac {\partial \log L}{\partial \beta } = \frac {n}{\beta } - (\alpha +1)\sum _{i=1}^{n}\frac {x_i}{1+\beta x_i}.\] The first equation solves explicitly for \(\alpha \) in terms of \(\beta \), \[\widehat {\alpha }(\beta ) = \frac {n}{\sum _{i=1}^{n}\log \left (1+\beta x_i\right )} .\] Substituting this into the second leaves a single equation in \(\beta \) alone, which has no closed-form solution and must be solved numerically — Newton’s method or a one-dimensional search. Once \(\widehat {\beta }\) is found, \(\widehat {\alpha }=\widehat {\alpha }(\widehat {\beta })\). This is the standard profile-likelihood device: eliminate the parameter that can be eliminated, and search over what is left.

Fisher information

The calculation is short once one notices that \[U = \log \left (1+\beta X\right ) \sim \text {EXP}(\alpha ),\] since \(P(X>x) = (1+\beta x)^{-\alpha }\) gives \(P(U>t)=e^{-\alpha t}\). Writing \(V=1+\beta X=e^{U}\), we get \(E\left (V^{-1}\right )=\alpha /(\alpha +1)\) and \(E\left (V^{-2}\right )=\alpha /(\alpha +2)\), whence \[E\left (\frac {X}{1+\beta X}\right ) = \frac {1}{\beta }\left (1-\frac {\alpha }{\alpha +1}\right ) = \frac {1}{\beta (\alpha +1)},\] \[E\left (\frac {X^{2}}{\left (1+\beta X\right )^{2}}\right ) = \frac {1}{\beta ^{2}}\left (1 - \frac {2\alpha }{\alpha +1} + \frac {\alpha }{\alpha +2}\right ) = \frac {2}{\beta ^{2}(\alpha +1)(\alpha +2)} .\] Taking second derivatives and expectations, \[J(\theta ) = n\begin {pmatrix} \dfrac {1}{\alpha ^{2}} & \dfrac {1}{\beta (\alpha +1)}\\[8pt] \dfrac {1}{\beta (\alpha +1)} & \dfrac {\alpha }{\beta ^{2}(\alpha +2)} \end {pmatrix}.\]

Questions on this section

Stuck on something here? Ask below and it stays attached to this topic.