2.1 Distribution-Free Statistic over a \(\mathscr {Z}\)

Let \(V_1, V_2, \cdots , V_n\) be random variables with a joint distribution denoted by \(D\) where \(D\) is member of \(\mathscr {Z}\) class of possible joint distribution. Let \(T(V_1, V_2, \cdots V_n)\) denote some statistic based on \(V_1, V_2, \cdots , V_n.\)

Definition 2.1.1. The statistic \(T(V_1, V_2, \cdots , V_n)\) is said to be distribution-free over \(\mathscr {Z}\) if the distribution of \(T(V_1, V_2, \cdots , V_n)\) is the same for every joint distribution in \(\mathscr {Z}\).

(i)
Let \(\mathscr {Z}_1\) denote the collection of joint distributions on \(n\, iid\, N(\mu _0, \sigma ^2)\), \(\, \mu _0\) is known while \(\sigma ^2\) is unknown.

\(D\in \mathscr {Z}_1\, \longrightarrow \, D\) has joint of \(n\)-normal with mean \(\mu _0, \, \sigma ^2\).

Let \(V_1, V_2, \cdots , V_n\) be a random sample from \(N(\mu _0, \sigma ^2)\).

Let \(\overline {V} = \dfrac {1}{n}\sum ^n_i V_{i=1}\hspace {0.3cm}\) and \(\, S^2 = \dfrac {1}{n - 1} \sum ^n_{i = 1} (V_i - \overline {V})^2\) sample mean and sample variance.

Let \(U_1 = \dfrac {\sqrt {n}\, (\overline {V} - \mu _0)}{S}\,\) and \(\, U_2 = \frac {(n-1)\, S^2}{\sigma ^2_0}\) \[U_1 = \frac {\sqrt {n}\,\, (\overline {V} - \mu _0)}{S} \, \thicksim \, t_{n-1}\] for whatever value of \(\sigma ^2\).
\(U_1\) is distribution-free over \(\mathscr {Z}_1\).

(ii)
Let \(\mathscr {Z}_2\) denote a collection of joint distribution of \(n\, iid \, N(\mu , \sigma ^2_0)\), \(\, \sigma ^2_0\) is known. Let \[U_2 = \frac {(n-1)\, S^2}{\sigma ^2_0}\, \thicksim \, \chi ^2_{n-1}\] \(U_2\) is distribution-free over \(\mathscr {Z}_2\).

Definition 2.1.2. Two random variables \(T_1\) and \(T_2\) are said to be equal in distribution denoted by \(T_1\overset {d}{=}T_2\) if they have the same cdf (applies to vectors, i.e we talk of joint pdf).

Remark 2.1.3.

(i)
\(\overset {d}{=}\) is read as “has the distribution as”
(ii)
\(X \overset {d}{=} X\)
(iii)
If \(\, X\overset {d}{=} Y\, \implies \, Y \overset {d}{=} X\)
(iv)
If \(X \overset {d}{=} Y\,\) and \(\, Y\overset {d}{=}Z\, \implies \, X\overset {d}{=} Z\).

Theorem 2.1.4. A random variable \(X\) has a distribution that is symmetric about \(\beta \) if and only if \(X- \beta \overset {d}{=}\beta - X\).

Proof. Suppose \(X\) is symmetrical about \(\beta \), we show that \(X - \beta \overset {d}{=} \beta - X\).
Let \(F(\cdot )\) denote the cdf of \(X\), then \(F_X(\beta + t) = 1 - F\left ((\beta - t)^-\right ) \hspace {0.3cm} \forall t\in \mathbb {R}\hspace {0.3cm}\cdots \hspace {0.3cm} (*)\).
The cdf of \(W = X - \beta \), \[G(w) = P(W\leq w) = P(X-\beta \leq w) = P(X\leq w +\beta ) = F(w + \beta )\] the cdf of \(U = \beta - X\), \[H(w) = P(U\leq w) = P(\beta - X \leq w) = 1 - P(X\leq \beta - w) = 1 - F(\beta - w).\] cdf of \(X - \beta = F(w + \beta )\,\) and cdf of \(\, \beta - X = 1 - F(\beta - w)\).

From \((*)\) we have \(F(w + \beta ) = 1 - F(\beta - w)\) \[\implies \hspace {0.4cm} X - \beta \overset {d}{=} \beta - X\] Conversely, we show that if \(X- \beta \overset {d}{=} \beta - X\) then \(X\) is symmetric about \(\beta \). \begin {align*} F_X(\beta + t) & = P(X\leq \beta + t) = P(X - \beta \leq t)\\ & = P(\beta - X \leq t ) = P(\beta - t \leq X)\\ & = 1 - P(X\leq \beta - t)\\ F_X(\beta + t) & = 1 - F(\beta - t) \end {align*}

\(\implies \hspace {0.3cm} X\) is symmetric about \(\beta \). □

Remark 2.1.5.

(i)
If a random variable \(Y\) has a cdf of the form \(F(Y - \theta )\), then \(\theta \) is known as the location parameter. Therefore if \(X\) has cdf \(F(t)\) and \(Y\) has cdf of \(F(t - \theta )\) then \(X\overset {d}{=} Y - \theta \).
(ii)
If a random variable \(Y\) has a cdf of the form \(F(t/\eta ) \,\hspace {0.2cm} \forall \, 0 < \eta < \infty \), then \(Y\) is said to have a distribution indexed by the sample parameter \(\eta \). Thus is \(X\) has cdf of the form \(F(t)\) and \(Y\) has cdf of the form \(F(t/\eta )\) then \(X\overset {d}{=}Y/\eta \).

Example 2.1.6. Let \(X_{(1)}, X_{(2)}, \cdots , X_{(n)}\) denote the order statistics for a random sample of size \(n\) from normal distribution with mean zero and variance \(\sigma ^2\hspace {0.3cm} (0 < \sigma ^2 < \infty )\) unknown.
Consider the statistic \(Q\) defined by \[Q = \frac {(X_{(n)} + X_{(1)})/2}{X_{(n)} - X_{(1)}}\] If \(Z_1, \, Z_2, \, \cdots ,\, Z_n\) are \(iid\, \, N(0, 1)\) \[Z_1, \, Z_2, \, \cdots , \, Z_n \, \overset {d}{=}\, \frac {X_1}{\sigma }\, , \, \frac {X_2}{\sigma }\, ,\, \cdots \, \, \, \frac {X_n}{\sigma }\] \[Z_{(1)}, \, Z_{(2)}, \, \cdots , \, Z_{(n)}\, \overset {d}{=}\, \frac {X_{(1)}}{\sigma }\, , \, \frac {X_{(2)}}{\sigma }\, ,\, \cdots \, \, \, \frac {X_{(n)}}{\sigma }\] \begin {align*} \frac {(X_{(n)}/\sigma \, + \, X_{(1)}/\sigma )/2}{X_{(n)}/\sigma \, - \, X_{(1)}/\sigma } & \overset {d}{=}\, \frac {(Z_{(n)} + Z_{(1)})/2}{Z_{(n)} - Z_{(1)}}\\\\ & \overset {d}{=} \frac {(X_{(n)} + X_{(1)})/2}{X_{(2)} - X_{(1)}}\\\\ & = \, Q \end {align*}

Thus for \(\sigma ^2 > 0\), the statistics \(Q\) has the same distribution as if it were constructed using observations from \(N(0,1)\). Therefore statistic \(Q\) is distribution free over the class \(\mathscr {Z}_3\) of the joint distribution of \(iid\, \, N(0, \sigma ^2)\).

Remark 2.1.7. Even though statistics \(U_1, \, U_2\) and \(Q\) are distribution-free over the classes \(\mathscr {Z}_1,\, \mathscr {Z}_2\) and \(\mathscr {Z}_3\) respectively we would not typically consider them to be non-parametric statistics. This is so because their distributions are different if the underlying distribution is something other than normal distribution.

Definition 2.1.8. A statistic \(T\) is said to be non-parametric distribution-free if the class \(\mathscr {Z}\) of joint distribution for which \(T\) is distribution-free includes more distribution form than one.

Questions on this section

Stuck on something here? Ask below and it stays attached to this topic.