2.5 Statistics Utilising Counting and Ranking
A third basic technique for constructing distribution-free statistics combines the two techniques
considered in sections 2.2 and 2.3.
Let \(Z_1, \, Z_2, \, \cdots \, , \, Z_n\) be a random sample from a continuous distribution that is symmetric about zero. Note that
the point of symmetry could be any known value. If \(X\) is symmetrically distributed about \(\theta _0\), then \(Z = X - \theta _0\) is
symmetrically distributed at zero.
Definition 2.5.1. Define \(\, \Psi _i = \Psi (Z_i)\), where \[\Psi (t) = \begin {cases} 1 & \text {if}\hspace {0.3cm} t > 0\\ 0 & \text {if}\hspace {0.3cm} t \leq 0\\ \end {cases}\]
Lemma 2.5.2. Let \(Z\) be a continuous random variable that is symmetric about zero, then the random variables \(|Z|\) and \(\Psi (Z)\) are stochasticaly independent.
Proof. \begin {align*} P\left (\Psi (Z) = 1\, , \, |Z| \leq t\right ) & = P\left (Z>0\, , \, |Z| \leq t\right )\\ & = P\left (0 < Z \leq t\right )\\ & = \frac {1}{2}\, P\left (-t \leq Z \leq t\right )\\ & = P\left (\Psi (Z) = 1\right )\, P\left (|Z|\leq t\right )\\ P\left (\Psi (Z) = 1\, , \, |Z| \leq t\right ) & = P\left (\Psi (Z) = 1\right )\, P\left (|Z|\leq t\right )\hspace {0.3cm}\cdots \cdots \hspace {0.3cm}* \end {align*}
\begin {align*} P\left (\Psi (Z) = 0\, , \, |Z| \leq t\right ) & = P\left (|Z|\leq t\right ) - P\left (\Psi (Z) = 1\, , \, |Z|\leq t\right )\\ & = P\left (|Z|\leq t\right ) - P\left (\Psi (Z) = 1\right )\, P\left (|Z| \leq t\right )\\ & = P\left (|Z|\leq t\right )\, \left (1 - P\left (\Psi (Z) = 1\right )\right )\\ & = P\left (|Z|\leq t\right )\, P\left (\Psi (Z) = 0\right ).\\ P\left (\Psi (Z) = 0\, , \, |Z| \leq t\right ) & = P\left (|Z|\leq t\right )\, P\left (\Psi (Z) = 0\right )\hspace {0.3cm}\cdots \cdots \hspace {0.3cm} ** \end {align*}
Thus from \(*\) and \(**\hspace {0.3cm} \Psi (Z)\) and \(|Z|\) are independent. □
Definition 2.5.3. For any random variables \(Z_1, \, Z_2, \, \cdots \, , \, Z_n\), the absolute rank of \(Z_i\), denoted by \(R^+_i\), is the rank
of \(|Z_i|\) among \(|Z_1|, \, |Z_2|, \, \cdots \, , \, |Z_n|\). The signed rank of \(Z_i\) is then \(\Psi (Z_i)\, R^+_i \, = \, \Psi _i\, R^+_i, \) where \(\Psi _i = \Psi (Z_i)\).
\(\Psi = \left (\Psi _1, \, \Psi _2, \, \cdots \, , \, \Psi _n\right )\) and \(R^+ = \left (R^+_1, \, R^+_2, \, \cdots \,, \, R_n^+\right )\). The statistics that involve \(\Psi _i\, 's\) and \(R^+_i\, 's\) are usually written in terms of signed ranks \(\Psi _1\, R^+_1, \, \Psi _2\,R^+_2, \, \cdots \cdots \, , \, \Psi _n\, R^+_n\), such statistics
are referred to as signed rank statistics.
The following theorem establishes an important joint distributed structure for \(\Psi \) and \(R^+\).
Theorem 2.5.4. Let \(Z_1, \, Z_2, \, \cdots \, , \, Z_n\) be a random sample from a continuous distribution that is symmetric at zero. Let \(R^+\) denote the vector of absolute ranks of the \(Z_i\, 's\). If \(\Psi _i = \Psi (Z_i), \hspace {0.2cm} i = 1, \, 2, \, \cdots \, , \, n\) then
- (a)
- \(\Psi _1, \, \Psi _2, \, \cdots \, , \, \Psi _n\, , \, R^+\) are mutually independent.
- (b)
- Each \(\Psi _i\) is a Bernoulli random variable with \(P = \frac {1}{2}\).
- (c)
- \(R^+\) is uniformly distributed over \(\mathcal {B}\) , the set of all permutations of \((1, \, 2, \, \cdots \, , \, n)\).
Proof.
- (a)
- Since \(Z_1, \, Z_2, \, \cdots \, , \, Z_n\) are independent it follows that \(|Z_1|\, \Psi _1, \, |Z_2|\,\Psi _2, \, \cdots \, , \, |Z_n|\,\Psi _n\) are \(2n\) mutually independent random variables from Lemma 2.4.2. The vector \(R^+\) is independent of the \(\Psi _i\, 's\) since it is a function of \(|Z_i|\) which are independent of the \(\Psi _i\, s\).
- (b)
- Each \(\Psi _i\) is a Bernoulli random variable with parameter \(P = P(Z_i > 0) = \frac {1}{2}\) because \(Z_i\) is continuous and symmetrically distributed about 0.
- (c)
- Since \(R^+\) is a rank vector of \(n\,\) iid continuous random variables, then \(R^+\) is uniformly distributed over \(\mathcal {B}\) by theorem 2.3.2.
Corollary 2.5.5. Let \(T(\Psi , R^+)\) be a statistic that depends on the observations \(Z_1, \, Z_2, \, \cdots \, , \, Z_n\) only through \(\Psi _1, \, \Psi _2, \, \cdots \, , \, \Psi _n\, , \, R^+\), then the statistic \(T(\cdot )\) is non-parametric distribution-free \(\mathscr {Z}_6\) (the collection of joint distributions symmetrically about zero continuous).
Proof. This result follows from Theorem 2.4.4 because \(\Psi \) and \(R^+\) have the same joint distribution \(F(Z_1, \, Z_2, \, \cdots \, , \, Z_n)\in \mathscr {Z}_6\). That is, the joint distribution of \(\Psi \) and \(R^+\) does not depend on the choices of \(F(Z_1, \, Z_2, \, \cdots \, , \, Z_n) \in \mathscr {Z}_6\). □
Theorem 2.5.6. Let \(Z_1, \, Z_2, \, \cdots \, , \, Z_n\) be a random sample from continuous distribution about zero. Let \(Q\) be the number of positive \(Z_i\, 's\) and \(Q = q\), let \(S_1 < S_2 < \cdots < S_q\) denote the ordered absolute ranks of those \(Z_i\, 's\) that are positive, then \[P_{pro}\left (Q = q, \, S_1 = s_1, \, S_2 = s_2, \, \cdots \, , \, S_q = s_q\right ) = \left (\dfrac {1}{2}\right )^n\] for \(q = 0, \, 1,\, 2, \, \cdots \,, \, n.\)
Proof. Let \(q\in \{0, \, 1, \, 2,\, \cdots \, , \, n\}\) and let \(\left (S_1,\, S_2, \, \cdots \, , \, S_q\right )\) be an arbitrary \(q-\)tuple such that \(s_i\) is an integer and \(1 \leq S_1 < S_2 < \cdots < S_q \leq n\). Define \(S^*_{q + 1} < S^*_{q + 2} < \cdots < S^*_n\) to be the
integers written in numerical order in \(\{1,\, 2,\, \cdots \,,\, n\}\) that are not included in \(\{s_1,\, s_2, \, \cdots \, ,\, s_n\}\), then theorem 2.4.4 we see
that the independence of \(\Psi = \left (\Psi _1, \, \Psi _2, \, \cdots \, , \, \Psi _n\right )\) and \(R^+\) and the corresponding distribution of each yields.
\[P\left (\Psi _1 = 1, R^+_1 = s_1, \, \cdots \, \Psi _q = 1\, , \, R^+_q = s_q\, , \, \Psi _{q + 1} = 0,\,R^+_{q + 1} = S^*_{q + 1},\, \cdots \, , \, \Psi _n = 0,\, R^+_n = S^*_n\right )\]
\[= \left (\frac {1}{2}\right )^n\cdot \frac {1}{n!}\]
However, this is only one \(\Phi \) and \(R^+\) combination that yields \(\left (Q = q, \, S_1 = s_1, \, S_2 = s_2, \, \cdots \, , \, S_q = s_q\right )\) there are \(\binom {n}{q}\,q!\, (n - q)! = n!\) such combinations.
Moreover, theorem 3.4.4 shows that each of these are equally likely \begin {align*} P\left (Q = q, \, S_1 = s_1, \, S_2 = s_2, \, \cdots \, , \, S_q = s_q\right ) & = \left (\frac {1}{2}\right )^n\, \left (\frac {1}{n!}\right )\, n! = \left (\frac {1}{2}\right )^n. \end {align*}
□
Remark 2.5.7. If \(X_1, \, X_2, \, \cdots \, , \, X_n\) is a random sample from a continuous distribution that is symmetric
about a known point \(\theta _0\), then the variables \(Z_i = X_i - \theta _0, \hspace {0.2cm} i = 1, \, 2, \, \cdots \, , \, n\) form a random sample from a continuous
distribution that is symmetric at zero. Thus all the theorems and corollaries in this section
2.4 apply equally well to the \((X_i - \theta _0)'s\). Consequently we talk about distribution-free signed rank
statistics over the class of continuous univariate distributions that are symmetric about \(\theta _0\),
where \(\theta _0\) is known real number.
If \(X_1, \, X_2, \,\cdots \, , \, X_n\) is a random sample from a continuous distribution that is symmetric about unknown,
known median \(\theta \in (-\infty ,\infty )\) and has distribution \(F(x)\). Let \(\theta _0\) be known, and consider testing \(H_0:\, \theta = \theta _0\) vs \(H_{a_1}:\, \theta > \theta _0\) or \(H_{a_2}:\, \theta < \theta _0\) or \(H_{a_3}:\, \theta \neq \theta _0\).
Define \(Z_i = X_i - \theta _0, \hspace {0.3cm} i = 1, \, 2, \, \cdots \, , \, n\) and let \(\Psi _1\, R^+_1, \, \Psi _2\, R^+_2, \, \cdots \, , \, \Psi _n\, R^+_n\) be the signed ranks of \(Z_1, \, Z_2, \, \cdots \, , \, Z_n\) respectively.
The Wilcoxon signed rank statistic is \[W^+ = \sum ^n_{i = 1} \Psi _i\, R^+_i, \] \(W^+\) is the sum of the signed ranks for the \(Z_i\, 's\) which are greater that zero for testing \(H_0:\, \theta = \theta _0\) vs \(H_a:\, \theta > \theta _0\). Reject \(H_0\) if \(W^+ \geq W^+(\alpha , n)\) where \(P\left (W^+\geq W^+(\alpha , n)\right ) = \alpha \). To obtain the necessary values of \(W^+(\alpha ,n)\) we need to establish a method of obtaining the form of the null distribution of \(W^+\).
Corollary 2.5.8. Let \(W^+\) be the Wilcoxon signed rank statistic for testing \(H_0: \theta = \theta _0\). For a random sample of size \(n\), the distribution of \(W^+\) under \(H_0\) is \[P(W^+ = k) = \begin {cases} \dfrac {C_n(k)}{2^n} & , \, k = 0, \, 1, \, \cdots \, , \, \dfrac {n(n+1)}{2}\\\\ 0 & \text {elswhere}\\ \end {cases}\] where \(C_n(k) = \)number of subsets of integers from \(\{1, \, 2, \, \cdots \, , \, n\}\) for which the sum is equal to \(k\). Note that these subsets can contain anywhere from 0 to \(n\) integers.
Remark 2.5.9. As was the case for the rank sum statistics \(W\)(Wilcoxon signed rank), the null
distribution of \(W^+\) given in corollary 2.4.7 has no closed form, i.e no closed for \(C_n(k)\). Therefore to
evaluate \(P_{H_0}(W^+ = k)\) for a particular sample size \(n\), we must calculate the value of \(W^+\) for each of \(2^n\) possible
subsets of \(\{1, \, 2, \, \cdots \, , \, n\}\) and tabulate \(C_n(k)\).
Let \(n = 5\). Then \(\{1, \, 2, \, 3, \, 4, \, 5\}\)
| \(w^+\) | \(w^+\) | \(w^+\) | |||||
| \(\{\}\) | 0 | \(\{2,\, 3\}\) | 5 | \(\{1,\, 3,\, 5\}\) | 9 | ||
| \(\{1\}\) | 1 | \(\{2, \, 4\}\) | 6 | \(\{1, \, 4, \, 5\}\) | 10 | ||
| \(\{2\}\) | 2 | \(\{2,\, 5\}\) | 7 | \(\{2, \, 3,\, 4\}\) | 9 | ||
| \(\{3\}\) | 3 | \(\{3,\, 4\}\) | 7 | \(\{2,\, 3, \, 5\}\) | 10 | ||
| \(\{4\}\) | 4 | \(\{3,\, 5\}\) | 8 | \(\{2, \, 4, \, 5\}\) | 11 | ||
| \(\{5\}\) | 5 | \(\{4,\, 5\}\) | 9 | \(\{3,\, 4,\, 5\}\) | 1 | ||
| \(\{1,\, 2\}\) | 3 | \(\{1, \, 2, \, 3\}\) | 6 | \(\{1, \, 2, \, 3, \, 4\}\) | 10 | ||
| \(\{1,\, 3\}\) | 4 | \(\{1, \, 2, \, 4\}\) | 7 | \(\{1, \, 2, \, 3, \, 5\}\) | 11 | ||
| \(\{1, \, 4\}\) | 5 | \(\{1, \, 2,\, 5\}\) | 8 | \(\{1, \, 2, \, 4, \, 5\}\) | 12 | ||
| \(\{1,\, 5\}\) | 6 | \(\{1, \, 3,\, 4\}\) | 8 | \(\{1, \, 3, \, 4, \, 5\}\) | 13 | ||
| \(\{2,\, 3, \, 4, \, 5\}\) | 14 | ||||||
| \(\{1, \, 2,\, 3, \, 4, \, 5\}\) | 15 |
| \(k\) | 0 | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 | 11 | 12 | 13 | 14 | 15 |
| \(C_5(k)\) | 1 | 1 | 1 | 2 | 2 | 3 | 3 | 3 | 3 | 3 | 3 | 2 | 2 | 1 | 1 | 1 |
| \(P_{H_0}(W^+ = k)\) | \(\frac {1}{32}\) | \(\frac {1}{32}\) | \(\frac {1}{32}\) | \(\frac {2}{32}\) | \(\frac {2}{32}\) | \(\frac {3}{32}\) | \(\frac {3}{32}\) | \(\frac {3}{32}\) | \(\frac {3}{32}\) | \(\frac {3}{32}\) | \(\frac {3}{32}\) | \(\frac {2}{32}\) | \(\frac {2}{32}\) | \(\frac {1}{32}\) | \(\frac {1}{32}\) | \(\frac {1}{32}\) |
The null distribution of \(W^+\) has been tabulated for various values of \(n\) upto \(n\leq 50\) by different Authors. For
sample sizes larger 15 there is a good approximation to the exact null distribution of
\(W^+\).
Let us consider the problem of paired observations. Let \((X_1,Y_1),\, (X_2,Y_2), \, \cdots \, , \, (X_n,Y_n)\) be a random sample from a continuous bivariate distribution. Define \(Z_i = Y_i - X_i\) \[H_0:\, \theta = \theta _0\hspace {0.3cm}\text {vs}\hspace {0.3cm} H_a:\, \theta > \theta _0, \hspace {0.3cm} \theta < \theta _0, \hspace {0.3cm} \theta \neq \theta _0\]
Example 2.5.10. In an investigation to determine the effect of asprin bleeding time and platelet adhesion, a study was done on the reaction of of 14 “normal” subjects to asprin. Let \(X\) observation for each subject is the bleeding tine (in seconds) before ingestion of asprin and \(Y\) observation is the bleeding time (in seconds) two hours after and administration of asprin.
| subject | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 | 11 | 12 | 13 | 14 |
| \(Y\) | 525 | 570 | 190 | 395 | 370 | 210 | 490 | 250 | 360 | 285 | 630 | 385 | 195 | 295 |
| \(X\) | 270 | 150 | 270 | 420 | 202 | 255 | 165 | 220 | 305 | 210 | 240 | 300 | 300 | 70 |
| \(|Y - X|\) | 255 | 420 | 80 | 25 | 168 | 45 | 325 | 30 | 55 | 75 | 390 | 85 | 105 | 225 |
| \(\Psi _i\) | 1 | 1 | 0 | 0 | 1 | 0 | 1 | 1 | 1 | 1 | 1 | 1 | 0 | 1 |
| \(R^+_i\) | 11 | 14 | 6 | 1 | 9 | 3 | 12 | 2 | 4 | 5 | 13 | 7 | 8 | 10 |
\begin {align*} W^+ & = 11 + 14 + 9 + 12 + 2 + 4 + 5 + 13 + 7 + 10\\ & =25 + 23 + 39\\ W^+ & = 87 \end {align*}
\(H_0:\, \theta = 0\,\) vs \(\, H_a:\, \theta > 0\) \[W^+ = \sum ^n_{i=1} \sum ^n_{i=1} \Psi _iR^+_i\] Reject \(H_0\) at \(\alpha = 0.05\), if \(W^+_{cal} \geq W_{(0.05, 14)}\) \[W^+_{cal} =87.\]
Theorem 2.5.11. Let \(W^+\) be the Wilcoxon signed rank statistic. When \(H_0:\, \theta = \theta _0\) is true, the distribution of \(W^+\) is symmetric about its mean \(\mu = \dfrac {n(n + 1)}{4}\).
Proof. Theorem 2.4.4 show that \(\Psi _1, \, \Psi _2, \, \cdots \, , \, \Psi _n\, , \, R^+\) are mutually independent random variables and that the \(\Psi _i\) is Bernoulli variate with \(P = \frac {1}{2}\) for \(i = 1, \, 2, \, \cdots \, , \, n\). Since \(1 - \Psi _i\) is also Bernoulli variate with \(P = \frac {1}{2}\), then \[\left (\Psi _1, \, \Psi _2, \, \cdots \, , \, \Psi _n\, , \, R^+\right ) \, \overset {d}{=}\, \left (1-\Psi _1, \, 1-\Psi _2, \, \cdots \, , \, 1-\Psi _n\, , \, R^+\right )\] computing \(W^+\) on both sides \[\sum ^n_{i = 1}\Psi _i\, R^+_i \, \overset {d}{=}\, \sum ^n_{i = 1} \left (1 - \Psi _i\right )\, R^+_i\] then apply the result used in proof of theorem 2.3.7 (if \(Z \overset {d}{=} V\) and \(K(\cdot )\) is a measurable function then \(K(Z) \overset {d}{=} K(V)\)) yields \[W^+ \, \overset {d}{=}\, \sum ^n_{i = 1} R^+_i - W^+\] \[W^+ \, \overset {d}{=}\, \frac {n(n + 1)}{2} - W^+\] \[W^+ - \frac {n(n+1)}{4}\, \overset {d}{=}\, \frac {n(n+1)}{2} - W^+ - \frac {n(n+1)}{4}\] \[W^+ - \frac {n(n+1)}{4}\, \overset {d}{=}\, \frac {n(n+1)}{4} - W^+.\] Thus \(W^+\) is symmetric about \(\dfrac {n(n+1)}{4}\) by theorem \(2.1.3\). □
- (i)
- \(E_{H_0}(W^+) = \dfrac {n(n+1)}{4}\)
- (ii)
- It can be shown that \(\hspace {0.2cm} var_{H_0}(W^+) = \dfrac {n(n+1)(2n+1)}{24}\)
Questions on this section
Stuck on something here? Ask below and it stays attached to this topic.