2.4 Ranking Statistics

We consider another technique for constructing non-parametric distribution free statistics. Let \(X_1, \, X_2, \, \cdots \, ,\, X_n\) be a random sample from a continuous distribution with cdf \(F(x)\), and let \(X_{(1)}, \, X_{(2)}, \cdots \,, \, X_{(n)}\) be the corresponding order statistics.

Definition 2.4.1. The sample observation \(X_i\) is said to have rank \(R_i\) among \(X_1, \, X_2, \, \cdots , X_n\) if \(X_i = X_{R_i}\) provided the \(R_i^{\text {th}}\) order statistic is uniquely defined.
Let \(R = (R_1,\, R_2, \, \cdots , \, R_n)\) where \(R_i\) is the rank of \(X_i\) \(\hspace {0.2cm} R\) is the vector of ranks.

Theorem 2.4.2. Let \(X_1, \, X_2, \, \cdots \, , \, X_n\) be a random sample from continuous distribution and let \(R = (R_1, \, R_2, \, \cdots \, , \, R_n)\) be a vector of ranks. If \(\mathcal {B} = \{r:\, r\) is a permutation of integers \(1,\, 2, \, \cdots ,\, n\}\), then \(R\) is uniformly distributed over \(\mathcal {B}\).

Proof. We note that there are \(n!\) elements in \(\mathcal {B}\). We need to show that \(R\) assumes each of the permutations of \((1,\, 2,\, \cdots ,\, n)\) with probability \(\dfrac {1}{n!}\).

Let \(r = (r_1, \, r_2, \, \cdots , \, r_n)\) be an arbitrary element in \(\mathcal {B}\). Then \begin {align*} P_{pro}(R = r) & = P\left ((X_1,\, X_2,\, \cdots ,\, X_n) = (X_{(r_1)}, \, X_{(r_2)},\, \cdots \,,\,X_{(r_n)}\right )\\ & = P(X_{d_1} < X_{d_2} < X_{d_3} < \cdots < X_{d_n}) \end {align*}

where \(d_i\) is the position of the number \(i\) in the permutation \(r\), for \(i = 1, \, 2, \, \cdots , \, n\).
Note that if \(Y_1, \, Y_2,\, \cdots \, ,\, Y_m\) are \(iid\) random variables and \((\alpha _1, \, \alpha _2, \, \cdots \,, \, \alpha _m)\) is any permutation of the integers \((1, \, 2,\, \cdots \, , \, m)\) \[\left (Y_1, \, Y_2, \, \cdots \cdots \, , \, Y_m\right )\, \overset {d}{=}\, \left (Y_{\alpha _1}, \, Y_{\alpha _2}, \, \cdots \cdots \, , \, Y_{\alpha _m}\right )\] Therefore \(\,\left (X_1, \, X_2, \, \cdots \cdots \, , \, X_m\right )\, \overset {d}{=}\, \left (X_{d_1}, \, X_{d_2}, \, \cdots \cdots \, , \, X_{d_m}\right )\,\) which implies that \[P(X_{d_1} < X_{d_2} < X_{d_3} < \cdots < X_{d_n}) = P(X_1 < X_2 < X_3 < \cdots < X_n).\] Thus \begin {align*} P(R = r) & = P(X_{d_1} < X_{d_2} < X_{d_3} < \cdots < X_{d_n})\\ & = P(X_1 < X_2 < X_3 < \cdots < X_n)\\ & = P(R = r_0)\hspace {0.3cm}, \hspace {0.5cm}\text {where}\hspace {0.2cm} r_0 = (1, \, 2,\, \cdots \, , \, n) \end {align*}

since there are \(n!\) elements in \(\mathcal {B}\), the result follows since \(r\) is an arbitrary element in \(\mathcal {B}\). \[\sum _{r\in \mathcal {B}} P(R = r) = 1.\hspace {1.5cm} n!\, (a) = 1 \, \implies \, a = \frac {1}{n!}\] □

Corollary 2.4.3. Let \(X_1, \, X_2, \, \cdots \, , \, X_n\) be a random sample from a continuous distribution, and let \(R\) be the corresponding vector of ranks, where \(R_i\) is the rank of \(X_i\) among the \(n\) random variables

    • \(P(R_i = r) = \begin {cases} \dfrac {1}{n} & \text {for}\hspace {0.2cm} r = 1, \, 2, \, \cdots \, , \, n\\ 0 & \text {elsewhere} \end {cases}\)
    • \(P(R_i = r\, , \, R_j = s)_{i\neq j} = \begin {cases} \dfrac {1}{n\, (n - 1)} & \text {if}\hspace {0.2cm} r\neq s\\ 0 & \text {elsewhere} \end {cases}\)
    • \(E(R_i) = \dfrac {n + 1}{2} \, \, , \,\, i = 1, \, 2, \, \cdots \cdots \, , \, n\)
    • \(var(R_i) = \dfrac {(n + 1)(n-1)}{12}\)
    • \(Cov(R_i, R_j) = \dfrac {-(n+1)}{12}\hspace {0.2cm}, \hspace {0.3cm} i \neq j.\)

Corollary 2.4.4. Let \(X_1, \, X_2,\, \cdots \, , \, X_n\) be a random sample from continuous distribution and let \(R = (R_1, \, R_2, \, \cdots \, , \, R_n)\) is a vector of ranks. If \(T(R)\) is a statistic based on \((X_1, \, X_2, \, \cdots \, , \, X_n)\), then \(T(R)\) is distribution-free statistics, over the class \(\mathscr {Z}_5\) of joint distributions of \(n\,\, iid\,\) continuous univariate random variables i.e \(T(R)\) is non-parametric distribution-free over \(\mathscr {Z}_5\).

Let \(X_1, \, X_2,\, \cdots \, , \, X_m\) and \(Y_1, \, Y_2,\, \cdots \, , \, Y_n\) be independent random samples from continuous distributions functions \(F(x)\) and \(G(x) = F(x - c)\) respectively where \(-\infty < c < \infty \) known as the shift parameter. \((X_i \overset {d}{=} Y_i - c)\)

\(H_0: \, c = 0\, \) vs \(\, \) \(\,H_{a_1}:\, c < 0\)
\(\,H_{a_2}: \, c > 0\)
\(\,H_{a_3}:\, c\neq 0\).

Let \(Q_i\, , \, \hspace {0.2cm} i = 1, \, 2, \, \cdots \, , \, m\) and \(R_j\, , \, \hspace {0.2cm} j = 1, \, 2, \, \cdots \, , \, n\) be the ranks of \(X_i\) and \(Y_j\) respectively, among \(N = m + n\) combined \(X\) and \(Y\) observations. Thus the rank vector \[R = \left (Q_1, \, Q_2, \, \cdots \, , \, Q_m\, , \, R_1, \, R_2,\, \cdots \, , \, R_n\right )\] is simply a permutation of \((1,\, 2,\, \cdots \, ,\, N)\), where \(N = m + n\) and though random must satisfy \[\sum ^m_{i = 1} Q_i \, + \, \sum ^n_{j = 1} R_j \, = \, \frac {(n + m)(n + m + 1)}{2}.\] To construct a suitable critical region for testing \(H_0 : c = 0\) vs \(H_{a_1}:\, c> 0\) we must decide which rank vector is in support of \(H_{a_1}\). Intuitively to the \(X_i\,'s\). One such statistic is the sum of the ranks for the \(Y_j\, 's\) proposed by Wilcoxon (1945) and studied independently by Mann and Whitney (1947). The test statistic proposed by Wilcoxon is \begin {equation} W_y = \sum ^n_{j = 1} R_j \end {equation} thus \(W_y\) is the sum of the ranks of \(Y_j\, 's\) when ranked among all the \(n + m\) observations. The statistic suggested by Mann and Whitney is \begin {equation} U_y = \sum ^m_{i = 1}\sum ^n_{j = 1} \Psi (Y_j - X_i) \end {equation} where \[\Psi (t) = \begin {cases} 1 & \text {if}\hspace {0.2cm} t > 0\\ 0 & \text {if}\hspace {0.2cm} t\leq 0\\ \end {cases} \]

\(U_y\) represents the total number of times a \(Y\) observation is larger than an \(X\) observation when they are no ties among \(X\) and \(Y\) observation \(W_y\) and \(U_y\) are linearly related by \begin {equation} W_y = U_y \, + \, \frac {n (n + 1)}{2} \end {equation}

Remark 2.4.5.

1.
If \(W_x = \sum ^m_{i = 1} Q_i\) and \(U_x = \sum ^m_{i = 1}\sum ^n_{j = 1} \Psi (X_i - Y_j)\) then \(\, W_x = U_x + \dfrac {m(m + 1)}{2}\)
2.
The statistic \(W_y\) is a function of the data only through the rank vector
\(R = \left (Q_1, \, Q_2, \, \cdots \, , \, Q_m\, , \, R_1, \, R_2,\, \cdots \, , \, R_n\right )\). Under \(H_0:\, c = 0\) observations
\(X_1, \, X_2, \, \cdots \, , \, X_m \, , \, Y_1, \, Y_2, \, \cdots \, , \, Y_n\) are iid continuous random variables and by corollary 2.3.4 the rank statistic \(W_y\) is non-parametric distribution-free.

Theorem 2.4.6. Let \(W_y\) be the Mann-Whitney/Wilcoxon rank sum statistics for testing \(H_0:\, c = 0\, \) vs \(\, H_a:\, c > 0\), when \(X_1, \, X_2, \, \cdots \, , \, X_m\) and \(\, Y_1, \, Y_2, \, \cdots \, , \, Y_n\) are independent random samples from \(F(x)\) and \(F(y - c)\), respectively. Under \(H_0:\, c = 0\) the distribution of \(W_y\) is given \[P_{H_0}\left (W_y = w\right ) = \begin {cases} \dfrac {t_{m,n}(w)}{\binom {N}{n}} & , \, w = \frac {n(n+1)}{2}\, , \, \frac {n(n+1)}{2} + 1,\, \cdots \,,\, \frac {n(n+1)}{2} + nm\\\\ 0 & ,\,\,\text {otherwise}\\ \end {cases}\] where \(t_{m,n}(w)\) is the number of unordered subsets of \(n\) integers taken (without replacement) from \(\{1, \, 2, \, \cdots \, , \, N\}\) for which the sum is equal to \(w, \, N = n + m\).

Example 2.4.7. For case \(m = 3\) and \(n = 2\) there are \(\binom {5}{2} = 10\) arrangements to be examined to obtain the distribution of \(W_y\)

Arrangement
Value of
1 2 3 4 5 \(w\)
\(x\) \(x\) \(x\) \(y\) \(y\) 9
\(x\) \(x\) \(y\) \(x\) \(y\) 8
\(x\) \(y\) \(x\) \(x\) \(y\) 7
\(y\) \(x\) \(x\) \(x\) \(y\) 6
\(x\) \(x\) \(y\) \(y\) \(x\) 7
\(x\) \(y\) \(x\) \(y\) \(x\) 6
\(x\) \(y\) \(y\) \(x\) \(x\) 5
\(y\) \(x\) \(y\) \(x\) \(x\) 4
\(y\) \(y\) \(x\) \(x\) \(x\) 3
\(y\) \(x\) \(x\) \(y\) \(x\) 5
\(w\) \(P\left (W_y = w\right )\)
3 \(1/10\)
4 \(1/10\)
5 \(2/10\)
6 \(2/10\)
7 \(2/10\)
8 \(1/10\)
9 \(1/10\)



\(\alpha = 0.2\). Reject \(H_0\) if \(W_{cal} \geq 8\)

Theorem 2.4.8. Let \(W_y = \sum ^n_{j = 1} R_j\) be the Wilcoxon rank statistic when \(H_0:\, c = 0\) is true, the distribution of \(W_y\) is symmetric about the value \[\mu = \frac {n(m + n + 1)}{2}.\] Given independent random samples \(X_1, \, X_2, \, \cdots \, ,\, X_m\) and \(Y_1, \, Y_2, \, \cdots \, ,\, Y_n\) with rank vector \(R = \left (Q_1, \, Q_2, \, \cdots \, , \, Q_m\, , \, R_1, \, R_2,\, \cdots \, , \, R_n\right )\) for combined sample.

Proof. Recall that under \(H_0\) the rank vector \(R\) is uniformly distributed (theorem 2.3.2), the set of permutations of integers \((1,\, 2, \, \cdots \, , N)\) where \(N= n + m\). We note that under \(H_0: c = 0,\) \(\, \left (-X_1, \,-X_2, \, \cdots \, , \, -X_m, \, -Y_1, \, -Y_2, \, \cdots \, , \, -Y_n\right )\) are \(iid\) continuous random variables and so their rank vector is also uniformly distributed on \(\mathcal {B}\). However multiplying by \(-1\) makes the largest be the smallest and so forth, and hence the reverse the order of ranking. Therefore the rank vector of \(\left (-X_1, \,-X_2, \, \cdots \, , \, -X_m, \, -Y_1, \, -Y_2, \, \cdots \, , \, -Y_n\right )\) written in terms of the components of the original \(R\) is \[\left (N + 1 - Q_1, \, N + 1 - Q_2, \, \cdots \, , \, N + 1 - Q_m\, , \, N + 1 - R_1, \, N + 1 - R_2,\, \cdots \, , \, N + 1 - R_n\right ).\] Therefore \[ \left (Q_1,\, \cdots \, , \, Q_m\, , \, R_1,\, \cdots \, , \, R_n\right ) \, \overset {d}{=}\left (N + 1 - Q_1, \, \cdots \, , \, N + 1 -Q_m\, , \, N + 1 - R_1, \, \cdots \, , \,N + 1 - R_n\right )\] computing the rank sum statistic on each side and applying the theorem which says that if \(Z\overset {d}{=}V\) and \(K(\cdot )\) a measurable function then \(K(Z) \overset {d}{=} K(V)\). \[W_y = \sum ^n_{j = 1} R_j \overset {d}{=} \sum ^n_{j = 1} (N + 1 - R_j)\] \[W_y \overset {d}{=} n(N+ 1) - \sum ^n_{j = 1} R_j = n(N + 1) - W_y\] \[W_y \overset {d}{=} n(N+1) - W_y\] \begin {align*} W_y - \frac {n(N+1)}{2}\, & \overset {d}{=}\, n(N+1)- W_y - \frac {n(N+1)}{2}\\\\ W_y - \frac {n(N+1)}{2}\, & \,\overset {d}{=} \frac {n(N+1)}{2} - W_y \end {align*}

\(\implies \hspace {0.3cm} W_y\,\) is symmetric at \(\, \dfrac {n (N + 1)}{2}\). □

Remark 2.4.9.

1.
For testing \(H_0:\, c = 0\) vs \(H_a:\, c < 0\), we want critical region of the form. Reject \(H_0\) if and only if \(W_y \leq W'(\alpha , m, n)\) value where \(\alpha \) is the level of significance and \(W'(\alpha , m, n)\) is the value such that \[P\left (W_y \leq W'(\alpha , m, n)\right ) = \alpha .\] From the proof of theorem 2.3.7 we see that \[W'(\alpha , m, n) = n(m + n + 1) - W(\alpha , m , n)\] where \(P\left (W_y\geq W(\alpha , m, n)\right ) = \alpha \).
2.
In the similar way, the critical region for testing \(H_0: c = 0\) vs \(H_a: c \neq 0\), reject \(H_0\) at \(\alpha \) level of significance if and only if \(W_y \geq W(\alpha /2, m , n)\) or \(\, W_y \leq W'(\alpha /2, m, n)\) or \(\, W_y \leq n(n + m + 1) - W(\alpha /2, m , n)\), or equivalently \[\begin {vmatrix} W_y - \dfrac {n(m + n + 1)}{2}\\ \end {vmatrix}\, \geq \, \begin {vmatrix} W(\alpha /2, m , n) - \dfrac {n(m + n + 1)}{2}\\ \end {vmatrix}.\]
3.
Theorem 2.3.7 also provides us with the mean of \(W_y\) under \(H_0: c = 0\)

Questions on this section

Stuck on something here? Ask below and it stays attached to this topic.