5.8 Practice Problems
Problem 5.4. Treatment of terminal renal failure involves surgical removal of a kidney. Mean diastolic blood pressure was recorded for five patients before and two months after surgery.
| Patient | 1 | 2 | 3 | 4 | 5 |
| Before surgery | 107 | 102 | 95 | 106 | 112 |
| After surgery | 87 | 97 | 101 | 113 | 80 |
The study aimed to establish whether diastolic pressure before surgery exceeds the reading two months after.
- (a).
- State hypotheses consistent with the aim of the study.
- (b).
- What is the study design?
- (c).
- Derive the null distribution of the Wilcoxon signed-rank statistic.
- (d).
- Test at the \(5\%\) level using the \(P\)-value.
Show solution
Solution. (a). With \(\theta \) the median of the differences \(d = \text {before} - \text {after}\), \[H_0 : \theta = 0 \qquad \text {against}\qquad H_1 : \theta > 0 ,\] one-sided, because the study asks specifically whether pressure falls after surgery.
(b). A paired design — each patient is measured twice and serves as his own control, which removes between-patient variation in blood pressure.
(c). The differences are \[20,\quad 5,\quad -6,\quad -7,\quad 32 ,\] with absolute values \(20, 5, 6, 7, 32\) ranking \(4, 1, 2, 3, 5\). Hence \[T_{+} = 4 + 1 + 5 = 10, \qquad T_{-} = 2 + 3 = 5, \qquad T_{+}+T_{-} = 15 = \tfrac {5(6)}{2}\ \checkmark \]
Under \(H_0\) each of the five ranks is equally likely to carry a \(+\) or a \(-\) independently, so \(T_{+}\) is the sum of a random subset of \(\{1,2,3,4,5\}\) and its null distribution is uniform over the \(2^{5} = 32\) subsets: \[\begin {array}{c|cccccc} t & 15 & 14 & 13 & 12 & 11 & 10\\\hline P(T_{+} \geq t) & \frac {1}{32} & \frac {2}{32} & \frac {3}{32} & \frac {5}{32} & \frac {7}{32} & \frac {10}{32} \end {array}\]
(d). The alternative \(\theta > 0\) makes \(T_{+}\) large, so \[P\text {-value} = P\left (T_{+} \geq 10\right ) = \dfrac {10}{32} = 0.3125 .\] Since \(0.3125 > 0.05\) we do not reject \(H_0\): these data give no evidence that surgery lowers diastolic pressure.
Why five pairs could hardly have shown anything. The smallest attainable one-sided \(P\)-value at \(n' = 5\) is \(\frac {1}{32} = 0.03125\), reached only when \(T_{+} = 15\) — that is, when all five differences are positive. With even one negative difference the test cannot reject at the \(5\%\) level, whatever the magnitudes. Here two of the five went the wrong way, so the outcome was settled before any arithmetic.
Problem 5.5. To compare two racing starts, the hole entry and the flat entry, ten college swimmers were timed to water entry with each.
| Swimmer | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 |
| Flat entry | 1.13 | 1.11 | 1.18 | 1.26 | 1.16 | 1.41 | 1.43 | 1.25 | 1.33 | 1.36 |
| Hole entry | 1.07 | 1.03 | 1.21 | 1.24 | 1.33 | 1.42 | 1.35 | 1.32 | 1.31 | 1.33 |
Determine whether the two methods differ, using (a) the sign test and (b) the Wilcoxon signed-rank test. State the hypotheses, the critical values and the decision in each case.
Show solution
Solution. The differences \(d = \text {flat} - \text {hole}\) are \[0.06,\ 0.08,\ -0.03,\ 0.02,\ -0.17,\ -0.01,\ 0.08,\ -0.07,\ 0.02,\ 0.03 .\] None is zero, so \(n' = 10\) throughout. Both tests take \[H_0 : \theta = 0 \qquad \text {against}\qquad H_1 : \theta \neq 0 .\]
(a) Sign test. Six differences are positive and four negative, so \(N_{+} = 6\). Under \(H_0\), \(N_{+} \sim B\!\left (10,\tfrac 12\right )\) and \[P\text {-value} = 2\,P\left (N_{+} \geq 6\right ) = 2\cdot \dfrac {210+120+45+10+1}{1024} = 2(0.3770) = 0.7539 .\] Nowhere near significant; do not reject.
(b) Wilcoxon signed-rank test. Ranking the \(\left |d\right |\) with ties averaged: \[\left |d\right |:\ 0.01,\ 0.02,\ 0.02,\ 0.03,\ 0.03,\ 0.06,\ 0.07,\ 0.08,\ 0.08,\ 0.17\] \[\text {ranks}:\ 1,\ 2.5,\ 2.5,\ 4.5,\ 4.5,\ 6,\ 7,\ 8.5,\ 8.5,\ 10\] Attaching signs, \[T_{+} = 6 + 8.5 + 2.5 + 8.5 + 2.5 + 4.5 = 32.5, \qquad T_{-} = 4.5 + 10 + 1 + 7 = 22.5,\] and \(32.5 + 22.5 = 55 = \frac {10(11)}{2}\) as required.
The two-sided \(5\%\) critical value at \(n' = 10\) is \(T \leq 8\). Here \(T = \min (32.5, 22.5) = 22.5\), far above it, so again do not reject; the exact two-sided \(P\)-value is about \(0.625\).
Conclusion. Neither test finds a difference in time to water entry.
What the comparison shows. Both tests agree, but not equally loudly: the sign test returns \(P = 0.754\) and the signed-rank test \(P = 0.625\). The signed-rank test uses the sizes of the differences as well as their directions, so it extracts more from the same data and its \(P\)-value is the smaller. When both are applicable the signed-rank test is the more powerful, and the sign test is reserved for data where only the direction is meaningful.
Problem 5.6. Wool yield (kg) of three groups of ewe lambs under different feeding regimes:
| Feed I | Feed II | Feed III |
| (pasture) | (pasture + concentrates) | (pasture + concentrates + minerals) |
| 8.78 | 5.16 | 9.92 |
| 6.23 | 9.64 | 14.74 |
| 7.65 | 6.52 | 7.93 |
| 5.10 | 5.38 | 11.90 |
| 5.95 | 6.80 | 10.77 |
| 8.50 | 7.37 | |
| 10.20 | 7.08 | |
| 5.67 | ||
- (a).
- State the study design and the appropriate test.
- (b).
- What assumptions does that test make?
- (c).
- State the hypotheses.
- (d).
- Do the three feeds affect wool yield? Use \(\alpha = 0.05\).
Show solution
Solution. (a). A completely randomised design with three independent groups of unequal size. The distribution-free test for \(k\) independent samples is Kruskal–Wallis. Friedman would be wrong here: there are no blocks, and the lambs in different groups are unrelated.
(b). Independent random samples; at least ordinal measurement; and populations that are identical under \(H_0\), so that in particular they have the same shape and spread and differ only in location if they differ at all.
(c). \(H_0\): the three feeds give identical yield distributions. \(H_1\): at least one differs.
(d). With \(n_1 = 7\), \(n_2 = 5\), \(n_3 = 8\) and \(N = 20\), ranking all twenty yields together gives rank sums \[R_1 = 67,\qquad R_2 = 35,\qquad R_3 = 108,\] and the check holds: \(67 + 35 + 108 = 210 = \frac {20(21)}{2}\). There are no ties, so no correction is needed. Then \[H = \dfrac {12}{20(21)}\left [\dfrac {67^{2}}{7} + \dfrac {35^{2}}{5} + \dfrac {108^{2}}{8}\right ] - 3(21) = 3.9796 .\] With \(k - 1 = 2\) degrees of freedom, \(\chi ^{2}_{0.05,2} = 5.9915\). Since \(3.9796 < 5.9915\) we do not reject \(H_0\); the exact \(P\)-value is \(0.1367\).
On this evidence the three feeds cannot be shown to differ.
A caution about reading that. The mean ranks are \(9.6\), \(7.0\) and \(13.5\), and Feed III looks the best of the three. The test says only that five, seven and eight animals are too few to establish it — not that the feeds are equivalent. Twenty lambs is a small experiment for a three-way comparison.
Problem 5.7. The following are the ranks awarded to seven debaters by two judges.
| Debater | A | B | C | D | E | F | G |
| Judge I | 3 | 2 | 1 | 6 | 7 | 4 | 5 |
| Judge II | 5 | 6 | 3 | 7 | 4 | 2 | 1 |
- (a).
- Explain the difference between association and correlation.
- (b).
- Which measure of association suits the agreement between the judges?
- (c).
- State its assumptions.
- (d).
- Test the significance of the relationship at the \(5\%\) level.
Show solution
Solution. (a). There is association between two variables when knowing one tells you something about the other. There is correlation when that association is monotone — in the parametric case, linear. Correlation is a special case; association is the wider idea, and two variables can be strongly associated with zero correlation.
(b). The data are ranks, so no scale of measurement is available and Pearson’s coefficient on the raw scores is not an option. Spearman’s rank correlation \(r_s\) is the natural measure. Kendall’s \(\tau \) would serve equally and is preferable if ties were present.
(c). The seven pairs are a random sample; the two members of each pair are measured on the same subject; and the data are at least ordinal.
(d). The differences and their squares: \[d = -2,\ -4,\ -2,\ -1,\ 3,\ 2,\ 4, \qquad \sum d^{2} = 54 .\] With \(n = 7\), \[r_s = 1 - \dfrac {6\sum d^{2}}{n^{3}-n} = 1 - \dfrac {6(54)}{343-7} = 1 - \dfrac {324}{336} = 0.0357 .\]
Testing \(H_0 : \rho _s = 0\) against \(H_1 : \rho _s \neq 0\), the exact null distribution over the \(7! = 5040\) equally likely rankings gives a two-sided \(P\)-value of \(0.906\). We do not reject \(H_0\).
The two judges agree almost not at all: \(r_s = 0.036\) is as close to independent ranking as this data could produce.
A misprint in the question. It asks to test “\(H_0 : \rho \neq 0\)”. A null hypothesis is a statement of no effect and must be an equality; the hypotheses are \(H_0 : \rho _s = 0\) against \(H_1 : \rho _s \neq 0\). Testing an inequality as a null is not something the machinery can do.
Problem 5.8. The Wall Street Journal has periodically run a stock-picking contest. Four Wall Street professionals each selected one stock, four randomly chosen readers each selected one, and four stocks were selected by throwing darts at a list. The percentage returns from March to August 2001 were:
| Experts | Readers | Darts |
| 39.5 | \(-31.0\) | 39.0 |
| \(-1.1\) | \(-20.7\) | 31.9 |
| \(-4.5\) | \(-45.0\) | 14.1 |
| \(-8.0\) | \(-73.3\) | 5.4 |
- (a).
- Which assumptions are appropriate here?
- (b).
- State the hypotheses.
- (c).
- Compute the test statistic and give the critical value at \(\alpha = 0.05\).
- (d).
- Is there evidence of a difference in median return between the three categories?
Show solution
Solution. (a). The choice of test is the first thing to get right. The three columns are independent groups: four different professionals, four different readers, twelve different stocks. Nothing pairs the first expert with the first reader, so the rows are not blocks and Friedman does not apply. The appropriate test is Kruskal–Wallis for \(k = 3\) independent samples.
The assumptions are then: independent random samples; at least ordinal measurement; and populations identical under \(H_0\).
(b). \(H_0\): the three selection methods give identical return distributions. \(H_1\): at least one differs.
(c). Ranking all \(N = 12\) returns together, smallest first: \[-73.3,\ -45.0,\ -31.0,\ -20.7,\ -8.0,\ -4.5,\ -1.1,\ 5.4,\ 14.1,\ 31.9,\ 39.0,\ 39.5\] gives rank sums \[R_{\text {experts}} = 12+7+6+5 = 30, \qquad R_{\text {readers}} = 3+4+2+1 = 10, \qquad R_{\text {darts}} = 11+10+9+8 = 38,\] and \(30+10+38 = 78 = \frac {12(13)}{2}\). There are no ties. Then \[H = \dfrac {12}{12(13)}\left [\dfrac {30^{2}}{4}+\dfrac {10^{2}}{4}+\dfrac {38^{2}}{4}\right ] - 3(13) = 47 - 39 = 8.0000 .\] With \(k-1 = 2\) degrees of freedom, \(\chi ^{2}_{0.05,2} = 5.9915\).
(d). Since \(8.0000 > 5.9915\) we reject \(H_0\); the exact \(P\)-value is \(0.0183\). There is evidence that the three methods differ.
Reading the result. The rank sums say where the difference lies: the darts scored highest (\(38\)), the experts next (\(30\)) and the readers lowest (\(10\)). Over this period the darts beat both the professionals and the readers. That is the finding the contest was designed to provoke, and Kruskal–Wallis confirms it is more than chance — though with four stocks per method it says nothing about why, and the Dow gained \(2.4\%\) over the same window, so only the darts beat the market at all.
Problem 5.9. Let \(X_1,\ldots ,X_n\) and \(Y_1,\ldots ,Y_m\) be random samples from continuous populations \(F\) and \(G\). The Mann–Whitney statistics are \[U_n = nm + \dfrac {n(n+1)}{2} - S_n, \qquad U_m = nm + \dfrac {m(m+1)}{2} - S_m ,\] where \(S_n\) and \(S_m\) are the rank sums of the \(X\)’s and \(Y\)’s in the combined sample. Show that \[U_n = nm - U_m .\]
Show solution
Solution. Add the two definitions: \[U_n + U_m = 2nm + \dfrac {n(n+1)}{2} + \dfrac {m(m+1)}{2} - \left (S_n + S_m\right ).\]
Every rank from \(1\) to \(N = n+m\) is used exactly once between the two samples, so \[S_n + S_m = \dfrac {N(N+1)}{2} = \dfrac {(n+m)(n+m+1)}{2}.\] Substituting, \[U_n + U_m = 2nm + \dfrac {n^{2}+n+m^{2}+m}{2} - \dfrac {(n+m)^{2}+(n+m)}{2}.\] Expanding the last numerator, \((n+m)^{2}+(n+m) = n^{2}+2nm+m^{2}+n+m\), so the two fractions differ only by the cross term: \[U_n + U_m = 2nm + \dfrac {-2nm}{2} = 2nm - nm = nm .\] Therefore \[\boxed {U_n = nm - U_m .} \qquad \blacksquare \]
What the identity says. \(U_n\) counts the pairs \((X_i, Y_j)\) with \(X_i > Y_j\) and \(U_m\) counts those with \(Y_j > X_i\). There are \(nm\) pairs altogether and, the populations being continuous, no ties — so the two counts must add to \(nm\). The algebra above is that observation done through the rank sums.
It also means only one statistic need ever be computed: the other follows by subtraction, and \(\min (U_n, U_m)\) — the value usually referred to tables — is available at once.
Problem 5.10. Explain what is meant by an order statistic, and name three non-parametric methods built on order statistics, illustrating with notation where useful.
Show solution
Solution. Definition. Given a sample \(X_1,\ldots ,X_n\), arrange the values in increasing order \[X_{(1)} \leq X_{(2)} \leq \cdots \leq X_{(n)} .\] The \(i\)-th of these, \(X_{(i)}\), is the \(i\)-th order statistic. It is a statistic in the ordinary sense — a function of the sample — but one that depends on the observations only through their relative magnitudes, which is precisely why order statistics underpin distribution-free methods: their distribution theory under \(H_0\) can be worked out without knowing \(F\).
Note \(X_{(1)} = \min _i X_i\) and \(X_{(n)} = \max _i X_i\), and the sample median is \(X_{((n+1)/2)}\) for odd \(n\).
Three methods built on them.
- (i).
- The confidence interval for a median. The interval \(\left (X_{(k)},\,X_{(n+1-k)}\right )\) has coverage \(1 - 2P\!\left (\text {Bin}(n,\tfrac 12) \leq k-1\right )\), free of \(F\) entirely. The endpoints are order statistics and nothing else is used.
- (ii).
- The Kolmogorov–Smirnov test. The empirical distribution function is built directly from the order statistics, \[\widehat {F}_n\left (X_{(i)}\right ) = \dfrac {i}{n},\] and the statistic \(D = \max \left |\widehat {F}_n - F_0\right |\) is evaluated at the \(X_{(i)}\).
- (iii).
- Rank tests generally — the sign test, Wilcoxon signed-rank, Wilcoxon rank-sum and Kruskal–Wallis. Ranking is exactly the operation of recording which order statistic each observation is, and every one of these tests uses that information and discards the rest.
The quartiles and percentiles are also order statistics, which is why the quartile test of the previous chapter belongs to the same family.
Questions on this section
Stuck on something here? Ask below and it stays attached to this topic.