4.5 The Hodges–Lehmann Bound
Table 4.1 raises an obvious question. Its worst entry for the Wilcoxon test is \(0.955\), at the normal. Is there some awkward distribution on which the Wilcoxon test performs arbitrarily badly? The answer is no, and it is the strongest single argument for rank methods.
Theorem 4.5.1 (Hodges and Lehmann, 1956). Over the class of all continuous distributions with finite variance, \[\inf _{F}\ \mathrm {ARE}(W,t) = \frac {108}{125} = 0.864 ,\] the infimum being attained at a particular density; and \[\sup _{F}\ \mathrm {ARE}(W,t) = \infty .\]
The asymmetry in Theorem 4.5.1 is the whole argument in one line. Using the Wilcoxon test when the \(t\)-test would have been appropriate can never cost more than about \(13.6\%\) in efficiency, whatever the distribution. Using the \(t\)-test when it is inappropriate can cost everything.
Note 4.5.2. It is worth being precise about what a bounded loss of this size means in practice. An ARE of \(0.864\) says that where the \(t\)-test would need \(100\) observations, the Wilcoxon needs about \(116\). The insurance premium against non-normality is sixteen extra observations in the worst case ever constructible, and about five extra observations in the normal case where the \(t\)-test is optimal. Set against that, the \(t\)-test on Cauchy-like data does not merely lose efficiency; it has no valid asymptotic theory at all.
This is why the classical advice to test for normality first and then choose the procedure is poor practice. The preliminary test has its own error rates, the choice it induces invalidates the nominal level of whatever follows, and the maximum gain from choosing correctly is a few percent while the maximum loss from choosing wrongly is unbounded.
Questions on this section
Stuck on something here? Ask below and it stays attached to this topic.