4.4 The Efficiency of the Wilcoxon Test

Combining Theorems 4.2.2 and 4.3.1 gives the central result of the chapter.

Corollary 4.4.1. \[\mathrm {ARE}(W,t) = 12\,\sigma ^{2} \left (\int ^{\infty }_{-\infty }f^{2}(x)\,dx\right )^{2} .\]

Everything now depends on the single functional \(\sigma ^{2}\left (\int f^{2}\right )^{2}\), which can be evaluated for any density of interest.

Example 4.4.2 (Normal data). Let \(f\) be the \(N(0,\sigma ^{2})\) density. Then \[\int ^{\infty }_{-\infty }f^{2}(x)\,dx = \frac {1}{2\sigma \sqrt {\pi }},\] so by Corollary 4.4.1 \[\mathrm {ARE}(W,t) = 12\,\sigma ^{2}\cdot \frac {1}{4\sigma ^{2}\pi } = \frac {3}{\pi } \approx 0.955 .\]

This is the number the subject is built on, and it deserves to be stated in words. On perfectly normal data — the case in which the \(t\)-test is exactly optimal, and the case most favourable to it — the Wilcoxon test, which throws away the numerical values entirely and keeps only their order, is \(95.5\%\) as efficient. Discarding the data’s magnitudes costs about one observation in twenty-two.

Example 4.4.3 (Heavy tails). For the double exponential (Laplace) density \(f(x)=\tfrac {1}{2b}e^{-|x|/b}\), which has \(\sigma ^{2}=2b^{2}\) and \(\int f^{2} = 1/(4b)\), \[\mathrm {ARE}(W,t) = 12\cdot 2b^{2}\cdot \frac {1}{16b^{2}} = \frac {3}{2}.\] The Wilcoxon test needs two thirds of the observations the \(t\)-test requires.

Distribution \(\mathrm {ARE}(W,t)\) \(\mathrm {ARE}(S,t)\)
Uniform \(1.000\) \(0.333\)
Normal \(3/\pi \approx 0.955\) \(2/\pi \approx 0.637\)
Logistic \(\pi ^{2}/9 \approx 1.097\) \(\pi ^{2}/12\approx 0.822\)
Double exponential \(1.500\) \(2.000\)
Cauchy \(\infty \) \(\infty \)
Table 4.1: Asymptotic relative efficiency of the Wilcoxon (\(W\)) and sign (\(S\)) tests with respect to the \(t\)-test, under location shift.

Two features of Table 4.1 deserve comment. The Wilcoxon test is never much worse than the \(t\)-test and is frequently better; already at the logistic — a distribution barely distinguishable from the normal by eye — it has overtaken it. And under the Cauchy the entry is not large but infinite, because the \(t\)-test has no finite variance to work with and fails entirely, while the rank test is unaffected.

The sign test tells a different story: it is markedly inefficient on well-behaved data, at roughly two thirds of the \(t\)-test under normality, but it is the best of the three under the double exponential. Nothing dominates everywhere.

Questions on this section

Stuck on something here? Ask below and it stays attached to this topic.