4.5 Practice Problems

Problem 4.1. Three systems were compared, each on six occasions:

System A 55 60 63 56 59 55
System B 54 57 53 64 49 62
System C 66 52 61 60 57 58

(a).
Obtain the analysis of variance table, showing all calculations.
(b).
Are the three systems equally effective? Use \(\alpha =0.05\).
(c).
Is it necessary to carry out pairwise tests? Give your reasons.

Show solution

Solution. (a). The totals are \(T_A=348\), \(T_B=339\), \(T_C=354\), so the grand total is \(T=1041\) with \(n=18\) observations in \(k=3\) groups of \(n_j=6\).

Start with the correction factor, which appears in every sum of squares: \[CF=\frac {T^2}{n}=\frac {1041^2}{18}=\frac {1083681}{18}=60204.5.\] With \(\sum \sum x_{ij}^2=60545\), \[SS_T=\sum \sum x_{ij}^2-CF=60545-60204.5=340.5,\] \[SS_B=\sum _{j}\frac {T_j^2}{n_j}-CF=\frac {348^2+339^2+354^2}{6}-60204.5=60223.5-60204.5=19.0,\] \[SS_W=SS_T-SS_B=340.5-19.0=321.5.\] Note \(SS_B\) uses each group’s own total \(T_j\), not the grand total.

Source \(SS\) df \(MS\) \(F\)
Between systems \(19.0\) \(k-1=2\) \(9.50\) \(0.44\)
Within systems \(321.5\) \(n-k=15\) \(21.43\)
Total \(340.5\) \(n-1=17\)

The degrees of freedom add up, \(2+15=17\), which is worth checking every time. Note that \(MS_W\) divides \(SS_W\) by \(n-k=15\), matching its own row.

(b). Testing \(H_0:\mu _A=\mu _B=\mu _C\) against the alternative that at least one differs: \[F=\frac {MS_B}{MS_W}=\frac {9.50}{21.43}=0.44.\] The critical value is \(F_{2,15,\,0.05}=3.68\), and \(0.44\) falls far short of it; the \(P\)-value is \(0.65\). We fail to reject \(H_0\). There is no evidence that the three systems differ in effectiveness.

Look at why. The group means are \(58.0\), \(56.5\) and \(59.0\) – a spread of only \(2.5\) – while individual readings within a single system run from \(49\) to \(66\). The variation between systems is far smaller than the variation within them, and \(F\) is precisely the ratio of those two. An \(F\) below 1 says the group means are closer together than one would expect from chance alone.

(c). No. Pairwise tests are what one does after a significant \(F\), to find which groups differ. Here \(F\) is not significant, so there is nothing to locate: we have no evidence that any of the systems differ from any other.

Running pairwise tests anyway would be a mistake, and not a harmless one. Three comparisons each at \(5\%\) give a much higher chance than \(5\%\) of at least one false positive, which is the whole reason for testing all groups together with a single \(F\) first.

Problem 4.2. A stimulus-response experiment with three treatments was laid out in a randomised block design using four subjects. The response is reaction time in seconds; the treatment applied is shown in brackets.

Subject 1
Subject 2
Subject 3
Subject 4
1.7 (1) 2.1 (3) 0.1 (1) 2.3 (2)
2.3 (3) 1.5 (1) 2.3 (2) 0.6 (1)
3.4 (2) 2.6 (2) 0.8 (3) 1.6 (3)

Do the data present sufficient evidence of a difference in mean response between the stimuli?

Show solution

Solution. The subjects are the blocks and the stimuli the treatments. As always in a randomised block design, the treatments appear in a different order in each block, so totals must be gathered by label and not by position in the column.

Subject 1 Subject 2 Subject 3 Subject 4 Treatment total
Treatment 1 1.7 1.5 0.1 0.6 3.9
Treatment 2 3.4 2.6 2.3 2.3 10.6
Treatment 3 2.3 2.1 0.8 1.6 6.8
Block total 7.4 6.2 3.2 4.5 21.3
Table 39: Reaction times rearranged by treatment. Both sets of totals are needed.

With \(N=12\) and \(T=21.3\), the correction factor is \[CF=\frac {T^2}{N}=\frac {453.69}{12}=37.8075.\] \begin {align*} SS_{\text {Total}} &=\sum x^2-CF=47.31-37.8075=9.5025\\ SS_{\text {Blocks}} &=\frac {7.4^2+6.2^2+3.2^2+4.5^2}{3}-CF=41.23-37.8075=3.4225\\ SS_{\text {Treat}} &=\frac {3.9^2+10.6^2+6.8^2}{4}-CF=43.4525-37.8075=5.6450\\ SS_{\text {Error}} &=9.5025-3.4225-5.6450=0.4350 \end {align*}

Each total is divided by the number of observations that went into it: block totals by the \(3\) treatments, treatment totals by the \(4\) blocks.

Source df \(SS\) \(MS\) \(F\)
Blocks (subjects) 3 3.4225 1.1408 15.74
Treatments (stimuli) 2 5.6450 2.8225 38.93
Error 6 0.4350 0.0725
Total 11 9.5025
Table 40: ANOVA table for the randomised block design.

The degrees of freedom add, \(3+2+6=11\), and so do the sums of squares.

At \(\alpha =0.05\), \(F_{2,6}=5.14\) for treatments and \(F_{3,6}=4.76\) for blocks. \[38.93>5.14\quad \text {and}\quad 15.74>4.76,\] so both are significant. There is very strong evidence of a difference in mean reaction time between the stimuli (\(P<0.001\)), and the subjects differ from one another too.

That the blocks are significant is worth noticing: blocking earned its keep here. Subject 3 is much faster than Subject 1 across every stimulus, and had that variation been left in the error term the treatment effect would have been far harder to see. Compare the randomised block example earlier in this chapter, where the blocks showed no difference and the blocking bought nothing.

Problem 4.3. Four groups of students were taught by different techniques and tested afterwards. Drop-outs left the groups unequal in size. Do the data indicate a difference in mean achievement between the four techniques?

Group 1 Group 2 Group 3 Group 4
65 75 59 94
87 69 78 89
73 83 67 80
79 81 62 88
81 72 82
69 79 73
90

Show solution

Solution. The groups have \(n_i=6, 7, 6, 4\), so \(N=23\). Unequal group sizes change nothing in principle — each treatment total is simply divided by its own \(n_i\).

Group 1 Group 2 Group 3 Group 4 Overall
\(n_i\) 6 7 6 4 23
\(T_i\) 454 549 421 351 1775
Mean 75.67 78.43 70.17 87.75 77.17
Table 41: Group totals and means for the four teaching techniques.

\[CF=\frac {T^2}{N}=\frac {1775^2}{23}=136\,983.7\] \begin {align*} SS_{\text {Total}} &=\sum x^2-CF=138\,899-136\,983.7=1915.30\\ SS_{\text {Between}} &=\sum _i\frac {T_i^2}{n_i}-CF =\frac {454^2}{6}+\frac {549^2}{7}+\frac {421^2}{6}+\frac {351^2}{4}-CF=766.67\\ SS_{\text {Within}} &=1915.30-766.67=1148.63 \end {align*}

Source df \(SS\) \(MS\) \(F\)
Between techniques \(k-1=3\) 766.67 255.56 4.23
Within (error) \(N-k=19\) 1148.63 60.45
Total 22 1915.30
Table 42: One-way ANOVA for the teaching techniques.

At \(\alpha =0.05\), \(F_{3,19}=3.13\). Since \(4.23>3.13\) we reject \(H_0\): there is evidence of a difference in mean achievement between the four techniques (\(P\approx 0.019\)).

The \(F\) test says only that the four means are not all equal. Reading the group means, Group 4 at \(87.75\) and Group 3 at \(70.17\) are the extremes and are the likely source, but establishing which pairs differ requires a follow-up procedure such as the least significant difference — not a series of unadjusted \(t\) tests.

Problem 4.4. Four assemblers each assembled four electronic devices, in the Latin square below. The response is assembly time in minutes and the device is shown in brackets. Two sources of unwanted variation are controlled: differences between people, and fatigue over the assembly sequence.

Position
Assembler 1
Assembler 2
Assembler 3
Assembler 4
1 44 (3) 41 (1) 30 (2) 40 (4)
2 41 (2) 42 (3) 49 (4) 49 (1)
3 59 (1) 41 (4) 59 (3) 34 (2)
4 58 (4) 37 (2) 53 (1) 59 (3)

Do the data indicate a difference in mean assembly time between the four devices?

Show solution

Solution. First check the design really is a Latin square: each device must appear exactly once in every row and every column. Reading across and down, each of \(1,2,3,4\) does. That property is what allows three effects — position, assembler and device — to be separated from only \(16\) observations.

Gathering the three sets of totals, with \(N=16\) and \(T=736\):

1 2 3 4
Row (position) totals 155 181 193 207
Column (assembler) totals 202 161 191 182
Treatment (device) totals 202 142 204 188
Table 43: The three sets of totals a Latin square requires.

\[CF=\frac {736^2}{16}=33\,856\] \begin {align*} SS_{\text {Total}} &=\sum x^2-CF=35\,186-33\,856=1330.0\\ SS_{\text {Rows}} &=\frac {155^2+181^2+193^2+207^2}{4}-CF=365.0\\ SS_{\text {Cols}} &=\frac {202^2+161^2+191^2+182^2}{4}-CF=226.5\\ SS_{\text {Treat}} &=\frac {202^2+142^2+204^2+188^2}{4}-CF=626.0\\ SS_{\text {Error}} &=1330.0-365.0-226.5-626.0=112.5 \end {align*}

The error degrees of freedom are \((p-1)(p-2)=3\times 2=6\).

Source df \(SS\) \(MS\) \(F\)
Rows (position) 3 365.0 121.67 6.49
Columns (assemblers) 3 226.5 75.50 4.03
Treatments (devices) 3 626.0 208.67 11.13
Error 6 112.5 18.75
Total 15 1330.0
Table 44: ANOVA table for the Latin square. Degrees of freedom: \(3+3+3+6=15\).

At \(\alpha =0.05\), \(F_{3,6}=4.76\) throughout.

Devices. \(11.13>4.76\): reject. There is strong evidence of a difference in mean assembly time between the four devices. Device 2 is quickest, with a total of \(142\) against \(204\) for device 3.
Position. \(6.49>4.76\): reject. Times lengthen down the sequence, totals running \(155, 181, 193, 207\) — which is the fatigue effect the design was built to control.
Assemblers. \(4.03<4.76\): fail to reject. No significant difference between the four people.

The design justified itself. Fatigue is real and significant; had it been ignored, its \(365\) would have gone into the error term, raising \(MS_{\text {error}}\) from \(18.75\) to \(53.1\) and cutting the device \(F\) from \(11.13\) to \(3.93\), against a critical \(F_{3,9}=3.86\) — barely significant instead of decisively so. Removing a known nuisance source is what makes the effect of interest visible.

Problem 4.5. The owner of a large company wanted to compare the mean daily output of a particular item for five plants. For each plant, a random sample of \(4\) days gave the data below:

Plants A B C D E
29 22 24 23 15
18 17 16 15 8
18 12 14 14 9
18 11 12 10 4

(a).
Write down a model for the above design. Define all the terms in your model.
(b).
Do the sample data indicate a difference in the population means for the five plants?

Show solution

Solution. (a). The model. Five plants are compared and the four days within each plant are a random sample from that plant alone, so this is a completely randomised design with one factor at five levels: \[Y_{ij}=\mu +\tau _i+\varepsilon _{ij},\qquad i=1,\ldots ,5;\quad j=1,\ldots ,4,\] where

\(\bullet \)
\(Y_{ij}\) is the output on the \(j\)th day at the \(i\)th plant;
\(\bullet \)
\(\mu \) is the overall mean output, common to all plants;
\(\bullet \)
\(\tau _i\) is the effect of the \(i\)th plant — the amount by which that plant’s mean departs from \(\mu \), subject to \(\sum \tau _i=0\);
\(\bullet \)
\(\varepsilon _{ij}\) is the random error, assumed independent and \(N(0,\sigma ^2)\) for every plant.

The hypothesis of no difference between plants is \(\tau _1=\tau _2=\cdots =\tau _5=0\).

(b). The analysis of variance. The plant totals are \(83\), \(62\), \(66\), \(62\) and \(36\), each from \(n_i=4\) days, with grand total \(G=309\) and \(N=20\) observations. \[CF=\frac {G^2}{N}=\frac {309^2}{20}=\frac {95{,}481}{20}=4774.05,\] \[SS_{\text {Total}}=\sum Y_{ij}^2-CF=5459-4774.05=684.95,\] \[SS_{\text {Plants}}=\sum \frac {T_i^2}{n_i}-CF =\frac {83^2+62^2+66^2+62^2+36^2}{4}-4774.05=5057.25-4774.05=283.20,\] \[SS_{\text {Error}}=684.95-283.20=401.75.\]

Source \(SS\) df \(MS\) \(F\)
Between plants 283.20 4 70.80 2.64
Error 401.75 15 26.78
Total 684.95 19

1.
\(H_0:\ \mu _A=\mu _B=\mu _C=\mu _D=\mu _E\); \(H_1:\) at least two plant means differ.
2.
\(\alpha =0.05\).
3.
\(F=\dfrac {MS_{\text {Plants}}}{MS_{\text {Error}}}=\dfrac {70.80}{26.78}=2.64\) on \((4,15)\) degrees of freedom.
4.
Critical region: \(F_{4,15,\,0.05}=3.06\), so reject if \(F>3.06\).
5.
\(2.64<3.06\), so we fail to reject \(H_0\). The data do not indicate a difference in the mean daily output of the five plants at the \(5\%\) level.

This is a close call — \(P=0.075\) — and the reason is visible in the data rather than the table. Every plant has one unusually high first day (\(29\), \(22\), \(24\), \(23\), \(15\)) and three lower ones, which inflates the within-plant variation and hence \(MS_{\text {Error}}\). That common pattern suggests the day is itself a source of variation: if the four days were the same four days at every plant, this is not a completely randomised design at all but a randomised block design with days as blocks. Treating them so gives \(SS_{\text {Days}}=377.75\) on \(3\) degrees of freedom, leaving \(SS_{\text {Error}}=24.00\) on \(12\) — so \(MS_{\text {Error}}\) falls from \(26.78\) to \(2.00\) and the plant \(F\) rises from \(2.64\) to \(35.4\), overwhelmingly significant. Problems 3 and 4 do exactly this. Which analysis is correct depends on a fact about how the data were collected that the question does not supply, and the two answers could hardly differ more.

Problem 4.6. The data in the following table represent final grades given by three professors in an advanced statistics course.

A B C
63 67 97
45 45 97
73 76 87
77 80 87
72 70 84
70 74
74
64

(a).
Write down a model for the above design, explaining all the terms in the model.
(b).
Is there sufficient evidence to indicate that not all population means are the same?

Show solution

Solution. (a). The model. The three professors have \(5\), \(6\) and \(8\) students, and there is nothing pairing one professor’s students with another’s. This is a completely randomised design with unequal group sizes: \[Y_{ij}=\mu +\tau _i+\varepsilon _{ij},\qquad i=1,2,3;\quad j=1,\ldots ,n_i,\] with \(n_1=5\), \(n_2=6\), \(n_3=8\), where \(Y_{ij}\) is the grade of the \(j\)th student of the \(i\)th professor, \(\mu \) the overall mean grade, \(\tau _i\) the effect of the \(i\)th professor, and \(\varepsilon _{ij}\sim N(0,\sigma ^2)\) independent errors. Unequal \(n_i\) change nothing in the structure of the model; they only change the divisors in the sums of squares.

(b). The analysis of variance. The professor totals are \(330\), \(408\) and \(664\), from \(5\), \(6\) and \(8\) students, with \(G=1402\) and \(N=19\). \[CF=\frac {1402^2}{19}=\frac {1{,}965{,}604}{19}=103{,}452.84,\] \[SS_{\text {Total}}=\sum Y_{ij}^2-CF=106{,}986-103{,}452.84=3533.16,\] \[SS_{\text {Professors}}=\frac {330^2}{5}+\frac {408^2}{6}+\frac {664^2}{8}-CF =21{,}780+27{,}744+55{,}112-103{,}452.84=1183.16,\] \[SS_{\text {Error}}=3533.16-1183.16=2350.00.\]

Source \(SS\) df \(MS\) \(F\)
Between professors 1183.16 2 591.58 4.03
Error 2350.00 16 146.88
Total 3533.16 18

1.
\(H_0:\ \mu _A=\mu _B=\mu _C\); \(H_1:\) at least two professor means differ.
2.
\(\alpha =0.05\).
3.
\(F=\dfrac {591.58}{146.88}=4.03\) on \((2,16)\) degrees of freedom.
4.
Critical region: \(F_{2,16,\,0.05}=3.63\), so reject if \(F>3.63\).
5.
\(4.03>3.63\), so we reject \(H_0\). There is sufficient evidence that not all three population means are the same.

Note where the degrees of freedom come from with unequal groups: \(k-1=2\) between, and \(N-k=19-3=16\) within — not \(k(n-1)\), which only applies when every group has the same \(n\).

The verdict is level-dependent again: \(P=0.038\), so this rejects at \(5\%\) but not at \(1\%\) (where the critical value is \(6.23\)). And the ANOVA says only that the three means are not all equal, not which differ. The group means are \(66.0\), \(68.0\) and \(83.0\), so professor C is the obvious candidate, but establishing that formally needs a multiple comparison procedure rather than the \(F\)-test.

Problem 4.7. In a certain study, \(3\) diets were assigned to each of \(6\) subjects in a randomised block design. At the end of the study each subject was put on a treadmill and the time to exhaustion, in seconds, was measured:

Subject
Diet 1 2 3 4 5 6
1 84 35 91 57 56 45
2 91 48 71 45 61 61
3 122 53 110 71 91 122

(a).
Write down a model for the above design. Explain each term in the model.
(b).
Perform the analysis of variance, separating out the diet, subject and error sum of squares. Use a \(0.01\) level of significance to determine if there are significant differences among the diets.

Show solution

Solution. (a). The model. Every subject received every diet, so subjects are blocks: \[Y_{ij}=\mu +\tau _i+\beta _j+\varepsilon _{ij},\qquad i=1,2,3;\quad j=1,\ldots ,6,\] where

\(\bullet \)
\(Y_{ij}\) is the time to exhaustion of subject \(j\) on diet \(i\);
\(\bullet \)
\(\mu \) is the overall mean time;
\(\bullet \)
\(\tau _i\) is the effect of the \(i\)th diet, with \(\sum \tau _i=0\) — the treatment effect, and the object of the study;
\(\bullet \)
\(\beta _j\) is the effect of the \(j\)th subject, with \(\sum \beta _j=0\) — the block effect, a nuisance to be removed rather than a question of interest;
\(\bullet \)
\(\varepsilon _{ij}\sim N(0,\sigma ^2)\) are independent errors.

There is one observation per cell, so no diet-by-subject interaction can be estimated; the model assumes there is none.

(b). The analysis of variance. The diet totals are \(368\), \(377\) and \(569\); the subject totals are \(297\), \(136\), \(272\), \(173\), \(208\) and \(228\). With \(G=1314\) and \(N=18\), \[CF=\frac {1314^2}{18}=\frac {1{,}726{,}596}{18}=95{,}922,\] \[SS_{\text {Total}}=\sum Y_{ij}^2-CF=108{,}064-95{,}922=12{,}142,\] \[SS_{\text {Diets}}=\frac {368^2+377^2+569^2}{6}-CF=100{,}219-95{,}922=4297,\] \[SS_{\text {Subjects}}=\frac {297^2+136^2+272^2+173^2+208^2+228^2}{3}-CF =101{,}955.33-95{,}922=6033.33,\] \[SS_{\text {Error}}=12{,}142-4297-6033.33=1811.67.\] The error degrees of freedom are \((r-1)(c-1)=(3-1)(6-1)=10\).

Source \(SS\) df \(MS\) \(F\)
Diets 4297.00 2 2148.50 11.86
Subjects (blocks) 6033.33 5 1206.67 6.66
Error 1811.67 10 181.17
Total 12,142.00 17

1.
\(H_0:\ \tau _1=\tau _2=\tau _3=0\) (the three diets have the same mean time); \(H_1:\) at least two diets differ.
2.
\(\alpha =0.01\).
3.
\(F=\dfrac {2148.50}{181.17}=11.86\) on \((2,10)\) degrees of freedom.
4.
Critical region: \(F_{2,10,\,0.01}=7.56\), so reject if \(F>7.56\).
5.
\(11.86>7.56\), so we reject \(H_0\). There are significant differences among the diets, and the diet totals (\(368\), \(377\), \(569\)) point to diet 3 as the one producing the longest times.

The blocking earned its place. Subjects differ enormously — subject 1 totals \(297\) seconds against subject 2’s \(136\) — and that variation has nothing to do with diet. Blocking sends it to its own line of the table, where \(F=6.66\) against \(F_{5,10,\,0.01}=5.64\) confirms it is real. Had subjects been ignored, that \(6033\) would have joined the error, giving \(MS_{\text {Error}}=(6033.33+1811.67)/15=523.0\) and a diet \(F\) of only \(4.11\) — short of \(F_{2,15,\,0.01}=6.36\), and the study would have concluded nothing. Pairing each subject with themselves across all three diets is what makes the experiment work.

Problem 4.8. Four kinds of fertilizer \(f_1\), \(f_2\), \(f_3\) and \(f_4\) are used to study the yield of beans. The soil is divided into \(3\) blocks each containing \(4\) homogeneous plots. The yields in kilograms per plot and the corresponding treatments are:

Block 1 Block 2 Block 3
\(f_1=42.7\) \(f_3=50.9\) \(f_4=51.1\)
\(f_3=48.5\) \(f_1=50.0\) \(f_2=46.3\)
\(f_4=32.8\) \(f_2=38.0\) \(f_1=51.9\)
\(f_2=39.3\) \(f_4=40.2\) \(f_3=53.5\)

(a).
Write down a model for the above design, explaining all the terms in the model.
(b).
Conduct an analysis of variance at the \(0.01\) level of significance using the randomised block model. What is your conclusion?

Show solution

Solution. (a). The model. Each block contains all four fertilizers, one per plot, and the order within a block is randomised — a randomised block design: \[Y_{ij}=\mu +\tau _i+\beta _j+\varepsilon _{ij},\qquad i=1,\ldots ,4;\quad j=1,2,3,\] where \(Y_{ij}\) is the yield of fertilizer \(i\) in block \(j\), \(\mu \) the overall mean yield, \(\tau _i\) the effect of the \(i\)th fertilizer (\(\sum \tau _i=0\)), \(\beta _j\) the effect of the \(j\)th block (\(\sum \beta _j=0\)), and \(\varepsilon _{ij}\sim N(0,\sigma ^2)\) independent errors. Blocks are strips of soil expected to differ in fertility; the design removes that difference from the comparison of fertilizers.

(b). The analysis of variance. Rearranged by treatment, the fertilizer totals are \[T_1=42.7+50.0+51.9=144.6,\qquad T_2=39.3+38.0+46.3=123.6,\] \[T_3=48.5+50.9+53.5=152.9,\qquad T_4=32.8+40.2+51.1=124.1,\] and the block totals are \(B_1=163.3\), \(B_2=179.1\), \(B_3=202.8\). With \(G=545.2\) and \(N=12\), \[CF=\frac {(545.2)^2}{12}=\frac {297{,}243.04}{12}=24{,}770.25,\] \[SS_{\text {Total}}=\sum Y_{ij}^2-CF=25{,}257.48-24{,}770.25=487.23,\] \[SS_{\text {Fertilizer}}=\frac {T_1^2+T_2^2+T_3^2+T_4^2}{3}-CF=24{,}988.45-24{,}770.25=218.19,\] \[SS_{\text {Blocks}}=\frac {B_1^2+B_2^2+B_3^2}{4}-CF=24{,}967.89-24{,}770.25=197.63,\] \[SS_{\text {Error}}=487.23-218.19-197.63=71.40,\] on \((4-1)(3-1)=6\) error degrees of freedom.

Source \(SS\) df \(MS\) \(F\)
Fertilizers 218.19 3 72.73 6.11
Blocks 197.63 2 98.82 8.30
Error 71.40 6 11.90
Total 487.23 11

1.
\(H_0:\ \tau _1=\tau _2=\tau _3=\tau _4=0\) (all four fertilizers give the same mean yield); \(H_1:\) at least two differ.
2.
\(\alpha =0.01\).
3.
\(F=\dfrac {72.73}{11.90}=6.11\) on \((3,6)\) degrees of freedom.
4.
Critical region: \(F_{3,6,\,0.01}=9.78\), so reject if \(F>9.78\).
5.
\(6.11<9.78\), so we fail to reject \(H_0\). At the \(1\%\) level there is no significant difference between the four fertilizers.

The conclusion turns entirely on the level demanded. At \(5\%\) the critical value is \(F_{3,6,\,0.05}=4.76\), and \(6.11\) exceeds it comfortably; the \(P\)-value is \(0.030\). So these data do distinguish the fertilizers at the conventional \(5\%\) level and fail only against the stricter \(1\%\) the question imposes. With six error degrees of freedom that is unsurprising: a \(1\%\) test on so little data is a demanding standard.

The blocks were worth having. Their \(F=8.30\) against \(F_{2,6,\,0.01}=10.92\) falls just short of significance at \(1\%\), but the block totals climb steadily — \(163.3\), \(179.1\), \(202.8\) — so block 3 is plainly the better soil. Ignoring blocks would put that \(197.63\) into the error term, giving \(MS_{\text {Error}}=(197.63+71.40)/8=33.63\) and a fertilizer \(F\) of only \(2.16\): nothing at all. The design rescued a real effect that a completely randomised experiment on the same twelve plots would have buried.

Problem 4.9. A biology professor was interested in whether there was a difference in the mean heart rate maximum after exercise on a treadmill between highly trained male and female runners. Do the following data indicate a difference in the population means? Use analysis of variance with a \(1\%\) level of significance.

sample size \(\overline {x}\) \(s\)
Males 10 180.3 7.2
Females 10 190.8 10.7

Show solution

Solution. Only summary statistics are given, so the sums of squares must be rebuilt from them. With \(k=2\) groups of \(n=10\), the grand mean is \[\overline {\overline {X}}=\frac {10(180.3)+10(190.8)}{20}=\frac {1803+1908}{20}=185.55.\] The between-groups sum of squares measures how far the group means sit from that: \[SS_{\text {Between}}=\sum n_i(\overline {X}_i-\overline {\overline {X}})^2 =10(-5.25)^2+10(5.25)^2=551.25.\] The within-groups sum of squares is recovered from the sample variances, since \(\hat {S}_i^2=SS_i/(n_i-1)\): \[SS_{\text {Within}}=\sum (n_i-1)\hat {S}_i^2=9(51.84)+9(114.49)=1496.97.\]

Source \(SS\) df \(MS\) \(F\)
Between sexes 551.25 1 551.25 6.63
Within (error) 1496.97 18 83.17
Total 2048.22 19

1.
\(H_0:\ \mu _M=\mu _F\); \(H_1:\ \mu _M\neq \mu _F\).
2.
\(\alpha =0.01\).
3.
\(F=\dfrac {551.25}{83.17}=6.63\) on \((1,18)\) degrees of freedom.
4.
Critical region: \(F_{1,18,\,0.01}=8.29\), so reject if \(F>8.29\).
5.
\(6.63<8.29\), so we fail to reject \(H_0\). At the \(1\%\) level the data do not indicate a difference in the mean maximum heart rate of male and female runners.

With exactly two groups, ANOVA and the pooled two-sample \(t\)-test are the same test. Here \[t=\frac {190.8-180.3}{\sqrt {83.17\left (\frac {1}{10}+\frac {1}{10}\right )}} =\frac {10.5}{4.078}=2.57,\] and \(t^2=(2.57)^2=6.63=F\), while the critical values obey the same relation: \(t_{18,\,0.005}^2=(2.878)^2=8.28=F_{1,18,\,0.01}\). The identity \(F_{1,v}=t_v^2\) holds generally, and it is why the \(F\)-test is described as the extension of the two-sample \(t\)-test to more than two groups. Note that the equivalence is with the two-tailed \(t\)-test: \(F\) is always one-tailed in appearance because squaring destroys the sign, so a large \(F\) covers a difference in either direction.

The verdict is again close: \(P=0.019\), significant at \(5\%\) but not at the \(1\%\) demanded. Note also that ANOVA assumes equal variances across groups, and the female runners’ \(s=10.7\) against the males’ \(7.2\) gives \(F=114.49/51.84=2.21\) on \((9,9)\) degrees of freedom against a critical \(F_{0.025,\,9,9}=4.03\) — so the assumption stands.

Questions on this section

Stuck on something here? Ask below and it stays attached to this topic.