1.16 Practice problems

Past tutorial-sheet questions on sampling and describing data. Try each one before opening the solution.

Problem 1.1. [Tutorial Sheet 1] Unaware that 45% of the 20,000 voters in his constituency support him, a politician decides to estimate his political strength. A sample of 200 voters shows that 40% support him.

(a).
What is the population?
(b).
What is the sample?
(c).
What is the parameter of interest? State its value.
(d).
What is the statistic of interest? State its value.
(e).
Compare your answers in (c) and (d). Is it surprising that they differ? If the politician were to sample another 200 voters, which of the two numbers would be likely to change? Explain.

Show solution

Solution. (a). The population is all \(20{,}000\) voters in the constituency – every unit we want a conclusion about, not merely those we happened to ask.

(b). The sample is the \(200\) voters actually questioned.

(c). The parameter is the proportion of all voters in the constituency who support him: \(p = 0.45\), or \(45\%\).

(d). The statistic is the proportion of the sample who support him: \(\hat {p} = 0.40\), or \(40\%\).

(e). It is not surprising. A parameter describes the population; a statistic describes a sample drawn from it. Because the sample is only \(200\) of \(20{,}000\) voters, it will rarely reproduce the population proportion exactly, and the gap here – five percentage points – is ordinary sampling variation rather than evidence of a mistake.

If a fresh sample of \(200\) were taken, the statistic would almost certainly change: a different \(200\) people give a different count. The parameter would not. It is a fixed property of the constituency, \(0.45\), whether or not anyone measures it.

That asymmetry is the whole basis of what follows in this course. The parameter is fixed and unknown; the statistic is known and varies. Everything in estimation is an argument about how far the second can be trusted to say something about the first.

Problem 1.2. [Tutorial Sheet 1] For the population \(\{0,\ 1,\ 2,\ 3,\ 5,\ 7\}\):

(a).
List all the simple random samples of size 5.
(b).
Give an example of a systematic sample of size 3, where the elements are listed in the order \(0, 1, 2, 3, 5, 7\).
(c).
Give an example of a proportional stratified sample of size 3, where the strata are \(\{0,1,2,3\}\) and \(\{5,7\}\).
(d).
Give an example of a cluster sample of size 2, where the clusters are \(\{0,1\}\), \(\{2,3\}\) and \(\{5,7\}\).

Show solution

Solution. (a). A simple random sample of size 5 is any subset of 5 of the 6 elements, so there are \(\binom {6}{5} = 6\) of them – one for each element left out: \[\{0,1,2,3,5\},\quad \{0,1,2,3,7\},\quad \{0,1,2,5,7\},\] \[\{0,1,3,5,7\},\quad \{0,2,3,5,7\},\quad \{1,2,3,5,7\}.\]

(b). In a systematic sample we take every \(k\)th element of the ordered list. For a sample of 3 from 6, \(k = 6/3 = 2\). Starting at the first element and taking every second one gives \[\{0,\ 2,\ 5\}.\] Starting at the second instead would give \(\{1,3,7\}\); either is a valid answer, and the starting point should be chosen at random.

(c). Proportional stratified sampling takes from each stratum in proportion to its size. The strata hold 4 and 2 of the 6 elements, so of a sample of 3 we take \[\frac {4}{6}\times 3 = 2 \text { from } \{0,1,2,3\}, \qquad \frac {2}{6}\times 3 = 1 \text { from } \{5,7\},\] choosing at random within each. For example \(\{1,\ 3,\ 5\}\).

(d). In cluster sampling we sample whole clusters, not individuals. A sample of size 2 is therefore one entire cluster, for example \[\{2,\ 3\}.\] This is the point that separates it from stratified sampling: there, we take some units from every stratum; here, we take all units from a few clusters. Stratifying aims to represent every group; clustering is usually done because reaching a whole group at once is cheaper.

Problem 1.3. [Tutorial Sheet 1] The recovery times, in days, for 36 patients are:

8 7 6 9 4 5 3 7 8 10 7 7
6 4 10 3 6 8 2 5 4 5 3 8
7 4 6 3 7 12 4 3 6 6 9 4

(a).
Construct a frequency distribution.
(b).
What percentage of recovery times are fewer than 6 days?
(c).
Find the mean and median number of days to recovery.
(d).
Find \(P_{20}\) and the sample variance \(s^{2}\).

Show solution

Solution. (a). Counting each distinct value:

Days \(x\) 2 3 4 5 6 7 8 9 10 12
Frequency \(f\) 1 5 6 3 6 6 4 2 2 1

The frequencies total \(36\), which is the first thing to check.

(b). Fewer than 6 days means \(x = 2,3,4,5\): \[1+5+6+3 = 15 \text { patients}, \qquad \frac {15}{36}\times 100 = 41.67\%.\]

(c). Using \(\sum fx\) rather than adding 36 numbers separately: \[\sum fx = 216, \qquad \bar {x} = \frac {216}{36} = 6 \text { days}.\] With \(n = 36\) even, the median is the average of the 18th and 19th ordered values. The cumulative frequencies reach 15 by \(x=5\) and 21 by \(x=6\), so both the 18th and 19th values are 6: \[\text {median} = 6 \text { days}.\]

(d). \(P_{20}\) sits at position \(\frac {20 \times 36}{100} = 7.2\), that is between the 7th and 8th ordered values. Cumulatively, 6 observations are at most 3 days and 12 are at most 4, so the 7th and 8th values are both 4: \[P_{20} = 4 \text { days}.\] For the sample variance, with \(\bar {x} = 6\), \[s^{2} = \frac {\sum f(x-\bar {x})^{2}}{n-1} = \frac {196}{35} = 5.6, \qquad s = \sqrt {5.6} \approx 2.37 \text { days}.\] Note the divisor \(n-1\), not \(n\): this is a sample variance, and dividing by \(n\) would underestimate the spread of the population it came from.

That the mean and median are both exactly 6 is a hint about shape – the distribution is close to symmetric, which the frequency table bears out: the counts rise to a plateau around 4 to 7 and fall away on both sides, with one long right tail at 12.

Problem 1.4. [Assignment] Which of the following variates are discrete and which are continuous?

(a).
The throw of a die.
(b).
The height of a person.
(c).
The weight of eggs in a basket.
(d).
The number of rooms in a house.
(e).
The age of a car.

Show solution

Solution. The test is whether the possible values can be listed, or whether the variate can take any value in an interval.

(a).
Discrete – only \(1,2,3,4,5,6\) are possible.
(b).
Continuous – height can take any value in a range.
(c).
Continuous – weight is a measurement.
(d).
Discrete – a count.
(e).
Continuous – age is elapsed time, which flows without gaps.

Part (e) is the one that causes argument, because age is usually stated as a whole number of years. But how a value is recorded is not the same as what it can be: a car three and a half years old has an age, whether or not anybody writes it down that way. Rounding a measurement does not make the underlying variate discrete – and the next question depends on exactly that point.

Problem 1.5. [Assignment] For the following classes, give the true class limits and the mid-class values.

(a).
Length in mm, measured to the nearest \(0.1\) mm: \(20.0\text {--}20.4\), \(20.5\text {--}20.9\), \(21.0\text {--}21.4\), \(21.5\text {--}21.9\).
(b).
Length in mm, measured to the nearest \(0.01\) mm: \(20.00\text {--}20.49\), \(20.50\text {--}20.99\), \(21.00\text {--}21.49\), \(21.50\text {--}21.99\).
(c).
Age in years: \(1, 2, 3, 4, 5\).
(d).
Weight in kg, to the nearest \(1\) kg: \(58, 59, 60, 61, 62, 63\).

Show solution

Solution. A continuous quantity recorded to the nearest unit could really have been anything within half a unit either side. The true class limits stretch the stated ones by that half unit, so that the classes meet with no gaps between them.

(a). Half of \(0.1\) is \(0.05\):

Stated True limits Mid-value
\(20.0\)–\(20.4\) \(19.95\)–\(20.45\) \(20.2\)
\(20.5\)–\(20.9\) \(20.45\)–\(20.95\) \(20.7\)
\(21.0\)–\(21.4\) \(20.95\)–\(21.45\) \(21.2\)
\(21.5\)–\(21.9\) \(21.45\)–\(21.95\) \(21.7\)

Each class ends exactly where the next begins, which is the check that the limits are right.

(b). Half of \(0.01\) is \(0.005\), so the limits are \(19.995\)–\(20.495\), \(20.495\)–\(20.995\), \(20.995\)–\(21.495\) and \(21.495\)–\(21.995\), with mid-values \(20.245\), \(20.745\), \(21.245\) and \(21.745\).

(c). Age behaves differently, and this is the trap. Age is not rounded, it is truncated: a person is 1 year old from their first birthday until the day before their second. So the class “1” runs from \(1.0\) to \(2.0\), not from \(0.5\) to \(1.5\), and the mid-value is \(1.5\). Likewise the classes are \(2.0\)–\(3.0\), \(3.0\)–\(4.0\), \(4.0\)–\(5.0\) and \(5.0\)–\(6.0\), with mid-values \(2.5\), \(3.5\), \(4.5\) and \(5.5\).

(d). Weight is rounded, so the half-unit rule applies again: \(57.5\)–\(58.5\), \(58.5\)–\(59.5\), and so on up to \(62.5\)–\(63.5\), with mid-values equal to the stated values \(58, 59, \ldots , 63\).

Parts (c) and (d) look alike and are not. Ask how the number was arrived at – rounded to the nearest unit, or counted down to the last one completed – because the mid-values feed straight into the mean, and getting them half a unit out shifts the answer.

Problem 1.6. [Assignment] For the data below, give the true class limits, mid-values and relative frequencies.

Height in cm (nearest cm) Frequency
\(120\)–\(129\) 2
\(130\)–\(139\) 12
\(140\)–\(149\) 17
\(150\)–\(159\) 18
\(160\)–\(169\) 7
\(170\)–\(179\) 1

Show solution

Solution. Heights are recorded to the nearest cm, so the true limits run half a centimetre either side of the stated ones. The frequencies total \(n = 57\).

Stated True limits Mid-value \(f\) Relative \(f\) Cumulative \(f\)
\(120\)–\(129\) \(119.5\)–\(129.5\) \(124.5\) 2 \(0.035\) 2
\(130\)–\(139\) \(129.5\)–\(139.5\) \(134.5\) 12 \(0.211\) 14
\(140\)–\(149\) \(139.5\)–\(149.5\) \(144.5\) 17 \(0.298\) 31
\(150\)–\(159\) \(149.5\)–\(159.5\) \(154.5\) 18 \(0.316\) 49
\(160\)–\(169\) \(159.5\)–\(169.5\) \(164.5\) 7 \(0.123\) 56
\(170\)–\(179\) \(169.5\)–\(179.5\) \(174.5\) 1 \(0.018\) 57

The relative frequencies sum to 1 and the cumulative frequency ends at 57; both are worth checking before going further.

For a histogram, the bars are drawn against the true limits so that they touch – gaps between bars would wrongly suggest heights that cannot occur. The cumulative frequency curve is plotted against the upper true limit of each class, since that is the point by which the accumulated count has been reached.

Using the mid-values, the mean height is estimated as \[\bar {x} = \frac {\sum fx}{\sum f} = \frac {8426.5}{57} \approx 147.8 \text { cm}.\] This is an estimate, not the mean of the original data: once measurements are grouped, the individual values are gone, and every observation in a class is treated as sitting at its mid-value.

Problem 1.7. [Tutorial Sheet 1] Given the data below:

32 37 52 42 42 42 40 61 48 45
49 32 49 44 44 49 44 33 37 32
33 36 32 39 46 45 41 22 39 39
43 40 40 49 42 36 51 44 43 46

(a).
Construct a grouped frequency distribution using the classes \(22-26\), \(27-31\), …
(b).
Construct a frequency histogram.
(c).
Construct a frequency polygon.
(d).
Describe the shape of the distribution.

Show solution

Solution. (a) The classes are given, so the width is \(5\) and they run from \(22-26\) up to whatever class contains the largest value, \(61\). Counting into each:

Class \(22-26\) \(27-31\) \(32-36\) \(37-41\) \(42-46\) \(47-51\) \(52-56\) \(57-61\)
\(f\) 1 0 8 9 14 6 1 1
Mid-point 24 29 34 39 44 49 54 59
Boundaries 21.5 26.5 31.5 36.5 41.5 46.5 51.5 56.5
Table 13: Grouped frequency distribution. The boundaries row gives the lower boundary of each class; the last class closes at \(61.5\).

The frequencies total \(40\), which is the check worth making before going further.

Notice the class \(27-31\) is empty. It must still be listed. Dropping a class with zero frequency would make the horizontal scale uneven and the histogram would lie about the shape.

(b) and (c) The histogram has bars of equal width standing on the class boundaries \(21.5, 26.5, 31.5, \dots \) with heights \(1, 0, 8, 9, 14, 6, 1, 1\), and the bars touch because the underlying scale is continuous. The frequency polygon joins the mid-points of the bar tops, \((24,1), (29,0), (34,8), (39,9), (44,14), (49,6), (54,1), (59,1)\), and is closed onto the axis at a mid-point one class below and one above.

(d) The bulk of the data sits in \(42-46\), with a short tail to the right (\(52-56\) and \(57-61\) hold one value each) and a longer, thinner tail to the left reaching the isolated \(22\). The distribution is roughly mound-shaped but skewed to the left: the mean, \(41.5\), sits below the median, \(42\), which is the signature of a left tail.

Problem 1.8. [Tutorial Sheet 1] The following data represent the miles driven each day by a salesman over a 30-day period:

33 37 43 44 44 55 58 65 65 66 71 74 75 75 78
81 81 81 82 84 86 86 87 89 89 92 92 93 93 95

(a).
Construct a dot plot.
(b).
Construct a stem and leaf plot.
(c).
Construct a boxplot.
(d).
Describe the shape of the distribution.

Show solution

Solution. The data are already in order, which is half the work.

(a) A dot plot puts one dot above each value on a number line running from \(33\) to \(95\). Repeated values stack: \(44\), \(65\), \(75\), \(86\), \(89\), \(92\) and \(93\) each get two dots, and \(81\) gets three.

(b) Taking the tens digit as the stem:

Stem Leaf
3 3 7
4 3 4 4
5 5 8
6 5 5 6
7 1 4 5 5 8
8 1 1 1 2 4 6 6 7 9 9
9 2 2 3 3 5
Table 14: Stem-and-leaf plot of the mileages. It has the shape of the dot plot but keeps every original value.

The leaf counts are \(2, 3, 2, 3, 5, 10, 5\), totalling \(30\).

(c) With \(n=30\), using the \(\tfrac {1}{4}(n+1)\) rule: \begin {align*} Q_1 &= 7.75\text {th value} = 58+0.75(65-58) = 63.25\\ Q_2 &= 15.5\text {th value} = \tfrac {1}{2}(78+81) = 79.5\\ Q_3 &= 23.25\text {th value} = 87+0.25(89-87) = 87.5 \end {align*}

so the five number summary is \(33,\ 63.25,\ 79.5,\ 87.5,\ 95\), and \(IQR = 24.25\).

(d) The box sits well to the right of the middle of the whiskers: the lower whisker runs \(33\) to \(63.25\), a span of \(30.25\), while the upper runs \(87.5\) to \(95\), a span of only \(7.5\). The stem plot says the same thing, piling up at stem \(8\) and thinning downwards. The distribution is skewed to the left, and the mean, \(73.1\), falls below the median, \(79.5\), as it must.

Problem 1.9. [Tutorial Sheet 1] Given the data below:

3.3 3.7 4.4 5.7 9.4 9.8 6.1 4.4 3.7 3.3
3.4 3.8 4.5 6.1 10.0 11.0 6.3 4.7 3.9 3.5
3.6 3.9 4.7 6.3 16.0 28.0 6.7 5.0 4.0 3.6
3.7 4.0 5.1 7.2 7.4 5.1 4.1 3.7 3.7 4.3
5.2 7.4 7.8 5.2 4.4 3.7

(a).
Find \(P_{40}\) and \(P_{88}\).
(b).
Find the percentile rank of \(5.7\).
(c).
Construct a boxplot.
(d).
Describe the shape of the distribution.

Show solution

Solution. There are \(n=46\) observations. Order them first.

(a) The \(k\)th percentile is the \(\tfrac {k}{100}(n+1)\)th value. \[P_{40}:\quad \tfrac {40}{100}(47)=18.8\text {th value}=4.2+0.8(4.3-4.2)=4.26\] \[P_{88}:\quad \tfrac {88}{100}(47)=41.36\text {th value}=9.4+0.36(9.8-9.4)=9.54\]

(b) Thirty values lie below \(5.7\) and one equals it, so \[\text {percentile rank of }5.7=\frac {30+\tfrac {1}{2}(1)}{46}\times 100=66.3.\] Counting half of the tied value is the usual convention; it keeps the rank of the smallest and largest observations symmetric.

(c) Again with the \(\tfrac {1}{4}(n+1)\) rule, \(n+1=47\): \[Q_1=11.75\text {th}=3.7,\qquad Q_2=23.5\text {th}=4.6,\qquad Q_3=35.25\text {th}=6.4\] so \(IQR=6.4-3.7=2.7\) and the fences sit at \[Q_1-1.5\,IQR=-0.35,\qquad Q_3+1.5\,IQR=10.45.\] Three values lie beyond the upper fence — \(11.0\), \(16.0\) and \(28.0\) — and are plotted individually as outliers. The whiskers then run only to the most extreme values inside the fences, \(3.3\) and \(10.0\).

(d) Strongly skewed to the right. The median is \(4.6\) but the mean is \(5.97\), dragged up by the three outliers, and \(28.0\) is more than four times the upper quartile. This is the case where the median describes the data honestly and the mean does not.

Problem 1.10. [Assignment] The lives of electric lamps in hours, to the nearest hour, are given below. Form a grouped frequency table and illustrate by (a) a histogram, (b) a cumulative frequency curve.

724 695 716 730 689 700 689 726 662 681 676 732 697 710 694
715 738 696 696 682 699 714 707 697 710 660 703 717 692 689
684 695 682 721 708 722 692 717 656 697 701 699 705 680 702
690 663 695 670

Show solution

Solution. The question refers to \(50\) lamps but only \(49\) values are listed, so the working below uses the \(49\) given. Say so in an answer rather than quietly inventing a fiftieth.

The values run from \(656\) to \(738\), a range of \(82\), so classes of width \(10\) give nine classes — a sensible number for \(49\) observations.

Class 655–664 665–674 675–684 685–694 695–704 705–714 715–724 725–734 735–744
\(f\) 4 1 6 7 14 6 7 3 1
\(Cf\) 4 5 11 18 32 38 45 48 49
Table 15: Grouped and cumulative frequencies for the lamp lives.

(a) The histogram has bars standing on the boundaries \(654.5, 664.5, 674.5,\ldots \) with heights \(4, 1, 6, 7, 14, 6, 7, 3, 1\). The lives were recorded to the nearest hour, so the true limits sit at \(.5\) and the bars touch.

(b) The cumulative frequency curve is plotted against the upper boundary of each class — \((664.5, 4), (674.5, 5), (684.5, 11), (694.5, 18), (704.5, 32), (714.5, 38), (724.5, 45), (734.5, 48), (744.5, 49)\) — joined by a smooth curve from \((654.5, 0)\). Plotting against mid-points instead is a common slip: the count \(4\) is the number of lamps lasting up to \(664.5\) hours, not up to \(659.5\).

Problem 1.11. [Assignment] A class of \(40\) students were weighed, the results recorded to the nearest kg:

Mass (kg) \(53-56\) \(57-60\) \(61-64\) \(65-68\) \(69-72\) \(73-76\) \(77-80\)
Frequency 3 5 10 11 5 4 2

Using an assumed mean of \(66.5\) kg, estimate (a) the mean mass and (b) the standard deviation. If two additional students of masses \(53\) kg and \(78\) kg were included, would you expect the standard deviation to increase, decrease or stay the same? Justify.

Show solution

Solution. The classes all have width \(h=4\), so code with \(d=\frac {X-66.5}{4}\).

C-I \(53-56\) \(57-60\) \(61-64\) \(65-68\) \(69-72\) \(73-76\) \(77-80\) Total
\(f\) 3 5 10 11 5 4 2 40
\(X\) 54.5 58.5 62.5 66.5 70.5 74.5 78.5
\(d\) \(-3\) \(-2\) \(-1\) 0 1 2 3
\(fd\) \(-9\) \(-10\) \(-10\) 0 5 8 6 \(-10\)
\(fd^2\) 27 20 10 0 5 16 18 96
Table 16: Coded working for the masses, with \(a=66.5\) and \(h=4\).

(a) \[\overline {X}=a+h\,\frac {\sum fd}{\sum f}=66.5+4\left (\frac {-10}{40}\right )=66.5-1=65.5\ \text {kg}.\]

(b) \[\sigma ^2=h^2\left [\frac {\sum fd^2}{\sum f}-\left (\frac {\sum fd}{\sum f}\right )^{2}\right ] =16\left [\frac {96}{40}-\left (\frac {-10}{40}\right )^{2}\right ]=16\,[\,2.4-0.0625\,]=37.4\] \[\implies \quad \sigma =\sqrt {37.4}=6.12\ \text {kg}.\]

The two extra students. It increases. Their masses, \(53\) and \(78\), lie \(12.5\) kg either side of the mean of \(65.5\) — about twice the current standard deviation. Being symmetric about the mean, they leave the mean itself unchanged, but variance measures squared distance from the mean and both sit far out. The standard deviation rises from \(6.12\) to \(6.56\).

Note what the symmetry does and does not do: it protects the mean, not the spread. A pair of values balanced about the centre can leave the mean untouched while widening the distribution considerably.

Problem 1.12. [Assignment] A tyre manufacturer tests \(100\) tyres and records the distance travelled before the legal wear limit is reached:

Distance (km) 5000–25000 25000–35000 35000–45000 45000–55000 55000–65000 65000–85000
Number of tyres 8 14 24 26 16 12

(a).
Plot these data as a histogram.
(b).
Calculate the mean distance.
(c).
Obtain the cumulative frequency distribution and estimate the median and the interquartile range.

Show solution

Solution. (a) The classes are not all the same width: the first and last span \(20\,000\) km, the rest \(10\,000\). Plotting the frequencies as bar heights would make those two classes look twice as important as they are. The height must be the frequency density, \[\text {density}=\frac {f}{\text {class width}},\] so that area, not height, represents frequency.

Class 5–25 25–35 35–45 45–55 55–65 65–85
\(f\) 8 14 24 26 16 12
Width (thousand km) 20 10 10 10 10 20
Density (per 10 000 km) 4 14 24 26 16 6
Table 17: Frequency densities for the tyre data. Distances in thousands of km.

Drawn on frequency the last class would stand at \(12\); on density it stands at \(6\), below the \(55\)–\(65\) class. That is the honest picture, and it changes the apparent shape of the right-hand tail completely.

(b) Using mid-points \(15\,000\), \(30\,000\), \(40\,000\), \(50\,000\), \(60\,000\), \(75\,000\): \[\overline {X}=\frac {\sum fX}{\sum f}=\frac {4\,660\,000}{100}=46\,600\ \text {km}.\]

(c) Cumulative frequencies, against the upper boundary of each class:

Distance \(\leq \) (km) 25000 35000 45000 55000 65000 85000
\(Cf\) 8 22 46 72 88 100
Table 18: Cumulative frequency distribution for the tyre data.

With \(n=100\), interpolating inside the class that contains each position: \begin {align*} Q_1\ \left (\tfrac {n}{4}=25\right ):&\quad 35\,000+\frac {25-22}{24}(10\,000)=36\,250\\ Q_2\ \left (\tfrac {n}{2}=50\right ):&\quad 45\,000+\frac {50-46}{26}(10\,000)=46\,538\\ Q_3\ \left (\tfrac {3n}{4}=75\right ):&\quad 55\,000+\frac {75-72}{16}(10\,000)=56\,875 \end {align*}

so the median is about \(46\,500\) km and \[IQR=56\,875-36\,250=20\,625\ \text {km}.\]

The two testing methods. Rollers at constant speed remove almost everything that makes real wear variable — road surface, cornering, braking, load, driver, weather. The distribution would be far more tightly clustered, with a much smaller standard deviation, and probably centred differently, since steady rolling is gentler than town driving. The roller test is repeatable and quick, and it is the right tool for comparing one compound against another under identical conditions. The road trial is the one that tells a customer what mileage to expect. Neither replaces the other.

Problem 1.13. [Assignment] Given that the mean and standard deviation of a set of figures are \(\mu \) and \(\sigma \), write down the new mean and standard deviation when (a) each figure is increased by a constant \(c\), (b) each figure is multiplied by a constant \(k\).

A group of students sat examinations in algebra and biology. The algebra marks were scaled linearly, a mark \(x\) becoming \(ax-b\), so that the means and standard deviations of both examinations became equal.

Algebra Biology
Mean mark 48 62
Standard deviation 12 10

Find \(a\) and \(b\). A particular student scored \(36\) in algebra and \(48\) in biology. In what sense, if any, has he done better in algebra?

Show solution

Solution. (a) Adding a constant shifts everything without changing any distance between observations: the mean becomes \(\mu +c\), the standard deviation stays \(\sigma \).

(b) Multiplying stretches both: the mean becomes \(k\mu \), the standard deviation \(|k|\sigma \). The absolute value matters — a negative \(k\) flips the data but a standard deviation cannot be negative.

Finding \(a\) and \(b\). Under \(x\mapsto ax-b\) the standard deviation is multiplied by \(a\) and the mean becomes \(a\mu -b\). Matching biology: \[12a=10\quad \implies \quad a=\frac {10}{12}=\frac {5}{6}\] \[48a-b=62\quad \implies \quad b=48\left (\frac {5}{6}\right )-62=40-62=-22.\] So \(b=-22\) and the scaling is \(y=\frac {5}{6}x+22\). That \(b\) comes out negative is worth noticing: the rule as written subtracts \(b\), so a negative \(b\) means marks are being scaled up, which is what a mean of \(48\) needs to reach \(62\).

Comparing the student. Scaling algebra puts both on one footing, and \(36\) becomes \[\frac {5}{6}(36)+22=52,\] against \(48\) in biology. The same conclusion follows from standardising: \[z_{\text {algebra}}=\frac {36-48}{12}=-1.0,\qquad z_{\text {biology}}=\frac {48-62}{10}=-1.4.\] He is one standard deviation below the mean in algebra and \(1.4\) below in biology. He is below average in both, but less far below in algebra — that, and only that, is the sense in which he did better. The raw marks, \(36\) against \(48\), suggest the opposite.

Problem 1.14. [Assignment] A teacher’s marks have mean \(36\), median \(42\), standard deviation \(12\), interquartile range \(15\) and greatest mark \(72\). The marks are scaled by \(y=cx+d\) with \(c,d>0\) chosen so that the new marks have mean \(50\) and standard deviation \(8\). Find \(c\) and \(d\), and hence the new median, interquartile range and greatest mark.

Show solution

Solution. Under \(y=cx+d\) the standard deviation is multiplied by \(c\) and unaffected by \(d\), while the mean is multiplied by \(c\) and shifted by \(d\). Take the standard deviation first, since it involves only one unknown: \[12c=8\quad \implies \quad c=\frac {2}{3}\] \[36c+d=50\quad \implies \quad d=50-36\left (\frac {2}{3}\right )=50-24=26.\] So \(y=\frac {2}{3}x+26\).

The median and the greatest mark are positions on the scale, so both parts of the transformation apply: \[\text {median}=\frac {2}{3}(42)+26=54,\qquad \text {greatest}=\frac {2}{3}(72)+26=74.\] The interquartile range is a distance, and a shift moves both quartiles equally, so \(d\) cancels: \[IQR=\frac {2}{3}(15)=10.\] Adding \(26\) to the new \(IQR\) is the standard mistake here. The test is always the same: ask whether the quantity is a location or a spread.

Problem 1.15. [Assignment] A population consists of \(n_1\) males and \(n_2\) females. The mean heights are \(\mu _1\) and \(\mu _2\), and the variances \(\sigma ^2_1\) and \(\sigma ^2_2\). Show that the mean height of the whole population is \(w_1\mu _1+w_2\mu _2\), where \(w_1\) and \(w_2\) are the proportions of males and females.

Show solution

Solution. Write \(T_1\) for the total of the male heights and \(T_2\) for the females’. By the definition of a mean, \(T_1=n_1\mu _1\) and \(T_2=n_2\mu _2\). The combined mean is the combined total over the combined count: \[\mu =\frac {T_1+T_2}{n_1+n_2}=\frac {n_1\mu _1+n_2\mu _2}{n_1+n_2} =\frac {n_1}{n_1+n_2}\,\mu _1+\frac {n_2}{n_1+n_2}\,\mu _2=w_1\mu _1+w_2\mu _2\] where \(w_1=\frac {n_1}{n_1+n_2}\) and \(w_2=\frac {n_2}{n_1+n_2}\) are the proportions, so \(w_1+w_2=1\).

The combined mean is therefore a weighted average of the two group means, not the plain average \(\frac {1}{2}(\mu _1+\mu _2)\) — those agree only when \(n_1=n_2\).

Worth adding, since the question supplies the variances and then does not use them: the combined variance is not \(w_1\sigma _1^2+w_2\sigma _2^2\). It is \[\sigma ^2=w_1\sigma _1^2+w_2\sigma _2^2+w_1(\mu _1-\mu )^2+w_2(\mu _2-\mu )^2,\] the average of the within-group variances plus a term for how far apart the group means are. Two groups can each be tightly clustered and still give a widely spread population if their centres differ. This is the same decomposition that the analysis of variance rests on.

Problem 1.16. [Assignment] The mark distributions of two groups of students taking the same examination:

Mark range 0–29 30–39 40–49 50–59 60–69 70–79 80–100
Group A 4 2 4 11 7 6 6
Group B 2 4 4 5 11 8 11

(a).
Plot these distributions so as to illustrate them separately and to show up the difference between them.
(b).
Estimate the median, interquartile range, \(P_{80}\), \(D_6\) and \(P_{90}\).

Show solution

Solution. Two things must be settled before plotting anything.

First, the groups are different sizes: A totals \(40\) and B totals \(45\). (The question as set printed \(34\) for both, which does not match either row; the frequencies are the data, so they are what is used here.) Comparing raw counts would therefore mislead, and the comparison must be made on relative frequencies.

Second, the classes are not of equal width — \(0\)–\(29\) spans \(30\) marks and \(80\)–\(100\) spans \(21\), while the rest span \(10\). So the histogram must use frequency density, or the two end classes will be given a weight they have not earned.

(a) Draw each group as its own density histogram, then overlay the two cumulative frequency curves — converted to percentages — on one pair of axes. The overlay is what shows the difference: B’s curve lies consistently to the right of A’s, which is exactly what ”B did better” looks like on a graph.

(b) Marks are integers, so the true boundaries are \(-0.5, 29.5, 39.5, \ldots , 100.5\). Interpolating within the class containing each position:

Group A (\(n=40\)) Group B (\(n=45\))
\(Q_1\) 49.5 52.0
Median 58.6 66.3
\(Q_3\) 72.8 79.2
\(IQR\) 23.3 27.2
\(P_{80}\) 76.2 83.3
\(D_6\) 63.8 70.8
\(P_{90}\) 86.5 91.9
Table 19: Positional measures for the two groups. Every one of B’s exceeds A’s.

Group B is ahead at every point of the distribution, not merely on average — its lower quartile (\(52.0\)) is close to A’s median (\(58.6\)). B is also more spread out, \(IQR\) \(27.2\) against \(23.3\).

Problem 1.17. [Assignment] The table shows the cumulative distribution of gross weekly earnings of male full-time workers, in thousands.

Gross weekly earnings (K) 35 40 45 50 55 60 70 80 100 200
Manual (000s) 0.1 0.3 0.7 1.3 2.0 2.7 4.0 4.9 5.7 6.1
Non-manual (000s) 0.1 0.2 0.4 0.6 0.9 1.2 1.8 2.4 3.2 4.1

Estimate (a) the median earnings of a manual worker, (b) the median earnings of a non-manual worker, (c) the proportion of manual workers earning less than the median earnings of a non-manual worker.

Show solution

Solution. These figures are already cumulative — the entry against \(50\) is the number earning up to K\(50\), not the number in a class. Treating them as ordinary frequencies is the error this question is set to catch.

(a) Manual workers total \(6.1\) thousand, so the median is at \(3.05\) thousand. The cumulative frequency passes \(3.05\) between K\(60\) (\(2.7\)) and K\(70\) (\(4.0\)): \[\text {median}=60+\frac {3.05-2.7}{4.0-2.7}\times 10=60+2.69=\text {K}62.69.\]

(b) Non-manual workers total \(4.1\) thousand, so the median is at \(2.05\), which falls between K\(70\) (\(1.8\)) and K\(80\) (\(2.4\)): \[\text {median}=70+\frac {2.05-1.8}{2.4-1.8}\times 10=70+4.17=\text {K}74.17.\]

(c) Now read the manual column at K\(74.17\), interpolating between K\(70\) (\(4.0\)) and K\(80\) (\(4.9\)): \[Cf=4.0+\frac {74.17-70}{10}\times 0.9=4.375\ \text {thousand}\] \[\implies \quad \frac {4.375}{6.1}=0.717,\ \text {about }72\%.\] So roughly seven in ten manual workers earn less than the typical non-manual worker — a sharper statement of the gap than the two medians alone give.

Problem 1.18. [Assignment] Prove that \(\sum _{i=1}^{n}(x-\overline {x})^2=\sum _{i=1}^{n}x^2-n\overline {x}^2\).

In a middle school there are \(253\) girls whose ages have mean \(11.8\) years and standard deviation \(1.9\) years, and \(312\) boys whose ages have mean \(12.3\) years and standard deviation \(1.9\) years. Calculate the mean and standard deviation of all \(565\) pupils.

Show solution

Solution. The identity. Expand the square and use \(\sum x=n\overline {x}\): \begin {align*} \sum (x-\overline {x})^2 &=\sum \left (x^2-2x\overline {x}+\overline {x}^2\right )\\ &=\sum x^2-2\overline {x}\sum x+n\overline {x}^2\\ &=\sum x^2-2\overline {x}(n\overline {x})+n\overline {x}^2\\ &=\sum x^2-n\overline {x}^2. \qedhere \end {align*}

\(\overline {x}\) is a constant, which is what lets it come outside the sum, and \(\sum \overline {x}^2\) is \(n\) copies of it.

The combined group. The mean is the weighted average: \[\overline {x}=\frac {253(11.8)+312(12.3)}{565}=\frac {2985.4+3837.6}{565}=\frac {6823.0}{565}=12.08\ \text {years}.\]

For the standard deviation, each group contributes its own spread and the distance of its mean from the combined mean: \[\sigma ^2=\frac {\sum n_i\left [\sigma _i^2+(\mu _i-\overline {x})^2\right ]}{\sum n_i}.\] With \(\mu _1-\overline {x}=-0.276\) and \(\mu _2-\overline {x}=0.224\): \[\sigma ^2=\frac {253\left [3.61+0.0762\right ]+312\left [3.61+0.0502\right ]}{565} =\frac {932.7+1142.1}{565}=3.672\] \[\implies \quad \sigma =1.92\ \text {years}.\]

Both groups have standard deviation \(1.9\), yet the combined figure is slightly larger. It has to be: the two groups are centred half a year apart, and that separation is extra spread in the combined population. Simply averaging the two standard deviations would have given \(1.9\) and missed it.

Problem 1.19. [Assignment] The table shows the mean and standard deviation of marks obtained by two classes.

Mean s.d. Class size
Class A 67 5 15
Class B 58 8 18

Calculate (a) the mean mark for the two classes, (b) the standard deviation for the two classes. (c) A third class of \(17\) pupils took the same test, with mean \(61\) and standard deviation \(7\). Calculate the mean and standard deviation for all three classes combined.

Show solution

Solution. (a) \[\overline {x}=\frac {15(67)+18(58)}{33}=\frac {1005+1044}{33}=\frac {2049}{33}=62.09.\]

(b) Using the same decomposition as before, with \(67-62.09=4.91\) and \(58-62.09=-4.09\): \[\sigma ^2=\frac {15\left [25+24.11\right ]+18\left [64+16.73\right ]}{33} =\frac {736.6+1453.1}{33}=66.36\] \[\implies \quad \sigma =8.15.\] Larger than either class’s own standard deviation of \(5\) or \(8\), because the class means are nine marks apart.

(c) With the third class, \[\overline {x}=\frac {1005+1044+17(61)}{50}=\frac {2049+1037}{50}=\frac {3086}{50}=61.72\] and, with deviations \(5.28\), \(-3.72\) and \(-0.72\) from this new mean, \[\sigma ^2=\frac {15\left [25+27.87\right ]+18\left [64+13.84\right ]+17\left [49+0.52\right ]}{50} =\frac {793.1+1401.1+841.8}{50}=60.72\] \[\implies \quad \sigma =7.79.\] Adding a class whose mean sits between the other two reduces the combined standard deviation, from \(8.15\) to \(7.79\), even though that class has a standard deviation of \(7\) of its own. Filling in the middle of a distribution makes it less spread out.

Problem 1.20. [Assignment] \(100\) pupils were tested to determine their IQ, all values given to the nearest integer:

IQ 45– 55– 65– 75– 85– 95– 105– 115– 125–134
No. of pupils 1 1 2 6 21 29 24 12 4

(a).
Calculate the mean, mean deviation and standard deviation of the IQs.
(b).
Draw a cumulative frequency curve and estimate how many pupils have IQs within one standard deviation of the mean.

Show solution

Solution. Every class has width \(10\), so the mid-points are \(50, 60, 70, \ldots , 130\).

(a) \[\overline {X}=\frac {\sum fX}{\sum f}=\frac {10\,120}{100}=101.2\] \[\text {M.D}=\frac {\sum f\,|X-\overline {X}|}{\sum f}=\frac {1104.0}{100}=11.04\] \[\sigma ^2=\frac {\sum f(X-\overline {X})^2}{\sum f}=\frac {21\,056}{100}=210.56 \quad \implies \quad \sigma =\sqrt {210.56}=14.51.\] The mean deviation is smaller than the standard deviation, as it always is: squaring gives the far-out values more weight than taking their size does.

(b) Cumulative frequencies, against the upper boundary of each class:

IQ \(\leq \) 55 65 75 85 95 105 115 125 135
\(Cf\) 1 2 4 10 31 60 84 96 100
Table 20: Cumulative frequency distribution of the IQ scores.

One standard deviation either side of the mean runs from \[101.2-14.51=86.69\quad \text {to}\quad 101.2+14.51=115.71.\] Reading the curve at each end and interpolating, \[Cf(86.69)=10+\frac {1.69}{10}(21)=13.5,\qquad Cf(115.71)=84+\frac {0.71}{10}(12)=84.9\] \[\implies \quad 84.9-13.5\approx 71\ \text {pupils}.\] About \(71\%\), close to the \(68\%\) a normal distribution would give — which is a reasonable check that treating IQ as roughly normal is not unreasonable here.

Problem 1.21. [Test] The table gives the concentration of antibody in blood serum (g/l) for \(50\) subjects.

11.6 8.2 9.9 13.7 8.4 7.2 17.3 15.5 19.2 11.2 10.8 8.6
17.5 11.5 9.6 14.8 15.3 14.2 9.4 12.0 15.7 10.5 14.2 13.8
11.5 15.8 17.0 15.7 13.2 10.0 9.9 16.0 12.0 10.7 7.5 15.2
16.2 12.2 7.8 15.0 14.0 10.2 16.1 12.6 17.3 17.0 5.2 14.0
12.5 12.1

(a).
Construct a frequency table using a class interval of length \(1.9\).
(b).
Draw a histogram and a frequency polygon.
(c).
Calculate the standard deviation from the frequency table.
(d).
Calculate the coefficient of skewness and comment.

Show solution

Solution. (a) The values run from \(5.2\) to \(19.2\). Starting the first class at the minimum and stepping by \(1.9\):

Class 5.2–7.1 7.1–9.0 9.0–10.9 10.9–12.8 12.8–14.7 14.7–16.6 16.6–18.5 18.5–20.4
Mid-point \(X\) 6.15 8.05 9.95 11.85 13.75 15.65 17.55 19.45
\(f\) 1 6 9 10 7 11 5 1
Table 21: Frequency distribution of antibody concentration, class width 1.9.

The frequencies total \(50\). Each class is taken as \([\,\text {lower},\ \text {upper})\), with the last closed at both ends, so that every value falls in exactly one class.

(b) The histogram has bars of equal width \(1.9\) standing on those boundaries with heights \(1, 6, 9, 10, 7, 11, 5, 1\); the classes are all the same width, so frequency may be plotted directly. The polygon joins the mid-points of the bar tops.

(c) From the grouped table, \[\overline {X}=\frac {\sum fX}{\sum f}=\frac {638.1}{50}=12.76\] \[\sigma ^2=\frac {\sum f(X-\overline {X})^2}{\sum f}=\frac {514.34}{50}=10.29 \quad \implies \quad \sigma =3.21\ \text {g/l}.\] The raw data give \(\overline {X}=12.74\) and \(\sigma =3.20\), so grouping has cost almost nothing here — the classes are narrow relative to the spread.

(d) With median \(12.55\), the Pearson coefficient is \[\text {Sk}=\frac {3(\overline {X}-\text {median})}{\sigma }=\frac {3(12.76-12.55)}{3.21}=0.20.\]

Comment. The value is close to zero, so the distribution is very nearly symmetric, with only the faintest lean to the right. The frequency column supports this: it rises to \(10\), dips to \(7\), rises again to \(11\), and falls away at both ends. If anything the shape is slightly bimodal rather than skewed — a dip in the middle that a single skewness figure cannot express. This is worth saying, because a coefficient near zero is often read as ”symmetric and mound-shaped”, and only the histogram shows that the second half of that description does not hold.

Questions on this section

Stuck on something here? Ask below and it stays attached to this topic.