9.3 Frequency distribution
Data as first collected is raw and unordered, and very little can be seen in it. The first job is always to organise it.
A shop recorded the number of loaves of bread sold in each of \(60\) successive one-hour periods:
| 22 | 14 | 24 | 39 | 13 | 23 | 19 | 22 | 9 | 14 |
| 24 | 21 | 32 | 16 | 22 | 39 | 16 | 15 | 21 | 21 |
| 26 | 22 | 16 | 22 | 38 | 29 | 5 | 20 | 15 | 25 |
| 20 | 34 | 27 | 27 | 22 | 21 | 23 | 29 | 24 | 19 |
| 29 | 22 | 20 | 16 | 17 | 17 | 5 | 28 | 25 | 20 |
| 28 | 24 | 23 | 5 | 19 | 24 | 31 | 33 | 30 | 13 |
Sixty numbers in no order tell us almost nothing. Grouping them into classes and counting how many fall into each gives a frequency distribution.
Note 9.6. Three rules guide the construction.
- (i).
- There should be neither too few classes nor too many. Too few hides the shape; too many leaves the table as scattered as the raw data. Between five and twelve is usually right.
- (ii).
- Equal class widths are preferred, so that frequencies can be compared directly. The first and last classes may be left open-ended to catch extreme values.
- (iii).
- Each class has a class mark, or midpoint, found by averaging the two class limits. It represents the whole class in later calculations.
Here the values run from \(5\) to \(39\). Taking a class width of \(5\) gives seven classes, from \(5\)–\(9\) up to \(35\)–\(39\). Tallying the data into them:
| Loaves sold | Class mark | Number of hours |
| 5 – 9 | 7 | 4 |
| 10 – 14 | 12 | 4 |
| 15 – 19 | 17 | 11 |
| 20 – 24 | 22 | 23 |
| 25 – 29 | 27 | 10 |
| 30 – 34 | 32 | 5 |
| 35 – 39 | 37 | 3 |
| Total | 60 | |
Note 9.7. Always total the frequency column and check it against the number of observations. If it does not come to \(60\), an item has been counted twice or missed, and everything built on the table afterwards will be wrong.
Example 9.8. The masses of \(24\) students, measured to the nearest kilogram, are
| 55 | 60 | 65 | 68 | 64 | 45 | 63 | 67 |
| 54 | 50 | 55 | 59 | 52 | 72 | 59 | 58 |
| 52 | 58 | 57 | 63 | 71 | 60 | 69 | 46 |
Construct a frequency table using a suitable class width.
Solution. Step 1 — find the range. \[\text {range}=\text {largest}-\text {smallest}=72-45=27.\] The classes must between them cover the whole of this spread.
Step 2 — choose the classes. A width of \(5\) divides the range into six classes, which is a reasonable number. Starting at \(45\): \[45\text {--}49,\quad 50\text {--}54,\quad 55\text {--}59,\quad 60\text {--}64,\quad 65\text {--}69,\quad 70\text {--}74.\]
Step 3 — tally and count.
| Mass (kg) | Frequency |
| 45 – 49 | 2 |
| 50 – 54 | 4 |
| 55 – 59 | 7 |
| 60 – 64 | 5 |
| 65 – 69 | 4 |
| 70 – 74 | 2 |
| Total | 24 |
9.3.1 Class boundaries and the width of an interval
Masses were recorded to the nearest kilogram, so a student in the \(50\)–\(54\) class actually has a mass anywhere from \(49.5\) kg up to \(54.5\) kg. These values are the class boundaries, and they are what a continuous variable really occupies.
The same set of classes can therefore be written in three ways, all meaning the same thing:
| Class limits | Class boundaries | Inequality form |
| 45 – 49 | 44.5 – 49.5 | \(44.5\leq m<49.5\) |
| 50 – 54 | 49.5 – 54.5 | \(49.5\leq m<54.5\) |
| 55 – 59 | 54.5 – 59.5 | \(54.5\leq m<59.5\) |
| 60 – 64 | 59.5 – 64.5 | \(59.5\leq m<64.5\) |
Definition 9.9. The width of a class is \[\text {width}=\text {upper class boundary}-\text {lower class boundary}.\]
Note 9.10. Widths are calculated from the boundaries, not the limits. For the class \(50\)–\(54\) the width is \(54.5-49.5=5\), not \(54-50=4\). Using the limits loses one unit from every class, which quietly corrupts every frequency density and every grouped average that follows.
Questions on this section
Stuck on something here? Ask below and it stays attached to this topic.