1.4 Data Collection

Our interest in this course is quantitative data. Before anything can be done with it, one question has to be settled: where did it come from? Data are either gathered by the investigator, or taken from someone who gathered it earlier.

Definition 1.14 (Primary data). Primary data are collected first-hand, by the investigator, for the specific purpose at hand.

The usual methods are:

1.
Questionnaires – reach many respondents cheaply, but only the questions asked get answered, and a badly worded one cannot be repaired afterwards.
2.
Oral or personal interview – allows a vague answer to be followed up, at the cost of time and of the interviewer’s own influence on the reply.
3.
Observation – recording what people do rather than what they say they do, which avoids the gap between the two.
4.
Experiment – the investigator sets the conditions rather than waiting to see what occurs, and this is the only method that can establish cause.

Definition 1.15 (Secondary data). Secondary data were collected by somebody else, for some other purpose, and are being reused here.

Common sources are newspapers, journals, government and agency reports, and published statistical abstracts.

Remark 1.16. The trade-off between them is worth stating plainly, because it decides which to use.

Primary data are collected to answer your question, so they measure what you want, on the population you care about, in the form you need. They are expensive and slow.

Secondary data are immediate and often free, and may be far larger than anything you could gather yourself. But they were defined by someone else’s purpose: the categories may not be the ones you want, the population may not be quite yours, the collection may be out of date, and you cannot ask how carefully it was done. Whenever secondary data are used, the source should be recorded along with the figures.

Questions on this section

Stuck on something here? Ask below and it stays attached to this topic.