11.3 Special Cases and Cautions
Note 11.4 (Special cases). Canonical correlation contains several familiar procedures. With \(p_1=p_2=1\) it is the ordinary correlation coefficient. With \(p_1=1\) and \(p_2>1\) the single squared canonical correlation is the coefficient of determination \(R^{2}\) from regressing that variable on the second set. And if the first set consists of dummy variables coding group membership, canonical correlation reproduces the discriminant analysis of Section 7. These are worth knowing: an unfamiliar method that contains regression and discriminant analysis as special cases is less unfamiliar than it looks.
Note 11.5 (Cautions). Canonical variates are chosen to maximise correlation, and with many variables and few observations a high canonical correlation can be produced from unrelated data by chance alone. Cross-validation, or Bartlett’s test of the significance of the remaining correlations, is essential before interpreting anything.
The variates are also frequently uninterpretable. Maximising a correlation places no premium on producing a combination that means anything, and a canonical variate weighting six variables with mixed signs may be real and still be useless. Report the correlation; interpret the weights only with care.
Questions on this section
Stuck on something here? Ask below and it stays attached to this topic.