6.2 Rationale Of K-S
If the maximum distance between the assumed cdf and the empirical distribution function is small, the assumed cdf is plausible; if that distance is large, the data are telling us the assumed distribution is wrong. The whole test is that one sentence — everything else is working out how large a distance chance alone can produce.
Note 6.1. The statistic is distribution-free for a reason worth seeing. If \(F_0\) is the true continuous cdf then \(F_0(X)\) is uniform on \([0,1]\), so the whole comparison can be transported to the uniform distribution whatever \(F_0\) was. The null distribution of \(D_n\) therefore does not depend on \(F_0\) at all — which is exactly what makes one table of critical values serve every goodness-of-fit problem.
This also shows the limitation: the argument needs \(F_0\) fully specified in advance. If parameters are estimated from the same data, the null distribution changes and the ordinary table no longer applies.
Questions on this section
Stuck on something here? Ask below and it stays attached to this topic.