Pretesting normality

One common misunderstanding demonstrated in medical publications is the testing of normal distribution as part of a decision process: If Shapiro-Wilk Test or Kolmogorov-Smirnov Test indicates that a variable has a statistically significant departure from normal distribution, group mean differences are tested using the Mann-Whitney Test instead of Student's t-test. This decision rule may sound rock solid, but it may be a serious mistake.

First, the practice treats absence of evidence as evidence of absence. A statistically nonsignificant normality test is simply not evidence of a normal distribution.

Second, for a two-sample t-test or a linear statistical model, the relevant distributional issue concerns the difference or model residual conditional on covariates, not whether each observed treatment group's raw outcome values are exactly normal.

Third, with small sample sizes, where the normality assumption may be important, the normality test has low power to detect non-normality. With large samples, where non-normality matters the least, the test has power to detect small, scientifically irrelevant departures from normality.

Fourth, the decision rule affects the error rates as it is a two-stage analysis, and the final p-value is interpreted as if the final test had been selected in advance, but it is selected using the same data. Consequently, the 5% type-1 error rate for the final test may not be 5% for the combined rule (1).

Fifth, Student's t-test is used to test mean values, and the Mann-Whitney Test is not useful for testing mean values. It is often believed that the test can be used to test medians, but, in fact, it tests the equality of distributions. Therefore, a nonsignificant result from a Mann-Whitney Test does not say anything about mean treatment effects. Perhaps even more surprising, a significant result does not necessarily mean that medians differ. The Mann-Whitney Test has undesirable properties and requires careful examination of distributional symmetry and variance homogeneity before the test results are interpreted (2). 

Sixth, inducing normality through an appropriate transformation or using an unequal variance t-test like Satterthwaite's or Welch's might be a better alternative than using the Mann-Whitney Test (3).

References

1. Rochon J, Gondan M, Kieser M. To test or not to test: Preliminary assessment of normality when comparing two independent samples. BMC Med Res Methodol. 2012 Jun 19;12:81. doi: 10.1186/1471-2288-12-81. PMID: 22712852; PMCID: PMC3444333.

2. Fagerland MW, Sandvik L. The Wilcoxon-Mann-Whitney test under scrutiny. Stat Med. 2009 May 1;28(10):1487-97. doi: 10.1002/sim.3561. PMID: 19247980.

3. Graeme D. Ruxton, The unequal variance t-test is an underused alternative to Student's t-test and the Mann–Whitney U test, Behavioral Ecology, Volume 17, Issue 4, July/August 2006, Pages 688–690, https://doi.org/10.1093/beheco/ark016

Comments

Popular posts from this blog

Randomisation and alternation in clinical trials

Conditional and marginal models