Statistical inference and data science

Statistical inference focuses on quantifying uncertainty. It is primarily used to generalise information from samples by estimating unobservable parameters and testing hypotheses about the properties of a population that cannot be observed in its entirety. Data science, which many statisticians consider to be part of statistical science (machine learning), is mainly about exploring patterns in data and developing predictions.

For example, to answer the question, "How well will this treatment work for the next patient?" a machine learning approach may be useful, and the answer should include the risk of an erroneous classification. If the question is instead, "What is the average effect of this treatment?" statistical inference is necessary, and the answer should include an assessment of the uncertainty of the presented treatment effect estimate. The two approaches often use the same statistical methods, such as logistic regression analysis, but the methods are typically used differently, like employing different principles for including covariates into the analysis. 

Comments

Popular posts from this blog

Randomisation and alternation in clinical trials

Conditional and marginal models

Pretesting normality