Biostatistics
What Is Biostatistics?
Biostatistics, historically called biometry, is the branch of statistics concerned with the design, analysis, and interpretation of studies in biology, medicine, and public health. It supplies the methods used to decide how many subjects a study needs, how they should be allocated to groups, which quantity actually answers the research question, and how much of an observed difference can be attributed to chance. Its subject matter is distinctive because biological measurements carry large person-to-person variation, outcomes are frequently observed incompletely, and the ethical constraints of human research limit what experiments may be run at all.
The discipline took shape around the turn of the twentieth century through the work of Francis Galton and Karl Pearson, who founded the journal Biometrika in 1901, and was reshaped by Ronald Fisher's formalization of randomization, replication, and blocking in experimental design. Austin Bradford Hill's 1948 streptomycin study for tuberculosis established the randomized controlled trial as the reference design for evaluating medical treatments, and that framework still organizes most of clinical biostatistics. An overview of biostatistics as a toolkit for exploring, validating, and interpreting clinical data traces how descriptive summaries, hypothesis tests, and modeling fit together across the stages of a study.
Study Design and Inference
Design decisions determine what can be concluded, and no analysis recovers what a flawed design has lost. Randomization balances measured and unmeasured confounders across arms in expectation, blinding prevents differential assessment, and stratification protects against imbalance in strong prognostic variables. Sample size calculation converts a clinically meaningful effect size, an assumed variability, and a target power into the number of subjects required. Analysis populations must be specified in advance, with the intention-to-treat principle keeping subjects in their assigned group regardless of what they actually received. Regulatory guidance has formalized this discipline: the ICH E9(R1) addendum on estimands and sensitivity analysis requires trials to define the targeted treatment effect through five attributes, including an explicit strategy for handling intercurrent events such as treatment discontinuation or rescue medication.
Time-to-Event and Longitudinal Methods
Many biomedical outcomes are not a single measurement but the time until an event occurs, and observation usually ends before every subject has experienced it. This censoring makes ordinary regression inappropriate and motivates a separate family of techniques. The Kaplan-Meier estimator produces a nonparametric survival curve, the log-rank test compares curves between groups, and the Cox proportional hazards model relates covariates to the hazard rate without specifying its baseline shape. An account of survival analysis in clinical trials covers these estimators along with the assumptions they rest on, including the requirement that censoring be independent of prognosis. Repeated measurements on the same subject bring a related problem, correlation within individuals, addressed through mixed effects models and generalized estimating equations.
High-Dimensional and Observational Data
Genomics, proteomics, and imaging generate data sets in which the number of measured variables far exceeds the number of subjects, so the classical concern with a single p-value gives way to controlling error across thousands of simultaneous tests. False discovery rate procedures, permutation-based null distributions, and penalized regression such as the lasso are standard responses. Where randomization is impossible, as in epidemiology, causal inference methods including propensity score matching, instrumental variables, and sensitivity analysis for unmeasured confounding are used to strengthen what can be claimed from observational cohorts and registries.
Applications
Biostatistics has applications in a range of fields, including:
- Clinical trial design, interim monitoring, and regulatory submission analysis
- Epidemiology and outbreak investigation, including infectious disease modeling
- Genetic association studies and genomic biomarker discovery
- Public health surveillance and health policy evaluation
- Diagnostic test evaluation using sensitivity, specificity, and receiver operating characteristic analysis
- Environmental and occupational exposure risk assessment
- Agricultural field experiments and animal breeding programs