Statistical foundations · start here if you are new

How a medical study becomes data.

Learn what the rows and variables represent, who the findings can apply to and how chance, bias and confounding can shape an estimate.

Use this when: rows, variables, study designs or target populations feel unfamiliar.
You will make: a one-page map from the research question to the recorded data.
Then choose: uncertainty, p-values, epidemiology or a worked course from the Knowledge Hub.

Why this matters

Before analysing numbers, understand where they came from.

A study begins with people, care and measurements, not a spreadsheet. Before selecting a test, decide how participants entered the study, when exposure and outcome were measured, and what each recorded value means.

Bring: a research question if you have one. If not, you can still use this guide to learn the ideas. No software or calculations are needed.

By the end you can

  • distinguish a cohort, cross-sectional study, case-control study and trial;
  • separate source, study, analysis and target populations;
  • identify exposure, outcome and covariate roles and common variable types;
  • explain chance, selection bias, information bias and confounding;
  • state how missing data may change the analysis population and conclusion.

1 · Study design

Ask how people and time entered the evidence.

DesignHow it beginsWhat it can usually estimateBeginner caution
Cross-sectionalA sample measured at one timePrevalence and contemporaneous associationsExposure–outcome time order may be unclear.
Case-controlPeople sampled by outcome statusExposure odds and odds ratiosOrdinary disease risk cannot usually be read directly from the sample.
CohortExposure defined before later follow-upOutcome occurrence, follow-up prevalence and temporal associationsObserved exposure groups may differ before follow-up.
Randomised trialAn intervention allocated before follow-upIntervention effects under the planned comparisonAllocation, adherence, missing outcomes and analysis still matter.

Standalone example: a respiratory-clinic cohort

A respiratory clinic follows 600 adults after hospital discharge. Inhaler technique is checked at baseline and readmission is recorded six months later. Technique was observed, not randomly assigned, so the analysis estimates an association and must consider baseline differences such as disease severity.

2 · Populations and time

Follow who could enter, who did enter and who was analysed.

  1. 01
    Source population

    The wider group from which eligible participants could arise.

  2. 02
    Study sample

    The 600 eligible clinic patients represented in the study file.

  3. 03
    Analysis population

    The participants with the information required for a particular estimate or model.

  4. 04
    Target population

    The people to whom the final interpretation is intended to apply.

Time zero

When does follow-up begin?

Here, inhaler technique and baseline health are recorded at discharge, before six-month readmission status.

Selection

Who is missing before analysis begins?

Eligibility, participation and retention can make the sample differ from the target population.

Generalisability

Where can the result travel?

A precise estimate for the observed cohort is not automatically valid for all older adults or health systems.

3 · Variables and missingness

Name the role and measurement before choosing the method.

Role or typeRespiratory-clinic exampleWhy it matters
ExposureCorrect inhaler technique: yes/noDefines the main observed comparison.
OutcomeReadmission within six months: yes/noDetermines the target estimate and primary model family.
CovariateBaseline disease severityMay help address confounding when its causal role is justified.
Nominal categorySmoking status: never/former/currentUse labelled counts and percentages; there is no numeric distance between categories.
Ordered categoryMild, moderate, severe symptomsOrder matters, but gaps between categories are not assumed equal.
Continuous or countAge; previous admissionsDistribution, units, bounds, zeros and time at risk affect description and modelling.
Missing is not “no”

A blank readmission value means follow-up status was not observed, not that the participant avoided hospital.

Complete-case analysis

Using only records complete for every model variable changes the analysis population and may introduce selection bias.

Imputation

Imputation uses an explicit model to represent missing information; it is not inventing a convenient value and requires justified assumptions.

Carry the decision

Record missing counts, plausible reasons, planned handling and sensitivity checks in the Methods decision log.

Designing a survey rather than analysing existing data? The Survey Methods Handbook follows these population and measurement decisions into sampling, questionnaire design, fieldwork and data processing.

4 · Why an estimate can mislead

Keep chance, bias and confounding separate.

Chance

Samples vary.

A different sample can produce a different estimate. Standard errors and confidence intervals describe sampling uncertainty under the chosen model.

Selection bias

Who contributes may distort the comparison.

Entry, retention or complete-case inclusion can depend on exposure and health in ways that move the estimate.

Information bias

What is recorded may be wrong or unequal.

Misclassification or measurement error in technique, readmission or covariates can alter the observed association.

Confounding

A third cause can create or mask an association.

Disease severity may affect both inhaler technique and later readmission. Adjustment requires time order and causal reasoning.

A larger sample can reduce random uncertainty. It does not automatically remove selection bias, measurement error or confounding.

5 · Beginner self-check

Decide what the evidence represents.

Twelve six-month readmission values are blank. Can the analysis still contain all 600 participants without further work?

No. A complete-case outcome analysis can include at most 588 before considering other model variables. State the analysis population and investigate why information is missing.

Inhaler technique was recorded before readmission. Does that prove technique caused any later difference?

No. Temporality is necessary for a causal effect, but the observational groups may differ because of confounding, selection and measurement.

A very large sample gives a narrow confidence interval. Does that remove bias?

No. Precision and bias are different. A narrow interval can surround a systematically biased estimate.

Continue Read estimates and confidence intervals → Learn how a sample estimate, its precision and clinically important values should be interpreted.
Go deeper · medical-study design and bias

Optional authoritative resources for fuller design, bias and reporting guidance.

Study design · BMJ

Statistics at Square One

Clinical examples connecting study questions, designs and statistical procedures.

Open the BMJ chapter
Bias · Cochrane

Bias in health-research evidence

Structured explanations of selection, measurement, missing outcomes and selective reporting.

Open Cochrane Handbook chapter 7
Observational reporting · STROBE

STROBE checklists

Items that should be transparent in cohort, case-control and cross-sectional reports.

Open STROBE
Epidemiology · CDC

Principles of Epidemiology

A public-health introduction to study populations, measurements, comparisons and inference.

Open the CDC course