Methods note · regression and confounding

How do you choose covariates?

Decide what each variable is doing before putting it in a model. Use the research question, timing and clinical knowledge, not a list of bivariate p-values.

Use this when: you have a list of possible regression variables.
Bring: the outcome, main exposure, study timeline and purpose of the model.
You will make: a short record explaining why each variable is included or excluded.

Start with purpose

First decide what you want the model to do.

Explain an association

Estimate an exposure–outcome relationship

Adjust for relevant common causes of the exposure and outcome. Avoid automatically adjusting for variables that occur after the exposure or are consequences of two other variables.

Predict an outcome

Make useful predictions for new people

Use information that would genuinely be available at the prediction time, then check performance on data not used to fit the model.

Do not apply causal-adjustment rules to a prediction model. Likewise, do not apply prediction-selection procedures to a causal question. Write down the model’s purpose before choosing any covariates.

Practical workflow

Make every covariate decision visible.

  1. Define the target effect. Name the exposure, comparator, outcome, timeframe, population and whether the target is a total or direct effect.
  2. List candidates before outcome-led modelling. Use the protocol, clinical knowledge, literature and data dictionary, not a table of bivariate p-values.
  3. Order variables in time. Mark each as pre-exposure, exposure-time or post-exposure.
  4. Assign a causal role. Plausible cause of exposure, cause of outcome, common cause, mediator, collider, design variable, precision variable or duplicate/proxy.
  5. Draw the assumed pathways. Use a simple causal diagram to identify which non-causal paths must be blocked.
  6. Check the data can support the set. Review measurement quality, missingness, sparse categories, collinearity, outcome events and parameters.
  7. Pre-specify and justify. Record the chosen primary set and any sensitivity sets before interpreting the adjusted exposure estimate.
CandidateTiming and assumed roleDecisionReason
Variable namePre/post exposure; cause, mediator, collider, precision or design variableInclude / exclude / sensitivitySubstantive and causal justification

Standalone worked example

Exercise-class attendance and falls.

A community service follows older adults for six months. Attendance at a weekly strength-and-balance class is observed, not assigned. The question is whether regular attendance is associated with having a fall. The table shows a plausible causal starting point, not a truth discovered by software.

CandidateWorking roleTeaching decisionWhy
AgePre-attendance common causeIncludeMay influence both class attendance and fall risk.
Falls in the previous yearPrior outcome and common causeIncludeMay prompt attendance and strongly predicts another fall.
Baseline mobilityPre-attendance common causeIncludeMay affect ability to attend and later fall risk.
Long-term conditionsPre-attendance health statusIncludeMay affect access, participation and fall risk.
Distance from the venueAccess factorConsider and justifyMay affect attendance; include only if the assumed pathways support adjustment.
Balance at three monthsPossible mediatorExclude from the total-effect modelThe class may improve balance, which may then reduce falls.
Enjoyment of the classPost-attendance consequenceExclude from the primary modelIt is measured after attendance begins and may sit on the pathway to continued attendance.

A real study must justify this structure from its protocol, setting and evidence. No table or p-value can discover the true causal diagram.

Keep bivariate analysis in its proper role

Crude associations matter, but answer a different question.

Bivariate work does

Describe the unadjusted exposure–outcome association, its direction, magnitude and uncertainty.

Bivariate work also does

Reveal sparse cells, distributions, coding problems, unexpected patterns and the crude estimate used for comparison.

Bivariate work does not

Decide which variables are confounders or admit covariates to the model according to p<0.05.

Adjustment does

Use the target effect, temporal order, subject knowledge and assumed causal structure to select a defensible set.

The analytical sequence

Describe the crude association → select covariates independently → fit the adjusted model → compare crude and adjusted estimates → explain the difference.

Previous falls differ strongly between attendance groups. Is a small bivariate p-value the reason to adjust for them?

No. The comparison documents baseline imbalance. Previous falls belong in the proposed adjustment set because they occur before attendance and plausibly affect both attendance and the chance of another fall.

ContinueInterpret the adjusted estimate and confidence interval →Keep magnitude, precision and clinical meaning ahead of a threshold.
Go deeper · causal diagrams and covariate selection
Causal diagrams · free tool

DAGitty

Draw assumed causal structures and identify minimally sufficient adjustment sets.

Open DAGitty
Epidemiology · open book

What If

Hernán and Robins explain causal questions, confounding and adjustment through epidemiological examples.

Open the book
Reporting · observational studies

STROBE

Report how confounders were defined, why variables entered models and how adjusted estimates were obtained.

Open STROBE