Skip to content

Biostatistics & Epidemiologic Principles

Diagnostic Testing & Contingency Metrics

Every Step 3 test taker encounters 15–20 biostatistics questions. The cornerstone is the standard 2×2 contingency table.

Interactive 2×2 Diagnostic & Epidemiologic Calculator

Adjust the counts below to dynamically compute Sensitivity, Specificity, PPV, NPV, Likelihood Ratios, and Odds Ratio.

Legend: A = TP (true positive), B = FP (false positive), C = FN (false negative), D = TN (true negative).

Counts are whole numbers 0 or higher: decimals round down, blanks count as 0.

Test result versus disease status Disease Present (+) Disease Absent (-) Row Totals
Test Positive (+) 0
Test Negative (-) 0
Column Totals 0 0 0

Accuracy

Sensitivity (TPR)
n/a
A / (A + C) = TP / All Sick
Specificity (TNR)
n/a
D / (B + D) = TN / All Healthy
PPV
n/a
A / (A + B) (Prevalence Dependent)
NPV
n/a
D / (C + D) (Prevalence Dependent)

Likelihood ratios

Positive Likelihood Ratio (LR+)
n/a
Sensitivity / (1 - Specificity); ∞ means specificity is 100%
Negative Likelihood Ratio (LR-)
n/a
(1 - Sensitivity) / Specificity

Epidemiology

Odds Ratio (OR)
n/a
(A × D) / (B × C)
Prevalence
n/a
(A + C) / Grand Total

Study Design Matrix

Study DesignDefinitionMeasure of AssociationPrimary Weakness / Bias
Randomized Controlled Trial (RCT)Experimental allocation into treatment vs placebo groupsRelative Risk (RR), ARR, NNTLoss to follow-up, ethical constraints
Cohort StudyObservational: groups defined by exposure followed forward in time for outcomeRelative Risk (RR), IncidenceConfounding, loss to follow-up
Case-Control StudyObservational: groups defined by outcome evaluated retrospectively for exposureOdds Ratio (OR)Recall bias, selection bias
Cross-Sectional StudySnapshot in time assessing both exposure and outcome simultaneouslyPrevalence, Odds RatioCannot establish temporality
Ecological StudyData analyzed at the population/group level, not individualCorrelationEcological fallacy (attributing group traits to individuals)

Statistical Hypothesis Testing & Errors

flowchart TD
Truth["Real-World Truth"] --> H0True["H0 is True (No real difference)"]
Truth --> H0False["H0 is False (Real difference exists)"]
H0True --> Reject1["Study Rejects H0: TYPE I ERROR (alpha)"]
H0True --> Fail1["Study Fails to Reject H0: Correct Decision (1 - alpha)"]
H0False --> Reject2["Study Rejects H0: POWER (1 - beta)"]
H0False --> Fail2["Study Fails to Reject H0: TYPE II ERROR (beta)"]

Statistical Test Selection Guide

  • Comparing Means:
    • 2 groups: Two-sample t-test (e.g., mean blood pressure in Drug A vs Placebo).
    • 3+ groups: ANOVA (Analysis of Variance).
  • Comparing Proportions / Categorical Variables:
    • Categorical outcomes: Chi-square (χ²) test (e.g., percentage of smokers vs non-smokers).
    • Small sample size (any cell < 5): Fisher’s exact test.
  • Correlation & Regression:
    • Pearson r: Linear correlation between two continuous variables (-1 to +1).
    • R² (Coefficient of Determination): Proportion of variance in dependent variable explained by independent variable.

High-Yield Official NBME Exam Traps & Biostats Pearls

1. Post-Hoc Subgroup Analysis & Multiple Comparisons (p-Hacking)

  • Exam Scenario: An RCT evaluates a new drug against placebo. The primary endpoint shows no statistically significant benefit (p = 0.12). The investigators then slice the dataset into 15 post-hoc subgroups (by age, sex, BMI, smoking status) and report that in “men aged 45–55 who exercise”, the drug reached p = 0.03.
  • The Core Biostatistical Principle:
    • Performing multiple statistical tests without adjusting alpha dramatically inflates the Family-Wise Error Rate (Type I Error / False Positive Rate).
    • If 20 independent statistical comparisons are made at alpha = 0.05, the probability of finding at least one false-positive “significant” result purely by chance is: P(at least 1 false positive) = 1 - (1 - 0.05)^20 ≈ 64%
  • Step 3 Rule: Post-hoc subgroup analyses are strictly hypothesis-generating, NEVER confirmatory. True evidence requires a pre-specified hypothesis and correction for multiple testing (e.g. Bonferroni correction).

2. Confounding vs Effect Modification

FeatureConfoundingEffect Modification (Interaction)
DefinitionAn extraneous variable is independently associated with both the exposure AND the outcome, distorting the apparent associationAn external variable changes the magnitude or direction of the true biological effect across strata
Is it a Bias?YES: A nuisance bias/distortion that must be controlled and eliminatedNO: A real natural biological phenomenon that must be reported, not eliminated
Stratified Analysis TestStrata-specific Relative Risks are EQUAL to each other, but DIFFERENT from the crude (unadjusted) RRStrata-specific Relative Risks are DIFFERENT from each other
Method of ControlStudy Design phase: Randomization, Restriction, Matching.
Data Analysis phase: Stratification, Multivariable Regression
Stratification (report the stratum-specific effect sizes separately)

3. Ascertainment & Detection Bias

  • Exam Scenario: A study compares pulmonary nodules found on lung cancer screening CT scans vs routine chest radiographs. Both groups must undergo identical baseline diagnostic evaluation.
  • Pearl: If the control group is evaluated less aggressively or with less sensitive testing than the intervention group, an ascertainment (detection) bias occurs, falsely attributing increased disease frequency or complications to the intervention.

4. Clinical Significance vs Statistical Significance

  • Exam Scenario: A massive study with n = 50,000 subjects demonstrates that Drug X lowers systolic blood pressure by 0.8 mmHg compared to placebo (p < 0.001).
  • Pearl: With enormous sample sizes, tiny, clinically meaningless differences achieve extreme statistical significance (p < 0.001). On Step 3 drug ads, look at the absolute magnitude of benefit, number needed to treat (NNT), and adverse effect profile, not just the p-value.