Examrix

Primary FRCA · Statistics

Study Design & Bias

Study design determines what conclusions can legitimately be drawn from clinical research. Randomised controlled trials provide the strongest protection against confounding when properly randomised, concealed, blinded and analysed. Observational studies are indispensable for prognosis, rare harms and real-world practice but require careful attention to confounding and selection bias. Diagnostic studies require appropriate patient spectra, independent reference standards and blinded assessment. Bias is systematic error and must be prevented by design whenever possible. For the Primary FRCA, mastery requires recognising study types, linking each to its effect measures, and identifying the specific biases that threaten validity.

What this note covers

  • Classify clinical study designs used in anaesthesia and peri-operative medicine and identify their strengths, limitations and appropriate uses.
  • Explain randomisation, allocation concealment, blinding, intention-to-treat analysis and other design features that protect internal validity.
  • Define, classify and recognise major forms of bias and confounding in trials, observational studies, diagnostic studies and systematic reviews.
  • Interpret exam-style descriptions of study design and bias in the context of Primary FRCA statistics questions.
  • Apply critical appraisal frameworks including CONSORT, STROBE and PRISMA to anaesthetic research.

Study Design and Bias for the Primary FRCA

Study design is the architecture by which a clinical question is converted into reliable evidence. Bias is systematic error: a predictable deviation of an estimate away from the truth due to the way a study is designed, conducted, analysed or reported. For the Primary FRCA, this topic is frequently examined because it underpins interpretation of trials in anaesthesia, critical care and peri-operative medicine. Candidates must be able to identify the type of study, the measure of association generated, the likely biases and the design features that reduce them.

The central distinction is between random error and systematic error. Random error reflects sampling variability and is reduced by increasing sample size; it is expressed by confidence intervals, standard errors and p values. Systematic error is bias; it is not corrected by a larger sample size and may make a very precise result wrong. A large, badly designed trial may provide a narrow confidence interval around a biased estimate.

Formulating the research question

A well-designed study begins with a structured question. In therapeutic research this is commonly framed as PICO: Population, Intervention, Comparator and Outcome. In diagnostic accuracy research, the equivalent structure is population, index test, reference standard and target condition. In prognostic research, it is population, prognostic factor or model, outcome and time horizon.

ComponentExample in anaesthesiaDesign implication
PopulationAdults aged 65 years and over undergoing urgent hip fracture surgeryDefines eligibility criteria and external validity
InterventionFascia iliaca block before spinal anaesthesiaMust be standardised: dose, timing, operator competence and ultrasound guidance
ComparatorSystemic opioid analgesia aloneShould reflect usual care or an active comparator
OutcomeDelirium within 5 postoperative daysRequires prespecified diagnostic criteria such as CAM or DSM-5

Good outcomes should be clinically important, valid, reliable and ideally patient-centred. Surrogate outcomes, such as a change in inflammatory markers or bispectral index values, may be useful mechanistically but can mislead if they do not predict outcomes that matter to patients, such as mortality, stroke, awareness, nausea, pain, disability or quality of recovery.

Hierarchy of evidence

The conventional hierarchy places systematic reviews and meta-analyses of high-quality randomised controlled trials at the top, followed by individual randomised trials, cohort studies, case-control studies, cross-sectional studies, case series and expert opinion. This hierarchy is useful but not absolute. A poorly conducted randomised controlled trial may be less credible than a well-designed prospective cohort study. Some questions cannot ethically or practically be answered by randomisation, such as the association between difficult airway management and aspiration risk.

Study typeTypical questionMeasure of effectMain vulnerabilities
Randomised controlled trialDoes intervention X cause outcome Y?Risk ratio, odds ratio, risk difference, mean difference, hazard ratioSelection bias if allocation concealment fails, performance bias, attrition bias, low power
Cohort studyDoes exposure predict later outcome?Risk ratio, incidence rate ratio, hazard ratioConfounding, loss to follow-up, immortal time bias
Case-control studyWhat exposures are associated with a rare outcome?Odds ratioRecall bias, selection bias, control selection bias
Cross-sectional studyWhat is the prevalence of a condition or association at one time?Prevalence, prevalence ratio, odds ratioCannot establish temporality, non-response bias
Diagnostic accuracy studyHow well does an index test identify disease?Sensitivity, specificity, likelihood ratios, ROC AUCSpectrum bias, verification bias, incorporation bias
Systematic review and meta-analysisWhat is the totality of evidence?Pooled effect estimate, heterogeneity statisticsPublication bias, poor search strategy, heterogeneity, selective outcome reporting

Randomised controlled trials

A randomised controlled trial is an experimental study in which participants are allocated to intervention or control by a chance mechanism. Randomisation aims to balance known and unknown confounders between groups. It does not guarantee balance in an individual trial, especially if small, but it ensures that any imbalance is due to chance rather than investigator choice.

Randomisation methods

  • Simple randomisation: analogous to repeated coin tosses. It is easy but may produce unequal group sizes in small trials.
  • Block randomisation: ensures balance after every block, such as blocks of 4, 6 or 8. If block size is fixed and known, allocation may become predictable; variable undisclosed block sizes reduce this risk.
  • Stratified randomisation: randomisation occurs within strata, such as ASA physical status I-II versus III-IV or centre. It helps balance strong prognostic variables.
  • Minimisation: a dynamic allocation method that assigns the next participant to reduce imbalance across multiple covariates. It often includes a random component.
  • Cluster randomisation: groups rather than individuals are randomised, for example operating theatres, hospitals or intensive care units. Analysis must account for intra-cluster correlation.

Allocation concealment is distinct from blinding and is crucial. It prevents investigators from knowing the next allocation before enrolling a participant. Adequate methods include central web-based randomisation or sequentially numbered, opaque, sealed envelopes prepared independently. Inadequate methods include alternation, date of birth, hospital number or unsealed envelopes. Failure of allocation concealment causes selection bias and often exaggerates treatment effects.

Blinding

Blinding prevents knowledge of allocation influencing care, outcome assessment or analysis. In drug trials, blinding may be straightforward if placebo is identical. In anaesthetic trials, blinding is often difficult because interventions such as regional anaesthesia, airway devices or ventilation strategies are visible to the clinician.

Type of blindingPurposeExample
Participant blindingReduces placebo effects and reporting biasPatient unaware whether antiemetic or placebo was administered
Clinician blindingReduces performance biasAnaesthetist unaware of study infusion contents
Outcome assessor blindingReduces detection biasResearch nurse assessing postoperative delirium unaware of intraoperative intervention
Statistician blindingReduces analytical biasGroups labelled A and B until analysis plan completed

Terms such as single-blind, double-blind and triple-blind are ambiguous and should be replaced by explicit statements of who was blinded. In exam answers, state the mechanism: blinding reduces performance and detection bias but does not correct inadequate randomisation.

Control groups and comparators

Controls may be placebo, no treatment, standard care, active treatment or sham procedure. Placebo control is ethically acceptable only when no proven effective therapy is withheld or when add-on placebo is used. Sham procedures may be methodologically attractive but ethically complex because they expose patients to risk without direct benefit. In anaesthesia, equipoise is essential: comparing a new analgesic technique against an obviously inadequate analgesic regimen is unethical and inflates effect size.

Parallel, crossover and factorial designs

DesignDescriptionAdvantagesLimitations
Parallel-group trialParticipants allocated to one treatment pathwayMost common; suitable for acute peri-operative outcomesRequires larger sample than crossover for continuous outcomes
Crossover trialEach participant receives both treatments in random order separated by washoutParticipant acts as own control; reduced between-subject variabilityNot suitable for curative interventions, unstable disease or outcomes with carryover; washout must exceed drug effect, commonly at least 5 half-lives
Factorial trialParticipants allocated to combinations of interventions, such as 2 x 2 designCan test two interventions and interaction efficientlyRequires assumption of no major interaction unless powered to detect it
Cluster trialRandomises groups rather than individualsUseful for service-level or educational interventionsNeeds larger sample due to design effect and intra-cluster correlation
Stepped-wedge trialClusters switch from control to intervention at different randomised timesAttractive when intervention expected to help and rollout is inevitableVulnerable to secular trends and complex analysis

Superiority, non-inferiority and equivalence trials

A superiority trial asks whether one intervention is better than another. A non-inferiority trial asks whether a new treatment is not unacceptably worse than standard therapy by more than a prespecified margin. An equivalence trial asks whether treatments differ by less than a prespecified margin in either direction. These designs are highly relevant to anaesthesia, for example comparing a newer airway device, vasopressor or reversal strategy with established practice.

  • Superiority: null hypothesis is no difference. Common two-sided alpha is 0.05, with power 80% or 90%.
  • Non-inferiority: requires a clinically justified non-inferiority margin, often expressed as an absolute risk difference or relative margin. Both intention-to-treat and per-protocol analyses are important; per-protocol is often more conservative for non-inferiority.
  • Equivalence: requires two-sided margins; confidence interval must lie wholly within the equivalence bounds.

Non-inferiority trials are particularly vulnerable to poor adherence, crossover and protocol violations, because these dilute differences and may falsely suggest non-inferiority.

Sample size, power and error

Sample size calculations should be prespecified and based on the primary outcome, expected event rate or standard deviation, clinically important effect size, significance level and power. Standard values are two-sided α = 0.05 and power 80% or 90%, corresponding to β = 0.20 or 0.10. A 95% confidence interval corresponds to repeated intervals containing the true value in 95% of comparable experiments, not a 95% probability that this particular interval contains the true value.

ConceptMeaningExam relevance
Type I errorFalse positive; concluding a difference exists when it does notProbability α, commonly 0.05
Type II errorFalse negative; missing a true differenceProbability β, commonly 0.20 or 0.10
PowerProbability of detecting a true effect of specified size1 - β; commonly 80% or 90%
Minimal clinically important differenceSmallest difference worth detectingShould be clinical, not merely statistical
Interim analysisAnalysis before completionRequires alpha spending or stopping rules to avoid inflated Type I error

Underpowered trials are common in anaesthesia. They may produce false-negative results or exaggerated estimates when positive. Multiplicity is also important: repeated analyses, multiple outcomes and subgroup testing increase the risk of spurious findings. Correction methods include Bonferroni adjustment, Holm procedures or prespecified hierarchical testing, but the best protection is a clear prespecified primary outcome and analysis plan.

Analysis populations

  • Intention-to-treat analysis: participants are analysed in the groups to which they were randomised, regardless of adherence. It preserves the benefits of randomisation and reflects pragmatic effectiveness.
  • Per-protocol analysis: includes only participants who adhered sufficiently to the protocol. It estimates biological efficacy but is vulnerable to selection bias.
  • As-treated analysis: analyses according to treatment actually received. It may be useful for safety but can destroy randomisation.
  • Modified intention-to-treat: a potentially problematic term unless clearly defined, for example excluding participants who never received any study drug.

For superiority trials, intention-to-treat is generally conservative because non-adherence dilutes treatment effects. For non-inferiority trials, intention-to-treat may falsely favour non-inferiority; both intention-to-treat and per-protocol analyses should be reported.

Observational study designs

Cohort studies

A cohort study follows exposed and unexposed groups over time to compare incidence of outcomes. Cohorts may be prospective or retrospective. Prospective cohorts allow careful exposure definition and outcome ascertainment but are expensive and slow. Retrospective cohorts use existing records and are efficient but depend on data quality.

Measures include risk ratio, risk difference, incidence rate ratio and hazard ratio. Hazard ratios from Cox proportional hazards models assume proportional hazards over time; if hazards cross, a single hazard ratio may be misleading. Cohort studies are preferred for rare exposures and for estimating incidence. They are poor for very rare outcomes unless very large.

Case-control studies

A case-control study begins with outcome status and looks backward for exposure. It is efficient for rare outcomes, such as peri-operative anaphylaxis, awareness under anaesthesia or malignant hyperthermia. It cannot directly estimate incidence or risk ratio unless nested in a defined cohort with incidence density sampling. The usual measure is the odds ratio, which approximates the risk ratio only when the outcome is rare, typically less than about 10%.

Control selection is critical. Controls should represent the exposure distribution in the population that produced the cases. Poor control selection produces severe selection bias. Matching can improve efficiency but matched variables cannot usually be evaluated as exposures unless special methods are used.

Cross-sectional studies

Cross-sectional studies measure exposure and outcome at a single time point. They estimate prevalence and are useful for surveys of practice, such as use of processed EEG monitoring or compliance with difficult airway guidelines. They cannot establish temporality, so causal inference is weak. Survivorship bias may occur if only those alive or available at the time of assessment are included.

Ecological studies

Ecological studies analyse groups rather than individuals, such as hospital-level anaesthetic staffing ratios and mortality. They are vulnerable to ecological fallacy: associations at group level may not hold at individual level.

Diagnostic accuracy studies

Diagnostic studies compare an index test with a reference standard. In anaesthesia, examples include airway ultrasound for predicting difficult laryngoscopy, gastric ultrasound for aspiration risk, or point-of-care tests for coagulopathy. Key indices are sensitivity, specificity, predictive values and likelihood ratios.

MetricFormulaInterpretation
SensitivityTrue positives / all with diseaseProbability test is positive when condition is present
SpecificityTrue negatives / all without diseaseProbability test is negative when condition is absent
Positive predictive valueTrue positives / all positive testsProbability of disease given positive test; depends strongly on prevalence
Negative predictive valueTrue negatives / all negative testsProbability of no disease given negative test; depends strongly on prevalence
Positive likelihood ratioSensitivity / (1 - specificity)How much a positive test increases odds of disease; values greater than 10 are usually strong
Negative likelihood ratio(1 - sensitivity) / specificityHow much a negative test reduces odds of disease; values less than 0.1 are usually strong

Biases in diagnostic studies include spectrum bias, when the study includes only obvious cases and healthy controls rather than the real clinical spectrum; verification bias, when not all participants receive the reference standard; differential verification bias, when different reference standards are used; incorporation bias, when the index test forms part of the reference standard; and review bias, when assessors are not blinded.

Systematic reviews and meta-analysis

A systematic review uses explicit methods to identify, select, appraise and synthesise studies. A meta-analysis statistically pools results. High-quality reviews should follow PRISMA guidance, currently the PRISMA 2020 statement with a 27-item checklist, and ideally have a registered protocol such as PROSPERO. Randomised trials should be reported according to CONSORT, which uses a 25-item checklist and participant flow diagram. Observational studies should follow STROBE, which has a 22-item checklist.

Meta-analysis may use fixed-effect or random-effects models. A fixed-effect model assumes one true treatment effect and that observed differences are due to sampling error. A random-effects model assumes that true effects vary across studies and incorporates between-study variance. Heterogeneity is commonly assessed with I squared. Values around 25%, 50% and 75% are often described as low, moderate and high heterogeneity, although interpretation depends on context.

Publication bias occurs when statistically significant or favourable studies are more likely to be published. Funnel plot asymmetry may suggest publication bias but also arises from heterogeneity, small-study effects and methodological quality. Selective outcome reporting occurs when measured outcomes are omitted or reported selectively according to results.

Bias: definition and classification

Bias is systematic deviation from truth. It may affect selection of participants, measurement of exposure or outcome, delivery of co-interventions, follow-up, analysis or reporting. Bias threatens internal validity, meaning whether the observed effect in the study is true for the participants studied. External validity or generalisability concerns whether results apply to other patients, clinicians, hospitals and health systems.

BiasMechanismExample in anaesthesiaPrevention or mitigation
Selection biasSystematic differences in baseline characteristics between groupsSicker patients preferentially allocated to invasive monitoring because allocation not concealedRandomisation, allocation concealment, clear inclusion criteria
Performance biasSystematic differences in care apart from interventionClinicians give more antiemetics to unblinded high-risk patients in one armBlinding, standardised protocols, equal co-interventions
Detection biasOutcome assessment differs between groupsUnblinded assessor more likely to diagnose agitation as emergence delirium in one groupBlinded outcome assessment, objective criteria
Attrition biasSystematic differences due to incomplete outcome dataPatients with severe postoperative nausea lost before 24-hour follow-upMinimise loss, intention-to-treat, sensitivity analyses
Reporting biasSelective reporting of outcomes or analysesOnly statistically significant pain scores reported from multiple time pointsProtocol registration, prespecified outcomes, CONSORT reporting
Recall biasDifferential accuracy of remembered exposurePatients with awareness recall peri-operative events more intensively than controlsUse records, prospective exposure measurement
Observer biasObserver expectations influence measurementInvestigator expects regional block to reduce pain and scores behaviour accordinglyBlinding, validated scales, training
Lead-time biasEarlier diagnosis appears to prolong survival without changing disease courseEarlier detection of postoperative cognitive disorder changes measured duration but not outcomeUse mortality or fixed-time outcomes
Length-time biasScreening preferentially detects slowly progressive diseaseLess common in anaesthesia but relevant to peri-operative cancer screening pathwaysUnderstand natural history, use randomised screening trials
Immortal time biasPeriod during which outcome cannot occur is misclassifiedPatients classified as receiving postoperative ICU therapy must survive long enough to receive itTime-dependent exposure modelling
Confounding by indicationIndication for treatment is itself associated with outcomeVasopressor use associated with mortality because shock severity caused both vasopressor use and deathRandomisation, adjustment, propensity methods, instrumental variables

Confounding

Confounding occurs when a third variable is associated with both exposure and outcome and is not on the causal pathway. It creates a non-causal association or distorts a causal association. For example, an observational study may show that epidural analgesia is associated with worse postoperative outcomes because patients receiving epidurals undergo larger operations and have greater comorbidity. Operation magnitude and baseline risk confound the association.

A variable is a confounder if it satisfies three conditions: it is associated with the exposure, it is independently associated with the outcome, and it is not an intermediate step between exposure and outcome. Adjusting for mediators can underestimate true effect; adjusting for colliders can create spurious associations.

Methods to control confounding

MethodStageAdvantagesLimitations
RandomisationDesignBalances known and unknown confounders on averageChance imbalance possible; not always feasible or ethical
RestrictionDesignLimits confounding by excluding variable levelsReduces generalisability
MatchingDesignImproves comparability and efficiencyCan overmatch; requires matched analysis
StratificationAnalysisShows effect within confounder strataLimited when many confounders exist
Multivariable regressionAnalysisAdjusts for multiple measured covariatesResidual confounding if variables measured poorly or omitted
Propensity scoreAnalysisModels probability of treatment given covariatesControls only measured confounders; poor overlap limits inference
Instrumental variableAnalysisCan address unmeasured confounding if assumptions holdAssumptions often untestable; estimates local average treatment effect

Regression adjustment is not magic. A model with too many covariates relative to outcome events overfits. A traditional rule of thumb is at least 10 outcome events per predictor variable in logistic regression, although modern penalised methods may differ. Model specification, missing data and collinearity matter. Multiple imputation is often preferable to complete-case analysis if data are missing at random, but it cannot rescue data missing not at random without sensitivity assumptions.

Effect modification and interaction

Effect modification occurs when the effect of an intervention differs across levels of another variable. It is not bias; it may be clinically important. For example, the benefit of processed EEG monitoring on awareness may differ between patients receiving total intravenous anaesthesia and volatile anaesthesia. Interaction should be tested directly rather than inferred from one subgroup being statistically significant and another not. Subgroup analyses should be prespecified, biologically plausible, few in number and interpreted cautiously.

Measures of association and clinical effect

Study design influences the appropriate measure. In a randomised trial with binary outcomes, absolute risk reduction and number needed to treat are often more clinically useful than relative risk reduction. A reduction in postoperative nausea from 40% to 30% gives an absolute risk reduction of 10%, relative risk reduction of 25%, and number needed to treat of 10. The same relative reduction from 4% to 3% gives an absolute reduction of only 1% and number needed to treat of 100.

MeasureCalculationComments
Absolute risk reductionControl event rate - treatment event rateDirectly clinically interpretable
Relative riskTreatment event rate / control event ratePreferred over odds ratio when risks are available
Relative risk reduction1 - relative riskOften appears impressive; depends on baseline risk
Odds ratioOdds in treatment / odds in controlUsed in case-control and logistic regression; overestimates risk ratio when outcome common
Number needed to treat1 / absolute risk reductionRound up to next whole patient; time horizon must be specified
Number needed to harm1 / absolute risk increaseUseful for adverse events

Pragmatic versus explanatory trials

Explanatory trials ask whether an intervention can work under ideal conditions. They tend to have strict eligibility criteria, highly standardised protocols and expert operators. Pragmatic trials ask whether an intervention works in usual practice. They have broader eligibility criteria, flexible delivery and clinically relevant outcomes. In anaesthesia, explanatory trials may demonstrate that a block reduces pain when performed by experts, whereas pragmatic trials test whether a regional analgesia pathway improves recovery across multiple hospitals.

Internal validity is usually prioritised in explanatory trials, external validity in pragmatic trials. Exam questions may ask why a highly controlled single-centre study fails to apply to a district general hospital population: operator expertise, case mix, peri-operative pathways and outcome definitions may differ.

Landmark peri-operative examples

Several peri-operative studies illustrate design issues. The POISE trial studied peri-operative extended-release metoprolol and found fewer myocardial infarctions but more stroke and death, highlighting that surrogate cardiac endpoints may not capture net patient harm. ENIGMA investigated nitrous oxide avoidance and raised questions around composite outcomes and biologic plausibility. B-Aware examined bispectral index monitoring and intraoperative awareness in high-risk patients, illustrating outcome rarity and the need for enriched populations. CRASH-2 in trauma demonstrated mortality benefit with early tranexamic acid, while WOMAN extended evidence into postpartum haemorrhage; these large pragmatic trials show the value of simple protocols and clinically important outcomes.

Reporting and governance standards are also examinable. The Declaration of Helsinki requires ethical conduct, informed consent and protocol review. ICH-GCP provides international standards for trial conduct. In the UK, clinical trials of investigational medicinal products require regulatory approval and safety reporting. Trial registration before enrolment is expected to reduce selective reporting.

Critical appraisal framework for Primary FRCA

When reading a paper or answering an exam stem, proceed systematically:

  1. Identify the question: therapeutic, diagnostic, prognostic, aetiological or service evaluation.
  2. Identify the design: randomised trial, cohort, case-control, cross-sectional, diagnostic study or systematic review.
  3. Assess internal validity: randomisation, allocation concealment, blinding, completeness of follow-up, objective outcomes, prespecified analysis.
  4. Assess bias and confounding: selection, performance, detection, attrition, reporting bias and confounding by indication.
  5. Interpret effect size: absolute and relative measures, confidence intervals, p values and clinical importance.
  6. Assess precision: width of confidence intervals, sample size, event numbers and power.
  7. Assess applicability: patient population, intervention expertise, comparator relevance, outcomes and setting.
  8. Balance benefits and harms: especially important in anaesthesia where interventions may have low-frequency but catastrophic adverse outcomes.

Common exam traps

The Primary FRCA often tests definitions with subtle distinctions. Randomisation is not the same as allocation concealment. Blinding after allocation does not prevent selection bias before allocation. A statistically significant result is not necessarily clinically important. A non-significant result does not prove no effect, especially if the confidence interval includes clinically important benefit or harm. Observational association is not causation. Odds ratio is not the same as risk ratio when the outcome is common. Intention-to-treat analysis protects randomisation in superiority trials but may favour non-inferiority. Increasing sample size reduces random error but not systematic bias.

Causation

Causal inference is strongest in randomised experiments but may also be considered in observational data using Bradford Hill considerations: strength of association, consistency, temporality, biological gradient, plausibility, coherence, experiment and analogy. Temporality is essential. However, Bradford Hill criteria are not a checklist that proves causation. Modern causal thinking emphasises directed acyclic graphs, confounders, mediators and colliders. In viva discussion, a sophisticated answer recognises that causation is a judgement based on design, bias, confounding, effect size, temporality and biological plausibility.

Test your knowledge on this topic

Reading is only half the work. Put this note into practice with exam-style Primary FRCA questions, worked explanations and analytics that show exactly which topics still need attention. Start free — no card required.

Not sure where this topic fits in your revision? The Primary FRCA preparation guide sets out the exam format, the syllabus and a revision plan. You can also read how the Primary FRCA pass mark is determined.

Related Primary FRCA resources

Chosen from the same subject and closely related concepts.