Primary FRCA · Statistics
Study Design & Bias
Study design determines what conclusions can legitimately be drawn from clinical research. Randomised controlled trials provide the strongest protection against confounding when properly randomised, concealed, blinded and analysed. Observational studies are indispensable for prognosis, rare harms and real-world practice but require careful attention to confounding and selection bias. Diagnostic studies require appropriate patient spectra, independent reference standards and blinded assessment. Bias is systematic error and must be prevented by design whenever possible. For the Primary FRCA, mastery requires recognising study types, linking each to its effect measures, and identifying the specific biases that threaten validity.
What this note covers
- Classify clinical study designs used in anaesthesia and peri-operative medicine and identify their strengths, limitations and appropriate uses.
- Explain randomisation, allocation concealment, blinding, intention-to-treat analysis and other design features that protect internal validity.
- Define, classify and recognise major forms of bias and confounding in trials, observational studies, diagnostic studies and systematic reviews.
- Interpret exam-style descriptions of study design and bias in the context of Primary FRCA statistics questions.
- Apply critical appraisal frameworks including CONSORT, STROBE and PRISMA to anaesthetic research.
Study Design and Bias for the Primary FRCA
Study design is the architecture by which a clinical question is converted into reliable evidence. Bias is systematic error: a predictable deviation of an estimate away from the truth due to the way a study is designed, conducted, analysed or reported. For the Primary FRCA, this topic is frequently examined because it underpins interpretation of trials in anaesthesia, critical care and peri-operative medicine. Candidates must be able to identify the type of study, the measure of association generated, the likely biases and the design features that reduce them.
The central distinction is between random error and systematic error. Random error reflects sampling variability and is reduced by increasing sample size; it is expressed by confidence intervals, standard errors and p values. Systematic error is bias; it is not corrected by a larger sample size and may make a very precise result wrong. A large, badly designed trial may provide a narrow confidence interval around a biased estimate.
Formulating the research question
A well-designed study begins with a structured question. In therapeutic research this is commonly framed as PICO: Population, Intervention, Comparator and Outcome. In diagnostic accuracy research, the equivalent structure is population, index test, reference standard and target condition. In prognostic research, it is population, prognostic factor or model, outcome and time horizon.
| Component | Example in anaesthesia | Design implication |
|---|---|---|
| Population | Adults aged 65 years and over undergoing urgent hip fracture surgery | Defines eligibility criteria and external validity |
| Intervention | Fascia iliaca block before spinal anaesthesia | Must be standardised: dose, timing, operator competence and ultrasound guidance |
| Comparator | Systemic opioid analgesia alone | Should reflect usual care or an active comparator |
| Outcome | Delirium within 5 postoperative days | Requires prespecified diagnostic criteria such as CAM or DSM-5 |
Good outcomes should be clinically important, valid, reliable and ideally patient-centred. Surrogate outcomes, such as a change in inflammatory markers or bispectral index values, may be useful mechanistically but can mislead if they do not predict outcomes that matter to patients, such as mortality, stroke, awareness, nausea, pain, disability or quality of recovery.
Hierarchy of evidence
The conventional hierarchy places systematic reviews and meta-analyses of high-quality randomised controlled trials at the top, followed by individual randomised trials, cohort studies, case-control studies, cross-sectional studies, case series and expert opinion. This hierarchy is useful but not absolute. A poorly conducted randomised controlled trial may be less credible than a well-designed prospective cohort study. Some questions cannot ethically or practically be answered by randomisation, such as the association between difficult airway management and aspiration risk.
| Study type | Typical question | Measure of effect | Main vulnerabilities |
|---|---|---|---|
| Randomised controlled trial | Does intervention X cause outcome Y? | Risk ratio, odds ratio, risk difference, mean difference, hazard ratio | Selection bias if allocation concealment fails, performance bias, attrition bias, low power |
| Cohort study | Does exposure predict later outcome? | Risk ratio, incidence rate ratio, hazard ratio | Confounding, loss to follow-up, immortal time bias |
| Case-control study | What exposures are associated with a rare outcome? | Odds ratio | Recall bias, selection bias, control selection bias |
| Cross-sectional study | What is the prevalence of a condition or association at one time? | Prevalence, prevalence ratio, odds ratio | Cannot establish temporality, non-response bias |
| Diagnostic accuracy study | How well does an index test identify disease? | Sensitivity, specificity, likelihood ratios, ROC AUC | Spectrum bias, verification bias, incorporation bias |
| Systematic review and meta-analysis | What is the totality of evidence? | Pooled effect estimate, heterogeneity statistics | Publication bias, poor search strategy, heterogeneity, selective outcome reporting |
Randomised controlled trials
A randomised controlled trial is an experimental study in which participants are allocated to intervention or control by a chance mechanism. Randomisation aims to balance known and unknown confounders between groups. It does not guarantee balance in an individual trial, especially if small, but it ensures that any imbalance is due to chance rather than investigator choice.
Randomisation methods
- Simple randomisation: analogous to repeated coin tosses. It is easy but may produce unequal group sizes in small trials.
- Block randomisation: ensures balance after every block, such as blocks of 4, 6 or 8. If block size is fixed and known, allocation may become predictable; variable undisclosed block sizes reduce this risk.
- Stratified randomisation: randomisation occurs within strata, such as ASA physical status I-II versus III-IV or centre. It helps balance strong prognostic variables.
- Minimisation: a dynamic allocation method that assigns the next participant to reduce imbalance across multiple covariates. It often includes a random component.
- Cluster randomisation: groups rather than individuals are randomised, for example operating theatres, hospitals or intensive care units. Analysis must account for intra-cluster correlation.
Allocation concealment is distinct from blinding and is crucial. It prevents investigators from knowing the next allocation before enrolling a participant. Adequate methods include central web-based randomisation or sequentially numbered, opaque, sealed envelopes prepared independently. Inadequate methods include alternation, date of birth, hospital number or unsealed envelopes. Failure of allocation concealment causes selection bias and often exaggerates treatment effects.
Blinding
Blinding prevents knowledge of allocation influencing care, outcome assessment or analysis. In drug trials, blinding may be straightforward if placebo is identical. In anaesthetic trials, blinding is often difficult because interventions such as regional anaesthesia, airway devices or ventilation strategies are visible to the clinician.
| Type of blinding | Purpose | Example |
|---|---|---|
| Participant blinding | Reduces placebo effects and reporting bias | Patient unaware whether antiemetic or placebo was administered |
| Clinician blinding | Reduces performance bias | Anaesthetist unaware of study infusion contents |
| Outcome assessor blinding | Reduces detection bias | Research nurse assessing postoperative delirium unaware of intraoperative intervention |
| Statistician blinding | Reduces analytical bias | Groups labelled A and B until analysis plan completed |
Terms such as single-blind, double-blind and triple-blind are ambiguous and should be replaced by explicit statements of who was blinded. In exam answers, state the mechanism: blinding reduces performance and detection bias but does not correct inadequate randomisation.
Control groups and comparators
Controls may be placebo, no treatment, standard care, active treatment or sham procedure. Placebo control is ethically acceptable only when no proven effective therapy is withheld or when add-on placebo is used. Sham procedures may be methodologically attractive but ethically complex because they expose patients to risk without direct benefit. In anaesthesia, equipoise is essential: comparing a new analgesic technique against an obviously inadequate analgesic regimen is unethical and inflates effect size.
Parallel, crossover and factorial designs
| Design | Description | Advantages | Limitations |
|---|---|---|---|
| Parallel-group trial | Participants allocated to one treatment pathway | Most common; suitable for acute peri-operative outcomes | Requires larger sample than crossover for continuous outcomes |
| Crossover trial | Each participant receives both treatments in random order separated by washout | Participant acts as own control; reduced between-subject variability | Not suitable for curative interventions, unstable disease or outcomes with carryover; washout must exceed drug effect, commonly at least 5 half-lives |
| Factorial trial | Participants allocated to combinations of interventions, such as 2 x 2 design | Can test two interventions and interaction efficiently | Requires assumption of no major interaction unless powered to detect it |
| Cluster trial | Randomises groups rather than individuals | Useful for service-level or educational interventions | Needs larger sample due to design effect and intra-cluster correlation |
| Stepped-wedge trial | Clusters switch from control to intervention at different randomised times | Attractive when intervention expected to help and rollout is inevitable | Vulnerable to secular trends and complex analysis |
Superiority, non-inferiority and equivalence trials
A superiority trial asks whether one intervention is better than another. A non-inferiority trial asks whether a new treatment is not unacceptably worse than standard therapy by more than a prespecified margin. An equivalence trial asks whether treatments differ by less than a prespecified margin in either direction. These designs are highly relevant to anaesthesia, for example comparing a newer airway device, vasopressor or reversal strategy with established practice.
- Superiority: null hypothesis is no difference. Common two-sided alpha is 0.05, with power 80% or 90%.
- Non-inferiority: requires a clinically justified non-inferiority margin, often expressed as an absolute risk difference or relative margin. Both intention-to-treat and per-protocol analyses are important; per-protocol is often more conservative for non-inferiority.
- Equivalence: requires two-sided margins; confidence interval must lie wholly within the equivalence bounds.
Non-inferiority trials are particularly vulnerable to poor adherence, crossover and protocol violations, because these dilute differences and may falsely suggest non-inferiority.
Sample size, power and error
Sample size calculations should be prespecified and based on the primary outcome, expected event rate or standard deviation, clinically important effect size, significance level and power. Standard values are two-sided α = 0.05 and power 80% or 90%, corresponding to β = 0.20 or 0.10. A 95% confidence interval corresponds to repeated intervals containing the true value in 95% of comparable experiments, not a 95% probability that this particular interval contains the true value.
| Concept | Meaning | Exam relevance |
|---|---|---|
| Type I error | False positive; concluding a difference exists when it does not | Probability α, commonly 0.05 |
| Type II error | False negative; missing a true difference | Probability β, commonly 0.20 or 0.10 |
| Power | Probability of detecting a true effect of specified size | 1 - β; commonly 80% or 90% |
| Minimal clinically important difference | Smallest difference worth detecting | Should be clinical, not merely statistical |
| Interim analysis | Analysis before completion | Requires alpha spending or stopping rules to avoid inflated Type I error |
Underpowered trials are common in anaesthesia. They may produce false-negative results or exaggerated estimates when positive. Multiplicity is also important: repeated analyses, multiple outcomes and subgroup testing increase the risk of spurious findings. Correction methods include Bonferroni adjustment, Holm procedures or prespecified hierarchical testing, but the best protection is a clear prespecified primary outcome and analysis plan.
Analysis populations
- Intention-to-treat analysis: participants are analysed in the groups to which they were randomised, regardless of adherence. It preserves the benefits of randomisation and reflects pragmatic effectiveness.
- Per-protocol analysis: includes only participants who adhered sufficiently to the protocol. It estimates biological efficacy but is vulnerable to selection bias.
- As-treated analysis: analyses according to treatment actually received. It may be useful for safety but can destroy randomisation.
- Modified intention-to-treat: a potentially problematic term unless clearly defined, for example excluding participants who never received any study drug.
For superiority trials, intention-to-treat is generally conservative because non-adherence dilutes treatment effects. For non-inferiority trials, intention-to-treat may falsely favour non-inferiority; both intention-to-treat and per-protocol analyses should be reported.
Observational study designs
Cohort studies
A cohort study follows exposed and unexposed groups over time to compare incidence of outcomes. Cohorts may be prospective or retrospective. Prospective cohorts allow careful exposure definition and outcome ascertainment but are expensive and slow. Retrospective cohorts use existing records and are efficient but depend on data quality.
Measures include risk ratio, risk difference, incidence rate ratio and hazard ratio. Hazard ratios from Cox proportional hazards models assume proportional hazards over time; if hazards cross, a single hazard ratio may be misleading. Cohort studies are preferred for rare exposures and for estimating incidence. They are poor for very rare outcomes unless very large.
Case-control studies
A case-control study begins with outcome status and looks backward for exposure. It is efficient for rare outcomes, such as peri-operative anaphylaxis, awareness under anaesthesia or malignant hyperthermia. It cannot directly estimate incidence or risk ratio unless nested in a defined cohort with incidence density sampling. The usual measure is the odds ratio, which approximates the risk ratio only when the outcome is rare, typically less than about 10%.
Control selection is critical. Controls should represent the exposure distribution in the population that produced the cases. Poor control selection produces severe selection bias. Matching can improve efficiency but matched variables cannot usually be evaluated as exposures unless special methods are used.
Cross-sectional studies
Cross-sectional studies measure exposure and outcome at a single time point. They estimate prevalence and are useful for surveys of practice, such as use of processed EEG monitoring or compliance with difficult airway guidelines. They cannot establish temporality, so causal inference is weak. Survivorship bias may occur if only those alive or available at the time of assessment are included.
Ecological studies
Ecological studies analyse groups rather than individuals, such as hospital-level anaesthetic staffing ratios and mortality. They are vulnerable to ecological fallacy: associations at group level may not hold at individual level.
Diagnostic accuracy studies
Diagnostic studies compare an index test with a reference standard. In anaesthesia, examples include airway ultrasound for predicting difficult laryngoscopy, gastric ultrasound for aspiration risk, or point-of-care tests for coagulopathy. Key indices are sensitivity, specificity, predictive values and likelihood ratios.
| Metric | Formula | Interpretation |
|---|---|---|
| Sensitivity | True positives / all with disease | Probability test is positive when condition is present |
| Specificity | True negatives / all without disease | Probability test is negative when condition is absent |
| Positive predictive value | True positives / all positive tests | Probability of disease given positive test; depends strongly on prevalence |
| Negative predictive value | True negatives / all negative tests | Probability of no disease given negative test; depends strongly on prevalence |
| Positive likelihood ratio | Sensitivity / (1 - specificity) | How much a positive test increases odds of disease; values greater than 10 are usually strong |
| Negative likelihood ratio | (1 - sensitivity) / specificity | How much a negative test reduces odds of disease; values less than 0.1 are usually strong |
Biases in diagnostic studies include spectrum bias, when the study includes only obvious cases and healthy controls rather than the real clinical spectrum; verification bias, when not all participants receive the reference standard; differential verification bias, when different reference standards are used; incorporation bias, when the index test forms part of the reference standard; and review bias, when assessors are not blinded.
Systematic reviews and meta-analysis
A systematic review uses explicit methods to identify, select, appraise and synthesise studies. A meta-analysis statistically pools results. High-quality reviews should follow PRISMA guidance, currently the PRISMA 2020 statement with a 27-item checklist, and ideally have a registered protocol such as PROSPERO. Randomised trials should be reported according to CONSORT, which uses a 25-item checklist and participant flow diagram. Observational studies should follow STROBE, which has a 22-item checklist.
Meta-analysis may use fixed-effect or random-effects models. A fixed-effect model assumes one true treatment effect and that observed differences are due to sampling error. A random-effects model assumes that true effects vary across studies and incorporates between-study variance. Heterogeneity is commonly assessed with I squared. Values around 25%, 50% and 75% are often described as low, moderate and high heterogeneity, although interpretation depends on context.
Publication bias occurs when statistically significant or favourable studies are more likely to be published. Funnel plot asymmetry may suggest publication bias but also arises from heterogeneity, small-study effects and methodological quality. Selective outcome reporting occurs when measured outcomes are omitted or reported selectively according to results.
Bias: definition and classification
Bias is systematic deviation from truth. It may affect selection of participants, measurement of exposure or outcome, delivery of co-interventions, follow-up, analysis or reporting. Bias threatens internal validity, meaning whether the observed effect in the study is true for the participants studied. External validity or generalisability concerns whether results apply to other patients, clinicians, hospitals and health systems.
| Bias | Mechanism | Example in anaesthesia | Prevention or mitigation |
|---|---|---|---|
| Selection bias | Systematic differences in baseline characteristics between groups | Sicker patients preferentially allocated to invasive monitoring because allocation not concealed | Randomisation, allocation concealment, clear inclusion criteria |
| Performance bias | Systematic differences in care apart from intervention | Clinicians give more antiemetics to unblinded high-risk patients in one arm | Blinding, standardised protocols, equal co-interventions |
| Detection bias | Outcome assessment differs between groups | Unblinded assessor more likely to diagnose agitation as emergence delirium in one group | Blinded outcome assessment, objective criteria |
| Attrition bias | Systematic differences due to incomplete outcome data | Patients with severe postoperative nausea lost before 24-hour follow-up | Minimise loss, intention-to-treat, sensitivity analyses |
| Reporting bias | Selective reporting of outcomes or analyses | Only statistically significant pain scores reported from multiple time points | Protocol registration, prespecified outcomes, CONSORT reporting |
| Recall bias | Differential accuracy of remembered exposure | Patients with awareness recall peri-operative events more intensively than controls | Use records, prospective exposure measurement |
| Observer bias | Observer expectations influence measurement | Investigator expects regional block to reduce pain and scores behaviour accordingly | Blinding, validated scales, training |
| Lead-time bias | Earlier diagnosis appears to prolong survival without changing disease course | Earlier detection of postoperative cognitive disorder changes measured duration but not outcome | Use mortality or fixed-time outcomes |
| Length-time bias | Screening preferentially detects slowly progressive disease | Less common in anaesthesia but relevant to peri-operative cancer screening pathways | Understand natural history, use randomised screening trials |
| Immortal time bias | Period during which outcome cannot occur is misclassified | Patients classified as receiving postoperative ICU therapy must survive long enough to receive it | Time-dependent exposure modelling |
| Confounding by indication | Indication for treatment is itself associated with outcome | Vasopressor use associated with mortality because shock severity caused both vasopressor use and death | Randomisation, adjustment, propensity methods, instrumental variables |
Confounding
Confounding occurs when a third variable is associated with both exposure and outcome and is not on the causal pathway. It creates a non-causal association or distorts a causal association. For example, an observational study may show that epidural analgesia is associated with worse postoperative outcomes because patients receiving epidurals undergo larger operations and have greater comorbidity. Operation magnitude and baseline risk confound the association.
A variable is a confounder if it satisfies three conditions: it is associated with the exposure, it is independently associated with the outcome, and it is not an intermediate step between exposure and outcome. Adjusting for mediators can underestimate true effect; adjusting for colliders can create spurious associations.
Methods to control confounding
| Method | Stage | Advantages | Limitations |
|---|---|---|---|
| Randomisation | Design | Balances known and unknown confounders on average | Chance imbalance possible; not always feasible or ethical |
| Restriction | Design | Limits confounding by excluding variable levels | Reduces generalisability |
| Matching | Design | Improves comparability and efficiency | Can overmatch; requires matched analysis |
| Stratification | Analysis | Shows effect within confounder strata | Limited when many confounders exist |
| Multivariable regression | Analysis | Adjusts for multiple measured covariates | Residual confounding if variables measured poorly or omitted |
| Propensity score | Analysis | Models probability of treatment given covariates | Controls only measured confounders; poor overlap limits inference |
| Instrumental variable | Analysis | Can address unmeasured confounding if assumptions hold | Assumptions often untestable; estimates local average treatment effect |
Regression adjustment is not magic. A model with too many covariates relative to outcome events overfits. A traditional rule of thumb is at least 10 outcome events per predictor variable in logistic regression, although modern penalised methods may differ. Model specification, missing data and collinearity matter. Multiple imputation is often preferable to complete-case analysis if data are missing at random, but it cannot rescue data missing not at random without sensitivity assumptions.
Effect modification and interaction
Effect modification occurs when the effect of an intervention differs across levels of another variable. It is not bias; it may be clinically important. For example, the benefit of processed EEG monitoring on awareness may differ between patients receiving total intravenous anaesthesia and volatile anaesthesia. Interaction should be tested directly rather than inferred from one subgroup being statistically significant and another not. Subgroup analyses should be prespecified, biologically plausible, few in number and interpreted cautiously.
Measures of association and clinical effect
Study design influences the appropriate measure. In a randomised trial with binary outcomes, absolute risk reduction and number needed to treat are often more clinically useful than relative risk reduction. A reduction in postoperative nausea from 40% to 30% gives an absolute risk reduction of 10%, relative risk reduction of 25%, and number needed to treat of 10. The same relative reduction from 4% to 3% gives an absolute reduction of only 1% and number needed to treat of 100.
| Measure | Calculation | Comments |
|---|---|---|
| Absolute risk reduction | Control event rate - treatment event rate | Directly clinically interpretable |
| Relative risk | Treatment event rate / control event rate | Preferred over odds ratio when risks are available |
| Relative risk reduction | 1 - relative risk | Often appears impressive; depends on baseline risk |
| Odds ratio | Odds in treatment / odds in control | Used in case-control and logistic regression; overestimates risk ratio when outcome common |
| Number needed to treat | 1 / absolute risk reduction | Round up to next whole patient; time horizon must be specified |
| Number needed to harm | 1 / absolute risk increase | Useful for adverse events |
Pragmatic versus explanatory trials
Explanatory trials ask whether an intervention can work under ideal conditions. They tend to have strict eligibility criteria, highly standardised protocols and expert operators. Pragmatic trials ask whether an intervention works in usual practice. They have broader eligibility criteria, flexible delivery and clinically relevant outcomes. In anaesthesia, explanatory trials may demonstrate that a block reduces pain when performed by experts, whereas pragmatic trials test whether a regional analgesia pathway improves recovery across multiple hospitals.
Internal validity is usually prioritised in explanatory trials, external validity in pragmatic trials. Exam questions may ask why a highly controlled single-centre study fails to apply to a district general hospital population: operator expertise, case mix, peri-operative pathways and outcome definitions may differ.
Landmark peri-operative examples
Several peri-operative studies illustrate design issues. The POISE trial studied peri-operative extended-release metoprolol and found fewer myocardial infarctions but more stroke and death, highlighting that surrogate cardiac endpoints may not capture net patient harm. ENIGMA investigated nitrous oxide avoidance and raised questions around composite outcomes and biologic plausibility. B-Aware examined bispectral index monitoring and intraoperative awareness in high-risk patients, illustrating outcome rarity and the need for enriched populations. CRASH-2 in trauma demonstrated mortality benefit with early tranexamic acid, while WOMAN extended evidence into postpartum haemorrhage; these large pragmatic trials show the value of simple protocols and clinically important outcomes.
Reporting and governance standards are also examinable. The Declaration of Helsinki requires ethical conduct, informed consent and protocol review. ICH-GCP provides international standards for trial conduct. In the UK, clinical trials of investigational medicinal products require regulatory approval and safety reporting. Trial registration before enrolment is expected to reduce selective reporting.
Critical appraisal framework for Primary FRCA
When reading a paper or answering an exam stem, proceed systematically:
- Identify the question: therapeutic, diagnostic, prognostic, aetiological or service evaluation.
- Identify the design: randomised trial, cohort, case-control, cross-sectional, diagnostic study or systematic review.
- Assess internal validity: randomisation, allocation concealment, blinding, completeness of follow-up, objective outcomes, prespecified analysis.
- Assess bias and confounding: selection, performance, detection, attrition, reporting bias and confounding by indication.
- Interpret effect size: absolute and relative measures, confidence intervals, p values and clinical importance.
- Assess precision: width of confidence intervals, sample size, event numbers and power.
- Assess applicability: patient population, intervention expertise, comparator relevance, outcomes and setting.
- Balance benefits and harms: especially important in anaesthesia where interventions may have low-frequency but catastrophic adverse outcomes.
Common exam traps
The Primary FRCA often tests definitions with subtle distinctions. Randomisation is not the same as allocation concealment. Blinding after allocation does not prevent selection bias before allocation. A statistically significant result is not necessarily clinically important. A non-significant result does not prove no effect, especially if the confidence interval includes clinically important benefit or harm. Observational association is not causation. Odds ratio is not the same as risk ratio when the outcome is common. Intention-to-treat analysis protects randomisation in superiority trials but may favour non-inferiority. Increasing sample size reduces random error but not systematic bias.
Causation
Causal inference is strongest in randomised experiments but may also be considered in observational data using Bradford Hill considerations: strength of association, consistency, temporality, biological gradient, plausibility, coherence, experiment and analogy. Temporality is essential. However, Bradford Hill criteria are not a checklist that proves causation. Modern causal thinking emphasises directed acyclic graphs, confounders, mediators and colliders. In viva discussion, a sophisticated answer recognises that causation is a judgement based on design, bias, confounding, effect size, temporality and biological plausibility.
Test your knowledge on this topic
Reading is only half the work. Put this note into practice with exam-style Primary FRCA questions, worked explanations and analytics that show exactly which topics still need attention. Start free — no card required.
Not sure where this topic fits in your revision? The Primary FRCA preparation guide sets out the exam format, the syllabus and a revision plan. You can also read how the Primary FRCA pass mark is determined.
