01 · Evaluation designs and causal inference
Randomised controlled trials
A randomised controlled trial assigns eligible units to treatment or control by lottery, making the two groups statistically identical before the programme starts — so the difference in their later outcomes is an unbiased estimate of impact. The design's strength is internal validity; its limits are feasibility, ethics, cost, and the fact that a trial alone says little about why an effect occurred or whether it will travel.
Last updated · Reviewed against 7 cited sources
The identification problem
People and places that end up in programmes are not like those that do not. Farmers who join an extension scheme are more motivated; clinics selected for an upgrade were chosen for a reason; districts that adopt a policy early differ from late adopters. Comparing participants with non-participants therefore mixes two things that cannot be separated after the fact: the effect of the programme and the pre-existing differences that drove selection [1].
Randomisation cuts the knot before it forms. When a lottery decides who receives the programme, assignment is independent of everything about the units — observed and unobserved alike — so the treatment and control groups are statistically identical in expectation at baseline. Whatever difference in outcomes appears later has one remaining explanation: the programme [1, 2]. This is a precise, narrow claim. Randomisation buys internal validity — an unbiased estimate of the average impact for this population, this implementation, this period — and nothing else. Treat “gold standard” rhetoric with suspicion; the design solves the selection problem, not every problem.
Design choices that carry the trial
Unit and level of randomisation. Randomise individuals when the intervention reaches individuals and spillovers between them are small. Randomise clusters — schools, facilities, villages — when the intervention operates at group level, when individual assignment is administratively or politically impossible, or when treated and untreated neighbours would contaminate each other [2, 3]. The choice is consequential: clustering shrinks effective sample size, and a trial with thirty clusters is a very different proposition from one with three hundred, regardless of how many people sit inside them [3].
Stratification. Randomising within strata — region, facility size, baseline outcome bands — guarantees balance on the stratifying variables and improves precision, which matters most in small samples [3, 4].
Phase-in designs. Where everyone will eventually receive the programme, randomise the order of rollout: later cohorts serve as controls for earlier ones. This fits real expansion budgets and softens the ethics of withholding, at the price of losing the control group once rollout completes — long-run impacts become unmeasurable [4].
Encouragement designs. Where access cannot be denied, randomise an encouragement — an invitation, a subsidy, an outreach visit — and estimate the effect on those the encouragement induces to take up the programme [3].
| Variant | What is randomised | Fits when | Main cost |
|---|---|---|---|
| Simple lottery | Who receives the programme | Demand exceeds supply | Withholding from controls |
| Phase-in | When units receive it | Rollout is staged anyway | No long-run control group |
| Encouragement | An offer or nudge | Access cannot be denied | Estimates a complier effect only |
| Clustered | Groups, not individuals | Group-level delivery or spillovers | Power depends on cluster count |
Power before fieldwork
A trial that cannot detect the effect it was built to find is an expensive way to produce a confidence interval that includes everything. Before any data collection, fix the minimum detectable effect — the smallest impact that would still justify the programme — and work backwards through outcome variance, expected take-up, expected attrition and, for clustered designs, the intraclass correlation (the degree to which outcomes within a cluster resemble each other) [3]. Intraclass correlation is the quiet killer: even modest within-cluster correlation means that adding people to existing clusters buys almost no power, while adding clusters does [3, 4]. J-PAL publishes open power-calculation tools and worked examples [5]; use them at design stage, and file the assumptions so the endline analysis can be judged against them. Sample-size mechanics for the surveys themselves belong to the sampling literature — see the sampling page in the data cluster.
Threats between randomisation and results
Randomisation guarantees comparability at baseline. Everything after that is exposed:
- Attrition. Losing units at endline is survivable if it is random; differential attrition — dropouts who differ between arms — quietly rebuilds the selection bias the lottery removed. Report attrition rates by arm and bound the estimates where they differ [1, 2].
- Spillovers and contamination. Controls that benefit from treated neighbours (information, market effects, disease protection) bias the estimate towards zero — or, with negative spillovers, away from it. Choose the randomisation level to contain the spillover, or measure it deliberately with multi-level designs [2, 3].
- Behavioural responses to being studied. Participants who alter behaviour because they are observed or know their assignment (Hawthorne and John Henry effects) confound the treatment with the trial itself [2].
- Non-compliance. Some assigned to treatment never take it up; some controls find their way in. Analyse by assignment, not receipt: the intention-to-treat estimate preserves the randomised comparison and answers the policy question of what offering the programme achieves. Where the question is the effect of actual participation, use random assignment as an instrument for take-up to recover the local average treatment effect for compliers — and say plainly that it applies to compliers only [1, 4].
The ethics of randomising
The strongest ethical case for a lottery is scarcity: when a programme cannot reach everyone at once, random allocation among the equally eligible is arguably the fairest rationing rule available, and more transparent than discretion [4]. The case collapses in two situations: entitlements — no one may be randomised out of something they hold a legal or humanitarian right to — and settled evidence, where withholding an intervention of proven benefit from controls fails the equipoise test. Randomised designs need the same consent, review and data-protection discipline as any human-subjects research; the ethics and consent page in the MEAL cluster covers the machinery.
Where the field disagrees
Two live positions, both worth reading in the original. The randomisation-first school, institutionalised by J-PAL and codified in Duflo, Glennerster and Kremer’s toolkit, holds that randomised designs sit atop a hierarchy of evidence because they alone remove selection bias by construction, and that development spending should be steered towards what trials validate [3, 5]. The theory-based school, stated sharply in White’s 3ie paper, replies that a bare trial answers “did it work” for one implementation in one place and is silent on why — and that without an explicit programme theory, mapped from inputs to outcomes and tested along the causal chain, trial results cannot be interpreted, generalised or transported [6]. The UK Magenta Book takes the pragmatic line most working evaluators land on: start from the question, not the method, and choose experimental designs when the question is average causal impact and assignment is genuinely controllable [7].
The synthesis this site recommends: randomise when you can and it is ethical; embed the trial in a programme theory so mechanisms are measured, not assumed; and when assignment cannot be controlled at all, use the quasi-experimental and theory-based designs elsewhere in this cluster — contribution analysis in particular addresses the questions a trial cannot [6].
Cost and timeline realism
For NGO-scale programmes in East Africa, the binding constraints are rarely statistical. Dedicated baseline and endline surveys are usually the dominant cost line, and outcomes take as long to emerge as they take — a livelihoods effect measured eighteen months early is a null result waiting to be misread. Three practices keep trials affordable: piggyback on phased rollout that was happening anyway; measure through routine and administrative systems where a data-quality assessment shows they can bear the weight; and power the trial for the primary outcome only, resisting the questionnaire bloat that turns endlines into censuses [4, 5]. A trial that a programme cannot afford to finish is worse than a well-executed quasi-experimental design it can.
Checklist before you commit to an RCT
- The evaluation question is average causal impact of a defined intervention on defined outcomes.
- Assignment is genuinely controllable, and a lottery is ethically defensible (scarcity or phased rollout; no entitlement; no settled evidence).
- The randomisation unit contains expected spillovers, and power calculations — with intraclass correlation for clustered designs — precede fieldwork.
- Attrition, compliance and contamination will be tracked by arm, with intention-to-treat as the primary analysis.
- A programme theory states the mechanism, and intermediate outcomes along it are measured.
- The budget and timeline survive to endline.
Sources
- Impact Evaluation in Practice, Second Edition — World Bank / Inter-American Development Bank, 2016.Gertler, Martinez, Premand, Rawlings & Vermeersch. Chapters 3–4 cover the counterfactual problem and randomised assignment.
- Randomized Controlled Trials (RCTs). Methodological Briefs: Impact Evaluation No. 7 — UNICEF Office of Research, Florence, 2014.White, Sabarwal & de Hoop — a compact practitioner brief on RCT logic and threats.
- Using Randomization in Development Economics Research: A Toolkit — NBER Technical Working Paper 333, 2006.Duflo, Glennerster & Kremer — the standard technical treatment of randomisation designs, power and inference.
- Running Randomized Evaluations: A Practical Guide — Princeton University Press, 2014.Glennerster & Takavarasha — field-level guidance on randomisation mechanics, ethics and implementation.
- J-PAL Research Resources — Abdul Latif Jameel Poverty Action Lab, MIT, updated continuously.Open handbooks and tools on power calculations, randomisation and measurement.
- Theory-Based Impact Evaluation: Principles and Practice. 3ie Working Paper 3 — International Initiative for Impact Evaluation (3ie), 2009.White — the case for embedding counterfactual designs in an explicit programme theory.
- The Magenta Book: Central Government Guidance on Evaluation — HM Treasury, United Kingdom, 2020.Question-first method selection; positions experimental designs among the alternatives rather than above them.