01 · Evaluation designs and causal inference
Realist evaluation
Realist evaluation, developed by Ray Pawson and Nick Tilley, starts from the observation that programmes do not work uniformly: they offer resources that people reason about and respond to, and those responses — the mechanisms — fire in some contexts and not others. The method builds and tests context–mechanism–outcome (CMO) configurations to answer not 'did it work?' but 'what works, for whom, in what circumstances, and why?'
Last updated · Reviewed against 4 cited sources
Generative causation: programmes do not work — people make them work
Realist evaluation begins with a claim about how causation operates in social programmes. Interventions are not treatments that act on passive subjects the way a drug acts on a body; they are offers of resources — information, skills, incentives, opportunities, sanctions — that enter people’s lives and get reasoned about. Whether anything changes depends on how participants respond to the offer, and that response depends on who they are and the circumstances they are in [1]. Pawson and Tilley’s formulation: programmes work by introducing resources that trigger mechanisms — changes in participants’ reasoning and behaviour — and mechanisms fire in some contexts and lie dormant in others, generating different outcomes in different settings [1].
This is generative causation, and it explains the pattern every experienced evaluator has met: the programme that succeeded brilliantly in one county and fell flat in the next, with identical activities and budgets. Under a purely counterfactual lens that pattern is noise — heterogeneity around an average effect. Under a realist lens it is the finding: the mechanism the programme relies on was enabled by something present in the first context and absent in the second, and the evaluation’s job is to identify what [1, 3]. Hence the method’s signature question, which replaces “did it work?”: what works, for whom, in what circumstances, and why? [1]
Writing CMO configurations that are not vacuous
The working unit of realist analysis is the CMO configuration: in context C, the programme triggers mechanism M, generating outcome O [1, 2]. The form is easy to write badly, and two disciplines separate a testable configuration from a truism:
The mechanism is the response, not the activity. “Training was delivered” is an activity; “SMS reminders were sent” is a resource. The mechanism is what happens in participants: the reasoning, motivation or capability change the resource provokes — “nurses gained confidence to challenge stock-card discrepancies”, “patients felt expected, so planned the journey”. If a claimed mechanism could be verified from the programme’s own activity reports alone, it is not a mechanism [1, 3].
The context must do explanatory work. Context is not background description — district names, population counts — but the specific conditions that switch the mechanism on or off: trust in the provider, phone ownership norms, staff turnover, supervision culture. A useful test: could a reader use the stated context to predict where else the programme would and would not work? If not, the C is decoration [1, 2].
Configurations at this altitude are falsifiable — a claim about where a mechanism fires can be checked in the next site — which is what makes realist evaluation a member of this cluster’s causal-inference family rather than a storytelling exercise.
The realist evaluation cycle
The method is theory-driven end to end [1, 3]:
- Elicit initial programme theories. Draw out, from designers, staff and prior research, the candidate CMO configurations — how is this programme supposed to change reasoning, for whom, under what conditions? These initial rough theories are the evaluation’s hypotheses.
- Test. Collect evidence designed to discriminate among the configurations: where should the programme work well, badly, differently? Sampling is purposive — sites and subgroups are chosen because the theories make different predictions about them, not to be representative.
- Refine. Revise the configurations against the evidence — mechanisms reformulated, contexts sharpened, outcomes differentiated — and, where the programme continues, test again. The output is a refined, transferable programme theory, not a single effect size.
Mixed methods are intrinsic, not a garnish: outcome patterns across contexts typically come from quantitative data, while mechanisms — reasoning inside people — are reached through qualitative work [1, 3]. The realist interview has its own craft, sometimes described as theory-gleaning: the evaluator puts the programme theory itself before informants, who confirm, refute and refine it from their position — a deliberate contrast with interviewing that fishes for unprompted themes [1].
Quality, cost, and choosing honestly
Realist evaluation has a quality infrastructure many theory-based approaches lack: the RAMESES II project publishes reporting standards, quality criteria and training materials for realist evaluation, developed for exactly the failure mode that follows fashionable methods — studies wearing the vocabulary without the discipline [2]. Commissioners should require conformance with the RAMESES II standards in the terms of reference and check reports against them.
The honest costs: realist evaluation demands substantial time, iterative fieldwork, and analysts who can move between quantitative outcome patterns and qualitative mechanism evidence without collapsing one into the other. It produces contingent, configurational findings that resist one-line summaries — a strength for programme design, a communications burden for reporting [2, 3]. And it is not always the right tool. Where the question is simply “what was the average effect of a uniform intervention?”, the counterfactual designs earlier in this cluster answer it more directly; UK government guidance places realist evaluation among the theory-based options selected when the question concerns how and for whom a complex intervention works [4]. Where results vary sharply by setting and the commissioning question is where to scale, adapt or stop — realist evaluation is often the only design actually addressing the question asked. It also pairs naturally with its neighbours: contribution analysis can frame the overall causal claim while realist configurations explain its variation, and process tracing supplies the within-case evidence tests for whether a claimed mechanism fired.
Checklist before you commit to realist evaluation
- The commissioning question is genuinely “what works, for whom, in what circumstances” — outcomes are known or expected to vary by context, and that variation matters for decisions.
- Initial programme theories can be elicited as candidate CMO configurations, with mechanisms stated as participant reasoning, not activities.
- Purposive site and subgroup selection is feasible, chosen to discriminate among the theories.
- The team can execute mixed methods and realist interviewing, and the budget covers iteration.
- Reporting will follow the RAMESES II standards, and findings will be presented as refined configurations — not compressed into a single average effect.
Sources
- Realist Evaluation — Paper prepared for the British Cabinet Office, 2004.Pawson & Tilley — the originators' own summary of the approach they first set out in Realistic Evaluation (SAGE, 1997).
- The RAMESES Projects: quality standards, reporting standards and training materials for realist evaluation and realist synthesis — RAMESES Projects, updated continuously.Home of the RAMESES II reporting and quality standards for realist evaluation.
- Realist Evaluation (approach page) — BetterEvaluation (Global Evaluation Initiative), updated continuously.Practitioner overview and resource collection.
- The Magenta Book: Central Government Guidance on Evaluation — HM Treasury, United Kingdom, 2020.Positions theory-based approaches, realist evaluation among them, within question-first method selection.