01 · Topic cluster

Evaluation designs and causal inference

Randomised trials, difference-in-differences, regression discontinuity, matching, interrupted time series, synthetic control, and the theory-based methods that work where none of those fit.

Every design in this section answers the same question — did the intervention cause the change? — and each buys its answer with a different assumption. Randomisation buys it with the design itself. Quasi-experimental designs buy it with a comparison group and a claim about what would have happened otherwise. Theory-based and case-based methods buy it with the strength of the causal argument and the evidence marshalled for and against it.

The practical skill is not knowing the mathematics of each estimator. It is recognising which assumption your situation can actually support, and being honest in the report about what that assumption is. A design chosen because it was feasible, with its assumption stated and tested, is worth more than a design chosen because it sounded rigorous and whose assumption nobody examined.

Start with the decision aid below if you are choosing a design; go straight to a page if you already know which one you are defending.

The aid assumes the programme-logic artefacts already exist — building a theory of change or a logical framework is covered on monival.com, not here. Start from the causal question you have to answer.

Design-selection decision aid for evaluation designs and causal inference A flowchart. Start: what causal question are you asking? Question one: can you randomise assignment? Yes leads to randomised controlled trials. No leads to question two: is there a credible comparison group or assignment rule? Yes branches by situation — an eligibility cutoff on a score leads to regression discontinuity; adoption at a known time with several comparison units leads to difference-in-differences; adoption at a known time with one treated unit leads to synthetic control; rich baseline covariates only leads to matching and propensity scores; one population with a long outcome series leads to interrupted time series. No leads to question three: is the question about contribution, within or across cases? A contribution story weighed against rivals leads to contribution analysis; a mechanism within one case leads to process tracing; what works, for whom, in what circumstances leads to realist evaluation; recipes across a medium number of cases leads to qualitative comparative analysis. Every terminal links to its reference page. What causal question are you asking? 1 · Can you randomise assignment? yes no 2 · Is there a credible comparison group or assignment rule? yes eligibility cutoff on a score known start date, several comparison units known start date, one treated unit rich baseline covariates only one population, long outcome series no 3 · Is the question about contribution, within or across cases? contribution story, weighed against rivals mechanism within one case what works, for whom, in what circumstances causal recipes across 10–50 cases Randomised controlled trials Regression discontinuity Difference-in-differences Synthetic control Matching & propensity scores Interrupted time series Contribution analysis Process tracing Realist evaluation Qualitative comparative analysis
Decision aid. Work down the three questions in order: each design's page states the assumption the choice commits you to, and how to test it. Groups two and three overlap in practice — combining a counterfactual design with theory-based analysis is normal, not a compromise.
  1. 00 · Updated 18 August 2026

    Randomised controlled trials

    How RCTs identify causal impact: randomisation logic, unit and level choices, power, threats to validity, ethics, and when not to randomise.

  2. 01 · Updated 18 August 2026

    Difference-in-differences

    What difference-in-differences estimates, the parallel trends assumption it rests on, how to test it, and the pitfalls that break it.

  3. 02 · Updated 18 August 2026

    Regression discontinuity designs

    When a cutoff assigns a programme, RDD compares units just above and below it. Sharp vs fuzzy designs, bandwidth choice, and validity checks.

  4. 03 · Updated 18 August 2026

    Matching and propensity score methods

    Constructing comparison groups from observables: propensity scores, matching estimators, balance diagnostics, and the unobservables caveat.

  5. 04 · Updated 18 August 2026

    Interrupted time series analysis

    Evaluating interventions with a long outcome series and a clear start date: segmented regression, level vs slope change, seasonality, autocorrelation.

  6. 05 · Updated 18 August 2026

    Synthetic control methods

    Building a weighted synthetic comparison for one treated region or policy: donor pools, pre-period fit, placebo inference, and feasibility limits.

  7. 06 · Updated 18 August 2026

    Contribution analysis

    Mayne's six-step approach to credible causal claims without a counterfactual: programme theory, evidence, rival explanations, contribution story.

  8. 07 · Updated 18 August 2026

    Process tracing for evaluation

    Within-case causal inference using evidence tests — straw-in-the-wind, hoop, smoking gun, doubly decisive — applied to programme evaluation.

  9. 08 · Updated 18 August 2026

    Realist evaluation

    What works, for whom, in what circumstances: context–mechanism–outcome configurations, realist programme theory, and RAMESES quality standards.

  10. 09 · Updated 18 August 2026

    Qualitative comparative analysis (QCA)

    Cross-case causal analysis with sets: necessary and sufficient conditions, truth tables, crisp vs fuzzy sets, and when medium-N beats regression.