01 · Evaluation designs and causal inference
Process tracing for evaluation
Process tracing establishes causation within a single case by spelling out the mechanism through which an intervention is claimed to have produced a result, deriving what evidence should exist if that mechanism operated — and what should exist if it did not — and then testing. Its core discipline is the probative value of evidence: four classic tests grade each piece of evidence by whether it is necessary and/or sufficient to sustain the causal claim.
Last updated · Reviewed against 4 cited sources
What process tracing is
Process tracing is within-case causal inference: instead of comparing outcomes across treated and untreated units, it asks whether the causal mechanism claimed to connect an intervention to a result actually operated in this case — and it answers by hunting for the observable traces that mechanism would have left [1, 2]. If an advocacy campaign changed a ministry’s policy, the change did not happen by magic: officials read briefs, meetings occurred, draft language moved, positions shifted in sequence. Each step of the claimed mechanism implies evidence that should exist if the mechanism ran, and evidence that should exist instead if a rival explanation ran. The method is the disciplined comparison of those expectations against what the case record actually contains [1, 2, 3].
This makes process tracing the natural design for advocacy, governance and institutional-change work — single-case, deeply contextual, saturated with documents and witnesses — where the counterfactual designs earlier in this cluster have nothing to compare [3]. It is also the tool that gives theory-based evaluation its teeth: a theory of change asserts links; process tracing is how a specific link is tested rather than assumed [4]. The interviewing and document-review craft it depends on is covered in the qualitative methods page of the data cluster.
The four evidence tests
The heart of the method, systematised by Collier (building on Van Evera’s formulation), is that pieces of evidence differ enormously in what they can prove. Each test is defined by two properties: whether passing it is necessary for the causal claim to survive, and whether passing it is sufficient to strongly confirm the claim [1].
Worked through an M&E example — the claim that a coalition’s evidence briefs caused a budget-line increase:
- Straw-in-the-wind. Officials say the coalition was “active and visible”. Consistent with the claim; consistent with rivals too. Neither passing nor failing settles anything, but each such straw adds or subtracts weight [1].
- Hoop test. Did the budget officials receive and read the briefs before the decision? If no — the claim fails outright, whatever else is true. If yes, the claim survives the hoop but is not yet confirmed, because officials read many things [1].
- Smoking gun. The budget circular reproduces the brief’s cost tables and its distinctive framing. Finding this strongly confirms influence; not finding it does not kill the claim, since influence often leaves no verbatim trace [1].
- Doubly decisive. A dated internal minute recording that the committee reversed its prior position after, and citing, the coalition’s presentation — evidence that both confirms the mechanism and excludes the main rival (that the increase was already decided). Such evidence is rare; a study is lucky to find one such observation, which is why the other tests carry most analyses [1].
The practical lesson is about probative value: one genuine smoking-gun observation outweighs ten straws, and a single failed hoop outweighs any accumulation of favourable colour. Evaluation reports that count consistent anecdotes as if all evidence weighed the same miss the method’s entire point [1, 2]. Beach and Pedersen’s development of the method makes this explicit by asking, for each piece of evidence, how likely it would be to exist if the mechanism ran and how likely if it did not — informal Bayesian reasoning that sharpens the same intuition [2].
Three variants
Beach and Pedersen distinguish three uses of the method, worth naming because they discipline scope [2]:
| Variant | Starting point | Ambition |
|---|---|---|
| Theory-testing | A specified mechanism, hypothesised in advance | Test whether that mechanism operated in this case |
| Theory-building | An outcome and a hunch, no specified mechanism | Derive a mechanism from the case that can then be tested |
| Explaining-outcome | One puzzling case that matters in itself | Assemble the best sufficient explanation of this outcome |
Most evaluation applications are theory-testing: the programme’s theory of change supplies the hypothesised mechanism, and the tracing tests its critical links [2, 4]. Be honest about which variant a study is running — sliding from theory-testing into explaining-outcome mid-study, without saying so, is how conclusions outrun designs.
Rival mechanisms and the evidence table
Process tracing confirms a mechanism only relative to the alternatives considered, so rival mechanisms get the same treatment as the favoured one: specify each rival, derive its expected evidence, and run the tests. The strongest studies design their evidence collection around discrimination — seeking the observations for which the favoured mechanism and its rivals predict different findings, since evidence both explanations predict equally cannot separate them [1, 2].
The documentation discipline that keeps all this honest is the evidence table: one row per test, recording the claim at stake, the test type, the evidence expected under the claim and under rivals, the evidence actually found (with its source), and the inference drawn. The table converts a narrative that must be taken on trust into an audit trail a sceptic can check — and it pairs naturally with contribution analysis, where it becomes the engine room of steps 4 and 5: the contribution story supplies the links, and process tracing supplies the tests that weigh them.
Failure modes
Two failure modes account for most bad process tracing. Cherry-picked evidence: assembling the observations consistent with the favoured mechanism and stopping — the method’s machinery run in one direction only. The antidote is the pre-specified evidence table, populated with expected evidence before the search, so that absences are recorded as findings. Untested rivals: treating the programme’s theory as the only candidate and “confirming” it with straws-in-the-wind that a dozen rival explanations predict equally well [1, 2]. A third, subtler failure is over-claiming the scope: a traced mechanism shows what happened in this case; generalising beyond it needs either theory or cross-case work — which is where qualitative comparative analysis picks up.
Checklist before you commit to process tracing
- The question is whether and how a mechanism operated within a specific case, and the case record (documents, witnesses, timelines) is rich enough to test it.
- The claimed mechanism is specified step by step in advance, from the programme’s theory of change.
- Rival mechanisms are specified with the same care.
- Expected evidence — under the claim and under each rival — is written down before collection, with test types assigned.
- The evidence table is maintained and published, absences included.
- Conclusions state their scope: this case, this mechanism, this confidence.
Sources
- Understanding Process Tracing — PS: Political Science & Politics, 44(4), 823–830, 2011.Collier — the standard exposition of the four evidence tests.
- Process-Tracing Methods: Foundations and Guidelines, 2nd edition — University of Michigan Press, 2019.Beach & Pedersen — the book-length treatment, including the theory-testing, theory-building and explaining-outcome variants.
- Process Tracing (method page) — BetterEvaluation (Global Evaluation Initiative), updated continuously.Practitioner overview of process tracing in evaluation.
- Theory-Based Impact Evaluation: Principles and Practice. 3ie Working Paper 3 — International Initiative for Impact Evaluation (3ie), 2009.White — the programme-theory frame within which mechanism-level evidence earns causal weight.