04 · Data quality, sampling and collection
Mixed methods and secondary data in evaluation
A mixed-methods evaluation combines quantitative and qualitative strands so that each covers the other's blind spots: numbers establish magnitude and pattern, qualitative work explains mechanism and meaning. The design choices are few — sequential explanatory, sequential exploratory, or convergent — and the discipline that separates genuine mixing from two stapled studies is integration: planned points where the strands must meet.
Last updated · Reviewed against 4 cited sources
Why mix at all
Every method has a blind spot, and evaluation questions rarely respect methodological boundaries. A well-identified impact estimate can say the programme raised enrolment by six percentage points and still be silent on why it worked in the lake counties and failed in the arid ones, what the enrolment felt like from inside a household weighing school against herding, and whether anything unmeasured — teacher morale, informal fees — moved too. Bamberger’s framing of the case for mixing is the standard one: quantitative strands establish magnitude, distribution and comparability, qualitative strands establish mechanism, context and meaning, and combining them offsets weaknesses that neither can fix from within [1]. The same logic underpins theory-based impact evaluation: an estimate becomes an explanation only when evidence is gathered along the causal chain, and much of that evidence is qualitative [2].
The corollary deserves equal emphasis: mixing is a design decision with costs — two skill sets, two timelines, integration work — and a single-method study done well beats a mixed one done thinly. Mix because the questions demand it.
The three core designs
Mixed-methods notation marks the dominant strand in capitals and sequence with an arrow: QUANT → qual is a quantitatively led design with a subsequent qualitative phase [1].
Sequential explanatory (QUANT → qual). The quantitative strand runs first; qualitative work then explains what the numbers cannot. Its distinctive strength is sampling on the results: interviews at the sites with the largest and smallest effects, the households that dropped out, the facilities that outperformed their inputs. It fits endline evaluations where estimates will exist and need interpreting [1].
Sequential exploratory (QUAL → quant). The qualitative strand runs first to map a poorly understood phenomenon — how households actually finance school, what “participation” means locally — and the quantitative instrument is then built from what was learned: locally grounded categories, response options and hypotheses. It fits baselines in new contexts and measurement of contested constructs [1].
Convergent (QUANT + QUAL in parallel). Both strands run in the same window and are merged at analysis and interpretation. It is the pragmatic default when the timeline allows one field visit — and the design most at risk of never actually integrating, because nothing structurally forces the strands to meet until the report deadline [1].
Integration: the discipline that makes it “mixed”
Bamberger’s test is the right one: mixing can happen at design, at data collection and analysis, and at interpretation — and a genuinely mixed evaluation integrates at more than one of these stages [1].
- At design: the strands are built to interrogate each other. The survey carries questions the qualitative scoping surfaced; the interview guide targets the constructs the survey will estimate; the sampling frames link (qualitative cases drawn from survey clusters, so the two datasets share identifiable units).
- At analysis: the strands exchange partway. Survey results select the interview sample (explanatory); qualitative typologies become quantitative coding categories; qualitative claims (“stockouts drive default”) are tested against the administrative series.
- At interpretation: findings are confronted jointly, claim by claim — a matrix with the evaluation questions as rows and each strand’s evidence as columns is the simplest tool that forces the confrontation.
An evaluation that collects both kinds of data but integrates at none of these stages is not mixed methods; it is two studies stapled together, and the staple is the cover page.
Triangulation, including when it fails
Triangulation — comparing evidence on the same question across methods or sources — has three legitimate outcomes, and a credible report shows all three where they occur [1][3]:
- Convergence: the strands agree; confidence rises.
- Complementarity: the strands answer adjacent parts of the question — the survey shows that uptake fell among adolescent girls, the groups show why — and together give a fuller account.
- Dissonance: the strands disagree. This is a finding, not a failure. The survey says satisfaction is high; the interviews are full of grievance. Investigate before adjudicating: courtesy bias in the survey, an unrepresentative interview sample, or two real populations behaving differently. Reporting only the strand that flattered the programme is selective reporting, whatever the methods section says.
Where the evaluation must support causal claims without an experiment, this triangulation discipline is the engine of theory-based approaches — assembling evidence for each link of the causal chain and confronting it with rival explanations [2] — developed further under contribution analysis.
Secondary and administrative data: appraise before use
Most evaluations inherit data they did not collect: HMIS and EMIS extracts, programme monitoring records, national survey microdata, partner reports. Secondary data can anchor baselines, supply comparison series and replace expensive primary collection — but it earns its place through appraisal, not availability. A workable checklist [1][4]:
- Provenance. Who collected it, under what incentives? Programme-reported results carry the reporting pressures of the programme; see data quality dimensions for the failure catalogue.
- Definition match. Does the source’s definition (of “attended”, “functional”, “household”) match the evaluation’s? Close-but-different definitions are more dangerous than obviously different ones.
- Coverage and period. Whom does the source not see — private providers, informal settlements, non-reporting sites — and do its periods align with the exposure being evaluated?
- Known quality flags. Reporting-completeness rates, revision history, documented breaks in series (a form redesign mid-way is a break, not a trend).
- Access and ethics. Individual-level administrative data engages the data-protection duties covered under data protection, including a lawful basis for the secondary use.
The distinction between routine monitoring data and survey data — cost, frequency, bias profile — shapes which questions each can carry; that architecture question is treated under M&E information systems.
Team, budget and the garnish problem
Mixed methods fail in the budget line before they fail in the field. The recurring pattern: the qualitative strand is costed as an afterthought, staffed by whoever is free, fielded after the survey exhausts the calendar, and “analysed” by pasting quotes into the draft. Bamberger’s counsel is structural: budget the qualitative strand for its full pipeline (design, skilled collection per the standards in qualitative data collection, transcription, coding, integration time), put a named analyst on it, and timetable the integration points as deliverables [1][3]. If the design cannot fund a real second strand, the honest choice is a single-method study with clearly stated limits — not garnish.
Checklist for a mixed-methods design
- The design (explanatory, exploratory, convergent) is named, with priority notation, and justified by the questions.
- Integration points are specified at design time, with a named owner and a deliverable at each.
- Qualitative and quantitative samples link where the design needs them to (shared clusters, results-based case selection).
- Triangulation outcomes — including dissonance — are reported per evaluation question.
- Every secondary source passed the appraisal checklist, documented in an annex.
- The qualitative strand has its own budget, timeline and analyst.
Sources
- Introduction to Mixed Methods in Impact Evaluation (Impact Evaluation Notes No. 3) — InterAction / The Rockefeller Foundation, 2012.Bamberger. The standard practitioner treatment of mixed-methods designs, priority notation and integration stages in evaluation.
- Theory-Based Impact Evaluation: Principles and Practice (3ie Working Paper 3) — International Initiative for Impact Evaluation (3ie), 2009.White. Why credible impact evaluation opens the causal chain with mixed evidence rather than reporting an estimate alone.
- Qualitative Research Methods: A Data Collector's Field Guide — Family Health International (FHI), 2005.Mack, Woodsong, MacQueen, Guest & Namey. The qualitative strand's collection standards, on which any mixed design depends.
- Impact Evaluation in Practice, Second Edition — World Bank / Inter-American Development Bank, 2016.Gertler, Martinez, Premand, Rawlings & Vermeersch. Situates qualitative and administrative data within counterfactual evaluation designs.