04 · Data quality, sampling and collection
Sampling for M&E: methods and sample size
The sampling design decides what a survey can claim before a single interview happens. Probability designs — simple random, systematic, stratified, multi-stage cluster — support statements about a population; purposive and quota designs support statements about the cases selected. Sample size follows from precision, variance and disaggregation demands, and clustered designs need to be larger than the textbook formula suggests.
Last updated · Reviewed against 5 cited sources
The claim comes before the sample
A sampling design is a statement about the claims a study will be allowed to make. If the endline report must say “42% of households in the programme area have access to safe water, ±4 percentage points”, the sample must be a probability sample of households in the programme area, drawn from a defensible frame, sized for that margin. If the report will instead say “across a deliberately varied set of communities, these are the mechanisms by which water committees fail”, a purposive sample chosen for variation is the correct instrument — and a random one may be worse. Most sampling failures in M&E are not calculation errors; they are mismatches between the design used and the claim made [1].
The dividing line is selection probability. In a probability design, every unit in the population has a known, non-zero chance of selection, which is what licenses inference from sample statistics to population values and gives meaning to confidence intervals [1]. In a non-probability design, selection depends on judgement, convenience or referral; the sample can be deeply informative about itself but supports no statistical statement about a wider population.
The probability designs, and when each earns its complexity
Simple random sampling draws units directly from a complete list with equal probability. It is the benchmark against which everything else is measured and almost never what large field surveys actually do, because a complete, current list of households rarely exists and a truly random national scatter of interviews is unaffordable to visit [1].
Systematic sampling selects every k-th unit from a list or a walk pattern after a random start. It is operationally simple and, when the list order is unrelated to the outcome, behaves like simple random sampling. Its known hazard is periodicity in the list — a fixed interval marching in step with some regular structure in the data [1].
Stratified sampling divides the population into groups (counties, urban/rural, facility type) and samples within each. Stratification never hurts precision and usually helps; more importantly for M&E, it is the only way to guarantee enough interviews in each group you must report on separately. Strata that are oversampled relative to their population share must be weighted back at analysis [1].
Multi-stage cluster sampling is the workhorse of household surveys in low-resource settings: sample enumeration areas or villages first (often with probability proportional to size), then sample households within selected clusters. It concentrates fieldwork geographically, which is why it is affordable — and it is less precise per interview than the alternatives, which is the design effect discussed next [1][5].
The design effect: why cluster samples must be bigger
Households in the same cluster resemble each other — they share water sources, markets, clinics, shocks. Each additional interview within a cluster therefore adds less new information than an interview drawn independently from the whole population. The design effect (DEFF) quantifies the penalty: it is the ratio of the variance under the actual design to the variance a simple random sample of the same size would have achieved. For single-stage cluster designs with equal cluster takes it is well approximated by 1 + (m − 1)ρ, where m is the number of interviews per cluster and ρ is the intraclass correlation — the degree of within-cluster resemblance [1].
Two practical consequences follow. First, the effective sample size is the nominal sample divided by DEFF: a survey of 1,200 households in large clusters can carry the information of far fewer independent interviews. Second, for a fixed budget, more clusters with fewer interviews each almost always beats fewer clusters with many interviews each, because the design effect grows with the cluster take m [1][2]. Evaluations that randomise at cluster level face the same arithmetic in their power calculations, where the intraclass correlation drives the number of clusters required [2].
Weighting is the other half of the bookkeeping: whenever selection probabilities differ — oversampled strata, probability-proportional-to-size cluster selection with fixed takes, household selection within compounds — analysis must apply weights equal to the inverse of each unit’s selection probability, adjusted for non-response. Unweighted analysis of a weighted design quietly estimates a different population than the one sampled [1].
What actually drives sample size
The familiar formula for estimating a proportion — n = z²p(1−p)/e² — answers only the narrowest version of the question. In practice four demands compete, and the largest wins [1][2]:
- Precision — the confidence-interval width the decision can tolerate. Halving the margin of error quadruples the sample.
- Variance — outcomes that vary more need more observations; for proportions, variance peaks at p = 0.5, which is why that value is the conservative default.
- Power for comparisons — detecting a difference (between arms, or between baseline and endline) needs substantially more than estimating a level; the minimum detectable effect, not the margin of error, is the binding constraint in evaluations [2].
- Disaggregation — the report that must state results for women, for youth and for each of four counties needs adequate precision in the smallest such cell, not in the total. This is the demand that most often dominates, and the one most often discovered after fieldwork.
To all of this add the design effect and expected non-response, then round up.
On Krejcie and Morgan. The 1970 table that circulates in countless theses and proposals gives, for a population of a given size, “the” required sample. It is a legitimate shortcut for exactly one case: estimating a population proportion at 95% confidence with a ±5% margin, assuming p = 0.5 and simple random sampling [3]. It embeds no design effect, no power for comparisons and no disaggregation. Citing it for a clustered baseline survey with subgroup reporting requirements is not a conservative simplification; it is the wrong tool.
Non-probability designs, used honestly
Purposive, quota and snowball sampling are not failed probability sampling; they are different instruments for different claims.
- Purposive sampling selects cases for what they can teach — maximum variation, typical cases, extreme cases, critical cases. It is the correct design for most qualitative components (see qualitative data collection) and for case-based methods generally. Its output is understanding of the cases studied and grounded hypotheses about the wider population — not estimates.
- Quota sampling fixes the sample’s composition on a few dimensions (sex, age band, location) and fills the quotas non-randomly. It buys demographic balance, not representativeness: within each quota cell, selection remains uncontrolled.
- Snowball / respondent-driven sampling reaches populations for which no frame can exist — undocumented migrants, sex workers, out-of-school youth. Referral chains introduce their own structure; specialised estimators exist, but for ordinary M&E purposes the honest posture is to report such samples as illustrative of the network reached [1][5].
The failure mode is not using these designs; it is the report that quietly converts “of the 60 purposively selected respondents, 70% said…” into “70% of beneficiaries say…”. A one-sentence scope statement under every figure — what population, what design, what claims — is cheap insurance.
Frames in low-resource settings
Every probability design presumes a frame — a list or map from which units can be drawn with known probability — and the frame is where field reality bites hardest. Census enumeration-area data may be a decade old; settlements grow, split and move; displacement makes yesterday’s list wrong by tomorrow [1][5]. UNHCR’s guidance for displacement settings is candid that frame construction is often the majority of the sampling work: registration databases where they exist and are current, fresh household listing of selected clusters where they do not, and satellite-imagery-assisted segmentation where even listing is infeasible [5]. Three practices travel well beyond humanitarian settings:
- Re-list selected clusters immediately before interviewing rather than trusting old counts; the listing also supplies the stage-two frame.
- Document frame coverage gaps — populations the frame cannot see (institutional, homeless, mobile pastoralist households) are excluded from every estimate, and the report must say so.
- Prefer probability-proportional-to-size selection with current measures of size; stale size measures reintroduce unequal probabilities that the weights must then repair [1].
The sampling annex the report owes its readers
A results claim is only auditable if the sampling is. The annex costs a page and should state: the target population and frame (with its date and known gaps); the stages, stratification and selection method at each stage; sample sizes designed and achieved, with response rates; the weighting procedure; and the design effects actually observed for headline indicators [1]. Evaluations add the power calculation and its assumptions [2]. This is the survey-side counterpart of the instrument-versioning discipline described under questionnaire design — and when a data quality assessment later traces reported figures back to their source, the annex is where verification starts (see data quality assessment).
Checklist before fieldwork
- The claim each headline figure must support is written down, and the design supports it.
- The frame’s date, source and coverage gaps are documented; selected clusters will be re-listed.
- Sample size covers the smallest reporting cell, the design effect and expected non-response — not just the total-sample formula.
- Selection probabilities are recorded at every stage so weights can be computed.
- Interviewer instructions make within-household respondent selection explicit (a common hidden stage).
- The sampling annex is drafted before fieldwork, not reconstructed after it.
Sources
- Designing Household Survey Samples: Practical Guidelines (Studies in Methods, Series F No. 98) — United Nations Statistics Division, 2005.The standard open reference for probability designs, design effects, weighting and frame problems in developing-country surveys.
- Impact Evaluation in Practice, Second Edition — World Bank / Inter-American Development Bank, 2016.Gertler, Martinez, Premand, Rawlings & Vermeersch. Part on sampling and power places sample-size decisions inside evaluation design.
- Determining Sample Size for Research Activities — Educational and Psychological Measurement, 30(3), 1970.Krejcie & Morgan — the source of the ubiquitous sample-size table, with the assumptions the table's users rarely restate.
- Designing Household Survey Questionnaires for Developing Countries: Lessons from 15 Years of the Living Standards Measurement Study — The World Bank / Oxford University Press, 2000.Grosh & Glewwe (eds.). The instrument side of the survey craft; sampling and questionnaire quality fail together.
- Sampling for Household Surveys — UNHCR Assessment and Monitoring Resource Centre, 2024.Guidance written for displacement settings, where sampling frames are weakest; publication date per the resource-centre record (October 2024).