04 · Data quality, sampling and collection
Data quality dimensions in M&E
Data quality in M&E is not one property but several: validity, reliability, integrity, precision and timeliness in the USAID tradition; completeness, timeliness and consistency in the WHO framing. Each dimension fails in a characteristic way, and each failure has a characteristic remedy — almost all of them applied at design time, not audit time.
Last updated · Reviewed against 4 cited sources
Two dimension families, one purpose
Ask three agencies what “good data” means and you will get three lists. USAID’s performance-monitoring practice assesses five standards — validity, integrity, precision, reliability and timeliness — and its indicator reference sheet template requires known weaknesses against them to be documented indicator by indicator [3]. WHO’s Data Quality Review framework, built for routine facility data, measures completeness, timeliness, internal consistency and external consistency [1]. The MEASURE Evaluation tool family, used across PEPFAR and Global Fund programmes, operationalises a near-identical set as verification and systems-assessment questions [2].
The overlap is large but not total, and the differences should be stated rather than smoothed over. WHO’s completeness (are all expected reports present, with their data elements filled?) has no single named counterpart in the USAID five; USAID’s integrity (is the data pipeline protected from manipulation?) is only implicit in WHO’s consistency checks. A programme reporting to more than one funder should map its quality checks to each framework explicitly, not assume one list covers the other.
The five working dimensions
Validity — does the data measure what the indicator claims to measure? The classic failure is a proxy quietly standing in for the construct: “people trained” reported against an indicator defined as “people who completed training”, with walk-outs and half-attendees counted in full. Validity failures are invisible in the numbers themselves; they live in the gap between the indicator definition and the collection instrument.
Reliability — would the same process, repeated, produce the same number? The classic failure is definition drift: two field offices counting “youth” with different age bands, or the same office changing its interpretation when staff turn over. A reliable process is documented tightly enough that two people recount the same figure from the same records.
Integrity — is the data protected from error and manipulation on its way through the pipeline? Failures range from transcription slips at each aggregation hop to deliberate inflation where targets carry incentives. Access controls, change logs and separation of recording from reporting are the standard defences.
Precision — is the data fine-grained enough for the decision it informs? A provincial average is imprecise for a district-level allocation decision; a survey estimate with a ±12-point confidence interval cannot support a claim of a 5-point improvement. Precision failures are usually design failures: the sample or the disaggregation was never sized for the question.
Timeliness — does the data arrive while the decision is still open? Data that is accurate but three quarters late has the same decision value as no data. WHO’s framework measures this directly as the share of reports submitted by deadline [1].
WHO’s additional lens, completeness, deserves separate attention in routine systems: a reporting rate of 60% does not mean the true value is the reported value scaled up, because non-reporting sites are rarely a random sample of all sites [1].
The failure-mode catalogue
Across audits and assessments the same handful of failures recurs, and almost none of them are arithmetic:
- Double counting — the same person or event counted at two sites, in two periods, or under two indicators; endemic wherever beneficiaries move between service points and no unique identifier exists.
- Phantom results — reported figures with no corresponding source records; sometimes fraud, more often a reporting form filled from memory after the register went missing.
- Unit confusion — households counted as people, cumulative totals reported as period totals, currency amounts mixing units. Unit errors survive aggregation and can dominate a national figure.
- Aggregation across incompatible definitions — summing “trained” from partners who each defined it differently; the total is arithmetic on apples and oranges.
- Transcription loss — each manual hop from register to tally to report to system introduces error; long paper chains can accumulate double-digit discrepancy before anyone computes anything [2].
Most of these are definitional or procedural. That is good news: they are preventable, cheaply, if prevention starts early enough.
Quality is won at design time
The moment of greatest leverage over data quality is before the first record is collected. Precise indicator definitions in a governed reference sheet, forms whose fields match those definitions one-to-one, entry validation that rejects impossible values, and unique identifiers that make double counting detectable — these controls cost little at design time and are nearly impossible to retrofit at audit time [3][4]. A data quality assessment can tell you how bad the data is; only design can make it good. The assessment side of the discipline — verification, system appraisal, action planning — is covered on the data quality assessment page.
Mapping dimension to remedy
| Dimension | Characteristic failure | Primary control |
|---|---|---|
| Validity | Instrument measures something other than the indicator definition | Definition-to-form mapping review before rollout |
| Reliability | Definition drift between sites, staff or periods | Governed reference sheet; enumerator training; recount checks |
| Integrity | Manipulation or transcription error in the pipeline | Access controls, audit trail, fewer manual hops |
| Precision | Data too coarse for the decision | Sample and disaggregation sized to the decision at design |
| Timeliness | Data arrives after the decision | Reporting calendar tied to decision calendar; completeness tracking |
| Completeness | Missing reports treated as zero or ignored | Reporting-rate monitoring; follow-up protocol for silent sites |
Checklist for a data quality plan
- Every indicator has one governed definition, and the collection form matches it field for field.
- The pipeline from source record to published figure is mapped, and each manual hop is either automated or verified.
- Entry validation rejects impossible values, and unique identifiers make duplicates detectable.
- The reporting calendar is derived from the decision calendar, not the other way round.
- Known weaknesses are documented per indicator — honesty in the reference sheet prevents findings in the audit [3].
Sources
- Data Quality Review (DQR): a toolkit for facility data quality assessment. Module 1: Framework and metrics — World Health Organization, 2017.The DQR dimension set — completeness, timeliness, internal consistency, external consistency — with standard metrics for each.
- Routine Data Quality Assessment (RDQA) Tool: User Manual — MEASURE Evaluation (USAID), 2017.Operationalises dimension checks as verification and systems-assessment questions.
- Recommended Performance Indicator Reference Sheet (PIRS): Guidance & Template — USAID, 2017.The reference-sheet template that requires known data limitations to be documented against USAID's five data quality standards.
- Ten Steps to a Results-Based Monitoring and Evaluation System — The World Bank, 2004.Kusek & Rist. Chapter on data collection places quality choices at system-design time.