03 · Indicator design and measurement

Baselines, targets and disaggregation

A target is only as meaningful as the baseline under it and the disaggregation behind it. Baselines establish where performance stands before the intervention; targets state the value sought by a date, set by methods that range from trend projection to open negotiation; disaggregation decides — at design time, not analysis time — whose outcomes will be visible separately. All three are design decisions with governance attached, and each has failure modes that surface years later as unexplainable results.

Last updated · Reviewed against 5 cited sources

Baselines: the options, ranked

A baseline is the value of an indicator before the intervention acts on it — the “where are we today” that Kusek and Rist place immediately before any talk of targets [1]. The options are not equivalent, and choosing among them is a cost-credibility trade that should be made per indicator, not per programme:

  1. Dedicated baseline study. Primary data collected for the purpose, with the same instrument and sampling approach the endline will use. Highest credibility, highest cost, and one unforgiving constraint: it must happen before implementation reaches the measured population. A baseline fielded in month nine of a delivery programme is a midline wearing the wrong label. Survey-based baselines stand or fall on their sampling design [4] — treated fully under sampling.
  2. Secondary or administrative baseline. National surveys, routine information systems, programme registers. Cheap and fast; credible exactly insofar as the source’s definition, population and timing match the indicator’s — mismatches documented in the reference sheet, not discovered by an evaluator.
  3. Retrospective (recall) baseline. Respondents reconstruct their pre-programme situation from memory. Sometimes the only option, and legitimately used — with its documented weaknesses stated: recall error grows with time and with the ordinariness of the fact recalled, and recall of the past is coloured by the present, particularly in a programme that has told participants a story about their own improvement.
  4. Rolling baseline. Where participants enter continuously, baseline data is collected from each cohort at entry. This is often the best design for demand-driven programmes — every participant has a true entry measurement — at the price of complicating cross-cohort comparison when context shifts between entries.

The “no baseline” pathology is discovered, predictably, at evaluation time. The honest remedies are reconstruction — from administrative records, from recall with its biases declared, or from a comparison population measured late — and the honest reporting rule is that a reconstructed baseline is labelled as one, with its method and direction of likely error, wherever the resulting change figure appears. Results-measurement audit regimes such as the DCED Standard exist in large part because undocumented baselines and projections could not otherwise be told apart from documented ones [2].

Target-setting: four methods, one honest admission

A target is the value of the indicator sought by a date [3]. Four methods produce one, and they differ in what the number means:

Target-setting methods compared
MethodThe target meansStrengthCharacteristic failure
Trend projectionWhat the past implies, plus programme effectAnchored in real dataAssumes history continues; needs several periods of comparable data
Benchmark transferWhat comparable programmes achievedExternal disciplineThe benchmark's context travels poorly; 'comparable' is doing unexamined work
Capacity-based build-upWhat resources and delivery capacity can produceForces implementation realismBottom-up sums drift toward what is comfortable
NegotiationWhat the parties agreed to promiseOwns the political realityUntethered from evidence, it produces ambition theatre — or sandbagging
Table 1. Each method answers a different question; the strongest targets are negotiated within a corridor bounded by the other three.

The honest admission, stated plainly because most guidance mumbles it: most real targets are negotiated. Proposal targets are set against funder expectations; government targets are set in performance-contract negotiations; and everyone in the room knows it. The response is not to pretend otherwise but to constrain the negotiation: bring a trend projection, a benchmark and a capacity estimate into the room, and require the negotiated number to land inside the corridor they bound — with the rationale recorded. Kusek and Rist’s formulation — target performance equals baseline plus desired improvement, given resources, capacity and timeframe — is precisely such a constraint: no baseline, no defensible target [1][3].

Two arithmetic disciplines follow from indicator behaviour. A cumulative target (“12,000 farmers by year 5”) needs period milestones that sum to it; a period target (“80% coverage in each year”) must never be summed. And slow-onset outcomes need a milestone profile that matches how change actually arrives — flat, then rising — because a straight line from baseline to endline target manufactures year-one “underperformance” that is nothing but geometry.

Target corridor chart with three trajectories, plus a disaggregation strip showing one lagging group

Upper panel: a chart with years zero to five on the horizontal axis and the indicator value on the vertical axis. A baseline dot at year zero and a target dot at year five are joined by three trajectories: a straight line, an S-curve that is flat early and steep in the middle, and a stepped milestone profile. A shaded band around them marks the credible corridor. Lower panel: four bars disaggregate the year-three value for women, men, youth and persons with disabilities; the persons-with-disabilities bar is visibly lower and flagged as lagging despite the aggregate being on track.

indicator valueyr 0yr 1yr 2yr 3yr 4yr 5credible corridorbaselineyear-5 targetstraight-lineS-curve (slow-onset outcome)milestone stepsthe year-3 value, disaggregatedwomenmenyouthPWDlagging — invisiblein the aggregate
Figure 1. Three trajectory shapes to the same year-5 target, and the credible corridor around them. Below: the year-3 value disaggregated — the aggregate is on track while one group is not.Trajectory logic follows Kusek & Rist (2004); disaggregation floor per the leave-no-one-behind agenda.

Disaggregation: deciding whose outcomes are visible

Disaggregation is the difference between “the programme is on track” and “the programme is on track for everyone it claims to serve”. The figure’s lower panel is the standard discovery: an aggregate comfortably inside the corridor, while one group’s value sits far below it — a gap the aggregate arithmetic is structurally incapable of showing.

Three design rules govern it:

  • The floor is sex, age, disability and geography — the leave-no-one-behind minimum that international results practice now expects, extended by whatever dimensions the programme’s own equity claims imply (wealth, displacement status, language) [3]. A programme whose proposal promises to reach “the most marginalised” and whose indicators disaggregate by nothing has made a claim it cannot ever evidence.
  • Dimensions are decided at design time, because they are properties of forms, registers and samples, not of analysis. A dataset that never recorded disability status cannot be disaggregated by disability later at any price; a sample sized for a county estimate will not support sub-county estimates just because someone asks — precision at each disaggregation level is a sampling-design decision made up front [4]. Each indicator’s dimensions and exact categories belong in its reference sheet.
  • Small cells are a privacy boundary. Disaggregation multiplies categories until cells contain a handful of identifiable people; publishing “2 women with disabilities in ward X reported abuse” is disclosure, not transparency. Suppression thresholds and their legal footing are treated under data protection.

Targets should be disaggregated wherever the disaggregated gap is the point: a coverage target with a “no group below X%” floor commits the programme to the distribution, not just the average — and changes what “on track” means in every review meeting.

Revising targets: governance, not guilt

Baselines turn out wrong; contexts collapse; programmes are cut or scaled. A target that has become fiction serves no one — reporting against it produces either demoralising failure or quiet redefinition, and the second is worse. The defensible path is a governed revision: proposed with evidence, approved by whoever owns the results framework, logged with its date and rationale, and reported openly — every report thereafter showing the revised target as revised, not as if it had always been so [1][3]. The change-log discipline is the same one that governs indicator definitions in the reference sheet; within the broader management system, revision authority and cadence are part of results-based management. What separates governance from goalpost-moving is entirely procedural: the paper trail, the approval, and the willingness to show both numbers side by side.

Sources

  1. Ten Steps to a Results-Based Monitoring and Evaluation System: A Handbook for Development Practitioners — The World Bank, 2004.Kusek & Rist — steps 4 and 5 are the canonical treatment of baselines and target-setting in results-based M&E.
  2. The DCED Standard for Results Measurement — Donor Committee for Enterprise Development (DCED), n.d..The results-measurement standard whose audit discipline includes defensible baselines and projections. Accessed 18 August 2026.
  3. Results-Based Management Handbook: Harmonizing RBM Concepts and Approaches for Improved Development Results at Country Level — United Nations Development Group (UNDG), 2011.The UN system's harmonised guidance on results, indicators, baselines and targets at country level.
  4. Designing Household Survey Samples: Practical Guidelines — United Nations Statistics Division, Studies in Methods, Series F No. 98, 2005.The sampling reference that governs whether a baseline survey — and its disaggregated estimates — can bear the weight put on them.
  5. DIME Wiki: practical impact-evaluation and measurement resource — World Bank Development Impact (DIME), n.d..Continuously updated practical guidance on baseline data collection and measurement in field research.