03 · Indicator design and measurement
Proxy indicators and composite indices
A proxy indicator measures something observable that stands in for a result that is unobservable, too slow or too costly to measure directly; a composite index bundles several indicators into one number. Both are legitimate and both are dangerous in the same way: the construction choices disappear into a value that looks like a measurement. The defence is the same for both — make the validity chain and the construction choices explicit, and publish them.
Last updated · Reviewed against 4 cited sources
When to proxy
The OECD glossary’s definition of an indicator — a variable providing a simple and reliable means to measure achievement or change [2] — quietly assumes the achievement is measurable at reasonable cost within a useful timeframe. Three situations break that assumption, and they are the legitimate occasions for a proxy [3]:
- The construct is unobservable. Women’s household decision-making power, social cohesion, institutional capacity — no register records them. Something observable must stand in.
- Direct measurement is unaffordable. Household consumption is the direct poverty measure; a full consumption module costs survey-hours most programmes do not have. Asset ownership and dwelling characteristics proxy it at a fraction of the cost — the logic behind proxy-means testing.
- The direct measure is too slow for the decision. Population-level mortality moves years after the programme acts and reaches the M&E system later still; service-contact and coverage measures move within the decision cycle.
A proxy is not a lesser indicator; it is a different claim. A direct indicator claims “this is the result”; a proxy claims “this observable thing moves with the result”. The second claim needs support the first does not, and the support is the validity chain: a written argument, in the indicator’s reference sheet, from proxy to construct.
Choosing a defensible proxy
Three elements make the validity chain inspectable rather than asserted:
- Association evidence. Why believe the proxy tracks the construct — published research, national survey correlations, the programme’s own baseline? “Widely used elsewhere” is citation, not evidence, if the elsewhere differs in the ways that matter.
- Direction-of-error reasoning. Every proxy fails somewhere; the design question is how it misleads when it misleads. School attendance as a proxy for learning fails upward — attendance can rise while learning stalls, so it flatters. Asset indices fail slowly — assets accumulate and persist, so they lag both improvement and shocks. Knowing the failure direction tells you which apparent results to distrust: be most sceptical of good news delivered by a proxy that fails toward good news.
- Triangulating companions. Two or three cheap proxies with different failure directions constrain each other; agreement across them is informative in a way no single proxy can be. This is portfolio design at the indicator level, and it is the practical alternative to over-trusting one clever measure [3][4].
One boundary worth policing: a proxy stands in for a result. An activity count used as a stand-in for an outcome (“trainings held” for “practice changed”) is not a proxy — it is a level confusion, treated under Types and levels.
Composite indices: the OECD/JRC discipline
A composite index is proxying at scale: many indicators compressed into one number meant to stand for a multidimensional construct — vulnerability, empowerment, capacity, readiness. The OECD/JRC Handbook on Constructing Composite Indicators is the standard of record, and its core contribution is a construction sequence in which every step is an explicit, revisable decision [1]: define the theoretical framework (what the construct is and what belongs in it — the step most indices skip and all bad indices skipped); select data against that framework; impute missing values deliberately rather than silently; run multivariate analysis to understand the structure and overlap of the components; normalise; weight and aggregate; and subject the result to uncertainty and sensitivity analysis before presenting it.
Three of those steps hide the decisions that most change the answer:
Normalisation. Components arrive on incommensurable scales — shillings, percentages, counts, minutes — and must be placed on a common one before any combination is meaningful. The handbook’s menu includes min–max rescaling, standardisation to z-scores, and distance to a reference point, and the choice is consequential: min–max is sensitive to the extremes in the data, z-scores let volatile components dominate movement, and both need “more is better” orientations agreed first (travel time must be inverted before it joins the index) [1]. Skipping this step — summing raw mixed scales — is not a simplification; it is an implicit weighting by unit of measure.
Weighting. Equal weights assert every component matters equally — a value judgement, not a neutral default. Expert or participatory weights (budget allocation, pairwise comparison) assert the panel’s priorities. Statistical weights from principal components or factor analysis assert that the variance structure of the data reflects importance — which rewards correlated components and can bury the dimension policymakers care about most. The handbook’s position, and this site’s: there is no correct weighting, only an explicit one, published with the index and defended in its documentation [1].
Aggregation. Linear (weighted-sum) aggregation is fully compensatory: a collapse in one dimension is offset by strength in another, and the index cannot see the collapse. Geometric aggregation punishes imbalance and limits compensability. The choice encodes what the construct means — whether “moderately good at everything” and “excellent at most things, failing at one” should score the same — and belongs to the framework step, not to whichever formula the spreadsheet made easy [1].
Communicating an index honestly
An index earns trust by exposing its insides, not by the confidence of its headline number:
- Always publish sub-scores. The component values are where diagnosis lives; a composite that ships without them converts a measurement instrument into a slogan. This is equally a dashboard-design rule — see dashboards and data visualisation.
- Report sensitivity as a band, not a caveat. The handbook’s uncertainty and sensitivity analysis asks how the index value or ranking moves under alternative defensible choices of normalisation, weights and aggregation [1]. If a county’s ranking swings by ten places across reasonable specifications, that instability is the finding, and presenting the central ranking alone misrepresents the measurement.
- Version the methodology. An index whose weights or components change between rounds without a logged, dated methodology note breaks its own time series — the same governance rule that applies to any indicator definition in a reference sheet.
Common failures
The same handful of failures account for most bad proxies and bad indices in programme practice:
- Double counting. Two highly correlated components — say, literacy rate and school completion — enter as if independent, and their shared dimension is silently weighted double. The handbook’s multivariate-analysis step exists to catch exactly this before the weights are set [1].
- Scale mixing. Raw values summed across units; or one unbounded component (income) swamping bounded ones after a careless normalisation.
- Proxy drift. A proxy adopted for one context is carried to another where the association no longer holds — asset indices calibrated on rural consumption patterns applied to urban populations. The validity chain is context-bound and must be re-argued, not transplanted.
- Index worship. Once a composite becomes a management target, its components stop being questioned and the organisation optimises the number — including by the cheapest available component, which is rarely the one that matters. The defence is the publication discipline above: sub-scores in every report, sensitivity bands on every ranking, and a standing willingness to say the index moved for an uninteresting reason. The data-quality dimensions the components inherit — and drag into the composite — are treated under data quality dimensions.
Sources
- Handbook on Constructing Composite Indicators: Methodology and User Guide — OECD Publishing & European Commission Joint Research Centre, 2008.The standard of record for composite-index construction: the step sequence, the normalisation and weighting menus, and the sensitivity-analysis obligation.
- Glossary of Key Terms in Evaluation and Results-Based Management for Sustainable Development (Second Edition) — OECD Publishing, 2023.The definition of record for indicator — the baseline against which proxy validity is argued.
- Ten Steps to a Results-Based Monitoring and Evaluation System: A Handbook for Development Practitioners — The World Bank, 2004.Kusek & Rist — proxy indicators as a practical answer where direct measurement fails the cost test, with the accompanying cautions.
- DIME Wiki: practical impact-evaluation and measurement resource — World Bank Development Impact (DIME), n.d..Continuously updated guidance on measurement construction and validation in field research.