03 · Indicator design and measurement
Indicator quality criteria: SMART, CREAM, SPICED
SMART, CREAM and SPICED are competing answers to the same question: how do you tell a usable indicator from a plausible-sounding one? They embody genuinely different philosophies of measurement — SMART and CREAM privilege objectivity and verifiability, SPICED privileges the perspective of the people whose change is being measured — and the disagreement between them is methodological, not cosmetic.
Last updated · Reviewed against 4 cited sources
Why criteria exist at all
An indicator fails long before its first data point. It fails when two field offices read the same words and count different things; when the data source it presumes turns out not to exist; when it measures activity because the outcome was hard; when it is technically flawless and nobody uses it for any decision. All of these are failures of definition and design, and they are what the quality-criteria acronyms are for: structured questions asked of an indicator on paper, before the collection machinery is built [1]. The precise wording of the indicator being tested is itself anchored by the OECD glossary definition — a variable providing a simple and reliable means to measure achievement or change [4].
Three frameworks dominate practice. They are usually presented as interchangeable checklists. They are not: each one catches failures the others let through, and SPICED disagrees with the other two about what a good indicator fundamentally is.
SMART — and its limits
SMART — Specific, Measurable, Achievable, Relevant, Time-bound — entered M&E from corporate performance management and is now the default test in donor templates and UN programme guidance [2]. Its virtues are real: it forces precision (Specific), forces the question “measured how, from what source?” (Measurable), and forces a time horizon (Time-bound).
Its limits are equally real, and worth stating plainly rather than footnoting:
- Measurability bias. An indicator set built to maximise the M is pulled toward the easily counted — training headcounts, items distributed — and away from the changes that matter but resist counting. The tool shapes the portfolio.
- Two letters test the wrong object. Achievable and Time-bound are properties of a target, not an indicator. “Proportion of children assessed within 24 hours” is neither achievable nor unachievable — it is a measurement. Applying SMART without noticing this pushes teams into writing targets into their indicators, the grammar error unpacked on Types and levels.
- Unstable letters. A stands for Achievable, Attainable or Attributable depending on the manual; R for Relevant or Realistic. A test whose content varies by publisher is a weak audit instrument.
- Silent on cost and verification. SMART never asks what the indicator costs to collect, or whether an outsider could independently verify the value. Those two omissions are exactly what CREAM adds.
CREAM — the World Bank’s test
Kusek and Rist’s Ten Steps handbook proposes that good performance indicators are CREAM: Clear (precise and unambiguous), Relevant (appropriate to the subject at hand), Economic (available at reasonable cost), Adequate (able to provide a sufficient basis to assess performance), and Monitorable (amenable to independent validation) — a formulation they credit to Schiavo-Campo (1999) [1].
Two of these earn their place immediately. Economic makes cost a first-class design criterion: an indicator that consumes the monitoring budget crowds out everything else the system should measure. Monitorable anticipates the auditor: if an independent reviewer could not re-derive the value from records, the indicator will not survive a data quality assessment — the standards-layer point; the audit-survival practice itself is the monival.com companion piece. Adequate does portfolio-level work the other frameworks skip entirely: it asks whether the indicator, together with its companions, gives enough basis to judge performance — a property of the set, not the single line.
SPICED — the participatory dissent
SPICED — Subjective, Participatory, Interpreted (and communicable), Cross-checked (and compared), Empowering, Diverse (and disaggregated) — emerged from NGO impact-assessment practice in the late 1990s and is documented for practitioners in INTRAC’s M&E Universe papers [3].
It is routinely misread as “SMART for qualitative indicators”. Read its first letter again: Subjective. SPICED asserts that informants’ own judgements are not a contaminant to be engineered out but a source of insight with special value precisely because of their position — and that indicators should be developed with, and interpreted by, the people whose change is being measured. That is a direct challenge to the SMART/CREAM premise that a good indicator minimises the observer’s judgement.
This is a genuine methodological divide, and this site will not flatten it. The objectivist position holds that accountability requires measures an outsider can verify without trusting anyone’s interpretation — Monitorable, in CREAM’s terms [1]. The participatory position holds that pre-specified, externally verifiable indicators systematically miss the changes that matter to communities, and that cross-checking (SPICED’s own C) supplies rigour by triangulation rather than by standardisation [3]. Both positions are coherent; they trade different risks. Mature practice runs them in portfolio: SMART/CREAM-tested indicators for the results the organisation must defend to funders, SPICED-style locally interpreted indicators where the change is complex, contested or belongs to the community — with each kind labelled honestly as what it is.
A structured indicator review protocol
Acronyms degrade into checkbox theatre. A review that changes indicator sets asks four concrete questions of every candidate, with evidence required for each answer:
- Definition test. Give the indicator’s written definition to two people who did not draft it, with the same three scenario cases. Do they classify the cases identically? Divergence here is the single strongest predictor of downstream data incomparability. The fix is a tighter definition in the reference sheet, not training.
- Data-source test. Name the specific register, form, dataset or instrument the value will come from — not a genre (“facility records”) but an artefact. If the source does not yet exist, the indicator’s real cost includes building it.
- Cost test. Estimate collection cost per reporting period, in staff-days and money — CREAM’s E made explicit [1]. Compare it against what the indicator’s answer is worth to any actual decision.
- Use test. Name the decision, report or review meeting that consumes this value, and the person who reads it. An indicator with no consumer is a standing tax on field staff. Cut it.
Fewer, better
Kusek and Rist’s advice on set size has aged well: take the time to arrive at the minimum number of indicators that answer the performance question, because every indicator kept is a permanent claim on collection, cleaning, storage and attention [1]. Overloaded indicator sets are not merely wasteful — they degrade quality across the whole set, because scarce verification effort spreads thinner per indicator. The disciplined follow-through — writing each surviving indicator’s full definition, source, method and limitations into a governed document — is the subject of the next page on reference sheets, and the data-quality dimensions those definitions protect are treated under data quality.
Sources
- Ten Steps to a Results-Based Monitoring and Evaluation System: A Handbook for Development Practitioners — The World Bank, 2004.Kusek & Rist — the source of the CREAM criteria (which they credit to Schiavo-Campo, 1999) and of the advice to keep indicator sets small.
- Results-Based Management Handbook: Harmonizing RBM Concepts and Approaches for Improved Development Results at Country Level — United Nations Development Group (UNDG), 2011.The UN system's harmonised RBM guidance, which carries the SMART tradition into UN programme practice.
- The M&E Universe — practitioner papers on monitoring and evaluation — INTRAC, n.d..Continuously maintained library of short practitioner papers, including the indicator papers that document the SPICED approach for participatory measurement.
- Glossary of Key Terms in Evaluation and Results-Based Management for Sustainable Development (Second Edition) — OECD Publishing, 2023.The definition of record for indicator — the baseline any quality criterion is testing against.