Glossary
Working definitions of the terms this reference uses, in plain language. Where a term has a page of its own, the definition links to it.
- Accountability
- The obligation to demonstrate that work has been conducted according to agreed rules and standards, and to report fairly and accurately on performance and the use of resources.
- Accountability to affected populations (AAP)
- The commitment by humanitarian and development actors to use power responsibly: sharing information with the people they serve, involving them in decisions, and providing channels to complain and receive a response. Read the full page →
- Activity
- An action taken or work performed through which inputs are mobilised to produce outputs — training sessions delivered, boreholes drilled, guidance drafted.
- Adaptive management
- Deliberate, structured adjustment of plans and delivery in response to monitoring evidence and changing context, with the reasoning for each adjustment documented. Read the full page →
- After-action review (AAR)
- A structured, facilitated discussion held after an event or implementation period that compares what was intended with what actually happened and records lessons while memory is fresh.
- Attribution
- The claim that an observed change was caused by the intervention rather than by other factors. Establishing attribution requires a counterfactual; asserting it without one is the field's most common overreach.
- Attrition
- The loss of study participants between measurement rounds. When attrition differs between treatment and comparison groups, the surviving samples are no longer comparable and impact estimates are biased.
- Balanced scorecard
- A strategy performance framework that tracks objectives and measures across financial, customer, internal-process and learning-and-growth perspectives rather than finances alone. Read the full page →
- Baseline
- The measured situation before an intervention begins, against which later change is assessed. Read the full page →
- Benchmark
- A reference point against which performance is compared — past performance, the performance of peers, or an external standard.
- Beneficiaries
- The individuals, groups or organisations whose situation an intervention is intended to improve, whether or not they are targeted directly.
- Bias
- Systematic error that pushes findings away from the truth in a consistent direction. It can enter through design, sampling, measurement, non-response or analysis, and unlike random error it does not shrink as samples grow.
- Cluster-randomised trial
- A randomised controlled trial that assigns intact groups — schools, health facilities, villages — rather than individuals. Statistical power depends on the number of clusters and the intraclass correlation, not just the number of people. Read the full page →
- Coherence
- One of the OECD-DAC evaluation criteria: how well the intervention is compatible with other interventions in the country, sector or institution, including its own organisation's wider policies. Read the full page →
- Comparison group
- A group not receiving the intervention, used to approximate the counterfactual. When membership is assigned by lottery it is called a control group; otherwise its credibility depends on how it was constructed.
- Composite indicator
- A single index built by normalising, weighting and aggregating several indicators. The weights are value judgements and should always be published alongside the index. Read the full page →
- Contribution analysis
- Mayne's six-step approach to credible causal claims without a counterfactual: articulate the programme theory, assemble evidence for and against it, address rival explanations, and refine a contribution story until it is plausible and robust. Read the full page →
- Core Humanitarian Standard (CHS)
- The standard on quality and accountability for organisations assisting people affected by crisis, structured as commitments made to affected people. The current edition was issued in 2024.
- Counterfactual
- What would have happened to the same population over the same period without the intervention. Impact is the difference between observed outcomes and the counterfactual. Read the full page →
- CREAM criteria
- An indicator quality test from Kusek and Rist's results-based management handbook: a good indicator is Clear, Relevant, Economic, Adequate and Monitorable. Read the full page →
- Data protection
- The legal and organisational safeguards governing the collection, storage, sharing and disposal of personal data — in Kenya under the Data Protection Act 2019, and in the EU under the GDPR. Read the full page →
- Data quality assessment (DQA)
- A structured exercise that verifies reported data against source records and appraises the data-management system that produced them. Read the full page →
- Developmental evaluation
- Patton's approach for innovations developing under complexity: the evaluator is embedded with the team and feeds back evidence in real time to shape what the initiative becomes, rather than judging a fixed model. Read the full page →
- DHIS2
- The open-source health information platform developed by the HISP programme at the University of Oslo, widely used by ministries of health to manage routine facility data. Read the full page →
- Difference-in-differences (DiD)
- A quasi-experimental design that estimates impact as the change in outcomes over time in the treated group minus the change in a comparison group, netting out fixed differences and common trends. Read the full page →
- Disaggregation
- Breaking indicator values down by sex, age, disability, geography or other dimensions so that averages cannot hide who is being left out. The dimensions must be decided at design time, not at analysis. Read the full page →
- Earned value management (EVM)
- A project-control technique that integrates scope, schedule and cost by comparing the budgeted value of work actually performed against planned value and actual cost. It is a delivery-analytics tool, not a value-for-money framework. Read the full page →
- Effectiveness
- One of the OECD-DAC evaluation criteria: the extent to which the intervention achieved, or is expected to achieve, its objectives and results, including any differential results across groups. Read the full page →
- Efficiency
- One of the OECD-DAC evaluation criteria: the extent to which the intervention delivers, or is likely to deliver, results in an economic and timely way. Read the full page →
- Empowerment evaluation
- Fetterman's approach in which the people running a programme conduct their own evaluation with an evaluator acting as coach and critical friend, aiming at self-determination as well as findings. Its rigour is a live debate in the field. Read the full page →
- Endline
- The measurement taken at or after the end of an intervention, compared against the baseline to describe change over the implementation period.
- Evaluability
- The extent to which an intervention can be evaluated reliably and credibly: a clear results logic, available data, and feasible methods. An evaluability assessment checks these before an evaluation is commissioned.
- Evaluation
- The systematic and objective assessment of an ongoing or completed intervention — its design, implementation and results — to determine relevance, effectiveness, efficiency, impact and sustainability.
- External validity
- The extent to which findings from a study hold beyond the setting, population and period in which they were produced — the question of whether a result travels.
- Feedback and complaints mechanism
- A safe, accessible, known channel through which affected people can give feedback or lodge complaints — including sensitive ones — and receive a response, with the loop closed visibly. Read the full page →
- Feminist evaluation
- Evaluation that treats gendered power relations as central to what is examined and how: knowledge is situated, methods attend to whose voices count, and findings are oriented to action on inequity. Read the full page →
- Formative evaluation
- Evaluation conducted during implementation to improve design and delivery, as distinct from summative evaluation, which judges a completed or maturing intervention.
- Gender-responsive evaluation
- Evaluation that assesses how an intervention affected gender equality, using gender analysis, disaggregated evidence and inclusive process throughout — beyond merely counting women and men. Read the full page →
- Goal
- The higher-order objective to which an intervention is intended to contribute, usually shared with other actors and not attributable to any single programme.
- Hawthorne effect
- A change in behaviour that occurs because people know they are being observed or studied, which can masquerade as a programme effect.
- Impact
- The change in outcomes attributable to an intervention — positive or negative, direct or indirect, intended or not.
- Impact evaluation
- An evaluation that estimates the changes attributable to an intervention by constructing a counterfactual, using experimental or quasi-experimental designs, or a theory-based causal analysis where those are infeasible.
- Indicator
- A quantitative or qualitative variable that provides a simple, reliable signal of change or performance against an intended result. Read the full page →
- Informed consent
- A participant's voluntary agreement to take part in data collection, given after genuinely understanding what participation involves, how data will be used, and the right to decline or withdraw without penalty. Read the full page →
- Inputs
- The financial, human and material resources used for an intervention — the first link of the results chain.
- Intention-to-treat (ITT)
- Analysing trial outcomes by original random assignment regardless of whether participants actually took up the intervention. It preserves the comparability randomisation created and estimates the effect of offering the programme. Read the full page →
- Internal validity
- The degree of confidence that an observed effect is attributable to the intervention rather than to confounding, selection, measurement error or chance within the study itself.
- Interrupted time series (ITS)
- A design for population-level interventions with a clear start date and a long outcome series: segmented regression tests whether the level or slope of the series changed at the interruption. Read the full page →
- Key performance indicator (KPI)
- An indicator selected as one of the few measures management actually steers by. The term signals prioritisation, not a different kind of indicator. Read the full page →
- Local average treatment effect (LATE)
- In a trial with partial compliance, the effect of the intervention on those induced to take it up by their assignment (the compliers), recovered by instrumenting take-up with random assignment. Read the full page →
- Logical framework (logframe)
- The matrix format that summarises an intervention's results chain with indicators, means of verification and assumptions at each level. Read the full page →
- Made in Africa Evaluation (MAE)
- The agenda, synthesised for AfrEA by Bagele Chilisa in 2015, holding that evaluation in Africa should be rooted in African contexts, philosophies and knowledge systems rather than imported wholesale — the intellectual foundation of the 2021 African Evaluation Principles. Read the full page →
- Management response
- The commissioning organisation's formal, published reply to an evaluation: acceptance or rejection of each recommendation with reasons, and the actions committed, tracked to completion. Read the full page →
- Matching
- Constructing a comparison group by pairing treated units with untreated units that have similar observed characteristics. Its central assumption — that no unobserved differences matter — cannot be tested. Read the full page →
- MEAL
- Monitoring, Evaluation, Accountability and Learning — the integrated discipline of tracking performance, judging merit, answering to stakeholders and adapting from evidence.
- Meta-evaluation
- The systematic evaluation of an evaluation, assessing it against standards for process and product. The JCSEE Program Evaluation Standards make a 'metaevaluative perspective' an explicit expectation. Read the full page →
- Mid-term review
- A structured review at the midpoint of implementation that checks progress, tests whether the design's assumptions still hold, and recommends course corrections while there is still time to act on them. Read the full page →
- Milestone
- A scheduled intermediate marker of progress toward a target, used to judge whether delivery is on track between baseline and endline.
- Mixed methods
- The deliberate integration of quantitative and qualitative approaches in one study so that each answers the questions it is suited to and the findings can be triangulated. Read the full page →
- Monitoring
- The continuous, routine collection and review of data on specified indicators to track progress against plans and budgets while implementation is under way.
- Most significant change (MSC)
- The Davies and Dart technique of collecting stories of change from the field and selecting the most significant through structured panel deliberation, with the reasons for each selection recorded. Read the full page →
- National M&E system
- The government-wide arrangements — policy, institutions, data systems and reporting cycles — through which a state monitors and evaluates public programmes, such as Kenya's NIMES. Read the full page →
- Objectives and key results (OKRs)
- A goal-setting method pairing a qualitative objective with a handful of measurable key results, reviewed on short cycles. Related to, but looser than, indicator-based performance frameworks. Read the full page →
- OECD-DAC evaluation criteria
- The six lenses through which the OECD-DAC recommends judging an intervention: relevance, coherence, effectiveness, efficiency, impact and sustainability. Distinct from the DAC Quality Standards, which govern the evaluation process itself. Read the full page →
- Outcome
- The short- and medium-term effects of an intervention's outputs — typically changes in behaviour, practice, capability or condition among the people or institutions reached.
- Outcome harvesting
- An approach that collects evidence of what changed and works backwards to how the intervention contributed, suited to complex settings where outcomes were not fully predefined. Read the full page →
- Outcome mapping
- IDRC's methodology that frames results as behaviour changes in the boundary partners a programme influences but does not control, monitored through graduated progress markers. Read the full page →
- Output
- The products, goods and services delivered by an intervention — within the programme's control, unlike the outcomes they are meant to enable.
- Parallel trends assumption
- The identifying assumption of difference-in-differences: absent the intervention, treated and comparison groups would have followed the same trend over time. Read the full page →
- Participatory evaluation
- Evaluation in which stakeholders — including intended beneficiaries — share control over questions, data collection, analysis or use, rather than serving only as data sources. Read the full page →
- Performance contracting
- A management regime in which public institutions sign agreements committing to measurable annual targets and are scored against them — in Kenya, a statutory regime run through the Public Service Performance Management Unit. Read the full page →
- Performance indicator reference sheet (PIRS)
- The governance document behind an indicator: its precise definition, unit, disaggregation, data source, collection method, frequency, responsible party and known limitations, with a change log. Read the full page →
- Process tracing
- A within-case method that tests a theorised causal mechanism against evidence using four probative tests — straw-in-the-wind, hoop, smoking gun and doubly decisive. Read the full page →
- Programme theory
- An explicit account of how an intervention's activities are expected to produce its intended results, including the assumptions that must hold. Theory-based evaluation methods test it against evidence. Read the full page →
- Propensity score
- The estimated probability that a unit receives treatment given its observed characteristics. Rosenbaum and Rubin showed that matching or weighting on this single score can balance many covariates at once. Read the full page →
- Proxy indicator
- An indirect measure used when the outcome of interest is unobservable, too costly or too slow to measure directly. Its value depends on the documented strength of the link between proxy and construct. Read the full page →
- Qualitative comparative analysis (QCA)
- A set-theoretic method for medium-N comparison that identifies combinations of conditions (causal recipes) associated with an outcome, using truth tables and measures of consistency and coverage. Read the full page →
- Qualitative data
- Non-numeric evidence — interviews, group discussions, observation, documents — analysed for meaning, mechanism and context rather than magnitude. Read the full page →
- Quasi-experimental design
- An impact evaluation design that constructs a comparison without random assignment — difference-in-differences, regression discontinuity, matching, interrupted time series or synthetic control.
- Randomised controlled trial (RCT)
- A design that assigns eligible units to treatment or control by lottery, making the groups statistically identical at baseline so that later outcome differences estimate the programme's impact. Read the full page →
- Realist evaluation
- Pawson and Tilley's approach asking what works, for whom, in what circumstances: programmes offer resources that fire mechanisms in some contexts and not others, expressed as context–mechanism–outcome configurations. Read the full page →
- Regression discontinuity design (RDD)
- A design for programmes assigned by a cutoff on a score: units just above and just below the threshold are compared, estimating the effect for the population near the cutoff. Read the full page →
- Relevance
- One of the OECD-DAC evaluation criteria: the extent to which the intervention's objectives and design respond to the needs, policies and priorities of beneficiaries and partners, and continue to do so as circumstances change. Read the full page →
- Reliability
- The consistency of a measurement process: the same method applied to the same situation returns the same result, across time, places and data collectors. Read the full page →
- Results chain
- The causal sequence from inputs through activities and outputs to outcomes and impact, with attributability decreasing at each step away from the programme's control.
- Results framework
- The donor-facing artefact that arranges a programme's intended results hierarchically with indicators at each level. Read the full page →
- Results-based management (RBM)
- A management discipline that directs resources, decisions and accountability toward defined results rather than activities, using a results chain, indicators and regular performance evidence. Read the full page →
- Risk register
- A governed list of identified risks with their likelihood, impact, owner and mitigation, reviewed on a fixed rhythm rather than filed at project start. Read the full page →
- Sampling
- Selecting units from a population so that measurements on the sample support conclusions about the whole, with the sampling method determining what inferences are defensible. Read the full page →
- Secondary data
- Data originally collected by others — censuses, administrative records, prior surveys — reused for monitoring or evaluation, with quality inherited from the original collection.
- Selection bias
- Distortion that arises when those who receive an intervention differ systematically from those who do not, so that outcome comparisons mix programme effects with pre-existing differences.
- SMART criteria
- The most widely used indicator quality test — Specific, Measurable, Achievable, Relevant, Time-bound. Its known weakness is a bias toward what is easily counted. Read the full page →
- SPICED criteria
- A quality test for participatory and qualitative indicators — Subjective, Participatory, Interpreted, Cross-checked, Empowering, Diverse — developed as a deliberate counterpoint to SMART. Read the full page →
- Spillover
- Intervention effects that reach people outside the treated group — neighbours adopting a practice, markets adjusting — which contaminate the comparison group and bias impact estimates.
- Stakeholders
- The agencies, organisations, groups and individuals with a direct or indirect interest in an intervention or its evaluation.
- Summative evaluation
- Evaluation conducted at or after completion to judge merit, worth or significance and inform decisions about continuation, scaling or termination.
- Sustainability
- One of the OECD-DAC evaluation criteria: the extent to which the net benefits of the intervention continue, or are likely to continue, after major assistance has ended. Read the full page →
- Synthetic control method
- A design for a single treated unit — a county, a country — that builds a weighted combination of untreated units to reproduce the treated unit's pre-intervention trajectory, then reads the post-intervention gap as the effect. Read the full page →
- Target
- The value an indicator is intended to reach by a specified date. Defensible targets state their basis — trend projection, benchmark or negotiated ambition — and distinguish cumulative from period values. Read the full page →
- Terms of reference (ToR)
- The commissioning document that defines an evaluation's purpose, scope, questions, methods, deliverables, timeline and quality standards — the contract the finished evaluation is judged against.
- Theory of change
- An explicit, diagrammed account of how and why an intervention is expected to produce its intended results, surfacing the assumptions between each step. Read the full page →
- Theory-based evaluation
- The family of approaches — contribution analysis, process tracing, realist evaluation — that assess causality by testing an explicit programme theory against evidence rather than by constructing a statistical counterfactual. Read the full page →
- Triangulation
- Using several data sources, methods or analysts to corroborate a finding and reduce the bias of any single approach.
- UNEG Norms and Standards
- The United Nations system's common reference for evaluation, adopted 2005 and revised 2016: general norms such as independence, impartiality, credibility and utility, paired with standards for how evaluation functions deliver them. Read the full page →
- Utilization-focused evaluation (UFE)
- Patton's approach that designs every step of an evaluation around its primary intended users and their intended uses, judging the evaluation by whether it is actually used. Read the full page →
- Validity
- The extent to which an indicator or measurement actually captures the concept it claims to measure — the first dimension checked in any data quality assessment. Read the full page →
- Value for money (VfM)
- The judgement of whether an intervention makes optimal use of resources, commonly analysed through the 4Es: economy, efficiency, effectiveness and equity. Read the full page →