09 · Norms, standards and quality

Evaluation Quality Assessment and Meta-Evaluation

Evaluation quality is assessable at three layers — the process that produced the evaluation, the report it produced, and the use anyone made of it — and each layer has its own instruments. The OECD DAC Quality Standards for Development Evaluation (2010) govern process, phase by phase; the UNEG Quality Checklist (2010) structures product review; the JCSEE evaluation accountability standards supply the warrant for meta-evaluation. A sound quality assurance system applies all three lenses in proportion to the stakes: a two-hour structured review for a routine study, a commissioned meta-evaluation for a flagship one.

Last updated · Reviewed against 5 cited sources

Three layers, three questions

“Is this a good evaluation?” is three questions wearing one coat, and quality assessment goes wrong the moment they are conflated:

  1. Process — was the evaluation commissioned, designed and conducted properly? Independence protected, stakeholders engaged, methods matched to questions, ethics observed.
  2. Product — is the report itself sound? Object and context described, methodology and limitations disclosed, findings traceable to evidence, conclusions following from findings, recommendations anchored in conclusions.
  3. Use — did anyone do anything? Management response issued, decisions traceably informed, follow-up tracked.

The layers move independently. A flawless process can produce a badly written report; an elegant report can conceal a captured process; and evaluations that pass both tests routinely sink without a ripple of use. Each layer has instruments of its own, and a quality assurance (QA) system is essentially a decision about which instruments to apply, when, and how deeply.

Three assessment lenses applied in sequence to an evaluation

An evaluation report icon on the left is examined by three lenses in sequence. Lens one, process, applies the DAC Quality Standards phase by phase. Lens two, product, applies the UNEG Quality Checklist section by section. Lens three, use, checks for a management response and tracked actions. An output arrow leads to a quality rating badge with the caveat that a rating is not usefulness — check lens three.

evaluation1 · processDAC Quality Standardsplanning ✓ design ✓conduct ✓ reporting ✓2 · productUNEG Quality Checklistfindingsmethodsconclusions3 · usemanagement responseactions tracked: 6/9quality ratingrating ≠ usefulness — check lens 3
Figure 1. Three lenses on one evaluation: process (DAC Quality Standards), product (UNEG Quality Checklist), use (management response evidence). A quality rating that stops after the second lens says nothing about whether the evaluation mattered.Instruments: OECD DAC (2010); UNEG (2010).

Lens 1 — process: the DAC Quality Standards

The OECD DAC’s Quality Standards for Development Evaluation, approved by the DAC Network on Development Evaluation in January 2010 and endorsed by the DAC that February, remain the most widely referenced process standard in development evaluation [1]. The document follows the shape of an evaluation’s life. It opens with overarching considerations that apply throughout — including a free and open evaluation process, evaluation ethics, partnership and coordination, capacity development and quality control — then sets standards for purpose, planning and design; for implementation and reporting; and for follow-up, use and learning [1].

Used as an assessment instrument, the standards convert into a phase review: for each phase, was the standard met, and what is the evidence? The strength of this lens is that it catches failures no report review can see — terms of reference written to foreclose findings, stakeholders consulted after conclusions were drafted, an evaluator’s independence quietly traded away mid-assignment. Its limit is symmetrical: process review cannot tell you whether the report’s conclusions actually hold.

One confusion to retire permanently: the DAC Quality Standards are not the DAC evaluation criteria. The criteria — relevance, coherence and their companions — are judgement lenses applied to the intervention being evaluated; the Quality Standards govern the evaluation process itself. They are different OECD instruments doing different jobs. This site does not explain the criteria; see the OECD-DAC evaluation criteria explainer on monival.com.

Lens 2 — product: the UNEG Quality Checklist and scored review grids

For the report itself, the standard instrument is the UNEG Quality Checklist for Evaluation Reports, approved at the 2010 UNEG Annual General Meeting [2]. Its logic is that a sound report is inspectable section by section: the evaluation object and its context are described; the purpose, criteria and questions are stated; the methodology is explained with its limitations; findings respond to the questions and rest on evidence; conclusions add analytical depth to findings rather than repeating them; recommendations are relevant, actionable and grounded in the conclusions. Reviewing against the checklist means asking, of each section, “is it there, and does it do its job?” — which is a different and better question than “do I like this report?”.

Many organisations convert the checklist into a scored review grid: each section rated on a defined scale, with an aggregate quality score. Two disciplines keep scored grids honest:

  • Anchor every rating in text. A score without a cited passage (or a noted absence) is a mood. The grid should force the reviewer to point at the report.
  • Calibrate reviewers. Have two reviewers score the same reports independently and compare, at least periodically. Inter-reviewer consistency is the difference between a measurement system and a lottery; where scores diverge, the scale descriptors — not the reviewers — usually need fixing.

Product review pairs naturally with the guidance on writing evaluation reports: the checklist is the reader’s side of the same contract.

Lens 3 — use: the layer most QA systems skip

An evaluation that changed nothing has, on the utility-first logic of the JCSEE standards, failed — however it scored on the other lenses. The use lens is assessed with mundane evidence: does a management response exist; does it accept, partially accept or reject each recommendation with reasons; is there an action tracker; do later decisions cite the evaluation? None of this requires methodological skill, which is perhaps why formal QA systems so often omit it — it belongs to administration, not method. Build it in anyway; the mechanics are covered under use and management response.

Meta-evaluation: assessing evaluations systematically

Meta-evaluation — the systematic evaluation of an evaluation — is not an academic indulgence; it is an explicit expectation of the JCSEE Program Evaluation Standards, whose evaluation accountability group (E1–E3) exists to “encourage adequate documentation of evaluations and a metaevaluative perspective” [3]. In practice, commission one when the stakes justify it: a flagship evaluation that will drive a funding or policy decision, a contested evaluation whose findings are under attack, or a periodic sample of an organisation’s evaluation portfolio to test whether the QA system itself works. A competent meta-evaluation applies the process and product lenses above with an independent reviewer, documents its own criteria, and reports on the QA system as well as the study.

At system level, some institutions publish the architecture that makes this routine. The World Bank Group’s Evaluation Principles (2019) state publicly the principles its evaluation function is governed and judged by, giving external assessors a fixed yardstick [4]. UNDP’s Evaluation Guidelines (revised 2021) embed a standing quality assessment of the decentralised evaluations its programme units commission, run through its Independent Evaluation Office [5]. The transferable lesson is not the scale but the publication: a QA system whose criteria are public can be held to them.

A proportionate QA architecture

Assessment effort should scale with stakes. A workable tiering:

Proportionate quality assurance by evaluation stakes
TierWhenWhat to runEffort
Structured reviewEvery evaluation, however smallUNEG-checklist product review with anchored ratings; confirm a management response exists≈ 2 hours per report
Phase QAEvaluations above a spend or visibility thresholdDAC-standards process checkpoints at ToR, inception and draft-report stage, plus the structured reviewHours spread across the lifecycle
Independent meta-evaluationFlagship, contested or precedent-setting studies; periodic portfolio samplesCommissioned external assessment against all three lenses, criteria documentedA short study in its own right
Table 1. Instruments referenced: DAC Quality Standards (2010), UNEG Quality Checklist (2010), JCSEE accountability standards (2010).

Two design rules make the tiers work. First, decide the tier at commissioning, not at delivery — a QA level chosen after the findings are known is itself a process failure. Second, feed results back into the system: recurring weaknesses across structured reviews (limitations sections missing, recommendations untraceable to conclusions — the categories the checklists themselves probe) are commissioning problems to fix in the next round of terms of reference, not reviewer observations to file.

What to take from this page

Quality assessment is three separate judgements — process, product, use — made with instruments that already exist and are free: the DAC Quality Standards, the UNEG Quality Checklist, and the meta-evaluation discipline the JCSEE standards mandate. The failure mode to design against is single-lens confidence: a “high quality” stamp derived from one lens, silently generalised to the other two. Say which lens produced every rating, and never publish a quality score for an evaluation without knowing whether anyone used it.

Sources

  1. Quality Standards for Development Evaluation — OECD DAC Network on Development Evaluation, Paris, 2010.Approved by EvalNet 8 January 2010, endorsed by the DAC 1 February 2010. The process standard: overarching considerations, then standards for each phase of an evaluation.
  2. UNEG Quality Checklist for Evaluation Reports — United Nations Evaluation Group (UNEG), 2010.Approved at the 2010 UNEG Annual General Meeting; the standard product-review instrument for evaluation reports.
  3. The Program Evaluation Standards, 3rd edition (summary of record) — Joint Committee on Standards for Educational Evaluation (JCSEE), 2010.The evaluation accountability group (E1–E3) documents the meta-evaluation expectation quoted on this page.
  4. World Bank Group Evaluation Principles — World Bank Group / Independent Evaluation Group, 2019.A system-level model: a large institution publishing the principles its evaluation function is governed and judged by.
  5. UNDP Evaluation Guidelines — UNDP Independent Evaluation Office, 2021.Revised edition, June 2021. A working example of an institutional QA architecture, including quality assessment of decentralised evaluations.