08 · Reporting, use and communication

Writing and judging evaluation reports

A credible evaluation report is a traceable argument: every finding rests on identifiable evidence, every conclusion on findings, every recommendation on conclusions. Structure, methodology honesty and a stand-alone executive summary serve that chain — and the chain, not the prose, is what quality reviewers test first.

Last updated · Reviewed against 4 cited sources

What a report is for

An evaluation report has one job: to carry an argument from evidence to judgement in a form a stranger can audit. Everything conventional about report structure follows from that job. The commissioning reader needs to know what was evaluated and why; the sceptical reader needs to know how the evidence was gathered and where it is weak; the acting reader needs judgements and recommendations they can take to a decision forum. The quality instruments that govern the genre — the UNEG Quality Checklist for Evaluation Reports and the OECD DAC Quality Standards for Development Evaluation — are, at bottom, itemised versions of those three readers’ entitlements [1] [4].

The canonical structure

Across the UNEG checklist, the DAC standards and current agency guidelines, the load-bearing sections are stable [1] [3] [4]:

The canonical evaluation report structure and what each section owes the reader
SectionWhat it must deliver
Executive summaryThe whole argument — object, questions, method in one line, key findings, conclusions, recommendations — readable on its own
Context and objectWhat was evaluated: the intervention, its logic, scale, timeframe, stakeholders, and the state of play at evaluation time
Purpose, scope and questionsWho the evaluation is for, what it covers and excludes, and the questions it will actually answer
Methodology and limitationsDesigns and methods per question, sampling, data sources — and the honest boundary of what the evidence can support
FindingsWhat the evidence shows, question by question, with the evidence visible or cited to annexes
ConclusionsThe evaluators' judgements, each explicitly resting on named findings
RecommendationsAddressed, specific, prioritised actions, each traceable to conclusions
AnnexesTerms of reference, instruments, sampling detail, evidence tables, persons consulted — the audit trail
Table 1. Section order varies by agency template; the obligations do not.

The annexes deserve more respect than they get. A report whose sampling strategy, instruments and evidence tables are inspectable invites the scrutiny that builds credibility; a report that asserts “interviews and document review were conducted” forecloses it. The ALNAP guide’s advice runs the same direction from the practical side: the annex is where the main text earns its brevity [2].

The evidence chain: the most-failed check

Quality reviewers who assess many reports converge on the same first test, and it is not prose quality. It is traceability: can each recommendation be followed back to a conclusion, each conclusion to findings, each finding to evidence presented or cited? The chain fails in two directions. Downstream, the orphan recommendation — a proposal reflecting the evaluator’s general convictions rather than anything this evaluation found — is the commonest single defect flagged in quality assessment [1] [3]. Upstream, data collected and never used signals an evaluation that gathered what was convenient rather than what its questions required.

Evidence-to-recommendation traceability in an evaluation report

Four rows of cards, bottom to top: evidence, findings, conclusions, recommendations. Solid lines connect specific evidence cards to the findings they support, findings to conclusions, and conclusions to recommendations. One recommendation card on the right has no upstream line and is drawn dashed, labelled orphan recommendation, reject. One evidence card has no downstream line and is labelled unused, report bloat.

RECOMMENDCONCLUDEFINDEVIDENCERec 1Rec 2Rec 3⚠ orphan — rejectConclusion AConclusion BFinding 1Finding 2Finding 3SurveyInterviewsRecordsSite visitsunused — report bloat
Figure 1. The evidence chain a quality reviewer traces: solid lines are the traceability a credible report makes visible. The orphan recommendation, with nothing upstream, fails the review; the unused evidence card is bloat that weakens confidence in the rest.Traceability requirement per the UNEG Quality Checklist for Evaluation Reports and DAC Quality Standards.

The discipline that produces a clean chain is unglamorous: an evidence table per evaluation question, maintained during analysis, mapping each emerging finding to its supporting data — and pruning both findings without evidence and evidence without findings before drafting. Theory-based designs make this explicit by construction; contribution analysis is essentially an evidence-chain method with the chain as the deliverable.

Methodology honesty

The methodology section is where reports most often protect themselves into vagueness. The standards require the opposite: designs and methods stated per question, sampling described concretely, data sources identified, and — critically — limitations stated plainly [1] [4]. A limitations section is not a confession; it is calibration. “Outcome claims rest on respondent recall over three years and should be read as indicative” tells the reader exactly how hard each claim can be leaned on, and a report that calibrates its claims earns more trust than one that asserts uniformly. The underlying data problems that limitations sections must disclose — completeness, consistency, verification status — are the subject of data quality assessment.

Language belongs to the same discipline. Findings sit in the indicative — what the evidence shows — and causal vocabulary must be sized to the design used: a pre-post comparison with no counterfactual supports “contributed to”, not “resulted in”. Overclaiming in the verbs is the subtlest way a competent evaluation becomes an incredible report.

Executive summaries and length

Most of a report’s decision-making readership reads only the executive summary, so the standards treat it as a stand-alone document: object, purpose, questions, method in a sentence, the principal findings and conclusions, and the recommendations — with no new material and no claims absent from the main text [1] [3]. Length discipline follows audience reality. A summary that runs to a dozen pages has simply relocated the report; the working test is whether a board member can absorb it in one sitting and accurately relay the conclusions to someone who has not read it.

Recommendations that can be actioned

Recommendations fail in implementation for defects visible at drafting time. The quality bar, consistent across the UNEG checklist and agency guidance [1] [3]:

  • Addressed — each recommendation names the unit or role expected to act; “stakeholders should…” is addressed to no one.
  • Specific — an action, not an aspiration: “revise the targeting criteria to include X” rather than “strengthen targeting”.
  • Prioritised and few — a report with thirty recommendations has made none; ranking forces the evaluators to judge.
  • Realistic and costed where possible — a recommendation whose cost is unknowable to its addressee will be deferred indefinitely.
  • Traceable — the chain again: each recommendation cites the conclusions it follows from.

What happens after the recommendations land — acceptance, rejection with reasons, action tracking — is a governed process of its own, treated at use and the management response.

Judging a report: the reviewer’s toolset

For readers on the receiving end, the same instruments that guide drafting serve as scoring tools. The UNEG checklist works as an item-by-item pass over structure, evidence and ethics [1]; the DAC standards frame the process-level questions — was the evaluation independent, were stakeholders consulted, is the report public [4]; agency guidelines such as UNDP’s add the house template and the quality-assessment machinery that scores completed reports [3]. Any figures and charts in the report answer to a further standard of their own — honest visual representation, covered at dashboards and data visualisation. A commissioner needs no evaluation training to run the traceability test in Figure 1; it can be done with a highlighter, and it is the single highest-value hour a report reviewer can spend.

Sources

  1. UNEG Quality Checklist for Evaluation Reports — United Nations Evaluation Group, 2010.The UN system's item-by-item checklist for judging a report's completeness and quality.
  2. Evaluation of Humanitarian Action Guide — ALNAP/ODI, 2016.Buchanan-Smith & Cosgrave. Practical guidance on reporting and communicating evaluations, written for real-world constraints.
  3. UNDP Evaluation Guidelines — UNDP Independent Evaluation Office, 2021.A full agency reporting standard in current use: report structure, quality assessment and the follow-up machinery.
  4. Quality Standards for Development Evaluation — OECD DAC Network on Development Evaluation, 2010.The DAC's process standards, including what a completed evaluation report must contain and disclose.