02 · Evaluation approaches

Feminist and Gender-Responsive Evaluation

Gender-responsive evaluation assesses how an intervention affects gender and power relations, and conducts the evaluation itself in a way that is inclusive and fair; feminist evaluation goes further, treating inequity as structural and the evaluation as an instrument of change. Sex-disaggregated data is the floor of this practice, not its content — the defining move is analysing power, and being explicit about whose knowledge counts.

Last updated · Reviewed against 4 cited sources

Four words that are not synonyms

The vocabulary in this corner of evaluation is used loosely, and the looseness has consequences: a terms of reference that says “gender-sensitive” when the commissioner means “gender-responsive” buys a different evaluation. The working distinctions, as they have settled in UN and donor practice [1]:

  • Gender-blind evaluation does not ask the question. It reports on “beneficiaries” and “households” as if these were undifferentiated units, and treats any gendered pattern in the data as noise.
  • Gender-sensitive evaluation notices gender: it disaggregates data by sex and remarks on differences. It stops at description.
  • Gender-responsive evaluation makes gender part of both what is evaluated and how. It examines the intervention’s effects on gender equality and on the power relations that produce inequality, and it conducts the evaluation process itself inclusively — in who is asked, who asks, and who participates in judgement [1].
  • Feminist evaluation shares the responsive practice and adds explicit commitments: gender inequity is structural and systemic, not an individual attribute; knowledge is situated — who produces it and from what position shapes what it contains; and the evaluation is oriented to action on inequality, not only to documenting it. Where gender-responsive evaluation can be a professional standard applied by any competent team, feminist evaluation is a declared stance.

The ladder below orders the practice by depth. Each layer is legitimate work; the failure mode is claiming an upper layer while doing a lower one.

Four-layer pyramid of gender responsiveness from counting women to shifting power

A pyramid of four layers, bottom to top: counting women through sex-disaggregation; comparing outcomes to find gender gaps; analysing causes in norms, power and access; and shifting power through transformative questions and use. An arrow up the side is labelled depth of gender responsiveness. Each layer carries an example evaluation question.

Counting women — sex-disaggregation“How many women and men were reached?”Comparing outcomes — gender gaps“Did women and men benefit equally, and where not?”Analysing causes — norms, power, access“What norms and constraints produced the gaps we found?”Shifting power — transformative questions & use“Did control over resources and decisions change — and for whom?”depth of gender responsiveness
Figure 1. Depth of gender responsiveness. Each layer presupposes the one below it; the example questions show what an evaluation at that depth actually asks.Synthesised from the UN Women Evaluation Handbook (2022).

What feminist evaluation commits to

Feminist evaluation’s distinctive commitments deserve stating precisely, because each one has a methodological consequence rather than being a mood [1][4]:

  • Power analysis is the core analytic task. The evaluation examines who holds resources, decision authority and voice, and how the intervention altered — or reproduced — that distribution. This is why an evaluation can report excellent sex-disaggregated outcomes and still fail a feminist reading: a livelihoods project can raise women’s incomes while leaving control of the income, and the labour burden, exactly where they were.
  • Knowledge is situated. Who frames the questions, who is believed as a source, and who interprets the findings are treated as analytic decisions with consequences, not logistics. In practice this pushes feminist evaluation toward the participatory end of the control spectrum, with women affected by the intervention involved in framing and interpretation.
  • Action orientation. The evaluation is designed so its findings can be acted on by and for those disadvantaged by current arrangements — which shapes reporting formats, dissemination and follow-through, not just recommendations. The evaluation profession’s own principles now name the common good and equity as commitments of practice, which places this orientation inside professional norms rather than outside them [4].

Gender-responsive practice end to end

The UN Women handbook’s contribution is operational: it treats gender-responsiveness as a property of the whole evaluation process, manageable step by step, rather than a topic to be covered [1]. The practice, compressed:

  • Questions. Gender analysis enters the evaluation questions themselves — effects on gender equality, on differently situated women and men, and on the causes of inequality — not a standalone “gender question” appended as number nine.
  • Team composition. Gender expertise on the team is a staffing requirement, and team composition (including language and gender balance of interviewers) is an instrument-design decision: who asks changes what is said.
  • Methods. Mixed methods as the default, because gendered effects live partly in what surveys count and partly in what only safe, well-facilitated qualitative work surfaces. Sampling and instruments must reach women in the conditions of their lives — timing, location, childcare, privacy from other household members.
  • Process standards. The evaluation is itself conducted inclusively and ethically — stakeholder engagement that includes rights-holders, informed consent that is real, and reference-group arrangements in which affected women have voice [1].
  • Use. Findings return to those with the most at stake in usable form, and management response mechanisms are tracked for whether gendered findings translate into changed practice.

The recognised anti-pattern is the gender annex: a conventional evaluation with a commissioned gender chapter, written by the team’s junior member from the disaggregated tables. The handbook’s whole architecture exists to prevent exactly this [1].

Disaggregation is the floor

Sex-disaggregated data is non-negotiable and insufficient — both halves matter. Without disaggregation, gendered effects are arithmetically invisible. With only disaggregation, the evaluation knows that outcomes differ but nothing about the mechanisms — norms, time use, control of assets, safety — that produced the difference, which is where an intervention can actually act. Disaggregation design (which dimensions, decided when, with what intersections — sex with age, disability and location, since “women” is not a homogeneous category) is treated with the rest of measurement design in baselines, targets and disaggregation; the analytic layers above it are this page’s subject [1].

African framings: the intersecting critique

Chilisa and Mertens press a challenge that intersects the gender lens rather than merely accompanying it: evaluation frameworks imported into African contexts carry their originating worldviews — individualist units of analysis, extractive data relationships, criteria set far from the evaluand — and applying them unexamined can do epistemic violence, misdescribing what communities value and how change happens in relational terms [2]. Their argument for indigenous, African-rooted frameworks is directly relevant here for two reasons. First, gender norms are precisely the kind of context where imported categories misfire; an analysis of “women’s empowerment” that ignores relational and communal conceptions of well-being can misread both the problem and the change. Second, the remedies converge: both feminist and Made-in-Africa framings insist that whose knowledge counts is an evaluation-design question, answered deliberately [2]. The continental principles that carry this agenda into practice standards are treated at African evaluation principles.

Safety and ethics in gender data collection

Gender-responsive evaluation routinely touches the most dangerous data in the sector: experience of violence, control of money, contested decision-making. The obligations are ethical-by-design, not review-stage formalities, and the UNEG ethical guidelines frame them for UN-commissioned work [3]:

  • Do-no-harm applies to the interview itself. Privacy from other household members is a safety condition, not a data-quality nicety; being seen to answer questions about violence or money can create risk after the team leaves.
  • Consent must be meaningful — in language and terms the respondent controls, with genuine ability to decline in front of no one whose opinion constrains her [3].
  • Interviewer selection and training are ethical controls: who asks, of what gender, with what referral knowledge when a disclosure occurs.
  • Small cells re-identify. Disaggregated results in small communities can expose individuals; suppression rules belong in the analysis plan.

These duties sit within the general framework treated in ethics and consent, and the accountability obligations owed to affected people — feedback, complaint channels, and closing the loop — in accountability to affected populations.

Sources

  1. UN Women Evaluation Handbook: How to Manage Gender-Responsive Evaluation — UN Women Independent Evaluation Service, 2022.The UN system's operational guidance for gender-responsive evaluation, end to end — questions, process, team, methods and use. 2022 edition; first issued 2015.
  2. Indigenous Made in Africa Evaluation Frameworks: Addressing Epistemic Violence and Contributing to Social Transformation — American Journal of Evaluation, 2021.Chilisa & Mertens — the case for African-rooted evaluation frameworks and the critique of imported paradigms.
  3. Ethical Guidelines for Evaluation — United Nations Evaluation Group (UNEG), 2020.The UN system's ethical obligations for evaluators — the frame for safety and dignity in gender data collection. 2020 revision.
  4. Guiding Principles for Evaluators — American Evaluation Association, 2025.The profession's principles, including its explicit commitment to the common good and equity. Most recently updated 2025.