07 · Project and strategy M&E

OKRs and KPI cascades in institutions

OKRs pair a qualitative, ambitious objective with a few measurable key results, graded honestly at the end of a short cycle. They work as a focus-and-alignment device, not an accountability system — and the moment key results are wired to appraisal or budget consequences, Goodhart's law begins converting every measure into a worse one.

Last updated · Reviewed against 5 cited sources

The mechanics, from the primary sources

The OKR method compresses into a sentence John Doerr uses as its formula: “I will (Objective) as measured by (Key Results)” [2]. The method’s lineage runs from Andy Grove’s management practice at Intel, through Doerr, to Google — which adopted OKRs in its first year and remains the reference implementation whose public documentation this page cites [1, 2].

The two components have deliberately different characters:

  • The objective is qualitative, significant, and directional — a statement of what the team intends to achieve that a person could be motivated by. It carries no numbers. “Make our routine data trustworthy enough to run the programme on” is an objective; “increase reporting completeness to 95%” is not.
  • Key results (typically three to five per objective) are the measurable claims by which achievement will be judged: specific, time-bound, and verifiable such that at the end of the cycle a grader can score each one without argument [1, 2]. Key results state outcomes, not activities — “completeness of facility reporting ≥ 95% for three consecutive months” rather than “conduct data-quality training”.

Two design norms distinguish OKRs from ordinary target-setting, and both come straight from the Google practice [1]:

Stretch, by policy. Objectives are set beyond comfortable reach. Google’s public guidance grades key results on a 0–1.0 scale and treats scores around 0.6–0.7 as success for ambitious (“moonshot”) goals: a team scoring 1.0 across the board is read as having sandbagged, not excelled. (Teams also mark some OKRs as committed rather than aspirational, where full delivery is genuinely expected — the distinction must be explicit at setting time.)

Short cycles, honest grading. OKRs run on a quarterly-to-annual cadence: set at cycle start, visible to everyone, graded at cycle end, and the grades used as input to the next cycle’s conversation. Grading is cheap and quick — the value is in the discussion of why a key result landed at 0.4, not in the decimal [1].

The precondition for the stretch norm is the one most institutional adopters violate on day one: grades must be decoupled from appraisal and pay. Scoring 0.7 can only be success in a system where 0.7 is safe to report. The moment OKR grades feed bonus arithmetic, every incentive Grove’s method relies on inverts — targets sandbag, grades inflate, and the instrument dies while its paperwork lives on. This is Goodhart’s territory, treated below.

OKRs, KPIs, performance-contract targets: a disambiguation

Institutions adopting OKRs rarely start from zero; they already run KPIs, and in the Kenyan public service they may also sit inside the statutory performance-contracting regime, in which public institutions negotiate annual performance targets with government and are scored against them under the guidelines issued for each cycle [3]. These are three different instruments, and cascading them into one undifferentiated “targets” pile is how organisations end up gaming themselves.

OKRs, KPIs and performance-contract targets
OKR key resultKPIPerformance-contract target
NatureAmbitious claim for one cycleStanding health metricNegotiated annual commitment
LifespanOne quarter to one year, then resetContinuous, year on yearOne contract cycle, formally reviewed
Ambition normStretch; ~0.6–0.7 is successThreshold; below target is a problemRealism; scored against agreed criteria
Consequence of a missA learning conversationInvestigation and correctionInstitutional score, published ranking
OwnerThe team that set itA function or process ownerThe accounting officer
Fails whenWired to appraisalNobody watches it between reviewsTargets negotiated soft
Table 1. The columns differ most on the dimension that matters: what happens when the number is missed.

The practical rule of thumb: KPIs monitor the machine; OKRs concentrate effort on changing it; performance contracts bind the institution externally [3, 4]. A quantity can graduate between roles — a key result achieved for three consecutive cycles should usually retire into a KPI with a threshold — but a quantity should never hold two roles at once, because the roles carry contradictory ambition norms.

Cascading without fragmenting

The naive cascade decomposes arithmetically: the organisation’s key result is split into departmental shares, each department splits its share across units, and so on down. This produces perfect paper alignment and a familiar disease: every unit optimises its fragment, the fragments do not add back up to the mission, and the seams between units — where most real work fails — belong to no one.

The alternative the OKR literature insists on is alignment by conversation: the organisational OKRs are published first; each team then drafts its own OKRs by answering “what is our best contribution to those?”, and the drafts are reconciled in cross-team discussion rather than imposed by division [1, 2]. Two structural devices make the difference in practice:

  • Shared key results. Where an outcome needs two units (a data directorate and a service delivery department; procurement and programmes), give both units the same key result rather than splitting it. Shared KRs are the single most effective anti-silo device in the method — the seam becomes jointly owned instead of unowned.
  • Local objectives, not distributed fractions. A county office told to “deliver 8% of the national target” has been handed an accounting entry. The same office asked to write the objective that best advances the national one from its position will usually produce something more ambitious and more locally intelligent — and will own it.

Depth matters too: a cascade deeper than two or three levels is a bureaucracy generator. Below that depth, teams need good KPIs and clear priorities, not their own ceremonial OKR sets.

OKR cascade lattice with a shared key result, beside a warning panel showing a gamed metric rising while the mission trend stays flat

Left: a lattice with one organisational objective node at the top holding three key-result chips, connected by solid lines to two department nodes below, each holding an objective and its own key-result chips. A horizontal dashed line links one key-result chip in each department, labelled shared key result — prevents silo optimisation. Right: a warning panel titled gaming signature contains two small lines — one, the reported metric, bending sharply upward; the other, labelled mission, staying flat.

Organisational objectiveKR 1KR 2KR 3alignment conversations, not decompositionDepartment A objectiveA · KR 1A · KR 2shared KRDepartment B objectiveB · KR 1B · KR 2shared KRshared key result — prevents silo optimisationGaming signaturemetric ↑mission flatmetric up, mission flat:audit the measure, not the team
Figure 1. A two-level OKR cascade. Departments align to the organisational objective through their own objectives; the dashed tie marks a key result shared by both departments — the anti-silo device. The side panel shows the gaming signature every grading review should look for: the metric bending upward while the mission line stays flat.

Institutional performance reporting

Where OKRs or cascaded KPIs feed an institutional reporting rhythm — the quarterly performance review a board, ministry or council actually sees — three disciplines keep the rhythm informative rather than ceremonial, all inherited from results-based reporting practice [4]:

  • Traffic lights with rules. Red/amber/green is legitimate compression, but only when each colour has a written, arithmetic definition per indicator (distance from trajectory, not distance from comfort). Unruled RAG ratings converge on amber, which conveys nothing.
  • Narrative with the number. Every off-track rating carries three sentences: what happened, why, and what changes next quarter. A number without its “so what” is an invitation to ignore it; the reporting cluster’s pages take this further.
  • Data of record. Reported figures trace to the monitoring system, with the same definitions the indicator reference sheets fix — not to figures re-derived in slide decks. Institutions that let the performance report and the M&E system diverge end up governed by whichever number is more flattering.

Goodhart’s warning, and living with it

The strongest empirical regularity in performance management is the one usually stated as Goodhart’s law, in Marilyn Strathern’s canonical phrasing: “When a measure becomes a target, it ceases to be a good measure” [5]. Strathern’s own case was the British university audit regime — ratings introduced to observe quality began, once consequential, to reorganise the activity they observed around producing ratings [5]. The mechanism is fully general: attach consequences to a proxy, and effort migrates from the construct to the proxy.

In OKR and KPI regimes the signatures are predictable. Target fixation: the measured thing crowds out the unmeasured mission (the tunnel-vision cousin of measurability bias that the indicator quality-criteria page dissects). Sandbagging: cycle-start targets negotiated down to what is already assured — endemic wherever grades carry consequences, which is why negotiated target regimes invest heavily in target vetting [3]. Metric inflation: definitions quietly loosened, denominators quietly narrowed, until the number rises without the world changing.

Mitigations that work are structural, not exhortative:

  1. Pair every consequential metric with a counter-metric that measures what gaming the first would damage — speed with error rate, enrolment with retention, reporting completeness with verification-failure rate. The pair is gameable only by actually doing the job.
  2. Keep grading honest by keeping it safe. The 0.6–0.7 norm survives only where grades are decoupled from individual consequence [1]; the counterpart in consequential regimes is independent verification of reported results — which is where the M&E function, with its data-quality assessment toolkit, walks back on stage.
  3. Watch for the signature. A metric improving faster than its counter-metric and faster than any plausible mechanism — metric up, mission flat — is an audit trigger on the measure, prior to any celebration of the team.

The layer OKRs occupy in an M&E-mature organisation

OKRs are a strategy-execution instrument, and in an organisation that already runs a results architecture they occupy the layer above it: concentrating institutional attention, each cycle, on the few changes that most need forcing. They do not replace the programme results layer — the indicators, baselines and targets that account for what programmes deliver and achieve — and an organisation that lets quarterly OKRs supplant its results framework has traded its accountability spine for a productivity technique. The framework layer itself is out of this page’s scope; see the results-framework explainer at monival.com. The border discipline runs in both directions: key results may draw on programme indicators, but programme indicators keep their own definitions, cadence and quality regime regardless of which key results come and go.

Checklist before rolling out OKRs

  • Objectives are qualitative and significant; key results are measurable outcomes a grader can score without argument.
  • Aspirational and committed OKRs are distinguished at setting time, and the ~0.6–0.7 success norm is stated policy for the former.
  • Grades are decoupled from appraisal and pay; reported results feeding consequential regimes are independently verifiable.
  • Cascading is by alignment conversation, at most two or three levels deep, with shared key results across every critical seam.
  • Every consequential metric has a paired counter-metric, and reviews look for the metric-up-mission-flat signature.
  • The programme results architecture keeps its own integrity beneath the OKR layer — OKRs concentrate effort; they do not account for results.

Sources

  1. Guide: Set Goals with OKRs — Google re:Work, 2016.Google's public account of its OKR practice: ambitious objectives, measurable key results, 0–1.0 grading with ~0.6–0.7 as the sweet spot for stretch goals.
  2. What is an OKR? Definition and Examples — WhatMatters.com (John Doerr's OKR resource), maintained continuously.The canonical formula — 'I will (Objective) as measured by (Key Results)' — and the method's lineage from Andy Grove at Intel through Doerr to Google.
  3. Performance Contracting Guidelines for FY 2025/26 (22nd Cycle) — Republic of Kenya, Executive Office of the President, 2025.The statutory cousin: Kenya's negotiated, scored public-service performance contracts — a target regime with consequences, and therefore a different instrument from OKRs.
  4. Results-Based Management Handbook — United Nations Development Group, 2011.The results-management tradition institutional KPI reporting descends from; the layer OKRs must sit above, not replace.
  5. 'Improving ratings': audit in the British university system — European Review, 5(3), 305–321, 1997.Strathern's formulation of Goodhart's law — 'When a measure becomes a target, it ceases to be a good measure' — the standing warning over every consequential metric regime.