04 · Data quality, sampling and collection

Survey and questionnaire design

Most measurement error in surveys enters between the indicator definition and the question a respondent actually hears. Questionnaire design is the discipline of closing that gap: precise wording, deliberate recall periods, response formats that match the analysis plan, careful translation, and a pretest-pilot-lock sequence that treats the instrument as a versioned artefact.

Last updated · Reviewed against 4 cited sources

The operationalisation gap

An indicator definition says “percentage of households with year-round access to an improved water source”. A questionnaire asks “Do you have good access to water?”. Everything wrong with the resulting data was decided at that moment: “good” is undefined, “you” may mean the respondent or the household, the season is unspecified, and “water” has quietly replaced “improved water source”. No sample size rescues an instrument that measures a different construct than the indicator names — and the survey-error literature is blunt that non-sampling error of this kind routinely dwarfs the sampling error that reports dutifully quantify [2].

The discipline, then, is traceability: every questionnaire item should be derivable from an indicator definition (or an explicit analysis need), and every term in the definition should be visible in the question or its interviewer instructions. Where indicator definitions live in a governed reference sheet, the mapping is a review step, not archaeology — see performance indicator reference sheets.

Wording rules with evidence behind them

The Living Standards Measurement Study programme spent years testing questionnaire choices across countries, and its lessons remain the standard [1]:

  • One question per question. “Do you have access to credit and savings services?” is two questions; any answer is uninterpretable. Split double-barrelled items, always.
  • Concrete beats abstract. Ask about behaviour and events (“How many times did a child in this household visit a health facility in the last 30 days?”) rather than dispositions (“Do you use health services regularly?”).
  • Specify the reference period — and choose it deliberately. Short recall periods (7 days for food consumption, 30 days for minor illness) reduce omission but miss rare events; long periods (12 months for hospitalisation, assets, shocks) capture rare events but invite telescoping — pulling earlier events into the window — and progressive forgetting [1]. The LSMS lesson is that the recall period is a per-topic design decision, not a house style.
  • Neutral formulations. “How satisfied are you with the training?” presumes satisfaction; offer the negative option in the question stem (“satisfied, dissatisfied, or neither?”).
  • The respondent’s vocabulary. Programme language (“beneficiary”, “intervention”, “value chain”) does not survive contact with a household interview. Terms must be the ones respondents use, established during pretesting.
  • Define the household, and every other unit, in the interviewer instructions. Membership rules (residency threshold, shared cooking) change measured household size and therefore every per-capita figure [1].

Sensitive topics — income, violence, sexual behaviour, opinions of authorities — need more than neutral wording: normalising preambles, self-completion or audio-assisted modes where feasible, interviewer matching and privacy protocols, all rehearsed in training [4]. The consent and referral obligations that apply are covered under ethics and consent.

Before-and-after repair of three defective survey questions

Three rows. Each row shows a defective question on the left — double-barrelled, vague recall, and leading — with the offending phrase underlined, an arrow, and the repaired version on the right with a label naming the fix. A footer strip shows the instrument lifecycle: draft, translate, cognitive test, pilot, then version 1.0 locked.

“Do you have access tocredit and savings services?”double-barrelledQ1 “…access to credit?”Q2 “…access to savings services?”fix: one question per question“Have you recently been ill?”vague recall period“In the last 30 days, have you beenill or injured?”fix: bounded, topic-appropriate recallHow satisfied are you withthe training?”leading — presumes satisfaction“Were you satisfied, dissatisfied, orneither, with the training?”fix: balanced stem offers all optionsInstrument lifecycle:drafttranslatecognitive testpilotv1.0 locked
Figure 1. Question repair: three common defects and their fixes. The footer shows the instrument lifecycle each repaired question then travels through.Defect taxonomy follows the LSMS design lessons in Grosh & Glewwe (2000).

Response formats: designed for the analysis plan

The response format is chosen at the moment of design but paid for at analysis. Working rules:

  • Categories must be exhaustive, mutually exclusive, and stable across rounds. A category set that changes between baseline and endline breaks the trend the survey exists to measure. Add “other (specify)” as the safety valve and review its contents at pilot [1].
  • Scales need anchored points, not just numbers. A 1–5 satisfaction scale means little unless each point carries a label, and label wording must survive translation into every fielded language.
  • Decide the “don’t know” and refusal policy per item, in advance. Suppressing “don’t know” manufactures false opinions; offering it too prominently invites satisficing. Whichever policy is chosen, code the two distinctly — they mean different things at analysis.
  • Record units explicitly wherever quantities are collected. Local units (bags, debes, acres vs. hectares), currencies and cumulative-versus-period amounts are classic sources of order-of-magnitude error; either fix the unit in the question or capture the unit alongside the value and convert with a documented table [1]. The data-quality consequences of unit confusion are catalogued under data quality dimensions.

Translation and cognitive testing in multilingual settings

In Kenyan and wider East African practice an instrument drafted in English will be administered in Kiswahili and often several other languages — which means the translated text is the instrument, and the English original is merely its specification. The working sequence [3]:

  1. Forward translation by a translator who knows the survey domain, not just the languages.
  2. Reconciliation or back-translation review — retranslating into the source language exposes drift, though the review discussion matters more than the mechanical back-translation itself.
  3. Cognitive interviewing in each fielded language: a handful of respondents from the study population think aloud while answering — What did that question mean to you? How did you arrive at your answer? — exposing comprehension failures that no desk review finds.
  4. Lock the translations together with the source: a wording change in one language is a change to the instrument and re-opens the others.

Where interviews will be conducted by bilingual enumerators translating on the fly from an English form, that fact should be reported honestly as a limitation — every enumerator becomes an uncontrolled translator, and reliability suffers accordingly [3].

Ordering, length and fatigue

Question order shapes answers. Early items set context for later ones; sensitive modules placed first depress cooperation, placed last they meet a tired respondent. The LSMS convention — household roster first, then progressively more sensitive and cognitively demanding modules, with the interview’s natural end reserved for interviewer observations — exists because it works [1]. Length is a quality variable, not just a cost variable: beyond a tolerable interview duration, item non-response and satisficing climb. The pilot, timed honestly, is where length is confronted; the remedy is cutting questions that no analysis needs — every item should have a named use before it earns its place [3].

Pretest, pilot, lock

Three distinct exercises, in order [3]:

  • Pretest (small, iterative): does each question work — comprehension, vocabulary, flow? Cognitive interviews and debriefs with a handful of respondents; expect to change wording repeatedly.
  • Pilot (dress rehearsal): does the whole system work — the instrument at full length, enumerator training, supervision, data flow, timing? Pilot data is inspected for item non-response, “other” overflow, scale clumping and outliers, then discarded.
  • Lock: the instrument becomes version 1.0. From this point, changes are breaking changes: they must be versioned, dated, logged against the indicators they affect, and flagged wherever trend comparisons cross the change. An unlogged mid-fieldwork wording “clarification” is a reliability failure waiting for a data quality assessment to find it — see data quality assessment.

Enumerator training belongs in the same breath: the interviewer is part of the instrument, and interviewer effects — differences in probing, pacing and tolerance for “don’t know” — are controlled by scripted instructions, practice interviews and field supervision, not by hope [3][4].

Digital form logic — enforced skips, range constraints, required fields — hardens all of the above at the point of entry; choosing a field data-collection platform and running it offline are practice topics covered by Monival’s guides to offline data collection and migrating between platforms, and are out of scope here.

Checklist before locking the instrument

  • Every item traces to an indicator definition or a named analysis use; every indicator in the plan is covered.
  • No double-barrelled, leading or undefined-term items survived pretesting.
  • Recall periods are set per topic and stated in every relevant question.
  • Category sets and scale anchors are stable across rounds and across languages.
  • All fielded language versions are cognitively tested and locked together.
  • The pilot ran end-to-end at full length, and its lessons were applied before v1.0.
  • A change log exists, and everyone who can edit the instrument knows a post-lock change re-opens it.

Sources

  1. Designing Household Survey Questionnaires for Developing Countries: Lessons from 15 Years of the Living Standards Measurement Study — The World Bank / Oxford University Press, 2000.Grosh & Glewwe (eds.). The deepest open evidence base on questionnaire design choices in developing-country surveys, module by module.
  2. Designing Household Survey Samples: Practical Guidelines (Studies in Methods, Series F No. 98) — United Nations Statistics Division, 2005.Places questionnaire quality within the survey-error framework: non-sampling error routinely dominates sampling error.
  3. DIME Wiki — survey design, piloting and questionnaire resources — World Bank Development Impact (DIME), n.d..Continuously updated practitioner reference (accessed August 2026) for questionnaire design, translation, piloting and enumerator training in impact evaluations.
  4. Qualitative Research Methods: A Data Collector's Field Guide — Family Health International (FHI), 2005.Mack, Woodsong, MacQueen, Guest & Namey. Interviewing technique and sensitive-topic practice that survey pretesting borrows from.