Survey Methods Handbook · Chapter 06 of 15

06 Devising survey questions

In this chapter…

This chapter shows how wording, response options and respondent cognition shape the answers a survey produces.

By the end of this chapter, you should be able to…

  • write clear questions that measure one intended concept
  • select response formats appropriate to the construct and mode
  • identify common wording, recall and response-option problems

In this chapter

By the end of this chapter, you should be able to:-

  • identify the communicative and scientific aims of quantitative surveys;
  • understand how principles of the psychology of respondent behaviour help to explain how questions about behaviour, attitudes and opinions are likely to be understood and answered;
  • demonstrate how an understanding of these principles can be used to develop questions that are most likely to yield valid, reliable, unbiased and discriminating data;
  • distinguish between open and closed questions, and describe the strengths and weaknesses of each;
  • describe key principles of question wording;
  • describe key principles of question ordering.

Communicative and scientific aims of quantitative surveys

People who are naturally articulate, intuitive and interested in the lives of others are often able to use their conversational skills to draw out their respondents on the required topics and elicit relevant information in a natural and easy-seeming way. Those who are gifted in this way may, with further training and experience, become good qualitative researchers.

Quantitative surveys also involve holding a special kind of ‘controlled conversation’ with respondents, in order to collect information on predetermined topics, and the questionnaire designer needs to have some of the same abilities to construct an artificial but natural-seeming conversation. The conversation is unnatural in that one person (the researcher or the interviewer) asks all the questions and the other person (the respondent) provides all the answers.  In a self-completion questionnaire, it is even more unnatural, since the questioner is not physically present and must instead correspond through the medium of the questionnaire.  Nevertheless, a good questionnaire can and should appear as a straightforward and logical sequence of requests for information, on clear and relevant topics and in a form and at a level of detail that the respondent is well able to provide. Providing that information may require some mental effort, but it should always be clear exactly what information is needed. These are the respondent-oriented aims of questionnaire design.

At the same time, however, the designer of quantitative instruments must also pursue another, quite different but equally important, scientific agenda. Questions asked on quantitative surveys should:-

  • provide ‘meaningful numbers’ for analysis (quantification);
  • measure the quantity or concept that they are intended to measure and no other (validity);
  • be as free as possible from systematic measurement bias (lack of bias);
  • be as free as possible from random measurement variability (reliability);
  • be able to detect relevant differences between respondents and groups of respondents (sensitivity or discriminatory power).

The art and science of questionnaire design consists in reconciling these two agendas. If, on the one hand, respondents are unable to relate the questions asked to their own concepts, knowledge and experience, the communication process will fail. If, on the other hand, the quantitative measures obtained fail seriously to meet the criteria just listed, the scientific aims of the survey will not be achieved. 

As information-gathering instruments, questionnaires have a carefully-calculated structure, but it is not possible to teach questionnaire design skills by defining a ‘questionnaire template’ (analogous to a statistical formula). The design of individual questionnaires must be determined by the particular topics and aims of the survey (including the exact form and level of detail required) and by the particular characteristics (knowledge of the subject matter, language skills, literacy, age etc.) of the respondent population. We will therefore proceed by giving examples of widely-used information-eliciting tactics that may help to satisfy the criteria mentioned above and by drawing attention to the many traps into which the unwary question and questionnaire designer can fall.


Some general principles, applying equally to interviewer-administered and self-completion questionnaires, are the following:-

  • Most survey respondents want to oblige and be polite.
  • But  they need to be told exactly what we require in terms of completeness, precision etc.
  • And there are limits to the amount of effort respondents will put in unless specially motivated.
  • Badly designed questions elicit inconsistent / incomplete / inaccurate / unreliable / biased responses.
  • But, when data collection has been completed, it is generally impossible to distinguish, without elaborate validation studies, between adequate and inadequate or misleading responses.
  • And even if it were possible to make such a distinction, poor responses once given cannot be converted into good responses.

Cognitive aspects of survey methodology

The term “Cognitive Aspects of Survey Methodology” (CASM) is applied to the growing interdisciplinary effort – involving survey methodologists, cognitive and social psychologists, anthropologists, socio-linguists and statisticians – to investigate and understand the cognitive processes employed by respondents in reading, comprehending and interpreting questions, and in formulating and providing answers to those questions.

Tourangeau and colleagues (2000) have proposed a model of the question response process, comprising four stages: comprehension; retrieval; judgement; response.  They point out that the model is not necessarily sequential, but may involve iteration through the stages. An understanding of these cognitive processes can help to inform the construction and wording of questions.

First, the respondent must perceive and attend to the question, infer its meaning  and thus identify the nature of the information sought by the researcher.  There is scope here for misunderstanding, with a consequent threat to response validity and reliability.


In the retrieval stage, the respondent has to: develop a strategy and a set of cues for retrieving the relevant material from memory; retrieve specific relevant information; fill in any gaps in memory through inference or the use of heuristics (short cuts).  Again, characteristics of the question and of the recalled material can affect the accuracy and comprehensiveness of this process (Jobe et al., 1993). 

Retrieval, however, does not necessarily yield a direct answer to the question posed.  Respondents may need to synthesise, supplement or otherwise process the information retrieved into a single overall judgement; they may also need to adjust an initial judgement to allow for omissions in retrieval (Touranegeau et al, 2000).  The type of question may influence judgement tasks and processes. 

The final stage in the model involves “mapping” the adjudged response on to one of the response categories offered.  This is not necessarily straightforward.  The response categories offered may involve “vague quantifiers” such as ‘usually’, ‘frequently’, ’not a lot’ etc. (Sudman et al, 1996), posing respondents with difficulties in choosing the most appropriate option.  Respondents also differ in the amount of effort they are prepared to put into selecting a response. Moreover, having made an initial choice of response category, they may consciously or subconsciously “edit” their response, in the interests of consistency, self-presentation or other criteria. 

Asking about behaviour

Most surveys include questions which ask about behaviour in some shape or form and some are mainly devoted to measuring behaviour (for example consultation behaviour, buying behaviour, travel behaviour etc.). The survey designer aims to identify an aspect of the behaviour of members of the study population that can be counted in a systematic and standardised way and which indicates in a consistent manner the amount or type of behaviour that those individuals do.  Estimates which may be derived from behavioural measurements on a sample include:-

  • proportions of the population engaging in the behaviour (e.g. percentage who smoke cigarettes);
  • average amount of the behaviour done or rate of engaging in the (e.g. average number of cigarettes smoked per day);
  • distributions of frequency, intensity or amount of behaviour across the population;
  • total numbers of people doing the behaviour (e.g. estimated total number of smokers in the population);
  • total amounts of the behaviour done (e.g. estimated total number of cigarettes consumed nationally per year);
  • money costs or expenditures associated with the behaviour (e.g. average expenditure per week per smoker on cigarettes);
  • and so on. 

Table 6 indicates the types of information that might be gathered and how the information might be used.

Measuring behaviour poses particular challenges in terms of collecting valid, reliable, unbiased and discriminating data, particularly where records of behaviour are compiled retrospectively and therefore rely on respondent’s recall of past events. It should be noted that answering questions about what one ‘usually’ or ‘normally’ does also involve recall. Adverbs of this kind when used in questions are known as ‘vague quantifiers’.

Distortions of recall and reporting of events and behaviour

There is evidence that recall and reporting of events and behaviour can be deficient or distorted in a number of ways. Some of them involve selective forgetting or rationalisation, others a tendency by respondents to reduce the amount of mental (cognitive) effort needed to arrive at an accepianswer by resorting to some mixture of recall and guessing. Different types of error can cause different biases and the same respondents may over-report in one context and under-report in another. Several types of defect in recall and reporting of events and behaviour occurring in the near or distant past are important in surveys.

Forgetting

Some events and behaviour are in themselves distinctive and memorable: for most people such events as having a baby or starting one’s first full-time paid job are in that category. For these types of event the respondent is likely not only to recall them, but also to be able (after a little thought) to date them fairly accurately, even if they occurred many years ago, because they have personal significance and can be mentally cross-referenced with other significant and dateable personal events. The less personal significance an event has, however, the more likely it is to be forgotten. Types of event or behaviour that are habitual or automatic are also much harder to recall, especially in relation to a specific reference period. This is particularly so where they are repetitive, so that individual occurrences are often hard to distinguish in memory and therefore to enumerate accurately and to place accurately in time.  Examples are routine (but not regular) small purchases and repeat visits to a therapist to have minor medical treatment for the same condition.

Table 6 Information collected by surveys about behaviour, life events and statuses

Topic areas (examples)

Typical survey aims

Examples of applications

Events and statuses occurring to individuals, households etc over time

Lifetime or shorter histories of:-

employment events and statuses

family events and statuses

housing events and statuses

health events and statuses

migration

accidents experienced

instances of crime victimisation

types and amounts of income received

To establish a complete and accurate record of relevant, defined events or episodes occurring to sample members over a specified back reference period, with identifiers, dates etc.

Studies of  (for example):

career trajectories of men and women

life patterns, family formation, fertility, demographic forecasting

housing careers/demand for housing.

health and disability careers

estimating crime rates

estimating income distributions

longitudinal research projects.

 

Behaviour

Details of current and recent:-

consumer behaviour

media consumption

expenditure amounts and patterns

uses and allocation of time

travel behaviour and transport choices

sports, leisure, exercise activities

health-related behaviour

dietary behaviour

educational choices and achievements

To quantify relevant behaviour and obtain detailed records of the frequency with which people engage in it, with details of amounts, circumstances etc of each occurrence

 

Needed (for example) to:-

guide marketing strategies

estimate aggregate expenditure for National Accounts

study household budgets

study patterns of time use

monitor dietary, smoking, travel, exercise, leisure behaviour etc.

study learning and skills acquisition

 

Recall performance can be improved by ‘cueing’ the respondent with significant events and incidents that may be associated in his or her mind with those of interest, but unfortunately in survey situations we seldom know, for each individual respondent, what specific cues will ‘ring the bell’.  Moreover, such personalised approaches are better suited to interviewer-administered surveys than to self-completion questionnaires.

Deliberate omissions and distortions

For many respondents, the events or behaviour that they are asked to recall may have little interest or significance and they may become embarrassed, bored or frustrated at being made to try and recall details. Such respondents may realise that they can escape from tedious questioning by saying ‘No’ or ‘Never’ to preliminary questions about occurrence of events or behaviour, even though the correct answer may be ‘Yes’.  This type of response pattern will obviously result in underestimates of the prevalence of the events or behaviour of interest. Once disaffection and the desire to escape sets in, the battle for good quality data is largely lost, since there are always many more ways that respondents can find to reduce the burden, for example by under-reporting or using short-cuts to give a ‘good enough’ answer (‘satisficing’), than there are ways in which the researcher can counteract such tendencies. The key factor is respondent motivation; this presents a particular challenge for postal and self-completion questionnaires, since there is no interviewer who can engage and encourage the respondent.

On the other hand, respondents’ desire to be helpful can also sometimes cause problems.  When asked to recall particular types of event or behaviour occurring within a reference period, respondents may assume that the researcher must want to know about some important in-scope item (say, a large purchase in a survey of expenditure, or a visit to a hospital accident and emergency department for an exacerbation of some health problem), even though they are more or less aware that it actually occurred (slightly) outside the reference period. This kind of motivated distortion is in practice hard to distinguish from (unconscious) ‘telescoping’, which is discussed below. It can be counteracted by stressing the importance of strictly observing the limits of the reference period.

Distortion can also occur if respondents find it embarrassing or ego-threatening to report accurately certain forms of behaviour in which they engage. Consumption of alcohol or illegal drugs, activities in the ‘black economy’, forms of sexual behaviour, and activities indicative of an unhealthy lifestyle might all tend to be under-reported for this reason, while socially desirable behaviour, such as eating fresh vegetables or taking exercise, might tend to be over-reported. It is hard to measure such mis-reporting tendencies accurately and unambiguously, since respondents cannot be relied upon to admit  to them if challenged and there are many other reasons why aggregate estimates from a survey might be too low or too high. For example, it may be known by comparison with other measures of total consumption that a survey produces under-estimates of the total amount of alcohol consumed by members of the population; but the under-estimation may be caused by any or all of sampling problems, differential non-response by heavier drinkers or forgetting of drinking occasions or what was consumed, as well as by deliberate under-reporting.

Researchers who wish to measure behaviour that they judge might be subject to deliberate mis-reporting, try to minimise such tendencies by various means. In general, the best strategy is likely to be: to convince each respondent that the survey is a serious, worthwhile and important one; that the accuracy of estimates for the population depends on all respondents doing their best to report accurately; that guarding the confidentiality of personal information is taken very seriously; and that there is no possibility of confidentiality being breached in reporting the survey results.  In an interviewer-administered survey, these assurances can be given verbally.  In a postal or other self-completion questionnaire, these points will usually need to be addressed in a covering letter, though further assurances can be given in the questionnaire itself, at the point when the sensitive questions are posed (e.g. We need to ask these questions about your income so that we can compare the findings from our survey with the population as a whole). Be aware, nonetheless, that the inter-personal psychology of reassurance is subtle. A barrage of assurances too early in the process of eliciting information can make respondents suspect that the survey must pose some hidden threat. It is important that both researchers and interviewers (if used) be themselves convinced that collecting the information required is reasonable, useful and justified and that they communicate that belief to respondents (through instructions, question wording and so on) in a way that is confident, unembarrassed and matter-of-fact. At the same time, they should be alert and ready to detect and address any worries that the respondent may have; in postal surveys, this may involve dealing with phone calls from sample members who are querying why a particular piece of information is being elicited.


Unconscious omissions and distortions

If pressed by survey questioning for detail of events or behaviour that they cannot really recall, respondents may resort to reconstructing what ‘must have happened’, what they ‘must have done’, or what they ‘usually do’. This is very common in conversation and we are often not clearly aware of when we are reconstructing or inferring, rather than directly recalling. Some of the behaviour question types discussed below implicitly rely on the assumption that respondents are good at reconstructing their activities in this way, but there is evidence that the results are often biased or unreliable.  In these situations, the form of questioning, including the words used, can influence what is ‘recalled’ in ways of which the researcher may be unaware. In one controlled experiment, subjects were shown film of a minor collision between two cars and then asked to recall details. Subjects were more likely to reply ‘Yes’ to the question ‘Was there any broken glass on the road...?’ when the second clause was ‘...when the two cars smashed into each other’ than when it was ‘...when the two cars collided.’ 

Rationalisation

Recall of material of high as well as low salience for the individual may be affected by the mainly subconscious cognitive process of rationalisation. Details that do not make satisfactory cognitive or emotional sense or appear irrelevant to the respondent tend to be unconsciously changed or discarded from memory, producing  a rationalised and simplified ‘story’ that is then produced in answer to questions. This tendency can distort even what seem to the respondent to be vivid memories. It is particularly likely to occur when there are no corrective factors at work, such as discussion with other persons who also directly recall the events and behaviour in question. In normal conversation, for example, spouses frequently challenge each other’s recall of family and domestic events, leading to discussion as a result of which an agreed and cross-referenced version emerges (‘It must have been before I changed jobs in September but after your mother broke her hip in June’).  But in survey research, respondents are generally exhorted to answer without help or input from others.

Distortion of the perception of elapsed time

An important way in which the perception of elapsed time may be distorted when respondents recall and report events in the past is known as ‘telescoping’. This is from the analogy with seeing things through a telescope as being closer than they really are. With typical back-reference periods where the later boundary is the present, telescoping causes the respondent, without realising it, to recall events that actually occurred outside the reference period as more recent than they really were and hence as falling inside the reference period.  In principle, the two effects of forgetting and ‘telescoping’ are countervailing in terms of the number of instances recalled, but of course there is no guarantee that the instances of forgetting, and the types of respondents prone to forget, are balanced by instances of ‘telescoping’ and the types of respondent who are prone to ‘telescope’.

Challenges of posing questions about behaviour and events

Respondent psychology and the resulting patterns of response discussed above give rise to particular challenges in posing questions about behaviour and events, in such a way as to gather valid, reliable, unbiased and discriminating data.

One problem relates to defining the behaviour of interest in a consistent and uniform manner. For some types of behaviour there is a ‘natural’ quantifiable, standard ‘unit of behaviour’ to ask about. For example, for smokers ‘a cigarette’ or ‘a pack of 10 (or 20) cigarettes’ fills this function. For use of health or social services ‘a consultation with a doctor’ or the equivalent is also useful, as is ‘a visit’ for use of leisure or shopping  facilities. These are not perfect solutions: concepts such as ‘consultation’, ‘visit’, ‘trip’ are much less standardised as a quantity than a pack of 20 cigarettes (e.g. does a consultation include a telephone call for advice; does a visit to the shops on which no purchase was made count?). But even ‘a pack of 20 cigarettes’ is not an ideal standard measure because cigarettes vary in their tar and nicotine content. 

For other forms of behaviour, defining a standard unit of behaviour can be much more problematic. For example, for measuring diet, the ideal, pursued by elaborate and costly dietary surveys, is to obtain weighed food intake reports classified and analysed for conversion into nutritional values. In terms of this ideal, asking respondents questions such as ‘How many times in the last seven days have you eaten fried food?’ is extremely crude, because of the vagueness and weak quantification of ‘a time’ as a unit of behaviour, the lack of information about the frying fat or oil used and the difficulty of defining, in terms that respondents can recognise and understand, what is meant by ‘fried’ (e.g. does it include stir-frying or stews in which the ingredients are fried before further cooking?). However, if the researcher thinks that it will still be analytically useful to have a crude measure – say one which divides respondents into those who report never eating fried food, who sometimes eat it and those who regularly eat it – then the data provided by the above question may serve the purpose.  Lateral thinking might suggest asking about food purchases over a reference period, rather than directly about food intakes. The question might also be improved by defining more clearly what is meant by ‘fried food’ (i.e. does it include stir-frying, pre-frying prior to using another method of cooking etc.? Does it include fried food purchased and / or eaten outside the home, or just home cooking?).

A second crux in producing standardised measures of behaviour is whether to pass to the respondent the task of providing frequency estimates, or to use a standard reference period. The first option typically involves asking questions containing words like ‘How often...’ or ‘How much...’ and response scales such as ‘Every day  / At least twice a week  / Once a week’ etc..  The second option involves asking respondents to recall how many times they have done the behaviour, or how much of the behaviour they have done (alternative methods of quantification), over a specified ‘back reference’ period appropriate to the behaviour (for hospitalisation or visits to the theatre it might be ‘the past 12 months’, for taking strenuous exercise ‘the past two weeks’).

Some points to consider about reference periods are summarised in Table 7.

In defining the most appropriate reference period, trade-offs must be made between:-

  1. a.the statistical desirability (in many applications) of using as long a reference period as possible, because this provides a larger and more reliable sample of each respondent’s behaviour and more instances of the type of event or behaviour of interest;
  2. b.the need to limit the burden on the respondent, so as to encourage co-operation and to ensure that the recall tasks imposed are ones that the respondent is actually able and willing to perform satisfactorily.
  3. c.the need to obtain records from as many cases as possible, so as to represent efficiently the distribution of the occurrence of events or behaviour across the survey population.

Table 7 Reference periods in measuring behaviour

Why do we need reference periods?

To standardise the estimate of behaviour across individuals

To enable rates and population totals to be calculated

To classify individuals reliably with respect to the behaviour

The reference period is a sample of individual behaviour

What is the ‘population of behaviour’?

How do we cope with variability?

What are the criteria for choosing a reference period to get good estimates?

Captures enough instances of behaviour

Gives a good sample of each respondent’s behaviour

Constraints on reference periods that can be used in data collection

Behaviour unmemorable

Errors of recall - omission and telescoping (misplacing in time)

Limits on respondent motivation

 

In selecting a reference period, careful thought and judgement on the part of the survey designer is therefore needed to strike the best compromise between possible aims in analysing the data obtained.

For example if, within fixed overall resources, the prime aim of the survey is to provide an overall picture of time use in the population of interest, it is better to obtain behavioural records for a short period for the largest possible number of sample members (point c above), than to obtain much longer records for a smaller sample of individuals (point a above), even though the total number of days on which activities are recorded by the sample members in aggregate may be the same. Because the behaviour of any given individual tends to follow a similar pattern from one day or week to the next, we learn more about behaviour in the population as a whole by selecting another individual than by taking twice as large a sample of the behaviour of the same individual.

On the other hand, if an important aim of the survey is to classify each sample member in terms of their (normal or typical) behaviour, then a larger sample of that behaviour for each individual (i.e. a longer reference period) is highly desirable. But the researcher’s ambitions in that direction will need to be set against the consequences for rates of response and data quality of imposing too heavy a response burden upon the individual respondent. 


To avoid bias the reference period to be used must be fixed by the researcher and not selected by the respondent. It must also be stressed to respondents that they must answer in terms of what actually happened in the reference week, even if that was an atypical week. This seems rather perverse to many respondents because they assume that a reflection of the typical behaviour of each sample member must be what is needed. But if all respondents tried to report their ‘typical’ behaviour in (say) a labour force survey, situations such as being temporarily unemployed, off sick or on holiday would be under-represented.

Another important factor to take into account in choosing a reference period is whether there is any natural periodicity in the occurrence of the events and behaviour of interest. In many cases there is, because many aspects of our lives are affected by weekly cycles of paid employment, domestic and leisure activity, by seasonal changes in the weather and the length of daylight, and by the occurrence of holidays and festivals. For this reason the reference period that will be most efficient in capturing the variability in behaviour (i.e. that provides most statistical information per day of recording) is one which takes account of the dominant activity cycle. There are many weekly activity cycles in time use, so a complete week might be chosen. However, within weeks the main differences in activity patterns for most individuals tends to be between weekdays and weekends. A good statistical choice might then be a reference period consisting of one randomly selected  weekday and one randomly selected weekend day. If that is too complicated to arrange, then one complete week of recording would probably be the best choice. But a survey with a one-week reference period would not capture seasonal patterns of behaviour unless continued for a whole year, or repeated in each season of the year.

The form and wording of questions to measure behaviour

In Table 8, we summarise the advantages and disadvantages of different question forms for behaviour questions.  The choice of question form is not trivial – different forms will elicit different information, and the choice must be informed by the objectives of data collection and the likely threats to data quality.  The differences are perhaps best illustrated by a series of example questions on the same topic – cinema attendance.

Table 8 Question forms for behaviour questions

Question form

Advantages

Disadvantages

Have you ever / about how many times have you (done behaviour)?

Easy and quick to ask.  For memorable, infrequent, easily defined and identified behaviour responses may be quite reliable

Does not tell us how much of behaviour currently goes on.  For frequent, unmemorable, not easily defined and identified behaviour, frequency estimate may be very unreliable.

On what date / how long ago did you last (do behaviour)?

Fairly easy and quick to answer.  Gives a crude estimate of average frequency over time of behaviour X in the population.

If frequency patterns are uneven or differ between individuals, comparisons will be distorted – this approach will give a biased representation of the relative positions, in terms of amount of behaviour done, of persons with regular and persons with irregular habits. For frequent, unmemorable, not easily defined and identified behaviour, frequency estimate may be very unreliable

How many times per (time period) do you usually / on average (do behaviour)?

Easy and quick to answer.  Gives a crude estimate of average frequency over time of behaviour X in the population.  Adequate for behaviour which follows a regular pattern

Either assumes that respondents interpret terms such as ‘on average’ consistently or in the same way as the survey designer or assumes that respondents know how to calculate an average and can do so accurately in their heads.  Neither assumption is correct, so estimates of frequent behaviour will be biased and inconsistent.  Tends to invite ‘conventional’ or ‘self-image’ answers which are systematically biased. Does not tell us anything about the pattern of individual behaviour over time.

 


Table 8 Question forms for behaviour questions (ctd)

Question form

Advantages

Disadvantages


How many times over the past (reference period) have you (done behaviour)?

Usually fairly easy and quick to ask, though not necessarily to answer.  This approach standardises by using a common reference period and is good for quantification if appropriate reference period is chosen. It gives an estimate of the frequency  with which individuals have done the behaviour over recent time and thus of recent frequency in the population.

Except for inherently very memorable and dateable behaviour, the reference period needs to be short to minimise forgetting and telescoping; hence data for rare behaviour will be sparse and ‘lumpy’.  For frequent, unmemorable behaviour there will be failures of recall even for short reference periods. There may also be context effects if behaviour varies (say) with time of year or place.

What is the total (amount / money value / duration) of (behaviour) that you have done within (reference period)?

This is often the exact type of information that the client requires – e.g. ‘What is the total amount of absolute alcohol consumed by members of the population over the past week?

Questions are often not readily askable in this form and have instead to be decomposed into a much more extensive and elaborate questioning routine.  Does not tell us anything about the pattern of individual activity over time.  Does not tell us anything about the circumstances in which the behaviour occurred.

Can you now please think back over the week ending last Sunday. On Monday of that week, how many times did you (do behaviour)? (Or other quantification such as ‘How much of (behaviour) did you do?’)

Repeat for other days.

Produces measure of recent frequency both for individuals and the population.  Captures individual pattern of behaviour over time.  Gives a basis for calculating aggregate measures (e.g. total consumed).  Can capture circumstantial evidence (who with, where etc.)

Involves a whole sequence of questioning and is therefore time consuming. In self-completion questionnaires, may be over-burdensome, leading to poor response rates.  Depends on respondent’s ability and willingness to recall and report in detail with accuracy. 

Version 1 Do you ever go to the cinema?

This is a conversational style question that implicitly requires the respondent to generalise about her or his behaviour in the past. However, it gives no explicit clue as to the duration of the retrospective period to which the response should refer – how should the respondent answer if s/he used to go to the cinema in the past but no longer does so?  Apart from this, it is quite easy to answer, but the information it yields is, in itself, of limited analytic use. It tells us nothing about how often the respondent goes to the cinema – yet most important quantitative applications require a measure of frequency of behaviour. Questions such as the above may therefore asked in surveys more by way of introducing a topic than as a means of obtaining the key information.

Version 2 When did you last go to the cinema?

This is a more demanding question to answer. It requires respondents to

  1. a)search their memories for instances of going to the cinema;
  2. b)identify the most recent instance;
  3. c)locate this instance in calendar time.

In this case,  step (a) may not be too difficult for most respondents, but step (b) and particularly step (c) are likely to be harder. When researchers ask such questions, they often tacitly recognise the difficulty of step (c) in particular by asking only for an approximate date, generally by using a bounded and defined response scale such as

Within the past week / 1 to 4 weeks ago / 1 to 6 months ago / 7  to 12 months ago / More than a year ago.

Although memory problems may occur – omission and telescoping  – such a question may obtain reasonably accurate answers. If the purpose is to divide respondents into broad groups in terms of frequency of cinema attendance, it may serve quite well. However, it does not permit the estimation of a population frequency distribution, since it does not measure the number of times that each respondent went to the cinema over a standard reference period – those making one visit in a given time interval would respond in the same way as those who attended twenty times in that period.


Version 3 Did you go the cinema last week?

This form of question about behaviour gives more focused and comparable information than Do you ever go to the cinema? because it uses the standardising device of a reference period (here ‘last week’) to which all respondents are asked to relate their answers. Such questions may be unproblematic for most respondents to answer if the reference period is short. However, like version 2, they do not satisfactorily capture frequency, since those who answer ‘Yes’ may have been to the cinema once only or several times over the reference period. Also, because the reference period is so short, it cannot distinguish between those who never go to the cinema and those who just happened not to go last week. 

Version 4 How often do you usually go to the cinema?

The wording of version 4 seeks to establish what individuals do typically, usually, or on average.  The aim of questions of this form is to obtain an approximate impression of how frequently the respondent does the behaviour of interest. However, the data provided by the use of these ‘vague quantifiers’ are not suitable for analysis in terms of numbers of occurrences of the behaviour, because they do not provide a complete enumeration of instances of the behaviour over a reference period. Instead, the respondent is left to interpret for him or herself (a) over what period the generalisation is to apply and (b) what ‘typical’, ‘usual’ or ‘normal’ corresponds to in terms of numeric frequency. 

The term ‘on average’ can have a precise meaning (arithmetic mean), but in most survey practice it must be regarded as another ‘vague quantifier’ term. Most respondents do not distinguish clearly between a measure of the average and an impression of what they ‘usually’ or ‘typically’ do , but what a person considers to be his typical  frequency of cinema-going  may not correspond to what has recently been his usual frequency and neither may correspond to his average frequency. In any case, few respondents are both able and willing to calculate an arithmetic mean correctly in their heads, particularly where no reference period is specified. Where behaviour is at all variable over time, the mean and the mode of the frequency distribution of their behaviour will usually not be identical. Therefore terms like ‘usually’ or ‘typically’ do not have any precise quantitative meaning that can be assumed consistent from one respondent to another and the above type of question is an inherently vague and inconsistent way of eliciting information on frequency.


Another inherent limitation of the data provided by such questions makes them unsuitable for many important types of analysis. For a survey user who intends to calculate, say, sample means and distributions of expenditure on cinema-going, or to use frequency data in regression models to predict cinema going, information on how often people ‘usually’ do things is not good enough. For such uses, an interval or ratio measure (i.e. an exact count) is required.

Data from questions using vague quantifiers may serve to construct broad classifications of respondents in terms of frequency of doing something (distinguishing, say, frequent, occasional and non-cinema goers), but even then there is a risk that allocation of individuals to categories may be seriously biased and/or error-prone because of differential interpretation of the vague quantifier terms.

Responding to such questions is also prone to psychological biases, such as social desirability bias, where respondents are influenced by a subconscious desire to suggest that they do more or less of the behaviour than they actually do.

Version 5 Over the last three months (or over the three months since [date]), how many times have you been to a cinema?

This form of question seeks information on the numeric frequency with which respondents have done the behaviour of interest over a defined reference period.  The request to enumerate instances of the behaviour of interest over a reference period dispenses with vague quantifiers or calculation of means.  Thus, on the face of things, such a question provides objective information on frequency of behaviour. If the question works as intended, the analyst can aggregate the three-month samples of cinema-going behaviour recorded across the sample and can then estimate distributions, means etc. representing the behaviour over the reference period for the population from which the sample was drawn.  However, there is an implicit assumption here that respondents can and will do the hard cognitive work of recalling, enumerating and counting instances of the behaviour accurately over the reference period. What generally seems to happen in practice is that respondents, unless specially motivated to do this hard work, resort to various low-effort guessing algorithms (heuristics) that will provide a reasonable and plausible answer while avoiding cognitively demanding recall and enumeration tasks. For example, when asked to recall instances of some behaviour over a year, they may think back over the past month, assume it is typical and multiply by 12 (not always accurately!). Krosnick (2000) has termed this tendency to minimise the amount of cognitive effort devoted to answering survey questions ‘satisficing’. Not surprisingly, when it is possible to check, the responses produced by ‘satisficing’ often perform poorly against quality criteria such as validity, reliability, completeness, precision, freedom from bias etc.

Enhancing accuracy and completeness of response

Experiments by survey researchers have shown that respondent performance can be much improved in terms of accuracy and completeness by careful use of particular interviewing techniques. These involve:-

  • stressing the importance of accurate data;
  • rewarding efforts at accurate recall with approval, but withholding approval when the respondent appears to give up too easily or to drift away from the point of the question;
  • giving more encouragement or trying further non-directive probes to encourage further effort.

In following this approach, it is very important for the interviewer to avoid appearing to encourage or approve particular responses (which could cause bias), but to make clear that what is approved is an effort to recall accurately. Instead of a single question and answer (with no indication of how the answer was arrived at), this mode of questioning involves multiple interventions by the interviewer and multiple responses by the respondent.

In postal and other self-completion questionnaires, this iterative process is not possible.  However, it can be stressed in instructions that precise answers are needed  to particular questions (the incentive effect is, of course, reduced if this appeal is applied to all  the questions).


Asking about assessments, attitudes, values, motives, intentions etc.

Social surveys also commonly include questions on attitudes and other non-factual or ‘psychological’ variables.  Table 9 summarises why such data might be collected.

Table 9 Why survey researchers ask questions to measure psychological variables

Aims of questions

Applications

To measure and understand public opinion

How popular is the Prime Minister? 

Are the public becoming more tolerant of abortion?

To monitor trends and patterns in opinions and attitudes

Are attitudes towards law and order becoming more lenient or more hard-line

To predict behaviour

Predicting voting behaviour

Predicting consumer choice

Predicting (health) service use

To influence behaviour through attitudes

What attitudes pre-dispose children to smoke? How might these be changed?

To understand motivation

Why do the long-term unemployed fail to take up training opportunities?

‘What if…’ issues (hypotheticals)

If the price of fuel to private motorists were to be raised by 50%, what would be the effect on the use of public transport?

To gauge satisfaction, with services, jobs, life aspects etc.

How satisfied are benefit claimants with the service they receive from the DWP?

To understand the formation of attitudes

How do social norms and personal values affect attitudes?

To gauge levels of knowledge

How well informed are people about a particular issue or topic?

 

Types of question used to measure psychological variables include:-

  • knowledge (e.g. ‘Could you please name five countries which use the Euro?’);
  • intentions (e.g. ‘If there was a general election tomorrow, what party would you vote for?’);
  • beliefs (e.g. ‘Do you believe that HIV can be contracted through giving  blood?’);
  • reasons (e.g. ‘Why did you decide to stay on at school after your GCSEs?’);
  • opinions (e.g. ‘In your opinion, how successful has the present government been in reducing hospital waiting lists?’; ‘Should the age of consent for sex be lowered to 16 for gay men?’);
  • attitudes and values (e.g. ‘What are your priorities in choosing a holiday destination?’; ‘People who break the law should always be punished severely.  Do you agree or disagree?’).

Actually, there is no sharp boundary in questionnaire survey practice between psychological variables and other types of variable such as those measuring attitudes and behaviour, since almost everything depends to some extent on the respondent’s interpretation of the question.  We are reliant on what respondents tell us for nearly all of the information collected and in practice apparently objective data on attributes and behaviour may be just as hard to measure in a valid, accurate and reliable manner (as we have discussed above).  Nonetheless there are particular considerations to be taken into account in asking questions about knowledge and attitudes.

Optimising response to knowledge questions

It can be argued that it is inappropriate to ask knowledge questions in self-completion surveys, since the respondent could consult others or use documentary sources (thereby providing an invalid measure of what they truly knew). However, if it is felt that knowledge questions should be included, the following strategies should be considered:-

  • Pitch the question at the appropriate level of difficulty for the aims of the survey.  Questions requiring a dichotomous (e.g. ‘Yes / No’) answer tend to be easier than more complex multiple choice questions or open-ended questions.  For issues that are likely to be ‘new’ to the respondent, simple questions may be more appropriate.  However, questions that are either too easy or too hard will be undiscriminating across different underlying levels of knowledge.
  • It may be appropriate to use filter questions to screen out respondents who lack insufficient information or experience (e.g. ‘Have you heard or read about the proposed changes to child benefits?) – going on to ask detailed questions of people who have no basic knowledge will lead to spurious accuracy in responses to those detailed questions.  Of course, in self-completion questionnaires, where the entire questionnaire can be previewed, there is a risk that a respondent will give a negative response to such a filter question, simply in order to skip the subsequent questions and thereby reduce the burden of responding.

  • If it is anticipated that a knowledge question will pose a threat to the self-esteem of the respondent (and may therefore result in item omission or falsification), consider posing the question as an opinion question (e.g. ‘In your opinion, what are the symptoms of bowel cancer?’).  Alternatively threat may be reduced by prefacing the question with ‘Do you happen to know…’ (e.g. ‘Do you happen to know who is the Chancellor of the Exchequer?’)
  • To minimise bias due to guessing, include a ‘Not sure’ or ‘Don’t know’ response category.
  • If numerical answers are needed, ask open-ended (e.g. ‘What is the current base interest rate?) rather than multiple choice questions; this reduces the bias that arises from a common tendency to opt for the apparently safe ‘middle’ response.
  • Use pictures and other non-verbal cues as well as standard questions (e.g. photographs of politicians, pictures of advertisements or products).
  • There are a number of techniques to control for guessing, which is a significant threat in knowledge questions.  These include: asking for additional information (e.g. what the person who is the subject of the question does as well as their identity); including ‘sleeper’ questions (e.g. asking about knowledge of a fictitious issues); and asking several questions on the same topic.  However, it may be impossible to totally eliminate guessing. For this reason, in computing knowledge scores across multiple items, a technique of ‘negative marking’ (in other words, computing a net score of the difference between the percentage of correct and percentage of incorrect responses endorsed) may be used.

 Obtaining valid and reliable measures of attitudes

Attitudes do not exist in the abstract; they are attitudes about something or someone.  This may be something or someone specific (e.g. genetically modified foods, Tony Blair) or about a more abstract concept or issue (e.g. local government).  It is important to be clear about what the attitude object  is.  There are three components of attitude: the cognitive – what the respondent knows about the attitude object; the affective / evaluative – the respondent’s disposition in favour or against the attitude object; and disposition to action – the respondent’s willingness or intention to act in relation to the attitude object.


In measuring the strength of attitudes, it is most common to have response categories which capture the strength as well as the direction of the attitude (e.g. ‘strongly agree, agree, uncertain, disagree, strongly disagree’). But the direction and strength dimensions can also be addressed in separate questions (e.g. ‘Do you approve, disapprove or have no opinion either way about …?’, followed by a question asked only of those who approve or disapprove ‘How strongly do you feel about …?’, with response categories of ‘Very strongly, strongly, not very strongly’).  Another approach is to ask multiple questions, each tapping different aspects of attitudes towards a particular issue, and to combine responses to the individual questions (e.g. by summation) to derive an overall measure of position on some dimension of attitude.

Questions about attitudes and values are often particularly prone to question wording effects (two wordings thought to be equivalent in meaning may in fact produce different response distributions). This is a particular danger when weight is to be placed upon assessments, made by members of survey samples, of their level of ‘satisfaction with services received’ and the like.

To achieve standardisation of response as between respondents, the researcher may present a concept or statement and asked to respond by choosing one of a set of labelled scale points (for example, ‘Very satisfied / Fairly satisfied / Neither satisfied nor dissatisfied / Fairly dissatisfied / Very dissatisfied’). Inexperienced survey designers often focus on issues such as ‘What is the correct number of scale points to offer?’  There is no standard answer to this, beyond saying that, for things which are important to them and central to their day-to-day experience, respondents can generally make finer meaningful distinctions than they can for issues about which they have, at best, vague and imperfectly formed views. In practice, it is unusual to use more than ten points and five are commonly used. Occasionally respondents may be asked to answer in terms of a ‘thermometer scale’ with 100 graduations, for example in assessing their own health, but the ground is usually prepared for this through ‘priming’ questions to ensure that all respondents are considering the whole domain of ‘health’ intended by the researchers.

Getting the form of the question and the presentation of the response categories right is generally more important than whether to use, say, five- or seven-point scales. However, the response scale must be seen as part of the question and whether to give the points labels at all is an issue that needs to be considered, given that respondents may be influenced by extraneous overtones of words and concepts such as (say) ‘I am satisfied’, rather than (say) ‘The service was excellent’.  Another important principle is to make it clear to respondents what are the poles of evaluation between which they are required to place themselves.

A device which addresses a number of these concerns is the format in which only the poles of the scale are labelled, as below.

Could not be better Could not be worse

 

Question wording effects occur particularly when respondents are asked to make judgements: about matters to which they have previously devoted little systematic thought and have no pre-formed views; where the concepts embodied in the question are vague and ill-defined; or where the words used in the question have evaluative overtones. One type of effect which occurs in these circumstances is the tendency of respondents to play safe and avoid impoliteness by choosing a neutral or mildly positive response from a range of labelled response categories offered (for example fairly satisfied).  A related effect is respondents’ tendency to avoid endorsing statements containing drastic or extreme-sounding expressions (such as forbid) where a less extreme sounding, though logically equivalent, expression (such as not allow) would have been endorsed.

In general, if evaluations are required, the more concrete and close to the respondent’s experience the concept offered by the question, the more likely it is that responses will be robust to extraneous wording effects. Thus, in a survey of inpatients’ experiences, a question about whether hospital meals were too cold, too hot or the right temperature when served will be less prone to unwanted response effects than more general and abstract questions about ‘satisfaction with hospital catering’. This response behaviour tends to create a tension with the desire of some survey sponsors and questionnaire designers to obtain overall or global assessments of the topic under investigation. It is possible to build up to overall assessments by asking about and analysing the responses to more concrete assessment questions, but a large number of such questions may be required to cover the domain.

Another potential source of bias in attitude questions is the ‘missing alternative’ – in the absence of an explicit option, respondents tend to implicitly supply their own (varying) alternatives.  In one German survey, 19% of women said that they would not like to have a job outside the home when asked ‘Would you like to have a job, if this were possible?’ while 68% said they would not like to have a job when the question contained the explicit alternative ‘Would you prefer to have a job, or do you prefer to just do your housework? .

Question forms

A major distinction in designing survey questions is between ‘open’ and ‘closed’ questions. Open questions invite a verbatim response from the respondent and should be phrased in a way which does not encourage the respondent to assume that one answer is more acceptable than another.  Closed questions present the respondent with a list of possible answers, from which s/he is asked to choose either just the one that best applies, or as many as seem applicable. Quite often question designers attempt a compromise by adding at the end of the pre-specified responses ‘Other answers – please specify’. Each type of question has its strengths and limitations. To a large extent the strengths of closed questions correspond to limitations or weaknesses of open questions and vice versa (see Box 3). 

In a survey of patient satisfaction with hospital services, an open question might read:

In what ways do you think this hospital’s services to patients need to be improved?

followed by a space, with or without ruled lines, on which answers are to be recorded verbatim.

The corresponding closed question might read:

In which of the following ways do you think this hospital’s services to patients need to be improved?

and would be followed by a list of  pre-specified response categories, which give an indication of the types of response that the researcher thinks are relevant.  These pre-specified response categories should be viewed as part of the question, since the selection of categories gives meaning to the question and indicates the frame of reference that the researcher has in mind.
Box 3 Open / closed questions

Open questions

Closed questions

Open questions avoid imposing the survey researcher’s perspective on respondents. If the researcher does provide a list (i.e. uses a closed question), respondents will focus on the listed items and not think of other possible items or frames of reference.

But closed questions focus the respondent on aspects that are relevant to the research. Without prompting, respondents may find it hard to think about vague and undefined domains (for example, ‘this council’s services to residents’).

Without a list of responses to choose from, different respondents may interpret the scope of the question differently. Only if the full range of responses is prompted can we be sure that all respondents have considered all the relevant response options.

Respondents will hopefully come up with open responses that reflect the issues of most importance to them.

But in practice less articulate respondents tend to give short, uninformative answers to open questions which do not necessarily do justice to more complex underlying motives etc. Many will put down the first answer that comes to mind and pass on, possibly failing to consider the full range of aspects that the researcher had in mind.

Open questions allow for unexpected responses. But it can be difficult to keep respondents focused on the issues which are of most relevant to the research (for example, an open question intended to elicit suggestions for improving a hospital’s services to in-patients may elicit responses about waiting times).

Closed questions with prompted response categories help to define the domain of interest to the research. But the listed answers may effectively steer the respondent towards certain types of response and away from others.

 

Open questions may provide vivid examples for inclusion in a report on the survey.

But it could be argued that this is really the province of qualitative, rather than quantitative, research.

Respondents often find answering open questions hard work and tend to feel that one answer (for example, mention of one reason for doing something) is enough.

Closed questions make it easy for the respondent  to choose several different answers if (s)he so wishes. However, the choices will be limited by the options offered. This applies even if an ‘Other answers – please specify’ option is provided.

Coding responses to open questions is time-consuming, laborious and unreliable (different coders may assign the same response to different categories).

Closed questions are precoded, so the data are classified (a form of standardisation) at source.

 


In the example below, the final prompted phrase of ‘Other answers (please specify)’ is an attempt to elicit spontaneous answers not covered by the pre-specified categories.

In which of the following ways do you think this hospital’s services to patients need to be improved?

Version A

Version B

More nurses on the wards 

More experienced doctors 

Shorter treatment waiting lists 

.

.

.

Other ways (please specify)

Wider choice of menus 

Better quality of food 

Longer visiting hours 

.

.

 

Other ways (please specify)

 

It might be thought that respondents presented with Version A would make heavy use of the ‘Other specify’ option to give responses of the kind shown in Version B, and vice versa, but in practice this does not happen. The provision of a pre-specified list, even with an ‘Other answers’ category,  tends to direct respondents’ thinking towards the domain represented by the pre-specified categories (in the case of A, major changes to the resourcing of the National Health Service; in the case of B, relatively minor improvements  to the hospital’s hotel and catering arrangements). Effectively, A and B are different questions!  Self-confident respondents with strong and well-formed views on a topic may give their views regardless, but most respondents tend to look for clues in the question wording or format as to what kinds of answers are expected.

For this reason, it is important to understand that open questions are not just non-leading versions of closed questions, but are essentially different. Repeated experiments have shown that, if the range of pre-specified responses used in a closed  version of the question is used to code open responses obtained from an equivalent sample of respondents, very different distributions of coded responses can result.

It is also important to recognise that in self-completion questionnaires, responses to open-ended questions are typically less full than in interview surveys. Moreover, better educated, articulate respondents tend to write more than the less educated (a subtle bias).  The amount of space provided on the questionnaire for the response is taken as an indication of the level of detail required. With larger samples, the time and labour required by office coding (covered in Chapter 13) is an important factor favouring the closed form of question.

It is recommended that open questions be kept to a minimum in self-completion questionnaires addressed to the general public.  Nonetheless, carefully focused open questions may have an important role, for instance in surveys of professionals who have well-developed views in their area of experience and expertise.

Question wording

Choosing between open and closed questions is not the only decision facing the survey researcher.  Other issues can be illustrated using the common case of questions intended to measure the frequency with which the respondent performs a particular type of behaviour. This may well be a ‘cognitively demanding task’ for the respondent – that is, it involves hard mental work by way of applying definitions, recalling and counting. In everyday conversation it is not normal to impose such tasks on the person you are talking to without warning. Instead, the parties often ‘negotiate meaning’, with the respondent first assuming that a quick and easy answer will suffice and the questioner ‘unfolding’ what s / he really wants to know through a series of ‘turns’ .

For example:-  Have you had a holiday this year?’

‘Well, we went to France for a fortnight over Easter.’

‘How about shorter breaks?’

‘Um, we had a long weekend in Scarborough in January.’

‘How about trips to visit friends or relations?...’ 

 

In designing a questionnaire, particularly for self-completion, the survey researcher must somehow focus on the precise type of information required, without several ‘turns’ of conversation. But respondents are not used to having to digest precise and elaborate definitions before attempting to answer a question. This is one of the things that makes the wording of survey questions difficult. 


The following problems are common:-

  • Respondents tend to revert to definitions which are familiar to them but do not necessarily correspond to the survey concept. For example, respondent and researcher definitions of ‘your family’ may differ - does it mean those living with you, or does it encompass the extended family?
  • Lay people and professionals may use words and concepts in quite different ways, or may not share a common vocabulary.  For example, to a health professional the word ‘chronic’ implies a long-term, ongoing health problem, but patients may interpret ‘chronic’ as ‘severe’, or ‘very painful’.  In a survey on women’s health, when asked ‘Do you have any problems with your menstrual cycle’, one respondent replied ‘I used to, but I’ve got a car now’!
  • If forced to do mental arithmetic, many respondents will guess or make serious mistakes. For example, as recognised in the section on measuring behaviour, many adults cannot add up accurately and do not know how to define or calculate an average (arithmetic mean) or a percentage.
  • The terms ‘on average’, ‘usually’, ‘normally’ are usually interpreted by the lay public to mean much the same thing. If behaviour is very regular, that may suffice. But if behaviour varies over the period you are interested in, ‘normally’ and ‘on average’, for example, should imply different answers. If you require accurate numeric information, it is best to collect it by getting the respondent to enumerate instances over a reference period and doing the summarising yourself in the office.
  • There are limits on how hard respondents are prepared to work. But if they are unable or unwilling to perform the cognitive tasks imposed by a question, they seldom say so. Instead they try a ‘short cut’ answer that requires little thought to see if it is acceptable. Therefore, if a precise and accurate answer is needed, you need to make a special point. For example, if you are interested in the number of times on which the respondent has used leisure facilities, you may need to ask:

Please include all occasions on which you attended an exercise class AND all occasions when you used the exercise equipment independently. Do not include occasions when you visited only the bar or café.


  • If asked how many times they have (say) suffered minor symptoms over past six months, few respondents will mentally enumerate the episodes.  Instead, they may assume that ‘last week’ was typical to work out a 6-month estimate (often making arithmetical errors!). Such short cuts may give seriously inaccurate or biased results.
  • Responses given are influenced by the choice of responses offered, especially if the concepts used in the question are vague. For example, responses to A below will produce more ‘2 times or less’ estimates than responses to B. This is because respondents will not really know what the researcher means by ‘problems with neighbours’ and will seek clues from the response categories offered.  In this example, the response categories offered in Version B suggest that ‘problems with neighbours’ are common, and therefore probably include quite trivial matters.  Version A, by contrast, suggests that ‘problems with neighbours’ are rather unusual, and therefore focus the mind on more serious disputes. Asking respondents to specify a number avoids bias arising from the response categories, but ‘problems with neighbours’ will still be differentially interpreted and needs to be defined.

How often have you had problems with your neighbours?

 Version A

Version B

Never

Once

Twice

More than twice

2 times or less

3-6 times

7-10 times

More than 10 times

 

Guidelines for wording and presenting questions include the following.

  • Study how the people you are addressing speak and use appropriate language.
  • Use simple, common words.  For example, ‘How often has this kind of thing happened?’, rather than ‘How frequently have incidents of this type occurred?
  • Keep the question short – a sentence of less than 20 words approximately is desirable. Two short and simple sentences (a preamble to clarify the reference of the question and then the question itself) are usually better than one long and complex one.

  • Avoid complex grammatical structures such as multiple or qualifying clauses (for example ‘Apart from…’, ‘Including…’, ‘Although…’). 
  • Avoid questions which are insufficiently specific. Often areas of vagueness only become apparent through piloting and testing questions.
  • Avoid over-generalised or ‘catch-all’ questions – such as Overall, how would you rate the service provided to you by the city council in the past year?  Respondents cannot easily summarise a series of disparate experiences or concepts; here it would be better to ask a series of specific questions about different aspects of the services – for example, refuse collection, library service etc. – with an appropriate rating scale for each.
  • Avoid ambiguity in wording and syntax – for example Has your child mentioned this problem to his/her teacher? IF SO What did he/she say?). Often such ambiguity is identified only through piloting.
  • Either avoid or define vague words and those with more than one meaning.  Words that have multiple meanings may be variably interpreted by different social groups (for example, ‘dinner’ alias ‘lunch’, ‘training’ alias ‘sports practice’ or ‘vocational education’, ‘book’ alias ‘magazine’), thus posing threats to reliability.  These are often identified only through piloting. If a vague term or word (for example, children) is used, define it in the context of the survey objectives (for example, children under the age of 16 years).
  • Avoid jargon and technical terms, including acronyms and abbreviations. As with vagueness, problematic technical terms are often identified only through piloting. Bear in mind that words commonly used and understood by professionals (for example, chronic, assessment, module, episode, semester) may be misunderstood by lay people. If it is necessary to use such terms, define or paraphrase them.
  • Avoid double-barrelled questions  - double-barrelled questions and those containing two or more different concepts or propositions to which respondents may have different reactions (for example, The rent office should open earlier and remain open longer with response categories of  Agree / Disagree). If an ‘and’ or an ‘or’ creeps into a question, it may be double-barrelled – beware!
  • Avoid double negatives – in  particular a negative statement followed by a negative response – for example, Do you agree or disagree that people under 21 should not be allowed to own air-guns?. Such question and response category combinations often confuse people. In this example, some people will want to say No, people under 21 should not be allowed to own air-guns and may therefore go for the negative response of Disagree, whereas they should actually reply Agree (i.e. I agree that they should not be allowed…).
  • Break down complex questions into parts.  Give important definitions as preambles to the question.  For example, We are interested in all visits you have made to the Apex shopping centre in the two weeks ending yesterday.  These could have been to do your own shopping or to accompany someone else on a shopping trip.  Include visits when you just ‘window shopped’ as well as those when you bought something.  In the last two weeks, how many visits have you made to the Apex shopping centre?
  • Avoid proverbs and clichés when measuring attitudes. 
  • Avoid leading questions  in which the viewpoint of the survey researcher or respondent is revealed, or which suggest, however subtly, that a certain response is expected – for example, Do you agree that the NHS is under-funded?
  • Beware of loaded words and concepts  - loaded words are those implying a value judgement or carrying emotive overtones. Loading can also result from an imbalanced set of response categories – for example, a response scale of completely satisfied – very satisfied – quite satisfied – not satisfied.
  • Beware of presuming questions  - these are questions which assume that a respondent has indulged in a specific type of behaviour or possesses a specific attribute or piece of knowledge – for example, What rate of interest do you earn on your savings? (in fact, at least two assumptions are being made here; first that the respondent has interest-bearing savings, and second that s/he knows the rate of interest).
  • Be cautious in the use of hypothetical questions. People in general are poor predictors of how they would behave in novel situations.
  • Do not over-tax respondents’ memories by asking, for example, for detailed recall of unmemorable events or behaviour.
  • If possible, provide cues to memory in terms of some easily recalled event or date.  For example, Over the past four months – that is, since last Christmas – have you …?.  Note that this approach is somewhat easier to implement in interviewer-administered questionnaires, since the interviewer can tailor the prompt on the spot, to take account of variable dates of questioning.

Since, as we have seen above,  response options give meaning to the associated question, careful attention also should be paid to the construction of response categories. Consider the principles of choosing an appropriate scale of measurement (nominal, ordinal, interval, ratio). 

The following points should be considered:-

  • Use open-ended questions sparingly – they are more resource-demanding, and are more subject to between-interviewer and between-coder variability and therefore are less reliable.
  • Ensure that response categories are mutually exclusive (i.e. do not overlap) and collectively exhaustive (i.e. all possible responses are catered for); 
  • Present response categories in a logical order.
  • Take the mode of questionnaire administration into account when ordering response categories. “Primacy” effects, whereby respondents select the first response that seems applicable, without considering the full range of alternatives, may be more common in self-completion questionnaires (Krosnick, 2000). “Recency” effects, whereby respondents choose response categories towards the end of multiple choice lists are more likely to occur in interview surveys (Krosnick, 2000) because of fatigue effects and memory effects (the respondent is more likely to remember the last options read out by the interviewer).
  • If some responses are more socially desirable than others, start with the least socially desirable option. This tends to counteract primacy effects.
  • Limit rating scales (for example, scales to measure satisfaction or strength of attitude) to not more than seven points when written descriptors are attached to each point.  A ten point scale (or even an eleven point scale 0-10) may be feasible when only the end points are labelled.
  • Include a middle, neutral alternative unless there are persuasive reasons not to do so (e.g. if you feel that respondents must have a view one way or the other).
  • For more than five response categories, use numeric scales.

  • Consider analogues such as ladders, clocks or thermometers for numerical scales with many points;
  • Ask respondents to respond to every item in a list rather than indicating only those that apply (i.e. ask them to respond Yes / No, or Applies / Does not apply, to each item rather than simply complying with an instruction to circle as many as apply).
  • Allow for ‘Don’t know’ and ‘Not applicable’ responses if appropriate. If you do not, respondents will omit the item and you will not know whether it is a ‘Don’t now’, a ‘Not applicable’, an ‘Overlooked the question’ or an ‘Unwilling to reply’.

Question ordering

A questionnaire is not just a list of topics and questions. The survey researcher also needs to pay attention to the order in which questions are presented and to logical dependencies between questions and sections of the questionnaire. 

The respondent should be led through the questionnaire in a logical sequence, from the general to the specific, although the choice and  the sequence of topics may not be the same for all respondents.  In a self-completion questionnaire, the ideal is to have all respondents answer all questions.  However, an exception may be made to this rule where whole sections of a questionnaire do not apply to certain respondents. For example, there may be a section on employment for those who are economically active, which will be omitted by those who are not. Within the employment section there may be special questions for those who work in particular industries.

In self-completion questionnaires addressed to the general public, complex routing and complicated skip and filter instructions will confuse and put off some respondents. If the respondent is required to make difficult judgements or to get their head around unfamiliar concepts, lead up to these through definition and by stages.  Does not apply response options may seem attractive in avoiding explicit routing instructions, but can back-fire in self-completion questionnaires (some respondents may use ‘Not applicable’ in an idiosyncratic way or as a synonym for ‘No’).

In self-completion questionnaires, there is no interviewer to aid the communication process.  Therefore particular care needs to be taken with instructions and layout.


When new topics are introduced, there should be a re-orientation statement or instruction. For example:

Questions 20-25 are about the last time you visited the doctor. If you have not visited a doctor in the past six months, please go to Question 26.

Guidelines for question ordering include:-

  • Start with a brief introduction.
  • Place some easy, non-threatening questions first.
  • Don’t start with an open question requiring a detailed response.
  • Place questions about personal circumstances (e.g. age, financial situation, family situation) last, since they can be seen as threatening or intrusive. (Exceptions must, however, be made where they are required to determine the respondent’s eligibility to complete the remainder of the questionnaire.)
  • Give a rationale for including potentially sensitive questions (e.g. We need to ask these questions about income so that we can describe our study population).
  • Use ‘funnelling’ procedures to minimise question order effects, starting with the general and moving to the specific.  Knowledge may be an important part of the process of qualifying and quantifying opinions and attitudes, so it is preferable to ask knowledge questions first.  However, if respondents are primed by knowledge questions about what is the correct thing to do, they may be more likely to report that their behaviour is in line with what this ideal.  Therefore, it may be better to ask about personal behaviour before asking about knowledge of the related issue.
  • In questions about chronologically ordered events, move forward in time if the starting date is highly salient or memorable.  Otherwise, it is probably easier to move back in time.
  • Complete questions on one topic before embarking on a new topic and avoid visiting the same topic twice.
  • Use transitional phrases and instructions when switching topic or frame of reference (for example, Questions 8, 9 and 10 are about your most recent holiday).

  • Respondents use context to interpret questions, so define question sequences and pilot the questions in sequence and context.
  • Order filter questions (those intended to establish who should answer what questions) in such a way to cover all contingencies and encourage complete responses.
  • Place branch questions as near as possible to the filter stem.
  • Avoid complex, multiple filters (e.g. If you have answered No to question 8 and Yes to question 9, go to …)
  • Finish the questionnaire with a Thank you and instructions in what to do with the completed questionnaire.

Summary of key points

  • In devising questions, the survey researcher is faced with the dual agenda of gathering high quality (valid, reliable, unbiased and discriminating) quantifiable data, while communicating with respondents in an artificially constrained manner.
  • An understanding of respondent psychology – how respondents comprehend what is being asked of them, retrieve the required information from memory, formulate an answer and choose a response category – is essential in anticipating and minimising threats to data quality.
  • Posing questions about behaviour – especially when there is a reliance on recall of past events – poses particular challenges in terms of collecting valid, reliable, unbiased and discriminating data.  The exact aims of data collection – the type of estimate of behaviour to be made – need to be taken into account when choosing  the form of question.
  • Surveys often involve collecting data on psychological variables – knowledge, beliefs, opinions, attitudes, motivations, expectations etc.  Threats to validity, reliability, lack of bias and discriminatory power also occur in measuring psychological variables; the choice of question form and of response categories are crucial in eliciting high quality data.
  • Open and closed questions are essentially different.  Each has its strengths and weaknesses and they are not interchangeable.  Open questions pose particular difficulties in self-completion questionnaires and should generally be kept to a minimum for that mode of administration.
  • Questions and their response categories together convey meaning and cannot be considered in isolation from each other.
  • In wording questions, consideration must be given to the linguistic and cognitive abilities of the target audience.  Care must be taken to: use appropriate language; avoid double negatives; avoid undue complexity, ambiguity and  jargon; avoid double-barrelled questions; avoid leading, loaded, presuming and hypothetical questions.
  • Questions involving memory are particularly prone to response error, and techniques to minimise the risk of such distortion – for example, cues to memory – should be utilised if possible.
  • The mode of administration, and the cognitive abilities of respondents should be taken into account in choosing and presenting response categories.
  • The survey researcher also needs to pay attention to the order in which questions are presented, ensuring a logical flow through the questionnaire.
  • In general, topics and questions should be ordered from the general to the specific.
  • Complex routing and branching patterns are more easily handled in interviewer-administered surveys, and should be kept to a minimum in self-completion questionnaires.

Further reading

Fowler FJ Junior. Improving survey questions - design and evaluation. Thousand Oaks: Sage Publications, 1995. (Applied Social Research Methods Series - Volume 38) (Chapter 4).


Current guidance and methodological literature for 2026–27:

References

Jobe JB, Tourangeau R and Smith AF. (1993). Contributions of survey research to the understanding of memory. Applied Cognitive Psychology, 7, 567-584.

Krosnick JA. The threat of satisficing in surveys: the shortcuts respondents take in answering questions. Survey Methods Centre Newsletter 2000, 20, 4-8.

Sudman S, Bradburn NM and Schwarz N. (1996). Thinking about answers. The application of cognitive processes to survey methodology. San Francisco:  Jossey-Bass Inc

Tourangeau R, Rips, LJ and Rasinski K. (2000). The psychology of survey response. Cambridge:  Cambridge University Press.