In this chapter…
This chapter follows survey responses from raw collection through coding, checking, documentation and an analysis-ready dataset.
By the end of this chapter, you should be able to…
- design consistent identifiers, codes and missing-value conventions
- apply reproducible data checking and editing procedures
- produce a codebook and processing record that support reuse and audit
In this chapter
By the end of this chapter, you should be able to:-
- Explain how coding and data processing issues should be anticipated in questionnaire design.
- State what ‘pre-codes’ are, and explain why a standardised approach to the definition of pre-codes is desirable.
- State what ‘post-codes’ or ‘office codes’ are, and draw up a coding frame for responses to open questions.
- Explain how data from questionnaires should be entered into a computer for analysis.
- Identify key steps in data validation and processing.
Anticipating coding and data processing in questionnaire design
A common method of entering data from paper-based questionnaires to computer for analysis purposes is data keying (specialised agencies, including the university’s Data Preparation Service, do this) to computer disk. For this, each potential item of information must be assigned a location (field) in the computer output file. Keying operators need:-
- a keying design (record and column guide), printed on the document (often on the right hand margin);
- clear and consistent layout for keyable data items (e.g. coding column (more usually used in interview schedules rather than self-completion questionnaires), circled numbers, ticked boxes);
- numeric codes to key (generally indicated by circling a number indicative of a particular response category, or by ticking a numbered box corresponding to that category); note that ‘qualitative’ responses such as gender or strength of opinion need to have numeric pre-codes or office codes allocated to each response category;
- prior numeric coding of all verbatim items (e.g. from open-ended questions) with codes inserted in pre-specified locations.
An alternative to keying is electronic scanning of documents – optical mark recognition (OMR) or optical character recognition (OCR). This must be planned into the document design and printing, and requires appropriate computer hardware and software. Respondents must be told how to record responses (e.g. tick box using a black pen). Scanning is less restrictive on the positioning of items than conventional manual data entry. It is particularly appropriate for closed questions, where a box or bubble can be ticked or filled in. However, it is less good for ‘write-in’ answers; special arrangements (optical character recognition) need to be made to capture numeric and alphabetic characters. Questionnaires may need to be dismantled for page by page scanning and then reassembled and restapled for future reference and some completed questionnaires will fail to scan correctly.
Computer-assisted approaches – CAPI, CATI and CASI, as well as internet-based surveys – generally involve data capture directly to disk, though verbatim responses to open questions may still need manual office coding (see below).
Identification of fields, and cards / records
Computerised data files can be viewed as a matrix of rows (records) and columns. Each variable – a response to a given question – must be allocated one or more fields (positions) in the computerised data file, in which that variable will be recorded and stored. In other words, the observed value of that variable must be stored in the corresponding position or field of the data file for all respondents so that the computer can retrieve and analyse the data.
A field is a location or position within the matrix of the data file. It can be uniquely designated by a record number (designating the row) and column number(s). For many questions, only one response will be allowed, and so only one field will need to be allocated. For example, in a question asking about the gender of the respondent, the response options will be ‘male’ (allocated, say, the code 1) or ‘female’ (allocated, say, the code 2). To store this variable, a field comprising a single column will be sufficient. For other variables, even though only a single response is allowed, multiple columns will need to be allocated to record and store the value. For example, in a question asking about the age of the respondent, allowance will have to be made to store responses of 100 years or more. To store this variable, a field consisting of three adjacent columns will need to be allocated. For yet other variables, multiple responses may be allowed and sufficient fields will need to be allocated to store the maximum number of responses permitted. For example, in a question asking respondents to indicate whether they do or do not read each of ten newspapers (where a ‘yes’ response is coded 1, and a ‘no’ response is coded 2 in each case), ten columns will need to be allocated.
A card or record is one line or row of data as entered into a data file (the term ‘card’ harks back to early days of computing, when 80-column punched cards were used). When the questionnaire is designed and the fields numbered, each record should be a maximum of 80 columns long (while data analysis packages can cope with longer records, they cannot be easily viewed on the standard computer screen). Once 80 columns have been reached and a new record is required, it should preferably begin at the top of a new page of the questionnaire. Each respondent may therefore have more than one record.
The first field(s) in the record will normally be used for the unique survey identifier. This can be structured to denote a hierarchical arrangement of respondent within area.
In self-completion questionnaires, it is not customary to repeat the identification (survey) number on the pages corresponding to the second and subsequent records. Instead, data preparation staff should be asked to punch the appropriate digits at the time of data entry, so that each record on the data file is correctly identified. The instructions might read something like this (assuming a six-digit identification number):-
Record 2
columns
1 - 5 as columns 1 - 5 of record 1 - survey number from front cover
6 always the digit 2.
However it is recommended that the identification number should be written in a second place on each questionnaire (for example, somewhere on the back page). Occasionally respondents detach the front page or cover in an effort to afford anonymity; if the identification number does not appear elsewhere, this can lead to problems.
Pre-codes
These are explicit codes (response category labels) that appear on the questionnaire to denote the codes corresponding to the various response options for closed questions (e.g. 1 for male, 2 for female). Numeric codes are preferred to alphabetic codes, since they are quicker to enter (using the numeric key pad on a computer) and easier to process using standard statistical packages.
For similar reasons, consistency in assigning codes to particular answers is also desirable – for example, the response category ‘Yes’ may always be assigned the code 1 and ‘No’ the code 2. (Equally, a code of 0 could be used to denote ‘No’ and 1 to denote ‘Yes’; the actual number used is not in itself important, since for categorical data, one would not aim to calculate means etc).
In allocating codes, however, one might wish to give consideration to the ‘logical’ interpretation of the relationship between codes and the corresponding verbal descriptors. For example, in measuring a construct such as ‘satisfaction’, respondents may instinctively feel that a high number should be indicative of greater satisfaction.
Pre-codes for missing data
Clear and consistent principles for the treatment of missing data (e.g. where a question has been skipped either appropriately or inadvertently, or where a response of ‘Don’t Know’ or ‘Not Applicable’ has been given). A decision as to whether to leave the relevant fields blank (and therefore declare and treat blanks as missing data – an option available in statistical packages such as SPSS) or to assign an explicit missing value code will need to be made. This decision will depend on how the package used for data analysis handles missing data and blanks, and whether, in analysis, there is a need to distinguish between different types of missing data (e.g. accidental omissions versus explicit ‘Don’t Know’ responses).
Office coding open-ended questions
When open-ended questions are used, it is usually desirable to group similar responses together, using a coding frame or system, and thereby to allocate ‘post-codes’ or ‘office codes’ to the verbatim responses prior to data entry and analysis. The principles and procedures to be considered in developing a coding frame are shown in Table 14.
Quality control of coders
- If possible every questionnaire should be scanned before being sent for punching and a sample should have a detailed check.
- Coders should write details in pencil on the questionnaire or on a post-it note inserted at the appropriate point in the questionnaire regarding any problems they encounter in coding and bring them to the attention of the checker (often the survey manager).
- Difficulties and differences in coding should be identified as soon as possible and resolved between the coder, the checker and the responsible researcher. A record should be kept of all decisions made, by whom, why and when.
Checking and editing questionnaires prior to data entry
Even in the best-designed questionnaires, with clear instructions and good routing, some respondents will not answer the questions as intended – pages will be skipped, routing instructions and branching patterns will be ignored, responses will be marked in the wrong place (e.g. the verbal descriptor will be ticked or circled, instead of the corresponding numeric code or box), comments will be written by the question to explain or qualify the answer given. Where open questions have been used, responses will generally need to be coded as described elsewhere in this chapter.
For these reasons, questionnaires need to be carefully checked (and edited if necessary) prior to data entry. To ensure consistency, and to minimise the risk of coding error and bias, staff involved in checking questionnaires need to be adequately trained and provided with written instructions on how to deal with the types of mistakes that are likely to be encountered. The implementation of research governance requires that raw data be retained and made available to external bodies, to guard against scientific fraud and falsification of data. An ‘audit trail’ from questionnaire responses to ‘data as analysed’ should be maintained. To this end, to distinguish between respondents’ original answers, and corrections made by the data checkers and coders, a distinctive colour should be used for any changes and additions made at this stage of data processing. Since respondents typically use blue or black pen in completing the questionnaire, we recommend the use of red pen for these changes. We also recommend that a careful record should be kept of the rationale underpinning any decisions about changes to respondents’ answers (for example, if a ‘No’ response is altered to ‘Not Applicable’ because comments made by the respondent and / or answers to other questions suggest that this is indeed the more appropriate response category). Keeping records of common mistakes is also useful in informing the design of subsequent questionnaires and surveys.
Table 14 Principles and procedures in developing a coding frame
|
Principle: |
Codes should be exclusive and exhaustive. |
|---|---|
|
Designers: |
The frame should be designed jointly by those survey researchers who have designed the questionnaire and if possible, the coders who will be carrying out the coding. |
|
Sample size: |
A representative sample of between 5% and 10% of the questionnaires with a minimum of 20 responses and a maximum of 50. (Obviously, in a very small survey, the lower limit may need to be relaxed.) |
|
Method: |
|
|
2) List all responses to one question individually, preferably by typing into a file in a word-processing package. Group responses as appropriate bearing in mind the above principle of exclusivity and exhaustiveness, the question asked, the objectives of the research and the underlying theory. Define codes. Discuss with other designers. |
|
|
Testing the frame: |
At least two coders code the original sample blind, the designers and coders then discuss and compare. The coding frame is revised if necessary. Then code a further sample (of similar size) blind and revise again if necessary. |
|
Number of coders: |
May vary from one to a pool of coders depending on size of study etc. There is an advantage of consistency if there is only one coder |
|
Coding ‘other’ responses: |
A category of ‘other’ may be used to cover any responses that do not fit within the devised codes. However, the ‘other’ category should not have a large number of responses. All responses given the code ‘other’ should be listed during coding and the list checked regularly by the designers of the coding frame to assess whether more codes are required (often denoted by patterns within it). If necessary more boxes may be added to the end of one record to allow for the post-coding of ‘other’ responses. |
|
Office coding: |
Most respondents will use blue or black pen when filling in a questionnaire. To distinguish respondents' original responses from those coded in the office, we recommend that a red pen should be used for office coding and any other changes made during this phase (e.g. adding in explicit missing value codes). Tippex should not be used as changes need to be legible in order to understand the rationale behind them. |
Data entry
Many universities employ data entry staff. Commercial data preparation or data entry agencies are to be found in many towns and cities. These departments and organisations have staff who are highly trained and experienced in entering data from questionnaires to computer; the quality of data entry is generally well worth the financial outlay.
Most data preparation agencies have facilities to set up range checks (and sometimes consistency checks) on the data entered, flagging out-of-range values. Providing a blank questionnaire (and coding scheme, if appropriate) to the agency in advance of sending the first batch of questionnaires for processing allows the agency to set up the data entry ‘templates’ accordingly. In addition, clear written instructions should be given to those carrying out data entry, especially if the data are not in a standard format or there if particular problems have been identified at the checking and editing phase. Where the data entry staff are unable to enter a code unambiguously (for example, because the respondent’s or coder’s writing is unclear or because more than the allowed number of responses have been endorsed), an agreed symbol should be entered – for example * or ? (but not a number and not a blank since those are likely to be legitimate responses) – and the offending questionnaire should be clearly marked.
It is also helpful to inform data preparation staff of how many of each type of questionnaire are being sent – this allows them to check that pages have not been missed or batches duplicated. A record should be kept of what questionnaires have been sent for data entry and when, and a check kept on the return of questionnaires from the data entry service. This helps to minimise the risk of some data being entered twice, or conversely, of some questionnaires being excluded completely.
Ideally, all data should be double entered (i.e. ‘punched and verified’); although this increases the cost, errors due to mis-typing are minimised and obvious problems are identified and rectified.
Data validation
Stages in validation
Validation checks should be run on each batch of data received from the data preparation service before the raw data are added to the master data set. Steps in the validation process should ideally be:
- Reconcile (cross-match) identification numbers (and other identifying data, such as age and gender) between the computerised data file of questionnaire responses and the survey data control system database to ensure that data has been entered for all those for whom questionnaires have been returned.
- List data and scan for data entry queries (designated by the agreed symbol), incorrect record lengths and incorrect number of records.
- Write a programme to check for range errors (i.e. values that are not possible).
- Write a programme to check for consistency and logical errors (e.g. stem and branch checks, correct handling of skipped sections, lie detector questions and IF THEN sets.
- After the master data set is produced, depending on the size of the questionnaire and the number of cases, run a "frequencies = all" check, using a statistical analysis procedure. This can be used to identify errors not already thought of and can pick up if there is more than one case with the same survey number (suggesting that the same questionnaire(s) have been entered more than once).
Making corrections
As with the manual checking stage, it is important to keep an ‘audit trail’ of changes made to computerised records. When a mistake or a discrepancy is identified, if the validator is able to show that an error has been made by the respondent, interviewer, coder or data preparation service, then the original data should be altered on the questionnaire using a distinctive colour of ink (e.g. green biro) and on the computer file using the edit facility. If a judgement has to be made in order to ensure internal consistency, the data itself should not be changed; rather, data transformation statements should be used (e.g. in an SPSS syntax file). A record should be kept of all alterations made to the data and of transformation statements added to the data description file with reasons for making or not making changes.
Under some circumstances, it may be desirable to leave errors in raw data (to compute item omission or error rates). In these circumstances, do not edit the data set to correct such errors. Instead, use data transformation statements to recode incorrect values to missing data, and note this in list of reasons for change.
Make regular back-ups of data sets (ensuring that the latest version is being backed up). Retain the original raw data (as entered from questionnaires) and the most recent back-up for audit trail purposes.
Summary of key points
- Options for data entry are data keying from a conventional ‘pencil-and-paper’ questionnaire, electronic scanning and direct capture to disk (CAPI, CATI, CASI and internet surveys).
- Data processing and analysis should be anticipated in designing the questionnaire, allocating codes to response categories and formatting the questionnaire.
- Each piece of information recorded on a questionnaire needs to be allocated a numeric code and a designated position in the computerised record.
- In allocating pre-codes, consistency and logic are key principles.
- If responses to open questions are to be quantified and statistically manipulated, numeric post-codes (office codes) will need to be allocated prior to data entry. A well-defined and systematic approach to developing, testing and applying codes is needed, with survey designers and coders contributing to the process.
- Careful checking, editing and validation of data is required both before and after data entry, to ensure that responses are appropriate, valid and consistent. Standardised protocols and rules need to be applied to ensure reliability and minimise the risk of bias. An ‘audit trail’ of changes made to raw questionnaire responses should be kept.
- Ideally, data keying should involve both data entry and data verification (i.e. should involve double entry of each questionnaire). It may be best to contract out this task to a specialised data preparation agency.
Further reading
Collins M and Kalton G. Coding verbatim answers to open questions. Journal of the Market Research Society 1980, 22, 239-247.
- de Vaus DA. Surveys in social research (4th edition). London: UCL Press, 1996. (Chapters 14 & 16)
- Fowler FJ Junior. Survey research methods (2nd edition). Newbury Park: Sage Publications (Applied Social Research Methods Series - Volume 1), 1993. (Chapter 8).
Mangione TW. Mail surveys - improving the quality. Thousand Oaks: Sage Publications (Applied Social Research Methods Series, Volume 40), 1995. (Chapter 9)
Moser CA and Kalton G. Survey methods in social investigation. London: Gower, 1971. (Chapter 16)
Oppenheim AN. Questionnaire design, interviewing and attitude measurement (2nd edition). London: Pinter Publications, 1992. (Chapter 14)
- The UK Data Archive have a useful resource on research data management at https://dam.ukdataservice.ac.uk/media/622416/trainingresourcespack.pdf
- A brief overview of code-books may be found at https://www.icpsr.umich.edu/icpsrweb/content/shared/ICPSR/faqs/what-is-a-codebook.html
- ICPSR’s codebook guide summarises the information a survey codebook should contain.
- The Understanding Society documentation portal provides current user guides, questionnaires, data documentation and information about its predecessor, the British Household Panel Survey.
- The UCL Centre for Longitudinal Studies’ questionnaire appendices provide a worked example of survey coding frames.
Current guidance and methodological literature for 2026–27:
- UK Data Service. Quality: ensuring quality of data at all stages. current guidance.
- DDI Alliance. DDI Common Core. Version 1, 2025.
- Wilkinson, M. D. et al.. The FAIR Guiding Principles for scientific data management and stewardship. 2016.
References
No specific references for this chapter.