← All articles
psychologyteachingassessmentmethodSeptember 17, 20263 min read

Why Are Some Questions Better Than Others? Writing a Test That Measures Something

By the BrainSnail editorial team. How these articles are written and checked, and how to tell us when one is wrong.

A question that everybody answers correctly and one that everybody gets wrong both measure nothing. Writing assessment that distinguishes between candidates is a technical craft with rules that are frequently ignored.

What a question is for

An assessment exists to produce information, which means distinguishing between candidates who know the material and candidates who do not, and a question contributes only insofar as it does that. A question everybody answers correctly and one nobody answers correctly both fail, since neither separates anybody, and a paper composed of such questions produces identical scores regardless of ability. That framing is uncomfortable because it implies a well-designed paper must contain questions many candidates get wrong, which conflicts with the instinct that a good test is one people can do. The resolution depends on what the assessment is for, since a test establishing whether a threshold has been reached has different requirements from one ranking candidates against each other.

How questions are evaluated

Standard measures tell an examiner which questions worked:

  • Facility, the proportion answering correctly, which should mostly sit in a middle range
  • Discrimination, how well performance on that question tracks performance overall
  • A question that strong candidates get wrong more often than weak ones is defective
  • Distractor analysis, examining which wrong options were chosen and whether any attracted nobody
  • Reliability across the whole paper, meaning how consistently it measures
  • These are calculated after the fact, which is why questions are trialled before being used

How questions go wrong

The failure modes are well catalogued and recur constantly. A question can be answerable from general reasoning without the knowledge it targets, which tests something else. It can contain a clue in its wording, with the correct option longer, more qualified or grammatically matching the stem when the others do not. It can be ambiguous, so that a defensible reading produces a different answer, which penalises careful candidates specifically. It can test reading comprehension or cultural familiarity rather than the subject. It can be so entangled with an earlier question that failing one guarantees failing both. And it can test recall of something trivial because that is easy to write unambiguously, which is the commonest failure of all.

Whether the marking is consistent

Anything requiring judgement introduces variation between markers, which is measured and is larger than most people expect. Studies giving the same scripts to multiple markers find substantial disagreement on extended writing, with the spread narrowing considerably where detailed criteria and training are used and never disappearing. The standard responses are published mark schemes specifying what earns credit, training markers against pre-marked scripts, double marking with reconciliation of differences, statistical monitoring of each marker against the cohort, and moderation of borderline scripts by senior examiners. Candidates almost never see any of that machinery. Objectively marked formats avoid the problem entirely, which is a genuine advantage and is why they persist despite what they cannot assess.

What different formats can and cannot do

Format constrains what can be assessed, and the trade is between coverage and depth. Multiple choice covers a lot of material quickly, marks objectively and at scale, and struggles to assess anything requiring construction, argument or original work, although well-written items can test reasoning far better than their reputation suggests. Short written answers assess construction and cost marking time. Extended writing assesses argument and organisation, and introduces marker variation that requires moderation, double marking and published criteria to control. Practical assessment measures what somebody can do and is expensive and hard to standardise. Most serious qualifications use several formats precisely because each covers what the others cannot, which is a deliberate design rather than administrative inertia.

The takeaway

A question everybody answers correctly and one nobody does both fail, since neither separates candidates, so facility and discrimination are calculated to find out which questions worked. Clues in wording, ambiguity and testing general reasoning rather than the subject are the standard defects. Format decides what can be assessed, which is why serious qualifications use several.

Practise this

Questions from Psychology in the Real World

Reading about something is not the same as being able to recall it. These are real questions from the Psychology in the Real World unit in our Psychology track, answers and explanations included. The unit has 119 in total across 23 steps.

  • Multiple choiceLevel 1

    1. What does sport psychology mainly study?

    • How the mind affects sports performancecorrect
    • How to sew clothes
    • How rockets fly
    • How to bake cakes

    Sport psychology looks at how thoughts and feelings affect how athletes perform.

  • Put in orderLevel 2

    2. Put these steps of hiring a new worker in the usual order.

    Answer: Post the job -> Read the applications -> Interview candidates -> Offer the job

    Employers usually post a job, review applicants, interview, and then make an offer.

  • Odd one outLevel 2

    3. Which of these is NOT usually a job of a forensic psychologist?

    • Baking bread for the courtroomcorrect
    • Helping judge if someone is fit for trial
    • Studying why people commit crimes
    • Understanding how eyewitnesses remember

    Baking is not a forensic job; the others all involve psychology and the law.