How to Write Good Multiple-Choice Questions (and Spot Bad AI Ones)

·10 min read
How to Write Good Multiple-Choice Questions (and Spot Bad AI Ones)

Share this article

Multiple-choice questions have a reputation for testing shallow recognition. Badly written ones do. A well-written one forces you to retrieve the idea and rule out its near neighbours, which is exactly the kind of thinking exams reward.

This guide is for two groups: students writing their own practice questions, and teachers writing quizzes or checking AI-generated ones. The rules are the same. At the end we take real questions produced by our text to quiz generator and grade them against the rules, because the fastest way to learn what makes a good question is to look closely at some that are good and some that are not.

Why the format is worth getting right

Multiple choice gets dismissed as "just recognition," but the research is more forgiving than that. Little, Bjork, Bjork and Angello (2012, Psychological Science) found that multiple-choice practice questions with competitive, plausible wrong answers improved later recall of both the tested information and related information attached to the wrong options. The quality of the distractors was the key ingredient.

There is one known risk. Roediger and Marsh (2005) showed that taking a multiple-choice test can lead students to later reproduce some of the wrong options they saw, which they called the negative suggestion effect. Butler and Roediger (2008) found that giving feedback after the test largely removed this. So: good distractors, and always check the answers afterwards. Every rule below serves one of those two goals.

The anatomy of a question

A multiple-choice item has three parts:

  • The stem: the question or incomplete statement.
  • The key: the correct answer.
  • The distractors: the wrong answers.

Most bad questions fail in the stem or the distractors. The key is usually fine.

Writing the stem

Put the whole problem in the stem

A reader should be able to answer the stem without seeing the options. "Glycolysis:" followed by four statements is not a question. It is four true-or-false questions stapled together. "Where in the cell does glycolysis take place?" is a question.

A quick test: cover the options. If you cannot form an answer in your head, rewrite the stem.

Ask one thing

"Which stage of respiration produces the most ATP and takes place in the matrix?" tests two facts at once, and one of the premises is false (the electron transport chain is on the inner membrane, not in the matrix). Split it.

Prefer positive wording

"Which of the following is NOT..." questions are sometimes necessary, but they test your ability to hold a negation in mind as much as the content. If you use one, capitalise the NOT and keep it rare.

Ask about relationships, not words

"What does ATP synthase do?" tests a definition. "What would happen to ATP production if the inner mitochondrial membrane became leaky to protons?" tests whether you understand why the definition matters. Butler (2010, Journal of Experimental Psychology: Learning, Memory, and Cognition) found that repeated retrieval practice improved performance on new inference questions, including ones in a different domain, compared with repeated studying. That transfer is what you want, and "what would happen if" stems practise it directly.

Writing distractors

This is where most of the value is, and where most questions fall apart.

Every distractor should be a real mistake

The best source of wrong answers is the specific misunderstanding a student might have. For cellular respiration, predictable confusions include:

  • mixing up where each stage happens (cytoplasm, matrix, inner membrane)
  • confusing total ATP per glucose with ATP from one stage
  • thinking fermentation produces ATP beyond glycolysis, when its job is regenerating NAD+
  • confusing what blocks the chain (cyanide) with what uncouples it (DNP)

Each of those is a distractor waiting to be written. "Photosynthesis" as an answer to a question about ATP synthase is not a distractor. Nobody picks it, so the question has effectively three options.

Three good options beat four with filler

Rodriguez (2005) reviewed decades of research on the number of options and concluded that three options (a key plus two good distractors) is usually enough; the fourth and fifth options tend to be ones hardly anyone chooses. For your own practice questions, if you cannot think of a third plausible wrong answer, do not invent a silly one.

Keep options parallel

All options should be the same kind of thing, the same grammatical form, and roughly the same length. If three options are single words and one is a careful sentence with a qualifying clause, the careful sentence is almost always right.

Cues that give the answer away

Haladyna, Downing and Rodriguez (2002) compiled a widely used set of item-writing guidelines, and a lot of them are about removing accidental cues. The ones that matter most for student-written and AI-written questions:

  • Longest-answer cue. The key is longer and more qualified because the writer was being careful to make it true.
  • Verbatim cue. The key copies a phrase straight from the notes, and the distractors are paraphrased. You match wording instead of understanding.
  • Grammatical cue. The stem ends in "an" and only one option starts with a vowel.
  • Absolute words. Distractors containing "always" or "never" are usually wrong, and experienced test-takers know it.
  • "All of the above." If you can spot two correct options, you know the answer without knowing the third. Avoid it.
  • Implausible outliers. One option is obviously absurd, so the question is really a three-way choice.

Judging AI-generated questions: a real example

AI generators are good at producing grammatical, well-structured questions quickly. They are less reliable at choosing what is worth testing and at keeping the key consistent with the source. Here is a real test.

We pasted about 250 words of notes on cellular respiration (glycolysis, the Krebs cycle, the electron transport chain, fermentation, cyanide and uncouplers) into the text to quiz generator, 10 questions, medium difficulty. The questions below were generated by the tool and are copied without edits.

What is the consequence of cyanide blocking cytochrome c oxidase?

  • A) Increased ATP production.
  • B) Stalled Krebs cycle.
  • C) No ATP from oxidative phosphorylation. (correct)
  • D) Enhanced fermentation.

Grade: good, with one weak option. "Stalled Krebs cycle" is a clever distractor, since it is actually a downstream consequence, so the student has to pick the most direct effect. "Enhanced fermentation" is also defensible. "Increased ATP production" is an implausible outlier nobody will choose.

How many ATP are produced from one glucose molecule during oxidative phosphorylation?

  • A) 2-4 ATP
  • B) 26-28 ATP (correct)
  • C) 30-32 ATP
  • D) 4-6 ATP

Grade: good distractor, low-value target. Offering 30-32 (the total for all of respiration) right next to 26-28 is exactly the "total versus one stage" confusion listed above. But the question tests a number rather than an idea. Keep it for revision, but do not let questions like this make up most of a quiz.

What is the effect of uncouplers like DNP on cellular respiration?

  • A) They increase ATP production.
  • B) They make the membrane leaky to protons. (correct)
  • C) They enhance the Krebs cycle.
  • D) They block glycolysis.

Grade: verbatim cue. The notes said "Uncouplers (e.g. DNP) make the membrane leaky to protons," and the key repeats that phrase almost word for word. A student who has seen the notes can match the wording without understanding why a leaky membrane stops ATP synthesis. A better stem: "Why does DNP cause cells to release more energy as heat?" with the mechanism as the key.

Mistakes to check for in any AI quiz

Beyond the cues above, look for these problems, all of which we have seen in real generator output:

  • Miskeyed answers. In a quiz generated from psychology notes, a question on proactive interference marked the definition of retroactive interference as correct, while its own explanation described the right answer. When the explanation and the key disagree, check the source.
  • Claims beyond the source. In a quiz from a history chapter, one question stated as fact something the chapter never said. The generator filled a gap with an inference. It may even be true, but you cannot verify it from your material.
  • Loaded words. "Immediate consequence" for an event that happened a month later, after several intervening steps. One word can make a key debatable.

A practical workflow: take the AI quiz first, as a student would. Then, only for the questions you got wrong or hesitated on, check the stem, the key and the explanation against your notes. That keeps checking time short and focuses it where errors matter to you.

A checklist for any multiple-choice question

Run each question through these seven checks:

  1. Can I answer it with the options covered?
  2. Does it test one idea?
  3. Is that idea worth testing, meaning a relationship or mechanism rather than trivia?
  4. Is every distractor a mistake a real student could make?
  5. Are the options parallel in length and form?
  6. Does the key avoid copying the source's wording?
  7. Does the key match the source and the explanation?

A question that passes all seven is a good practice question, whether you wrote it, a teacher did, or a tool did.

The takeaway

The value of a multiple-choice question lives in its distractors. Write stems that stand alone, build wrong answers from real misconceptions, remove length, wording and grammar cues, and always review answers after the quiz so wrong options do not stick. Use AI to generate a first draft fast, then apply the checklist to the questions you missed. If you would rather write your own, draft them from the same notes and compare: the gaps between your questions and the tool's are often the most useful thing you learn.

If your material is a textbook chapter or a set of slides rather than notes, see how to quiz yourself from a textbook chapter. For the broader method behind all of this, our guide to active recall covers why testing beats rereading.

Frequently asked questions

How many options should a multiple-choice question have?

Three or four. Research on option counts suggests three well-written options (one key and two plausible distractors) usually works as well as four or five, because extra options are often ones nobody picks. Four is fine if the fourth is a genuine misconception.

How do you write good distractors?

Base each one on a specific mistake a learner might make: a confused term, a reversed cause and effect, a number from the wrong step, or a related concept from the same topic. If nobody would realistically choose an option, replace it or drop it.

Can multiple-choice questions test deep understanding?

Yes, if the stem asks about a mechanism, prediction or application rather than a definition, and the distractors represent plausible misunderstandings. "What would happen if..." stems are one of the easiest ways to move from recognition to reasoning.

Should I trust AI-generated multiple-choice questions?

Use them as a draft. They are usually well formed, but they can miskey an answer, copy wording from the source, or add claims the source does not support. Check any question you got wrong or were unsure about against your own material.

Studying from YouTube lectures?

Paste a lecture link and Notiq turns it into chapter notes, flashcards and exam questions in about two minutes.

Free · no card needed · lectures up to 20 minor browse real notebooks

Studying from notes or slides instead? Use the free text to quiz generator — paste text or upload photos, PDFs and PowerPoints.

Share this article

Related Articles