Research · AI & Tutoring
AI & Tutoring · Automatic Item Generation

Can AI write exam-quality questions?

Language models can draft a thousand plausible questions before lunch. Whether any of them measures anything is a different claim — one the psychometric literature has been testing since 1970. What the evidence says about machine-written items, and why Future Proof™’s Question Studio keeps a human between the generator and the learner.

TL;DR

The finding: Inside disciplined, template-based pipelines, machine-generated test items have matched human-written ones in expert quality ratings and psychometric behaviour. Free-form LLM generation makes volume essentially free — but a meaningful share of generated items still fails expert review, so acceptance is earned item by item, not assumed.

The mechanism: In template-based AIG, a validated cognitive model — not the surface text — carries the measurement properties, so generated siblings inherit them. LLMs drop the template: fluency comes free, but difficulty, discrimination, and a single defensible key are empirical properties that only review and field data can confirm.

The product: Future Proof’s Question Studio generates items at scale — and every item passes human review and automated quality gates before a learner ever sees it.

Every serious assessment programme eventually hits the same wall: it runs out of questions. Items get exposed, retired, or leaked; curricula shift; every new test form needs fresh material that behaves like the old material. And writing a good item is slow, skilled work. The multiple-choice question — deceptively simple on the page — is governed by a thick rulebook of craft knowledge, which one influential review distilled into 31 separate item-writing guidelines (Haladyna, Downing & Rodriguez, 2002).

It is also expensive work. Cost analyses in the measurement literature put a professionally written, reviewed, and field-tested item anywhere from several hundred to a few thousand dollars (Kosh et al., 2019). Multiply that by the thousands of items an adaptive test or a large course catalogue consumes each year, and the appeal of automation is obvious. So the question worth asking is not whether machines can produce questions — they demonstrably can, at near-zero marginal cost — but whether those questions measure anything.

A sixty-year-old idea

Automatic item generation did not start with large language models. In 1970, John Bormuth proposed deriving test questions directly from instructional text through explicit linguistic transformations — in effect, an algorithm for turning prose into probes (Bormuth, 1970). The proposal proved too rigid for operational testing, but its framing stuck: if item writing follows rules, then in principle a machine can follow them too.

The field that grew from that idea matured under the label AIG — automatic item generation — and was codified in the volume edited by Gierl and Haladyna (Gierl & Haladyna, 2012). Modern template-based AIG is a three-step discipline. First, subject-matter experts build a cognitive model: an explicit map of the knowledge and reasoning the item is meant to probe. Second, they build an item model — a template in which some elements are fixed and others vary within constrained ranges. Third, a generation engine instantiates the template, producing tens or hundreds of sibling items from a single validated design.

The template era’s report card

The evidence from that programme is genuinely strong — within its lane. In medical education, Gierl, Lai and Turner showed that a single well-built item model could yield large families of usable multiple-choice questions (Gierl, Lai & Turner, 2012). In a follow-up evaluation, expert reviewers rated machine-generated and traditionally written items and judged the generated items comparable in quality (Gierl & Lai, 2013). On the economics, a cost–benefit analysis concluded that item-model-based generation can cut per-item development costs substantially once the up-front modelling work is paid for (Kosh et al., 2019).

The psychometric case runs deeper still. Embretson showed that when items are generated from a common structural model, their statistical properties are not a lottery: item difficulty can be partly predicted from the generative features themselves, and item families can be calibrated as families rather than one orphan at a time (Embretson, 1999). That is the quiet superpower of template-based AIG. Because the variation between siblings is controlled, the measurement properties travel with the template. The items are narrow, unglamorous, and expensive to design — and they behave.

The scarce resource in assessment was never the items; it was the validated reasoning behind them. Model the reasoning once, and the items follow. The central argument of Gierl & Haladyna (2012), paraphrased

Then the generators learned to talk

Natural-language generation loosened the template’s grip. The most comprehensive map of that transition is the systematic review by Kurdi and colleagues, which surveyed the automatic question generation literature through 2019 and found a field with impressive machinery and immature evaluation: systems overwhelmingly generated shallow, fact-oriented questions; evaluation methods varied so widely that results could rarely be compared across studies; and questions were far more often judged by a handful of experts than tested on real learners in real settings (Kurdi et al., 2020). Control over item difficulty — the property psychometricians care about most — was rarely attempted at all.

Neural generation reached measurement circles at about the same time. Von Davier trained a recurrent network on a pool of personality items and showed it could emit novel, human-plausible items — while being explicit that the output was raw material for human curation, not a finished instrument (von Davier, 2018). By 2022, transformer-based generation had reached operational, high-stakes use: the Duolingo English Test’s interactive reading items are drafted by language models, then passed through human review and psychometric screening before any test-taker sees them (Attali et al., 2022). In parallel, evaluation studies comparing LLM-generated with human-written educational questions found the machines approaching human quality on several rated dimensions — while still producing a share of items that needed filtering or repair before use (Wang et al., 2022).

Psychometric equivalence is the bar

Here the vocabulary matters. A test item is not a piece of content; it is a measurement instrument with empirical properties — a difficulty, a discrimination (how sharply it separates stronger from weaker candidates), a single defensible key, and distractors that actually attract real misconceptions. None of these properties is visible on the item’s surface. A question can be grammatical, plausible, even elegant, and still be psychometrically dead: two defensible answers, a giveaway in the stem, distractors nobody chooses.

This is why “the model writes good questions” and “the model writes exam-quality questions” are different claims. For template-based AIG, equivalence evidence exists because the templates were engineered to preserve it (Embretson, 1999), and expert review backed it up (Gierl & Lai, 2013). For free-form LLM generation, that guarantee vanishes with the template: every generated item is a fresh hypothesis about difficulty and discrimination, and hypotheses about items are settled by review and field data, not by fluency. A review of AIG in medical assessment reached essentially this verdict — feasibility is well demonstrated; the validity evidence is still being assembled (Falcão, Costa & Pêgo, 2022).

Why human review is a hard gate, not a courtesy

Across five decades of this literature, one design constant survives every technology change: a qualified human stands between the generator and the examinee. In the systematic-review evidence, expert judgment is the dominant way generated questions get evaluated at all (Kurdi et al., 2020). In the operational deployments, human review plus empirical screening is built into the pipeline, not bolted on (Attali et al., 2022). And the review is not rubber-stamping. Published pipelines consistently report that a non-trivial share of generated items is revised or rejected — for cueing the key, for stems with more than one defensible answer, for distractors that collapse, for difficulty far from intent. Acceptance rates vary with the domain, the prompt, and the model, which is precisely why no serious programme has removed the gate.

The economics still favour the machine. Reviewing an item is far cheaper than writing one, so a pipeline that generates ten candidates and keeps six beats hand-writing six on cost and speed (Kosh et al., 2019). And the reviewer does not have to improvise a standard: the rubric already exists, in the item-writing guidelines the field has been refining for decades (Haladyna, Downing & Rodriguez, 2002). What generation really changes is the human’s role — from author to editor, from producing items to adjudicating them.

Applied research → product

How Future Proof™ applies this.

Question Studio generates items at scale from your source content and skills map — but nothing it drafts is learner-visible by default. Every candidate item runs through automated quality gates (guideline linting, single-key verification, distractor checks, difficulty estimates) and then through human subject-matter review, exactly the hard gate the literature says cannot be skipped. Items that pass are calibrated against real response data over time; items that don’t never reach a learner.

See Question Studio

What the evidence doesn’t show

An honest reading of this literature has to mark its edges, because several load-bearing claims have not been established:

  • No evidence generated items can skip pretesting. Item parameters can be partly predicted from generative structure (Embretson, 1999) — but “partly” is doing real work in that sentence. The residual is large enough that no one has shown a generated item can be certified for high-stakes use without empirical data.
  • Little live-classroom evidence. Most generation studies evaluate items with expert raters, not with learners in authentic settings (Kurdi et al., 2020). Expert approval is a proxy, and the field knows it.
  • Higher-order items remain hard. The bulk of the generation literature targets recall and comprehension; evidence that machines can reliably generate items probing analysis, evaluation, or transfer is thin (Kurdi et al., 2020).
  • Fairness at scale is largely untested. Systematic evidence on whether generated item banks behave equivalently across demographic groups is scarce; the operational deployments that do audit this are the exception, not the rule (Attali et al., 2022).
  • The showcases are not the average case. The strongest results come from well-resourced testing organisations with psychometric staff (Gierl, Lai & Turner, 2012) (Attali et al., 2022). How the same tooling performs in a typical training team without those safeguards is an open question.

The bottom line for practice

Read as a whole, the literature supports a narrower — and more useful — conclusion than either the hype or the backlash. Generation solves the volume problem; only the validation pipeline solves the quality problem. The best current systems treat the model as a junior item writer with infinite stamina and incomplete judgment: prolific, cheap, frequently good, never trusted alone. The rate-limiting step in assessment has moved from writing items to adjudicating them. That is a real shift in the economics of measurement — it is just not the same claim as “the AI writes your exam.”

References

Selected papers.

This is not an exhaustive bibliography — these are the studies cited above, in chronological order.

  1. Bormuth, J.R. (1970). On the Theory of Achievement Test Items. University of Chicago Press, Chicago. PDF
  2. Embretson, S.E. (1999). Generating items during testing: Psychometric issues and models. Psychometrika 64(4): 407–433. PDF
  3. Haladyna, T.M., Downing, S.M., & Rodriguez, M.C. (2002). A review of multiple-choice item-writing guidelines for classroom assessment. Applied Measurement in Education 15(3): 309–334. DOIPDF
  4. Gierl, M.J., & Haladyna, T.M. (Eds.) (2012). Automatic Item Generation: Theory and Practice. Routledge, New York. PDF
  5. Gierl, M.J., Lai, H., & Turner, S.R. (2012). Using automatic item generation to create multiple-choice test items. Medical Education 46(8): 757–765. PDF
  6. Gierl, M.J., & Lai, H. (2013). Evaluating the quality of medical multiple-choice items created with automated processes. Medical Education 47(7): 726–733. PDF
  7. von Davier, M. (2018). Automated item generation with recurrent neural networks. Psychometrika 83(4): 847–857. PDF
  8. Kosh, A.E., Simpson, M.A., Bickel, L., Kellogg, M., & Sanford-Moore, E. (2019). A cost–benefit analysis of automatic item generation. Educational Measurement: Issues and Practice 38(1): 48–53. PDF
  9. Kurdi, G., Leo, J., Parsia, B., Sattler, U., & Al-Emari, S. (2020). A systematic review of automatic question generation for educational purposes. International Journal of Artificial Intelligence in Education 30(1): 121–204. DOIPDF
  10. Attali, Y., Runge, A., LaFlair, G.T., Yancey, K., Goodwin, S., Park, Y., & von Davier, A.A. (2022). The interactive reading task: Transformer-based automatic item generation. Frontiers in Artificial Intelligence 5: 903077. PDF
  11. Falcão, F., Costa, P., & Pêgo, J.M. (2022). Feasibility assurance: A review of automatic item generation in medical assessment. Advances in Health Sciences Education 27. PDF
  12. Wang, Z., Valdez, J., Basu Mallick, D., & Baraniuk, R.G. (2022). Towards human-like educational question generation with large language models. In Proceedings of Artificial Intelligence in Education (AIED 2022), Lecture Notes in Computer Science 13355. PDF
See it in the product

Question Studio drafts. Humans decide.

Book a 20-minute demo with your own content. We’ll generate a set of items live, run them through Future Proof’s quality gates, and show you exactly what the human reviewer sees before anything reaches a learner.

12 citations Reviewed July 2026 Open peer review welcomed