Can AI write exam-quality questions?
Language models can draft a thousand plausible questions before lunch. Whether any of them measures anything is a different claim — one the psychometric literature has been testing since 1970. What the evidence says about machine-written items, and why Future Proof™’s Question Studio keeps a human between the generator and the learner.
The finding: Inside disciplined, template-based pipelines, machine-generated test items have matched human-written ones in expert quality ratings and psychometric behaviour. Free-form LLM generation makes volume essentially free — but a meaningful share of generated items still fails expert review, so acceptance is earned item by item, not assumed.
The mechanism: In template-based AIG, a validated cognitive model — not the surface text — carries the measurement properties, so generated siblings inherit them. LLMs drop the template: fluency comes free, but difficulty, discrimination, and a single defensible key are empirical properties that only review and field data can confirm.
The product: Future Proof’s Question Studio generates items at scale — and every item passes human review and automated quality gates before a learner ever sees it.
In this article
- 01A sixty-year-old idea
- 02The template era’s report card
- 03Then the generators learned to talk
- 04Psychometric equivalence is the bar
- 05Difficulty on demand: the unsolved control problem
- 06Why human review is a hard gate, not a courtesy
- 07What the evidence doesn’t show
- 08The bottom line for practice
Of all the tasks language models have been pointed at, writing test questions looks like the most clearly solved. The output is short, the format is rigid, the training data holds millions of examples, and the results read beautifully. Which is exactly why this corner of the literature teaches so much. It is the clearest case we have of a task where surface skill and working skill come apart. The thing the model produces looks identical to the thing an expert produces. The difference only shows up when real examinees start answering.
Every serious assessment programme eventually hits the same wall: it runs out of questions. Items get exposed, retired, or leaked; curricula shift; every new test form needs fresh material that behaves like the old material. And writing a good item is slow, skilled work. The multiple-choice question looks simple on the page. In fact it is governed by a thick rulebook of craft knowledge, which one influential review distilled into 31 separate item-writing guidelines (Haladyna, Downing & Rodriguez, 2002).
31 Item-writing guidelines in the field’s standard rulebook for a “simple” multiple-choice question — the craft knowledge any generator’s output is graded against, and the ready-made rubric its reviewers inherit (Haladyna, Downing & Rodriguez, 2002).
It is also costly work — startlingly so, to people outside the field. Cost analyses in the measurement literature put a professionally written, reviewed, and field-tested item anywhere from several hundred to a few thousand dollars (Kosh et al., 2019). Multiply that by the thousands of items an adaptive test or a large course catalogue burns each year, and the appeal of automation is obvious. So the question worth asking is not whether machines can produce questions. They plainly can, at near-zero marginal cost. The question is whether those questions measure anything.
A sixty-year-old idea
The first thing worth knowing about this field is its age — because the age is what supplies the evidence standards.
Automatic item generation did not start with large language models. In 1970, John Bormuth proposed deriving test questions directly from teaching text through explicit linguistic transformations — in effect, an algorithm for turning prose into probes (Bormuth, 1970). The proposal proved too rigid for operational testing. But its framing stuck: if item writing follows rules, then in principle a machine can follow them too.
The field that grew from that idea matured under the label AIG — automatic item generation — and was codified in the volume edited by Gierl and Haladyna (Gierl & Haladyna, 2012). Modern template-based AIG is a three-step discipline. First, subject experts build a cognitive model: an explicit map of the knowledge and reasoning the item is meant to probe. Second, they build an item model — a template in which some elements are fixed and others vary within set ranges. Third, a generation engine fills in the template, producing tens or hundreds of sibling items from a single validated design.
The template era’s report card
The discipline sounds laborious because it is — and the labour is the point. Everything downstream of the cognitive model inherits its validity, so the field concentrated its scrutiny where the leverage was.
The evidence from that programme is genuinely strong — within its lane. In medical education, Gierl, Lai and Turner showed that a single well-built item model could yield large families of usable multiple-choice questions (Gierl, Lai & Turner, 2012). In a follow-up study, expert reviewers rated machine-made items against hand-written ones, and judged the machine’s items comparable in quality (Gierl & Lai, 2013). On the economics, a cost–benefit analysis found that item-model-based generation can cut per-item costs a great deal once the up-front modelling work is paid for (Kosh et al., 2019).
The psychometric case runs deeper still. Embretson showed that when items come from a common structural model, their statistical properties are not a lottery. Item difficulty can be partly predicted from the generative features themselves, and item families can be calibrated as families rather than one orphan at a time (Embretson, 1999). That is the quiet superpower of template-based AIG. Because the variation between siblings is controlled, the measurement properties travel with the template. The items are narrow, unglamorous, and costly to design — and they behave.
The scarce resource in assessment was never the items; it was the validated reasoning behind them. Model the reasoning once, and the items follow.The central argument of Gierl & Haladyna (2012), paraphrased
Then the generators learned to talk
The template era’s limitation was never quality; it was reach. A cognitive model takes experts weeks to build and covers one narrow slice of one domain, which confines disciplined AIG to the programmes that can afford it — licensure boards, testing companies, medical schools. The obvious question was whether generation could escape the template without losing what the template guaranteed. That is the question the neural era reopened.
Natural-language generation loosened the template’s grip. The best map of that shift is the systematic review by Kurdi and colleagues, which surveyed the automatic question generation literature through 2019. It found a field with impressive machinery and immature evaluation. Systems almost always generated shallow, fact-based questions, and grading methods varied so widely that results could rarely be compared across studies. Questions were far more often judged by a handful of experts than tested on real learners in real settings (Kurdi et al., 2020). Control over item difficulty — the property psychometricians care about most — was rarely attempted at all.
Neural generation reached measurement circles at about the same time. Von Davier trained a recurrent network on a pool of personality items and showed it could emit novel, human-plausible items. He was explicit that the output was raw material for human curation, not a finished instrument (von Davier, 2018).
By 2022, transformer-based generation had reached live, high-stakes use. The Duolingo English Test’s interactive reading items are drafted by language models, then passed through human review and psychometric screening before any test-taker sees them (Attali et al., 2022). In parallel, studies compared LLM-written questions with human-written ones. The machines neared human quality on several rated dimensions — while still producing a share of items that needed filtering or repair before use (Wang et al., 2022).
Psychometric equivalence is the bar
Before weighing any of these results, it is worth pausing on what the bar actually is — because most public discussion of AI-written questions is conducted against the wrong one.
Here the vocabulary matters. A test item is not a piece of content. It is a measuring instrument with empirical properties: a difficulty, a discrimination — how sharply it separates stronger from weaker candidates — a single defensible key, and distractors that actually attract real misconceptions. None of these properties is visible on the item’s surface. A question can be grammatical, plausible, even elegant, and still be psychometrically dead: two defensible answers, a giveaway in the stem, distractors nobody picks.
This is why “the model writes good questions” and “the model writes exam-quality questions” are different claims. For template-based AIG, equivalence evidence exists because the templates were engineered to preserve it (Embretson, 1999), and expert review backed it up (Gierl & Lai, 2013). For free-form LLM generation, that guarantee vanishes with the template. Every generated item is a fresh hypothesis about difficulty and discrimination. Hypotheses about items are settled by review and field data, not by fluency. A review of AIG in medical assessment reached essentially this verdict: feasibility is well shown; the validity evidence is still being assembled (Falcão, Costa & Pêgo, 2022).
Difficulty on demand: the unsolved control problem
Among an item’s empirical properties, difficulty deserves its own section. It is the one generation pipelines most need to control, and the one they control least. An adaptive test does not want “questions about photosynthesis”; it wants a question about photosynthesis that roughly 60 percent of intermediate learners will answer correctly. The reason: the information an item yields peaks when it is matched to the examinee’s level. An item bank without difficulty coverage leaves the algorithm nothing to pick from.
Ask a generator for an easy item and a hard item on the same concept and it will comply — with surface gestures: it lengthens the stem, adds jargon, or tacks on a negation. Whether those gestures move empirical difficulty, and by how much, is precisely what the surface cannot reveal.
The template tradition earned partial control the hard way. When siblings vary only along modelled features, difficulty can be partly predicted from the features themselves (Embretson, 1999). Partly — the residual is big enough to keep field-testing in the loop. The free-form literature has mostly not earned it at all; Kurdi and colleagues found difficulty control rarely even attempted across the systems they reviewed (Kurdi et al., 2020). The operational answer, for now, is empirical: generate, screen, then let real response data assign each survivor its measured difficulty before it counts toward anything (Attali et al., 2022). The generator proposes; the data disposes.
Fluency comes free; difficulty does not. A generator asked for an “easy” and a “hard” item complies with surface gestures — longer stems, jargon, a negation — and whether those gestures move empirical difficulty is exactly what the surface cannot reveal. Across the systems in the field’s systematic review, difficulty control was rarely even attempted (Kurdi et al., 2020).
Why human review is a hard gate, not a courtesy
Across five decades of this literature, one design constant survives every technology change: a qualified human stands between the generator and the examinee. In the systematic-review evidence, expert judgment is the dominant way generated questions get evaluated at all (Kurdi et al., 2020). In the live deployments, human review plus empirical screening is built into the pipeline, not bolted on (Attali et al., 2022).
And the review is not rubber-stamping. Published pipelines consistently report that a real share of generated items is revised or rejected. The reasons repeat: cueing the key, two defensible answers, distractors that collapse, difficulty far from intent. Acceptance rates vary with the domain, the prompt, and the model, which is precisely why no serious programme has removed the gate.
The economics still favour the machine, and by a comfortable margin. Reviewing an item is far cheaper than writing one. So a pipeline that generates ten candidates and keeps six beats hand-writing six on cost and speed (Kosh et al., 2019). And the reviewer does not have to improvise a standard: the rubric already exists, in the item-writing guidelines the field has been refining for decades (Haladyna, Downing & Rodriguez, 2002). What generation really changes is the human’s role — from author to editor, from producing items to judging them.
How Future Proof™ applies this.
Question Studio generates items at scale from your source content and skills map — but nothing it drafts is learner-visible by default. Every candidate item runs through automated quality gates (guideline linting, single-key verification, distractor checks, difficulty estimates) and then through human subject-matter review, exactly the hard gate the literature says cannot be skipped. Items that pass are calibrated against real response data over time; items that don’t never reach a learner.
See Question Studio →What the evidence doesn’t show
An honest reading of this literature has to mark its edges, because several load-bearing claims have not been established. Each gap below is a place where a confident vendor pitch runs ahead of the published record:
- No evidence generated items can skip pretesting. Item parameters can be partly predicted from generative structure (Embretson, 1999) — but “partly” is doing real work in that sentence. The residual is large enough that no one has shown a generated item can be certified for high-stakes use without empirical data.
- Little live-classroom evidence. Most generation studies evaluate items with expert raters, not with learners in authentic settings (Kurdi et al., 2020). Expert approval is a proxy, and the field knows it.
- Higher-order items remain hard. The bulk of the generation literature targets recall and comprehension; evidence that machines can reliably generate items probing analysis, evaluation, or transfer is thin (Kurdi et al., 2020).
- Fairness at scale is largely untested. Systematic evidence on whether generated item banks behave equivalently across demographic groups is scarce; the operational deployments that do audit this are the exception, not the rule (Attali et al., 2022).
- The showcases are not the average case. The strongest results come from well-resourced testing organisations with psychometric staff (Gierl, Lai & Turner, 2012) (Attali et al., 2022). How the same tooling performs in a typical training team without those safeguards is an open question.
Where the evidence stops
- 1No evidence generated items can skip pretesting
- 2Little live-classroom evidence
- 3Higher-order items remain hard
- 4Fairness at scale is largely untested
- 5The showcases are not the average case
The bottom line for practice
It helps to state plainly what the two failure modes cost, because they are not symmetric. An item bank that is too small is an inconvenience: forms repeat, exposure rises, refresh cycles strain. An item bank full of unvalidated items is a measurement failure that looks like success. Scores are produced and decisions are made, yet nothing on the score report reveals that some numbers rest on questions with two right answers or none. Volume problems announce themselves; validity problems have to be hunted. Any pipeline that trades the second risk for relief from the first has made a bad bargain, however impressive the generation demo.
Read as a whole, the literature supports a narrower — and more useful — conclusion than either the hype or the backlash. Generation solves the volume problem; only the validation pipeline solves the quality problem. The best current systems treat the model as a junior item writer with infinite stamina and incomplete judgment: prolific, cheap, frequently good, never trusted alone. The rate-limiting step in assessment has moved from writing items to judging them — a real shift in the economics of measurement, an order-of-magnitude change in what a small team can build and maintain. It is just not the same claim as “the AI writes your exam”. The organisations getting the most from generation are the ones that understood the distinction first.
Selected papers.
This is not an exhaustive bibliography — these are the studies cited above, in chronological order.
The evidence, by year
- 1970Bormuth
- 1999Embretson
- 2002Haladyna
- 2012Gierl
- 2012Gierl
- 2013Gierl
- 2018Davier
- 2019Kosh
- 2020Kurdi
- 2022Attali
- 2022Falcão
- 2022Wang
- Bormuth, J.R. (1970). On the Theory of Achievement Test Items. University of Chicago Press, Chicago. PDF
- Embretson, S.E. (1999). Generating items during testing: Psychometric issues and models. Psychometrika 64(4): 407–433. PDF
- Gierl, M.J., & Haladyna, T.M. (Eds.) (2012). Automatic Item Generation: Theory and Practice. Routledge, New York. PDF
- Gierl, M.J., Lai, H., & Turner, S.R. (2012). Using automatic item generation to create multiple-choice test items. Medical Education 46(8): 757–765. PDF
- Gierl, M.J., & Lai, H. (2013). Evaluating the quality of medical multiple-choice items created with automated processes. Medical Education 47(7): 726–733. PDF
- von Davier, M. (2018). Automated item generation with recurrent neural networks. Psychometrika 83(4): 847–857. PDF
- Kosh, A.E., Simpson, M.A., Bickel, L., Kellogg, M., & Sanford-Moore, E. (2019). A cost–benefit analysis of automatic item generation. Educational Measurement: Issues and Practice 38(1): 48–53. PDF
- Attali, Y., Runge, A., LaFlair, G.T., Yancey, K., Goodwin, S., Park, Y., & von Davier, A.A. (2022). The interactive reading task: Transformer-based automatic item generation. Frontiers in Artificial Intelligence 5: 903077. PDF
- Falcão, F., Costa, P., & Pêgo, J.M. (2022). Feasibility assurance: A review of automatic item generation in medical assessment. Advances in Health Sciences Education 27. PDF
- Wang, Z., Valdez, J., Basu Mallick, D., & Baraniuk, R.G. (2022). Towards human-like educational question generation with large language models. In Proceedings of Artificial Intelligence in Education (AIED 2022), Lecture Notes in Computer Science 13355. PDF
Question Studio drafts. Humans decide.
Book a 20-minute demo with your own content. We’ll generate a set of items live, run them through Future Proof’s quality gates, and show you exactly what the human reviewer sees before anything reaches a learner.