Bloom’s Taxonomy, Seventy Years On
Nearly everyone in training can recite the pyramid; almost no assessment climbs it. What the 1956 handbook actually claimed, what the 2001 revision fixed, which parts survived seventy years of testing — and how Future Proof™ builds quizzes that leave recall behind.
The finding: Bloom’s 1956 taxonomy and its 2001 revision remain the field’s shared language for the cognitive demand of a task. The evidence supports the ordering of its lower levels and, above all, its value as an audit tool. When researchers rate real course assessments against it, the overwhelming majority of items sit in the bottom one or two levels.
The mechanism: Recall items are cheap to write, unambiguous to grade, and easy to defend, so unmanaged item banks drift to the bottom of the taxonomy. And learners mostly improve at the level they practice — drilling facts does not automatically produce analysis.
The product: Future Proof’s quiz engine ships five Bloom levels with every chapter and draws randomized per-learner question sets, so assessment climbs past recall by design rather than by author discipline.
In this article
- 01What the 1956 handbook actually claimed
- 02The 2001 revision: nouns become verbs
- 03What the evidence supports
- 04Why most LMS assessment never leaves recall
- 05The economics just changed
- 06What the evidence doesn’t show
- 07What this means for assessment design
Frameworks that survive seventy years in daily professional use are rare enough to deserve suspicion. Long life can mean validity. It can also mean the framework became furniture — recited, diagrammed, and never again held against evidence. Bloom’s taxonomy turns out to be both at once. Knowing which parts are which is worth more than either the reverence or the debunking.
In 1948, at an American Psychological Association convention in Boston, a group of college examiners agreed to attempt something unglamorous. They wanted a shared way to classify the things examinations ask students to do. Eight years later the group, led by Benjamin Bloom of the University of Chicago, published Taxonomy of Educational Objectives, Handbook I: Cognitive Domain (Bloom et al., 1956). It became one of the most cited works in the history of education. Its six-layer pyramid is now the most recognizable diagram in instructional design — so recognizable that few of its users have ever read the handbook it summarizes.
Seventy years on, the taxonomy sits in a strange position. Nearly everyone who builds training can recite the levels; almost nobody’s assessments climb them. Audits of real course exams keep finding the same thing: the overwhelming majority of questions sit in the bottom layer or two (Momsen et al., 2010). This article treats the taxonomy as a piece of science rather than a poster. What did it claim, and what did the 2001 revision change? Which claims held up when tested — and why does assessment in practice stay at recall anyway?
What the 1956 handbook actually claimed
The taxonomy was born as a measurement tool, not a theory of learning. Bloom’s committee wanted university examiners to be able to swap test items and compare results across institutions. That required a common vocabulary for what an item demands of a student (Krathwohl, 2002). The result was six categories: Knowledge, Comprehension, Application, Analysis, Synthesis, and Evaluation. Each was defined not just in prose but with sample objectives and, crucially, sample test items.
The handbook made two claims that matter for evidence. The first is ordering: the categories run from simple to complex. The second is stronger — a cumulative hierarchy. Each category presupposes mastery of the ones below it. On the strong reading, a learner cannot genuinely analyze material they cannot first recall and comprehend. Patterns of test performance should then reflect that ladder (Bloom et al., 1956).
It is worth noticing what the handbook did not claim. It classified the demands of tasks, not the sequence of teaching. The later folklore — “cover the facts first, save the analysis for the end of the course” — was read into the pyramid, not out of the book. And the attempt-first literatures elsewhere in this library suggest the folklore sequence is often backwards.
The 2001 revision: nouns become verbs
Frameworks this widely used rarely get a formal revision; folklore just builds up around them. That makes the taxonomy’s 2001 overhaul a notable act of intellectual upkeep, done with a living link to the original committee. By the 1990s the taxonomy’s limits were well rehearsed. A group co-chaired by Lorin Anderson and David Krathwohl — one of the original 1956 authors — undertook a full revision (Anderson & Krathwohl, 2001). Three changes matter.
First, the category names became verbs: Remember, Understand, Apply, Analyze, Evaluate, Create. The change marks a shift from classifying test content to classifying mental processes. Second, the single scale became a two-dimensional table. Every objective is now classified twice — by process and by the type of knowledge involved: factual, conceptual, procedural, or (new in the revision) metacognitive. “List the stages of the sales process” and “diagnose why this deal stalled” involve the same subject matter but occupy very different cells. Third, the top of the pyramid changed: Synthesis, renamed Create, moved above Evaluate.
The quietest change was the most important. The revision relaxed the strict cumulative hierarchy. The categories are still held to differ in complexity, but they are now allowed to overlap (Krathwohl, 2002). That was an admission that four decades of data had not been kind to the strong ordering claim.
The Taxonomy of Educational Objectives is a framework for classifying statements of what we expect or intend students to learn as a result of instruction.Krathwohl 2002, Theory Into Practice
What the evidence supports
A classification scheme is not obviously something you can test — one does not test the Dewey Decimal System. But Bloom’s committee made claims about the world, and claims can be scored. Two are separable. Do the levels form a genuine ladder? And does the scheme sort real assessment items reliably enough to be useful? The evidence treats the two very differently, which is this article’s central finding.
The testable claim is the hierarchy, and it has been tested. The first major attempt came in the 1960s. Kropp and Stoker gave tests written at each taxonomy level, across several school subjects, to large samples of secondary students (Kropp & Stoker, 1966). Difficulty broadly increased up the levels, consistent with the simple-to-complex ordering. But the correlations between levels — how scores at one level tracked scores at another — only partly matched a strict cumulative ladder.
Madaus, Woods and Nuttall fit causal models to those data. They found a branching structure rather than a single chain, with a general-ability factor feeding the upper levels directly (Madaus et al., 1973). Hill and McGaw reanalyzed the same data with stronger methods and recovered clearer support for a cumulative ordering (Hill & McGaw, 1981). Even then the fit was imperfect, and it held only once the lowest category was set apart from the rest. The most thorough review of this literature appears in the taxonomy’s own forty-year retrospective. Its verdict has not changed since: the lower levels behave roughly cumulatively; the upper levels are distinct but do not stack neatly (Kreitzer & Madaus, 1994).
If the ladder is only half true, why keep the taxonomy? Because its second use — as an audit and design instrument — has fared much better. Crowe, Dirks and Wenderoth built the “Blooming Biology Tool,” a rubric for rating exam items by level. They used it to realign university courses so that in-class practice matched the cognitive level of the exams (Crowe et al., 2008). Momsen and colleagues applied the same kind of rating at scale to introductory biology courses. They found assessment concentrated almost entirely at the lowest levels (Momsen et al., 2010) — turning “our tests are shallow” from a vague worry into a measurable, fixable property of a course.
levels 1–2 Where the overwhelming majority of real exam questions sit when courses are audited against the taxonomy — and that is in university courses run by trained educators, not in workplace quiz banks (Momsen et al., 2010).
The audit instrument’s value compounds with our smile-sheet and testing-effect reviews. An organization that measures outcomes at a delay, using an item bank whose level mix it has audited, finally knows two things. It knows whether its people learned, and what kind of capability the learning was — recognition of pages, or judgment on cases. Neither measurement alone answers the question executives actually ask.
And the level of assessment appears to matter for outcomes in its own right. Jensen and colleagues compared course sections whose regular exams demanded higher-order thinking with sections whose exams stayed at recall. On a common final, the higher-order sections did better on the higher-order items and at least as well on the factual ones (Jensen et al., 2014). Agarwal’s experiments sharpen the point from the other direction. Retrieval practice on facts alone did not improve performance on higher-order test questions, while practice on higher-order questions did (Agarwal, 2019). To a first approximation, learners improve at the level they practice.
Why most LMS assessment never leaves recall
The audit findings pose the question this section answers. If everyone can recite the pyramid, why do real assessments hug its floor? The answer is neither ignorance nor laziness. Getting the diagnosis right matters, because the wrong one — “authors need more training” — has been prescribed for decades without moving the numbers. Nothing in the taxonomy explains why recall dominates practice; economics does.
A recall item can be written in a minute from any sentence of source material. It has an unambiguous answer key, is trivially auto-graded, and is easy to defend when a learner disputes it. An analysis or evaluation item needs more. It needs a scenario the learner has not seen before, wrong options that encode plausible misconceptions, and grading criteria that survive argument. Every incentive in course production points downhill.
The audit literature shows how strong that pull is even under favorable conditions. In university courses run by trained educators, the large majority of assessment items still tested the lowest levels (Momsen et al., 2010). Typical workplace quizzes are built with less assessment expertise and tighter deadlines, by authors who are subject-matter experts rather than test designers. And question banks, once written, are copied forward for years. The predictable result is the industry’s open secret. Most “knowledge checks” verify short-term recognition of the page the learner just read — a measurement of the scroll wheel, certified as a measurement of the mind.
The fix the evidence suggests is unglamorous. Treat level coverage as a property of the assessment itself. Decide the mix deliberately, and hold the item bank to it. That is the same discipline Crowe and colleagues applied when they realigned their courses (Crowe et al., 2008).
The economics just changed
Everything in the previous section described a cost structure, and cost structures can move. The reason banks drifted to recall was never pedagogy. A scenario-based analysis item cost an hour of skilled authoring where a recall item cost a minute. Grading an explanation required a human where grading a recognition click did not. Both prices have now collapsed. Generative systems draft plausible scenario items at recall-item speeds, and automated scoring makes open-response formats gradeable at scale — the two shifts this library’s AI cluster examines in detail, with the human-review gate both literatures insist on.
The collapse does not make the taxonomy obsolete. It makes it operational in a way its authors could not have imagined. For the first time, an organization can set a level mix as policy — say, no objective certified without application-level items on novel material — and have the item supply meet the spec. The marginal cost of the upper levels no longer forbids it. The audit studies documented a ceiling built of economics (Momsen et al., 2010). The economics were the constraint, and the constraint has moved.
What remains fixed is the design judgment the verb lists could never automate. Someone must decide which cell of the two-dimensional table each objective belongs in, and verify that a generated item genuinely demands the level it claims (Stanny, 2016). The taxonomy’s seventy-year career as a wall poster may finally be ending. In its place: the career its authors intended — a working specification for what assessments demand.
What the evidence doesn’t show
Fluency with a framework breeds strong claims in its name. Several of the strongest circulating claims are exactly what the evidence declines to support. Four honest limits. First, the strict cumulative hierarchy is not supported at the top of the taxonomy; the empirical work consistently finds the upper levels branching rather than stacking (Madaus et al., 1973) (Kreitzer & Madaus, 1994). Second, the taxonomy is a classification scheme, not an intervention. There is no meaningful effect size for “using Bloom’s taxonomy” — studies like Jensen’s test the effect of higher-order assessment, not of the framework itself (Jensen et al., 2014).
Third, the popular verb lists are weak stand-ins for cognitive level. Stanny’s review of published lists found the same verb assigned to different levels on different lists (Stanny, 2016). And a verb alone cannot tell you whether a task is novel to the learner — “explain” is pure recall if the explanation was given in the course. Classifying an item correctly requires knowing what the learner has already seen. That is why audit studies rely on trained raters rather than keyword matching. Fourth, the taxonomy does not license “facts first” curricula: the evidence that fact drill automatically prepares learners for higher-order performance is, at best, missing (Agarwal, 2019).
Classify items by what is novel to the learner, not by the verb in the stem. The same verb lands on different levels across published lists, and “explain” is pure recall if the explanation was given in the course (Stanny, 2016) — which is why credible audits use trained raters and the learner’s history, never keyword matching.
What this means for assessment design
- Audit before you author. Rate the existing item bank by level. Most teams that do this discover they are running a recognition test, not an assessment.
- Decide the mix, per objective. The revision’s two-dimensional table is the practical tool here: pick the cell each objective lives in, then write items for that cell.
- Write above recall against novel material. An application item must present a situation not covered verbatim in the content, or it collapses into recall regardless of its verb.
- Align practice with the target level. If the goal is analysis, learners need retrieval practice at analysis — not a diet of fact checks with one hard question at the end.
Where the evidence stops
- 1Audit before you author
- 2Decide the mix, per objective
- 3Write above recall against novel material
- 4Align practice with the target level
How Future Proof™ applies this.
Every chapter ships with a quiz set spanning five Bloom levels — from remembering through creating — so the level mix is a property of the platform, not of author discipline. Each learner receives a randomized draw from the leveled item bank, which means no two learners see the same set and no one can pass on recall alone. Admins see performance broken out by level, so “knows the facts but can’t apply them” shows up in the data instead of on the job.
See the quiz engine →Selected papers.
This is not an exhaustive bibliography — these are the studies cited above. The full reading list is in the downloadable Science Library PDF.
The evidence, by year
- 1956Bloom
- 1966Kropp
- 1973Madaus
- 1981Hill
- 1994Kreitzer
- 2001Anderson
- 2002Krathwohl
- 2008Crowe
- 2010Momsen
- 2014Jensen
- 2016Stanny
- 2019Agarwal
- Bloom, B.S., Engelhart, M.D., Furst, E.J., Hill, W.H., & Krathwohl, D.R. (1956). Taxonomy of Educational Objectives: The Classification of Educational Goals. Handbook I: Cognitive Domain. New York: David McKay. PDF
- Anderson, L.W., & Krathwohl, D.R. (Eds.) (2001). A Taxonomy for Learning, Teaching, and Assessing: A Revision of Bloom’s Taxonomy of Educational Objectives. New York: Longman. PDF
- Krathwohl, D.R. (2002). A Revision of Bloom’s Taxonomy: An Overview. Theory Into Practice 41(4): 212–218. DOI
- Kropp, R.P., & Stoker, H.W. (1966). The Construction and Validation of Tests of the Cognitive Processes as Described in the Taxonomy of Educational Objectives. Institute of Human Learning, Florida State University. PDF
- Madaus, G.F., Woods, E.M., & Nuttall, R.L. (1973). A Causal Model Analysis of Bloom’s Taxonomy. American Educational Research Journal 10(4): 253–262. PDF
- Hill, P.W., & McGaw, B. (1981). Testing the Simplex Assumption Underlying Bloom’s Taxonomy. American Educational Research Journal 18(1): 93–101. PDF
- Kreitzer, A.E., & Madaus, G.F. (1994). Empirical Investigations of the Hierarchical Structure of the Taxonomy. In L.W. Anderson & L.A. Sosniak (Eds.), Bloom’s Taxonomy: A Forty-Year Retrospective. Chicago: University of Chicago Press. PDF
- Crowe, A., Dirks, C., & Wenderoth, M.P. (2008). Biology in Bloom: Implementing Bloom’s Taxonomy to Enhance Student Learning in Biology. CBE—Life Sciences Education 7(4): 368–381. PDF
- Momsen, J.L., Long, T.M., Wyse, S.A., & Ebert-May, D. (2010). Just the Facts? Introductory Undergraduate Biology Courses Focus on Low-Level Cognitive Skills. CBE—Life Sciences Education 9(4): 435–440. PDF
- Jensen, J.L., McDaniel, M.A., Woodard, S.M., & Kummer, T.A. (2014). Teaching to the Test… or Testing to Teach: Exams Requiring Higher Order Thinking Skills Encourage Greater Conceptual Understanding. Educational Psychology Review 26(2): 307–329. PDF
- Agarwal, P.K. (2019). Retrieval Practice and Bloom’s Taxonomy: Do Students Need Fact Knowledge Before Higher Order Learning? Journal of Educational Psychology 111(2): 189–209. PDF
- Stanny, C.J. (2016). Reevaluating Bloom’s Taxonomy: What Measurable Verbs Can and Cannot Say About Student Learning. Education Sciences 6(4): 37. PDF
Assessment that climbs the taxonomy by default.
Book a 20-minute demo using your team’s actual material. We’ll show you the five-level quiz set Future Proof generates for one of your real chapters — and the randomized mix each learner actually receives.