Research · Assessment Science
Assessment Science · Bloom’s Taxonomy

Bloom’s Taxonomy, Seventy Years On

Nearly everyone in training can recite the pyramid; almost no assessment climbs it. What the 1956 handbook actually claimed, what the 2001 revision fixed, which parts survived seventy years of testing — and how Future Proof™ builds quizzes that leave recall behind.

TL;DR

The finding: Bloom’s 1956 taxonomy and its 2001 revision remain the field’s shared language for the cognitive demand of a task. The evidence supports the ordering of its lower levels and, above all, its value as an audit tool: when researchers rate real course assessments against it, the overwhelming majority of items sit in the bottom one or two levels.

The mechanism: Recall items are cheap to write, unambiguous to grade, and easy to defend, so unmanaged item banks drift to the bottom of the taxonomy. And learners mostly improve at the level they practice — drilling facts does not automatically produce analysis.

The product: Future Proof’s quiz engine ships five Bloom levels with every chapter and draws randomized per-learner question sets, so assessment climbs past recall by design rather than by author discipline.

In 1948, at an American Psychological Association convention in Boston, a group of college examiners agreed to attempt something unglamorous: a shared classification of the things examinations ask students to do. Eight years later the group, led by Benjamin Bloom of the University of Chicago, published Taxonomy of Educational Objectives, Handbook I: Cognitive Domain (Bloom et al., 1956). It became one of the most cited works in the history of education, and its six-layer pyramid is now the most recognizable diagram in instructional design.

Seventy years on, the taxonomy occupies a strange position. Nearly everyone who builds training can recite the levels; almost nobody’s assessments climb them. Audits of real course exams keep finding the same thing: the overwhelming majority of questions sit in the bottom layer or two (Momsen et al., 2010). This article treats the taxonomy as a piece of science rather than a poster — what it originally claimed, what the 2001 revision changed, which claims held up under empirical testing, and why assessment in practice stays at recall anyway.

What the 1956 handbook actually claimed

The taxonomy was born as a measurement tool, not a theory of learning. Bloom’s committee wanted university examiners to be able to exchange test items and compare results across institutions, which required a common vocabulary for what an item demands of a student (Krathwohl, 2002). The result was six categories — Knowledge, Comprehension, Application, Analysis, Synthesis, and Evaluation — each defined not just in prose but with sample objectives and, crucially, sample test items.

The handbook made two claims that matter for evidence. The first is ordering: the categories run from simple to complex. The second is stronger — a cumulative hierarchy, in which mastery of each category presupposes mastery of the ones below it. On the strong reading, a learner cannot genuinely analyze material they cannot first recall and comprehend, and patterns of test performance should reflect that ladder (Bloom et al., 1956).

It is worth noticing what the handbook did not claim. It classified the demands of tasks, not the sequence of teaching. The later folklore — “cover the facts first, save the analysis for the end of the course” — was read into the pyramid, not out of the book.

The 2001 revision: nouns become verbs

By the 1990s the taxonomy’s limitations were well rehearsed, and a group co-chaired by Lorin Anderson and David Krathwohl — one of the original 1956 authors — undertook a full revision (Anderson & Krathwohl, 2001). Three changes matter.

First, the category names became verbs, reflecting a shift from classifying test content to classifying cognitive processes: Remember, Understand, Apply, Analyze, Evaluate, Create. Second, the single scale became a two-dimensional table. Every objective is now classified twice — by cognitive process and by the type of knowledge involved: factual, conceptual, procedural, or (new in the revision) metacognitive. “List the stages of the sales process” and “diagnose why this deal stalled” both involve the same subject matter but occupy very different cells. Third, the top of the pyramid changed: Synthesis, renamed Create, moved above Evaluate.

The quietest change was the most important. The revision relaxed the strict cumulative hierarchy: the categories are still held to differ in complexity, but they are allowed to overlap (Krathwohl, 2002) — an acknowledgment that four decades of data had not been kind to the strong ordering claim.

1956 · NOUNS 2001 · VERBS Evaluation Synthesis Analysis Application Comprehension Knowledge Create Evaluate Analyze Apply Understand Remember 2001 ADDS A SECOND AXIS — KNOWLEDGE TYPE: FACTUAL · CONCEPTUAL · PROCEDURAL · METACOGNITIVE
Figure 1. The 1956 categories (nouns) and the 2001 revision (verbs). The top two levels swapped order, and the revision added a second, knowledge dimension. After Krathwohl 2002.
The Taxonomy of Educational Objectives is a framework for classifying statements of what we expect or intend students to learn as a result of instruction. Krathwohl 2002, Theory Into Practice

What the evidence supports

The taxonomy’s testable claim is the hierarchy, and it has been tested. The first major attempt came in the 1960s, when Kropp and Stoker administered tests written at each taxonomy level, across several school subjects, to large samples of secondary students (Kropp & Stoker, 1966). Difficulty broadly increased up the levels — consistent with the simple-to-complex ordering — but the pattern of correlations between levels only partially matched a strict cumulative ladder.

Madaus, Woods and Nuttall fit causal models to those data and found a branching structure rather than a single chain, with a general-ability factor feeding the upper levels directly (Madaus et al., 1973). Hill and McGaw reanalyzed the same data with stronger methods and recovered clearer support for a cumulative ordering, though the fit was still imperfect and only held once the lowest category was set apart from the rest (Hill & McGaw, 1981). The most thorough review of this literature, in the taxonomy’s own forty-year retrospective, reaches a verdict that has not changed since: the lower levels behave roughly cumulatively; the upper levels are distinguishable but do not stack neatly (Kreitzer & Madaus, 1994).

If the ladder is only half true, why keep the taxonomy? Because its second use — as an audit and design instrument — has fared much better. Crowe, Dirks and Wenderoth built the “Blooming Biology Tool,” a rubric for rating exam items by level, and used it to realign university courses so that in-class practice matched the cognitive level of the exams (Crowe et al., 2008). Momsen and colleagues applied the same kind of rating at scale to introductory biology courses and found assessment concentrated almost entirely at the lowest levels (Momsen et al., 2010) — turning “our tests are shallow” from a vague worry into a measurable, fixable property of a course.

And the level of assessment appears to matter for outcomes. Jensen and colleagues compared course sections whose regular exams demanded higher-order thinking with sections whose exams stayed at recall; on a common final, the higher-order sections did better on the higher-order items and at least as well on the factual ones (Jensen et al., 2014). Agarwal’s experiments sharpen the point from the other direction: retrieval practice on facts alone did not improve performance on higher-order test questions, while practice on higher-order questions did (Agarwal, 2019). To a first approximation, learners improve at the level they practice.

Why most LMS assessment never leaves recall

Nothing in the taxonomy explains why recall dominates practice; economics does. A recall item can be written in a minute from any sentence of source material, has an unambiguous answer key, is trivially auto-graded, and is easy to defend when a learner disputes it. An analysis or evaluation item needs a scenario the learner has not seen before, distractors that encode plausible misconceptions, and grading criteria that survive argument. Every incentive in course production points downhill.

The audit literature shows how strong that pull is even under favorable conditions: in university courses run by trained educators, the large majority of assessment items still tested the lowest levels (Momsen et al., 2010). Typical workplace quizzes are built with less assessment expertise and tighter deadlines, by authors who are subject-matter experts rather than test designers — and question banks, once written, are copied forward for years. The predictable result is the industry’s open secret: most “knowledge checks” verify short-term recognition of the page the learner just read.

The fix suggested by the evidence is unglamorous: treat level coverage as a property of the assessment itself, decide the mix deliberately, and hold the item bank to it — the same discipline Crowe and colleagues applied when they realigned their courses (Crowe et al., 2008).

What the evidence doesn’t show

Four honest limits. First, the strict cumulative hierarchy is not supported at the top of the taxonomy; the empirical work consistently finds the upper levels branching rather than stacking (Madaus et al., 1973) (Kreitzer & Madaus, 1994). Second, the taxonomy is a classification scheme, not an intervention — there is no meaningful effect size for “using Bloom’s taxonomy,” and studies like Jensen’s test the effect of higher-order assessment, not of the framework itself (Jensen et al., 2014).

Third, the popular verb lists are weak proxies for cognitive level. Stanny’s review of published lists found the same verb assigned to different levels on different lists, and a verb alone cannot tell you whether a task is novel to the learner — “explain” is pure recall if the explanation was given in the course (Stanny, 2016). Classifying an item correctly requires knowing what the learner has already seen, which is why audit studies rely on trained raters rather than keyword matching. Fourth, the taxonomy does not license “facts first” curricula: the evidence that fact drill automatically prepares learners for higher-order performance is, at best, missing (Agarwal, 2019).

What this means for assessment design

  • Audit before you author. Rate the existing item bank by level. Most teams that do this discover they are running a recognition test, not an assessment.
  • Decide the mix, per objective. The revision’s two-dimensional table is the practical tool here: pick the cell each objective lives in, then write items for that cell.
  • Write above recall against novel material. An application item must present a situation not covered verbatim in the content, or it collapses into recall regardless of its verb.
  • Align practice with the target level. If the goal is analysis, learners need retrieval practice at analysis — not a diet of fact checks with one hard question at the end.
Applied research

How Future Proof™ applies this.

Every chapter ships with a quiz set spanning five Bloom levels — from remembering through creating — so the level mix is a property of the platform, not of author discipline. Each learner receives a randomized draw from the leveled item bank, which means no two learners see the same set and no one can pass on recall alone. Admins see performance broken out by level, so “knows the facts but can’t apply them” shows up in the data instead of on the job.

See the quiz engine
References

Selected papers.

This is not an exhaustive bibliography — these are the studies cited above. The full reading list is in the downloadable Science Library PDF.

  1. Bloom, B.S., Engelhart, M.D., Furst, E.J., Hill, W.H., & Krathwohl, D.R. (1956). Taxonomy of Educational Objectives: The Classification of Educational Goals. Handbook I: Cognitive Domain. New York: David McKay. PDF
  2. Anderson, L.W., & Krathwohl, D.R. (Eds.) (2001). A Taxonomy for Learning, Teaching, and Assessing: A Revision of Bloom’s Taxonomy of Educational Objectives. New York: Longman. PDF
  3. Krathwohl, D.R. (2002). A Revision of Bloom’s Taxonomy: An Overview. Theory Into Practice 41(4): 212–218. DOI
  4. Kropp, R.P., & Stoker, H.W. (1966). The Construction and Validation of Tests of the Cognitive Processes as Described in the Taxonomy of Educational Objectives. Institute of Human Learning, Florida State University. PDF
  5. Madaus, G.F., Woods, E.M., & Nuttall, R.L. (1973). A Causal Model Analysis of Bloom’s Taxonomy. American Educational Research Journal 10(4): 253–262. PDF
  6. Hill, P.W., & McGaw, B. (1981). Testing the Simplex Assumption Underlying Bloom’s Taxonomy. American Educational Research Journal 18(1): 93–101. PDF
  7. Kreitzer, A.E., & Madaus, G.F. (1994). Empirical Investigations of the Hierarchical Structure of the Taxonomy. In L.W. Anderson & L.A. Sosniak (Eds.), Bloom’s Taxonomy: A Forty-Year Retrospective. Chicago: University of Chicago Press. PDF
  8. Crowe, A., Dirks, C., & Wenderoth, M.P. (2008). Biology in Bloom: Implementing Bloom’s Taxonomy to Enhance Student Learning in Biology. CBE—Life Sciences Education 7(4): 368–381. PDF
  9. Momsen, J.L., Long, T.M., Wyse, S.A., & Ebert-May, D. (2010). Just the Facts? Introductory Undergraduate Biology Courses Focus on Low-Level Cognitive Skills. CBE—Life Sciences Education 9(4): 435–440. PDF
  10. Jensen, J.L., McDaniel, M.A., Woodard, S.M., & Kummer, T.A. (2014). Teaching to the Test… or Testing to Teach: Exams Requiring Higher Order Thinking Skills Encourage Greater Conceptual Understanding. Educational Psychology Review 26(2): 307–329. PDF
  11. Agarwal, P.K. (2019). Retrieval Practice and Bloom’s Taxonomy: Do Students Need Fact Knowledge Before Higher Order Learning? Journal of Educational Psychology 111(2): 189–209. PDF
  12. Stanny, C.J. (2016). Reevaluating Bloom’s Taxonomy: What Measurable Verbs Can and Cannot Say About Student Learning. Education Sciences 6(4): 37. PDF
See it on your content

Assessment that climbs the taxonomy by default.

Book a 20-minute demo using your team’s actual material. We’ll show you the five-level quiz set Future Proof generates for one of your real chapters — and the randomized mix each learner actually receives.

12 citations Reviewed July 2026 Open peer review welcomed