Early-warning systems that actually warn early.
A decade after Purdue’s Course Signals lit up the field, we know what at-risk prediction can do, where its most famous claim fell apart, and how it quietly fails. It is the research programme behind Future Proof™’s AI Risk Forecaster — which flags drift weeks early, from practice signals, so teams intervene before failure rather than after.
The finding: Modern learning-analytics systems can flag academically at-risk learners well before a course ends, using effort signals — logins, submissions, activity — alongside grades and history. When those flags are paired with a real intervention, the intervention can help. The prediction is the easy half; acting on it well is the hard half.
The mechanism: Risk is legible early because disengagement leaves a trail before it reaches a gradebook. A learner who stops submitting, or whose activity decays relative to peers, is drifting weeks before the failing grade lands. A model that watches the trajectory — not just the current score — can raise the flag while there is still time to act.
The product: Future Proof’s AI Risk Forecaster flags drift weeks early from practice signals so teams intervene before failure, not after — a trajectory watcher, not a post-mortem dashboard.
The modern field of learning analytics has a founding artifact, and it is a traffic light. In 2012, Kimberly Arnold and Matthew Pistilli presented Course Signals, a system built at Purdue University that mined the learning-management system for signals of effort — logins, submissions, activity relative to classmates — combined them with grades and prior academic history through a proprietary risk algorithm, and rendered the output to each student as a red, yellow, or green light, alongside a personalized note from the instructor (Arnold & Pistilli, 2012). It was a simple idea with an enormous premise: that a learner about to fail is legible before the failure, and that surfacing that fact — early, and to a human who can act — changes the outcome.
That premise has held up better than the specific claims first made for it. In the years since, at-risk prediction has become one of the most studied — and most oversold — applications in education technology. This piece is about what the evidence actually supports: that early flagging works when it is built on the right signals, that the field’s most cited retention number did not survive scrutiny, and that the interesting failures are not in the models but in what surrounds them.
What early prediction can actually do
Start with the part that is genuinely robust. A model trained on activity and grade data can identify a large share of the learners who will struggle, and it can do so while the term is young enough for the flag to matter. The clearest early demonstration of this at scale came from the Open Academic Analytics Initiative, whose flagship study — Jayaprakash and colleagues — built an open-source early-alert model at one institution and then deliberately tested whether it ported to others (Jayaprakash et al., 2014). This matters because a model that only works where it was born is a research demo, not an early-warning system.
Their finding was two-sided and honest. The predictive model identified a substantial majority of at-risk students and retained useful accuracy when deployed at partner colleges with different populations — but that accuracy degraded on transfer, and the interventions that followed the flag produced mixed results: flagged students who received an intervention withdrew from courses at a higher rate than a control group even as their content mastery held or improved (Jayaprakash et al., 2014). Read carefully, that is not a failure of prediction. It is the first clear signal that predicting risk and fixing it are different problems.
The raw materials of prediction are now well understood, in part because of open resources like the Open University Learning Analytics Dataset, which pairs demographic data with daily clickstream summaries for tens of thousands of learners and has become a common benchmark for this work (Kuzilek, Hlosta & Zdrahal, 2017). The lesson from a decade of building on such data is consistent: the trajectory carries the signal. A single low score is ambiguous. A downward slope in submissions or activity, relative to a learner’s own baseline and their cohort, is what separates a bad week from a learner drifting toward the exit.
The retention claim that unravelled
No account of this field is honest without the Course Signals controversy, because it is a case study in how a good tool can be attached to a bad number. Purdue publicized a striking result: students who took two or more Course Signals–enabled courses graduated at markedly higher rates than students who took none — a headline retention lift that circulated widely and helped make the case for learning analytics across higher education (Arnold & Pistilli, 2012).
In 2013, that claim came apart under external scrutiny. Michael Caulfield first flagged an anomaly, and Alfred Essa built it into a formal simulation: the retention comparison was vulnerable to reverse causality (Caulfield, 2013) (Essa, 2013). Students who dropped out did so partway through their studies, which mechanically limited how many Course Signals courses they could have accumulated — so taking fewer of these courses was partly a consequence of leaving early, not a cause of it. Sort students by exposure to the treatment when exposure is itself a function of how long they stayed enrolled, and you will manufacture a retention effect out of nothing but the calendar. Essa’s simulation reproduced Purdue’s reported pattern without any causal benefit at all (Essa, 2013).
The correct lesson is narrow and important. What unravelled was a specific retention statistic, produced by a correlational design that confounded treatment with tenure. The underlying capability — flagging at-risk learners early from behavioural signals — was never what failed. The episode is a permanent reminder that in this field the model is rarely the weakest link; the causal claim wrapped around it usually is.
Learning analytics should not promote one size fits all.Gašević, Dawson, Rogers & Gašević, 2016
Why the same model fails in the next course
The most consequential failure mode is not fraud or hype; it is portability. A risk model learns the behavioural fingerprint of struggle in the courses it was trained on — and courses differ enormously in how, and how much, learners are expected to touch the system. Dragan Gašević and colleagues made this concrete across nine undergraduate courses: predictors that mattered in one course were irrelevant or reversed in another, and a single generalized model built by pooling everything performed worse for guiding practice than course-specific models did (Gašević et al., 2016). Their title has become the field’s watchword — learning analytics should not promote one size fits all (Gašević et al., 2016).
The mechanism is mundane and unavoidable. In a course where all coursework runs through the LMS, low activity is a strong danger sign. In a course that meets in person and uses the LMS only to post slides, the same low activity means nothing — the diligent student and the absent one look identical to the log. A model that does not know the instructional context of a course will confidently mislabel learners in it. This is why an early-warning system is not a model you buy once; it is a model that has to be re-fit and re-validated against the actual shape of each learning environment it watches.
The bias hiding in the flags
The second failure mode is fairness, and it is easy to miss precisely because aggregate accuracy can look fine while subgroup accuracy does not. When Renzhe Yu and colleagues systematically compared data sources for predicting college success, they found that institutional and LMS data both carried real predictive power — and that including personal-background features could make predictions less fair for some groups, in one case pushing international students toward worse predicted outcomes than their actual performance warranted; models built on behavioural clickstream data were fairer for those learners than models leaning on demographics (Yu et al., 2020). Related work on college-dropout prediction has found that racialized disparities in error rates can persist across model designs, and are not reliably eliminated simply by adding or removing sensitive attributes (Yu et al., 2021).
The practical implication is uncomfortable and specific: an early-warning system can be right on average and systematically wrong for a subgroup, over-flagging one population as at-risk while missing genuine risk in another. If those flags route real resources — an advisor’s time, a mandatory check-in, a place in an intervention programme — the errors are not abstract. This is why fairness in early warning is not a compliance checkbox appended after training; it has to be measured per subgroup, on the outcome that will actually be acted on.
What the evidence doesn’t show
It is worth being precise about the limits of this literature, because the field’s own history is a warning against overclaiming.
- Prediction is not intervention. A model that flags at-risk learners accurately tells you nothing, on its own, about whether the follow-up action helps. In the OAAI trials, flagged-and-treated students sometimes withdrew more even as their learning held — a reminder that a well-aimed flag attached to a blunt intervention can backfire (Jayaprakash et al., 2014).
- The headline retention numbers are fragile. The most-cited retention figure in the field’s history did not survive causal scrutiny (Caulfield, 2013) (Essa, 2013). Correlational “students who used it did better” comparisons are the default trap; only designs that handle self-selection and tenure should be trusted.
- Models do not transfer for free. Predictive power is course- and context-specific, and pooling data into one universal model can degrade the guidance it offers (Gašević et al., 2016). Accuracy reported on one population is not a promise about the next.
- Aggregate accuracy can hide subgroup harm. A model can be broadly accurate and still misclassify particular groups at higher rates, and adding or dropping demographic features does not reliably fix it (Yu et al., 2020) (Yu et al., 2021).
- Effect sizes at scale are modest and implementation-bound. Even mature analytics deployments tend to show gains that depend heavily on how the surrounding intervention is run, not on the raw quality of the prediction. The prediction is necessary; it is nowhere near sufficient.
None of this argues against early-warning systems. It bounds what they are for. Prediction buys time — the weeks between the first sign of drift and the grade that confirms it. What a team does with that time is a separate, harder question, and it is the question that determines whether an early-warning system warns early or merely records failure faster.
What “early” has to mean
The word doing the work in “early-warning system” is early. A dashboard that turns a learner’s cell red the week after they fail an assessment is not an early-warning system; it is a late-warning system with better graphics. The whole value is in the lead time — and lead time comes only from watching trajectory, validated against the specific environment, checked for who the flags fall on, and wired to an intervention someone has actually tested. A system that gets the prediction right and the other three wrong will produce confident, timely, unfair flags that no one acts on well. The literature’s clearest verdict is that all four have to be right at once.
How Future Proof™ applies this — the AI Risk Forecaster.
The AI Risk Forecaster is built for the one thing this literature says matters: lead time. It watches trajectory, not snapshots — decay in practice signals against each learner’s own baseline and their cohort — so drift surfaces weeks before it reaches a completion metric. It is fit per environment, not a single pooled model, because a signal that means struggle in one program means nothing in another. Flags are checked for who they fall on, monitored across groups so the system does not route attention unfairly. And every flag lands as an action, not a red cell: a specific suggested intervention tied to the concept that is slipping, so a manager or coach can step in while the window is still open. Prediction is the easy half; Future Proof is built around the hard half.
See the AI Risk Forecaster →Selected papers.
This is not an exhaustive bibliography — these are the studies cited above. The full reading list is in the downloadable Research Library PDF.
-
Kuzilek, J., Hlosta, M., & Zdrahal, Z. (2017). Open University Learning Analytics dataset. Scientific Data 4: 170171. DOI
-
Essa, A. (2013). Can we improve retention rates by giving students chocolates? A simulation of the Course Signals reverse-causality problem. Analysis and simulation of the Purdue Course Signals retention data, 2013. PDF
-
Gašević, D., Dawson, S., Rogers, T., & Gasevic, D. (2016). Learning analytics should not promote one size fits all: The effects of instructional conditions in predicting academic success. The Internet and Higher Education 28: 68–84. DOI
-
Yu, R., Li, Q., Fischer, C., Doroudi, S., & Xu, D. (2020). Towards Accurate and Fair Prediction of College Success: Evaluating Different Sources of Student Data. Proceedings of the 13th International Conference on Educational Data Mining (EDM 2020): 292–301. PDF
The failing grade is the one signal that comes too late.
Book a 20-minute demo using your team’s actual content. Watch the AI Risk Forecaster read a real learner’s trajectory, flag the drift weeks before a completion metric would, and hand your team a specific action while the window is still open.