
Ask most institutions how their video is performing and you will get a number: total plays, minutes watched, maybe a completion percentage. Those figures feel like insight. They are not. A play count tells you a file was requested. Minutes watched tell you a browser tab stayed open. Neither tells you whether a student understood the material, got stuck, or quietly gave up at minute nine. The gap between measuring activity and measuring learning is the single biggest reason video analytics disappoint the people who buy them.

The 2025 EDUCAUSE Horizon Report on Data and Analytics makes the point in institutional terms: the value of data now depends less on collecting more of it and more on the maturity of how it is governed, interpreted, and acted on. The report urges teams to move past counting "how many" and "how much" toward the "how" and "why" behind student outcomes. This article translates that idea into a practical maturity model for video, so you can locate where your program sits today and what the next rung actually requires.
A view is a request, not a result. It fires whether the learner watched attentively, muted the tab and walked away, or scrubbed to the end to mark the item done. Watch time is only marginally better, because time in front of a video is not the same as cognitive effort. Research on instructional video keeps landing on the same conclusion: what predicts learning is not how long students watch but how they engage while watching.
A study in Computers & Education, "Students' active cognitive engagement with instructional videos predicts STEM learning," used digital trace data to show that behaviors like pausing and skipping within a video, and moments of cognitive disequilibrium, related to learning outcomes. The framing draws on Chi and Wylie's ICAP model, which sorts engagement into passive, active, constructive, and interactive modes of increasing effort. The lesson for analytics is direct: if your dashboard cannot distinguish passive playback from active manipulation of the material, it cannot tell you anything about learning. It is measuring the wrong layer.
Maturity models exist so teams can stop debating tools and start assessing capability honestly. For video, a useful ladder has five levels, each defined by the question it can answer rather than the report it can produce.
The point of the ladder is not that Level 5 is the goal for every course. It is that each level answers a different kind of question, and most institutions believe they are higher on the ladder than their data actually supports.
The first two levels are where the majority of video programs live, often without realizing it. Standard hosting platforms hand you plays and a retention graph, and those are genuinely useful for spotting a badly edited intro or a fifteen-minute video that loses everyone at minute four. But descriptive behavior has a ceiling. It can show you that a cohort drops off at 6:20. It cannot tell you whether they left confused, bored, or simply because they had already grasped the concept.
Level 2 also quietly rewards vanity. A retention curve that looks strong can hide the fact that nobody engaged with anything, they just left the video playing. Programs stall here because the next rung requires a different kind of data, not a prettier chart of the same data.
Moving up means capturing what students do inside the timeline, not just how far they scrolled. This is the leap from watch data to engagement data. When learners can ask questions, leave notes, reply to peers, and answer checks at the exact second a concept appears, every one of those actions is a timestamped signal of cognitive engagement. That is the layer ICAP research says actually matters.
Annoto's analytics are built for this level. Because discussions, notes, and quizzes are anchored to specific moments, the data distinguishes a video that was played from one that was worked through. A cluster of questions at 6:20 is diagnostic in a way a drop-off point never is: it tells you the concept, not just the coordinate. Aggregate those signals and Level 4 comes into reach. Silence where you expected discussion, questions that never got answered, and progress that stalls after Week 1 all function as early indicators of who needs help, well before a low grade confirms it.
Analytics that never change anything are an expensive hobby. The top of the ladder is where a signal reliably produces an action and you can see the action's effect. In practice that means routing an in-video question to a TA or to Lumo, prompting an instructor to reach out to students who went quiet, flagging a confusing segment for a re-record, and then checking whether the next cohort's engagement at that timestamp improved.
Closing the loop is also where video analytics connect to institutional systems. Engagement signals that feed back into the LMS gradebook and into advising turn an isolated dashboard into part of a student-success workflow. The EDUCAUSE report's emphasis on mixed methods applies here too: pair the quantitative drop-off with a qualitative read of what students actually wrote, and you understand the why, not just the what.
You cannot climb the ladder on a shaky foundation. The EDUCAUSE report is blunt that institutions which fail to build strong data literacy and governance risk falling behind. For video analytics that means clear ownership of the data, consistent definitions of what an "engaged" learner is, defensible privacy practices, and instructors who are trained to read the signals rather than drown in them. A Level 5 capability layered on Level 1 governance produces confident, wrong decisions. Maturity is the whole stack moving together.
Use a short self-assessment. For a representative course, ask which of these you can answer with evidence:
The first level you cannot answer honestly is your real position. Most programs discover they can produce Level 2 charts but cannot answer a single Level 3 question, because the engagement layer that would generate those answers does not exist in their video yet.
That is the practical takeaway. More dashboards will not move you up the ladder. Richer signals will. The move from counting views to understanding learning starts by making video something students act inside, then reading those actions as the evidence they are. If your team is ready to see what Level 3 and beyond looks like against your own courses, explore how Annoto turns video assignments into measurable engagement across higher education programs.
Get the best resources and research on active video learning — in your inbox, once a month.