Back to blog
Education Research

Formative vs. Summative: Where AI Helps and Where It Gets Out of the Way

Graded paper quiz returned to student beside an open notebook

The distinction between formative and summative assessment is one of those concepts that gets cited frequently in education technology conversations without always being applied precisely. Vendors routinely claim their tools "support both formative and summative assessment" as if the two are variations on the same task. They are not. They have different purposes, different design requirements, and very different implications for what AI can and cannot usefully contribute.

The purpose distinction

Formative assessment happens during the learning process. Its purpose is to generate information that can change what the teacher does or what the student does next. A formative assessment that reveals a misconception is only valuable if someone acts on that information before the next summative event. The whole point is actionability -- if the result cannot influence instruction, it is not really serving a formative function even if it looks like a formative tool.

Summative assessment happens at the end of a learning period. Its purpose is to evaluate what a student has learned -- to assign a grade, certify a level, or produce a record of achievement. A summative assessment is not supposed to change what the teacher does in the current unit; the unit is over. Its value is in providing a defensible, accurate account of what the student knows and can do as of a particular point in time.

These different purposes create different validity requirements. A formative assessment needs to be sensitive enough to detect a gap that instruction can still address. A summative assessment needs to be accurate enough to fairly represent a student's level for grading or credentialing purposes. A tool that optimizes for one of these goals will generally not serve the other equally well.

Why AI is well-suited for formative diagnostics

AI-based tools have clear strengths in the formative assessment context. The formative loop requires frequent, low-stakes measurement -- ideally embedded in practice so that students do not experience it as a test. AI can generate adaptive question sequences, score responses instantly, identify error patterns across multiple items, and surface the resulting information for the teacher in real time. It can also do this at a frequency that would be impractical with teacher-created assessments, because generating and scoring formative check items is expensive in teacher time.

The diagnostic use case is a particular strength. A teacher who wants to know whether five students have a prerequisite gap before next week's lesson cannot realistically design, administer, score, and interpret a custom diagnostic for those students in the available time. An adaptive diagnostic system can do this continuously and present the result as a simple summary. The teacher sees the output, not the underlying assessment mechanics.

AI also handles the frequency requirement of formative assessment in a way that avoids assessment fatigue. Because adaptive diagnostics can be embedded in practice sequences that students perceive as learning activities rather than tests, the measurement happens without creating the psychological weight of a formal assessment event. This is particularly important in grades where students have developed test anxiety.

Where AI runs into limits with summative assessment

Summative assessment is a different story. The validity requirements are stricter because the stakes are higher -- a summative grade goes into a student's permanent record and may affect placement, graduation, or post-secondary opportunities. The question of whether a particular assessment instrument accurately and fairly measures what it claims to measure is a technical, legal, and ethical question that requires human expertise and judgment to answer.

AI-generated summative assessments raise several concerns that formative diagnostics do not. First, the content validity question: does the assessment actually cover the full domain that was taught, in appropriate proportions? An AI that selects questions from a bank based on observed student error patterns may produce an idiosyncratic selection that is valid as a diagnostic but does not representatively sample the summative domain. A summative test has to be interpretable as evidence about the full course objective, not just the gaps the system happened to identify.

Second, the bias and fairness question is more consequential when stakes are higher. An AI that learned from historical student performance data may reflect historical patterns in how different groups of students have performed on particular question types, potentially producing assessments that systematically disadvantage certain students in ways that are not visible from the score alone. For formative use, a biased item might at worst cause a teacher to spend an extra session on a topic. For summative use, it might affect a student's grade in a course.

The AI-human handoff in summative contexts

This does not mean AI has no role in summative assessment. It means the role is different. AI can be useful for item generation that human experts then review and curate, for automated scoring of straightforward item types where human and automated scoring agree, and for flagging items that show unusual statistical properties (too easy, too hard, discriminating in unexpected ways) for review before results are used.

What AI should not be doing in summative contexts without human oversight is making final grading decisions, calibrating difficulty against student populations in ways that could introduce bias, or generating entire assessment instruments from scratch without expert review of content validity. These are exactly the points where human judgment about fairness, context, and educational values is not optional.

The practical implication for schools considering AI assessment tools is to be clear about which function the tool is serving. "AI-assisted formative assessment" is a meaningfully different product category from "AI-generated summative assessment," and the appropriate level of vendor scrutiny is very different for each. For formative tools, the relevant questions are about accuracy, actionability, and time-to-insight. For summative tools, they are about validity, bias, and the role of human review in final grading decisions.

See gap detection in action

Join the early-access program and run a pilot with your classroom or program.