Two students receive 85% on the same assessment.
On the spreadsheet, they look almost identical.
One of them was struggling several months ago and has made substantial progress. She now understands most of the core ideas, although she still finds it difficult to explain her reasoning independently.
The other usually performs above 90%. This time, he has misunderstood an important concept that he previously seemed to understand.
Same assessment. Same score. Very different learning stories.
There is nothing inherently wrong with grades. Test scores and standardized assessments give schools useful information, and we need ways to record achievement, communicate results and monitor performance.
The problem comes when we ask a number to tell us more than it can.
A score tells us something about performance. It does not, on its own, explain what a student understands, how much progress has been made, why something is difficult, what should happen next, or whether the teaching that follows actually helps.
For that, we have to look more closely.
The Compression Problem
Learning is messy in ways that numbers are not.
A student can understand a concept but struggle to explain it. Another can answer familiar questions confidently and then struggle when the same idea appears in an unfamiliar context. A student who is still below the expected standard may, in fact, have made more progress than almost anyone else in the class.
Once all of that is compressed into a number, much of the story disappears.
Take a student who receives 76% on a reading assessment. Perhaps the student understands explicit information but struggles with inference. Perhaps the inference is perfectly reasonable but the supporting evidence is weak. Vocabulary might be getting in the way. The student might understand the text in discussion but struggle to put that understanding into writing.
Or perhaps it was just a bad day.
The 76% is still useful information. It just isn’t an explanation.
What Are We Actually Measuring?
Before we interpret a result, we need to be clear about what the task was supposed to measure in the first place.
Imagine a science lesson in which students are expected to understand why a particular reaction occurs.
A multilingual student understands the science but struggles to write an extended explanation in English. The final response is weak.
What does that result represent?
It may tell us something about scientific understanding. It may tell us something about academic writing or English proficiency. Quite legitimately, it may tell us something about all three.
What matters is whether that is what we intended to assess.
We cannot observe understanding directly. We see what students say, write, create or do, and from that we make a judgment about what they know or can do. There is always some inference involved.
That is why the task itself matters. Students need a reasonable opportunity to demonstrate the learning we are interested in, and we need to be careful not to make claims that go further than the evidence allows.
This may sound like an issue for assessment specialists, but it is an everyday classroom question.
What should students know or be able to do?
What would genuinely show me whether they can do it?
Those two questions can prevent a surprising amount of confusion.
They also remind us that not everything worth learning fits neatly into a spreadsheet.
Achievement Is Not Progress
Achievement and progress are related, but they are not the same thing.
Achievement tells us where a learner is now. Progress tells us how that position has changed.
A student can remain below the expected standard while making substantial progress. Another can remain one of the strongest students in the class while making relatively little progress.
Both pieces of information matter. Neither gives us the whole picture.
This is one reason assessments such as MAP Growth include information about achievement as well as growth. MAP Growth uses its RIT scale to describe achievement, while results across administrations allow educators to examine change over time. Growth norms provide additional context for interpreting that change.
Useful information, certainly. But still another lens on learning rather than a complete picture of the learner.
From Information to Understanding
Most schools are not short of information.
We have tests, quizzes, rubrics, standardized assessments, exit tickets, written work, projects, presentations, observations and conversations with students.
Often the harder part is deciding what all of it means.
A score, response or observation can tell us something, but only in relation to the learning we are trying to understand. Several relevant pieces of evidence pointing in the same direction may strengthen our interpretation. Simply collecting more, however, does not automatically make that interpretation better.
There is also an important difference between what we observe and what we conclude from it.
Suppose a student selects weak supporting evidence in three different reading tasks.
That is an observation.
We might conclude that the student understands inference but has difficulty judging which evidence best supports it.
That is an interpretation.
It may be a good interpretation. It may even turn out to be correct. But it is still an interpretation, and another explanation may fit the evidence.
So instead of asking only, What does this test tell me?, it is often more useful to ask:
What does the evidence seem to suggest, and how confident am I in that interpretation?
Sometimes the answer is simply: I don’t know yet.
That is not a failure of assessment. Sometimes it is the most responsible conclusion we can reach.
The picture can change over time too. Evidence that described a learner well six weeks ago may no longer describe what that student can do today. Students learn. Difficulties change. Strengths develop.
Our judgments need to be allowed to change with them.
Good evidence should make us curious before it makes us certain.
When Evidence Actually Changes Something
One of the strange things about assessment is how much of it can happen without changing very much.
We teach. We assess. We grade. We record the result. Then the class moves on.
There may be a beautifully maintained spreadsheet at the end of the process, but the more useful question is what happened because of what we learned.
Evidence → Interpretation → Action → New Evidence
In practice, that means asking four questions:
What is happening?
What might it mean?
What should we do?
Did it work?
That last question is easy to neglect.
If I change my teaching because I believe I have identified a problem, I eventually need some indication of whether that change helped. Otherwise, I have responded to evidence without learning anything about the quality of my response.
Assessment can therefore tell us something about our teaching as well as about our students.
Evidence does not make the decision for us. It gives us better grounds for making one.
A Practical Way to Think It Through
For more complicated situations, I find it useful to make the reasoning visible:
Intended Learning → Learning Experience → Evidence → Interpretation → Current Level and Progress → Pattern → Learning Need or Barrier → Instructional Response → Next Goal → New Evidence
The sequence looks detailed on paper, but the questions behind it are quite ordinary.
Intended Learning
What should students know, understand or be able to do?
Learning Experience
Have students actually had sufficient and appropriate opportunities to develop that learning?
Evidence
What have students said, written, created or done, and how well does it relate to what we wanted them to learn?
Interpretation
What might this mean? Is there another reasonable explanation? How confident am I?
Current Level and Progress
What can students do now, and what appears to have changed over time?
Pattern
Is the same strength or difficulty appearing again? Repetition makes a pattern more interesting, but it still does not tell us why it is happening.
Learning Need or Barrier
What seems to be helping or limiting further progress? The answer may lie in prior knowledge, language, the task, instruction, opportunity to learn, the learning environment, or somewhere else.
Instructional Response
What would be a sensible response? If the evidence is weak, perhaps the next step is to find out more before changing anything.
Next Goal
What would be the most useful next step for the learner?
New Evidence
What happened after the response? Did it help? Has it changed our understanding?

This is not a psychometric model, and I certainly would not suggest turning it into another form teachers have to complete.
It is simply a way of slowing down the reasoning when the answer is not obvious.
Most classroom decisions do not require this level of analysis. Teachers make hundreds of judgments every day, often quickly and effectively. The framework becomes useful when evidence conflicts, a difficulty keeps appearing, the reason is unclear, or the decision matters enough to warrant a closer look.
And it is a loop, not a checklist. New evidence may confirm what we thought, or it may send us back to reconsider the whole thing.
A Simple Classroom Example
Imagine the learning objective is:
Students can make an inference and support it with relevant textual evidence.
After a short reading task, several students receive weak scores. The teacher also has an exit ticket, notes from class discussion and some previous written work.
The obvious response might be to reteach inference.
But when the work is examined more closely, most of those students are actually making reasonable inferences. What keeps going wrong is the evidence they choose to support them.
That changes the lesson.
Instead of starting inference again from the beginning, the teacher might show three pieces of textual evidence and ask:
Which one best supports this inference, and why?
Students compare them, explain why one is stronger than another, and then try again independently.
For some students, inference itself may still be the problem. They need something different.
That is precisely the point. The same weak score does not necessarily call for the same teaching response.
A short task at the end of the lesson then gives the teacher something new to look at.
Did the change help?
Now assessment is doing more than recording what happened. It is shaping what happens next.
Where AI Might Help
Doing this well takes time.
A teacher with 100 students can quickly accumulate quizzes, writing samples, observations, presentations, exit tickets, assessment results and informal notes. Even when the evidence is useful, there may simply be too much of it to examine closely every time.
This is where AI becomes interesting.
We often talk about AI making assessment faster: generating questions, marking work or drafting comments. Those uses may save time, but I think there is another possibility that matters more.
AI can help us look at evidence before we decide what it means.
For example, it can organize information from several sources, compare evidence over time, highlight recurring patterns, point out contradictions or suggest explanations we may not have considered.
Suppose a student repeatedly performs poorly on written reading responses. It is easy to conclude that the student has weak reading comprehension.
Perhaps that is true.
But perhaps vocabulary is limiting understanding. Perhaps the student understands the text when speaking but struggles to express the same reasoning in writing. Perhaps the student misunderstands what the question is asking.
AI can help put those possibilities on the table.
It cannot tell us which one is true simply because it can describe each one convincingly.
Useful Support, Not the Final Judgment
This distinction matters.
AI can organize evidence and help us notice things. The teacher still needs to return to the actual student work, consider the context and decide whether the suggested interpretation makes sense.
There is another risk here that is easy to underestimate. AI is very good at turning scattered information into a coherent explanation. Sometimes the explanation sounds more certain than the evidence deserves.
A plausible story is still a hypothesis.
One useful approach is to ask AI to:
- identify patterns across several sources of evidence;
- separate observations from interpretations;
- show what supports and contradicts each interpretation;
- suggest other possible explanations;
- point out what information is missing;
- identify conclusions that are not well supported;
- suggest possible teaching responses, or say when more evidence would be useful first;
- suggest what to look for next.
I would add one instruction explicitly:
Do not make a final judgment about the student. Give me possibilities to check against the original evidence.
And there is a limit that matters regardless of how sophisticated the AI becomes.
AI cannot rescue weak evidence.
If the original evidence is incomplete, biased, poorly matched to the intended learning or missing important context, AI may simply produce a more polished interpretation of a weak foundation.
Its role should be to make working with evidence more manageable, not to lower the standard of evidence we expect.
Any use of student information with AI also needs to follow school policy, privacy requirements and approved systems.
Looking Beyond One Classroom
The same way of thinking can be useful beyond an individual student.
A pattern might appear in one class, across a grade, within a subject, across several subjects, or eventually at school level:
Learner ↔ Classroom ↔ Subject/Grade ↔ Across Subjects ↔ School
The questions do not change very much as the scale grows.
What are we seeing? What might explain it? What should we change? And what will we look at later to see whether the change helped?
When Subjects Start Seeing the Same Thing
Imagine English teachers noticing that students struggle to support interpretations with textual evidence.
Science teachers find that students can state conclusions but struggle to justify them using experimental evidence.
History teachers see claims that are not well supported by sources.
Mathematics teachers find students who can calculate answers but have difficulty explaining why their solutions work.
There might be something worth investigating here.
Perhaps students are struggling more broadly to construct and communicate reasoning supported by evidence.
Perhaps not.
Reasoning in mathematics is not the same as reasoning in history, and experimental evidence in science does not work in the same way as textual evidence in English. Similar language can make different disciplinary problems look more alike than they really are.
There is another check worth making too. Are the same students showing this difficulty in different subjects, or are school averages creating what looks like a common pattern?
We do not need a complicated data exercise to begin investigating it.
Ask teachers from several subjects to bring two or three anonymized examples of student work. Put them on the table and look together.
What does successful performance look like in each subject? Where do we genuinely see common difficulties? Where are the differences important? Is there anything worth reinforcing more consistently?
Then look again several weeks later.
That conversation is likely to tell us more than another meeting spent comparing averages on a screen.
If Your School Uses MAP Growth
MAP Growth adds another useful source of information.
Its achievement results help describe current performance, while results across administrations allow educators to examine growth over time. Recent MAP Growth reporting can also connect performance with standards and help inform conversations about instructional next steps.
It should still be treated as one part of a larger picture.
When a MAP result catches our attention, useful questions are:
What does MAP seem to suggest?
Do I see the same thing in the student’s classroom work?
Is there evidence that complicates that interpretation?
Do I know enough to act yet?
What should I look for next?
Small differences or changes should also be treated cautiously rather than read as precise indicators of learning.
MAP becomes much more useful when it is considered alongside student work, classroom assessment, observation, discussion and what students themselves say about their learning.
An assessment can tell us something important without telling us everything we need to know about the learner.
From Classroom Evidence to School Improvement
At school level, the questions become bigger, but the temptation to jump to conclusions remains much the same.
Reading growth may have been weaker for several years. Students may know a great deal of factual content but struggle when they have to transfer it. One group may appear to be progressing differently from another. A new intervention may seem to be working.
Each of those observations is a reason to investigate, not a conclusion.
Lower results do not automatically mean poor teaching. Higher results do not prove excellent teaching. One disappointing assessment does not demonstrate curriculum failure. Similar difficulties across several subjects do not automatically share the same cause.
At school level, the process might look something like this:
School Outcomes → Evidence → Patterns → Investigation → Priority → Response → Implementation → New Evidence → Impact
Schools are usually quite good at documenting implementation.
We introduced a new curriculum. Teachers completed professional development. We launched an intervention. We adopted a new assessment system.
All of that may be worthwhile.
But implementation is not impact.
Eventually we need to ask what changed for students and what evidence gives us reasonable grounds to think our actions contributed to that change.
That is a much harder question, but it is the one that matters.
Keep the Learner in the Conversation
With all this talk of evidence, interpretation and systems, it is surprisingly easy for the student to disappear.
Assessment should not become something adults do about learners while the learners themselves wait for the result.
Students can increasingly learn to recognize what they are working toward, where they have improved and what they need to do next.
There is a considerable difference between a student saying:
“I got 76%.”
and:
“I can make an inference, but I still need to choose evidence that directly supports it. That’s what I’m working on next.”
The second does not make the score irrelevant. It makes the learning more visible.
Student reflection is not automatically accurate, and it does not replace other evidence. But students can tell us things about their experience of learning that a spreadsheet cannot.
They should be part of the conversation.
Try It With One Outcome
None of this requires a new assessment system.
Pick one important learning outcome next week.
Ask yourself:
What do I actually want students to learn?
What would give me useful evidence of that learning?
What seems to be happening?
What will I do because of what I found?
What will I look at next to see whether it helped?
If AI is useful, let it help you organize or question the evidence. Keep the interpretation and teaching decision with the teacher.
Then, at the end, ask one more question:
Do I understand my students’ learning better than I would have from the score alone?
If you do, the assessment has already become more useful.
Back to the Two Students
Both students received 85%.
The school may still need the percentage. Parents may want to see it. The students may want to know it.
That’s fine.
But we now know that the same number can sit on top of two very different stories.
For the first student, 85% sits alongside substantial progress, growing understanding and a clear next step.
For the second, it raises a question. Something that previously appeared secure may no longer be consistently demonstrated, and we need more evidence before deciding why.
The number has not changed. What we understand about the learning has.
I do not think the future of assessment requires us to get rid of grades, tests or standardized measures. It requires us to be clearer about what they can tell us and what they cannot.
A useful assessment should eventually lead somewhere.
Sometimes it confirms what we thought. Sometimes it changes the next lesson. Sometimes it makes us look again. Sometimes the most responsible conclusion is that we do not know enough yet.
But the questions remain much the same:
What is happening?
What might it mean?
What should we do?
Did it work?
Collecting evidence is only the beginning.
Its real value lies in what happens next.
References and Further Reading
National Research Council. Knowing What Students Know: The Science and Design of Educational Assessment. National Academies Press, 2001.
OECD. Unlocking High-Quality Teaching. OECD Publishing, 2025.
NWEA. 2025 MAP Growth Norms Technical Manual. 2026.
NWEA. “Get Guidance on Standards-Aligned Instruction with the New State Standards View in the MAP Growth Class Profile Report.” 2026.
UNESCO. AI Competency Framework for Teachers. 2024, updated 2026.
UNESCO. Guidance for Generative AI in Education and Research. 2023, updated 2026.
Accrediting Commission for Schools, WASC. Guidance on effective progress reporting and continuous school improvement.
