Key ideas
Core concepts
- Question-level data invites invalid inferences: difficulty varies by question, surface structure confuses, and samples of a few marks are unreliable.
- The data is mostly collected for reporting to leadership and rarely changes instruction, while students’ emotional reactions to marks block engagement with feedback.
Question level analysis is the practice of recording marks for each exam question in spreadsheets and reviewing the entire exam question-by-question in class. It carries a heavy time cost for minimal learning benefit.
Connected To
Feedback | Surface and Deep Structure | Practice | Formative Assessment
Why the inferences fail
Question-level data leads to false conclusions about student understanding. Questions vary in complexity, so an 80% failure rate on a Band 6 question may be appropriate rather than evidence of poor teaching. Surface structure can confuse students even when they understand the underlying concept (Chi et al., 1981). A cricket chirping scenario for data analysis might obscure whether students struggle with statistics or merely with an unfamiliar context. Small sample sizes compound these problems: four marks of trigonometry questions cannot reliably assess understanding of the topic.
There is also a gap between collection and action. Teachers spend hours recording individual question performance, but the data serves mainly for reporting to leadership. Despite knowing the results, teachers often don’t change their instruction (Wiliam, 2011).
How students respond
When students receive marks, their emotional reactions interfere with learning. Low achievers become disheartened and assume they won’t understand feedback. Students focused on marks pester teachers for additional points rather than learning from mistakes. High achievers zone out because their performance suggests they need no attention. Few students, then, enter the mindset needed to receive and act on feedback (Butler, 1988; Kluger & DeNisi, 1996).
Question-by-question review also gives only one opportunity to address each error, yet learning requires sustained practice. Students may correct their immediate mistake during class review but lack the practice needed to automate the correct approach, and often repeat the same mistakes within a week.
The administrative overhead is substantial. Teachers record individual question marks for entire classes, create spreadsheets and analysis documents, spend class time going through each question systematically, and prepare individual feedback.
Better alternatives
Instead of comprehensive question-by-question analysis, teachers can identify 2-3 common errors affecting many students, plan specific reteaching for these misconceptions, and provide focused practice on the identified problem areas (Black & Wiliam, 1998).
Rather than dwelling on test performance, errors can inform future teaching: planning additional practice opportunities for difficult concepts and integrating remediation into ongoing instruction (Hattie & Timperley, 2007).
Efficient reteaching starts by identifying error patterns across the class: which errors occurred across multiple students, and what conceptual misunderstandings they reveal. Teachers then decide which misunderstandings need addressing, provide targeted instruction for the gaps, and build remediation into future lessons.
Students can also identify their own error patterns, explain their mistakes to themselves, set goals for improvement areas, and monitor their progress regularly. This self-directed approach develops metacognitive skills alongside content knowledge.
References
Black, P., & Wiliam, D. (1998). Assessment and classroom learning. Assessment in Education: Principles, Policy & Practice, 5(1), 7-74. https://doi.org/10.1080/0969595980050102
Butler, R. (1988). Enhancing and undermining intrinsic motivation: The effects of task-involving and ego-involving evaluation on interest and performance. British Journal of Educational Psychology, 58(1), 1-14. https://doi.org/10.1111/j.2044-8279.1988.tb00874.x
Chi, M. T. H., Feltovich, P. J., & Glaser, R. (1981). Categorization and representation of physics problems by experts and novices. Cognitive Science, 5(2), 121-152. https://doi.org/10.1207/s15516709cog0502_2
Hattie, J., & Timperley, H. (2007). The power of feedback. Review of Educational Research, 77(1), 81-112. https://doi.org/10.3102/003465430298487
Kluger, A. N., & DeNisi, A. (1996). The effects of feedback interventions on performance: A historical review, a meta-analysis, and a preliminary feedback intervention theory. Psychological Bulletin, 119(2), 254-284. https://doi.org/10.1037/0033-2909.119.2.254
Wiliam, D. (2011). Embedded formative assessment. Solution Tree Press.