Case Dossier — UK Ofqual Grading Algorithm (2020)
491 words
Case Dossier — UK Ofqual Grading Algorithm (2020)
Opens Session 2. Work the case cold before naming any principle.
What happened
With 2020 A-level exams cancelled during the COVID-19 pandemic, England's exams regulator Ofqual used a standardisation algorithm to award grades. Teachers submitted centre-assessed grades (CAGs) and student rankings; the algorithm adjusted them primarily against each school's historical results distribution.
When results were released on 13 August 2020, roughly 39% of teacher-assessed grades were downgraded (about 35% by one grade, some by two or more). The adjustment pattern was regressive: large state-school cohorts — where the algorithm leaned hardest on historical school performance — absorbed the downgrades, while small cohorts (typical of private schools) defaulted heavily to teacher assessment; the proportion of A/A* grades rose fastest at independent schools. High-achieving students at historically weak schools were the canonical victims: predicted A, awarded C, university place lost.
After days of protests and legal challenge threats, the government U-turned on 17 August 2020: grades reverted to teacher assessment. Scotland had already made an equivalent reversal for its SQA results. The head of Ofqual and the top civil servant at the Department for Education subsequently left their posts.
Why it opens Session 2
Nobody in the room disputes something went wrong. The work is naming what:
- The algorithm graded students partly on which school they attended — group history determined individual outcomes (justice).
- Individual students could not see, contest, or appeal the mechanism in time (explicability, autonomy).
- Harm was concentrated and real: lost university places (non-maleficence).
- The system was built to defend a statistic (grade inflation) rather than promote student outcomes (beneficence).
- Accountability was diffuse until it was forced — regulator, ministers, algorithm designers all pointed elsewhere (the responsibility triad).
The five principles are not taught; they are extracted — labels for what participants already said.
Discussion prompts
- What exactly failed: the algorithm, its objective function, or the decision to use it at all?
- Who was harmed, and was the harm distributed fairly?
- Who was responsible — and who was held responsible?
- Your institution deploys an automated system touching grades or admissions tomorrow. What would have to be true for it to be legitimate?
Facilitator notes
- Ofqual is not an "AI" system in the LLM sense — that's a feature. It shows the ethics is about automated judgment over students, not about any particular technology. Pair with St George's (same failure, 40 years earlier) to prove the pattern predates the hype.
- Keep numbers conservative: ~39% downgraded, U-turn on day 4, disproportionate protection of small (independent-school) cohorts. Details verified against contemporaneous reporting; do not invent precision beyond this.
- Bridge line to principles: "You just named five things. Oxford's AI4People synthesis calls them beneficence, non-maleficence, autonomy, justice, explicability. You didn't need the framework to see the failure — the framework is a memory aid for what you already know."