Case Dossier — Texas A&M–Commerce "ChatGPT Detector" Incident (May 2023)
500 words
Case Dossier — Texas A&M–Commerce "ChatGPT Detector" Incident (May 2023)
Opens Session 3. Work the case cold before the Stanford detector evidence lands.
What happened
In May 2023, an instructor at Texas A&M University–Commerce suspected students of using ChatGPT in a final assignment. His method: he pasted student essays into ChatGPT itself and asked it whether it had written them. ChatGPT — which has no capability to recognise its own past outputs — obligingly "confirmed" authorship for essay after essay.
On that basis the instructor told the class he was failing the submissions, and graduating seniors had diplomas temporarily withheld while the university investigated. Students protested innocence; at least one produced timestamped drafts as evidence. The story went viral. The university stated that no students ultimately failed the class or were barred from graduating, and that it was developing policies for AI use — which, at the time of the incident, it lacked.
Why it opens Session 3
Every failure the session addresses is present in one incident:
- Evidentiary failure: the "detector" had zero validity — ChatGPT cannot identify its own output, and even purpose-built detectors carry high false-positive rates. The verdict was noise delivered with confidence.
- Burden inversion: students had to prove innocence against a machine's unfalsifiable accusation.
- Consequence asymmetry: grades and degrees moved on evidence that would survive no appeals process.
- The policy vacuum: no institutional guidance existed, so one instructor's improvisation became the policy — for an entire class, at the worst possible moment.
- Process evidence worked: the student with timestamped drafts had the only defensible evidence in the room. That is the successor regime, demonstrated by accident.
Discussion prompts
- List every point where this could have been stopped. Which was cheapest?
- Would this survive your institution's appeals process? Are you sure — is the detector's evidentiary status written down anywhere?
- The instructor is not a villain; he was improvising in a vacuum. What did the institution owe him that it failed to provide?
- A student in that class honestly never used AI. Describe their week.
Facilitator notes
- Sequence: convict the method first, then generalise with Liang et al. (Patterns 2023): purpose-built detectors falsely flagged 61%+ of real TOEFL essays by non-native English writers while scoring near-perfect on US eighth-grader text — and one "rewrite like a native speaker" prompt flips the verdict. Biased one way, evadable the other: the worst possible combination for evidence.
- Then land the local stakes: in a Malaysian university, second-language English writers are not the edge case — they are the main case.
- Keep claims conservative: instructor pasted essays into ChatGPT and asked; diplomas temporarily held; university said no one ultimately failed or was blocked from graduation. Avoid naming the instructor on slides (the lesson is systemic, not personal).
- Balance against denial swing: misuse is real (Anthropic Education Report documents answer-seeking and detector-evasion requests at scale). The problem is genuine; the tool is wrong.