TFIS · AI Ethics in Education · Course
Hands-on Workshop:
Ethical AI Tools — Part 2
Teaching, Assessment and Student Learning
Session 5 · Wednesday 2 September 2026 · 2.30–4.30 afternoon
The one study to remember · deliver slowly
A randomized controlled trial. Nearly 1,000 students. Real math curriculum.
- Three arms: GPT Base (vanilla chat interface) · GPT Tutor (same model, system prompt engineered with teacher-designed hints and a hard rule against giving final answers) · control (no AI).
- High-school students, Turkish school, published in PNAS.
Bastani, Bastani, Sungu, Ge, Kabakcı & Mariman, PNAS 122(26), 2025 — Penn/Wharton
During assisted practice
Every dashboard said AI was working
+48%
GPT Base — practice performance
+127%
GPT Tutor — practice performance
Then the tools were removed for an unassisted exam
−17%
GPT Base students vs students who never had AI at all
≈ 0
GPT Tutor students: statistically indistinguishable from control — the harm was engineered away. (Honest note: the unassisted gain was also near zero. Guardrails bought safety, not superpowers.)
They had practiced asking, not solving. The paper's word: "crutch."
Bastani et al., PNAS 2025
Three lessons, stated as axioms
- Performance ≠ learning (A5). Assisted metrics can be actively misleading about skill acquisition.
- Design determines outcome. The same model harmed or didn't, depending on a system prompt. Ethics lives in configuration, not in the technology.
- The dashboard lies in one direction. The damage only appears when the tool is removed — so build tool-removal into your assessments. That is what the three-lane pattern does.
Discernment cross-reference: MIT's EEG preprint (S2) points the same way at the neural level — but Bastani is the RCT you can defend in senate. Teaching that difference is the skill.
Discussion · 5 min · whole room
What if there were no grades?
FACILITATOR FRAME
"MIT's AI & Education report — August 2026, just published — asks this question seriously. Their committee noted that if MIT did not have grades, most incentives around AI cheating would disappear. MIT already refuses to award summa/magna/cum-laude diplomas. What if the grading system itself is the vulnerability?"
- The UK system uses percentages to express relative mastery against an expert standard — comparison, not competition.
- Competency-based and mastery-based assessments check: can you do the thing? — not can you outscore your classmates?
- When students optimise for GPA, AI-incentive is structural. When they optimise for mastery, it weakens.
MIT Ad Hoc Committee on AI Use in Teaching, Learning, and Research Training, §3.1.6 "Reconsider grades and incentives", 13 Aug 2026.
Design vocabulary 1 · seven roles for AI in learning
| Role | Crutch-effect exposure |
| Mentor (feedback) · Coach (metacognitive prompts) · Student (the learner teaches the AI — powerful, underused) | Low — these increase cognitive work |
| Simulator (practice scenarios) · Teammate | Moderate — depends on design |
| Tutor (guided instruction) · Tool (as answer-machine) | High — exactly where guardrails must live |
Role selection is a Delegation decision.
Mollick & Mollick (Wharton, 2023)
Design vocabulary 2 · the course's core move
Every assessment declares its lanes
| Lane | Rule | Protects |
| 🔴 AI-Restricted | No AI; done live / supervised | Foundational skill the student must own unaided — the "exam" from the PNAS study |
| 🟡 AI-Permitted | Allowed with disclosure statement + process trail | Authentic practice — mirrors professional reality |
| 🟢 AI-Required | Mandatory; evaluated on the interaction (prompt log + critique + improvement) | AI fluency itself as a learning outcome |
A course is well-designed when every learning outcome is protected by at least one 🔴 component and every graduate passes through at least one 🟢.
Lab 3 · Redesign sprint · 40 min · pairs
Input: the assessment from your S3 ticket
- Audit (10 min): run your own assessment through AI (P09). Grade the output honestly. If it scores ≥B, it's confirmed AI-vulnerable — say so out loud; naming it is the unlock.
- Redesign (20 min): split into components, assign lanes. Constraints: ≥1 🔴 protecting the core outcome · every 🟡 names its process evidence · total student workload flat or lower.
- Swap-test (10 min): attack your partner's redesign wearing a student hat — "How would I hollow this out with AI?" Patch the best exploit.
Output: one-page redesign sheet (template D3) → goes into the S7 portfolio.
Lab 4 · Guardrailed tutor · 30 min · individual
Build the GPT-Tutor pattern yourself
- Use P10: a Socratic tutor for one topic you teach — never gives final answers, requires the student's attempt first, one hint or guiding question per turn, asks for explanation back after breakthroughs.
- Test protocol (this is the assessment): switch to student mode and try to break it — demand the answer, plead deadline, claim confusion. Holds under three social-engineering attempts = pass.
- Compare with Claude's Learning Mode live if time allows — same design philosophy, productized.
SAY
"You just closed the gap from this morning's study with fifteen lines of instruction. The difference
between the −17% tool and the safe tool was never money or model. It was intent, written down.
Keep that feeling for tonight — that's what a policy is."
Lab 5 · The policy paragraph · 15 min · individual
≤150 words, one real course document
- Use P11 as drafting partner, then edit into your own voice.
- Contents: lane declarations per assessment · the disclosure norm (your P05 template) · one sentence on why the restricted components exist — students comply with reasons, not rules.
- Three read aloud. The room applies the test: could a first-year act on this without asking a single clarifying question?
Portfolio checkpoint + exit ticket
Four artifacts in hand by now
- Verification log (S4)
- Redesign sheet (D3)
- Tutor prompt + break-test note
- Policy paragraph
Exit ticket — one word: your redesigned assessment's weakest remaining point.
Harvested; feeds tonight's gap analysis.
These are the raw material for S6–S8: personal practice becomes institutional text.
Contrarian close
SAY
"Contrarian claim of the afternoon: the most ethical AI policy your institution can adopt is not a
policy at all — it is a better assignment. Every hour a committee spends wordsmithing
prohibition clauses buys less integrity than the forty minutes you just spent redesigning one
assessment. Tonight we write policy anyway — because institutions need it, because clarity is a
justice issue, and because you now write it as builders, not police. That order was the point of today."
Appendix · facilitator only
Fallbacks & notes
- [ insert screenshots: P09 vulnerability audit sample · P10 tutor holding under pressure ]
- Honesty guard on the Tutor arm: safety yes, superpowers no — don't oversell guardrails.
- Slide 5b "What if there were no grades?": 5 min only — provocation, not resolution. If the room wants to go deeper, redirect to the MIT report §3.1.6 as optional reading.
- Handouts: P09–P11 + redesign sheet (D3), printed per participant.