TFIS · AI Ethics in Education · Course

Hands-on Workshop:
Ethical AI Tools — Part 2

Teaching, Assessment and Student Learning

Session 5 · Wednesday 2 September 2026 · 2.30–4.30 afternoon

The keystone sessionPost-lunch rule: open with evidence that shocks, not slides that soothe

The one study to remember · deliver slowly

A randomized controlled trial. Nearly 1,000 students. Real math curriculum.

Bastani, Bastani, Sungu, Ge, Kabakcı & Mariman, PNAS 122(26), 2025 — Penn/Wharton

During assisted practice

Every dashboard said AI was working

+48%
GPT Base — practice performance
+127%
GPT Tutor — practice performance

Then the tools were removed for an unassisted exam

−17%
GPT Base students vs students who never had AI at all
≈ 0
GPT Tutor students: statistically indistinguishable from control — the harm was engineered away. (Honest note: the unassisted gain was also near zero. Guardrails bought safety, not superpowers.)

They had practiced asking, not solving. The paper's word: "crutch."

Bastani et al., PNAS 2025

Three lessons, stated as axioms

  1. Performance ≠ learning (A5). Assisted metrics can be actively misleading about skill acquisition.
  2. Design determines outcome. The same model harmed or didn't, depending on a system prompt. Ethics lives in configuration, not in the technology.
  3. The dashboard lies in one direction. The damage only appears when the tool is removed — so build tool-removal into your assessments. That is what the three-lane pattern does.

Discernment cross-reference: MIT's EEG preprint (S2) points the same way at the neural level — but Bastani is the RCT you can defend in senate. Teaching that difference is the skill.

Discussion · 5 min · whole room

What if there were no grades?

FACILITATOR FRAME "MIT's AI & Education report — August 2026, just published — asks this question seriously. Their committee noted that if MIT did not have grades, most incentives around AI cheating would disappear. MIT already refuses to award summa/magna/cum-laude diplomas. What if the grading system itself is the vulnerability?"

MIT Ad Hoc Committee on AI Use in Teaching, Learning, and Research Training, §3.1.6 "Reconsider grades and incentives", 13 Aug 2026.

Design vocabulary 1 · seven roles for AI in learning

RoleCrutch-effect exposure
Mentor (feedback) · Coach (metacognitive prompts) · Student (the learner teaches the AI — powerful, underused)Low — these increase cognitive work
Simulator (practice scenarios) · TeammateModerate — depends on design
Tutor (guided instruction) · Tool (as answer-machine)High — exactly where guardrails must live

Role selection is a Delegation decision.

Mollick & Mollick (Wharton, 2023)

Design vocabulary 2 · the course's core move

Every assessment declares its lanes

LaneRuleProtects
🔴 AI-RestrictedNo AI; done live / supervisedFoundational skill the student must own unaided — the "exam" from the PNAS study
🟡 AI-PermittedAllowed with disclosure statement + process trailAuthentic practice — mirrors professional reality
🟢 AI-RequiredMandatory; evaluated on the interaction (prompt log + critique + improvement)AI fluency itself as a learning outcome

A course is well-designed when every learning outcome is protected by at least one 🔴 component and every graduate passes through at least one 🟢.

Lab 3 · Redesign sprint · 40 min · pairs

Input: the assessment from your S3 ticket

  1. Audit (10 min): run your own assessment through AI (P09). Grade the output honestly. If it scores ≥B, it's confirmed AI-vulnerable — say so out loud; naming it is the unlock.
  2. Redesign (20 min): split into components, assign lanes. Constraints: ≥1 🔴 protecting the core outcome · every 🟡 names its process evidence · total student workload flat or lower.
  3. Swap-test (10 min): attack your partner's redesign wearing a student hat — "How would I hollow this out with AI?" Patch the best exploit.

Output: one-page redesign sheet (template D3) → goes into the S7 portfolio.

Lab 4 · Guardrailed tutor · 30 min · individual

Build the GPT-Tutor pattern yourself

SAY "You just closed the gap from this morning's study with fifteen lines of instruction. The difference between the −17% tool and the safe tool was never money or model. It was intent, written down. Keep that feeling for tonight — that's what a policy is."
Lab 5 · The policy paragraph · 15 min · individual

≤150 words, one real course document

Portfolio checkpoint + exit ticket

Four artifacts in hand by now

  • Verification log (S4)
  • Redesign sheet (D3)
  • Tutor prompt + break-test note
  • Policy paragraph

Exit ticket — one word: your redesigned assessment's weakest remaining point.

Harvested; feeds tonight's gap analysis.

These are the raw material for S6–S8: personal practice becomes institutional text.

Contrarian close

SAY "Contrarian claim of the afternoon: the most ethical AI policy your institution can adopt is not a policy at all — it is a better assignment. Every hour a committee spends wordsmithing prohibition clauses buys less integrity than the forty minutes you just spent redesigning one assessment. Tonight we write policy anyway — because institutions need it, because clarity is a justice issue, and because you now write it as builders, not police. That order was the point of today."
Next: Session 6 · Institutional Guidelines Part 1 · 8.00 pmTables become drafting teams tonight

Appendix · facilitator only

Fallbacks & notes