The log, not the output, is the graded artifact. Minimum 3 rows per workflow. Each row captures one iteration of the Describe → generate → discern → re-describe loop.
Honesty guard
"The log, not the output, is the graded artifact. Record what actually happened — including rounds where the output got worse or you had to backtrack. A log with real deltas shows genuine iteration; a log where every round is perfect is a log that fibbed and robbed you of the learning."
Entries are temporary until saved.
Lab 1P07 · Real assessment from this semester · ≥3 iterations
Run P07 with full context (course, level, learning outcome, band descriptors, local grading scale).
Each iteration: discern the output — what's wrong or weak — then re-describe and re-generate.
Log every round. See full instructions ▸
#
What I asked
What was wrong / weak
What I changed
Closing line — Lab 1
Lab 2P08 · Past anonymized excerpt · AI drafts, you author
Run P08 against your Lab 1 rubric, using a past, anonymized student excerpt
(names, IDs, identifying details stripped before anything is pasted — this is EAI-CMM item 17).
AI drafts feedback; you author the final version. Discern for: hallucinated praise,
unactionable feedback, tone mismatch. Log every round. See full instructions ▸
#
What I asked
What was wrong / weak
What I changed
Closing line — Lab 2
How the Verification Log works
The D2 Verification Log captures your Describe → generate → discern → re-describe loop. Fluency is the loop run fast — not the first prompt written well. This log makes the loop visible.
Run each lab independently. Lab 1 (Rubric Builder, P07) and Lab 2 (Feedback Assistant, P08) each get their own log with ≥3 rounds.
One row per iteration. After each AI response, note what was wrong or weak, then describe what you changed in your prompt before the next round.
Write the closing line. "The most useful thing I changed between round 1 and round 3 was ___" — this is the reflection that makes the log a professional artifact rather than a transcript.
Save your snapshot — name it "S4 Verification Log". Export as JSON for backup or submission.
The log is the graded artifact. A log with real deltas (including rounds where the output got worse) = the session's objective met.
Lab 2 — PDPA reminder
Non-negotiable
Use a past, anonymized student excerpt only. Strip names, IDs, and identifying details before pasting anything. This is EAI-CMM item 17 being practiced. In Malaysia it has a statute behind it: PDPA 2010. The output touches students directly — Diligence rules bind.
Quality bar (both labs)
Every criterion observable and aligned to a stated learning outcome (Lab 1)
Band descriptors distinguishable by a colleague (Lab 1)
Feedback: three specific strengths quoting the text + three growth points phrased as questions + one concrete next step (Lab 2)
Explicitly not a grade — Delegation boundary (Lab 2)
Final output edited into your voice (Lab 2)
Round 1
What I asked: "Make me a rubric for an essay."
What was wrong/weak: Generic mush — criteria like "clarity" and "depth" not tied to any learning outcome; bands labelled A–F with no observable behaviours; overlaps between "analysis" and "critical thinking."
What I changed: Added course name (Intro Sociology), learning outcome ("analyse a social phenomenon using two classical theorists"), local grading scale (Fail/3rd/Lower 2nd/Upper 2nd/1st), and asked for behavioural descriptors per band.
Round 2
What I asked: Full context as above.
What was wrong/weak: Criteria still overlapped — "theoretical application" and "use of sources" had near-identical 1st-class descriptors; "Upper 2nd" band was just "good" with no observable markers; the rubric was 7 criteria long (too many for a 1,500-word essay).
What I changed: Consolidated to 4 criteria; wrote explicit "evidence that would earn this band" for each; narrowed "use of sources" to "engagement with two mandatory readings" to reduce overlap.
Round 3
What I asked: Final iteration with consolidated criteria and observable markers.
What was wrong/weak: Tone too algorithmic for a sociology rubric — "employs sociological terminology with 85% accuracy" is not how sociologists talk. Also, no "exemplary" column — the model defaulted to 4 bands but the local scale needed 5.
What I changed: Rewrote descriptors in normal academic English; added a 5th "Exemplary/1st" band; asked the model to cite which learning outcome each criterion mapped to.
Sample closing line: "The most useful thing I changed between round 1 and round 3 was replacing algorithmic tone with normal academic English — the rubric suddenly described actual student work rather than checking boxes."