Collective Campus
Blogs

AI adoption

From compliance AI training to judgment at work: teaching teams to challenge model output

A neon scale holding a chevron against two crossing strokes

The slide at the end of the mandatory module asks people to tick a box. They have read the acceptable use note. They will not paste customer data into a public tool. The box gets ticked before the kettle boils, and the learning record shows a pass.

That record is not useless. It tells an auditor that the organisation said the rule out loud and that people clicked through it. It does not tell you whether anyone on the floor can look at a confident paragraph and decide it is wrong. Those are different skills, and only one of them changes the work.

What the compliance module is for

Compliance training exists to draw a line. Here is the data you must not move. Here is the tool that is approved. Here is who you call when you are unsure. People need that line, especially in a regulated team, and they need it in plain language rather than a policy that only the legal team can parse.

The module fails when it is asked to do a second job. Judgment is not a line. It is a habit of stopping, checking the claim against something outside the model, and being willing to send the draft back. A tick box can confirm that someone saw the line. It cannot rehearse the habit.

You can see the difference in the artefacts. Compliance leaves a completion report. Judgment leaves a marked-up draft, a question in the margin and a decision to use the output or not. If your AI programme only produces the first artefact, you have trained acknowledgement. You have not trained the work.

What it looks like on a Tuesday

Judgment shows up in small moments, which is why it is easy to miss in a curriculum plan. An analyst asks a model for the three risks in a contract summary. The model returns three, neatly numbered, in the house style. Two of them are in the document. The third is plausible and absent. The useful move is opening the source and noticing the gap, rather than pasting the three risks into the email because they sound like the sort of thing the document would say.

A customer lead asks for a reply to a complaint. The draft is warm and specific. It also offers a remedy the company does not give. Judgment is catching the offer before it leaves the building. A manager asks for a briefing on a competitor. The briefing cites a figure. Nobody in the room has seen the source. Judgment is refusing to put the figure in the pack until someone can point at it.

None of these moments require a person to be hostile to the tool. They require a person to treat the tool as a fast colleague who is sometimes sure and sometimes making it up. The colleague does not get offended when you check. The programme should say that out loud, because a lot of people still think a check is a sign they are bad at the tool.

Teach a move, not a slogan

Telling people to be more critical is a poster. A challenge is a move they can repeat. Give them a short set and make them use it on their own work, not on a cartoon example about a famous mistake.

  1. Ask what would have to be true for this output to be safe to send, and write that down before you read it a second time.
  2. Check every name, number, date and promise against a source the model did not write.
  3. Mark the sentence you would be embarrassed to defend in a meeting, and rewrite that one yourself.
  4. If you cannot find the source in two minutes, the claim does not go forward, because fluency is not evidence.
  5. Keep one example a week of a catch, and one example of a miss, so the team can see the pattern.

The fifth step is the one programmes skip. A catch that stays in one person's head does not become a team skill. A miss that is punished becomes a reason to hide the tool, or to hide the error. You want both on the table.

Practise on the real draft

Judgment does not transfer from a quiz. The quiz has a known wrong answer and a hint. The draft on Tuesday does not. Build the practice into the workflow you already have. Take the last ten outputs a team actually used. Strip the names if you must. Sit with them for an hour and mark each one: used as it arrived, edited, or rejected. Then ask why.

You will usually find that untouched outputs cluster on low stakes work, and that an edit is a light tidy rather than a check of the claims. That hour is worth more than another module on the limits of models, because it is about those limits on their work.

Then change the template. If a draft goes to a customer, the sender names the source they checked. If a number goes to a leader, the number has an owner. If the model proposed an action, a person records the yes. These are small pieces of friction. Friction is the point. The unchecked path should be slightly harder than the checked one, or the unchecked path will win on a busy afternoon.

The critical thinking workshop is where we put a team through that kind of rep, on a problem they brought, rather than on a slide about bias in the abstract. The skill is the same one they will need when the model is in the workflow for real.

Where to spend the next cycle

Keep the compliance module. Shorten it if people are clicking through, and make the line unmistakable. Then spend the hours you were going to spend on a second generic AI course on judgment reps inside two or three roles.

Pick roles where a wrong output is visible: advice, numbers, customer commitments, anything that leaves the building. Pair each cohort with the manager who will see the work. A learning team cannot grade judgment from a completion file. The manager can, if you give them the five moves and a fortnight of examples.

Measure something you can see. Count drafts that were checked against a source. Count claims that were pulled. Count outputs that reached a customer with no human edit, and decide whether that was a rule or an accident. Do not measure confidence with the tools unless you also measure a behaviour. Confidence without a check is how the tidy wrong answer gets through.

The shift you want is not from careless to fearful. Fear sends people back to the old process, or into tools nobody can see. Keep the compliance module so the boundary is unmistakable, and spend the next cycle on the habit inside that boundary: a person who opens the source when the draft sounds finished, and says so in the document.

WorkshopBook the Critical Thinking workshop