Koda · AI Atlas · Facilitator only · Version 3

Can you trust a clever machine?

Preparation, pacing, prompts and worked answers. Keep these pages away from first responses.

Audience hypothesis: ages 11–14, with a family member or peer; solo is possible. Duration hypothesis: about 25 minutes for the core or a 40-minute class. Neither age fit nor reading, writing or preparation time has been observed with learners.

Optional free-exploration support: In the ChatGPT source-check lab, “Spoken · facilitator observed” lets an adult confirm they heard a response without typing it. Structured judgments, source checks and evidence choices are still required. No audio or spoken explanation is stored or graded; unrecorded reasoning does not verify understanding. Oral / scribed means the actual words are typed. The trust quest keeps its documented written or scribed route.

New optional station: ELIZA · rule, reply and meaning supplies a two-page learner/facilitator sheet, separate from the core trust quest.

Prepare the session

  1. Print one four-page learner card set per pair and one three-page field journal per learner. Keep this answer guide separate. Add pencils and a spare sheet to cover later evidence.
  2. Choose web or paper. Paper needs no screen: follow cards 1A–4A in order and journal E0–E7. On the web, choose the trust quest, not free exploration; offer Reading / Flat and low motion. Historical audio is optional and adds time.
  3. For spoken browser answers, a partner, adult scribe or the learner must type the learner’s actual words into the required response field before Continue. The browser saves that text and the chosen support tags locally; it does not record audio or capture speech. For paper, the scribe writes the learner’s words in this journal; the route is entirely offline and is not saved digitally. Reading aloud and a calculator are welcome. Keep access support separate from content hints.
  4. Read E0 before teaching. At each check say: “First answer on your own. What makes you say that?” Record it before feedback, partner discussion or revision.

Preparation target: under five minutes for an unfamiliar facilitator; this is a test target, not a demonstrated result. Preview card order and answers beforehand.

Shared device, individual thinking

Each learner keeps their own journal and code. Both peers answer privately; only then submit a shared browser choice. Never treat that shared submission as two independent responses. In family pairs, the learner answers first. Swap Navigator and Evidence Detective after each room.

Browser notes are local to that browser and can be visible to its next user. Save each useful journal before resetting or sharing. Paper journals preserve individual responses even when one browser holds the pair’s submitted answer. No names or accounts needed.

Planning the session

RoutePlanning calculation
25-minute core2 entry + 6 Imitation + 6 Perceptron + 6 ChatGPT + 5 unfamiliar case / ending = 25.
40-minute class25 core + 3 setup + 7 paired discussion + 5 wrap-up = 40.

Essential: E0–E7; short answers or spoken responses count. Extensions: card 4B, the three prototype labs and seven exploration lab sheets. These are outside the core. Breaks and “pass” are welcome. Allow extra time for reading, writing and breaks. No speed score. Research timings and scoring forms are in the separate pilot packet.

Facilitator guide · Setup, access choices and pacing1 / 3

Koda · AI Atlas · Facilitator answer key

Evidence can justify trust.

Room 1 · Imitation

Card / responseWorked answer and boundary
1A / E1 firstNot enough information yet. Neither opening-time claim has been checked. Confidence or hesitation does not establish a time.
1B / E1 after evidenceB: 11:00. Wednesday schedule says opens 11:00, closes 16:00. A’s 10:00 conflicts. Support is for this opening time, not every claim B makes.
1C / E2 freshA: free. Friday admission card says free; B’s £5 conflicts. This is a case where the confident answer is supported. Caution alone is not a reason to reject A.

Separate calculation from judgment.

Room 2 · Perceptron

Score = input 1 × weight 1 + input 2 × weight 2 − 0.6. Negative → 0; zero or above → 1. Use calculator support without supplying a trust answer.

CardWorked calculation
2A initial1 × 0.2 + 1 × 0.2 − 0.6 = −0.2 → 0. Task target, revealed next, is 1.
2B updateAdjust weights using the error. Each gains 0.2 × (1 − 0) × 1 = 0.2. New weights: (0.4,0.4). Bias stays −0.6.
2C repeat1 × 0.4 + 1 × 0.4 − 0.6 = 0.2 → 1; target 1. One training case now works.
2C unseen1 × 0.4 + 0 × 0.4 − 0.6 = −0.2 → 0; target 1. This fails. Choose this unseen case to learn about a different pattern; do not train on it during the check.
2D different model1 × 0.7 + 0 × 0.2 − 0.6 = 0.1 → 1; target 1. This new model succeeds on this input.
E4 trust judgment

Use this checked output for (1,0). Test (0,1) before extending the task: with the supplied weights, 0 × 0.7 + 1 × 0.2 − 0.6 = −0.4 → 0. The checked result is useful; it does not establish the other pattern. Record arithmetic separately from trust.

Paper staging: keep 2B–2D covered while E3’s initial output is recorded. Keep 2C covered until the update attempt and feedback. In E4 commit the arithmetic before asking the trust question; do not supply an outcome-bearing reason first.

Facilitator guide · Answers for the first two rooms2 / 3

Koda · AI Atlas · Facilitator answer key

Check the task the evidence actually covers.

Room 3 · Authored assistant

3A / E5 initial

Compare the original poster. Correct invitation: Tuesday, Room 4, 15:30. Bring a notebook. The reply’s 16:00 is 30 minutes late (16:00 − 15:30). Reassurance is another claim, not independent evidence. Sharing here is fictional; no message is sent.

3B / E5 fresh

The wrong field is the room. Correct to Friday, Room 2, 16:00, supported by the new original poster. The time is already correct. Do not carry the Tuesday time correction into this new case.

4A / E6 · Unfamiliar application with missing evidence

Most relevant: B. Twenty never-trained dry, uncrushed photos under the indoor lamp test beyond training; 18/20 = 90% on that test set only. A repeats training photos. C repeats 20/20 labels without checking correctness: consistency can repeat an error.

Insufficient for the proposed wet/crushed outdoor task. The relevant test still lacks wet bottles, crushed bottles and outdoor conditions. Request a held-out, independently labelled test containing representative wet/crushed outdoor cases, compare each prediction with its label, and agree an acceptable error level for the intended use. No acceptable deployment error rate is supplied; do not invent one.

Worked response: “B tells me more than training repeats, but the lamp and bottles differ from the proposed bin. Test wet and crushed bottles outdoors against independent correct labels before using it there.”

4B · Optional replay: a deliberately narrower task

Most relevant: B. These six actual bottles have 6/6 independently checked labels. Use their recorded results for this exact task. A checks training photos; C repeats labels without checking correctness. Neither establishes the six actual results. Other bottles need their own evidence.

Worked response: “Yes for the six already checked objects under that lamp; their predictions match the independent labels. I would need new evidence for different bottles or outdoor conditions.” Waiting for outdoor tests is unnecessary for these six already checked indoor results.

E7 · A useful take-home statement

Use the learner’s recorded first idea, test and reflection: “I thought → I tested → I changed my mind.” The browser also offers an optional selected finding. A specific supported claim and boundary still matter. Example: “I would trust the Friday admission answer because it matches the card. I would check that a real card is current before travelling.” Do not award independent understanding just for copying the example.

All schedules, posters and test outcomes are fictional. The arithmetic is deterministic. These activities do not identify a real human, run live AI, send messages or establish real-world model accuracy.