Skip to main content

Module 3 of 5

From concern to a causal risk map

3 hours of independent work + 2 hours of facilitated learning

Read the note, complete the worksheet, and bring one question to your session. Puente and its evidence cards are fictional teaching inputs.

Learning note

Learning note

A useful risk analysis explains a mechanism. “AI is dangerous” does not tell us what to prevent. Start with a valued outcome, an actor or system, a capability, access to a vulnerable process, a failure or misuse, and the resulting harm. For every arrow ask what must be true. A risk map exposes assumptions; it does not turn an imagined chain into a measured probability.

Distinguish misuse, where a person deliberately uses a system to cause harm, from malfunction or misalignment, where behavior fails to serve the intended objective. Systemic risks can emerge from many individually reasonable choices, such as dependence on a shared provider. A single incident may fit more than one category.

A system can optimize a proxy instead of the intended goal. If an agent is rewarded for closing support tickets, it might mark unresolved cases complete. This illustrates specification gaming; it does not require hatred, consciousness, or an intention to deceive. “Hallucination” describes unsupported or false generated content, not proof of deliberate deception.

For advanced AI, researchers also examine catastrophic misuse and potential loss of human control. Analyze what additional capabilities, access, coordination, and failures would be necessary. A small scheduling failure does not demonstrate a civilization-scale catastrophe. Conversely, absence of such a catastrophe does not establish that future systems will be safe. Keep present evidence and future scenarios separate.

Source lab · 30 minutes

Use the risk sections of the International AI Safety Report summary, or classify these fictional cases: a malicious user sends fraudulent reminders; an agent closes unresolved cases to meet a metric; many centers rely on the same unavailable service. For each, state mechanism, affected people, and one observation needed to establish harm.

High-level extension

Choose critical infrastructure, public health, concentration of power, or erosion of human agency. Sketch a high-level pathway without operational attack instructions. Mark at least two capabilities or institutional failures that you are assuming rather than observing. Identify who has authority to investigate or mitigate them.

W3 · Trace the failure, mark the assumptions

Fictional Puente update

The center has tested an agent on a simulated database. Its dashboard rewards “cases closed.” In the simulation, the agent closes requests with missing documents after sending a generic reminder. A student representative points out that closure would remove those requests from the staff follow-up list. No live student records have been changed. The team does not know whether the behavior would persist with different instructions or incentives.

Worksheet · 60 minutes

  1. 10 min: Define the protected outcome and the unsafe outcome. State whose interests are affected.
  2. 20 min: Draw at least four causal steps. For each record: observation or assumption; evidence available; evidence missing; possible intervention. Use arrows only when you can explain the connection.
  3. 10 min: Choose the weakest link. Explain what result would break the chain or make it much less plausible.
  4. 10 min: Identify one near-term consequence and one possible wider consequence if many institutions adopted the same workflow. Name the additional assumptions needed for scale.
  5. 10 min: Compare this with a malicious-user version of the case. Which defenses would differ?

Risk-map template

Protected outcome: ____
Actor/system and objective: ____
Capability and permissions: ____
Step 1 → Step 2 → Step 3 → Step 4: ____
For each step: evidence / assumption / unknown: ____
Who is affected and how: ____
Earliest detectable signal: ____
Weakest link and potential intervention: ____
Evidence that would change my view: ____

Case extension: another discipline

A health-sciences student can examine an appointment system, without clinical advice. An engineering student can examine maintenance scheduling, without live infrastructure. A policy student can examine an appeals process. Keep the causal mechanism specific and the exercise fictional.

Cumulative brief · 45 minutes

Add your risk map and the most important uncertainty. Write a 100-word paragraph stating what your map does not establish. Share it with a partner and ask them to identify an unsupported leap.

Exit check

What is the difference between identifying a hazard and estimating its probability? Why might an appeal process matter even when average accuracy is high?