Learning note
Learning note
Defense in depth uses multiple safeguards so that one failure need not lead to harm. BlueDot's AGI Strategy framework distinguishes preventing dangerous training, constraining dangerous capabilities, and withstanding dangerous actions. HASI applies that broad structure to decisions students and institutions can analyze. A campus pilot cannot by itself control frontier-model training; it can set procurement and deployment conditions, constrain its agent's access, and prepare recovery.
At the development or procurement stage, require relevant evaluations and clearly bounded uses. At the capability and access stage, restrict tools and permissions, test intended behavior, and separate drafting from consequential action. At the consequence stage, preserve records, provide human correction, and prepare a manual fallback. These are example interventions to evaluate, not a guarantee of safety.
Three versions of the same check may share one failure. If the drafting model also reviews itself using the same missing information, both can agree on a wrong answer. Independent evidence, different authority, and a tested fallback can be more useful than adding another warning message. Do not multiply failure probabilities unless independence and the individual probabilities are justified.
“Human oversight” needs an owner with time, information, authority, and a usable intervention. A reviewer who sees hundreds of outputs without source documents may become a rubber stamp. A stop button is useful only if someone notices the problem, can activate it, and can prevent or reverse consequences. Some harms are irreversible.
Governance turns intentions into responsibilities: who approves, who monitors, who can stop, who can appeal, and who pays for errors? Every safeguard also has costs. More checking can delay support; collecting more records can increase privacy exposure. Compare these tradeoffs explicitly.
Source lab · 30 minutes
Read the risk-management section of the International AI Safety Report summary, or compare W4's fictional safeguard cards. Explain how a technical control and an institutional procedure depend on each other. Optional: BlueDot defense framework.
Retrieval
Name one prevention measure, one permission limit, and one recovery measure. Explain a shared failure that could defeat all three.
W4 · A defense plan with a real owner
Fictional planning constraint
Puente has ten staff-hours for pilot setup this week. These are teaching estimates, not real prices or validated implementation times. Choose a package at or below ten hours and justify excluded options.
A · 2 hours: Keep the agent draft-only; disable live database changes and sending. Owner: technical lead. Limitation: drafts can still contain wrong information.
B · 3 hours: Create a source-check procedure for deadlines and eligibility. Owner: staff lead. Limitation: ongoing review time is additional.
C · 2 hours: Run representative fictional test cases with severe errors reported separately. Owner: evaluation lead. Limitation: tests may miss new conditions.
D · 3 hours: Test a manual fallback and correction/appeal route. Owner: operations lead. Limitation: a correction cannot always undo a missed deadline.
E · 1 hour: Add a second automated reviewer using the same model. Owner: vendor. Limitation: correlated failures and missing source data.
F · 4 hours: Interview a small group about accessibility and notification clarity. Owner: engagement lead. Limitation: findings may not represent all participants.
Worksheet · 60 minutes
- 15 min: Select your package. Show the sum and explain each measure's place in the risk map.
- 15 min: For each chosen defense, state owner, trigger, action, failure mode, and evidence that it works. Separate setup hours from ongoing staffing.
- 10 min: Identify the shared dependency most likely to defeat the package.
- 10 min: Write a stopping rule with a measurable trigger, an authorized person, a fallback, and a restart condition.
- 10 min: Respond to a new constraint: the reviewer is absent and volume doubles. Continue, narrow, or suspend? Explain.
Stress-test prompt
You cannot buy every safeguard. Which missing measure creates the greatest remaining risk? Could you narrow the pilot so fewer safeguards are needed? A defensible decision may be to delay or to retain the manual process.
Cumulative brief · 45 minutes
Add the selected defenses, rejected alternative, operational owner, and remaining risk to your final brief. Compare consequences for a student with reliable internet and one who depends on an in-person visit, without assuming language or income from ethnicity.
Exit check
What makes a safeguard independent? Why is “we have a human in the loop” insufficient by itself?
