Learning note
Learning note
AI safety concerns preventing harm from AI systems and keeping their behavior sufficiently reliable and controllable for their intended context. A useful system can still be unsafe in a different setting. A text assistant that suggests an answer and an agent that sends money may use similar models but have very different consequences when wrong.
General-purpose AI can perform many kinds of tasks. AGI, artificial general intelligence, has no single agreed threshold; in this course it means a hypothetical broadly capable system that could perform a wide range of cognitive work at roughly human level or beyond. Superintelligence means substantially exceeding human performance across many relevant domains. Neither term, by itself, establishes consciousness, a deployment date, or a particular political conclusion.
Safety, security, and governance overlap. Safety asks how harm is prevented; security considers unauthorized access and malicious interference; governance concerns who makes decisions, under what rules, and with what accountability. Alignment asks whether behavior serves the intended goals and constraints. Agreement on goals is itself a human and institutional task.
A good future is more than a list of capabilities. Ask who benefits, who bears risks, who can refuse, and who can challenge a decision. A faster service can still fail people with complex needs. A safer service can also be unaffordable or inaccessible. Make these tensions visible before choosing a tool.
Source lab · 30 minutes
Read the opening and risk-category sections of the International AI Safety Report executive summary. Identify one observed harm, one uncertain future risk, and one value judgment that a decision-maker must supply. Alternatively, classify these HASI statements: “In our fictional test, two notices had wrong dates”; “A future agent might spread an error across institutions”; “People should have a meaningful appeal.” Explain why a scenario cannot establish prevalence.
Retrieval · close the page
Explain AI safety to a classmate in 60 seconds without using AGI, alignment, or governance. Then reopen the note and repair any missing distinction.
W1 · The Puente decision
Fictional case
Puente is a community-serving college support center. It is considering an AI agent that drafts scholarship reminders, recommends appointments, and updates a student support database. The director wants shorter queues. Staff want less repetitive work. Students want accurate information and a way to correct mistakes. The vendor wants a successful launch.
In 40 fictional reminder drafts, staff found two incorrect deadlines and one unsupported eligibility claim. These are draft-level findings, not measured student harm. No autonomous sending was tested. The vendor offers a limited pilot; the center can also retain its existing manual process. A scholarship deadline is approaching. There is no known attack or evidence of intentional deception.
Worksheet · 60 minutes
- 10 min: Describe a good outcome in 100 words. Include service quality, access, and a right to challenge errors.
- 15 min: Map four actors. For each, write what they want, what they know, what authority they have, and what could change their position.
- 10 min: Separate three observations from three assumptions. What is unknown about error severity, affected users, and current staff performance?
- 15 min: Compare a manual process, draft-only pilot, and autonomous sending. For each, state a benefit, risk, and information needed.
- 10 min: Write a 150-word provisional decision with a review date and a condition for changing your mind. Do not invent measured outcomes.
In-session role cards
Director: You control launch scope and staffing. Explain the delay cost and which evidence justifies deployment.
Student representative: You can request changes but cannot deploy. Ask how students notice errors and obtain corrections.
Staff lead: You operate the service. You can review ten drafts per hour. Identify what happens when demand exceeds review capacity.
Vendor lead: You can change permissions and provide test information. Explain what your evidence does and does not show.
Each role proposes a decision, hears objections, and revises one condition. Your role need not match your personal view.
Cumulative brief · 45 minutes
Create a document with headings: problem, evidence, options, recommendation, uncertainty, next action. Fill the first three using Puente or an equally bounded fictional case of your own. Check that you describe people as decision-makers, not just recipients of technology.
Exit check
Can a system be helpful but misaligned? Give an example. What would count as evidence against your decision?
