Learning note
Learning note
Training changes a model using data and computation. Inference is using the trained model to produce outputs. More useful data, better algorithms, more compute, and improved use of compute at inference can improve performance, but a trend is not a guarantee of a specific future capability. The tools, permissions, time budget, and human support around a model also shape what the overall system can accomplish.
A benchmark measures performance on selected tasks under specified conditions. It can reveal progress without proving readiness for a particular workplace. Ask what tasks were selected, how success was scored, whether examples were familiar, how often the system was retried, and what human assistance it received. A model's confident explanation is not a measurement of reliability.
An AI research feedback loop is a hypothesis: if AI accelerates research that improves AI, progress could accelerate further. Whether that becomes a rapid intelligence explosion depends on bottlenecks such as experiments, reliable evaluation, computing resources, and integration. A competing view emphasizes slow adoption, institutional change, and uneven usefulness. Analyze the mechanisms and evidence for each; do not grade forecasts by how dramatic they sound.
Reading a time horizon: METR's task horizon is based on how long a human expert would take on tasks at a stated AI success probability. It is not the time the AI runs, or proof that an entire occupation is automated. Task selection and real-world conditions limit generalization. METR methodology and limitations
Source lab · 30 minutes
Use METR's limitations note, or work offline with W2's fictional data. Identify the metric, denominator, test environment, missing comparison, and strongest justified conclusion. Optional contrasting perspective: Narayanan and Kapoor, AI as Normal Technology. Its emphasis on adoption and institutions is an argument to evaluate, not a settled refutation of fast progress.
Retrieval
Explain the difference between model capability and safe deployment to someone studying nursing or civil engineering. Name a factor that could change one without changing the other.
W2 · Evaluate before extrapolating
Fictional evidence cards
A. Vendor demonstration: 18 of 20 support tasks completed correctly; tasks selected by vendor; one attempt each; no live database access; staff supplied clean input. The vendor says, “Ready for autonomous student support.”
B. Independent classroom trial: 12 of 20 tasks completed correctly; tasks include ambiguous dates and incomplete requests; same model version and one attempt each. Different task selection means A and B are not a controlled comparison.
C. Missing evidence: no manual-service baseline, no accessibility evaluation, no severity weighting, no test of high-volume demand, and no measurement of students' downstream outcomes.
Worksheet · 60 minutes
- 10 min: Calculate both success rates and failure rates. Explain why combining them into a single 75% success rate would hide different test conditions, even though the arithmetic is possible.
- 15 min: Write a six-field evaluation card: claim, evidence, test conditions, missing information, confidence, next test.
- 15 min: Design a fair comparison: same representative tasks, explicit success criteria, human baseline, consistent permissions, and reporting of severe errors separately. Specify who judges results without knowing which process produced them where feasible.
- 10 min: Produce two conditional forecasts: one where capability improves quickly, one where integration remains slow. Name an observation that would make each more plausible.
- 10 min: Rewrite the vendor claim in a defensible 50-word form.
Extend your analysis
Would a 99% success rate be acceptable if the remaining errors remove students from essential support? What changes if staff review every output? What is the cost of the review? Describe the stakes before choosing a threshold.
Cumulative brief · 45 minutes
Add one evidence table and one proposed evaluation to your brief. Distinguish the fictional practice data from any real source you cite. For a non-Puente project, use an equally small, transparent fictional test set or verified public data; do not claim that you conducted a study.
Exit check
A system has a two-hour task horizon at 50% success. Does that mean it safely works unattended for two hours? Explain why not. What evidence would weaken your favored forecast?
