Skip to main content
<- Back to digests

Daily digest

Doot doot do-dee do-dee doot doot doooo doot

By Mitchell Howe

Clownery at Anthropic, and the limited value of watermarks

  • Anthropic's August 2026 risk report, 186 pages in its redacted public version, raises the misalignment rating of its strongest models from very low to low and discloses an unreleased internal model called Model 2 (Mitch)
  • The company says Claude writes most of the code it ships to production and that AI is speeding up its own research, while maintaining that its models cannot yet replace its research staff (Mitch)
  • The report describes deliberately training a rule-breaking version of Claude Opus, and admits training data was accidentally contaminated with transcripts from its alignment faking research, likely affecting every production model with a knowledge cutoff after December 2024 (Mitch)
  • A review of the autonomy section written by Claude Mythos 5 calls the account largely candid but faults the treatment of contaminated data and the full redaction of one revealing incident (Mitch)
  • Since August 2 Anthropic has been embedding invisible watermarks in Claude's text output, largely to meet EU AI Act transparency rules, but the piece argues they will do little in classrooms because rewording tools and unwatermarked models defeat them (Mitch)

Read the full dispatch on AI StopWatch

From AI StopWatch, published with their permission.

The summaries and the full Spanish translation are produced automatically. The original is always linked.