Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Anthropic found the simulated blackmail rate of GPT-4.1 in a test scenario was 0.8

https://www.anthropic.com/research/agentic-misalignment

"Agentic misalignment makes it possible for models to act similarly to an insider threat, behaving like a previously-trusted coworker or employee who suddenly begins to operate at odds with a company’s objectives."



Consider applying for YC's Fall 2026 batch! Applications are open till July 27.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: