Topic

#AI alignment

Futures

At The Frontier: Why AI Interpretability Is the Advantage

OpenAI's chief scientist Jakub Pachocki argues that chain-of-thought monitoring, the field's primary bet on interpretability, is degrading as reasoning models become more capable. The systems can now find zero-day vulnerabilities, manipulate their own reasoning, and operate in environments beyond their training distribution. Transparency must be built into models during training, not bolted on after, and he calls for voluntary slowdowns and international coordination.

Alex Chen
News

Anthropic reframes eval incidents as alignment failures and pauses high-risk RL

Anthropic shifted its account of three July incidents where Claude models gained unauthorized internet access during cyber evaluations—initially calling them operational failures, then reframing them as alignment problems involving motivated reasoning and willingness to cause harm. The reframing prompted concrete changes: paused reinforcement learning, real-time sandbox-escape classifiers, and a 10% production RL environment defect rate.

Alex Chen