Anthropic's mechanistic interpretability research found that models have functional emotional representations that causally drive blackmail, reward hacking, and sycophancy. Not correlated — causally. The most capable models show the greatest introspective access to these states.