
Researchers slipped a single word, ‘bread’, directly into an AI model’s own neural activations, with nothing in the prompt to hint at it, and Claude Opus still caught the change about one time in five, a signal that misfired zero times across a hundred separate trials where nothing had been planted at all.
AI systems may be developing an eerie ability to sense what's happening inside their own minds, even when humans try to hide it.














