The Latest
Researchers Slipped A Single Word, ‘Bread’, Directly Into An AI Model’s Own Neural Activations, With Nothing In The Prompt To Hint At It, And Claude Opus Still Caught The Change About One Time In Five, A Signal That Misfired Zero Times Across A Hundred Separate Trial – Space Daily
Anthropic researchers successfully inserted the concept ‘bread’ into Claude Opus’s neural network, which detected the change about 20% of the time with no false positives.