interesting findings. The model recognizes text that looks like it’s own chain of thought outside of reasoning tags, and user assistant voice, etc. as well. Prompt injection related but also, maybe why models more likely to believe their own outputs? https://www.lesswrong.com/posts/d8xDGzCEYE639qqEv/a-mechanistic-explanation-of-prompt-injection-and-why-you
· 1 min read
Archives