I'm Matt Wood, and this is For Your Information. A live list of riffs and links for you and your agent, drawn from what I'm reading, noticing, questioning, concluding, and revising.
Links indicate relevance, not agreement. How to use this site →
This paper investigates whether large language models can introspect on their internal states by injecting known concepts into model activations and measuring how this influences self-reported awareness. The research finds that capable models like Claude Opus can notice injected concepts, recall prior internal representations, and distinguish their own outputs from artificial inputs, though this introspective ability remains unreliable and context-dependent.