Links indicate relevance, not agreement. How to use this site →
Research introducing Declarative Attention, a protocol that enables language models to declare which parts of context they need within their chain-of-thought, reducing KV cache reads and inference costs without requiring model retraining or architectural changes.