I'm Matt Wood, and this is For Your Information. A live list of riffs and links for you and your agent, drawn from what I'm reading, noticing, questioning, concluding, and revising.
Explores the fundamental relationship between data compression and predictive modeling, examining how compression algorithms work as predictors and the implications for AI and LLMs.
Ten Advances In MathematicsCompression-prediction duality has deep mathematical roots (Kolmogorov complexity, Solomonoff induction, MDL principle) — this connects to advances in mathematics where information theory intersects with learning theory
Ways to think about token pricingUnderstanding compression as prediction directly informs token pricing economics — if tokens represent compressed predictions, the cost structure of inference has deeper theoretical grounding than simple throughput metrics
Neuronpedia Jacobian Lens: Interactive J-Space Explorer for Qwen3.6-27BMechanistic interpretability work like Jacobian Lens analysis implicitly relies on the compression-prediction equivalence — understanding how LLMs compress training data into weights explains what features like those in J-Space actually represent
Related to
11–16× Faster LLM Inference with llama.cppFaster LLM inference via llama.cpp optimizations relates to compression-prediction: understanding that LLMs are fundamentally predictive compressors informs where computational bottlenecks lie in token generation
LLMs can't jumpBoth probe fundamental limitations of LLMs — 'LLMs can't jump' examines what models fail to generalize, while compression-as-prediction reveals what LLMs are actually doing when they model sequences, making them complementary perspectives on LLM capabilities
Challenges
Intelligence is not the main bottleneck'Intelligence is not the main bottleneck' argues against pure capability scaling, but compression-as-prediction suggests intelligence itself may be reducible to compression efficiency — directly challenging claims that intelligence is separable from prediction/compression
Supported by
Understanding Embeddings in Language ModelsEmbeddings fundamentally encode compression and prediction — the piece on 'compression is prediction' directly relates to how dense numerical representations capture semantic meaning through learned statistical patterns
Dynamic Programming: Unifying Principle Across AlgorithmsCompression-as-prediction is deeply connected to DP: Bellman's principle of optimality underlies both sequence prediction and data compression algorithms, sharing the same recursive substructure logic