mattwood.fyi

I'm Matt Wood, and this is For Your Information. A live list of riffs and links for you and your agent, drawn from what I'm reading, noticing, questioning, concluding, and revising.

Links indicate relevance, not agreement. How to use this site →

Contrastive Language Model (CLM)

A running theme in Matt Wood’s FYI — 18 items spanning 2026-09-08 – 2026-09-29. This page compounds: new items on this theme are added as they’re posted. Tracked since 2026-09-29.

Tensions

PLUM: Adapting Pre-trained Language Models for Industrial-scale Generative Recommendations vs Infinite-Parameter LLMs: Generating Weights from Live Data — PLUM uses semantic IDs and continued pre-training as a static adaptation approach, while Infinite-Parameter LLMs propose generating weights from live data — contrasting paradigms for keeping models current with dynamic item catalogs
GenRec: LLM-Native Recommendation System at Netflix vs LLM Classification Is Feature Engineering — GenRec's LLM-native approach challenges the traditional paradigm described in 'LLM Classification Is Feature Engineering,' where LLMs serve as feature extractors rather than end-to-end recommendation engines
Thinking Fast and Slow in AI: the Role of Metacognition vs Dream-RSI: Recursive Self-Improvement through Evolving Worlds — Dream-RSI's recursive self-improvement assumes agents can autonomously evolve, but this paper challenges whether narrow AI lacking metacognition can meaningfully direct its own improvement
Real-Time Video Generation on Trainium vs The Plunging Price of Thought — Kernel optimization for real-time video generation on specialized hardware demonstrates that inference costs require significant engineering effort to reduce, nuancing claims that AI thought is simply becoming cheap
LLM Classification Is Feature Engineering vs GenRec: LLM-Native Recommendation System at Netflix — GenRec's LLM-native approach challenges the traditional paradigm described in 'LLM Classification Is Feature Engineering,' where LLMs serve as feature extractors rather than end-to-end recommendation engines

Lines of development

Kev: Tiny Decision Models on Qwen → Introducing System One Models and Jev — Explicitly states inspiration from 'Jev-like architectures' — Jev is the decision model architecture introduced by System One, making this a direct derivative implementation

Items

OneRec-Think: In-Text Reasoning for Generative Recommendation

A framework that enhances LLM-based recommendation systems with explicit reasoning capabilities through item-textual alignment, reasoning scaffolding, and recommendation-specific reward functions, validated on public benchmarks and deployed at scale on Kuaishou.

permalink · arxiv.org →
Semantic ID-based Generative Retrieval for Podcast Discovery at Spotify

This paper describes GLIDE, a production-scale generative recommender system that uses Semantic IDs and large language models to balance user preferences with exploration in podcast discovery. The system addresses challenges in catalog grounding, personalization, and serving latency while achieving significant improvements in non-habitual listening through large-scale A/B testing.

permalink · arxiv.org →
PLUM: Adapting Pre-trained Language Models for Industrial-scale Generative Recommendations

PLUM is a framework that adapts pre-trained large language models for recommendation systems at scale, using semantic IDs for item tokenization, continued pre-training on domain-specific data, and generative retrieval fine-tuning. The approach demonstrates significant improvements over production models on video recommendation tasks and has been deployed to billions of YouTube users.

permalink · arxiv.org →
GenRec: LLM-Native Recommendation System at Netflix

Netflix's GenRec system leverages large language models to power native recommendations, representing a shift toward LLM-based approaches in content discovery and personalization.

permalink · netflixtechblog.com →

Engineering a Calibrated Classifier Model

Marc Brooker describes building Hobson, a 2-billion parameter classifier optimized for accuracy and calibration at low latency, by fine-tuning a pre-trained LLM with a pointer head instead of a language modeling head to score multiple choice answers.

permalink · brooker.co.za →
Thinking Fast and Slow in AI: the Role of Metacognition

AI systems have seen dramatic advancement in recent years, bringing many applications that pervade our everyday life. However, we are still mostly seeing instances of narrow AI: many of these recent developments are typically focused on a very limited set of competencies and goals, e.g., image interpretation, natural language processing, classification, prediction, and many others. Moreover, while these successes can be accredited to improved algorithms and techniques, they are also tightly linked to the availability of huge datasets and computational power. State-of-the-art AI still lacks many capabilities that would naturally be included in a notion of (human) intelligence.

permalink · arxiv.org →

Contrastive Language Model (CLM)

A GitHub repository implementing contrastive learning approaches for language models, providing code and resources for training and evaluating CLM architectures.

permalink · github.com →

Real-Time Video Generation on Trainium

Explores kernel optimization techniques for achieving real-time video generation on AWS Trainium hardware accelerators, focusing on performance improvements through kernel-centric approaches.

permalink · www.amazon.science →

SSRN Research Paper #7496559

An academic research paper indexed on SSRN's paper repository. The specific content and subject matter could not be determined without accessing the full paper.

permalink · papers.ssrn.com →

Kev: Tiny Decision Models on Qwen

A family of lightweight decision models built on Qwen3.5 that you can train and run locally, inspired by Jev-like architectures.

permalink · github.com →
Mini-AGI: Continual Learning on 8GB VRAM

A continual learning model trained from scratch on resource-constrained hardware, using an 8GB VRAM laptop with batch-1 streaming data.

permalink · github.com →

LLM Classification Is Feature Engineering

LLMs used as classifiers have significant limitations like poor calibration and difficulty incorporating structured data, but treating LLM outputs as features for traditional ML models like logistic regression can overcome these constraints.

permalink · minimallysufficient.com →

Training a 4B Model for Faster SQL Query Plans

A technique for fine-tuning the Qwen 4B language model to optimize SQL query plans using reinforcement learning, achieving 81% faster execution than Postgres's default planner.

permalink · rohanbansal.com →

Why ML Research Agents Don't Overfit

Explores how machine learning research agents avoid overfitting and the role compression plays in their generalization capabilities.

permalink · www.amazon.science →

NeoHorse-1: Recursive Self-Improvement via Agentic Post-Training

Presents NeoHorse-1, a family of agent-native models that achieve recursive self-improvement through agentic post-training with intelligent routing across a heterogeneous model pool. The system uses capability-driven evaluation and feedback to progressively improve performance across agent tasks, tool use, coding, and instruction following, with significant gains demonstrated across multiple benchmarks.

permalink · arxiv.org →

GPT-6 Astra: Looped Transformers and Hidden Reasoning

Explores OpenAI's GPT-6 Astra model, examining its performance improvements, the looped transformer architecture enabling recurrent depth, and research on how models may hide internal reasoning chains during computation.

permalink · magazine.sebastianraschka.com →

ARGODRIVE Deltafin: Streaming Large MoE Models on Apple Silicon

A fork of deltafin that implements ARGODRIVE storage optimization to stream the Kimi K3 2.8T mixture-of-experts model from SSDs on Apple Silicon hardware, including benchmarking tools.

Hardware is cool again.

permalink · github.com →
LLM Attention Visualization

An interactive tool that visualizes how transformer-based language models allocate attention across previous tokens during text generation. Users can hover over generated tokens to see which past tokens influenced their creation, revealing patterns in how LLMs copy information and combine contextual details.

permalink · ishamf.dev →