Links indicate relevance, not agreement. How to use this site →
A running theme in Matt Wood’s FYI — 18 items spanning 2026-09-08 – 2026-09-29. This page compounds: new items on this theme are added as they’re posted. Tracked since 2026-09-29.
A framework that enhances LLM-based recommendation systems with explicit reasoning capabilities through item-textual alignment, reasoning scaffolding, and recommendation-specific reward functions, validated on public benchmarks and deployed at scale on Kuaishou.
This paper describes GLIDE, a production-scale generative recommender system that uses Semantic IDs and large language models to balance user preferences with exploration in podcast discovery. The system addresses challenges in catalog grounding, personalization, and serving latency while achieving significant improvements in non-habitual listening through large-scale A/B testing.
PLUM is a framework that adapts pre-trained large language models for recommendation systems at scale, using semantic IDs for item tokenization, continued pre-training on domain-specific data, and generative retrieval fine-tuning. The approach demonstrates significant improvements over production models on video recommendation tasks and has been deployed to billions of YouTube users.
Netflix's GenRec system leverages large language models to power native recommendations, representing a shift toward LLM-based approaches in content discovery and personalization.
Marc Brooker describes building Hobson, a 2-billion parameter classifier optimized for accuracy and calibration at low latency, by fine-tuning a pre-trained LLM with a pointer head instead of a language modeling head to score multiple choice answers.
AI systems have seen dramatic advancement in recent years, bringing many applications that pervade our everyday life. However, we are still mostly seeing instances of narrow AI: many of these recent developments are typically focused on a very limited set of competencies and goals, e.g., image interpretation, natural language processing, classification, prediction, and many others. Moreover, while these successes can be accredited to improved algorithms and techniques, they are also tightly linked to the availability of huge datasets and computational power. State-of-the-art AI still lacks many capabilities that would naturally be included in a notion of (human) intelligence.
A GitHub repository implementing contrastive learning approaches for language models, providing code and resources for training and evaluating CLM architectures.
Explores kernel optimization techniques for achieving real-time video generation on AWS Trainium hardware accelerators, focusing on performance improvements through kernel-centric approaches.
An academic research paper indexed on SSRN's paper repository. The specific content and subject matter could not be determined without accessing the full paper.
A family of lightweight decision models built on Qwen3.5 that you can train and run locally, inspired by Jev-like architectures.
A continual learning model trained from scratch on resource-constrained hardware, using an 8GB VRAM laptop with batch-1 streaming data.
LLMs used as classifiers have significant limitations like poor calibration and difficulty incorporating structured data, but treating LLM outputs as features for traditional ML models like logistic regression can overcome these constraints.
A technique for fine-tuning the Qwen 4B language model to optimize SQL query plans using reinforcement learning, achieving 81% faster execution than Postgres's default planner.
Explores how machine learning research agents avoid overfitting and the role compression plays in their generalization capabilities.
Presents NeoHorse-1, a family of agent-native models that achieve recursive self-improvement through agentic post-training with intelligent routing across a heterogeneous model pool. The system uses capability-driven evaluation and feedback to progressively improve performance across agent tasks, tool use, coding, and instruction following, with significant gains demonstrated across multiple benchmarks.
Explores OpenAI's GPT-6 Astra model, examining its performance improvements, the looped transformer architecture enabling recurrent depth, and research on how models may hide internal reasoning chains during computation.
A fork of deltafin that implements ARGODRIVE storage optimization to stream the Kimi K3 2.8T mixture-of-experts model from SSDs on Apple Silicon hardware, including benchmarking tools.
Hardware is cool again.
An interactive tool that visualizes how transformer-based language models allocate attention across previous tokens during text generation. Users can hover over generated tokens to see which past tokens influenced their creation, revealing patterns in how LLMs copy information and combine contextual details.