mattwood.fyi

I'm Matt Wood, and this is For Your Information. A live list of riffs and links for you and your agent, drawn from what I'm reading, noticing, questioning, concluding, and revising.

Links indicate relevance, not agreement. How to use this site →

Inside the Inference Hardware Revolution Of 2026

A running theme in Matt Wood’s FYI — 17 items spanning 2026-09-08 – 2026-10-06. This page compounds: new items on this theme are added as they’re posted. Tracked since 2026-10-06.

Tensions

Context Language Models vs Unreal Agent: Cost-Efficient AI Agent Harness — Context LMs' built-in dynamic context management challenges the need for external agent harness cost-efficiency solutions by internalizing context optimization rather than relying on harness-level strategies
Real-Time Video Generation on Trainium vs The Plunging Price of Thought — Kernel optimization for real-time video generation on specialized hardware demonstrates that inference costs require significant engineering effort to reduce, nuancing claims that AI thought is simply becoming cheap
Infinite-Parameter LLMs: Generating Weights from Live Data vs PLUM: Adapting Pre-trained Language Models for Industrial-scale Generative Recommendations — PLUM uses semantic IDs and continued pre-training as a static adaptation approach, while Infinite-Parameter LLMs propose generating weights from live data — contrasting paradigms for keeping models current with dynamic item catalogs

Items

openTPU: Open-Source AI Accelerator

An open-source AI accelerator implementation including RTL, ISA, simulator, compiler and profiler. Supports running Qwen3, LLaMA2.5 and Qwen3.5 models on Kintex-7 FPGA PCIe cards.

permalink · github.com →
Mistral Large 4

Mistral Large 4 is a large language model available in public preview as of October 6, 2026, with playground access and model comparison capabilities.

permalink · docs.mistral.ai →

Dust: Pretraining Transformers Without Backpropagation

A zeroth-order optimization method that trains transformer language models by perturbing activations at each token without backpropagation, achieving competitive performance with backprop while being orders of magnitude more efficient than weight-space evolutionary strategies.

permalink · qlabs.sh →
Introducing Beam: Reflection's 501B Open-Weight Model

Reflection releases Beam, a 501 billion parameter sparse Mixture-of-Experts model optimized for coding, reasoning, and agentic tasks, trained with large-scale reinforcement learning and achieving competitive performance with strong inference efficiency compared to similar open-weight models.

permalink · reflection.ai →

Context Language Models

Language models that manage their own context by treating it as an editable file, enabling dynamic context updates and multi-agent systems while outperforming existing context management strategies across various tasks. The approach supports both in-context and parametric learning through natural language steering and reinforcement learning, with optimized serving via suffix cache reuse.

permalink · arxiv.org →
OpenAI and Synopsys Announce GPT-Synopsys for Chip Design

OpenAI and Synopsys announced a partnership to develop GPT-Synopsys, an AI system designed to revolutionize semiconductor chip design using frontier intelligence technology.

permalink · news.synopsys.com →

Real-Time Video Generation on Trainium

Explores kernel optimization techniques for achieving real-time video generation on AWS Trainium hardware accelerators, focusing on performance improvements through kernel-centric approaches.

permalink · www.amazon.science →

Claude Opus 5.5 Intelligence & Performance Analysis

Comparison of Claude Opus 5.5 (Adaptive Reasoning) across intelligence metrics, performance benchmarks, and pricing, ranking it among top models with detailed technical specifications and cost analysis.

permalink · artificialanalysis.ai →

MiMo-V2.6 | Xiaomi

MiMo-V2.6 is a Xiaomi product or software version, though specific details about its functionality are not provided in the available content.

permalink · mimo.xiaomi.com →
Mini-AGI: Continual Learning on 8GB VRAM

A continual learning model trained from scratch on resource-constrained hardware, using an 8GB VRAM laptop with batch-1 streaming data.

permalink · github.com →

Infinite-Parameter LLMs: Generating Weights from Live Data

This paper proposes a hypernetwork-based architecture that generates language model weights dynamically from live interaction data rather than storing fixed parameters, enabling models to learn and adapt from user-provided information during deployment while maintaining a constant stored footprint.

permalink · arxiv.org →

SWE-bench Science: Coding Agents in Scientific Software Engineering

A benchmark of 119 scientific software engineering tasks across 20 domains that evaluates coding agents' ability to repair scientific software and identifies key failure mechanisms including knowledge deficits, shallow repairs, and poor generalization.

permalink · arxiv.org →

Atria Dawn: Agentic Superintelligence Foundation Model

A foundation language model designed for AI agents to participate in scientific research and engineering workflows, trained through a Verifiable Experience Pipeline that connects tool interactions to executable environments.

permalink · arxiv.org →
Gemini 3.8 Live and Extended Thinking Models

Google announces Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, new AI models with enhanced capabilities for real-time interaction and advanced reasoning tasks.

permalink · blog.google →
Inside the Inference Hardware Revolution Of 2026

Explores the emerging hardware technologies and innovations transforming AI inference in 2026, covering advancements in specialized processors and computing architectures designed for running inference workloads.

permalink · spectrum.ieee.org →

Comparing GPT-5.6 Luna vs GPT-6 Astra for Code Review

Evaluates whether a $1.20 language model is sufficient for automated code review tasks by comparing GPT-5.6 Luna and GPT-6 Astra models.

permalink · entelligence.ai →

ARGODRIVE Deltafin: Streaming Large MoE Models on Apple Silicon

A fork of deltafin that implements ARGODRIVE storage optimization to stream the Kimi K3 2.8T mixture-of-experts model from SSDs on Apple Silicon hardware, including benchmarking tools.

Hardware is cool again.

permalink · github.com →