Links indicate relevance, not agreement. How to use this site →
A running theme in Matt Wood’s FYI — 17 items spanning 2026-09-08 – 2026-10-06. This page compounds: new items on this theme are added as they’re posted. Tracked since 2026-10-06.
An open-source AI accelerator implementation including RTL, ISA, simulator, compiler and profiler. Supports running Qwen3, LLaMA2.5 and Qwen3.5 models on Kintex-7 FPGA PCIe cards.
Mistral Large 4 is a large language model available in public preview as of October 6, 2026, with playground access and model comparison capabilities.
A zeroth-order optimization method that trains transformer language models by perturbing activations at each token without backpropagation, achieving competitive performance with backprop while being orders of magnitude more efficient than weight-space evolutionary strategies.
Reflection releases Beam, a 501 billion parameter sparse Mixture-of-Experts model optimized for coding, reasoning, and agentic tasks, trained with large-scale reinforcement learning and achieving competitive performance with strong inference efficiency compared to similar open-weight models.
Language models that manage their own context by treating it as an editable file, enabling dynamic context updates and multi-agent systems while outperforming existing context management strategies across various tasks. The approach supports both in-context and parametric learning through natural language steering and reinforcement learning, with optimized serving via suffix cache reuse.
OpenAI and Synopsys announced a partnership to develop GPT-Synopsys, an AI system designed to revolutionize semiconductor chip design using frontier intelligence technology.
Explores kernel optimization techniques for achieving real-time video generation on AWS Trainium hardware accelerators, focusing on performance improvements through kernel-centric approaches.
Comparison of Claude Opus 5.5 (Adaptive Reasoning) across intelligence metrics, performance benchmarks, and pricing, ranking it among top models with detailed technical specifications and cost analysis.
MiMo-V2.6 is a Xiaomi product or software version, though specific details about its functionality are not provided in the available content.
A continual learning model trained from scratch on resource-constrained hardware, using an 8GB VRAM laptop with batch-1 streaming data.
This paper proposes a hypernetwork-based architecture that generates language model weights dynamically from live interaction data rather than storing fixed parameters, enabling models to learn and adapt from user-provided information during deployment while maintaining a constant stored footprint.
A benchmark of 119 scientific software engineering tasks across 20 domains that evaluates coding agents' ability to repair scientific software and identifies key failure mechanisms including knowledge deficits, shallow repairs, and poor generalization.
A foundation language model designed for AI agents to participate in scientific research and engineering workflows, trained through a Verifiable Experience Pipeline that connects tool interactions to executable environments.
Google announces Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, new AI models with enhanced capabilities for real-time interaction and advanced reasoning tasks.
Explores the emerging hardware technologies and innovations transforming AI inference in 2026, covering advancements in specialized processors and computing architectures designed for running inference workloads.
Evaluates whether a $1.20 language model is sufficient for automated code review tasks by comparing GPT-5.6 Luna and GPT-6 Astra models.
A fork of deltafin that implements ARGODRIVE storage optimization to stream the Kimi K3 2.8T mixture-of-experts model from SSDs on Apple Silicon hardware, including benchmarking tools.
Hardware is cool again.