Links indicate relevance, not agreement. How to use this site →
A running theme in Matt Wood’s FYI — 74 items spanning 2026-04-24 – 2026-09-21. This page compounds: new items on this theme are added as they’re posted. Tracked since 2026-08-28.
A family of lightweight decision models built on Qwen3.5 that you can train and run locally, inspired by Jev-like architectures.
A continual learning model trained from scratch on resource-constrained hardware, using an 8GB VRAM laptop with batch-1 streaming data.
LLMs used as classifiers have significant limitations like poor calibration and difficulty incorporating structured data, but treating LLM outputs as features for traditional ML models like logistic regression can overcome these constraints.
Nari Labs achieves top rankings in Coval's voice AI benchmarks for both Speech-to-Text and Text-to-Speech models, leading on the quality-latency Pareto Frontier while offering competitive pricing compared to other publicly available endpoints.
DeepSeek announces V4.1-Flash, a compact model in their new architecture family featuring native visual understanding, designed for faster inference and higher throughput while maintaining greater capability.
Announces the release of Mercury 2.5, a framework or tool update from Inception Labs with new features and improvements.
Comprehensive benchmark testing of different quantization levels (1-bit to 8-bit) for the Qwen3.8 27B model, showing that 4-bit quantization maintains performance while 1-bit quantization severely degrades results across GPQA Diamond, IFBench, and Terminal-Bench 2.1.
OpenAI announces ChatGPT Images 2.5, an update to their image generation capabilities within ChatGPT. The release details new features and improvements to image creation functionality.
This is a research paper from OpenAI examining gaps between primes.
Not an expert, but this looks to be a novel solution from GPT 6 Astra.
OpenAI's safety documentation for GPT-6 Astra covering model training data, internal deployment, safe completions evaluations, and robustness against jailbreaks.
Meta's Muse Spark is an AI model for generating creative content, featuring advanced capabilities for text-to-image synthesis and multimodal understanding.
Google announces two new AI models in the Gemini family: Gemini 3.8 Flash and 3.8 Flash Cyber, representing updates to their fast, efficient model offerings.
OpenAI's overview of the development trajectory and roadmap toward Astra, their advanced AI system.
Release notes for vLLM version 0.28.0, detailing new features, improvements, and changes to the LLM inference optimization library.
OpenAI's custom-designed Jalapeño inference chip, developed in partnership with Broadcom, outperforms through hardware-software codesign and general-purpose optimization rather than model-specific specialization.
Trismik is a platform for switching between different models based on evidence and data-driven decisions, enabling users to select the most appropriate model for their specific needs.
DiffusionGemma is an open-weight language model that uses discrete diffusion to generate text in parallel blocks of 256 tokens, achieving around 1,500 output tokens per second on a single GPU—substantially faster than autoregressive models. The model is created by fine-tuning Gemma 4 with a two-stage training pipeline combining supervised fine-tuning and reinforcement learning, while retaining support for thinking mode, multimodal inputs, and long contexts.
We started OpenRouter in early 2023 on a simple belief: intelligence will be multi-model. No single model will win every task, and the frontier will move rapidly. That freedom is critical infrastructure for the industry. AI is too important for its future to be decided by whichever single model gets embedded first. AI has become the single largest driver of economic growth in the US, and inference is quickly becoming the largest line item for every company.
Chef's kiss. No notes.
Explores how sending LLM requests twice and taking the faster response can reduce tail latency more effectively than paying for premium service tiers, with real benchmarks from a voice agent service.
Comprehensive intelligence, performance, and price analysis of Alibaba's Qwen3.8 27B open-weights model, including benchmark scores, technical specifications, and comparison with similar models.
An analysis of GPT 5.6 Sol and its capabilities as OpenAI's most advanced vision model for computer vision tasks.
A comprehensive guide to Kimi K3, covering its features and capabilities for developers in 2026.
Cerebras and OpenAI introduce Ultrafast Mode, a new service tier delivering up to 750 output tokens per second for GPT-5.6 Sol, enabling 11x faster performance than competing models while maintaining frontier-level intelligence for time-sensitive applications.
Google announces Gemini 3.7 Flash, a high-performance AI model designed to deliver intelligent results with efficient processing capabilities for practical applications.
Exploration of 11 AI models including DeepSeek, Qwen, Kimi and open-source alternatives, examining their different outputs and capabilities for the same prompts.
A research benchmark evaluating how well large language models can discover new thermally conductive dielectric materials for advanced semiconductor applications, with a leaderboard tracking computational discoveries and synthesis feasibility.
Grok 4.6 achieves frontier-level AI intelligence with a score of 61 on the Artificial Analysis Intelligence Index, matching GPT-5.6 Sol while offering significantly lower costs and excelling in agentic tasks like customer service and terminal-based work.
A large-scale language model from Qwen featuring 512 experts with a mixture-of-experts architecture, designed for text generation tasks and available on Hugging Face.
DeepSeek V4 Pro 0813 is a large-scale mixture-of-experts model available through OpenRouter with pricing of $0.435/$0.87 per 1M tokens and 1M context window support.
OpenAI's latest specialized cybersecurity models (Daybreak Red and Blue) are now available on Amazon Bedrock with enterprise security features, data governance, and pricing that matches OpenAI's first-party rates.
NVIDIA introduces a lightweight open model and routing library that enables faster, more efficient AI agents with greater control over data and workflows across edge devices, PCs, workstations, data centers, and cloud environments.
Explores the fundamental relationship between data compression and predictive modeling, examining how compression algorithms work as predictors and the implications for AI and LLMs.
An H3 inference engine implementation for Mac computers, providing MiniMax-based AI inference capabilities.
GPU passthrough in macOS virtual machines, covering how Apple Silicon's architecture handles GPU virtualization, the technical challenges involved, and how the Virtualization framework enables GPU resource sharing between host and guest macOS environments.
LFM2.5-2.6B is a 2.6 billion parameter language model developed by Liquid AI, available on Hugging Face, featuring instruction-following capabilities with a chat template supporting system prompts and tool use.
Needle 2 is a 45-million-parameter, 14MB open-source language model designed for tool calling, device control, and structured extraction on low-cost edge hardware like microcontrollers, budget phones, and Raspberry Pis. It runs a full session in 28MB of RAM using CQ2-bit compression, achieving competitive performance against much larger small models on mobile device use benchmarks.
Feels like lots is all happening at once with small mobile-friendly models, but this is a space which has been making steady progress for months now. Encouraging.
Methods for inferring hidden details about how large language models like GPT-5 and Claude were trained, including estimating parameter counts, dataset mixtures, and training timelines by probing models with carefully curated questions. Covers the three main stages of frontier model training (pre-training, capability fine-tuning, and post-training) as context for understanding what these probing techniques reveal.
Muse Glimmer is a 30-billion-parameter open-source AI model from Meta Superintelligence Labs, optimized for local agentic workflows including function calling, coding, and LLM-as-a-judge tasks, designed to run on consumer hardware without requiring cloud infrastructure. The model was developed using a combination of logit distillation from a larger teacher model, mid-training on agent-heavy data, and post-training with reinforcement learning, and is released under an Apache 2.0 license.
AMD's acquisition of AI chip startup Taalas aims to enhance inference performance by physically encoding AI models directly into silicon hardware. The deal represents AMD's effort to compete in the AI accelerator market by leveraging Taalas's approach of optimizing chips at the hardware level for specific AI workloads.
OpenAI's work on improving the GPT-4.5 or a related model's performance on solving problems, likely focusing on enhancements to reasoning, accuracy, or problem-solving capabilities within ChatGPT. The content likely details technical improvements, benchmark results, or methodology changes made to advance the model's abilities.
An exploration of how the economic concept of Pareto efficiency can be applied to optimizing character and kart builds in Mario Kart 8, using the game's multiple competing statistics (speed, acceleration, handling, etc.) to identify dominant choices and eliminate suboptimal ones. The piece uses interactive visualizations to demonstrate how the Pareto front helps narrow down thousands of possible build combinations to a set of objectively non-dominated options.
Important work.
Marin is an open collaborative lab for building foundation models from scratch, sharing all code, data, experiments, and results transparently in real-time. It invites open-source contributors to participate in model training, architecture research, and experiments, with publicly documented models like Marin-8B and Marin-32B that compete with leading open-weight models.
Open Athena is a nonprofit organization that partners with academic institutions to develop open-source AI foundation models by providing engineering talent, compute resources, and coordination. Current projects include large language models with Stanford, plant genomics DNA models with Cornell, and protein generation models with MIT.
Bonsai is an OCaml library by Jane Street for building dynamic web applications using Js_of_ocaml. It provides a framework for creating interactive front-end UIs compiled from OCaml to JavaScript.
Maintainers are now facing an assault on two fronts. The barrier to entry for generating code has dropped to zero... flooded repository gates with an alarming volume of low-quality, AI-generated pull requests. Maintainers who once spent their time writing code are now forced to become full-time, unpaid code reviewers.
Key research highlights at ICML 2026: - FlashAttention-4: Algorithm and kernel pipelining co-design for asymmetric hardware scaling (Tri Dao et al.) - Mamba-3: Next-generation state space model (CMU + Princeton + Together + Cartesia) - DeepSWE: Fully open-sourced state-of-the-art coding agent trained by scaling RL - Cache-aware prefill-decode disag
Key positioning: "Stop relying on generic AI models. Databricks has the tools to build agent systems that deliver accurate, data-driven results."
In the run that used GPT-5.5 for both planners and workers, the workers alone cost $9,373. In the run where Opus 4.8 did the planning and Composer 2.5 did the work, the entire worker fleet cost $411.
From the homepage: "Build something Lovable. Create apps and websites by chatting with AI." Templates include spatial canvas tools, blogs, habit trackers, code-powered presentation builders, and e-commerce stores. Key metrics highlighted: millions of projects built, with substantial new projects created per week.
Jensen Huang positioning: "Fireworks is the TSMC of AI Factories..."
Baseten delivers the infrastructure, tooling, and expertise needed to bring the most performant AI products to market, fast.
Valkey is an open source (BSD) high-performance key/value datastore that supports a variety of workloads such as caching, message queues, and can act as a primary database. The project is backed by the Linux Foundation, ensuring it will remain open source forever.
"We are building an intelligence layer for the entire aviation system with Zee. The company that owns the data and the foundation model will help lead the aviation industry into the next era of flight," said Adam Goldstein, founder and CEO of Archer.
Model specs: - 975B total params, 41B active (Mixture-of-Experts) - Inkling-Small: 12B active params - 1M token context window - Pretrained on 45T tokens (text, images, audio, video) - Controllable thinking effort (0.2 to 0.99 sweep)
There are only two things you can say with certainty about token prices: we're in a supply crunch, and this is unstable. All of the variables are in play, and the market will get shaken out over the next few years to arrive at a new equilibrium.
gog mcpruns a typed MCP server over stdio for agent clients that need a permissioned Google Workspace tool surface. It intentionally does not expose a generic shell/argv bridge. Each MCP tool has a fixed schema and maps to a specific gog operation. MCP defaults are read-only. Write tools are hidden unless the server is started with `--allow-wri
Executing deep neural networks on GPUs has dominated AI for a decade, but we think the next jump in energy efficiency demands a fundamentally different computer, one where physics does the computing. We built Un-0, an image generator powered by a simulated system of coupled oscillators, an example of an emerging physical computing substrate.
I am now in the camp who believe that we are only at most two or three model generations away from AI finally being controlled like nuclear weapons. Only a few will have access to superintelligence above the classes of models we're seeing this year.
The 1997 analogy: We are in an era of radical uncertainty where the technology is transformative but most high-value use cases haven't been built yet and the "winners" are not yet clear. Just as it was impossible in 1997 to predict that a search engine with a quirky logo would reshape the world, we cannot yet see the final shape of the AI-driven ec
Problem: Adapting foundation models to a morphologically rich language (Azerbaijani) with limited training data, no existing blueprint for efficient LLM training in that language.
Mollick's post (May 30, 2026): "It does seem like meaningfully better AI releases are accelerating, especially from OpenAI & Anthropic. To illustrate, I caused this timeline to be created. It only lists new models that scored 3 points or higher over previous models in the Artificial Analysis index."
Key quotes from Levie (via X/LinkedIn, reported across multiple outlets):
(LinkedIn source blocked; synthesized from Columbia Business School event, Fortune reporting, and Forbes coverage of Chatterji's research)
Stats: 34.7k stars, 2.8k forks, TypeScript 70.4%, MIT License