Links indicate relevance, not agreement. How to use this site →
A running theme in Matt Wood’s FYI — 17 items spanning 2026-09-10 – 2026-09-25. This page compounds: new items on this theme are added as they’re posted. Tracked since 2026-09-25.
The Model Context Protocol (MCP) enables AI agents to interact with external tools more efficiently through code execution instead of direct tool calls, reducing token consumption and context window overhead when managing hundreds or thousands of tools.
Explores trends in the declining costs of AI computation and cognitive processing, examining how price reductions are reshaping the economics of artificial intelligence capabilities.
Explores how rapidly decreasing AI token costs across orders of magnitude will transform machine learning from a premium product to ubiquitous computing infrastructure within 1-3 years, examining improvements in GPUs, models, and the implications of supply and demand-side effects.
Unreal Agent is an agent harness that achieves up to 40% cost savings compared to Codex by managing tool calls asynchronously, allowing agents to respond quickly to users while reducing model overhead without sacrificing performance.
MiMo-V2.6 is a Xiaomi product or software version, though specific details about its functionality are not provided in the available content.
Strands introduces their state-of-the-art agent harness that delivers frontier-level performance while reducing token costs by 28%.
A family of lightweight decision models built on Qwen3.5 that you can train and run locally, inspired by Jev-like architectures.
A continual learning model trained from scratch on resource-constrained hardware, using an 8GB VRAM laptop with batch-1 streaming data.
AX is a platform for declaring and running agentic tasks at scale with sandboxed execution, workspace management, network policies, and model configuration. It enables billions of concurrent agent sessions per cluster with sub-second task resumption and dense resource multiplexing.
Presents SoL-Pi, a method for improving coding agents through recursive self-improvement loops that reduces token usage by 44.7-49.0% and API costs by about one third while maintaining performance on code generation tasks.
LLMs used as classifiers have significant limitations like poor calibration and difficulty incorporating structured data, but treating LLM outputs as features for traditional ML models like logistic regression can overcome these constraints.
Explores the emerging hardware technologies and innovations transforming AI inference in 2026, covering advancements in specialized processors and computing architectures designed for running inference workloads.
Guide on creating a custom agent harness for building and deploying intelligent agents with LangChain, covering frameworks and infrastructure options.
Kiro Crew 0.6.0 introduces selectable agent harnesses, remote crew capabilities, and an Apps Launchpad, while requiring Python 3.12 and improving chat interface organization for long-running tasks.
Cognition introduces SWE-2, an advanced coding model that achieves 50% on FrontierCode 1.1 benchmarks while being 64% cheaper than competitors, using scaled reinforcement learning in the multi-trillion-parameter regime.
DeepSeek announces V4.1-Flash, a compact model in their new architecture family featuring native visual understanding, designed for faster inference and higher throughput while maintaining greater capability.
A guide for developers on how to fundamentally change their workflow to leverage AI agents for autonomous code generation and iteration, covering ten core principles and the mindset shift required to achieve significant productivity gains.