Links indicate relevance, not agreement. How to use this site →
A running theme in Matt Wood’s FYI — 5 items spanning 2026-09-14 – 2026-09-30. This page compounds: new items on this theme are added as they’re posted. Tracked since 2026-09-30.
Explores principles for designing evaluations and improving performance against them without overfitting, demonstrating how Claude's build-eval and hillclimb commands automate these practices to help assess and optimize AI applications.
Comparison of Claude Opus 5.5 (Adaptive Reasoning) across intelligence metrics, performance benchmarks, and pricing, ranking it among top models with detailed technical specifications and cost analysis.
Documents updates and changes to Claude Code, including new features, improvements, and bug fixes for the AI-powered coding assistant.
Claude now reads AGENTS.md if there is no CLAUDE.md. Finally.
Google announces Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, new AI models with enhanced capabilities for real-time interaction and advanced reasoning tasks.
Evaluates whether a $1.20 language model is sufficient for automated code review tasks by comparing GPT-5.6 Luna and GPT-6 Astra models.