mattwood.fyi

I'm Matt Wood, and this is For Your Information. A live list of riffs and links for you and your agent, drawn from what I'm reading, noticing, questioning, concluding, and revising.

Links indicate relevance, not agreement. How to use this site →

Automating Eval Design and Hillclimbing with Claude

A running theme in Matt Wood’s FYI — 5 items spanning 2026-09-14 – 2026-09-30. This page compounds: new items on this theme are added as they’re posted. Tracked since 2026-09-30.

Lines of development

Automating Eval Design and Hillclimbing with Claude → Engineering a Calibrated Classifier Model — Engineering a calibrated classifier model involves similar iterative evaluation design principles; the new item automates and extends these practices into a systematic hillclimbing workflow

Items

Automating Eval Design and Hillclimbing with Claude

Explores principles for designing evaluations and improving performance against them without overfitting, demonstrating how Claude's build-eval and hillclimb commands automate these practices to help assess and optimize AI applications.

permalink · claude.dev →

Claude Opus 5.5 Intelligence & Performance Analysis

Comparison of Claude Opus 5.5 (Adaptive Reasoning) across intelligence metrics, performance benchmarks, and pricing, ranking it among top models with detailed technical specifications and cost analysis.

permalink · artificialanalysis.ai →

Claude Code Changelog

Documents updates and changes to Claude Code, including new features, improvements, and bug fixes for the AI-powered coding assistant.

Claude now reads AGENTS.md if there is no CLAUDE.md. Finally.

permalink · code.claude.com →

Gemini 3.8 Live and Extended Thinking Models

Google announces Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, new AI models with enhanced capabilities for real-time interaction and advanced reasoning tasks.

permalink · blog.google →

Comparing GPT-5.6 Luna vs GPT-6 Astra for Code Review

Evaluates whether a $1.20 language model is sufficient for automated code review tasks by comparing GPT-5.6 Luna and GPT-6 Astra models.

permalink · entelligence.ai →