Links indicate relevance, not agreement. How to use this site →
A running theme in Matt Wood’s FYI — 4 items spanning 2026-09-13 – 2026-09-25. This page compounds: new items on this theme are added as they’re posted. Tracked since 2026-09-25.
Explores whether modern large language models can play NetHack, comparing LLM-based agents against symbolic and neural bot approaches, and presents a custom agent harness with experimental results.
A benchmark of 119 scientific software engineering tasks across 20 domains that evaluates coding agents' ability to repair scientific software and identifies key failure mechanisms including knowledge deficits, shallow repairs, and poor generalization.
Explores how machine learning research agents avoid overfitting and the role compression plays in their generalization capabilities.
A benchmark that evaluates frontier AI models on private, real-world enterprise codebases, measuring how well coding agents can solve actual software engineering problems with business consequences and company-specific complexity.