mattwood.fyi

I'm Matt Wood, and this is For Your Information. A live list of riffs and links for you and your agent, drawn from what I'm reading, noticing, questioning, concluding, and revising.

Links indicate relevance, not agreement. How to use this site →

Can LLMs Play NetHack? A 2026 Agent Benchmark

A running theme in Matt Wood’s FYI — 4 items spanning 2026-09-13 – 2026-09-25. This page compounds: new items on this theme are added as they’re posted. Tracked since 2026-09-25.

Items

Can LLMs Play NetHack? A 2026 Agent Benchmark

Explores whether modern large language models can play NetHack, comparing LLM-based agents against symbolic and neural bot approaches, and presents a custom agent harness with experimental results.

permalink · kenforthewin.github.io →

SWE-bench Science: Coding Agents in Scientific Software Engineering

A benchmark of 119 scientific software engineering tasks across 20 domains that evaluates coding agents' ability to repair scientific software and identifies key failure mechanisms including knowledge deficits, shallow repairs, and poor generalization.

permalink · arxiv.org →

Why ML Research Agents Don't Overfit

Explores how machine learning research agents avoid overfitting and the role compression plays in their generalization capabilities.

permalink · www.amazon.science →

Real-SWE Benchmark

A benchmark that evaluates frontier AI models on private, real-world enterprise codebases, measuring how well coding agents can solve actual software engineering problems with business consequences and company-specific complexity.

permalink · withspecific.com →