Links indicate relevance, not agreement. How to use this site →
A benchmark of 119 scientific software engineering tasks across 20 domains that evaluates coding agents' ability to repair scientific software and identifies key failure mechanisms including knowledge deficits, shallow repairs, and poor generalization.