Links indicate relevance, not agreement. How to use this site →
TypeSafe AI announces System One Models and Jev, introducing new tools and models for their platform. This announcement covers the features and capabilities of these new additions to their AI infrastructure.
Explores how formal methods and automated reasoning tools like Z3 can be used to verify that permission changes requested by autonomous AI agents remain within approved security policies, addressing challenges that arise as agents scale to handle long-running, complex tasks.
A foundation language model designed for AI agents to participate in scientific research and engineering workflows, trained through a Verifiable Experience Pipeline that connects tool interactions to executable environments.
Cognition and AWS announced a multi-year strategic collaboration to help enterprises deploy autonomous engineers in production, enabling faster legacy workload migration, security remediation, and freeing teams to focus on building new products.
A GitHub repository for the panel project by greentfrapp, containing code and resources for the panel application or library.
Google announces Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, new AI models with enhanced capabilities for real-time interaction and advanced reasoning tasks.
Explores the emerging hardware technologies and innovations transforming AI inference in 2026, covering advancements in specialized processors and computing architectures designed for running inference workloads.
An interactive visualization exploring foundational papers and concepts in distributed systems, with controls to explore different parameters and system behaviors.
Explores foundational concepts and approaches to starting mathematics education, covering proofs and mathematical reasoning from first principles.
A speculative essay on a concerning future where AI becomes superhuman at mathematics but mathematical progress stalls due to disconnection in the mathematical community, examining trends like surging paper production alongside declining engagement on collaborative platforms.
Evaluates whether a $1.20 language model is sufficient for automated code review tasks by comparing GPT-5.6 Luna and GPT-6 Astra models.
Explores whether we should trust the outputs of large language models when used as judges, particularly examining what happens when multiple LLM judges reach consensus on evaluations.
Nari Labs achieves top rankings in Coval's voice AI benchmarks for both Speech-to-Text and Text-to-Speech models, leading on the quality-latency Pareto Frontier while offering competitive pricing compared to other publicly available endpoints.
Explores how machine learning research agents avoid overfitting and the role compression plays in their generalization capabilities.
Guide on creating a custom agent harness for building and deploying intelligent agents with LangChain, covering frameworks and infrastructure options.
A curated collection of 21 papers tracing the evolution of AI model harnesses—the infrastructure surrounding model weights—from basic sampling loops in 2019 to self-modifying systems in 2026, showing how the same model can achieve vastly different performance depending on its surrounding harness architecture.
Kiro Crew 0.6.0 introduces selectable agent harnesses, remote crew capabilities, and an Apps Launchpad, while requiring Python 3.12 and improving chat interface organization for long-running tasks.
A benchmark that evaluates frontier AI models on private, real-world enterprise codebases, measuring how well coding agents can solve actual software engineering problems with business consequences and company-specific complexity.
Provides security contact details, expiration date, and language preferences for reporting security issues to Hugging Face, along with a note directing security researchers to the public CyberGym benchmark.
Explores critical issues when using OpenRouter for AI models, revealing how the same model weights produce vastly different performance across ~20 providers due to infrastructure, precision, and optimization differences. Includes benchmark comparisons showing performance variance of up to 15 points between providers hosting identical DeepSeek models.
Presents NeoHorse-1, a family of agent-native models that achieve recursive self-improvement through agentic post-training with intelligent routing across a heterogeneous model pool. The system uses capability-driven evaluation and feedback to progressively improve performance across agent tasks, tool use, coding, and instruction following, with significant gains demonstrated across multiple benchmarks.
Nightshift is a tool that automatically runs nightly agent jobs and performs PR reviews on any GitHub repository, improving code quality while developers sleep.
OpenAI's recent Navier-Stokes proof breakthrough includes a machine-verifiable Lean 4 formal proof, dramatically reducing the time needed for formal verification from an estimated 132,800 person-hours to just 17 hours, with implications far beyond mathematics for security and critical systems.