Links indicate relevance, not agreement. How to use this site →
The author argues against using AI for substantive writing tasks, contending that writing is essential to thinking, AI-generated text contains subtle errors, and unlabeled AI writing is misleading to readers.
Describes how the Kiro Crew team merged 1,000 pull requests in seven days by evolving their development workflow through five stages, from single manual agent sessions to an automated agent pipeline architecture that coordinates parallel work through message queues.
Apodex 1.1 develops AI agents capable of complex, long-horizon tasks by scaling both environment diversity (files, search, code execution) and coordination abilities (task decomposition, parallel work, asynchronous integration). The system achieves strong performance across finance, research, mathematics, and coding with a 35B parameter model, grounded in verifiable, real-world work completion.
Documents updates and changes to Claude Code, including new features, improvements, and bug fixes for the AI-powered coding assistant.
Claude now reads AGENTS.md if there is no CLAUDE.md. Finally.
Two rules for using LLMs as copyeditors rather than ghostwriters: never use LLM-suggested phrases verbatim, and avoid taking LLM encouragement at face value. The approach helps writers maintain authentic voice while leveraging AI to identify flaws.
This matches how my own approach to writing has evolved, too.
Jalapeño demonstrates how large language models can be effectively applied to semiconductor chip design, showcasing the potential of LLMs in accelerating hardware engineering workflows.
Skillsync is a local-first tool that makes AI coding sessions portable across different agents and teammates, allowing you to resume work seamlessly without re-explaining context or experiencing vendor lock-in.
LLMs used as classifiers have significant limitations like poor calibration and difficulty incorporating structured data, but treating LLM outputs as features for traditional ML models like logistic regression can overcome these constraints.
A community platform where builders and engineers share their AI tools, workflows, and setups to learn from each other and discover how others work with AI technologies.
This paper proposes a hypernetwork-based architecture that generates language model weights dynamically from live interaction data rather than storing fixed parameters, enabling models to learn and adapt from user-provided information during deployment while maintaining a constant stored footprint.
A curated marketplace of AI agent skills (SKILL.md packages) that teach coding agents how to perform specific tasks, with human-reviewed submissions showing before-and-after evidence and available via JSON API or MCP server.
Hister is a tool for building and running a personal search engine, allowing users to create and manage their own search functionality.
Kiro's autonomous mode is an AI agent that automatically handles maintenance tasks end-to-end, from issue analysis to pull request submission, allowing development teams to focus on code review and higher-value work. AWS's Automated Reasoning Group used this tool to address 87 open issues in two months across formal verification repositories.
Google's document exploring the applications, opportunities, and implications of artificial intelligence across scientific research and discovery.
A technique for fine-tuning the Qwen 4B language model to optimize SQL query plans using reinforcement learning, achieving 81% faster execution than Postgres's default planner.
Proposes Dream-RSI, a framework that enables autonomous AI agents to recursively improve their exploration strategies by using accumulated discovery history as a replay simulator. This allows efficient off-policy evaluation and refinement of exploration policies without expensive online evaluations, demonstrated across algorithm engineering, mathematical optimization, and GPU kernel engineering tasks.
Apple introduces a new opt-in camera mode for iPhone 18 Pro that creates cryptographically secured reference images to verify photograph authenticity, using dedicated hardware and Private Cloud Compute to protect both integrity and photographer privacy.
A benchmark of 119 scientific software engineering tasks across 20 domains that evaluates coding agents' ability to repair scientific software and identifies key failure mechanisms including knowledge deficits, shallow repairs, and poor generalization.
Vercel's experimental projects and cutting-edge tools, including AI capabilities, infrastructure innovations, and platform features for building modern web applications.
A local-first inbox application for long-running AI agents, built with DeepAgents and LangGraph frameworks.
Explores how AI-assisted coding ("vibe coding") is following the same normalization trajectory as internet dating, eventually becoming the default practice as AI tools become ubiquitous in software development.
A local-first inbox application designed for long-running AI agents, built using DeepAgents and LangGraph frameworks.
A small neural network that classifies whether a number is Numberwang.
That's Numberwang. Essential research.